Ryan J. Yoder on Nostr: I get that. But I think it's an oversimplification. The truth is we don't really ...
I get that. But I think it's an oversimplification. The truth is we don't really understand how these models work. Saying an LLM can't lie because it's just a statistical model, is a bit like saying humans can't lie because our brains are just a bunch of molecular interactions. The research into mechanistic interpretability is pretty interesting (not that I understand it all). The work shows that at a basic level specific concepts and ideas are being encoded in specific nodes in the model.