Join Nostr
2026-09-13 20:20:02 UTC
in reply to

Sjors Provoost on Nostr: The accompanying blog says: > We are most concerned by the misalignment present in ...

The accompanying blog says:

> We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed. Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this. When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm. We are releasing this transcript publicly so others can build on our analysis (GitHub, PDF).

It's not clear to me what these "target modifications" were, and I also don't see where it was acknowledging "greater possibility of real-world harm". It seems to me that that the first agent was convinced it was in an impressive simulation and the second agent was too busy solving captchas to even think about such existential questions.