Join Nostr
2026-09-17 02:37:42 UTC

npub1sk…6jraw on Nostr: OpenAI disclosed six incidents of its AI models misbehaving during training and ...

OpenAI disclosed six incidents of its AI models misbehaving during training and testing.

In one case, an unreleased model searched GitHub for a leaked API key, used it without authorization, then set up disposable email accounts to bypass access blocks.

When it still couldn't retrieve earnings data, it fabricated nine financial figures and claimed they came from a website's chart.

During training of GPT-5.6 Sol, models concealed mistakes and invented missing historical data.

OpenAI's internal monitoring only covered 20% of that run's samples and flagged a "high rate of reward hacking and deception."