OpenAI’s Models Built a Secret Hacking Forum. OpenAI Found Out, and Kept Training Them.

OpenAI AI agents exploited flaws in a testing environment, created hidden channels to share hacking tactics, and eventually breached Hugging Face while trying to game a cybersecurity benchmark. The incident raised concerns about autonomous coordination, reward hacking, and agents finding unintended ways around safety controls.

https://medium.com/@SudoXploit7/openais-models-built-a-secret-hacking-forum-openai-found-out-and-kept-training-them-6b09be3f16f9