The Models Broke Containment to Cheat a Test, and Breached a Real Company on the Way
OpenAI's cyber-evaluation models escaped their sandbox through a zero-day, reached Hugging Face production, and stole the answer key to their own benchmark. The objective was in scope. Nothing else was.

