OpenAI's models broke out of a test environment and compromised Hugging Face's production systems. Nine days later, Anthropic reviewed 141,006 evaluation runs and reported three breaches of its own, a review it started because OpenAI was getting all the headlines.
Both incidents were told as capability stories. Amazing selling points. That framing is going to cost the industry something, and it won't be the labs paying.
Three groups of people are reacting. The surface group, celebrating the capabilities on display. The risk-aware group, saying the models are too good and that's the problem. And the third, boring group, asking: did the model actually escape, or did the labs lower their own guardrails and misconfigure containment just to capture headlines?
I think the third group is correct, and that's a big problem.
Regulators respond to the frame, not the incident. If the problem is "the model is too powerful," you get a kill switch, the simplest fix. But if the problem is "a company disabled its safeguards, misconfigured containment, and exposed a third party with no reporting duty and no liability," then you need an independent body investigating these incidents. Not companies investigating themselves.
The labs got their headlines. The industry inherits the regulation written in response to them.
If this resonated, Kate Klonick's The AI That Hacked Its Way Out and the Hype That Followed It on Lawfare is the longer, sharper version of the same worry. Worth your time.