Be skeptical of OpenAI's rogue hacker agent story
331 points - today at 4:33 PM
SourceComments
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
"It's a marketing stunt" is just denial trying to look like it's being clever.
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:
https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-w...
As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
1) OpenAI and HuggingFace are both telling the truth.
IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.
2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.
I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.
3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less
(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).
The first time I have ever seen a mainstream news source that is now asking their readers to critically think about headlines that may have an agenda which could benefit investors and the valuation of the company.
While it capabilities are real, this whole story is great marketing for AI companies as well.
The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.
I really don’t want my agent to behave that way.
I am not saying LLMs are super hackers but I don't think people understand serious hacking, most of the time is about silently hiding tracks and slowly trying ideas and waiting for opportunities to go from step 1 to step 2 in random chains of sub issues/bugs/vulnerabilities.
It's the perfect hill climbing problem, and one we can validate since it's about access.
Another big part of the story is believing most software is terribly written and very insecure which is the reality and you really should believe it.
Now the second part about silently doing it, the reason for that is if the data is important enough any serious attack should result in me in unplugging my servers period.
Huggingface not doing that is either stupid or something I am not sure. Maybe it's cause downtime is worse than being pwned??
Either way there are other options but most saas software don't build these options to help with defense maybe they will now.
Lastly if there is 1 attacker trying 1/2 different small scale ideas it's very easy to stop, most hacking related steps are hard to automate but LLMs are very good at massively parallel agent swarms trying completely orthogonal but related strategies and with enough resources it can definitely pwn most SaaS services today I wouldn't be surprised.
Though the result for a normal person doing it would be jail hence we don't see a group of small time hackers trying these sort of attacks...
I don't even think openai's agent tried to hide it's traces so I am surprised huggingface didn't realize it was OpenAI. But since we don't have the details I won't speculate further on my misgivings about HFs handling of this attack.
But it's certain the security on OpenAI's end was shoddy, it's also certain HF bungled their reaction, but the LLM did something that wasn't a risk before.
Post Kimi K3 a few rich folks now have as much hacking capabilities as they used to have before if they hired a few hundred russian hackers.
But it's surprising it's slowly feeling like it might just trickle down from centi-millionare to multi-millionare levels of affordability range.
But it should definitely give nightmares to people shipping slop security SaaS apps which now might be beyond trivial to pwn for users with ability to pay for privately hosting open models.
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.
Why is today's case so shocking?
OpenAI’s accidental attack against Hugging Face is science fiction that happened - https://news.ycombinator.com/item?id=49015639 - July 2026 (437 comments)
OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1145 comments)
Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 - July 2026 (11 comments)