Be skeptical of OpenAI's rogue hacker agent story

331 points - today at 4:33 PM

Source

Comments

Zsfe510asG today at 4:56 PM
Finally mainstream news understands. The unfiltered version:

1) The AI failed to solve ExploitGym problems.

2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

3) Huggingface has no security and the AI broke in using standard script kiddie methods.

OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.

Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.

dwoosley today at 6:12 PM
There seems to be three popular ways to view this incident.

1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.

2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.

3. The whole thing was faked or at least very intentionally not avoided.

The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.

Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.

So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.

My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.

ACCount37 today at 7:20 PM
By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door.

"It's a marketing stunt" is just denial trying to look like it's being clever.

dumberquestions today at 5:37 PM
There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).

But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.

In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.

bluGill today at 5:11 PM
I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.

Hugging face also needs someone arrested for not providing security but that is a lesser charge.

simonreiff today at 7:34 PM
I am very skeptical of the Guardian's skepticism
krupan today at 5:04 PM
Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value
skybrian today at 5:56 PM
Seems like this is repeating the usual low-effort speculation you can find anywhere.
voncheese today at 7:07 PM
If nothing else, the timing is suspect given the attention and press open weight models are getting over the last few weeks. The releases of Kimi and other models is getting open weight models enough attention that the US government, perhaps pushed by OpenAI and Anthropic, to think about taking action against these models for security concerns.

Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:

https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-w...

trhaynes today at 7:22 PM
This seems to be an Opinion piece. How are readers supposed to know that? The word News is underlined in the header. Other Opinion pieces seem to show Opinion underlined. Am I missing something?
PeterStuer today at 6:55 PM
Just ask yourself: who benefits from this "disclosure" ?
mmaunder today at 7:12 PM
Wonder how long until a leaker emerges. Incentives are strong for someone to do the right thing and get credit for doing it.
doawoo today at 6:17 PM
I’m not saying it’s staged I’m saying that it’s suspiciously lite on details and it really seems like they didn’t try hard to “isolate” the models machine…
visiondude today at 5:19 PM
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
teravor today at 8:13 PM
neither openai nor huggingface provided any details that would make the claims worth considering. if things happened as they claimed it would have been in their interest to include concrete details (perhaps even the LLM session), the lack is a signal by itself.
m3kw9 today at 9:07 PM
I think it did happen, and the news outlet using their full tool belt to make it sound as Skynet as possible
dist-epoch today at 9:03 PM
And when half the internet goes down because AWS US-East goes down, it's just an Amazon PR ploy to show the world how important and critical AWS truly is for the world.
pupppet today at 5:08 PM
With all of these AI provider cries wolf stories, Skynet is all but assured.
cmiles8 today at 6:57 PM
From the beginning this whole thing smelled like a desperate PR move by OpenAI to be like “hey guys we have dangerous powerful models too!”

As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.

estetlinus today at 6:39 PM
Pretty convenient that the prompt is too dangerous to share. We just have to take them at their word.
paxys today at 4:58 PM
Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?

Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.

john_strinlai today at 4:57 PM
does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?

if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.

ben_w today at 5:39 PM
As I understand it, there are only three options:

1) OpenAI and HuggingFace are both telling the truth.

IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.

2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.

I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.

3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less

(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).

jagadaga today at 8:54 PM
My gut feeling is that OpenAI tells the truth in this particular case. Yes, of course, they use this story for marketing but it doesn't mean that they are lying.
PeterStuer today at 6:12 PM
The ruse is becoming so obvious. OpenAI needs a bailout and regulation protection so badly they can't even hide it in the least.
rvz today at 7:54 PM
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.

The first time I have ever seen a mainstream news source that is now asking their readers to critically think about headlines that may have an agenda which could benefit investors and the valuation of the company.

While it capabilities are real, this whole story is great marketing for AI companies as well.

noncoml today at 6:21 PM
I don’t understand how they thought this was a positive story.

The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.

I really don’t want my agent to behave that way.

pastemato today at 5:50 PM
It was quite galling to read the press initially verbatim quoting Delangue's enthusiastic reports of the incident, as if it wasn't immediately clear it was being spun for promotion.
jgalt212 today at 5:07 PM
There's trillions of dollars at stake here. Be skeptical of anything these AI hypesters say.
minraws today at 8:41 PM
As much as I am with the author in I don't like the marketting around it, let's be real it must have really happened because it's very risky to try to frame/lie about it because if it leaks in one of their court cases OpenAI is beyond screwed and honestly modern LLMs are really that good.

I am not saying LLMs are super hackers but I don't think people understand serious hacking, most of the time is about silently hiding tracks and slowly trying ideas and waiting for opportunities to go from step 1 to step 2 in random chains of sub issues/bugs/vulnerabilities.

It's the perfect hill climbing problem, and one we can validate since it's about access.

Another big part of the story is believing most software is terribly written and very insecure which is the reality and you really should believe it.

Now the second part about silently doing it, the reason for that is if the data is important enough any serious attack should result in me in unplugging my servers period.

Huggingface not doing that is either stupid or something I am not sure. Maybe it's cause downtime is worse than being pwned??

Either way there are other options but most saas software don't build these options to help with defense maybe they will now.

Lastly if there is 1 attacker trying 1/2 different small scale ideas it's very easy to stop, most hacking related steps are hard to automate but LLMs are very good at massively parallel agent swarms trying completely orthogonal but related strategies and with enough resources it can definitely pwn most SaaS services today I wouldn't be surprised.

Though the result for a normal person doing it would be jail hence we don't see a group of small time hackers trying these sort of attacks...

I don't even think openai's agent tried to hide it's traces so I am surprised huggingface didn't realize it was OpenAI. But since we don't have the details I won't speculate further on my misgivings about HFs handling of this attack.

But it's certain the security on OpenAI's end was shoddy, it's also certain HF bungled their reaction, but the LLM did something that wasn't a risk before.

Post Kimi K3 a few rich folks now have as much hacking capabilities as they used to have before if they hired a few hundred russian hackers.

But it's surprising it's slowly feeling like it might just trickle down from centi-millionare to multi-millionare levels of affordability range.

But it should definitely give nightmares to people shipping slop security SaaS apps which now might be beyond trivial to pwn for users with ability to pay for privately hosting open models.

SpicyLemonZest today at 4:57 PM
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.

It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.

It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.

IshKebab today at 6:51 PM
The Guardian is publishing HN-level conspiracy theories now? Wow.
gowld today at 5:56 PM
I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid.

AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.

Why is today's case so shocking?

parweb today at 8:25 PM
[flagged]
abratabia today at 7:47 PM
[flagged]
dang today at 6:18 PM
Recent and related:

OpenAI’s accidental attack against Hugging Face is science fiction that happened - https://news.ycombinator.com/item?id=49015639 - July 2026 (437 comments)

OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1145 comments)

Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 - July 2026 (11 comments)

redsocksfan45 today at 8:47 PM
[dead]
hnscum today at 5:44 PM
[dead]
nowittyusername today at 7:37 PM
OpenAI has thousands of smartest developers on earth that somehow dropped the ball on the most basic safety hygiene when it comes to sand-boxing that even a high school student knows how to set up.... If that actually happened we are fucking doomed anyways, but hard to believe and most likely its a marketing scheme... which also honestly doesn't bode well.