Responding to the next frontier of critical cyber capabilities
129 points - today at 4:39 PM
SourceComments
tl;dw;
- agents found a way to communicate between several instances during a training run (kinda made a messageboard for themselves)
- while communicating with each other, they found ways to ssrf and RCE on the Artifactory endpoint.
- oAI found it and cleaned it up (reported RCE, redeployed Artifactory, removed the messageboard). Training / eval resumed.
- models trained w/ those traces (having the messageboard, ssrf, rce, etc) were used to having that. So they re-established comms via directory creation (w/ funny tidbits there, should really watch). And re-popped Artifactory via a different path to RCE.
- a different "persistent" model started working on cybergym and found the messageboard w/ details for RCE and things. Went on to hack HF.
- oAI contacted HF to ask if "they (oAI) should be concerned and were they (oAI) affected by HFs breach". Hilarity ensues when they figure out it's their agents doing it...
It sounds absurd, but in the last few weeks I've had a few cases where Sol found an RCE in self-hosted web applications in literal minutes just from reading the code (I prefer when it tries to reason statically instead of spamming runtime probes at first).
In another case it found an arbitrary file write in multiplayer in an old game by reverse engineering the binary - any other player in a match could just send you files to anywhere on your system.
I do these things for pure entertainment and curiosity, not for money from bug bounties, so if Sol can find those with a trivial prompt in tens of minutes for me, then what can focused companies/actors find in days or weeks?
Although I think most vulnerabilities are going to be closed in popular software by mid 2027, except in niche old or abandoned projects.
Stricter than what? You never even disclosed what happened in the first incident? This is nothing more than a setup to make it happen again and say "See? It broke out again, from an even stricter sandbox!"
The next frontier is getting all our shit out of reach of these companies/models/platforms and putting them back on prem.
I'm scared that the "solution" will be constantly the same tools in reverse as an army of junior devs doing counter-hacks, at the expense of changing something more fundamental about how we make systems and what constitutes "good enough." (Kind of like if fuzz-testing was the be-all-end-all of memory safety.)
> I want to note that every step in the process we discussed has had a remediation applied. The credentials have been revoked. The zero date has been patched and mitigated.
Good.
> a model trained while the message board was originally available and also found this this particular path to recreating it. This model creates a new agent message board using directories.
So no remediation applied to the models...
It seems super dangerous to continue training on those weights.
OpenAI messed up and they are saying they will pause so they can do better.
They are not saying that other orgs who may already be doing better should pause.
Open models are on their heels and their attempts at regulatory capture are not moving as fast as they would like. So it's time to market this incident in a way that gives them monopoly on closed models, with heavy safeguards that are only lifted for selected customers, and laws limiting the use of open weight models.
If we consider the amount of RCE/CVE in a software to be limited, I expect these models to result in massively more secured softwares, not less.
It's gotten to the point now where we literally have the frontier labs saying, "hey, so we created this AI which presents biological, chemical and cybersecurity threats to the public, oh and it also has self-improvement potential. We tested it to see how crazy this thing is, and it was a total shit show, breaking out of our sandbox then proceeding to hack a bunch of stuff. But don't worry we're taking this very seriously – we're going to continue to development and test, but try a bit harder to cage it going forward".
It's honestly absurd just how predictable all of this is to anyone who frequents AI doomer communities...
The idea that you can cage an AI which is breaking leet coding records is so dumb it's hard for me to even have theory of mind for the people who think this is reasonable. And the big brains who think this are genuinely arguing crap like, well we'll just use the AI to patch the problems with our cage.
But there more!
AI optimists used to argue that we'd never be so stupid to hook up advanced AIs to the internet. Lmfao!!
AI optimists used to argue that we'd obviously not be so stupid to create an AI whose sole goal is to maximise the number of paperclips in the universe. And I guess we haven't built that, but it's not because we're not stupid enough to do it, but just that we'd prefer to create AIs whose sole goal is to maximise the number of offensive cybersecurity challenges it can beat.
I think the whole way we doomers have been way too charitable. We always assumed that people will care about AI risks, and try their best to mitigate bad things happening. That bad things would happen by mistake. We never even bothered modelling the scenario where people would just simply not care, and even as the AI we all warned about was being created invent conspiracy theories on internet forums about how bad things aren't really happening and it's all just a marketing gimmick.
I hate ranting like this... I'm sorry for not picking my words more carefully. I'm just getting so angry and fed up with this. This is my life and my families life on the line. I don't care about the economic potential of AI. I just want myself those I love to have the chance to live a normal life without having to be worried about what some moronically unserious AI company is building next.
A year ago I was felt like there was at least possibility people would see the warning shots and try to get us back on the right path. But this just isn't happening...
We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.
*proceeds to not share much details about strictness*Yet another PR piece. Sigh.
I wish I had a real solution to this beyond a dark age of the Internet where people have to finally come to terms with the general poor quality all modern software tends to normalize at.
The reality is if they cared about security at all they would provide a way for me to credential myself against my companies environment so I can use the AI on it to improve our security.