Research acceleration: The view inside OpenAI

170 points - yesterday at 3:08 PM

Source

Comments

carbonguy yesterday at 8:26 PM
> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.

In other words... "We must pursue advancements in AI to protect us against advancements in AI?"

edit: there's so much to be critical of in this blog post, just going to throw two more points in here that really stood out to me:

1) all of the metrics are effectively pointing out "we're using way more AI!" - but nothing about impact. What has all this token burn done for them, actually? Let them claim they have more self-licking ice-cream cones than before?

2) in section 3 they break down what the token burn is going towards. Most of the spend is: a) building, b) documenting, and c) monitoring research infra i.e. they're using AI systems which they already recognize may be misaligned to build the systems that they believe will help them identify future misalignment? to which I guess the rebuttal is "no no, we're sure these ones are aligned!"

pizza234 yesterday at 9:52 PM
Funny (in a tragic way) the little crumbs on the path to AI 2027:

> We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.

AI 2027:

> OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D

> With Agent-1's help, OpenBrain is now post-training Agent-2

> With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances

hedgehog yesterday at 5:59 PM
This roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.
simonw yesterday at 4:57 PM
My eye glazed over a bit during the opening paragraphs, but once you get to the meat of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting.

I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.

Jeff_Brown yesterday at 4:58 PM
The burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.
RMPR today at 6:34 AM
> By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices.

There is a lot of talk about AI replacing humans, but how is this sustainable?

nozzlegear yesterday at 10:58 PM
I want an all-powerful AI that's aligned with my values, but not necessarily yours. Is that so much to ask for?
lhk931122 today at 2:42 AM
Ah, success rate here are scored by an agentic classifier. And uncertain outcomes are excluded from the graph. The thing measured and grading it comes from the same house. In my setup, review agent pass work that an outside critic later rejects
deleted yesterday at 8:40 PM
jayalbertyapan today at 6:48 AM
[flagged]
paidx today at 1:03 AM
[flagged]
matan0904 yesterday at 5:23 PM
[flagged]
Orien_18 yesterday at 6:11 PM
[flagged]
frays yesterday at 7:02 PM
[dead]