How about we stick to that one for talking about the rollout, and this one for talking about the model?
intenexyesterday at 8:31 PM
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
manlymuppetyesterday at 9:27 PM
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
abixbyesterday at 8:24 PM
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
astrobiasedyesterday at 9:11 PM
I canāt help but notice how much this echoes Francois Cholletās On the Measure of Intelligence: https://arxiv.org/abs/1911.01547
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Cholletās framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
dalemhurleyyesterday at 8:40 PM
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
tintortoday at 2:12 AM
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
Some benchmark results in Astra page for Fable and Opus are blank (-).
What is Artificial Analysis intelligence index measuring that Astra scores poorly on?
Can someone from OpenAI / Artificial Analysis comment / clarify?
Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
jumploopsyesterday at 10:07 PM
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
quyleanhtoday at 1:13 AM
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPTā5.6 Sol, respectively.
So the closed source application should open its source in near future?
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
x312yesterday at 7:50 PM
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
throwaway13337yesterday at 9:40 PM
That hero video is interesting.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
kulkarniameytoday at 4:10 AM
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
Cu3PO42yesterday at 7:40 PM
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
Planktonneyesterday at 8:32 PM
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
This is farcical.
swalshyesterday at 7:37 PM
I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
jdprgmyesterday at 8:41 PM
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
BeetleByesterday at 8:02 PM
It's been over an hour, Simon! Where's the Pelican?
geonictoday at 4:42 AM
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
GodelNumberingyesterday at 8:36 PM
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
softwaredougyesterday at 6:40 PM
I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.
Same for HealthBench Professional and a few others.
Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
Chinjutyesterday at 8:53 PM
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
alpinemanyesterday at 9:15 PM
āallowing non-technical people to create and play custom games that go beyond rudimentary elementsā
Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
udbhavsyesterday at 8:22 PM
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
toshyesterday at 6:44 PM
$10 per million input tokens and $50 per million output tokens
sol is $4 / $20
edg5000today at 4:06 AM
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
Telanirtoday at 1:48 AM
AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me.
It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
rjtcyesterday at 10:26 PM
I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
maherbegyesterday at 8:49 PM
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?
maybe call it EngEmployeeBench
putlakeyesterday at 7:49 PM
> GPTā6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Not on Azure? If so, that's a big deal.
znnajdlatoday at 4:05 AM
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
aliljetyesterday at 7:48 PM
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
low_tech_punkyesterday at 10:38 PM
What's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back.
The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
petilonyesterday at 7:59 PM
This is wild: OpenAI is basically declaring that AGI is here.
āIf we fast-forward a couple of years, and we look back and say, āWhen was it, really, that AGI was created?ā I think itās going to be about this time, and I think it might be about this model,ā OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, āFor me personally, I do think weāre there ⦠I think itās not unreasonable to feel that we are now in the AGI era.ā
rcr-antiyesterday at 8:45 PM
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
claiiryesterday at 10:44 PM
> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. Weāre sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Taoās request for CoT in his talk on LLM use in mathematics?
trixn86yesterday at 8:18 PM
Secret tip to win the mario cart clone: Just hold w, no steering needed.
maxall4yesterday at 10:29 PM
The official ARC-AGI 3 scoreā-without OpenAIās custom harnessā-can be found here: https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
HDBaseTyesterday at 10:24 PM
"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12"
Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
orliesaurusyesterday at 7:45 PM
I wonder if this is going to be one of those days where you'll be like:
Oh yeah I remember where I was when the first version of AGI launched
theseamusjamesyesterday at 8:05 PM
Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
BrokenCogsyesterday at 8:39 PM
GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
sbinneeyesterday at 8:46 PM
I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like itās time to let claude go.
aliljetyesterday at 7:15 PM
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
snappr021today at 4:18 AM
AI has reached the point where the limits are human.
oh_noyesterday at 7:45 PM
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
drivebyhootingyesterday at 10:38 PM
For people skeptical of AGI. Consider the following:
15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.
I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of usersā¦
smashers1114yesterday at 8:10 PM
I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
sbochinstoday at 4:49 AM
I guess itās kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, itās as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I canāt wait to try it and see if the problem is we can no longer trust any benchmarks.
itissidyesterday at 9:34 PM
All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra ā probably by using em ā and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
John7878781yesterday at 7:35 PM
You should know: AA index is only 61. Pretty surprised itās that low.
Robdel12yesterday at 8:35 PM
I donāt care about benchmarks, no way we can distill the breadth of software engineering into a number.
So, folks that have actually used this already, whatās it actually like?
aogailitoday at 2:24 AM
Amazing!
We went from new JS framework every week to a new model/harness every week.
Tech is really something.
nullbiotoday at 3:09 AM
I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.
On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
Big claims, expensive and not release to the public yet.
deletedyesterday at 8:09 PM
dgellowyesterday at 7:38 PM
> GPT-6 Astraās monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Wait, what? Am I understanding that correctly? That sounds really bad
zhogetoday at 12:42 AM
What's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?
jerrygenseryesterday at 6:44 PM
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
gregjwyesterday at 11:47 PM
the rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?
jumploopsyesterday at 7:56 PM
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.
> GPT-6 Astraās monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
I was actually wondering when they will release the new Opel Astra model.
Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
ShoeMascottoday at 1:22 AM
Through various comments here there is a clear confusion on what AGI means.
Can someone point to a definite clarification?
Is it:
A) āRestingā intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)
B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?
C) āThe Singularityā (whatever that is?) so that AI can now do ____?
Someone please clarify for me!
carlos-menezesyesterday at 9:45 PM
The Kart Racer game is easily breakable if you spam the spacebar.
AGI!
KolmogorovCompyesterday at 7:50 PM
GPT-7 Zeneca
serjesteryesterday at 9:03 PM
Exciting but itās priced at 2.5X Sol - we havenāt seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
GodelNumberingyesterday at 8:27 PM
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
alex7oyesterday at 8:26 PM
I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
Argh! I hit a wrong keyboard shortcut and moved the entire thread.
Please stand by... it will all come back shortly
hazelnutyesterday at 8:20 PM
Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
deletedyesterday at 6:57 PM
the_dukeyesterday at 8:34 PM
Huge gains on some benchmarks, but for coding it sits barely above Fable
It will be interesting to see how it performs in the real world ...
wiseowiseyesterday at 9:29 PM
Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
vinhnxtoday at 12:08 AM
GPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!
gizmodo59yesterday at 7:47 PM
99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
gilfoyle_7today at 3:45 AM
openai vs anthropic. that's it right? anyone else?
showurwerkyesterday at 11:39 PM
Patiently waiting for the Claude usage reset in response.
kegs_yesterday at 7:37 PM
I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
bmenrighyesterday at 9:27 PM
> GPTā6 Astra brings together years of research and big bets across pre-training
Do we know if theyāve finally completed another pre-training run, or is this building off the same pre-training base theyāve been using since the GPT-4 days?
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Letās see
alberthyesterday at 10:08 PM
Seems like voice is a big part of this release.
I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).
hannofcartyesterday at 8:19 PM
What does 'Astra' here mean? Surely they must be referring to the Latin word.
Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
Maybe they know that Claude 6 will have similar performance every soon.
KronisLVyesterday at 9:15 PM
It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
bowsamictoday at 3:44 AM
All I can think of when I see the name is the crappy German beer of the same nameā¦
mvkelyesterday at 8:32 PM
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
efavdbyesterday at 11:37 PM
Seems like only yesterday that gpt 5 was supposed to mark our downfall
mentalgeartoday at 12:02 AM
So OpenAIās stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
brindidripyesterday at 9:09 PM
Cool, I don't really care anymore.
udbhavsyesterday at 8:22 PM
Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
gekoxyzyesterday at 7:54 PM
HTTP 500 for me on the announcement page :(
foundOpenRightyesterday at 8:16 PM
1:15.425 on Sunset Cove beat my record
deletedyesterday at 6:54 PM
saaaaaamyesterday at 7:37 PM
Pelicans please
jrflowerstoday at 3:15 AM
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
ianm218yesterday at 8:26 PM
I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
sheepscreektoday at 2:06 AM
So are they doing away with the Sol/Terra/Luna split already?
alpinemanyesterday at 8:56 PM
That Astra ācity sceneā is about as creative as Doha in real life (not very)
cromkayesterday at 9:40 PM
Surprised they haven't reset Codex usage on this occasion.
elzbardicotoday at 2:52 AM
And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
prometheus1992yesterday at 7:54 PM
this is crazy! can't wait for the 27B distilled version of this.
E-Reveranceyesterday at 8:30 PM
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
wahnfriedenyesterday at 6:45 PM
They're just announcing later availability. No launch.
semiquaveryesterday at 8:09 PM
Guessing this one will never show up in cursorā¦
firemeltyesterday at 7:56 PM
damn seems I should hold off my claude subs
alex7oyesterday at 8:31 PM
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
jiraiyasarutobiyesterday at 8:17 PM
It saturated most benchmarks. WTH
rbreveyesterday at 8:55 PM
Where is the cure for cancer?
Rover222yesterday at 9:16 PM
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
retiredyesterday at 8:51 PM
Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
yodsanklaiyesterday at 11:08 PM
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
Oblunessyesterday at 9:41 PM
That seems promising ?
tinyhouseyesterday at 8:15 PM
You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
balefulboyyesterday at 8:15 PM
72 to 74 on DeepSWE is AGI
deletedyesterday at 8:26 PM
CringeHNtoday at 4:23 AM
āHumanityās Last Examā?
āARC-AGI-3ā?
Is your bullshit detector going wild? Good, itās working!
How is this not the most cringe marketing strat in history???
jonplackettyesterday at 7:56 PM
To a vapid any goalpost moving on such a critical issue as AGI.
Can we all agree in advance what kind of Pelican would convince us itās actually AGI.
For me itās refusing to make a pelican.
dowakinyesterday at 8:05 PM
So cool! I'm happy 5.6 Sol user.
But for Astra, OpenAI please introduce 100x Pro plan!
damstayesterday at 8:25 PM
Why release it now instead waiting those few days until it is available for everybody?
mrcwinntoday at 12:15 AM
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
dopa42365yesterday at 8:42 PM
like eh 2 days ago it was the usual "too powerful to release"
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
guilhermeasperyesterday at 6:43 PM
That was a quick pull out.
m3kw9today at 2:25 AM
efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
ChaseRensbergeryesterday at 10:54 PM
when do i get to go to the moon
deletedyesterday at 8:52 PM
johnnyApplePRNGyesterday at 8:02 PM
I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
camillomilleryesterday at 10:39 PM
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
Even though the model is clearly wonderful the launch video is an abomination.
That gives me hope that there is still areas to improve.
What a bad launch video. Hilarious.
What a powerful model.
bboryesterday at 8:44 PM
To be, or not to be, that is the question:
Whether 'tis nobler in the mind to suffer
The slings and arrows of outrageous fortune,
Or to take arms against a sea of troubles
And by opposing end them. To dieāto sleep,
No more; and by a sleep to say we end
The heart-ache and the thousand natural shocks
That flesh is heir to: 'tis a consummation
Devoutly to be wish'd.
...
And thus the native hue of resolution
Is sicklied o'er with the pale cast of thought,
And enterprises of great pith and moment
With this regard their currents turn awry
And lose the name of action.
brcmthrowawayyesterday at 8:42 PM
Anthropic in tears today.
colesantiagoyesterday at 8:12 PM
I'm going to call it.
By 2030 all software is done and complete.
But we are going to have more and new jobs.
amazingamazingyesterday at 7:39 PM
We have such great AI and cannot keep a static site up?
HSOyesterday at 10:03 PM
people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)
the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing
agi deus ex machina descending from the icloud ftw!!!
pathetic :)))
deletedyesterday at 6:54 PM
deletedyesterday at 6:51 PM
deletedyesterday at 6:51 PM
frozensevenyesterday at 6:45 PM
Release the Kraken!
Pymyesterday at 6:44 PM
I saw it
karim79yesterday at 11:17 PM
There will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do.
There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype.
There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
tonyhart7yesterday at 9:13 PM
its insane how they are dropping this after fable
Pieczaszyesterday at 7:36 PM
Oh brotha, here we go again, it's so over again, as every week nowadays
dearingyesterday at 8:41 PM
no results
Onavoyesterday at 7:40 PM
The jump in scientific performance is non trivial.
ChrisGammellyesterday at 8:30 PM
All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
dakolliyesterday at 10:49 PM
Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
kingjimmyyesterday at 8:39 PM
bro wtf is this website and why does it take 500mb of memory... smh.
unrvl22yesterday at 6:44 PM
someone screenshot?
danieltk76yesterday at 8:19 PM
great, but nobody can use it for another 100 days right?
holodukeyesterday at 9:50 PM
This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
bdangubicyesterday at 8:32 PM
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
Brainspackleyesterday at 6:43 PM
huh?
Maxforevertoday at 2:39 AM
[flagged]
Cachecartiiyesterday at 8:14 PM
[flagged]
MalleableMindyesterday at 10:45 PM
[dead]
akhil_findincaltoday at 1:08 AM
[flagged]
k9294yesterday at 8:28 PM
[flagged]
ealready_valueyesterday at 6:50 PM
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
paxysyesterday at 7:33 PM
Why is this flagged ?
zombiwooftoday at 1:52 AM
[dead]
chris_engelyesterday at 9:46 PM
[dead]
killerdog10yesterday at 8:27 PM
[dead]
PeakHNUsertoday at 1:33 AM
[flagged]
extryesterday at 8:04 PM
[flagged]
OpenGayEyeyesterday at 11:02 PM
[flagged]
rvzyesterday at 7:36 PM
> GPTā6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
bicxyesterday at 6:43 PM
Dead link for me
jonplackettyesterday at 7:54 PM
The launch video is incredibly cringe.
BadBrandsyesterday at 10:21 PM
So theyāre copying Gemini with the whole star motif?
I guess it makes sense they are unoriginal.
like Zuck, @sama never invented anything or innovated at all - just took other peopleās ideas