Accelerating GPT-5.6 Sol Ultrafast

590 points - yesterday at 6:10 PM

Source

Comments

iamcoder18 yesterday at 6:55 PM
I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration.

> In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.

This is actually insane.

Hopefully the release ultrafast of Terra and Luna too.

csallen yesterday at 9:05 PM
People underestimate the importance of speed on quality of thought, because people underestimate just how much quality is a result of simple iteration.

When an LLM thinks, it typically just makes one pass. It outputs tokens from top to bottom, beginning to end, and then it's done. But when people think, especially strong thinkers, we typically iterate and revise our thoughts on the fly. We do numerous passes. We stop and restart, we reconsider, we review, we reevaluate. Sometimes we do this so quickly and automatically that we don't even realize we're doing it. I think a lot of what separates a highly intelligent or effective person from others has less to do with the quality of their first pass and more to do with just how many additional passes they're able to do in the same amount of time, and of course what kind of criteria they're habituated to consider during their review passes.

Introspecting about this is difficult, but experimenting with LLMs is easy. First, simply ask an LLM to do something complex. For example, to come up with a new business idea, or to plan the next month of your life, etc. After it finishes, tell it:

"Review what you just wrote, according to some appropriate list of evaluation criteria that you come up with first. And then, based on the results, iterate and generate a better response if warranted."

It's insane how much better the next answer will usually to be. Often it'll catch and erase tons of hallucinations, logical errors, and inefficiencies. And you can simply copy-paste this again and again until you begin to hit diminishing returns. Or, in a harness like Claude Code, for example, I might shortcut this whole process by saying, "Use sub-agents to iteratively review and iterate on your work until convergence."

The reason why most people don't prompt LLMs to do this (besides simply not thinking of it) is that it takes time.

But what if it didn't?

What if the LLM's response came back in milliseconds rather than minutes? Then there would be almost no reason NOT to do this. In fact, one could almost imagine it baked into the assistant/harness -- a massive step change in practical quality, enabled by nothing more than speed.

GodelNumbering yesterday at 6:37 PM
The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/

There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding

Topfi yesterday at 7:20 PM
Unless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance.

Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops. A line such as "this is the same performance, just faster, with no downsides" would go a long way in clarity and communication. Along with no pricing information, I'll hold out on further information.

[0] https://openai.com/index/previewing-ultrafast/

wxw yesterday at 6:37 PM
> Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

Awesome work. I'm personally very excited for faster models/inference.

I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast it was.

ricardobeat yesterday at 7:11 PM
The omission of Mimo v2.5-Pro Ultraspeed, released in June, which can achieve 1000tok/s is an interesting flaw in the comparison graphs.

It is a bit outdated (scores ± 40% lower), but smart enough for a lot of coding tasks, and can cost under 1/10th of Sol.

https://mimo.mi.com/models/en-US/mimo-v2.5-pro-ultraspeed

tristanMatthias yesterday at 8:31 PM
> GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second

https://taalas.com/products/

> delivering 17k tokens per second per user on Llama 3.1 8B model.

Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.

Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.

thraway3837 yesterday at 6:43 PM
This is really cool. Someone here commented about similarity between this and hardware advancements for AV encode/decode.

I think it's only a matter of time before miniaturization can have a thumbnail sized user-replaceable accessory that contains the LLM built onto the hardware. I admit I don't know how any of that works, but would be amazing to experience. Fully local, fully offline, ultra fast local inference better than any personal computing product.

aenis yesterday at 8:33 PM
Good news for Intel and AMD.

Rught now on large scale codebases the bottleneck is both claude/codex inference, as well as time it takes to run tens of thousands of tests. We put those workloads on dedicated epyc 9005 build machines - but it still takes minutes per run. Those who can afford the fast tokens will be in the market for faster CPU that money can buy today.

johnfn yesterday at 8:55 PM
This does look pretty incredible, but don't forget that incredible token thoroughput can only necessarily solve certain bottlenecks. If your e2e tests take an hour, they'll still take an hour after Ultracode. If the agent runs a 10 minute typecheck after a change, that will still take 10 minutes. grep over a massive codebase is still just as slow, etc. I say this not to take away from this accomplishment but just to ensure everyone here keeps a clear head about what it means - 14x faster tokens does not mean it completes every task 14x faster.

I suspect Humanity's Last Exam is without tool-calls, making it kind of the perfect benchmark to highlight how fast Ultrafast is, but not really the same as the everyday work you or I do.

owentbrown yesterday at 7:16 PM
Whoa. This looks both powerful and expensive.

My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference.

Within labs, I've heard spend is already far beyond this per developer.

anthonypasq yesterday at 7:15 PM
I'd just like to point out that the largest model Cerebras has ever served is Kimi K2.6 which is 1T parameters, so that either means that theyve had a breakthrough on the hardware engineering side of things, or GPT-5.6 Sol is likely a lot smaller than people think.

If it truly is only ~1-2T parameters, then this kinda kills 2 narratives for me.

1. all the handwringing about open source catching up via Kimi K3 (3T params) is complete nonsense. All that matters imo for determining which labs are leading is intelligence per parameter. Anyone with a enough compute can train a giant model, but being able to squeeze capabilities into smaller models gives you a massive inference and training edge.

2. Inference margins are clearly insane, and this explains why OpenAI was able to lower the price of Luna by 80%. Id guess that thing is probably 120b params based on the TPS they are serving it at.

huey77 yesterday at 11:09 PM
I think token output speed is going to be one of the biggest fundamental shifts for AI this year. In my experience, models figure out problems after enough turns (or in agent swarms if its a lower tier model). Compressing that time horizon could take days of agentic coding into minutes. How ever will my monkey brain keep up?
ttul today at 1:07 AM
With this level of intelligence offered at this level of speed, new real-time applications become possible, such as providing expert advice during a phone call or court hearing. Current SOTA models are too slow in many cases to provide the kind of insights that we expect to receive from an intelligent human colleague, such as a sales coach or lawyer handing us a note or writing a Slack message during a difficult call. For these real-time applications, even a 10x increase in per-token cost would often be tolerable.
exabrial today at 1:25 AM
Right now, the "economic model" of AI is "who has the best model", or really weights.

That'll go away eventually, just like operating systems eventually became free.

Instead, it's going to come down to selling inference hardware. We'll likely see the "apple" model where a custom OS runs on their hardware, but we'll probably also see more things like Cerebras become commodity hardware instead of kilowatt-class datacenter only hardware.

crazysim yesterday at 6:37 PM
GPT 5.6 Luna Ultrafast when?
johnnyApplePRNG today at 12:45 AM
Now I can blast through my weekly 20x pro codex credit in like an hour, great!

The amount of usage you receive on Codex these days is dismal compared to what it was a few months ago, FYI.

And they charge more for going faster.

As a Codex customer, I am not impressed with their shenanigans over the past few months and I have resolved to master the art of Pi Coding Harness creation and loving it.

Thanks for all the fish, Sam!

stillpointlab yesterday at 8:12 PM
I haven't wrapped my head around what level of reasoning this involves. Is it equivalent to max?

I didn't like Sol initially but it is growing on me the more I use it. Its personality is a bit flat and I caught it taking shortcuts a few times. But once I learned how to interact with it, I'm genuinely warming up to it. I find that it writes code that has fewer bugs even than Fable (although, to be fair I reach for Fable when the task is less well defined).

If this has similar performance to Sol at max reasoning level, this would be a compelling reason to shift even more of my work (maybe the majority) to this model.

z_rho_one today at 12:18 AM
Wonder how many X usage this would consume when it becomes available for everyone. Fast mode already consumes 1.5x usage for 2.5x speed. Hopefully, this does not mean 8x usage for 14x speed, but rather something more reasonable such as 4x usage.
buybackoff yesterday at 8:09 PM
This is something I'm ready to pay for. Not more per token, but I will be happy to burn through 20x Pro subscription as fast as I consume my Plus weekly limit now, with 10x more tokens per unit of time. I've learned how to deal with and steer Sol medium quite efficiently, but at the same time I realize it's so slow for the small tasks it can do well, and still so unreliable for open-ended tasks.
equinumerous yesterday at 11:39 PM
This is an amazing result. Can't wait until they release this to the general public, and I hope it's only a matter of time before other models are accelerated. I long for the day that regular consumers can run such models locally on specialized hardware.
sashank_1509 yesterday at 8:38 PM
I don’t know if this is that useful for coding. In some autonomous world, where no one check the code and the agent can just spend 10X more time checking its work and leading to better results, yes maybe it is useful.

But if humans need to check its work, then 10X speed doesn’t really matter I guess.

jdthedisciple today at 8:46 AM
now let's have this thing iterate away nonstop 24/7/365 solving humanity's major challenges and see how far we get, shall we?
fg137 yesterday at 7:10 PM
> allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work

Curious, what are some of the use cases?

HawtAds yesterday at 6:33 PM
Their dinner plate chips are impressive.
storus yesterday at 6:57 PM
Wow, that's even faster than diffusion LLMs but with the Fable-level quality! Congrats!
dewarrn1 yesterday at 9:48 PM
In light of recent news, it is hard not to think about the 4.5-day hack on Hugging Face's systems happening ~10 times faster and be slightly concerned.
Nevin1901 yesterday at 11:56 PM
Would gladly switch over to OpenAI and pay them 2x what I'm paying Claude if this becomes generally available
damsta yesterday at 9:28 PM
So if Fast mode is 1.5x faster at 2x the price, will Ultrafast cost 20x as much? $100/$900 per 1M tokens?
andrethegiant yesterday at 10:21 PM
Why did OpenAI partner with Cerebras when they've already built their own chip, Jalapeño?
minraws yesterday at 11:09 PM
I will happily pay 2x luna pricing for this speed with Luna :P
ilaksh yesterday at 8:19 PM
Did Cerebras get rid of their like $1500 per month plans for open models?
scotty79 yesterday at 6:57 PM
I swear that now frontier AI stuff comes out few times a week.
pingou yesterday at 7:28 PM
Meanwhile they are down 12,68% today because of disappointing earnings.
behnamoh yesterday at 7:08 PM
Fast mode is already 1.5 times faster and 2x more expensive in the Codex subscription plan. If this thing is 14 times faster, then I can imagine running out of my quota in one session.
poly2it yesterday at 6:39 PM
I guess Gemini 3.7 Flash is no longer at the pareto frontier of speed to intelligence.
yieldcrv yesterday at 11:51 PM
I'm loving this competition (now that I can see how workflows keep us employed)

All the frontier labs go seemingly dormant for a month or two, while another one has its flurry of press releases, and people start to question whether the other lab is doing anything and then boom, the other lab finishes baking its next thing and releases its flurry of press releases

lostmsu yesterday at 8:45 PM
Still no KV caching?
Marciplan yesterday at 8:23 PM
“our stock price went down today, here’s something to feed it”
applfanboysbgon yesterday at 7:07 PM
This kills the crab.

Compilation time will be a genuine bottleneck for slop coding if this becomes the standard generation rate over the next few years. Go, Zig or even C99 with TCC for dev builds, any language that can get you systems-level performance (or close to it) in a dev environment where you can iterate in ms rather than minutes is going to be immensely more appealing than generating a potential prototype in 10 seconds and waiting 15 minutes for it to compile.

madhu_ghalame today at 8:05 AM
[dead]
fenestella today at 2:22 AM
[dead]
stephencoyner yesterday at 7:43 PM
[dead]
Jr23_xd yesterday at 10:06 PM
[flagged]
deleted today at 8:02 AM
huflungdung yesterday at 6:33 PM
[dead]
_345 today at 5:42 AM
Hope to be a part of this someday...