I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
paxystoday at 5:11 PM
Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
6thbittoday at 5:33 PM
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
Dibestoday at 5:21 PM
I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code.
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
not_a9today at 5:11 PM
> Opus 5’s safeguards match
those of Claude Fable 5’s, with one change: it now permits source-code vulnerability
discovery at all access levels. This means that the model can support defensive
cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
theHocineSaadtoday at 6:36 PM
Opus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2].
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
abroszka33today at 5:16 PM
What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
Sol-today at 5:11 PM
How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
ealready_valuetoday at 5:18 PM
I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
yewenjietoday at 5:35 PM
Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
wuhhhtoday at 8:06 PM
It really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
thewebguydtoday at 5:09 PM
> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
MasterScrattoday at 9:17 PM
Damn the pelican guy can’t get no sleep
pyridinestoday at 5:17 PM
The wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions.
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
ceberttoday at 5:11 PM
I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
artninja1988today at 5:14 PM
That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
guybedotoday at 6:17 PM
Looking at intelligence vs cost:
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost.
- Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
hrpnktoday at 7:23 PM
The breaking changes vs. Opus 4.8 are interesting [1]
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
guess_who_istoday at 8:42 PM
I have started distilling
ddxvtoday at 5:11 PM
"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
alasanotoday at 5:23 PM
Half the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems.
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
tekacstoday at 5:53 PM
Something fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
trunnelltoday at 7:35 PM
The chaos appears to be tamed for now.
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
irthomasthomastoday at 6:11 PM
Changelog
- fixed issue where model acts like qwen when prompted in chinese
consumer451today at 7:15 PM
I have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
bottlepalmtoday at 6:04 PM
Page 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
paxystoday at 7:02 PM
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
born-jretoday at 7:07 PM
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
albert_etoday at 5:20 PM
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
lucamarktoday at 6:07 PM
But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.
I've never trusted on model cards though. I'm sorry.
luciana1utoday at 8:04 PM
473 comments in 3 hours. people are speedrunning having opinions about it
the_lucifertoday at 5:09 PM
Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
williamsteintoday at 5:43 PM
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Annoyingly, this is a concrete argument that open source software may be easier to attack.
itissidtoday at 7:41 PM
I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
skybriantoday at 5:51 PM
Looks like the API price in tokens is same as previous Opus or Sol, double the price of Terra.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
destringtoday at 5:17 PM
Google is having their Meta moment where they failed to stay at the frontier
skerittoday at 5:31 PM
Interesting, they finally support `system` messages anywhere in a chat conversation:
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud.
>
> This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird.
How old is Sonnet 5 really?
deletedtoday at 6:18 PM
jatinstoday at 5:12 PM
Better than Fable 5 on all but 3 evals.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
vatsachaktoday at 5:57 PM
GPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
dehuggertoday at 5:16 PM
Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
6thbittoday at 5:35 PM
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
beydogantoday at 7:18 PM
my early and non scientific feeling:
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
markasoftwaretoday at 5:21 PM
Soo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
rad_valtoday at 6:37 PM
After Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
vinhnxtoday at 6:24 PM
The benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
geooff_today at 5:18 PM
FYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
stevefan1999today at 5:51 PM
Where's the reset...
arjietoday at 6:48 PM
I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
theplumbertoday at 6:51 PM
The most important thing is it has the same drama queen mode on safety “guards” like Fable.
alvistoday at 5:02 PM
What really impress me is opus 5 is better in alignment than fable 5!
tyretoday at 5:32 PM
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
boctoday at 5:43 PM
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
bovermyertoday at 5:48 PM
This stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
holoduketoday at 8:47 PM
Is it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.
briandolltoday at 5:00 PM
Very interesting to see such a focus on cost for performance here
shockemboppertoday at 5:56 PM
I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
m_w_today at 5:03 PM
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
pmg1991today at 5:10 PM
Same cost as 4.8 but better that 4.8. Happy to get more efficient model.
But is there any reason all companies are releasing models back to back after GLM 5.2.
twothreeonetoday at 5:06 PM
It starts at page 148.
arrowleaftoday at 5:27 PM
I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
toephu2today at 6:06 PM
How does it score on DeepSWE?
throwaw12today at 5:12 PM
is coding and engineering solved yet?
_pdp_today at 6:55 PM
Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.
abc42today at 6:18 PM
Are we getting to singularity or something? This seems a bit crazy.
zuzululutoday at 5:06 PM
so almost fable 5 with 50% cheaper cost? sign me up
deletedtoday at 5:12 PM
hmontazeritoday at 5:30 PM
Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
mihautoday at 5:29 PM
30% on ARC-AGI-3
LoganDarktoday at 6:15 PM
These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
throwaway23597today at 5:55 PM
The truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.
Maybe I'm wrong and Opus 5 is a real unlock?
simianwordstoday at 5:42 PM
My thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
sbochinstoday at 5:25 PM
Quick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
alvistoday at 4:56 PM
here we go
ismailmajtoday at 6:39 PM
I'd pay good money to see OpenAI "oh fuck" war rooms.
sudohalttoday at 6:06 PM
Anthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
mrcwinntoday at 5:56 PM
Can someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
alex1138today at 7:55 PM
Am I misreading anything or are comparisons to Fable (and/or Mythos although AFAICT it was only a crackdown on Fable) always going to be a bit missing the mark now due to what the Trump admin did?
StrauXXtoday at 5:17 PM
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
wyretoday at 5:14 PM
In the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
mnky9800ntoday at 5:48 PM
Yay just in time for neurips lol
justindotdevtoday at 5:11 PM
> . Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
ffs just keep it man.
marsven_422today at 9:14 PM
[dead]
deletedtoday at 8:19 PM
gorkemyildirimtoday at 7:48 PM
[flagged]
vilmiretoday at 6:05 PM
[flagged]
nee_oo_rutoday at 5:43 PM
[dead]
emunovatoday at 6:33 PM
[dead]
Nevin1901today at 5:08 PM
Excited to use it? Will we be seeing Haiku 5 next? /s
datakantoday at 5:02 PM
> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5
Ok then so what's the point?
aleenz1102today at 5:15 PM
this claude fable & opus 5 should be cheaper and can compete in pricing with chatgpt latest models
TheJCDentontoday at 6:08 PM
I think it's the first time Anthropic release a model without any meaningful disruptions while doing it
midnightbobaruntoday at 5:12 PM
It looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')