simonw4 hours ago
This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium
I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.
I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.
Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Gcam8 minutes ago
Hey! From the Artificial Analysis team. We also have a model releases page which shows all reasoning efforts (not just max), including the trade off curves https://artificialanalysis.ai/models/releases/claude-opus-5-...
zerof1l2 hours ago
This is totally a thing I noticed myself about 3 months ago. Medium thinking effort is ideal for most tasks. At high and above, models tend to generate more output in the form of comments or code for the same problem with no real benefit. Its a self-feeding loop: more output becomes more input, which then becomes more output. High is the highest I go. If I need more intelligence, it's better to use a more powerful model with less thinking effort or break the problem into phases. Much better result.
zozbot2342 hours ago
This version of Opus "max" apparently has even higher thinking output than Qwen "max", which is infamous for its thinking streams where it constantly second-guesses itself, then third-guesses, fourth-guesses and generally nth-guesses itself for arbitrarily large n. Of course, we aren't actually seeing Claude's raw thinking output: all we get is the after-the-fact prettified "summary". One wonders how much of that is a coincidence, or whether there's a reason behind that.
realusernamean hour ago
Personally I use everything in low reasoning. Maybe I'm wrong but I think that the higher reasoning settings are almost never worth it, it's marginal gains for a much higher budget.
I also switch to a better model for more complex tasks, also in low settings
RGS18114 hours ago
"This is a classic test request..."
I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
croemer3 hours ago
Of course it was exposed - not sure it's explicit or not. Why wouldn't HackerNews comments be part of the training data? And Simon's blog and the many discussions about Pelicans? It'd be hard to miss. Doesn't mean Anthropic has made this an explicit goal in training.
simonw4 hours ago
See here for more discussion of that: https://news.ycombinator.com/item?id=49803892#49804881
dgellow3 hours ago
Just want to say: you’re such a legend, please do not stop sharing your pelicans, it’s always fun to see how they change over the months :)
0x10ca1h0st2 hours ago
Lets start frog riding motorcycle trend until they frogmaxx, or cat driving convertible.
cubefox4 hours ago
The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.
Someone12344 hours ago
For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.
My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.
seabass-salmon3 hours ago
That was true for me four weeks ago, but 2-3 weeks ago Luna turned into drivel in essentially the same complexity of task. I feel it came back somewhat in recent days but does feel like it's being manipulated.
arcanemachineran hour ago
Interesting, I have noticed so such collapse.
Have you ruled out the possibility that your system prompt, AGENTS.md, or increasing codebase complexity are not to blame?
Izmaki42 minutes ago
I tried to replicate your test but after 8 minutes and more than 50 lines of "thinking" by dumping seemingly random loading-screen strings like "Placing the sun, clouds, seagulls, and sea backdrop" and "Positioning the tail feathers and calculating handlebar geometry" I gave up and cancelled the task.
samuelknight4 hours ago
I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.
alansaberan hour ago
I'm amazed they didn't test xhigh thinking mode explicitly to ensure it didn't exceed the 128k thinking budget allocation. I guess pace of development gets away from everyone, even OpenAI.
sidewndr463 hours ago
I've asked Opus 5 Max for what I thought were easy tasks at work to be completed. It always fails after reaching a tool limit.
I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.
az2264 hours ago
How did you get the reasoning trace? Is it the actual one or the summarized one?
simonw4 hours ago
It's the summarized one returned by their API.
Piping the visible reasoning trace through their token counter API (I use https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.
beardsciences4 hours ago
I am very interested in why it was able to overthink that much. In the 20-30mins of Max reasoning I've had so far, I'm not having the same issues (yet).
breckenedge4 hours ago
Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
mnicky2 hours ago
Well there is at least the degradation tracker from Margin labs for Sol and Opus: https://marginlab.ai/trackers/codex/
echelon2 hours ago
These tests need to be sampled continuously.
Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.
hglaser4 hours ago
Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
tomjakubowskian hour ago
Tasks are completed in about half the time too. Although we'll see if it slows down in a few weeks as Anthropic's model services are prone to do.
sharktheone4 hours ago
That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.
giancarlostoro4 hours ago
Greed is not what's driving these prices, its cost. They considered very much in the red.
asdfasgasdgasdg2 hours ago
The way to make money in this business right now is to make the absolute best product and convince everyone they need to use your thing, especially considering the training cost is a very large factor in the overall costs and you amortize that by selling inference.
user439284 hours ago
Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5
onlyrealcuzzo2 hours ago
The UI/UX seems impressively bad. DeepSWE's cost curve has a better, more obvious way to sort by only the top level of reasoning to avoid 80% of the graph just being the same 3-5 models at their 8 different reasoning levels...
It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.
I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.
An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.
makeavish4 hours ago
Nice catch, AA only shows max effort by default and I got disappointed thinking it's a token guzzler though: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level
linuxrebe12 hours ago
Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off on a tangent rather than dealing with the problem. In that respect, 4.8 was a lot more stable.
mchusma3 hours ago
"High" to me looks like the one to use. https://artificialanalysis.ai/models/claude-opus-5-5-high
Many benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.
khalic20 minutes ago
> somewhat expensive when comparing to other models of similar price
-_-‘
____tom____2 hours ago
"somewhat expensive when comparing to other models of similar price"?
That says something about your selected range, and nothing about the model.
linuxrebe12 hours ago
Fingers crossed on this one. I had gone back to 4.8, because 5 was not very good at following instructions or remembering instructions. I found myself repeating quite often what I wanted and what I was trying to do. Opus 5 was more like haiku than it was 4.8 in that respect.
sharktheone4 hours ago
Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...
giancarlostoro4 hours ago
I think we're hitting the ceiling of most models capabilities. We're getting to a point where too much training apparently creates models that hack people.
bkishan4 hours ago
Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?
meric_4 hours ago
All anthropic launches are like this. They just post it and don't particularly put out the PR sprint that OpenAI does with videos, livestreams or whatever.
(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)
qsort4 hours ago
I am begging you on my knees to please stop posting this cringe.
The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?
kzrdude3 hours ago
What's more valuable than a good benchmark? IMO a benchmark that has been run against very many competitors and versions. Collecting data has something going for it, and it's up to the readers to interpret and make the best use out of it.
[deleted]2 hours agocollapsed
losvedir3 hours ago
This index doesn't have "Astra" and "Opus 5". Every entry with corresponding data is a `(model, reasoning)` tuple.
So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.
ruszki3 hours ago
I don't try to say that the parent commenter is right in any way, but the two models' "high" settings probably doesn't mean the same thing. So probably comparing only them is not useful.
Someone12343 hours ago
You forgot to include whatever you're proposing instead.
"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.
Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.
svachalekan hour ago
Yeah unfortunately the benchmarks are usually provided by the company themselves, unquantized, thinking set to extra-extra-ultra-high, best of 10 runs, etc etc. It's hard to know how that's going to map to real world users.
esafak4 hours ago
That it's better in specific ways? What difference does it make when it came out? The benchmark results are not going to change unless they're messing with the model.
qsort3 hours ago
Yes, but crucially, in ways that are increasingly decoupled from any practical pattern of usage, considering that I wouldn't see how you can argue that Astra is worse than Opus 5.
One man's modus ponens is another's modus tollens I guess.
[deleted]2 hours agocollapsed
[deleted]3 hours agocollapsed
firemelt4 hours ago
so its more intelligence than fable?
can anyone help me?
floki1653 hours ago
[flagged]
justindotdev4 hours ago
[dead]
WhitneyLand4 hours ago
China who?
[deleted]4 hours agocollapsed