simonw6 hours ago
I posted this in the other Astra thread but it's just fallen off the homepage, so...
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
threatripperan hour ago
It seems to start understanding how bicycles work. Look at the front fork. There are reasons why it's shaped the way it is in real bicycles. If you do physical simulations in 3D with reinforcement learning you will start understanding that most of the arrangements that the other agents use just won't work.
epihelix7 minutes ago
Yes, the curve of the front fork on the max pelican was what I noticed also. It's impressive (if not an accident), and a remarkably accurate bike overall. The pelican just needs to raise their seat a little.
steve-atx-76003 hours ago
You would not expect the developers of the model to optimize for a well known benchmark?
simonw24 minutes ago
Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
CamperBob24 minutes ago
If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.
y1n02 hours ago
I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
kibae2 hours ago
There was a HN post that tested this hypothesis a few weeks ago: https://news.ycombinator.com/item?id=49010129
Kranar2 hours ago
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
mudkipdev2 hours ago
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
awakeasleep2 hours ago
You gotta look up how RLHF works before you ask a demanding question like this.
mi_lk2 hours ago
It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it
Treat it like a bit as is
benatkinan hour ago
It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.
jsdalton4 hours ago
Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
simonw4 hours ago
For the moment the best place to see them all is to browse the tag on my blog - 140 posts now! https://simonwillison.net/tags/pelican-riding-a-bicycle/
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
bahmboo37 minutes ago
Would be great to spend some money (other people's money) on paying some illustrators and artists to do the same for comparison. It could be straight images converted to SVG or direct vector graphics.
vb-84484 hours ago
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
thimabi4 hours ago
This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.
pizza2344 hours ago
Astra is the only model that correctly depicts occlusion of crank and leg, although interestingly, at max and medium levels (not in between).
vessenes3 hours ago
Looked to me like it missed the chain though.
CamperBob27 minutes ago
Crazy how much the Astra pelicans resemble GLM 5.3's (https://crimson-jeri-74.tiiny.site/), down to the color of the bike and the scarf.
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
petilon3 hours ago
Astra pelican looks amazing: it looks like it was made by a professional artist. All the others look like they were made by kindergartners.
BrokenCogs5 hours ago
Interesting that all of the bikes are turquoise colored, except for the medium effort
andai4 hours ago
Oh my goodness. I was not prepared for Luna on "none".
Reminded me of https://clocks.brianmoore.com/
tyre4 hours ago
Haiku the GOAT
[deleted]5 hours agocollapsed
[deleted]5 hours agocollapsed
Xunjin5 hours ago
Do you have other ideas of "combinations" for this kind of benchmark?
I'm wondering if this is being trained on by the models today.
r_lee4 hours ago
I have a feeling they've been doing that for a while now, even if unintentional, as it's such a well known benchmark
kart233 hours ago
radial spokes on the back wheel is not a thing irl. i wonder if any model has gotten that right?
bnorton3 hours ago
You’re closer to this than I am but do you find it odd that helmets are almost never included? Biking is almost always accompanied by helmets
maxlapdev3 hours ago
But pelicans are almost never accompanied by helmets, so it cancels out.
sumedh3 hours ago
> Biking is almost always accompanied by helmets
Isnt Netherlands the leader in bike riders and they dont wear helmets.
neutronicus3 hours ago
At least one reasoning trace I saw considered a helmet and discarded the idea because it was worried about obscuring some detail
PacificSpecifican hour ago
Fallen off the homepage so what? I genuinely don't understand
fHr4 hours ago
Luna is my go to daily model it's great value
tuo-lei3 hours ago
Me toooo, for all my personal projects I need to pay for the tokens~ At workplace I use sol because I don't need to pay
ghthor3 hours ago
Mine as well, it’s fast and keeps me in flow; and cheap!
lspears5 hours ago
Luna's price seems off
droidjj3 hours ago
The price is probably coming from when Luna was first released. OpenAI slashed the price by 80% at the end of July.
simonw3 hours ago
Yes! Good catch, thanks - I'll fix that.
simonw3 hours ago
I had the prices wrong on Sol and Terra as well - they've all had price drops:
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
Luna — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 7.83 | 1.57 |
| xhigh | 4.24 | 0.85 |
| high | 2.46 | 0.49 |
| medium | 1.26 | 0.25 |
| low | 0.76 | 0.15 |
| none | 0.71 | 0.14 |
+--------+--------+---------+
Per million tokens:
Before: $1 input / $6 output
After: $0.20 input / $1.20 output
Sol — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 48.55 | 32.37 |
| xhigh | 24.11 | 16.08 |
| high | 10.38 | 6.92 |
| medium | 10.55 | 7.03 |
| low | 8.33 | 5.55 |
| none | 5.90 | 3.93 |
+--------+--------+---------+
Per million tokens:
Before: $5 input / $30 output
After: $4 input / $20 output
Terra — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 32.09 | 25.67 |
| xhigh | 14.67 | 11.74 |
| high | 3.74 | 2.99 |
| medium | 3.46 | 2.77 |
| low | 3.47 | 2.78 |
| none | 2.60 | 2.08 |
+--------+--------+---------+
Per million tokens:
Before: $2.50 input / $15 output
After: $2 input / $12 outputman42 hours ago
[dead]
jedjjfjf2 hours ago
[dead]
leoqa6 hours ago
[flagged]
samuelknight5 hours ago
How are we supposed to know if Astra is frontier without the pelican?
satvikpendem5 hours ago
It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.
ComplexSystems6 hours ago
It's absolutely related.
StopTheCringe5 hours ago
[flagged]
jjcman hour ago
It's ability to handle non-90 degree cutouts and shapes for web dev is one of the best I've seen. The vision model on this is VERY capable.
Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...
And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
XCSme6 hours ago
That's some crazy SVG generation:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
satvikpendem5 hours ago
That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.
dprkh2 hours ago
I thought all the designers have been using vector graphics for a long time now.
mceachen2 hours ago
Aldus Freehand 1.0 was released in 1988. Adobe Illustrator was first released in 1987.
kulahan5 hours ago
That photo looks pretty hilariously stupid, so this appears to be more of a first toe dip rather than some indication we can one-shot a previously difficult process.
appplication4 hours ago
Sure, it is a bit cartoonish but it’s relatively impressive. I do wonder what you would get if you asked for photorealism
redox9934 minutes ago
The problem is not photorealism. The SVG is outright dumb, its on the wrong side of the table, the table has fucked up geometry (its tilted) and many more minor flaws.
embedding-shape5 hours ago
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.
XCSme5 hours ago
Do you prefer the fable one?
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
kulahan5 hours ago
No, both of them look awful. I genuinely am unsure if this is a meme you're making that's going over my head...?
XCSme4 hours ago
Well, this was the state of SVG generation until a few months ago:
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
XCSme5 hours ago
Yeah, that's confusing, the score is for the entire benchmark, not for SVG generation only.
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
XCSme5 hours ago
I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!
readams4 hours ago
The Astra one looks pretty good except it's standing on the wrong side of the table
[deleted]5 hours agocollapsed
kingstnap7 hours ago
Its also available finally to Pro users! Just took 24 hours.
InsideOutSanta6 hours ago
They gave out bankable resets for every day people on pro plans didn't get Astra. Given that, I wish they'd waited a few more days before activating it on my account :-D
wincy6 hours ago
They haven’t activated Astra for me yet, I have two resets now. I’ve been using the opportunity to test out how good 5.6 Sol is at computer use asking it to generate stuff in Blender which has been… interesting
Edit: nevermind it JUST gave me a notification to use it!
paxys6 hours ago
That's a pretty genius internal incentive to move fast.
embedding-shape6 hours ago
> They gave out bankable resets for every day people on pro plans didn't get Astra.
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
sumedh3 hours ago
Just got access to it on Plus plan in Australia. 2 Banked resets as well.
vb-84484 hours ago
Played in codex app a couple of hours today: it feels much faster than SOL, even if the TPS is half of it.
algoth16 hours ago
Just got it. European plus user here. Only codex, no chatgpt
jaesonaras4 hours ago
Anyone had success using Astra as a Foundry model via Github Copilot? The error I get is that tooling is not available if reasoning has a value.
[deleted]6 hours agocollapsed
gavinray6 hours ago
I have GPT-6 access in Codex and OpenAI API now
I'm a Business plan user with Cyber verification enabled, FWIW.
embedding-shape6 hours ago
Same just got access literally this minute, Pro user here, no Cyber verification but have passed my ID over to them back in 2024 or something, maybe at the ChatGPT 3 API launch or something?
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
r_lee6 hours ago
is Azure for this actually ZDR?
starik366 hours ago
What is the actual utility of using this model on Azure? It's twice as expensive, according to the link.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?
hhh4 hours ago
It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.
claiir5 hours ago
They’re ZDR and the OAI ones aren’t
[deleted]4 hours agocollapsed
olalonde4 hours ago
WTHIT?
spdustin3 hours ago
ZDR = Zero Data Retention — they don't store your inputs/outputs.
itsjustkev5 hours ago
Compared to OpenAI flex? I'm pretty sure that is their batch processing endpoint, which is naturally cheaper.
cute_boi6 hours ago
I got it, but sadly no resets...
vatsachak6 hours ago
Damn I am so hyped
marsven_4224 hours ago
[dead]
StopTheCringe5 hours ago
[flagged]
OpenGayEye_4 hours ago
[flagged]
noob00536 hours ago
[dead]
taywrobel6 hours ago
You made an account just to post this?
6thbit5 hours ago
Invoking simonw for pelicans pretty please.
drivers995 hours ago
That's over here: https://news.ycombinator.com/item?id=49570643
[deleted]6 hours agocollapsed
Osama04567 hours ago
is it finally available on openrouter and what about fable 5.1
r_lee6 hours ago
that's literally what the link is for...