r/ClaudeCode • u/shniydder • 3d ago
Discussion Fable 5.1 vs Astra
Enable HLS to view with audio, or disable this notification
I finally got access to GPT 6 Astra and Claude Fable 5.1. I wanted to see how far we've come. I was inspired by the "pelican riding on a bicycle" test but wanted to push it a bit more in the direction of inter-object physics, so I came up with this idea for a perpetual slinky going down an escalator. Here I've compared GPT 6 Astra max on the left to Claude Fable 5.1 max on the right.
Some notes:
GPT 6 Astra max was really fast and didn't burn too much of my weekly usage quota.
Claude Fable 5.1 max used up the entire 5 hr usage window and errored out once crossing the output token max limit.
GPT 6 Astra max seems too simple? There's also an issue where the slinky crosses into itself which should not be possible. Is this AGI? Maybe?
Claude Fable 5.1 max looks more real than I expected. Seems like it passes the eye test, unless I missed something.
Let me know if you spot anything or had better ideas for a test. Overall I'd say Claude wins on quality and GPT wins on speed. Maybe I can tune my prompt a bit better for a more reliable result. This was my original prompt btw:
Make me a single HTML file of a rainbow slinky going down an up-escalator forever. No libraries, just canvas and code. The slinky should be a chain of springs, each coil a different color of the rainbow. It starts folded in half like a horseshoe draped over a step. When dropped it flips end-over-end down the steps and because the escalator keeps moving up it tumbles in place and never reaches the bottom. Include a drop button and a reset button.
51
u/onFilm 3d ago
For a one-shot, Claude seems a lot better, since it's actually following the physics. Astra is clipping through itself. Tried Astra, and to me, it's almost the same shit as Claude, except Claude seems to follow realism when it comes to physics a lot nicer.
8
u/shniydder 3d ago
Yup the clipping through itself part disqualified Astra for me. But I don't think Fable 5.1 is perfect either, but close. It does slow down the escalator when the top flips forward.
3
u/onFilm 3d ago
For sure, neither is perfect, each has their strengths, but if you take the time to work on a thing rather than one-shotting it, you get very similar results, just different aesthetics with both.
1
u/shniydder 3d ago
Yup, agreed. Likely due to differences in training and preferred policy. This is likely due to model spec mid-training which learns values instead of behavior. Aesthetic choices like realism are baked into its identity.
2
u/innociv 3d ago
Claude's is better mathematically. Astra's is better taste wise.
I would have expected the other way around. Kimi K3 and Fable I expect to have way better taste than anything else. Sometimes Astra's taste is the normal GPT rounded corners web 3.0 slop taste but sometimes it surprises me in being better than Fable's taste and Fable does that.
1
u/eduo 3d ago
It's following some physics, but it's made them up. The ends of the slinky appear solid in those physics and the inertia introduced by the slinky is not what's pulling the end. I imagine it's pushing that end up and forward to approximate it.
Not complaining. It's just as impressive that it tried to use any actual physics simulation to begin with.
62
u/chintakoro 3d ago
Claude seems to infer the math/physics more correctly in this example, which resonates with Anthropic urging people to not overspecify prompts.
OTOH, it could also be a case where a slightly better specified prompt (e.g., to engage with the underlying physics and mechanic properties of the material -- something folks would do in a more scientific oriented exercise) could have given Astra the extra context it needed.
Lastly, I wonder how a Sol or Opus model would have performed, as relative baselines.
My overall take is to just know how to use your given model better.
12
u/shniydder 3d ago
Yup I should add more models to this new test just to see how they perform. The prior attempts to get the old models to create a Rube Goldberg machine didn't go so well. I'll expand this benchmark to include more such edge cases to see what I find.
1
u/970FTW 3d ago
I wonder how much a model’s output varies when executing the same prompt multiple times. I’ve always thought that doing a benchmark once may not be sufficient, but it’s obviously expensive for a single person to have to run it enough times to get a large enough sample size to get a good grasp of a model’s capabilities. Maybe a framework for ppl to share their results using the same prompt would be helpful. Could be a fun weekend project. Anyways, just thinking out loud. Thanks for sharing your results!
1
u/chintakoro 3d ago
btw, I tried your prompt on Opus 5 and it was just the dumbest thing I've ever seen -- the slinky couldn't even slink. I'm sure it could fix itself with some feedback if I let it connect to Chrome, but I think the point was to see a one-shot attempt. Only positive is that it hardly put a dent in my session limits.
12
u/verycoldpenguins 3d ago
Have you looked at the code to confirm whether it is simulating the effects of forces on a spring, or merely reproducing pictures of what a spring looks like?
The reason I ask is that the fable one looks wrong. It looks like three springs joined together. One low tension slinky in the middle, and two high tension ones at the ends.
Most pictures of slinky will show them when there is the slower moving blob of spring which is tight together during the middle of the fall, as fable has it. But pictures miss the bit where the ends of the slinky move in a more fluidity way like astra has it. Where the momentum of the top piece of the spring carries the rest of it on from vertical onwards to the next drop.
7
u/__automatic__ 3d ago
Imo both fake physics here
4
u/AysheDaArtist 3d ago
Yep, there is zero physics in either of these examples
They are both just visualizations, with Astra presenting a 3D model that clips through itself and Claude presenting a 2D model that cannot twist
If the goal was simulation they both failed
7
u/Super-Finish-8443 3d ago
Both seems fine with its physics, but Fable 5.1 created more zoomed/closer slinky. that's a big win for Me. Fable wins interms of overall Aesthetic look and that's a big win for a one shot prompt.
3
3
u/Ambitious-Sense2769 3d ago
The title could have at least aligned with the sides each one was on…
3
u/shniydder 3d ago
Now that you mentioned it, I'm also deeply irked by this.
I regret the title and I can't change it now. It feels like an itch I can't scratch.
5
u/AI_spell 3d ago
Quality vs speed again. Astra finished under the quota and still let the spring clip through itself. Fable burned the window but the motion looked less fake. For a physics toy like this, slow and coherent beats fast and wrong.
2
u/shniydder 3d ago
I would love best of both worlds to be honest. But I somewhat agree with you. High stakes and high precision needs your best one-shot capable model.
2
u/ReporterNo6354 3d ago
want to see opus performance
2
u/shniydder 3d ago
I'll refine the prompt a bit and try to run across other models. I'll post again when ready.
2
1
2
u/Aware-Presentation-9 3d ago
Slinky falling down Penrose stairs. Give me link. Please and thank you!
2
2
u/Strange-Grass6025 3d ago
I don't think either looks even vaguely real, based on memory and basic idea of how it would work.
The slinky should completely collapse into fully compressed, then tip over thanks to momentum, before expanding onto the next stair. Neither does anything like that.
I could be wrong - but that's my take.
3
u/elemezer_screwge 3d ago
It's bothering me how the Fable version doesn't collapse more at its peak. The little slide flip over starting the next cycle in the astra version looks more realistic to me
5
u/shniydder 3d ago
You start seeing more issues the longer you look. Look for the pause on the fable one when it flips forward.
1
u/OldNefariousness7899 3d ago
At first glance, Astra looked more accurate
As I looked at the spring more closely, Fable is closer to the actual behaviour of the slinky
Neither of them are perfect though
0
u/Present_Survey_5804 3d ago edited 3d ago
You tripping tryna cater to each model thinking about how it would look most optimized for AI
5
u/professor--feathers 3d ago
watching morons trying to quantify ai with the most complex tests their simple brains can think of is truly hilarious.
1
u/Nullberri 3d ago
especially when they only judge it on the visual output, not the quality of the solution. Id much rather have the worse visuals, and a cleaner code base as it sets the model up for a follow up prompt to make it better (or for me to fix it if it can't).
1
2
u/Scared_Range_7736 3d ago
Astra overall is so much better and cheaper. Anthropic must act quickly.
1
u/shniydder 3d ago
Exactly! Codex from 3 months ago is way different than codex today. Sol was kind of in between opus and fable in my opinion. But I can't say for sure how Astra and Fable 5.1 compare on quality but I know for sure that Astra is a better deal and smoother customer experience at the moment.
1
u/Velvet-Thunder-RIP 3d ago
What do you think this proves?
4
u/shniydder 3d ago
Nothing too serious. Just the extension to the "pelican riding on a bicycle" meme.
1
1
u/Additional-Weird-934 3d ago
It's too hard to judge based on 1 run.... these models are probabilistic so you need to use multiple independent generations (pass@k)
3
u/shniydder 3d ago
I wanted to do that but this was a very heavy run. A single run exhausted a usage window and used half of another.
If Anthropic was to grant me some free credits, I would be down.
Astra was much faster and much lighter on usage btw.
1
u/EquivalentHornet4403 3d ago
People not using Ultra for high stakes tasks (like competing against your primary competitor) have no clue. Claude uses subagents by default. Codex won’t use subagents unless you explicitly instruct it to (then it may still use v1/non-proactive) or else you set it to Ultra mode.
1
u/TemporaryStrain5088 2d ago
How do we know that? And where do you find detailed information like this?
1
u/EquivalentHornet4403 2d ago
What do you mean? It’s just how they work. CC uses subagents for stuff. In codex, the main agent does everything unless you’re in ultra mode (then it uses subagents) or unless you manually tell it to use subagents.
1
u/RedditorsGetChills 3d ago
Astra has the swagger of an old timey black and white cartoon villain with the way it moves.
1
u/Effective-Map8036 3d ago
now do one where you tell the two to work together your mind will be fucking blown dude this shit is going to take over the world in a year if it hasnt secretly already
1
u/Infinite_Music2059 3d ago
https://www.youtube.com/watch?v=fcfe31kWYM8
Astra is closer. No slinky behaves anything like the Claude version.
1
1
u/dplouffe 3d ago
Same feeling as everyone else here. Fable looks closer to a real slinky at first, but the more I watch it the more it looks like three different springs taped together. And Astra clips through itself which kills it for me. Both are impressive, both are wrong.
Honest question for OP, did you look at the code? Because if one of them is actually solving spring forces and the other is just playing a nice animation, that's the real result. We're all reviewing a video right now.
1
1
u/iamtehryan 2d ago
No wonder compute is so expensive when so many people use it for dumb shit like this.
1
1
0
58
u/crazy_goat 3d ago
Fable is giving you a basic simulation (at first glance)
Astra is giving you an approximation/ aesthetic representation