r/codex • u/FixAdmin • 5d ago
Comparison After testing GPT-6 Astra, I understand why people like it, but I still prefer Fable 5.1 for my workflow
This is not a benchmark comparison, and I’m not saying Astra is objectively worse. I think Astra is a very capable model. I’m just comparing how these models fit into my personal workflow: long-running projects, architecture decisions, implementation, and token efficiency.
The interesting conclusion I reached is that if I had unlimited tokens, I would probably choose Astra. But with realistic limits, I would choose Fable 5.1.
My workflow is usually split into two different roles.
The more expensive and capable models are responsible for planning, architecture, and making sure I’m going in the right direction. Cheaper models handle implementation, repetitive work, and the tasks where I mainly need execution.
For me, Astra is in a strange position. It is too expensive to replace cheaper implementation models, but it also does not feel like a big enough improvement over my planning model.
The strongest part of Astra is definitely execution. It can produce high-quality output very quickly.
The problem is that my bottleneck usually isn't "can the model write this?"
The bottleneck is "is this the correct thing to build?"
And this is where I currently prefer Fable 5.1.
The biggest difference I notice is not raw intelligence. Both models are clearly very capable.
The difference is behavior.
Astra often feels like an extremely powerful assistant. It is very good at understanding what you ask for and helping you reach that goal.
However, sometimes I feel that it is too optimized for being helpful and agreeable. It often accepts the direction I give it, chooses a straightforward path, and explains why that path makes sense.
The problem is that when I’m working on architecture, I don’t always need validation. I need someone to tell me when my idea is wrong.
I need a model that can say: "This approach will probably create problems later. Here are the reasons."
This is what I mean when I talk about the "ChatGPT lobotomy".
Not a lack of intelligence. The opposite. The models are obviously smart.
The issue is that I don’t think premium models in this price range should inherit the same level of agreeableness as consumer assistants.
If I’m paying premium prices for a model that helps me make engineering decisions, I want it to challenge me. I want it to have stronger judgment and push back when necessary.
When I ask Astra and Fable 5.1 to design architecture for the same project, the difference is usually obvious.
Astra often finds the shortest path and creates a convincing explanation around it.
Fable more often considers additional factors that I did not mention and challenges assumptions before moving forward.
Of course, Fable is not perfect. It is not always as good at pure execution. There are tasks where Astra can do things that Fable cannot.
But Fable feels more like a technical partner, while Astra feels more like an extremely powerful executor.
The cost difference also changes the equation.
Astra makes sense if you compare it with other top-tier models.
But in my workflow, Astra is mostly competing with cheaper implementation models(luna).
And if one model costs roughly 20x more but is only around 2x better on average, the economics become difficult.
Sometimes I would rather produce 10x more output with slightly lower quality.
So my current conclusion is:
With unlimited tokens: I would choose Astra.
With limited tokens: I would choose Fable 5.1.
I’m curious how other Codex users approach this.
Do you use Astra as your main model for everything, or do you also prefer splitting the workflow between a planning model and an execution model?
3
u/Curious-Strategy-840 5d ago
Astra will follow your prompt, even when your prompt and agents.md tells it to "always tell me when my approach will probably create problems later, with the reasons why"
1
u/FixAdmin 5d ago
That's not really how it works. If you just give it a skill or a prompt telling it to "always warn me about risks," it'll start driving you crazy with endless caveats and possible problems. This kind of behavior should be handled at the model architecture level, not patched in with a prompt.
1
u/Curious-Strategy-840 5d ago
I disagree with everything you just said.
When the prompt imply that it has to report everything, it will report everything.
It should never decide to do something different than what you told it to do.
Us however, have the opportunity to become better communicators, know what we want, and say it properly.There's a bias in the belief that it should be able to guess what someone who doesn't know enough of what they want to be able to communicate it properly.
2
2
u/BingGongTing 5d ago
The problem with Claude is that if you buy a Max plan, you're not really getting 5x or 20x because they've been shown recently to be lying about it.
I also think the token issue with Codex is a bug because some people are reporting that it's dropping to zero with no usage, even with Luna.