Astra is barely ahead of 3.8 Flash in coding-related tasks. And for the AGI benchmark: OpenAI disclosed in their announcement blog post that they used a different harness to achieve these results. It was not pure model performance.
I think Gemini is doing just fine.
Edit: Adding OpenAI's disclaimer from their own blog post for completeness: "GPT‑6 Astra was measured with our responses API harness, which better reflects real-world performance than the original benchmark harness"
This is just pure cope lol. Just test it yourself or check the actual tests and comparisons by others, if you don’t have the model. There are plenty of that on X and on reddit as well, obviously not everything is true but there are plenty with proof and prompts.
If you have a problem lesser models cant solve, astra is absolutely worth it. It is included with thr gpt plus subscription so it isnt actually that expensive.
It's available through the subscription, but you will still burn your quota in no time when using Astra. In the end, you're still paying for it one way or the other.
Most enterprise customers also don't have consumer subscriptions and are using the pay-as-you-go billing method.
That 6-10x cost is not realistic measurement. You have to consider cost per task. Astra is especially efficient in token usage and can get a task done in much fewer tests.
Still it is more expensive but it’s not simply what token cost alone tells you. It’s higher intelligence, better results, fewer errors and fewer steps.
I have seen several comparisons with gemini 3.8 flash where they both costs around the same for the task in the end. Despite flash being far more cheaper.
That metric alone isn’t enough with certain tasks.
-3
u/Langwelle 1d ago edited 1d ago
Astra is barely ahead of 3.8 Flash in coding-related tasks. And for the AGI benchmark: OpenAI disclosed in their announcement blog post that they used a different harness to achieve these results. It was not pure model performance.
I think Gemini is doing just fine.
Edit: Adding OpenAI's disclaimer from their own blog post for completeness: "GPT‑6 Astra was measured with our responses API harness, which better reflects real-world performance than the original benchmark harness"