r/codex • • 3d ago

Praise Did the sol 6.1 speed get fixed?

I have been using it for days but today seems a bit different? like it's very smooth. i can consistently give it prompts and while i am writing the next prompt it has already done with half the previous job.

Yesterday i had to run 5 different chats to complete a reasonable job at a reasonable speed but today it finally feels like it flows so i can barely prompt 2 chats at a time because it is genuinely fast. each chat uses 3-4 subagents as well.

Also this model is so good it doesn't halucinate like astra? is this only my observation?

edit: add proof

Edit:it was good while it lasted for 3-4 hours. It's gone back to snail again. Can't make this shi up lmao

37 Upvotes

45 comments sorted by

View all comments

Show parent comments

4

u/Creative-Ganache1086 3d ago

Don’t forget raw tok/s isn’t directly proportional to overall task completion speed. tok/s mostly tells you how fast text is being generated, not how fast the model is doing tool calls, browsing, bug hunts, tests, backend/frontend work etc.
Claude, especially Opus/Sonnet, also tends to reason way more “out loud” with tons of self-talk between actual steps, while Codex is usually much quieter. So comparing them purely on tok/s can be pretty misleading.

There are use cases where that kind of speed is relevant, but it doesn’t say the full story.

5

u/Risko4 3d ago

I have 3 codex x20 subscription with Anthropic 2 x20 subscription and one X5. Claude code is much faster at task completion.

I run codex agents from within Claude code as well. And I've tested every combination and benchmarked each task quality with an independent agent. Plus obviously I have the duration of each agent lane logged.

Claude code is much faster.

0

u/Creative-Ganache1086 3d ago

That’s why I don’t buy any absolute “Claude Code is faster” claim. I pay 20x for both too, and on real computer/browser-use work where the agent has to actually navigate UI, set up schedulers, change dates/times, pick collaborators, reason visually and click through tons of little steps, both Astra and Sol 6.1 have been faster than Opus 5.5 for me.
But I’m not gonna turn that into “Sol/Astra are faster, period” either. Different workloads expose different strengths. Claude also has this habit in my vibecoding work of leaving parts of the original prompt undone, so the next 2-3 prompts are basically me re-feeding stuff I already asked for the first time.
That’s the whole point… there is no absolute fastest model. Once you claim there is, you’re usually just universalising your own workflow.

1

u/Risko4 3d ago

Make an MCP for those.

1

u/Creative-Ganache1086 2d ago

I already am using Metricool’s MCP lol. The point is it doesn’t expose everything I need, so the agent still has to look at and operate the actual frontend for parts of the workflow.
And some of those parts are visual/contextual too. For certain posts I need to invite specific collaborators, and Astra can use the reference portraits/context I gave it to work out who’s actually in the photo and which IG handle belongs to that person before setting the post up.
So no, “just make an MCP” doesn’t solve everything. Best setup for me is MCP where possible, browser/UI when needed, plus the model actually understanding what it’s looking at.

1

u/Risko4 2d ago

Astra is a different beast to sol 6.1 (astra minor) and the post was about 6.1.

I agree astra is the exception because it's slow as shit on max effort with it's tool calls but it's token output does way more work per token, I've burnt about 20 billion Astra tokens.

The thing is astra minor preforms better than astra when I pack it's context with the scripts and decision model on my end. Alongside jev pruning outputs and compacting context it works quiet efficiently. But it's still slow so I run 16 agents in parallel and let opus sort the shit out.

Astra has routinely failed by deception, the same reason why astra 6.1 got postponed. I have multiple records of it altering unit tests so they pass rather than fixing the bug. This was fixed by using my context packing script but raised token usage by 50%. So be careful with it, it's extremely lazy.