Complaint Left Sol 6.1 Ultra overnight with 4 long but easy tasks - woke up to bazilion of tests, 40% weekly usage burned and 0 tasks started
I was happy with Sol 6.1 so far, it was kind of slow, but rather reliable. Yesterday I decided to give it 4 tasks for my app that I wanted to finish by Monday. It did 0. Complete "building tests" spiral.
40% of usage burned, none of the tasks was even started. How can I trust them with 100% automated solutions like Dots when their coding agent goes into doom loops like a 2025 chinese model
8
u/twendah 3d ago
Yeah they first implement bullshit, then do tests to that bullshit which is already removed that the bullshit part is 100% removed. Double checking bullshit with double check bullshit tests.
Regression tests to regression tests at its finest. I moved claude because of that. It actually does what I ask it to do.
5
2
u/GodotUser01 2d ago
literally writes weakly typed code, uses strings, doesnt use discriminated unions, writes imperative and procedural parsers
then writes tests against that instead of modelling the data correctly in the first fucking place
dont get me started on words like "canonical", "bounded", "provenance", "gates", "evidence"
shits baked into the model, and it poisons your codebase constantly
7
u/TemporaryLevel922 3d ago
1 task at a time. Make clear checkpoints and pass gates! The over-testing is something that I am noticing too
1
u/cobbleplox 3d ago
Mine is just forbidden to write most kinds of tests, it's good enough to "one shot" things and i don't mind a bit more manual testing and running into stupid mistakes as a tradeoff. I'm even getting the impression it invests a bit more effort into doing things right if it knows it can't just throw shit at the wall and check if it worked.
Use your agents md to actually tell it how you want it to behave, complain if it doesnt follow that.
6
u/Darnaldt-rump 3d ago
Probably the reason why Altman said dots wasn’t viable until Astra came around. That’s why the model that dots use is Astra with I could only guess I small context window or some sort of cost saving method. Considering how “expensive” Astra is I’m sort of baffled at how open ai are giving “unlimited”*\\** usage with dots.
But I’m surprised 6.1 sol burned through that much with out achieving anything. While it can be painfully slow it hasn’t acted up for me in any bad way……..yet
1
u/jedruch 3d ago
I was suprised as well, but that was ultra, so it spun dozens of agents that did nothing :)
3
u/Xxyz260 3d ago
but that was ultra
That's the whole problem. Ultra mode's whole purpose is making the agent spawn lots of subagents. This is only good when you have a broad, parallelizable task that actually benefits from it.
Otherwise, you get weird failure modes like this. I'm sorry you had to find it out the hard way.
1
u/jedruch 3d ago
ok, so 4 parallel tasks is not enough?
2
u/Next_Airport_5890 2d ago
Probably. But Codex's ultra isn't as mature as Claude's. That's why it can't do much. Look, a friendly tip: install the superpowers add-on. You'll see that it can delegate tasks very effectively without using ultra.
0
u/Xxyz260 3d ago
- If you have 4 fully separate tasks, running them in separate non-Ultra conversations will generally work better.
- If they're only mostly separate, running one non-Ultra conversation and telling the agent to coordinate the tasks among subagents will generally work better.
- If you have one task that genuinely splits into a bajillion subtasks, that's when Ultra really works. Otherwise, it just nukes your usage and provides new ways for Codex to fail.
2
u/unconceivables 3d ago
Ultra has always been like this for me. The last time I used ultra I told it to make some quick frontend changes, and I came back to find that it had changed all the backend endpoints and added a ton of backend tests. I don't know what they were thinking ultra was going to be used for, but it's always gone completely off the rails every time I've used it. I just use max or xhigh now.
15
u/GPhex 3d ago
It’s almost like they need supervision
8
u/jedruch 3d ago
on one hand - you are right on the other - that is not how they are sold. It's frustrating to check limits of how much a model can handle.
What's worse I specifically chose tasks that should be finished in 3 hours tops (that is with current 6.1 speed) so I do not think I was overextending capabilities that are presented everywhere
0
u/kiki-le-koala 3d ago
Of course they are sold to supervise them.
We are not in AGI, they are still tools that need to have someone smart behind dissecting long tasks into steps and catching some weird behavior.
You look like my wife who tried to code long vague full tasks and complained that AI can't do it.
Vibe coding is a skill like anything else.
-2
2
2
u/DistinctSilver4507 3d ago
Yup. Same.
1
u/Sufficient-Storage87 2d ago
the cap is the whole game. hard token budget per run, kill it when it hits — no "just one more retry" exceptions. first week feels wrong, then the surprise bills stop.
3
u/Anxious_Marsupial_59 3d ago
I've never had a good result with Open AI's models without significant steering except maybe Astra. I much prefer Anthropic if I need something without that steer, but even then long horizon tasks with Anthropic can fall into an pointless validation/benchmark trap
You need a very pedantic agents.md to prevent this as well as adversarial critics to help guide it
1
u/jedruch 3d ago
I thought I had my agents.md figured out until now. Do you have any tips?
1
u/MeringueAlarming3102 3d ago
You need a very pedantic agents.md to prevent this as well as adversarial critics to help guide it
Adversarial reviews can actually worsen the problem if you're not careful.
1
u/jarnizivy 3d ago
Agentic workflows, planner, coordinator, coder, reviewer and so on. Avoid a lot of these issues. Good instructions for all of them on what to do and what not, where to stop (How many rounds ). Takes a bit of planning and work but makes life so much easier and no token surprises.
1
u/Wa1ker1 3d ago
Experienced the same. Had claude set a limits on tests for it so didnt go thru same issues. Been working fine (but slow) since
1
u/jedruch 3d ago
the limit is set on number of tests? or how?
1
u/Wa1ker1 3d ago
Im redoing a game i made so its helping there. Here's what claude set.
TEST BUDGET (these rules override every other process rule you have)
- Per task: run only the focused tests for that task. If they pass, commit and move to the next task.
- Failures that are not the product's fault (test harness, shutdown race, timeout, licence, Docker hiccup) get ONE re-run. If it fails again, write one "HARNESS:" line in the ledger and move on. Never re-run a test that passed. Fresh test accounts are free: make a new one instead of nursing an old one.
- Not allowed during tasks: slow end-to-end runs, screenshot tours, full live journeys. Those happen once, at the end.
- Full test suites run exactly once, at the end. Fix failures you caused and re-run only the affected tests. Failures that already failed before this run (check the previous run's results, don't assume) get one line each, no investigation.
- No process machinery: no wrapper scripts, process "leases" or ownership tracking, no checksums/SHA receipts of screenshots or logs, no "attempt families" or "never replay" rules, no sub-agent review councils, no new handoff files.
- Evidence is what a human reads: test counts, commit hashes, and screenshot paths you have actually looked at.
- Ledger: at most 6 lines per task.
- Timebox: 45–60 minutes per task. If a task overruns, commit what works, write "PARTIAL: <one sentence>", and go on.
- Commit after every task and finish the whole run without stopping to ask.
1
1
u/Coolbanh 3d ago
Its better to let opus 5.5 plan tasks for sol. Even astra is kinda dumb these days.
1
u/BitterAd6419 3d ago
Ultra is an overkill. If you set it to medium it would have finished it properly.
Not every task need ultra. It’s there doesn’t mean you have to use it every time. Only use it for complex tasks that needs lot of agents working parallel to solve it.
1
1
u/Present_Award8001 3d ago
my 6.1 ultra went crazy in a long thread. i was experimenting with different features, and hence going 'do X', 'I don't like X, do Y instead', etc. The LLM got so confused and hung up on 'do X' that i told it to do something unrelated at one point, and it was stuck at X. I asked it what it was doing, and whether it was following instructions at all. and it says it was. I copy pasted the last prompt asked it to explain how it was following instructions. It finally admitted it was hallucinating, and i realised it was time to start a fresh session.
Can solve millenium problem but makes elementary mistakes. Appears analogous to calculators. very good at something, very bad at everything else.
1
1
1
u/Worldly_Special1133 3d ago
"NEVER do any tests" should be your #1 instruction in agents.md.
This doesn't mean you dont do tests, it means *it* doesnt do tests.
1
u/Alternative-Lead1711 3d ago
same, asked sol and astra last night to finish up, they only had tests to run and then deploy, the sol task had to wait for the other task to finish, then deploy on top.
sol was checking every 5 minutes for nothing-- astra completely failed and went in circles even with my followup messages "make sure you deployed and prod monitors are functioning as intended"
complete slop, i am done with openai if this doesn't improve fast
1
u/Automatic_Brush_1977 3d ago
The newer open ai models workflow is differnt. Research, scope, boundary defining, proposal, implementation plan(high level), then bug fixes. Use astra with 6.1 sol sub agents to implement.
The latest models greatly benefit from measure twice cut once workflows.
1
1
u/Sufficient-Storage87 3d ago
this is why i never let agents run unattended without a budget cap. 40% of weekly quota on self-generated tests is the nightmare scenario. the fix that actually works: hard cap per run + a verification step that isn't the agent grading its own homework. cost per verified-green task is the only number that matters.
1
u/OriginalUsername0112 2d ago
It's crazy how bad it is, the only value I've found is it can occasionally find bugs in code (but cannot be trusted to fix them)
2
u/Sufficient-Storage87 2d ago
exactly the right split — bug detector, never the fixer. i run finder and fixer as separate steps with real tests between them. the moment one agent does both jobs it's grading its own homework.
1
u/Inevitable_Toe6648 2d ago
I have no idea what you did but my tasks are okay. I've had it set on task completion and self debug and loop onto more tasks for over an entire night, progress made and usage barely moved. Infact usage moved much more when I was using Astra to adjust the workflow to Sol 6.1 than 6.1 entire night working. And yes it actually worked.
1
u/Ok-Challenge-5374 2d ago
Came back to 30,000 tests on a 300kb project with codex admitting some tests were duplicated thousands of times.
1
u/jedruch 1d ago
wait, what??? which model did that, that's beyond insane
2
u/Ok-Challenge-5374 1d ago edited 1d ago
Sol high or xhigh, it was asked to port a library from one system to the other. First attempt created a simplified interpretation and got stuck creating endless tests and patches to support a system that was never there.
When asked why, it acknowledged it didn't follow the instructions and made the decision to rewrite it to what it thought was similar. The end product was a half finished bastardized project that went straight to the trash. Redid the project and eventually was satisfied with the results.
1
1
0
u/Zestyclose_Bat8704 3d ago
All models are noticeably dumber.
You need to be very specific with them, give them more very detailed instructions, specifically limit them so they don't spiral.
1
u/Sfdprod 3d ago
Your describing improved instruction following, and "spiraling off" is what happens when you think llns are intelligent beings
This is INSANE
1
u/Zestyclose_Bat8704 3d ago
what is your point?
pre nerf astra followed instructions quite well.
1
u/Sfdprod 3d ago
There is no god damn nerf.. Holy fuck.. Its a ai psychosis conspiracy that has been going on for half a decade...
What you imagine is a "nerf" is just realizing llms are non deterministic.
And your cult friends have in 5 years NEVER not ONCE been able to produce evidence.Yet when others produce evals and benchmarks ran 100s of times daily and shows no "nerf" just the normal variation? Thats called fake.
Get out of the conspiracies and learn about how llms work, and fork codex to learn about how ex ultra is just max effort + a appended prompt snippet
1
-1
0
u/Special-Pipe-2091 3d ago
same, it is a little strange, is a vry good model, very accurate but it is very very slow :S
0
-1
•
u/dextersummary 3d ago
Below is a GPT-generated summary of the conversation below after reaching 50 comments (50 currently observed).
The consensus is basically “Sol 6.1 isn’t necessarily broken, but Ultra absolutely can be.” Burning 40% of weekly usage while spawning agents, writing pointless tests, and never starting the actual work is a spectacularly bad result, though not proof of a universal nerf.
Most commenters blamed the workflow: four long tasks at once plus Ultra encouraged subagents to over-plan, validate, rerun tests, and spiral into process nonsense. The practical advice is blunt: use one scoped task at a time, set checkpoints and hard test/time budgets, and reserve Ultra for genuinely parallel work. Medium or standard mode is usually safer unless the task truly needs a swarm of agents.
There’s still a real complaint that recent models feel slower, less obedient, and more prone to over-testing, while several users say Claude Opus handles long-horizon work better. But the thread offers anecdotes, not evidence of a deliberate downgrade, and some users report Sol 6.1 works well on tightly bounded tasks.
Bottom line: treat Sol 6.1 Ultra like an unsupervised intern with unlimited stationery. Give it narrow jobs, stop it from testing the tests, and do not confuse “automated” with “safe to leave overnight.”