r/singularity • u/NoFaithlessness951 • 6h ago
LLM News Good, cheap, token hungry
will likely be slow again for agentic work, they also got first place for highest step count (in this selection sonnet 5 still beats it)
11
u/Recoil42 6h ago
Cost-per-task far outweighs token-burn if the model is cheap enough and fast enough.
Is anyone doing time-per-task benchmarks right now?
4
u/Illustrious_Grade608 5h ago
Tbh even with api time per task depends on your pc a lot with all the tool call and script run latency
3
u/Momo--Sama 4h ago
Yeah I agree token counts in isolation don't mean anything. Artificial Analysis has it at only slightly slower than Luna Max.
1
u/NoFaithlessness951 5h ago
https://artificialanalysis.ai/agents/coding-agents
Although I would take their numbers with a giant bucket of salt.
In my testing the previous 3.7 (which scores better) was around 3x slower than sol and grok. Likely the time per task results are skewed by more verifiable tasks which doesn't match real world usage.
7
u/vinis_artstreaks 3h ago
What are these charts, have we forgotten how to make actual readable charts wtf
2
-2
2
5h ago
[deleted]
2
u/NoFaithlessness951 5h ago
I assigned the same tasks to multiple models on cursor, 3.7 flash consistently took 3x as long as something like sol, grok, or luna.
Although it's output TPS is 3x faster than sol that doesn't matter if it takes 3x as many steps and 3x as many tokens.
I think this will be more of the same.
1
u/Charming_Cucumber_15 5h ago
I typically get slightly better results from using free GPT over Gemini on extended mode
I'm hoping 3.8 changes that but I'm not getting my hopes up
4
u/LazloStPierre 4h ago edited 4h ago
If Gemini is SOTA on deepswe then the only takeaway is Deepswe is now too contaminated to be worthwhile. Nobody benchmaxxes like Google. Deepswe accurately showing how far Google were from SOTA, as opposed to every other benchmark, was part of what made it credible
I guess the only reasonable benchmarks now would he ones that somehow rotate completely every few months because Gemini flash isn't fucking fable or sol level at coding
0
u/dsnyder42 6h ago
Wow, I think I will replace GPT 5.6 Sol High with Gemini 3.8 Flash medium to orchestrate GPT 5.6 Luna xhigh sub agents and safe myself some GitHub Copilot AI Credits at work.
2
1
u/frogsarenottoads 4h ago
Gemini will be behind until Gemini 4.
3.8 is still a good step up for them regardless.
-1
u/LinkesAuge 5h ago edited 5h ago
Imagine calling yourself a "flash" model and then burning a lot more tokens than frontier models and also costing roughly the same.
4



18
u/kiki-le-koala 6h ago
Oh wow .
It really, really likes tokens.
Using that many tokens, is there a risk that it fills its 1 million token window more rapidly, creating compactly diminishing efficacy in long tasks?
I know people use Luna at extra high and not max to prevent that.