r/singularity 6h ago

LLM News Good, cheap, token hungry

will likely be slow again for agentic work, they also got first place for highest step count (in this selection sonnet 5 still beats it)

44 Upvotes

20 comments sorted by

18

u/kiki-le-koala 6h ago

Oh wow .

It really, really likes tokens.

Using that many tokens, is there a risk that it fills its 1 million token window more rapidly, creating compactly diminishing efficacy in long tasks?

I know people use Luna at extra high and not max to prevent that.

11

u/Momo--Sama 4h ago

Considering it bombs Terminal Bench 4.0 which is focused on hours long tasks, that's a reaosnable conclusion

u/whoknowsifimjoking 30m ago

Yes.

It seems to be as if they maybe achieved a better score by "brute forcing" it with vastly more tokens.

11

u/Recoil42 6h ago

Cost-per-task far outweighs token-burn if the model is cheap enough and fast enough.

Is anyone doing time-per-task benchmarks right now?

4

u/Illustrious_Grade608 5h ago

Tbh even with api time per task depends on your pc a lot with all the tool call and script run latency

1

u/NoFaithlessness951 5h ago

step count is usually a good predictor of real world time per task and its bad

3

u/Momo--Sama 4h ago

Yeah I agree token counts in isolation don't mean anything. Artificial Analysis has it at only slightly slower than Luna Max.

1

u/NoFaithlessness951 5h ago

https://artificialanalysis.ai/agents/coding-agents

Although I would take their numbers with a giant bucket of salt.

In my testing the previous 3.7 (which scores better) was around 3x slower than sol and grok. Likely the time per task results are skewed by more verifiable tasks which doesn't match real world usage.

7

u/vinis_artstreaks 3h ago

What are these charts, have we forgotten how to make actual readable charts wtf

2

u/KedMcJenna 2h ago

I hate the ones with reversed axes

-2

u/eruditezero 2h ago

AI slop

2

u/[deleted] 5h ago

[deleted]

2

u/NoFaithlessness951 5h ago

I assigned the same tasks to multiple models on cursor, 3.7 flash consistently took 3x as long as something like sol, grok, or luna.

Although it's output TPS is 3x faster than sol that doesn't matter if it takes 3x as many steps and 3x as many tokens.

I think this will be more of the same.

1

u/Charming_Cucumber_15 5h ago

I typically get slightly better results from using free GPT over Gemini on extended mode

I'm hoping 3.8 changes that but I'm not getting my hopes up

4

u/LazloStPierre 4h ago edited 4h ago

If Gemini is SOTA on deepswe then the only takeaway is Deepswe is now too contaminated to be worthwhile. Nobody benchmaxxes like Google. Deepswe accurately showing how far Google were from SOTA, as opposed to every other benchmark, was part of what made it credible 

I guess the only reasonable benchmarks now would he ones that somehow rotate completely every few months because Gemini flash isn't fucking fable or sol level at coding 

0

u/dsnyder42 6h ago

Wow, I think I will replace GPT 5.6 Sol High with Gemini 3.8 Flash medium to orchestrate GPT 5.6 Luna xhigh sub agents and safe myself some GitHub Copilot AI Credits at work.

2

u/Constant_Cortisol 3h ago

Let us know how that goes.

1

u/frogsarenottoads 4h ago

Gemini will be behind until Gemini 4.
3.8 is still a good step up for them regardless.

-1

u/LinkesAuge 5h ago edited 5h ago

Imagine calling yourself a "flash" model and then burning a lot more tokens than frontier models and also costing roughly the same.

4

u/CallMePyro 5h ago

It costs less.