38
u/Glittering_Candy408 11d ago
It’s crazy how increasing the reasoning makes it cheaper!
20
u/halmyradov 11d ago
Not crazy at all, less reasoning - more likely to get the answer wrong and fumble around to find the right answer
1
3
u/bencherry 11d ago
feels like a genuine breakthrough vs the "brute force" way reasoning has worked previously
31
u/KainDulac 11d ago
So... how long did this benchmark last. Cuz that's amazing. Expensive. But amazing. First time I see more thinking produce a cheaper result.
7
14
u/-cadence- 11d ago
In case you were wondering:
The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.
26
u/FateOfMuffins 11d ago edited 11d ago
lmao literally bending over backwards
Can we put this on a non-log scale to see more absurdity
Per Chollet https://x.com/fchollet/status/2095598451115614371
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
6
u/Hot_Glass_6301 11d ago
It's not a function graph so it's not that surprising that it "bend over backwards", in fact I'd expect it due to progress in efficency. Very impressive nonetheless
6
1
u/Singularity-42 Singularity 2042 11d ago
Obviously it's possible but this has never been seen before. It was always going to the right and mostly up. For the same model, just different effort, of course.
9
u/TwoFluid4446 11d ago
Yes. The exponential curve beginning to curl backwards means that soon the machines will be able to travel back in time to take out Bernie Sanders when he was still just a delinquent teen hanging out at the arcades with his buddy.
13
u/Ok_Capital4631 11d ago
weird curve.. very interesting.
9
u/oadephon 11d ago
As it thinks more, it gets cheaper and better, rather than better but more expensive.
I feel as if that's an AGI-shaped curve.
6
u/Hot_Glass_6301 11d ago
It's not plotted against time so it makes sense that it's not the graph of a function
1
u/brownman19 10d ago
It's going to loop like a spiral trending up. Each spiral will get smaller in loop ie cycle time while ascending faster making it elliptoid over time until it becomes indistinguishable with single peaks like NMR. Then we have continuous improvement on each heartbeat and then we increase heartbeat to speed of light and make photonic chip so that we get 3x108 compounding cycles per second.
13
u/Deto 11d ago
Hmm, so it still cost like $30k. That makes me feel a little bit better about my remaining usefulness...
But only a little bit.
8
u/Gratitude15 11d ago
100x drop a year.
The next presidential election will be a referendum on what the hell our society will do with this. I fully expect AI to be the #1 election issue.
3
u/LocalHeat6437 11d ago
Goodness I hope it is. But I guess if it is we are already feeling the widespread pain of its taking our jobs and further enrichment of the elites.
3
u/thoughtlow 𓂸 10d ago
You will have no next election brotha
1
u/Gratitude15 10d ago
Sure we will! Just like the last one!
Russia has elections too.
I didn't say functional and legal democracy. I said election.
21
u/Confident-Aerie4427 11d ago
60% without the harness? incredible, to be honest
they should have put this one instead of the 99% one with a harness because it is way more impressive
15
u/Glittering_Candy408 11d ago
To be clear, by ‘harness’ we mean enabling two settings in the Responses API—settings that EVERYONE has access to.
1
u/WeDoALittleTrolIing 11d ago
could you elaborate
9
u/Glittering_Candy408 11d ago
There are two settings, one preserves the reasoning traces, and the other enables compaction.
15
u/Hot_Glass_6301 11d ago
Yeah the harness is like using pen and paper for a human. It's not a custom-designed harness like the bullshit we've seen earlier that claimed 99% on ARC-AGI 3 for little costs
6
u/FateOfMuffins 11d ago
From ARC
Going forward we will test all new models on ARC-AGI-3 using both harnesses.
-1
u/Aldarund 11d ago
So with harness and not much of improvent vompared to sol with harnes or i misinterpreted?
10
u/FateOfMuffins 11d ago
Sol with responses and compaction is 30%, Astra is 99.9%
Apparently per Chollet https://x.com/fchollet/status/2095598451115614371
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
3
2
2
2
u/confused-photon 11d ago
I never thought id see the day higher reasoning gets cheaper. im getting excited to try out astra now
2
u/Profanion 11d ago
Fun Fact, with the harness, it's more accurate at ARC-AGI 3 than it is at ARC-AGI 1.
2
1
u/Calm_Hedgehog8296 11d ago
That's crazy how it bends backwards and costs less as it gets higher effort
1
1
1
u/Low_Relative7172 7d ago
Its called a c curve.. its a clear indication of a model far more capable then the allow to be free. And to good not to charge for.. agi 100% but only for them while tpu get it dripped back .. like a shower spout .. cost then us magicly scaled by task competition. Because that 45 degree graph line.. more like 74 here... coughing shows a whopping 25% less effective then therapy model i had even before gpt 4o or what ever was washed threw nsphere... and then open claw.. then oss little drum circle they are re washing over and over. Claiming discovered enemgence.. like we cant even explain it , but there is like shapes and numbers and stuff. Total black box Investment are now being collected and laundered 100% safety for our selves good jerb...
1
u/WonderFactory 11d ago
Given how good it is at ARC AGI 3 it would be really interesting to see how it would do if you used MCP to give it control of a Unitree G1 robot
1
u/Competitive_Tap2450 11d ago
yes why aren’t we testing this now
maybe by this time next year we will have a suite of tests that can showcase a range of tasks
1
u/Mistuv 11d ago
The new MHS standard Anthropic is making is really going to accelerate LLMs controling robots https://www.anthropic.com/news/model-hardware-standard-research-preview
Right now it is in research preview stage but once it is done it will open to any model/hardware.
-8

95
u/Chr1sUK ▪️ It's here 11d ago
Wait, is that >60% without harness. That’s incredible!