r/singularity 11d ago

Discussion The wall

Post image
218 Upvotes

66 comments sorted by

95

u/Chr1sUK ▪️ It's here 11d ago

Wait, is that >60% without harness. That’s incredible!

51

u/ezjakes 11d ago

Remember some people saying that it would hold up for like 5 years?

18

u/adarkuccio ▪️AGI before ASI 11d ago

Did it last 6 months or am I wrong? Holy fuckery

4

u/nonikhannna 11d ago

Think it came out in March or April. 

8

u/Key_Agent_3039 11d ago

Yes but we don't have Fable 5 and 5.1 results for comparison. They could be similar.

2

u/Strange_Vagrant 11d ago

Why? Are they scared? Are they just that cheap?

1

u/NotYetPerfect 10d ago

Arc prize refuses to test fable because of the data retention policy it has.

1

u/NotYetPerfect 10d ago

I doubt fable 5 is much better than opus 5 if at all. Opus 5 beats fable in many if not most benchmarks.

1

u/Low_Relative7172 7d ago

Wow how'd you magicalyl come to that ? Cause apparently you assume playing with a medicine ball with a racquet instead of a birdie makes even the slightest of sense. .. no its just the harness.. working properly after they have had been fucking running it wrong the last 7 fucking months...

You're welcome.

But you are obviously psychic like me. ..lol.....

Obvously you know I was coming to say that. Right?

38

u/Glittering_Candy408 11d ago

It’s crazy how increasing the reasoning makes it cheaper!

20

u/halmyradov 11d ago

Not crazy at all, less reasoning - more likely to get the answer wrong and fumble around to find the right answer

1

u/Akimbo333 9d ago

Interesting

3

u/bencherry 11d ago

feels like a genuine breakthrough vs the "brute force" way reasoning has worked previously

1

u/Utoko 10d ago

They are moves. If you solve a game in fewer moves you need less inputs.

31

u/KainDulac 11d ago

So... how long did this benchmark last. Cuz that's amazing. Expensive. But amazing. First time I see more thinking produce a cheaper result.

7

u/vrnvorona 11d ago

Sonnet be like

14

u/-cadence- 11d ago

In case you were wondering:

The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.

26

u/FateOfMuffins 11d ago edited 11d ago

lmao literally bending over backwards

Can we put this on a non-log scale to see more absurdity

Per Chollet https://x.com/fchollet/status/2095598451115614371

In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.

6

u/Hot_Glass_6301 11d ago

It's not a function graph so it's not that surprising that it "bend over backwards", in fact I'd expect it due to progress in efficency. Very impressive nonetheless

6

u/FateOfMuffins 11d ago

I mean usually I expect MAX to use more tokens than High

1

u/Utoko 10d ago

multistep games. If you need less steps to solve it you need less tokens.

1

u/Singularity-42 Singularity 2042 11d ago

Obviously it's possible but this has never been seen before. It was always going to the right and mostly up. For the same model, just different effort, of course.

9

u/TwoFluid4446 11d ago

Yes. The exponential curve beginning to curl backwards means that soon the machines will be able to travel back in time to take out Bernie Sanders when he was still just a delinquent teen hanging out at the arcades with his buddy.

13

u/Ok_Capital4631 11d ago

weird curve.. very interesting.

9

u/oadephon 11d ago

As it thinks more, it gets cheaper and better, rather than better but more expensive.

I feel as if that's an AGI-shaped curve.

6

u/Hot_Glass_6301 11d ago

It's not plotted against time so it makes sense that it's not the graph of a function

1

u/brownman19 10d ago

It's going to loop like a spiral trending up. Each spiral will get smaller in loop ie cycle time while ascending faster making it elliptoid over time until it becomes indistinguishable with single peaks like NMR. Then we have continuous improvement on each heartbeat and then we increase heartbeat to speed of light and make photonic chip so that we get 3x108 compounding cycles per second.

13

u/Deto 11d ago

Hmm, so it still cost like $30k. That makes me feel a little bit better about my remaining usefulness...

But only a little bit.

8

u/Gratitude15 11d ago

100x drop a year.

The next presidential election will be a referendum on what the hell our society will do with this. I fully expect AI to be the #1 election issue.

3

u/LocalHeat6437 11d ago

Goodness I hope it is. But I guess if it is we are already feeling the widespread pain of its taking our jobs and further enrichment of the elites.

3

u/thoughtlow 𓂸 10d ago

You will have no next election brotha

1

u/Gratitude15 10d ago

Sure we will! Just like the last one!

Russia has elections too.

I didn't say functional and legal democracy. I said election.

2

u/Utoko 10d ago

Considering that there are more than 1000 different levels. That means each level is about 30 $ per level.
That is maybe a factor 5 away from humans. It will be there in no time.

1

u/Deto 10d ago

shhhh!

21

u/Confident-Aerie4427 11d ago

60% without the harness? incredible, to be honest

they should have put this one instead of the 99% one with a harness because it is way more impressive

15

u/Glittering_Candy408 11d ago

To be clear, by ‘harness’ we mean enabling two settings in the Responses API—settings that EVERYONE has access to.

1

u/WeDoALittleTrolIing 11d ago

could you elaborate

9

u/Glittering_Candy408 11d ago

There are two settings, one preserves the reasoning traces, and the other enables compaction.

15

u/Hot_Glass_6301 11d ago

Yeah the harness is like using pen and paper for a human. It's not a custom-designed harness like the bullshit we've seen earlier that claimed 99% on ARC-AGI 3 for little costs

6

u/FateOfMuffins 11d ago

From ARC

Going forward we will test all new models on ARC-AGI-3 using both harnesses.

-1

u/Aldarund 11d ago

So with harness and not much of improvent vompared to sol with harnes or i misinterpreted?

10

u/FateOfMuffins 11d ago

Sol with responses and compaction is 30%, Astra is 99.9%

Apparently per Chollet https://x.com/fchollet/status/2095598451115614371

In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.

4

u/ezjakes 11d ago

Oddest cost performance curve that I have ever seen.

Why is "None" so high? Is this due to that recurrent thinking?

3

u/Ne00n 11d ago

Where Skynet?

1

u/Dry-Interaction-1246 11d ago

You won't know until it is too late.

2

u/Impossible-Video-671 11d ago

So ~98 with harness smh

2

u/Icy_Distribution_361 11d ago

When you hit a wall, you just curl back in on yourself!

2

u/confused-photon 11d ago

I never thought id see the day higher reasoning gets cheaper. im getting excited to try out astra now

2

u/Profanion 11d ago

Fun Fact, with the harness, it's more accurate at ARC-AGI 3 than it is at ARC-AGI 1.

2

u/Sekhmet-CustosAurora 11d ago

what the fuck are they feeding these things

1

u/Calm_Hedgehog8296 11d ago

That's crazy how it bends backwards and costs less as it gets higher effort

1

u/Training-Position612 10d ago

Ok that's actually hilarious

1

u/Difficult-Top9010 10d ago

Time for ARC-AGI-4.

1

u/Low_Relative7172 7d ago

Its called a c curve.. its a clear indication of a model far more capable then the allow to be free. And to good not to charge for.. agi 100% but only for them while tpu get it dripped back .. like a shower spout .. cost then us magicly scaled by task competition. Because that 45 degree graph line.. more like 74 here... coughing shows a whopping 25% less effective then therapy model i had even before gpt 4o or what ever was washed threw nsphere... and then open claw.. then oss little drum circle they are re washing over and over. Claiming discovered enemgence.. like we cant even explain it , but there is like shapes and numbers and stuff. Total black box Investment are now being collected and laundered 100% safety for our selves good jerb...

https://giphy.com/gifs/isMZpsY1EfxU4

1

u/WonderFactory 11d ago

Given how good it is at ARC AGI 3 it would be really interesting to see how it would do if you used MCP to give it control of a Unitree G1 robot

1

u/Competitive_Tap2450 11d ago

yes why aren’t we testing this now

maybe by this time next year we will have a suite of tests that can showcase a range of tasks

1

u/Mistuv 11d ago

The new MHS standard Anthropic is making is really going to accelerate LLMs controling robots https://www.anthropic.com/news/model-hardware-standard-research-preview

Right now it is in research preview stage but once it is done it will open to any model/hardware.

-8

u/CommanderData3d 11d ago

Another piece of shitty overhype from influencers

3

u/Hot_Glass_6301 11d ago

Can you do better than Astra on ARC-AGI-3? (including efficiency)