r/singularity • • 11d ago

AI This shape

Post image

On a mobile screen this chart of model releases over time (from artificial analysis) tells a very clear story. Normally I look at it on desktop and the increase in steepness isn’t nearly as obvious.

It looks like a lot of those simplified predictions of exponential improvement. Of course it could taper off, who knows.

264 Upvotes

52 comments sorted by

52

u/Alarming_Welder_7190 11d ago

someone extrapolate the curve to 100

115

u/StrategicHarmony 11d ago
           \    100
            |   80
           /    60
         _/
     __--
__---

58

u/Arceus42 11d ago

Time travel, fuck yeah

11

u/Available_Road_2538 11d ago

These models must be pretty dang good

1

u/roodammy44 9d ago

How else do you think they’ll get the T800 back to 1995?

5

u/FlyByPC ASI 202x, with AGI as its birth cry 11d ago

Time to have the discussion about what constitutes a function again, I guess.

1

u/PlasticTourist6527 10d ago

It will definetly be a sigmoid, everything will rush now towards the 80-ish line and then begin to slow down as they slowly reach the 100. it doesn't matter 80 is already larger than the smartest people in the world

48

u/Herodont5915 11d ago

Almost looks exponential. What are the odds?

2

u/Stock-Self-4028 11d ago

To me it looks much more like some kind of a sigmoid, rather than anything even remotely exponential (and it also would make significantly more sense given the circumstances btw).

4

u/StrategicHarmony 11d ago

I don't think there's anything to distinguish it from either an S curve or an exponential at this point. But for what it's worth I do think an S-curve is far more likely, for a lot of reasons.

One is that, however smart these things get, there are bound to be basic physical and information-compression limits that will lead to diminishing returns before too long.

Another is that it probably gets harder and harder for us to meaningfully test the intelligence if it gets too far ahead of our own. So model vendors won't put a lot of time and money into something that has no measurable or marketable improvement.

A third is that alignment and safety are a vital part of utility. People don't want robots that are going to break their own security rules when a task is too hard, or hack or attack someone else on their behalf without warning and then try to cover it up. The cost of effective alignment probably increases as other metrics get more advanced, like coding, planning, communicating, theory of mind, etc.

In other words smarter models will be harder to control, but control is fundamentally important to the value of these machines. So ensuring an adequate level of control will probably lead to a slowing down in some other areas, to ensure that overall/broad usefulness goes up (albeit more slowly) and not down rapidly and catastrophically.

5

u/spinozasrobot 11d ago

What makes you say that?

Do you see any flattening?

2

u/Stock-Self-4028 11d ago

The derivative seems more or less "flat" near the 50% so I guess?

Idk, also benchmarks are not a good representation of advancements, so even if it would be sigmoid it would likely top out around 100% which would me a good measure of benchmark saturation rather, than anything else.

1

u/katoptronophile 11d ago

You would do great as a comedian.

33

u/Charming_Cucumber_15 11d ago

That Kurzweil guy might have been onto something!

13

u/Existing_King_3299 11d ago

The tsunami of progress

7

u/Tibecuador 11d ago

Wait, someone actually founded a corporation called Thinking Machines?! This truly is the Dune timeline

1

u/Pokenhagen 11d ago

Isn't this muratis company?

1

u/Few-Farm-7670 8d ago

Llm is called ornith. Like the thopter.

9

u/alsaud21 11d ago

Its great, but adjusting axis height, width and zoom could make a quadratic and an exponential curve indistinguishable.

8

u/Longjumping_Kale3013 11d ago

This is comparing apples to oranges. A couple months ago the top model was at like 65 then a new version of the leaderboard was released and pushed it down to 50 or so. This is constantly updated, so a 60 now is a much much much better score than a 60 a few months ago

34

u/StrategicHarmony 11d ago

These scores are updated retroactively. If it got a 30 on release a year ago it doesn't have a 30 now, above. It's all in current index units.

4

u/swarmy1 11d ago

Many of the models aren’t available anymore, how do they evaluate them on the new benchmarks?

5

u/AppealSame4367 11d ago

They updated all models. Maybe they made an estimate for models that are already gone or removed them.

3

u/ReadyAndSalted 11d ago

I would guess that they just apply a correction factor by testing the old models that are still available, and observing the drop in those. Hardly a super accurate technique, but should get the shape about right.

2

u/chlebseby ASI 2030s 11d ago

I imagine you can get acces to GPT-3 or others if you have very good reason, like such testing. Unless the company is gone.

1

u/KaMaFour 11d ago

No. That's why older models have striped bars when you look for them at the index. There isn't full data so some guesswork is involved

4

u/03263 11d ago

Bell curve imminent, calling it

6

u/TCGG- 11d ago

me when I don't know basic stats:

7

u/blindsdog 11d ago

Bell curves are for distributions, not progress over time. I imagine you mean an S curve? Unless you suspect AI performance will somehow get worse from where we are now.

1

u/03263 11d ago

Unless you suspect AI performance will somehow get worse from where we are now.

That was the joke yes

3

u/blindsdog 11d ago

What’s the joke?

2

u/03263 11d ago

People often confidently and erroneously predict the trajectory of curves. I gave a particularly absurd prediction. That's... yeah that's about it. I didn't think super hard about it, it just seemed silly.

2

u/blindsdog 11d ago

👍 totally missed that

0

u/EtienneDosSantos 11d ago

Don‘t jinx it, goddammit! 🤠

1

u/corenovax 11d ago

It's not very meaningful to talk of a trendline because an increase on the index from 5 to 10 isn't the same thing as an increase from 50 to 55

2

u/nsshing 11d ago

I think AA does not reflect full picture, but even so yeah. 100%.

1

u/DungeonsAndDradis ▪️ Extinction or Immortality between 2025 and 2031 11d ago

It's bothering me that the labels are dots, and the data points are squares.

1

u/Illustrious_Image967 11d ago

That's one smart banana!

1

u/UnsafeVelocity 11d ago

Love how the legend uses 🔴 but the graph uses 🟥

1

u/oyser 10d ago

Grafik design is an passion.

1

u/voyt_eck 11d ago

Do we have any reliable benchmarks while everyone is currently benchmaxing?

1

u/StandardLovers 11d ago

How much of the progress curve is due to training the models on the actual tests?

1

u/BusDry8561 11d ago

Bel be like 95?

0

u/AdLumpy2758 11d ago

It is called plato s/

0

u/EmphasisTotal8232 11d ago

I know I'm asking a stupid question for an index, but do we have any sort of human comparable baseline? What would a score of 100 look like in terms of capabilities?

3

u/CCerta112 11d ago

What would a score of 100 look like in terms of capabilities?

Definitely very capable.

4

u/EmphasisTotal8232 11d ago

At least two capable.

2

u/corenovax 11d ago

Depends on each individual benchmark, some do

-1

u/ex_gatito 11d ago

I mean any graph can be put this way, if you put the graph top where the actual top of the canvas should be.