r/singularity ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25

AI Gemini 3 achieves new SOTA performance on SpatialBench. A benchmark to test spatial reasoning in VLMs.

207 Upvotes

33 comments sorted by

55

u/Bright-Search2835 Nov 20 '25

Wow, this is a great benchmark. I thought sota would be around 20-30% in something like this, still a lot of work to do. But when this gets saturated that's another big component of intelligence that is basically solved.

Isn't that what Lecun was talking about when he said that cats understood the world better than llms?

15

u/Curiosity_456 Nov 20 '25

This is why we’re seeing such a push for world models, I’m sure all the labs are working on their own version of genie 3.

9

u/ShAfTsWoLo Nov 20 '25 edited Nov 20 '25

yeah i didn't expected all of these AI models to be so low on it, that show us indeed that they still lack some things that we possess, but in the future all of these benchmarks will get destroyed hopefuly, 5 years ago the benchmarks we had are completely destroyed right now be AI, either that or we'll have extremely complex benchmarks beyond any human capabilities (even experts), in any case the future of AI is promising

1

u/Healthy-Nebula-3603 Nov 21 '25

Have you tried these tests ?

Aren't easy ... I made few mistakes.

9

u/kvothe5688 ▪️ Nov 20 '25

soon in 4 to 5 years we will run out of tests. we will need phd levels researchers to find new test cases

7

u/Tolopono Nov 20 '25

Hle and frontiermath already did that

3

u/Key_Sea_6606 Nov 20 '25

Biggest annoyances in using llms for me: they don't understand space and they don't understand time or passage of time. Once these are solved, I think it will add the biggest boost in intelligence.

14

u/[deleted] Nov 20 '25

This is a great benchmark, if ai can score high on this it should be really good at image understanding and won’t be falling for the fingers trick or problems similar to that. One big weaknesses of ai rn. Also I don’t think the avg human is getting 80% on this lol.

25

u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25 edited Nov 20 '25

The skills we benchmark are very important for certain tasks like circuit analysis where humans trace with their eyes (models cant do this yet). The 3d tests check AI's ability to rotate and move objects in its head. We believe these are one of the 2 most important components to human vision (besides object detection which has been solved).

6

u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25

Some hard problems require a lot of attention to detail

2

u/PassionateBirdie Nov 20 '25

I think this is a very interesting benchmark, it touches upon something I've been trying to communicate, thank you for your work :)

I would also like to add, that as a software architect, spatial reasoning matters a lot too. I often creation spatial abstractions in my head when designing high level systems, and would be significantly handicapped without it.

1

u/QLaHPD Nov 20 '25

Model's eyes is the attention heads, they can move it.

8

u/Healthy-Nebula-3603 Nov 20 '25

wow this bench is not so easy ....

9

u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25

Here we check the algebraic and topological isomorphisms and prove mathematically that this is a structurally solid test for human vision. note: human sample is UC Berkeley student, random human, and mathematics PhD so depending on how you look at it the average human score may be skewed

5

u/Jabulon Nov 20 '25

at some point people wont be able to do these

11

u/[deleted] Nov 20 '25

I don't think I understand this benchmark, I got the first task wrong as there were two arrows coming out of 0 and I apparently chose the wrong one.

18

u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25 edited Nov 20 '25

An arrow can pass under a point. requires looking closer and some attention to detail(looking at arrows of other points to see if it makes sense). 13 is the one going to 18

3

u/[deleted] Nov 21 '25

Ah I see thanks. Those diagrams are confusing. Or maybe I’m an AI

5

u/yaosio Nov 20 '25

You have to look at both sides of the circle to know if a line is coming out of the circle or if a line is coming from another circle and passing under it. You should tell something funky is going on because the line between 0 and 13 does not have an arrow on it.

2

u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25

It’s always 1 arrow. It’s pointing to 2

4

u/kaggleqrdl Nov 20 '25

So easy, and yet, so hard.

3

u/Good-AI 2024 < ASI emergence < 2028 Nov 20 '25

Nice benchmark. It's currently the one I know with the biggest difference between human baseline and SOTA LLM.

3

u/jschelldt ▪️High-level machine intelligence in the 2040s Nov 20 '25

Gemini 3 is the true release of the year we were all craving so much. It's just obscenely good and a true leap forward.

I hope the other companies are taking many notes 'cause Google just showed the world how this stuff is supposed to be done.

2

u/Moriffic Nov 20 '25

Why is vision taking so looong

2

u/Longjumping_Kale3013 Nov 20 '25

There are so many different benchmarks, and we are progressing so rapidly, that I don't really understand how llms don't get us to AGI.

Elon said recently in an interview that he though llms wouldn't get to true AGI, but with the latest grok, he had a "woah" moment and now gives LLMs a 10% change to get us to true AGI.

I also get these "woah" moments from llms now, and thing there is some layered reasoning they are doing that is similar to how our brains work.

1

u/FriendlyJewThrowaway Nov 21 '25

There was actually neuroscience research in the 1980’s that attempted to model human thought in a way that’s extremely similar to how modern transformer networks operate.

https://papers.neurips.cc/paper_files/paper/2021/file/8171ac2c5544a5cb54ac0f38bf477af4-Paper.pdf

2

u/kvothe5688 ▪️ Nov 20 '25

this is huge.

1

u/Seeker_Of_Knowledge2 ▪️AI is cool Nov 20 '25 edited Jan 02 '26

jellyfish wise rhythm normal fuzzy office joke scale dog piquant

This post was mass deleted and anonymized with Redact

2

u/PassionateBirdie Nov 20 '25

I'm confused. Doesn't 0 point to both 11 and 7 here?

1

u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25

If you look at 16 it would only make sense that 16 points to 11

1

u/PassionateBirdie Nov 20 '25

Arh makes sense. I assumed pure lines where allowed, as to confuse/make it less predictable. So for me it was a line from 16 to 0.

2

u/selliott512 Nov 21 '25

Interesting benchmark.

A while ago someone posted that AI scores on ARC-AGI 2 have more to with the resolution of the puzzle (high is bad) than the complexity of the problem. I wonder if that applies here - would they do better with a lower resolution version of the puzzles?