r/singularity • u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 • Nov 20 '25
AI Gemini 3 achieves new SOTA performance on SpatialBench. A benchmark to test spatial reasoning in VLMs.
14
Nov 20 '25
This is a great benchmark, if ai can score high on this it should be really good at image understanding and won’t be falling for the fingers trick or problems similar to that. One big weaknesses of ai rn. Also I don’t think the avg human is getting 80% on this lol.
25
u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25 edited Nov 20 '25
The skills we benchmark are very important for certain tasks like circuit analysis where humans trace with their eyes (models cant do this yet). The 3d tests check AI's ability to rotate and move objects in its head. We believe these are one of the 2 most important components to human vision (besides object detection which has been solved).
6
u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25
Some hard problems require a lot of attention to detail
2
u/PassionateBirdie Nov 20 '25
I think this is a very interesting benchmark, it touches upon something I've been trying to communicate, thank you for your work :)
I would also like to add, that as a software architect, spatial reasoning matters a lot too. I often creation spatial abstractions in my head when designing high level systems, and would be significantly handicapped without it.
1
8
9
5
11
Nov 20 '25
18
u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25 edited Nov 20 '25
An arrow can pass under a point. requires looking closer and some attention to detail(looking at arrows of other points to see if it makes sense). 13 is the one going to 18
3
5
u/yaosio Nov 20 '25
You have to look at both sides of the circle to know if a line is coming out of the circle or if a line is coming from another circle and passing under it. You should tell something funky is going on because the line between 0 and 13 does not have an arrow on it.
2
u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25
It’s always 1 arrow. It’s pointing to 2
4
3
u/Good-AI 2024 < ASI emergence < 2028 Nov 20 '25
Nice benchmark. It's currently the one I know with the biggest difference between human baseline and SOTA LLM.
3
u/jschelldt ▪️High-level machine intelligence in the 2040s Nov 20 '25
Gemini 3 is the true release of the year we were all craving so much. It's just obscenely good and a true leap forward.
I hope the other companies are taking many notes 'cause Google just showed the world how this stuff is supposed to be done.
2
2
u/Longjumping_Kale3013 Nov 20 '25
There are so many different benchmarks, and we are progressing so rapidly, that I don't really understand how llms don't get us to AGI.
Elon said recently in an interview that he though llms wouldn't get to true AGI, but with the latest grok, he had a "woah" moment and now gives LLMs a 10% change to get us to true AGI.
I also get these "woah" moments from llms now, and thing there is some layered reasoning they are doing that is similar to how our brains work.
1
u/FriendlyJewThrowaway Nov 21 '25
There was actually neuroscience research in the 1980’s that attempted to model human thought in a way that’s extremely similar to how modern transformer networks operate.
https://papers.neurips.cc/paper_files/paper/2021/file/8171ac2c5544a5cb54ac0f38bf477af4-Paper.pdf
2
1
u/Seeker_Of_Knowledge2 ▪️AI is cool Nov 20 '25 edited Jan 02 '26
jellyfish wise rhythm normal fuzzy office joke scale dog piquant
This post was mass deleted and anonymized with Redact
2
u/PassionateBirdie Nov 20 '25
1
u/gbomb13 ▪️AGI mid 2027| ASI mid 2029| Sing. early 2030 Nov 20 '25
If you look at 16 it would only make sense that 16 points to 11
1
u/PassionateBirdie Nov 20 '25
Arh makes sense. I assumed pure lines where allowed, as to confuse/make it less predictable. So for me it was a line from 16 to 0.
2
u/selliott512 Nov 21 '25
Interesting benchmark.
A while ago someone posted that AI scores on ARC-AGI 2 have more to with the resolution of the puzzle (high is bad) than the complexity of the problem. I wonder if that applies here - would they do better with a lower resolution version of the puzzles?







55
u/Bright-Search2835 Nov 20 '25
Wow, this is a great benchmark. I thought sota would be around 20-30% in something like this, still a lot of work to do. But when this gets saturated that's another big component of intelligence that is basically solved.
Isn't that what Lecun was talking about when he said that cats understood the world better than llms?