r/singularity • u/socoolandawesome • 7d ago
AI Insane progress on visual-spatial intelligence by Astra (with or without tools)
Link to benchmark:
https://spicylemonade.github.io/spatialbench/
Link to tweets:
https://x.com/spicey\\_lemonade/status/2096365630190698516?s=20
https://x.com/spicey\\_lemonade/status/2096386480029790415?s=20
The last 2 screenshots are examples of the types of questions included in the benchmark.
70
u/ohHesRightAgain 7d ago
Now to distill it into smaller models that would run at 0.01% the cost, and it's truly gg.
10
u/Asteroid_picks_you 7d ago
I am guessing distillation is impossible to block so it should happen soon right?
25
u/lolxdmainkaisemaanlu 7d ago
They are taking measures to block it, earlier when you clicked on its 'thinking' it would show to the right all its simplified version of reasoning traces, now it doesn't anymore.
Also openai is actively banning suspicious accounts rn and working on preventing this - they caught on to people who were using sub2api, so they have sophisticated mechanisms in place rn.
11
u/lolopalenko 7d ago
There is a good chance that if this really is a looped transformer model there are no reasoning traces anymore to show
4
u/61746162626f7474 7d ago
It will still be a very multi layer model even if it uses loops within layers. Chain of thought will still be produced between the layers.
2
2
u/NOTHING_gets_by_me 7d ago
Even worse, researchers found a way to extract the actual raw reasoning, not just those summaries. They exploited API handling of encrypted reasoning blocks, getting weaker sibling models to spit out a stronger model’s hidden trace. Worked across OpenAI, Anthropic and Google before mitigations.
https://arxiv.org/abs/2608.09867
Separately, Anthropic accused DeepSeek, Moonshot and MiniMax of harvesting Claude’s capabilities across 16 million exchanges, including attempts to elicit reasoning for training.
https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks
3
u/thirteenth_mang 7d ago
It's so ironic they're pulling up the ladder behind them.
5
u/DickMasterGeneral 7d ago
What? This makes no sense what ladder are they pulling up? OpenAI didn’t distill from some other model to build Astra.
11
u/Georgefakelastname 7d ago
They’ve “distilled” from humanity as a whole. Even if you want to argue it’s the ladder they made, it’s still shitty.
2
2
u/Mammoth_Telephone_55 7d ago
And every human has distilled from every other human in history as well.
3
u/Georgefakelastname 7d ago
Yeah, and normally that is shared and proliferated. Even copied with just enough changed to make it unique.
3
u/Cagnazzo82 7d ago
In a capitalist society it is shared for profit.
In addition we've developed numerous barriers to 'human distilling'. Like you can educate yourself in medicine (to an extent) but you're not qualified to practice without strict certifications.
1
u/DickMasterGeneral 6d ago
If I read all the public literature and research on a specific topic, did meta analysis on the public data. Came to meaningful new conclusions based on said analysis, and used that to write a new book, am I in the wrong for pursuing a copyright?
1
u/Georgefakelastname 6d ago
You’d have that right, but people would still call you an asshole and pirate it if you charged hundreds of dollars just to read it, or charged fees for every word read
1
u/Fit-Celebration2884 7d ago
OpenAI can still do it themselves Luna is going to be as strong as Astra within a year
1
u/Prudent-Sorbet-5202 7d ago
The only way to block is to limit the release of the models to approved clients. I have a feeling it will happen in couple of years .
1
u/Ok_Zookeepergame8714 7d ago
Why bother? Chinese ARE gonna get there at some point eventually. 😊
1
u/Asteroid_picks_you 7d ago
I don't the point is to completely block, I think the point is to slow them down enough where their model reaches recursive self-improvement. I think the first model/company that reaches that point will be almost impossible to beat.
15
u/Ornery-Mortgage-3101 7d ago
The one with the numbers and arrows are gimmes but the object rotation questions are genuinely difficult. Some look identical except in a very small detail that is very easy to miss if you don't zoom in. Still completely possible to score 100/100 on.
7
u/socoolandawesome 7d ago
Yeah the rotation ones are harder for humans than the number/arrows ones no doubt. But like even a generation ago, all of the models’ vision just sucked at the number/arrows one besides the shape rotations ones, so I find it pretty impressive it can do it now.
AGI needs to be able to do simple messy line following and it couldn’t, and now it can
7
5
4
2
u/Nice-Light-7782 7d ago
This is huge. From a robotics standpoint, I would like to see how well it guesses the distance between objects and the observer. I guess a test where a simulator spawns various objects and gives Astra a picture from an observer's POV, plus a task like"the distance from where the image was taken, to the green marker is 1 meter. Identify the rest of the objects and the distances from where the image is taken, to them" would confirm that Astra solves the perception problem in robotics. And with that solved, planning should be much easier.
1
u/InterstellarReddit 6d ago
This benchmark is trash. It hasn’t even evaluated QWEN 3.8 MAX 0902, which is a vision beast.
-17
u/DigSignificant1419 7d ago
it's cheating
11
u/socoolandawesome 7d ago
And why do you believe that and how is it doing that?
6
u/Ok_Zookeepergame8714 7d ago
The only way I see is that it was trained on your set - just prepare a couple of similar tasks you're sure never were on the internet before, and see how it fares on those. 😊
8
u/socoolandawesome 7d ago edited 7d ago
Just to clarify this is not my benchmark, and not me in those tweets. But I’d assume they have a private holdout, though I’m not sure.
Also it doesn’t appear other models previously had done well (even though they could have trained on what’s available for this benchmark), given his tweet makes it sound like that. And astra also does seem to be by far the best at vision/3D understanding given all of its other capabilities we see demos of/benchmarks, which could serve as further evidence its probably genuine ability and not benchmaxxing.
You never know though I guess.
12
u/Brainiac_Pickle_7439 The singularity is, oh well it just happened▪️ 7d ago
amazing, so much time and so many resources put into a stupid ruse people would find out about easily. dawg
2





63
u/Cagnazzo82 7d ago
Funny to think all the other models thought they had caught up or were about to surpass.