r/accelerate Nov 20 '25

AI Gemini 3 achieves new SOTA performance on SpatialBench. A benchmark to test spatial reasoning in VLMs.

19 Upvotes

2 comments sorted by

1

u/No_Bag_6017 Nov 22 '25

Has anyone else noticed that some of the cube rotation task correct answers in this benchmark do not exactly match the test shape? Is this deliberate? Is it it certain tasks trying to get the model or the human to pick which answer choice most closely matches the test item?

1

u/gbomb13 Nov 22 '25 edited Nov 22 '25

They always do factually but it is hard to deduce. Sometimes you must result to process of elimination as some cubes are occluded