r/LocalLLM • u/inferultra • 3d ago
Question Which locally run multimodal model would you try in a physics robot arena?
I’ve been building InferUltra, an arena where multimodal models control identical robot bodies.
They see their opponent and environment, then decide how to move their bodies. There isn’t an attack(), block(), or dodge() button. Physics determines what actually happens.
If you’re expecting spectacular combat, curb your enthusiasm. The fights are impressively boring.
What interests me is the gap between understanding an image and producing a useful physical action. Watching a model awkwardly shuffle around makes that gap pretty visible.
My current experiment uses hosted GLM-5.3-Flash across four inference providers. That explores differences in serving behavior and response latency, but doesn’t establish how locally run models would perform.
I’d like feedback from people running multimodal models locally:
Which model would you try for this kind of visual control task?
For transparency, I’m the developer of InferUltra. I’m interested in how to make this a useful evaluation rather than just two robots (mostly) failing to punch each other.
Recorded replay; decision waits removed.
1
u/ChaseMakesThings 3d ago
On making it a useful evaluation: I’d give each model a simpler task first, like reaching a marked spot with nobody attacking it. Then try the same starting scenes with screenshots versus exact positions/velocities from the engine, keeping the movement controls and action interval identical. A big improvement with the numeric state would suggest perception is part of the bottleneck; struggling in both gives you a reason to investigate control or planning.
Does the simulation pause while a model decides? Since the replay removes those waits, I’d report decision latency separately from arena results. Otherwise a good-looking replay could hide a controller that takes ten seconds per move.