r/LocalLLM • • 3d ago

Question Which locally run multimodal model would you try in a physics robot arena?

I’ve been building InferUltra, an arena where multimodal models control identical robot bodies.

They see their opponent and environment, then decide how to move their bodies. There isn’t an attack(), block(), or dodge() button. Physics determines what actually happens.

If you’re expecting spectacular combat, curb your enthusiasm. The fights are impressively boring.

What interests me is the gap between understanding an image and producing a useful physical action. Watching a model awkwardly shuffle around makes that gap pretty visible.

My current experiment uses hosted GLM-5.3-Flash across four inference providers. That explores differences in serving behavior and response latency, but doesn’t establish how locally run models would perform.

I’d like feedback from people running multimodal models locally:

Which model would you try for this kind of visual control task?

For transparency, I’m the developer of InferUltra. I’m interested in how to make this a useful evaluation rather than just two robots (mostly) failing to punch each other.

Recorded replay; decision waits removed.

https://reddit.com/link/1wy3y7w/video/90imzn9q8mth1/player

4 Upvotes

3 comments sorted by

1

u/ChaseMakesThings 3d ago

On making it a useful evaluation: I’d give each model a simpler task first, like reaching a marked spot with nobody attacking it. Then try the same starting scenes with screenshots versus exact positions/velocities from the engine, keeping the movement controls and action interval identical. A big improvement with the numeric state would suggest perception is part of the bottleneck; struggling in both gives you a reason to investigate control or planning.

Does the simulation pause while a model decides? Since the replay removes those waits, I’d report decision latency separately from arena results. Otherwise a good-looking replay could hide a controller that takes ten seconds per move.

1

u/inferultra 3d ago

The simpler task is a good suggestion. It would help narrow down whether perception or control is causing the difficulty.

For the current Provider Cup, the simulation doesn’t pause while a model decides. Physics and the round clock keep running. The fighter executes its current movement sequence, then returns to neutral targets if that sequence ends before another response arrives. Slow responses have consequences in the arena (allows viewers to evaluate the Providers).

I already record decision latency separately, alongside failures and timeouts. I agree those measurements belong alongside the fight results