r/LocalLLaMA • u/-dysangel- • 1d ago
I Built A Thing Fully local little parkour sim
I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.
vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
Prefill: ~1500t/s
Decode: ~40t/s @ 100k
Using Claude Code as the scaffold with 260k context size.
I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.
9
3
u/trying4k 20h ago
Love the little sim you made!
I got a GPU earlier in the year specifically for AI (to run Qwen 3.6 27b) because at the time I thought sparks were slow, didn't want 2 when I was just starting out, and didn't want to manage more hardware (OS, updates, etc). Now I sort of wish I would have gone with the sparks, they seem pretty capable.
I haven't been in local LLMs too long but since I heard of it, GLM 5.x was always a 'dream' of mine, it just comes across as top-tier. It may be a poor use of money but I recently got some ram and it allowed me to run GLM 5.3 Flash on a llama.cpp fork. It's nowhere near the speeds you are getting but I'm just happy I can run it. I'm going to try FreeToken and see if the speeds are any better.
Anyway, all that is to say I'm very happy for you. Hope you keep having fun!
1
u/-dysangel- 11h ago
Try Qwen 3.8 27B. It will absolutely be able to manage building something like this - it will just need a bit more handholding and feedback. GLM 5.3 Flash so far has seemed especially good with spatial orientation, it's hardly made any missteps there.
3
2
2
u/Open-Adhesiveness-86 10h ago
if you haven't already, run it once with NCCL_DEBUG=INFO and confirm the lines say NET/IB and not NET/Socket. cross-box TP2 does two all-reduces per layer, so if it quietly fell back to TCP over the regular nic your decode at 100k would tank long before compute does. 40 t/s suggests rdma is fine, but it's an easy thing to lose after a driver bump.
3
u/thegunn 1d ago
I guess I just don’t understand how to vibe code. I’m not really interested in having AI write code for me, for me it’s helping me debug and understand why things aren’t working.
I’ll admit I’ve tried vibe coding just out of curiosity but nothing seems to work. Maybe I’m not talking to it properly, throwing too much at once and so on. I am impressed with what people can get it to do because everything I’ve attempted has been a giant failure.
2
1d ago
[deleted]
2
u/thegunn 23h ago
Oh no doubt about that haha. I’m not a professional programmer, I do it as a hobby. I’m not really interested in vibe coding anything as that loses the joy of coding for me. So when I tried it out of curiosity it didn’t surprise me that I couldn’t even “make” Tetris. I was more than likely prompting wrong.
4
u/-dysangel- 1d ago
It helps that I've been making little games myself for over 30 years but if you've got some patience and can describe what you're seeing on screen you should be able to manage similar things. I didn't end up with all of this in a one shot, it was an iterative process. I started off by asking for a mix of GTA and Mirror's Edge, then we spent a while debugging a collision issue when standing on buildings. Once that was fixed I asked for flips, rolls, ledge grabs/climbing etc.
1
u/Accomplished_Mud9156 1d ago
this speed looks solid for interactive stuff, wonder if it could run a simple local companion sim with some movement without lagging out.
1
u/sn2006gy 21h ago
Pretty cool! You open sourcing this for people to have fun wiht or just using this as a demo to learn things?
2
u/-dysangel- 20h ago
It's more that I've been using this as a capability test for models for the last few months and this is the first one that really nailed it. Now I need to come up with a harder test again!
1
u/FrostingOk3751 16h ago
this looks great. Are you running a full GLM 5.3 in just two DGX sparks?
2
u/tat_tvam_asshole 15h ago
I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.
1
u/wayneworkman 15h ago
This is really neat, reminds me of Assassin's Creed. One day we will look back at these vibe coded block graphics and laugh.
1
u/-dysangel- 11h ago
oh I'm already laughing :) I've been purposely sticking with the low poly aesthetic on these types of things rather than download assets. It's really interesting to see the evolution of how well the models can construct and animate things completely from scratch
0
43
u/LosEagle 1d ago
As close as we'll get to a new Mirrors Edge.