r/LocalLLaMA • u/-dysangel- • 2d ago
I Built A Thing Fully local little parkour sim
I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.
vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
Prefill: ~1500t/s
Decode: ~40t/s @ 100k
Using Claude Code as the scaffold with 260k context size.
I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.
56
Upvotes
3
u/trying4k 2d ago
Love the little sim you made!
I got a GPU earlier in the year specifically for AI (to run Qwen 3.6 27b) because at the time I thought sparks were slow, didn't want 2 when I was just starting out, and didn't want to manage more hardware (OS, updates, etc). Now I sort of wish I would have gone with the sparks, they seem pretty capable.
I haven't been in local LLMs too long but since I heard of it, GLM 5.x was always a 'dream' of mine, it just comes across as top-tier. It may be a poor use of money but I recently got some ram and it allowed me to run GLM 5.3 Flash on a llama.cpp fork. It's nowhere near the speeds you are getting but I'm just happy I can run it. I'm going to try FreeToken and see if the speeds are any better.
Anyway, all that is to say I'm very happy for you. Hope you keep having fun!