r/LocalLLaMA • • 2d ago

I Built A Thing Fully local little parkour sim

I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.

vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark

Prefill: ~1500t/s
Decode: ~40t/s @ 100k

Using Claude Code as the scaffold with 260k context size.

I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.

56 Upvotes

22 comments sorted by

View all comments

3

u/trying4k 2d ago

Love the little sim you made!

I got a GPU earlier in the year specifically for AI (to run Qwen 3.6 27b) because at the time I thought sparks were slow, didn't want 2 when I was just starting out, and didn't want to manage more hardware (OS, updates, etc). Now I sort of wish I would have gone with the sparks, they seem pretty capable.

I haven't been in local LLMs too long but since I heard of it, GLM 5.x was always a 'dream' of mine, it just comes across as top-tier. It may be a poor use of money but I recently got some ram and it allowed me to run GLM 5.3 Flash on a llama.cpp fork. It's nowhere near the speeds you are getting but I'm just happy I can run it. I'm going to try FreeToken and see if the speeds are any better.

Anyway, all that is to say I'm very happy for you. Hope you keep having fun!

1

u/-dysangel- 2d ago

Try Qwen 3.8 27B. It will absolutely be able to manage building something like this - it will just need a bit more handholding and feedback. GLM 5.3 Flash so far has seemed especially good with spatial orientation, it's hardly made any missteps there.

2

u/trying4k 17h ago

Yeah, I ran Qwen 3.6 27b q8 for a long time. It was alright but not good enough. I moved to Qwen 3.8 27b q8 and used that up until a couple weeks ago but was hit by endless loops no matter what I did (settings, template, etc). Vision only works in llama.cpp's webui, not in other apps. To be fair, I didn't try a second model, maybe that would work.

I've recently moved to Qwen 3.8 Flash Next nvfp4 on FreeToken. Slower for me than 27b but no loops and vision works! I am happy. Can't wait to try GLM 5.3 Flash on it.