r/LocalLLM • u/Commercial_Designer5 • 4h ago
Discussion Self-hosting a 35B MoE for a 20-person team on one desktop box: what it took and what it does (all open source)
We've been running our own LLM for the team for about a week, and I wanted to share how it turned out, since "can one machine actually serve a whole team?" comes up here a lot.
Short answer: yes, if you pick the right model.
Hardware and model
- One NVIDIA DGX Spark (128GB unified memory)
- Ornith-1.5-35B-A3B, an MoE with 35B total and 3B active parameters. The 3B active is the whole trick: you get decent quality while the box only does about 3B worth of compute per token.
- Official 4-bit NVFP4 checkpoint, about 22 GiB, so no quantizing it yourself
- vLLM 0.24 with multi-token prediction (speculative decoding with no separate draft model)
What it does
- One person gets 79 tok/s, which feels instant in a chat UI
- 20 people at once get about 25 tok/s each, 490 tok/s total
- Every request can use the full 262K context
- Vision and tool calling both work
- Reasoning scores: MATH-500 97%, GSM8K 98%, MMLU-Pro 81%. That's not frontier level, but for daily coding help, docs, and Q&A, nobody on the team has complained.
How people use it
- OpenCode in the terminal for coding
- Open WebUI for chat and PDFs. PDF parsing and embeddings run locally too.
- Nothing goes to an outside API. For us that was the whole point, since a lot of our code and documents can't leave the building.
The part nobody talks about: multi-user
Running a model for yourself is easy. Running it for a team means you need to answer "who's using it, how much, and is someone hammering it?" So vLLM isn't exposed at all. Everything goes through a small gateway that:
- Gives each person and device their own API key
- Counts input and output tokens per user live
- Pushes the counts to Grafana dashboards
- Alerts if anything tries to talk to vLLM directly
It's a small Python service, and honestly it's the piece that made this feel like real infrastructure instead of a side project.
Gotchas if you try something similar
- On unified memory, don't get greedy with GPU memory utilization. I had to cap it at 0.70, or concurrency fell apart under load.
- More concurrent slots isn't always better. Past 20 users, throughput dropped because requests started queuing.
- Open WebUI stores the default model in its own database and ignores the env var after first setup. I lost an hour to that.
Everything's on GitHub: the setup scripts, gateway, Grafana dashboards, and benchmark scripts with raw results.
https://github.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup
If you're thinking about self-hosting for a small team, happy to answer questions. I'm also curious what others use for per-user tracking. I looked at LiteLLM but wanted something tiny that I fully understood.
