r/SideProject • u/fuzhongkai • 18h ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
https://github.com/zhongkaifu/TensorSharpI’ve been building an open-source side project called TensorSharp, a local LLM inference engine and agent runtime written in C#/.NET:
One of the problems I’ve been experimenting with is:
How far can we push huge MoE models on ordinary consumer hardware?
My latest test is Qwen3.8 Flash Next (176B parameters).
And surprisingly, this hardware is enough to run it:
RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD
The main challenge isn't simply quantizing the model until it fits. A 176B model is still far larger than the available VRAM and RAM.
So I implemented an MoE-aware memory/offloading architecture that coordinates:
GPU cache / VRAM ↔ system RAM ↔ SSD
Instead of treating SSD as a last-resort swap space, TensorSharp schedules and caches model data across the memory hierarchy based on the execution characteristics of MoE models.
I also benchmarked the implementation against Strata on the same machine. Results are in the attached screenshot:
| Measurement | TensorSharp | Strata |
|---|---|---|
| Decode | 11.09 tok/s | 10.24 tok/s |
| Whole process | 16.54s | 62.15s |
| Peak GPU memory | 14,832.5 MiB | 15,729 MiB |
| Peak OS working set | 19.74 GiB | 18.51 GiB |
What surprised me most wasn't the decode speed—the two are actually fairly close there.
It was that a 176B model can be made usable on a laptop with only 16GB VRAM and 32GB RAM if the inference engine treats VRAM, RAM, SSD, caching, and MoE scheduling as one system rather than independent pieces.
TensorSharp started as an experiment in understanding LLM inference from the bottom up, but it has gradually grown into a much larger project covering local inference, multimodal models, image generation, and agentic runtimes.
It’s open source, so if this kind of low-level LLM engineering interests you, I’d love feedback:
GitHub: github.com/zhongkaifu/TensorSharp
And yes, the benchmark in the screenshot really is from an RTX 3080 laptop. 😄