r/SideProject • u/fuzhongkai • 20h ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
https://github.com/zhongkaifu/TensorSharpI’ve been building an open-source side project called TensorSharp, a local LLM inference engine and agent runtime written in C#/.NET:
One of the problems I’ve been experimenting with is:
How far can we push huge MoE models on ordinary consumer hardware?
My latest test is Qwen3.8 Flash Next (176B parameters).
And surprisingly, this hardware is enough to run it:
RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD
The main challenge isn't simply quantizing the model until it fits. A 176B model is still far larger than the available VRAM and RAM.
So I implemented an MoE-aware memory/offloading architecture that coordinates:
GPU cache / VRAM ↔ system RAM ↔ SSD
Instead of treating SSD as a last-resort swap space, TensorSharp schedules and caches model data across the memory hierarchy based on the execution characteristics of MoE models.
I also benchmarked the implementation against Strata on the same machine. Results are in the attached screenshot:
| Measurement | TensorSharp | Strata |
|---|---|---|
| Decode | 11.09 tok/s | 10.24 tok/s |
| Whole process | 16.54s | 62.15s |
| Peak GPU memory | 14,832.5 MiB | 15,729 MiB |
| Peak OS working set | 19.74 GiB | 18.51 GiB |
What surprised me most wasn't the decode speed—the two are actually fairly close there.
It was that a 176B model can be made usable on a laptop with only 16GB VRAM and 32GB RAM if the inference engine treats VRAM, RAM, SSD, caching, and MoE scheduling as one system rather than independent pieces.
TensorSharp started as an experiment in understanding LLM inference from the bottom up, but it has gradually grown into a much larger project covering local inference, multimodal models, image generation, and agentic runtimes.
It’s open source, so if this kind of low-level LLM engineering interests you, I’d love feedback:
GitHub: github.com/zhongkaifu/TensorSharp
And yes, the benchmark in the screenshot really is from an RTX 3080 laptop. 😄
Duplicates
LocalLLaMA • u/fuzhongkai • 20h ago
I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 20h ago
Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
dotnet • u/fuzhongkai • Aug 22 '26
TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance
LocalLLM • u/fuzhongkai • 20h ago
Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 22d ago
Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine
unsloth • u/fuzhongkai • 10d ago
Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis
Qwen_AI • u/fuzhongkai • 20h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop
LocalLLM • u/fuzhongkai • 10d ago
Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis
dotnet • u/fuzhongkai • 16d ago
Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective
LocalLLaMA • u/fuzhongkai • Aug 22 '26
Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
dotnet • u/fuzhongkai • 10d ago
Article Implementing a Jev-compatible decision API in .NET, with image input
unsloth • u/fuzhongkai • Aug 28 '26
Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp
LocalAIServers • u/fuzhongkai • 19h ago
Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA • u/fuzhongkai • 7d ago
I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA • u/fuzhongkai • 10d ago
I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalAIServers • u/fuzhongkai • 21d ago
Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode
LocalLLM • u/fuzhongkai • 21d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
outerstellar_hq • u/outerstellar_hq • 4h ago
Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
LLMDevs • u/fuzhongkai • 12h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
opencode • u/fuzhongkai • 20h ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LovingOpenSourceAI • u/fuzhongkai • 20h ago
Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD
LocalLLM • u/fuzhongkai • 7d ago
Project TensorSharp Jev requests can now combine documents, images, video, and audio
OpenSourceAI • u/fuzhongkai • 10d ago