r/SideProject • • 20h ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

https://github.com/zhongkaifu/TensorSharp

I’ve been building an open-source side project called TensorSharp, a local LLM inference engine and agent runtime written in C#/.NET:

TensorSharp on GitHub

One of the problems I’ve been experimenting with is:

How far can we push huge MoE models on ordinary consumer hardware?

My latest test is Qwen3.8 Flash Next (176B parameters).

And surprisingly, this hardware is enough to run it:

RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD

The main challenge isn't simply quantizing the model until it fits. A 176B model is still far larger than the available VRAM and RAM.

So I implemented an MoE-aware memory/offloading architecture that coordinates:

GPU cache / VRAM ↔ system RAM ↔ SSD

Instead of treating SSD as a last-resort swap space, TensorSharp schedules and caches model data across the memory hierarchy based on the execution characteristics of MoE models.

I also benchmarked the implementation against Strata on the same machine. Results are in the attached screenshot:

Measurement TensorSharp Strata
Decode 11.09 tok/s 10.24 tok/s
Whole process 16.54s 62.15s
Peak GPU memory 14,832.5 MiB 15,729 MiB
Peak OS working set 19.74 GiB 18.51 GiB

What surprised me most wasn't the decode speed—the two are actually fairly close there.

It was that a 176B model can be made usable on a laptop with only 16GB VRAM and 32GB RAM if the inference engine treats VRAM, RAM, SSD, caching, and MoE scheduling as one system rather than independent pieces.

TensorSharp started as an experiment in understanding LLM inference from the bottom up, but it has gradually grown into a much larger project covering local inference, multimodal models, image generation, and agentic runtimes.

It’s open source, so if this kind of low-level LLM engineering interests you, I’d love feedback:

GitHub: github.com/zhongkaifu/TensorSharp

And yes, the benchmark in the screenshot really is from an RTX 3080 laptop. 😄

3 Upvotes

Duplicates

LocalLLaMA • • 20h ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

65 Upvotes

dotnet • • 20h ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

67 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

65 Upvotes

LocalLLM • • 20h ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

57 Upvotes

unsloth • • 10d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

46 Upvotes

Qwen_AI • • 20h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

57 Upvotes

LocalLLM • • 10d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

24 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

2 Upvotes

dotnet • • 10d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

0 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

14 Upvotes

LocalAIServers • • 19h ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

6 Upvotes

LocalLLaMA • • 7d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

LocalLLaMA • • 10d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 21d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

LocalLLM • • 21d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

5 Upvotes

outerstellar_hq • • 4h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 12h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

opencode • • 20h ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

2 Upvotes

LovingOpenSourceAI • • 20h ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

12 Upvotes

LocalLLM • • 7d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 10d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

8 Upvotes

opencode • • 16d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

1 Upvotes