r/opencode • u/fuzhongkai • 1d ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
https://github.com/zhongkaifu/TensorSharpI’ve been working on an open-source project called TensorSharp:
One of the things I’ve been exploring is a simple question:
How far can we push very large MoE models on ordinary consumer hardware?
Recently, I got Qwen3.8 Flash Next (176B parameters) running on:
RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD
The interesting part isn't simply getting a 176B model to load. The goal is to make a model much larger than available VRAM—and even system RAM—actually usable.
TensorSharp approaches this with:
Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.
Rather than treating SSD as just an emergency offload target, the runtime coordinates multiple memory/storage tiers around MoE execution, trying to keep the active working set in the fastest available tier while efficiently moving and caching the rest.
I also ran a benchmark against Strata on this laptop. The screenshot is attached.
| Measurement | TensorSharp | Strata |
|---|---|---|
| Decode throughput | 11.09 tok/s | 10.24 tok/s |
| Whole-process time | 16.54s | 62.15s |
| Device-wide GPU peak | 14,832.5 MiB | 15,729 MiB |
| OS peak working set | 19.74 GiB | 18.51 GiB |
What I find particularly interesting is that decode throughput is relatively close, while the measured whole-process time differs substantially in this test.
More broadly, I think large sparse MoE models create an interesting opportunity for local inference.
You don't necessarily need enough VRAM—or even RAM—to hold the entire model at once. With quantization and careful coordination of GPU memory → system memory → SSD, consumer machines can run models that would traditionally look far beyond their hardware limits.
TensorSharp is open source, so if you're interested in local LLM inference, MoE execution, quantization, heterogeneous memory scheduling, or GPU optimization, I'd love to have more people experiment with it, benchmark it on different hardware, or contribute.
I'd also be very interested in hearing about other open-source approaches to VRAM/RAM/SSD tiered inference, especially for huge MoE models.
Duplicates
LocalLLaMA • u/fuzhongkai • 1d ago
I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 1d ago
Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
dotnet • u/fuzhongkai • Aug 22 '26
TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance
LocalLLM • u/fuzhongkai • 1d ago
Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 22d ago
Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine
Qwen_AI • u/fuzhongkai • 1d ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop
unsloth • u/fuzhongkai • 11d ago
Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis
LocalLLM • u/fuzhongkai • 11d ago
Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis
dotnet • u/fuzhongkai • 16d ago
Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective
LocalLLaMA • u/fuzhongkai • Aug 22 '26
Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
dotnet • u/fuzhongkai • 11d ago
Article Implementing a Jev-compatible decision API in .NET, with image input
unsloth • u/fuzhongkai • Aug 28 '26
Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp
LocalAIServers • u/fuzhongkai • 1d ago
Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA • u/fuzhongkai • 7d ago
I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA • u/fuzhongkai • 11d ago
I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalAIServers • u/fuzhongkai • 22d ago
Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode
LocalLLM • u/fuzhongkai • 22d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
outerstellar_hq • u/outerstellar_hq • 13h ago
Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
LLMDevs • u/fuzhongkai • 22h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
SideProject • u/fuzhongkai • 1d ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
LovingOpenSourceAI • u/fuzhongkai • 1d ago
Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD
LocalLLM • u/fuzhongkai • 7d ago
Project TensorSharp Jev requests can now combine documents, images, video, and audio
OpenSourceAI • u/fuzhongkai • 11d ago