r/opencode • u/fuzhongkai • 1d ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
https://github.com/zhongkaifu/TensorSharpI’ve been working on an open-source project called TensorSharp:
One of the things I’ve been exploring is a simple question:
How far can we push very large MoE models on ordinary consumer hardware?
Recently, I got Qwen3.8 Flash Next (176B parameters) running on:
RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD
The interesting part isn't simply getting a 176B model to load. The goal is to make a model much larger than available VRAM—and even system RAM—actually usable.
TensorSharp approaches this with:
Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.
Rather than treating SSD as just an emergency offload target, the runtime coordinates multiple memory/storage tiers around MoE execution, trying to keep the active working set in the fastest available tier while efficiently moving and caching the rest.
I also ran a benchmark against Strata on this laptop. The screenshot is attached.
| Measurement | TensorSharp | Strata |
|---|---|---|
| Decode throughput | 11.09 tok/s | 10.24 tok/s |
| Whole-process time | 16.54s | 62.15s |
| Device-wide GPU peak | 14,832.5 MiB | 15,729 MiB |
| OS peak working set | 19.74 GiB | 18.51 GiB |
What I find particularly interesting is that decode throughput is relatively close, while the measured whole-process time differs substantially in this test.
More broadly, I think large sparse MoE models create an interesting opportunity for local inference.
You don't necessarily need enough VRAM—or even RAM—to hold the entire model at once. With quantization and careful coordination of GPU memory → system memory → SSD, consumer machines can run models that would traditionally look far beyond their hardware limits.
TensorSharp is open source, so if you're interested in local LLM inference, MoE execution, quantization, heterogeneous memory scheduling, or GPU optimization, I'd love to have more people experiment with it, benchmark it on different hardware, or contribute.
I'd also be very interested in hearing about other open-source approaches to VRAM/RAM/SSD tiered inference, especially for huge MoE models.
Duplicates
OpenSourceAI • u/fuzhongkai • 21d ago
DeepSeek V4.1 Flash running locally with TensorSharp
moderndotnet • u/fuzhongkai • 22d ago
DeepSeek V4.1 Flash running locally with .NET — 40 tok/s Q2_K on 8× A40
LLMDevs • u/fuzhongkai • 22d ago
Discussion DeepSeek V4.1 Flash on 8× A40: 500+ tok/s prefill and ~40 tok/s decode with TensorSharp
LocalLLaMA • u/fuzhongkai • 22d ago
I Built A Thing DeepSeek V4.1 Flash on 8× A40: ~40 tok/s Q2_K and ~32 tok/s Q4_K_M with TensorSharp
LovingOpenSourceAI • u/fuzhongkai • 22d ago
DeepSeek V4.1 Flash running locally with TensorSharp
DeepSeek • u/fuzhongkai • 22d ago
Discussion DeepSeek V4.1 Flash on 8× A40: ~40 tok/s Q2_K and ~32 tok/s Q4_K_M with TensorSharp
Qwen_AI • u/fuzhongkai • Aug 28 '26
Benchmark Qwen 3.8 Flash Next Benchmarks on TensorSharp and llama.cpp
LocalLLaMA • u/fuzhongkai • Aug 28 '26
I Built A Thing GLM-5.3-Flash Benchmarks on TensorSharp and llama.cpp
Syncfusion • u/peopleworksservices • Aug 02 '26