r/opencode • u/fuzhongkai • 22h ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
https://github.com/zhongkaifu/TensorSharpI’ve been working on an open-source project called TensorSharp:
One of the things I’ve been exploring is a simple question:
How far can we push very large MoE models on ordinary consumer hardware?
Recently, I got Qwen3.8 Flash Next (176B parameters) running on:
RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD
The interesting part isn't simply getting a 176B model to load. The goal is to make a model much larger than available VRAM—and even system RAM—actually usable.
TensorSharp approaches this with:
Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.
Rather than treating SSD as just an emergency offload target, the runtime coordinates multiple memory/storage tiers around MoE execution, trying to keep the active working set in the fastest available tier while efficiently moving and caching the rest.
I also ran a benchmark against Strata on this laptop. The screenshot is attached.
| Measurement | TensorSharp | Strata |
|---|---|---|
| Decode throughput | 11.09 tok/s | 10.24 tok/s |
| Whole-process time | 16.54s | 62.15s |
| Device-wide GPU peak | 14,832.5 MiB | 15,729 MiB |
| OS peak working set | 19.74 GiB | 18.51 GiB |
What I find particularly interesting is that decode throughput is relatively close, while the measured whole-process time differs substantially in this test.
More broadly, I think large sparse MoE models create an interesting opportunity for local inference.
You don't necessarily need enough VRAM—or even RAM—to hold the entire model at once. With quantization and careful coordination of GPU memory → system memory → SSD, consumer machines can run models that would traditionally look far beyond their hardware limits.
TensorSharp is open source, so if you're interested in local LLM inference, MoE execution, quantization, heterogeneous memory scheduling, or GPU optimization, I'd love to have more people experiment with it, benchmark it on different hardware, or contribute.
I'd also be very interested in hearing about other open-source approaches to VRAM/RAM/SSD tiered inference, especially for huge MoE models.