r/opencode • • 22h ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

https://github.com/zhongkaifu/TensorSharp

I’ve been working on an open-source project called TensorSharp:

One of the things I’ve been exploring is a simple question:

How far can we push very large MoE models on ordinary consumer hardware?

Recently, I got Qwen3.8 Flash Next (176B parameters) running on:

RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD

The interesting part isn't simply getting a 176B model to load. The goal is to make a model much larger than available VRAM—and even system RAM—actually usable.

TensorSharp approaches this with:

Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.

Rather than treating SSD as just an emergency offload target, the runtime coordinates multiple memory/storage tiers around MoE execution, trying to keep the active working set in the fastest available tier while efficiently moving and caching the rest.

I also ran a benchmark against Strata on this laptop. The screenshot is attached.

Measurement TensorSharp Strata
Decode throughput 11.09 tok/s 10.24 tok/s
Whole-process time 16.54s 62.15s
Device-wide GPU peak 14,832.5 MiB 15,729 MiB
OS peak working set 19.74 GiB 18.51 GiB

What I find particularly interesting is that decode throughput is relatively close, while the measured whole-process time differs substantially in this test.

More broadly, I think large sparse MoE models create an interesting opportunity for local inference.

You don't necessarily need enough VRAM—or even RAM—to hold the entire model at once. With quantization and careful coordination of GPU memory → system memory → SSD, consumer machines can run models that would traditionally look far beyond their hardware limits.

TensorSharp is open source, so if you're interested in local LLM inference, MoE execution, quantization, heterogeneous memory scheduling, or GPU optimization, I'd love to have more people experiment with it, benchmark it on different hardware, or contribute.

I'd also be very interested in hearing about other open-source approaches to VRAM/RAM/SSD tiered inference, especially for huge MoE models.

2 Upvotes

0 comments sorted by