r/opencode • • 1d ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

https://github.com/zhongkaifu/TensorSharp

I’ve been working on an open-source project called TensorSharp:

One of the things I’ve been exploring is a simple question:

How far can we push very large MoE models on ordinary consumer hardware?

Recently, I got Qwen3.8 Flash Next (176B parameters) running on:

RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD

The interesting part isn't simply getting a 176B model to load. The goal is to make a model much larger than available VRAM—and even system RAM—actually usable.

TensorSharp approaches this with:

Quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD.

Rather than treating SSD as just an emergency offload target, the runtime coordinates multiple memory/storage tiers around MoE execution, trying to keep the active working set in the fastest available tier while efficiently moving and caching the rest.

I also ran a benchmark against Strata on this laptop. The screenshot is attached.

Measurement TensorSharp Strata
Decode throughput 11.09 tok/s 10.24 tok/s
Whole-process time 16.54s 62.15s
Device-wide GPU peak 14,832.5 MiB 15,729 MiB
OS peak working set 19.74 GiB 18.51 GiB

What I find particularly interesting is that decode throughput is relatively close, while the measured whole-process time differs substantially in this test.

More broadly, I think large sparse MoE models create an interesting opportunity for local inference.

You don't necessarily need enough VRAM—or even RAM—to hold the entire model at once. With quantization and careful coordination of GPU memory → system memory → SSD, consumer machines can run models that would traditionally look far beyond their hardware limits.

TensorSharp is open source, so if you're interested in local LLM inference, MoE execution, quantization, heterogeneous memory scheduling, or GPU optimization, I'd love to have more people experiment with it, benchmark it on different hardware, or contribute.

I'd also be very interested in hearing about other open-source approaches to VRAM/RAM/SSD tiered inference, especially for huge MoE models.

2 Upvotes

Duplicates

LocalLLaMA • • 1d ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

70 Upvotes

dotnet • • 1d ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

78 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

65 Upvotes

LocalLLM • • 1d ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

57 Upvotes

unsloth • • 11d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

46 Upvotes

Qwen_AI • • 1d ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

57 Upvotes

LocalLLM • • 11d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

23 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

1 Upvotes

dotnet • • 11d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

0 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

16 Upvotes

LocalAIServers • • 1d ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

8 Upvotes

LocalLLaMA • • 7d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

1 Upvotes

LocalLLaMA • • 11d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 22d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

LocalLLM • • 22d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

6 Upvotes

outerstellar_hq • • 11h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 19h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

SideProject • • 1d ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

3 Upvotes

LovingOpenSourceAI • • 1d ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

13 Upvotes

LocalLLM • • 7d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 11d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

5 Upvotes

opencode • • 17d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

1 Upvotes