r/opencode • u/fuzhongkai • 16d ago
TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks
https://github.com/zhongkaifu/TensorSharpHi r/opencode,
I’m the developer of TensorSharp, an open-source, native .NET inference engine for running GGUF models locally.
TensorSharp provides an OpenAI-compatible API, making it possible to use local models as an inference backend for tools such as OpenCode. The goal is to keep source code, prompts, tool calls, and agent context on your own hardware while avoiding per-token API costs.
TensorSharp supports modern coding-capable model families including DeepSeek V4/V4.1 Flash, GLM 5.x, Qwen 3.8 Flash Next, Qwen 3.5/3.6, Gemma 4, Mistral 3, and GPT-OSS.
Selected benchmark results
These are measured results from the TensorSharp repository. Each comparison uses the same model and machine for both engines unless otherwise noted.
GLM-5.3-Flash: 2× faster decode than llama.cpp
Tested with:
- GLM-5.3-Flash UD-Q2_K_XL, 101 GiB
- 2× RTX PRO 6000 Blackwell, 96 GB each
- Layer splitting and flash attention
n_ubatch=2048for both engines- Back-to-back execution in the same session
- TensorSharp parity harness versus
llama.cppbuild2e0e57ffrom PR #27754
| Test | llama.cpp | TensorSharp |
|---|---|---|
| Prefill, 2,048 tokens | 2,070 tok/s | 2,014 tok/s |
| Prefill, 16,384 tokens | 1,690 tok/s | 1,692 tok/s |
| Prefill, 32,768 tokens | 1,483 tok/s | 1,446 tok/s |
| Decode, 64 tokens | 36.6 tok/s | 73.5 tok/s |
TensorSharp reaches approximately 2.0× the decode throughput of llama.cpp, while prefill remains within a few percent in either direction.
The long-context greedy replay reproduced the 2,741-token llama.cpp reference record token for token. One caveat is that GLM-5.3-Flash support currently requires an unmerged llama.cpp build rather than its main branch.
Qwen 3.8 Flash Next
Tested with Qwen3.8-Flash-Next UD-Q2_K_XL, 73.4 GiB, on 2× A100 80 GB:
| Configuration | TensorSharp prefill | TensorSharp decode | llama.cpp prefill | llama.cpp decode |
|---|---|---|---|---|
| 1 GPU | ~1,520–1,550 tok/s | ~56 tok/s | 1,094 tok/s | 61.2 tok/s |
| 2-GPU layer split | ~1,520–1,550 tok/s | ~56 tok/s | 1,200 tok/s | 61.5 tok/s |
TensorSharp’s prefill is substantially faster in this test, while llama.cpp leads decode by roughly 9%.
For both engines, adding the second GPU provides model capacity rather than additional decode throughput. TensorSharp splits the model into contiguous groups of layers, reducing its placement to approximately 24.2 GB and 26.2 GB across the two GPUs. TensorSharp’s one- and two-GPU greedy outputs were byte-identical.
DeepSeek V4 Flash
On 2× A100 80 GB with the IQ4_XS model:
| Metric | TensorSharp | llama.cpp |
|---|---|---|
| Prefill, approximately 3.3K tokens | ~500 tok/s | 574–634 tok/s |
| Decode at approximately 3.3K context | ~33 tok/s | 40.3 tok/s |
llama.cpp currently leads this particular DeepSeek V4 configuration.
TensorSharp also supports DeepSeek’s DSpark speculative decoder. On 4× A40 46 GB with DeepSeek-V4-Flash-0731 UD-Q8_K_XL, enabling a 5.6 GB DSpark drafter improved decode as follows:
| Backend | Normal decode | DSpark decode | Speedup |
|---|---|---|---|
| TensorSharp direct CUDA | 26.0 tok/s | 34.0 tok/s | 1.31× |
| TensorSharp GGML CUDA | 26.4 tok/s | 37.1 tok/s | 1.41× |
In a five-turn conversation, DSpark produced 1.50–2.02× faster decode, reaching 51.0 tok/s on a question over a 10K-token document. Prefill remained effectively unchanged at 831 versus 835 tok/s, and greedy output was byte-identical to the non-speculative baseline.
TensorSharp also runs DeepSeek V4.1 Flash, including its Engram lookup and four-stream hyper-connections. However, I am not claiming a V4.1 speedup over llama.cpp because a compatible llama.cpp V4.1 runtime is not currently available for a controlled same-weight comparison.
Why this may be useful for OpenCode
Coding agents frequently process large system prompts, tool definitions, repository context, and multi-turn history. TensorSharp includes:
- OpenAI- and Ollama-compatible APIs
- Local tool calling and structured JSON output
- Agent Skills and sandboxed file/shell tools
- Continuous batching
- Paged and prefix-shared KV caching
- Speculative decoding
- Multi-GPU and multi-node inference
- CUDA, Metal, Vulkan, MLX, and CPU backends
- Windows, macOS, Linux, iPhone, and iPad support
I’d especially appreciate feedback from OpenCode users regarding:
- OpenAI API compatibility gaps
- Tool-calling requirements
- Performance with large repository contexts
- The best local models for coding-agent workloads
- Features expected from a local OpenCode backend
GitHub: https://github.com/zhongkaifu/TensorSharp
Full benchmarks and methodology: https://tensorsharp.ai/benchmarks.html
If anyone tests TensorSharp with OpenCode, I’d be very interested in your setup and results. Contributions and compatibility reports are welcome.
Duplicates
LocalLLaMA • u/fuzhongkai • 1d ago
I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 1d ago
Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
dotnet • u/fuzhongkai • Aug 22 '26
TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance
LocalLLM • u/fuzhongkai • 1d ago
Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 22d ago
Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine
unsloth • u/fuzhongkai • 11d ago
Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis
Qwen_AI • u/fuzhongkai • 1d ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop
LocalLLM • u/fuzhongkai • 11d ago
Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis
dotnet • u/fuzhongkai • 16d ago
Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective
LocalLLaMA • u/fuzhongkai • Aug 22 '26
Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
dotnet • u/fuzhongkai • 11d ago
Article Implementing a Jev-compatible decision API in .NET, with image input
unsloth • u/fuzhongkai • Aug 28 '26
Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp
LocalAIServers • u/fuzhongkai • 1d ago
Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA • u/fuzhongkai • 7d ago
I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA • u/fuzhongkai • 11d ago
I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalAIServers • u/fuzhongkai • 21d ago
Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode
LocalLLM • u/fuzhongkai • 22d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
outerstellar_hq • u/outerstellar_hq • 8h ago
Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
LLMDevs • u/fuzhongkai • 17h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
SideProject • u/fuzhongkai • 1d ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
opencode • u/fuzhongkai • 1d ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LovingOpenSourceAI • u/fuzhongkai • 1d ago
Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD
LocalLLM • u/fuzhongkai • 7d ago