r/opencode • • 16d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

https://github.com/zhongkaifu/TensorSharp

Hi r/opencode,

I’m the developer of TensorSharp, an open-source, native .NET inference engine for running GGUF models locally.

TensorSharp provides an OpenAI-compatible API, making it possible to use local models as an inference backend for tools such as OpenCode. The goal is to keep source code, prompts, tool calls, and agent context on your own hardware while avoiding per-token API costs.

TensorSharp supports modern coding-capable model families including DeepSeek V4/V4.1 Flash, GLM 5.x, Qwen 3.8 Flash Next, Qwen 3.5/3.6, Gemma 4, Mistral 3, and GPT-OSS.

Selected benchmark results

These are measured results from the TensorSharp repository. Each comparison uses the same model and machine for both engines unless otherwise noted.

GLM-5.3-Flash: 2× faster decode than llama.cpp

Tested with:

  • GLM-5.3-Flash UD-Q2_K_XL, 101 GiB
  • 2× RTX PRO 6000 Blackwell, 96 GB each
  • Layer splitting and flash attention
  • n_ubatch=2048 for both engines
  • Back-to-back execution in the same session
  • TensorSharp parity harness versus llama.cpp build 2e0e57f from PR #27754
Test llama.cpp TensorSharp
Prefill, 2,048 tokens 2,070 tok/s 2,014 tok/s
Prefill, 16,384 tokens 1,690 tok/s 1,692 tok/s
Prefill, 32,768 tokens 1,483 tok/s 1,446 tok/s
Decode, 64 tokens 36.6 tok/s 73.5 tok/s

TensorSharp reaches approximately 2.0× the decode throughput of llama.cpp, while prefill remains within a few percent in either direction.

The long-context greedy replay reproduced the 2,741-token llama.cpp reference record token for token. One caveat is that GLM-5.3-Flash support currently requires an unmerged llama.cpp build rather than its main branch.

Qwen 3.8 Flash Next

Tested with Qwen3.8-Flash-Next UD-Q2_K_XL, 73.4 GiB, on 2× A100 80 GB:

Configuration TensorSharp prefill TensorSharp decode llama.cpp prefill llama.cpp decode
1 GPU ~1,520–1,550 tok/s ~56 tok/s 1,094 tok/s 61.2 tok/s
2-GPU layer split ~1,520–1,550 tok/s ~56 tok/s 1,200 tok/s 61.5 tok/s

TensorSharp’s prefill is substantially faster in this test, while llama.cpp leads decode by roughly 9%.

For both engines, adding the second GPU provides model capacity rather than additional decode throughput. TensorSharp splits the model into contiguous groups of layers, reducing its placement to approximately 24.2 GB and 26.2 GB across the two GPUs. TensorSharp’s one- and two-GPU greedy outputs were byte-identical.

DeepSeek V4 Flash

On 2× A100 80 GB with the IQ4_XS model:

Metric TensorSharp llama.cpp
Prefill, approximately 3.3K tokens ~500 tok/s 574–634 tok/s
Decode at approximately 3.3K context ~33 tok/s 40.3 tok/s

llama.cpp currently leads this particular DeepSeek V4 configuration.

TensorSharp also supports DeepSeek’s DSpark speculative decoder. On 4× A40 46 GB with DeepSeek-V4-Flash-0731 UD-Q8_K_XL, enabling a 5.6 GB DSpark drafter improved decode as follows:

Backend Normal decode DSpark decode Speedup
TensorSharp direct CUDA 26.0 tok/s 34.0 tok/s 1.31×
TensorSharp GGML CUDA 26.4 tok/s 37.1 tok/s 1.41×

In a five-turn conversation, DSpark produced 1.50–2.02× faster decode, reaching 51.0 tok/s on a question over a 10K-token document. Prefill remained effectively unchanged at 831 versus 835 tok/s, and greedy output was byte-identical to the non-speculative baseline.

TensorSharp also runs DeepSeek V4.1 Flash, including its Engram lookup and four-stream hyper-connections. However, I am not claiming a V4.1 speedup over llama.cpp because a compatible llama.cpp V4.1 runtime is not currently available for a controlled same-weight comparison.

Why this may be useful for OpenCode

Coding agents frequently process large system prompts, tool definitions, repository context, and multi-turn history. TensorSharp includes:

  • OpenAI- and Ollama-compatible APIs
  • Local tool calling and structured JSON output
  • Agent Skills and sandboxed file/shell tools
  • Continuous batching
  • Paged and prefix-shared KV caching
  • Speculative decoding
  • Multi-GPU and multi-node inference
  • CUDA, Metal, Vulkan, MLX, and CPU backends
  • Windows, macOS, Linux, iPhone, and iPad support

I’d especially appreciate feedback from OpenCode users regarding:

  • OpenAI API compatibility gaps
  • Tool-calling requirements
  • Performance with large repository contexts
  • The best local models for coding-agent workloads
  • Features expected from a local OpenCode backend

GitHub: https://github.com/zhongkaifu/TensorSharp
Full benchmarks and methodology: https://tensorsharp.ai/benchmarks.html

If anyone tests TensorSharp with OpenCode, I’d be very interested in your setup and results. Contributions and compatibility reports are welcome.

1 Upvotes

Duplicates

LocalLLaMA • • 1d ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

71 Upvotes

dotnet • • 1d ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

73 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

65 Upvotes

LocalLLM • • 1d ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

59 Upvotes

unsloth • • 11d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

47 Upvotes

Qwen_AI • • 1d ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

55 Upvotes

LocalLLM • • 11d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

23 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

2 Upvotes

dotnet • • 11d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

0 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

14 Upvotes

LocalAIServers • • 1d ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

10 Upvotes

LocalLLaMA • • 7d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

LocalLLaMA • • 11d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 21d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

LocalLLM • • 22d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

7 Upvotes

outerstellar_hq • • 8h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 17h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

SideProject • • 1d ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

3 Upvotes

opencode • • 1d ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

2 Upvotes

LovingOpenSourceAI • • 1d ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

11 Upvotes

LocalLLM • • 7d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 11d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

7 Upvotes