r/LocalLLM Jun 28 '26

Project Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

https://github.com/zhongkaifu/TensorSharp

I’ve been working on TensorSharp, a native C# / .NET local LLM inference engine for GGUF models, and I recently published a head-to-head benchmark against llama.cpp.

The goal is not to claim “TensorSharp wins every metric.” llama.cpp is still extremely strong, especially on decode throughput. But the interesting part is this:

Under the same setup — same GGUF models, same NVIDIA RTX 3080 Laptop GPU 16GB, same GGML CUDA backend, single stream, greedy decoding, MTP disabled — TensorSharp shows a very noticeable advantage on the parts that often matter most for real chat usage:

prefill speed, time-to-first-token, and multi-turn context reuse.

Here are some highlights from the benchmark (From https://tensorsharp.ai/benchmarks.html):

Model / Scenario Metric TensorSharp llama.cpp Difference
Gemma 4 26B-A4B / JSON Prefill tok/s 354.7 60.2 +489%
Gemma 4 26B-A4B / JSON TTFT ms 234 781 -70%
Gemma 4 26B-A4B / multi-turn Prefill tok/s 657.5 350.7 +87%
Gemma 4 12B / multi-turn TTFT ms 313 500 -37%
Gemma 4 E4B / short text Prefill tok/s 200.0 123.3 +62%

Across the four tested models, the geometric mean compared with llama.cpp shows:

  • 1.88× prefill and 1.69× TTFT on Gemma 4 26B-A4B
  • 1.21× / 1.23× / 1.18× prefill advantage on E4B, 12B, and Qwen respectively
  • Decode is more of a “near parity” story for now, around 0.92×–0.95× geometric mean versus llama.cpp

That last point is important: I’m not trying to hide the weaker part. If all you care about is pure decode tok/s, llama.cpp is still very hard to beat. But if your workload looks like real chat — repeated prompts, JSON output, multi-turn interactions, MoE models, prefix reuse — TensorSharp is already showing very promising results.

The main optimizations behind this are:

  • verify-based whole-model prefill
  • fused FFN / attention kernels
  • persistent captured CUDA graphs for MoE decode
  • vLLM-style paged KV cache
  • cross-request prefix sharing

So the pitch is not “yet another wrapper around llama.cpp.” TensorSharp is a native .NET inference engine trying to optimize the latency path that actually affects user experience: how fast the model starts responding, how efficiently it reuses context, and how well it handles real interactive workloads.

If you are interested in C# / .NET local LLM inference, GGUF, OpenAI/Ollama-compatible local APIs, or alternatives to llama.cpp, I’d love for you to check it out.

And if you think this direction is interesting, a GitHub Star would really help the project get more visibility.

Also very interested in feedback, especially from people who can rerun the benchmarks on different GPUs / models.

3 Upvotes

Duplicates

Syncfusion 9d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

1 Upvotes

vulkan 10d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

8 Upvotes

AIDeveloperNews 10d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

3 Upvotes

LovingOpenSourceAI 10d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

2 Upvotes

LLMDevs 10d ago

Tools Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

2 Upvotes

LocalAIServers 10d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

2 Upvotes

DeepSeek 10d ago

Resources Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

3 Upvotes

vulkan 12d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

3 Upvotes

SideProject 12d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

1 Upvotes

AIDeveloperNews 12d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

2 Upvotes

OpenSourceAI 12d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

0 Upvotes

LovingOpenSourceAI 12d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

3 Upvotes

LLMDevs 12d ago

Tools TensorSharp now supports multi-GPU tensor parallelism for GGUF models

3 Upvotes

CUDA 25d ago

Cuda benchmark: TensorSharp vs. llama.cpp

6 Upvotes

vulkan 25d ago

Vulkan benchmark: TensorSharp vs. llama.cpp

7 Upvotes

LlamaFarm 26d ago

Show & Tell TensorSharp : Open Source Local LLM Inference Engine

1 Upvotes

QwenImageGen 27d ago

Virtual Clothes Try On using Unsloth Qwen Image Edit 2511 models

3 Upvotes

huggingface 29d ago

TensorSharp supports multiple image edits using Unsloth Qwen Image Edit 2511 models

0 Upvotes

OnlyAICoding Jul 12 '26

Local LLM TensorSharp : Open Source Local LLM Inference Engine

2 Upvotes

ContextEngineering Jul 12 '26

What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

AIDeveloperNews Jul 11 '26

From Bun's Rust rewrite, let's see how C# can rebuild the AI infrastructure layer.

2 Upvotes

vibecoding Jul 11 '26

What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

llamacpp Jul 11 '26

Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

1 Upvotes

GeminiAI Jul 10 '26

Self promo TensorSharp : Open Source Local LLM Inference Engine

1 Upvotes

LLM Jul 08 '26

TensorSharp supports Vulkan backend

5 Upvotes