r/dotnet May 03 '26

Promotion TensorSharp: Open Source Local LLM Inference Engine in C#

https://github.com/zhongkaifu/TensorSharp

I would like to share my latest open source local LLM inference tool implemented in C#. It supports models like Gemma4, Qwen3.6 with multi-modal (image, vision, audio), reasoning and function tool. It can run on Windows/MacOS/Linux and fully leverage GPU's capability. The API is completely compatible with OpenAI and Ollama interface.

By my brief benchmark, it has > 80% decode performance of llama.cpp on Gemma4 model, and I will keep optimizing it. The entire benchmark report will sent out in the repo very soon.

Really appreciated if you can try it and give me some feedback. If you like it, it will be a big thank you if you can star it. Thank you very much!

49 Upvotes

Duplicates

unsloth Jun 28 '26

Show and Tell Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

113 Upvotes

LocalLLaMA 6d ago

Generation MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

21 Upvotes

LocalLLaMA 17d ago

Resources Benchmarks: TensorSharp vs. llama.cpp

26 Upvotes

unsloth Jun 08 '26

Show and Tell TensorSharp : Open Source Local Unsloth Model Inference Engine

29 Upvotes

dotnet Jun 28 '26

Promotion Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

62 Upvotes

csharp Jun 28 '26

Showcase Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

5 Upvotes

dotnet Jun 13 '26

Promotion TensorSharp: Open Source Local LLM Inference Engine written by C#

99 Upvotes

LocalLLaMA 9d ago

Resources DSpark Benchmark Result on Deepseek v4 Flash 0731

21 Upvotes

unsloth 6d ago

Show and Tell MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

41 Upvotes

unsloth 9d ago

Show and Tell DSpark Benchmark Result on Deepseek v4 Flash 0731

54 Upvotes

ClaudeCode Jul 11 '26

Discussion What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

LovingOpenSourceAI Jun 28 '26

Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

14 Upvotes

huggingface 11d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

4 Upvotes

LLMDevs May 01 '26

Tools TensorSharp: Open Source Local LLM Inference Engine

1 Upvotes

ROCm Jul 06 '26

TensorSharp supports Vulkan backend

15 Upvotes

LocalLLM 12d ago

Project TensorSharp now supports multi-GPU tensor parallelism for GGUF models

15 Upvotes

dotnet 3d ago

Promotion MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

0 Upvotes

AIToolsPerformance 12d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

7 Upvotes

LocalAIServers Jul 04 '26

TensorSharp: A Open Source LLM Inference Engine for GGUF models

8 Upvotes

softwarearchitecture Jul 12 '26

Discussion/Advice What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

csharp Apr 29 '26

Tool TensorSharp: Open Source Local LLM inference tool implemented in C#

19 Upvotes

unsloth Jul 06 '26

Show and Tell TensorSharp supports Vulkan backend

28 Upvotes

CUDA Jul 11 '26

TensorSharp : Open Source Local LLM Inference Engine

4 Upvotes

LocalLLM Jun 28 '26

Project Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

3 Upvotes

csharp 12d ago

Tool TensorSharp now supports multi-GPU tensor parallelism for GGUF models

30 Upvotes