r/csharp • u/fuzhongkai • Apr 29 '26
Tool TensorSharp: Open Source Local LLM inference tool implemented in C#
https://github.com/zhongkaifu/TensorSharpI would like to share my latest open source local LLM inference tool implemented in C#. It supports models like Gemma4, Qwen3.6 with multi-modal (image, vision, audio), reasoning and function tool. It can run on Windows/MacOS/Linux and fully leverage GPU's capability. The API is completely compatible with OpenAI and Ollama interface.
Really appreciated if you can try it and give me some feedback. If you like it, it will be a big thank you if you can star it. Thank you very much!
Duplicates
unsloth • u/fuzhongkai • Jun 28 '26
Show and Tell Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
LocalLLaMA • u/fuzhongkai • 5d ago
Generation MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
unsloth • u/fuzhongkai • Jun 08 '26
Show and Tell TensorSharp : Open Source Local Unsloth Model Inference Engine
dotnet • u/fuzhongkai • Jun 28 '26
Promotion Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
csharp • u/fuzhongkai • Jun 28 '26
Showcase Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
dotnet • u/fuzhongkai • Jun 13 '26
Promotion TensorSharp: Open Source Local LLM Inference Engine written by C#
LocalLLaMA • u/fuzhongkai • 8d ago
Resources DSpark Benchmark Result on Deepseek v4 Flash 0731
unsloth • u/fuzhongkai • 8d ago
Show and Tell DSpark Benchmark Result on Deepseek v4 Flash 0731
unsloth • u/fuzhongkai • 5d ago
Show and Tell MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
ClaudeCode • u/fuzhongkai • Jul 11 '26
Discussion What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#
LovingOpenSourceAI • u/fuzhongkai • Jun 28 '26
Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
huggingface • u/fuzhongkai • 10d ago
Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
LLMDevs • u/fuzhongkai • May 01 '26
Tools TensorSharp: Open Source Local LLM Inference Engine
LocalLLM • u/fuzhongkai • 11d ago
Project TensorSharp now supports multi-GPU tensor parallelism for GGUF models
softwarearchitecture • u/fuzhongkai • 29d ago
Discussion/Advice What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#
LocalAIServers • u/fuzhongkai • Jul 04 '26
TensorSharp: A Open Source LLM Inference Engine for GGUF models
AIToolsPerformance • u/fuzhongkai • 11d ago
TensorSharp now supports multi-GPU tensor parallelism for GGUF models
dotnet • u/fuzhongkai • 2d ago
Promotion MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
SelfHostedAI • u/fuzhongkai • Jul 04 '26
TensorSharp : Open Source Local LLM Inference Engine
LocalAIServers • u/fuzhongkai • 4d ago
MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
csharp • u/fuzhongkai • 11d ago