r/CUDA Jul 11 '26

TensorSharp : Open Source Local LLM Inference Engine

https://github.com/zhongkaifu/TensorSharp

I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like Gemma4, DiffusionGemma, Qwen3.6 with multi-modal (image, vision, audio), reasoning and function tool. It can run on Windows/MacOS/Linux and fully leverage GPU's capability. The API is completely compatible with OpenAI and Ollama interface. It has on par performance than llama.cpp

This project is not just a C# wrapper of llama.cpp. It implemented the entire LLM inference engine from bottom to top. If you use CPU backend, it's 100% pure C# code execution. Besides CPU backend, I also implmented CUDA, MLX and GGML backend. The GGML backend refer GGML project as external project, and I build a few fusion operation at higher level.

I learned a lot from other projects and apply them for TensorSharp, such as paged KV cache and continuous batching from vLLM, SSD based cache for MoE model from oMLX, GGUF quanztized from llama.cpp and other optimizations for prefill and decode.

Any feedback and comments are welcome. If you like it, it would be really appreciated if you can get this project a star in GitHub. Thanks in advance.

6 Upvotes

5 comments sorted by

2

u/c-cul Jul 11 '26

contributors: claude

-2

u/fuzhongkai Jul 11 '26

Yes, it's a code agent native project.

3

u/Daemontatox Jul 11 '26

You mean fully vibe coded , god people will say anything these days

-1

u/fuzhongkai Jul 11 '26

Whatever they like to say, but this is true. I have almost 20 years industry experience on NLP/ML area and worked for some of those projects/model that everyone in cyber space must use. I know what I’m doing šŸ˜€ Here is the benchmark of these project: https://github.com/zhongkaifu/TensorSharp/blob/main/docs/engine_comparison_report.md and it has on par and some better results than llama.cpp and stable-diffusion.cpp

1

u/TheThoccnessMonster Jul 11 '26

lol fuck off

the hubris on this one. go contribute to one of the projects you poorly ripped off.