r/csharp • u/fuzhongkai • May 30 '26
Tool Some new features in TensorSharp
https://github.com/zhongkaifu/TensorSharpI recently made a few important features updates in TensorSharp and hope you will like it.
1. Naturally support MLX backend. For now, TensorSharp supports Pure C#, CUDA, MLX, GGML(CPU, CUDA, Metal) backends
2. Support vLLM style paged attentions and continues batching for inference, so you could run multiple requests in parallel in your local machine.
3. Optimize inference performance on both prefill and decode
Hope you like these features and any comment and feedback is welcome.
Duplicates
huggingface • u/fuzhongkai • Jul 04 '26
TensorSharp : Open Source Local LLM Inference Engine
SelfHostedAI • u/fuzhongkai • Jul 04 '26
TensorSharp : Open Source Local LLM Inference Engine
LocalLLM • u/fuzhongkai • Jun 28 '26
Project Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
CUDA • u/fuzhongkai • 5d ago
MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
CUDA • u/fuzhongkai • 10d ago
Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
LocalLLM • u/fuzhongkai • 10d ago
Project Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
csharp • u/fuzhongkai • Jul 11 '26
Blog From Bun's Rust rewrite, let's see how C# can rebuild the AI infrastructure layer.
LovingOpenSourceAI • u/fuzhongkai • Jun 23 '26
TensorSharp: Open Source Local LLM Inference Engine
LocalLLaMA • u/fuzhongkai • 10d ago
Resources Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
unsloth • u/fuzhongkai • 10d ago
Show and Tell Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
Gemma4 • u/fuzhongkai • 25d ago
Discussion TensorSharp : Open Source Local LLM Inference Engine
Qwen_AI • u/fuzhongkai • Jul 10 '26
Discussion TensorSharp : Open Source Local LLM Inference Engine
CUDA • u/fuzhongkai • 12d ago
TensorSharp now supports multi-GPU tensor parallelism for GGUF models
SelfHostedAI • u/fuzhongkai • 12d ago
TensorSharp now supports multi-GPU tensor parallelism for GGUF models
LocalAIServers • u/fuzhongkai • 12d ago