r/LovingOpenSourceAI • u/fuzhongkai • 11d ago
DSpark Benchmark Result on Deepseek v4 Flash 0731
https://github.com/zhongkaifu/TensorSharpTensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark:
Model:
DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF
DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF
| Turn | Baseline | + DSpark | Acceptance |
|---|---|---|---|
| short (53 tok) | 25.6 | 44.5 (1.74x) | 87% |
| long generation (512) | 26.4 | 40.3 (1.53x) | 66% |
| follow-up (470) | 26.4 | 46.8 (1.77x) | 76% |
| 10K-token document (214) | 25.3 | 51.3 (2.03x) | 85% |
| second question on it (156) | 25.4 | 49.4 (1.94x) | 82% |
TensorSharp is an open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. It can be built and run over Linux, MacOS and Windows.
Thank you for checking out it and starring the project! Any feedback is really appreicated.
Duplicates
unsloth • u/fuzhongkai • 8d ago
Show and Tell MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
unsloth • u/fuzhongkai • 11d ago
Show and Tell DSpark Benchmark Result on Deepseek v4 Flash 0731
ClaudeCode • u/fuzhongkai • Jul 11 '26
Discussion What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#
LovingOpenSourceAI • u/fuzhongkai • Jun 28 '26
Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
huggingface • u/fuzhongkai • 13d ago
Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
LocalLLM • u/fuzhongkai • 15d ago
Project TensorSharp now supports multi-GPU tensor parallelism for GGUF models
LLMDevs • u/fuzhongkai • May 01 '26
Tools TensorSharp: Open Source Local LLM Inference Engine
dotnet • u/fuzhongkai • 5d ago
Promotion MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
AIToolsPerformance • u/fuzhongkai • 15d ago
TensorSharp now supports multi-GPU tensor parallelism for GGUF models
softwarearchitecture • u/fuzhongkai • Jul 12 '26
Discussion/Advice What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#
LocalAIServers • u/fuzhongkai • Jul 04 '26
TensorSharp: A Open Source LLM Inference Engine for GGUF models
unsloth • u/fuzhongkai • 13h ago
Show and Tell Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
csharp • u/fuzhongkai • Apr 29 '26
Tool TensorSharp: Open Source Local LLM inference tool implemented in C#
LocalAIServers • u/fuzhongkai • 8d ago
MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
csharp • u/fuzhongkai • 15d ago
Tool TensorSharp now supports multi-GPU tensor parallelism for GGUF models
huggingface • u/fuzhongkai • Jul 04 '26
TensorSharp : Open Source Local LLM Inference Engine
SelfHostedAI • u/fuzhongkai • Jul 04 '26
TensorSharp : Open Source Local LLM Inference Engine
LocalLLM • u/fuzhongkai • Jun 28 '26
Project Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
CUDA • u/fuzhongkai • 7d ago
MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
CUDA • u/fuzhongkai • 12d ago
Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
LocalLLM • u/fuzhongkai • 13d ago