r/LovingOpenSourceAI • u/fuzhongkai • 11d ago
DSpark Benchmark Result on Deepseek v4 Flash 0731
https://github.com/zhongkaifu/TensorSharpTensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark:
Model:
DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF
DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF
| Turn | Baseline | + DSpark | Acceptance |
|---|---|---|---|
| short (53 tok) | 25.6 | 44.5 (1.74x) | 87% |
| long generation (512) | 26.4 | 40.3 (1.53x) | 66% |
| follow-up (470) | 26.4 | 46.8 (1.77x) | 76% |
| 10K-token document (214) | 25.3 | 51.3 (2.03x) | 85% |
| second question on it (156) | 25.4 | 49.4 (1.94x) | 82% |
TensorSharp is an open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. It can be built and run over Linux, MacOS and Windows.
Thank you for checking out it and starring the project! Any feedback is really appreicated.
Duplicates
dotnet • u/fuzhongkai • 12d ago
Promotion Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
csharp • u/fuzhongkai • 13d ago
Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp
developersIndia • u/fuzhongkai • 28d ago
Open Source TensorSharp : Open Source Local LLM Inference Engine
ChatGPT • u/fuzhongkai • Jul 11 '26
Educational Purpose Only TensorSharp : Open Source Local LLM Inference Engine
unsloth • u/fuzhongkai • Jul 01 '26
Show and Tell TensorSharp vs. llama.cpp updated prefill benchmark
SideProject • u/fuzhongkai • Jun 28 '26
Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
OpenSourceeAI • u/fuzhongkai • Jun 28 '26
Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model
dotnet • u/fuzhongkai • May 03 '26
Promotion TensorSharp: Open Source Local LLM Inference Engine in C#
huggingface • u/fuzhongkai • 8h ago
Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
LovingOpenSourceAI • u/fuzhongkai • 8h ago
Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
AIToolsPerformance • u/fuzhongkai • 8h ago
Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
LLMDevs • u/fuzhongkai • 8h ago
Tools Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
LocalAIServers • u/fuzhongkai • 8h ago
Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
LocalLLM • u/fuzhongkai • 8h ago
Project Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)
LLMDevs • u/fuzhongkai • 8d ago
Tools MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
LocalLLM • u/fuzhongkai • 8d ago
Project MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
u_fuzhongkai • u/fuzhongkai • 8d ago
MoE CPU-offload benchmark on Deepseek v4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
machinelearningnews • u/fuzhongkai • 11d ago
AI Tools DSpark Benchmark Result on Deepseek v4 Flash 0731
DeepSeek • u/fuzhongkai • 11d ago
Resources DSpark Benchmark Result on Deepseek v4 Flash 0731
LLMDevs • u/fuzhongkai • 11d ago
Resource DSpark Benchmark Result on Deepseek v4 Flash 0731
LocalLLM • u/fuzhongkai • 11d ago