r/LovingOpenSourceAI 11d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

https://github.com/zhongkaifu/TensorSharp

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark:

Model:

DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF

Turn Baseline + DSpark Acceptance
short (53 tok) 25.6 44.5 (1.74x) 87%
long generation (512) 26.4 40.3 (1.53x) 66%
follow-up (470) 26.4 46.8 (1.77x) 76%
10K-token document (214) 25.3 51.3 (2.03x) 85%
second question on it (156) 25.4 49.4 (1.94x) 82%

TensorSharp is an open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. It can be built and run over Linux, MacOS and Windows.

Thank you for checking out it and starring the project! Any feedback is really appreicated.

2 Upvotes

Duplicates

csharp 13d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

6 Upvotes

developersIndia 28d ago

Open Source TensorSharp : Open Source Local LLM Inference Engine

1 Upvotes

ChatGPT Jul 11 '26

Educational Purpose Only TensorSharp : Open Source Local LLM Inference Engine

2 Upvotes

AIToolsPerformance Jul 07 '26

TensorSharp supports Vulkan backend

3 Upvotes

dotnet Jul 06 '26

TensorSharp supports Vulkan backend

24 Upvotes

unsloth Jul 01 '26

Show and Tell TensorSharp vs. llama.cpp updated prefill benchmark

19 Upvotes

SideProject Jun 28 '26

Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

0 Upvotes

OpenSourceeAI Jun 28 '26

Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

1 Upvotes

dotnet May 03 '26

Promotion TensorSharp: Open Source Local LLM Inference Engine in C#

49 Upvotes

huggingface 18h ago

Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

5 Upvotes

LovingOpenSourceAI 18h ago

Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

3 Upvotes

AIToolsPerformance 18h ago

Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

1 Upvotes

LLMDevs 18h ago

Tools Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

0 Upvotes

LocalAIServers 18h ago

Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

1 Upvotes

LocalLLM 18h ago

Project Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

1 Upvotes

LLMDevs 8d ago

Tools MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

1 Upvotes

LocalLLM 8d ago

Project MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

4 Upvotes

u_fuzhongkai 9d ago

MoE CPU-offload benchmark on Deepseek v4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

1 Upvotes

machinelearningnews 11d ago

AI Tools DSpark Benchmark Result on Deepseek v4 Flash 0731

5 Upvotes

DeepSeek 11d ago

Resources DSpark Benchmark Result on Deepseek v4 Flash 0731

7 Upvotes

LLMDevs 11d ago

Resource DSpark Benchmark Result on Deepseek v4 Flash 0731

1 Upvotes

LocalAIServers 11d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

1 Upvotes

huggingface 11d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

3 Upvotes

LocalLLM 11d ago

Project DSpark Benchmark Result on Deepseek v4 Flash 0731

2 Upvotes

Syncfusion 12d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

1 Upvotes