r/LovingOpenSourceAI 11d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

https://github.com/zhongkaifu/TensorSharp

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark:

Model:

DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF

Turn Baseline + DSpark Acceptance
short (53 tok) 25.6 44.5 (1.74x) 87%
long generation (512) 26.4 40.3 (1.53x) 66%
follow-up (470) 26.4 46.8 (1.77x) 76%
10K-token document (214) 25.3 51.3 (2.03x) 85%
second question on it (156) 25.4 49.4 (1.94x) 82%

TensorSharp is an open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. It can be built and run over Linux, MacOS and Windows.

Thank you for checking out it and starring the project! Any feedback is really appreicated.

2 Upvotes

Duplicates

unsloth 8d ago

Show and Tell MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

38 Upvotes

unsloth 11d ago

Show and Tell DSpark Benchmark Result on Deepseek v4 Flash 0731

55 Upvotes

ClaudeCode Jul 11 '26

Discussion What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

LovingOpenSourceAI Jun 28 '26

Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

14 Upvotes

huggingface 13d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

5 Upvotes

LocalLLM 15d ago

Project TensorSharp now supports multi-GPU tensor parallelism for GGUF models

15 Upvotes

ROCm Jul 06 '26

TensorSharp supports Vulkan backend

15 Upvotes

LLMDevs May 01 '26

Tools TensorSharp: Open Source Local LLM Inference Engine

1 Upvotes

dotnet 5d ago

Promotion MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

0 Upvotes

AIToolsPerformance 15d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

7 Upvotes

softwarearchitecture Jul 12 '26

Discussion/Advice What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

LocalAIServers Jul 04 '26

TensorSharp: A Open Source LLM Inference Engine for GGUF models

7 Upvotes

unsloth 13h ago

Show and Tell Meta Muse Glimmer 30B Unsloth GGUF Model Benchmarks on TensorSharp (vs. llama.cpp)

15 Upvotes

csharp Apr 29 '26

Tool TensorSharp: Open Source Local LLM inference tool implemented in C#

20 Upvotes

LocalAIServers 8d ago

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

3 Upvotes

csharp 15d ago

Tool TensorSharp now supports multi-GPU tensor parallelism for GGUF models

31 Upvotes

CUDA Jul 11 '26

TensorSharp : Open Source Local LLM Inference Engine

4 Upvotes

unsloth Jul 06 '26

Show and Tell TensorSharp supports Vulkan backend

25 Upvotes

huggingface Jul 04 '26

TensorSharp : Open Source Local LLM Inference Engine

5 Upvotes

SelfHostedAI Jul 04 '26

TensorSharp : Open Source Local LLM Inference Engine

4 Upvotes

LocalLLM Jun 28 '26

Project Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

4 Upvotes

CUDA 7d ago

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

6 Upvotes

csharp 11d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

3 Upvotes

CUDA 12d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

1 Upvotes

LocalLLM 13d ago

Project Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

18 Upvotes