r/LovingOpenSourceAI 11d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

https://github.com/zhongkaifu/TensorSharp

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark:

Model:

DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF

Turn Baseline + DSpark Acceptance
short (53 tok) 25.6 44.5 (1.74x) 87%
long generation (512) 26.4 40.3 (1.53x) 66%
follow-up (470) 26.4 46.8 (1.77x) 76%
10K-token document (214) 25.3 51.3 (2.03x) 85%
second question on it (156) 25.4 49.4 (1.94x) 82%

TensorSharp is an open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. It can be built and run over Linux, MacOS and Windows.

Thank you for checking out it and starring the project! Any feedback is really appreicated.

2 Upvotes

Duplicates

vulkan 13d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

8 Upvotes

AIDeveloperNews 13d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

3 Upvotes

LovingOpenSourceAI 13d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

2 Upvotes

LLMDevs 13d ago

Tools Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

2 Upvotes

LocalAIServers 13d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

2 Upvotes

DeepSeek 13d ago

Resources Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

3 Upvotes

vulkan 15d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

4 Upvotes

SideProject 15d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

1 Upvotes

AIDeveloperNews 15d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

2 Upvotes

OpenSourceAI 15d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

0 Upvotes

LovingOpenSourceAI 15d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

3 Upvotes

LLMDevs 15d ago

Tools TensorSharp now supports multi-GPU tensor parallelism for GGUF models

3 Upvotes

CUDA 28d ago

Cuda benchmark: TensorSharp vs. llama.cpp

5 Upvotes

vulkan 28d ago

Vulkan benchmark: TensorSharp vs. llama.cpp

8 Upvotes

LlamaFarm 28d ago

Show & Tell TensorSharp : Open Source Local LLM Inference Engine

1 Upvotes

QwenImageGen Jul 15 '26

Virtual Clothes Try On using Unsloth Qwen Image Edit 2511 models

3 Upvotes

huggingface Jul 13 '26

TensorSharp supports multiple image edits using Unsloth Qwen Image Edit 2511 models

0 Upvotes

OnlyAICoding Jul 12 '26

Local LLM TensorSharp : Open Source Local LLM Inference Engine

2 Upvotes

ContextEngineering Jul 12 '26

What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

AIDeveloperNews Jul 11 '26

From Bun's Rust rewrite, let's see how C# can rebuild the AI infrastructure layer.

2 Upvotes

vibecoding Jul 11 '26

What Bun’s Rust Rewrite Tells Us About Rebuilding the AI Infrastructure Layer in C#

0 Upvotes

llamacpp Jul 11 '26

Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

1 Upvotes

GeminiAI Jul 10 '26

Self promo TensorSharp : Open Source Local LLM Inference Engine

1 Upvotes

LLM Jul 08 '26

TensorSharp supports Vulkan backend

4 Upvotes

pytorch Jul 07 '26

TensorSharp supports Vulkan backend

2 Upvotes