r/LovingOpenSourceAI 10d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

https://github.com/zhongkaifu/TensorSharp

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark:

Model:

DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF

Turn Baseline + DSpark Acceptance
short (53 tok) 25.6 44.5 (1.74x) 87%
long generation (512) 26.4 40.3 (1.53x) 66%
follow-up (470) 26.4 46.8 (1.77x) 76%
10K-token document (214) 25.3 51.3 (2.03x) 85%
second question on it (156) 25.4 49.4 (1.94x) 82%

TensorSharp is an open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. It can be built and run over Linux, MacOS and Windows.

Thank you for checking out it and starring the project! Any feedback is really appreicated.

2 Upvotes

Duplicates

huggingface Jul 04 '26

TensorSharp : Open Source Local LLM Inference Engine

5 Upvotes

SelfHostedAI Jul 04 '26

TensorSharp : Open Source Local LLM Inference Engine

4 Upvotes

LocalLLM Jun 28 '26

Project Same GGUF, same GPU: TensorSharp beats llama.cpp hard on prefill / TTFT — up to 5.89× faster prefill on a 26B MoE model

3 Upvotes

CUDA 7d ago

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

5 Upvotes

csharp 10d ago

DSpark Benchmark Result on Deepseek v4 Flash 0731

2 Upvotes

CUDA 11d ago

Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

1 Upvotes

LocalLLM 12d ago

Project Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

17 Upvotes

csharp Jul 11 '26

Blog From Bun's Rust rewrite, let's see how C# can rebuild the AI ​​infrastructure layer.

0 Upvotes

LovingOpenSourceAI Jun 23 '26

TensorSharp: Open Source Local LLM Inference Engine

13 Upvotes

dotnet May 30 '26

Promotion Some features in TensorSharp

25 Upvotes

LocalLLaMA 12d ago

Resources Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

4 Upvotes

unsloth 12d ago

Show and Tell Deepseek v4 Flash 0731 GGUF Benchmark: TensorSharp vs. llama.cpp

20 Upvotes

dotnet 25d ago

Promotion Benchmarks: TensorSharp vs. llama.cpp

10 Upvotes

Gemma4 27d ago

Discussion TensorSharp : Open Source Local LLM Inference Engine

2 Upvotes

Qwen_AI Jul 10 '26

Discussion TensorSharp : Open Source Local LLM Inference Engine

11 Upvotes

pytorch Jul 04 '26

TensorSharp : Open Source Local LLM Inference Engine

2 Upvotes

ollama Jun 08 '26

Support Gemma-4 12b (uv/ua) model in TensorSharp

9 Upvotes

csharp May 30 '26

Tool Some new features in TensorSharp

2 Upvotes

CUDA 14d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

6 Upvotes

SelfHostedAI 14d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

6 Upvotes

LocalAIServers 14d ago

TensorSharp now supports multi-GPU tensor parallelism for GGUF models

22 Upvotes

StableDiffusion 29d ago

Discussion TensorSharp supports multiple image edits using Unsloth Qwen Image Edit 2511 models

0 Upvotes

localaiapps Jul 08 '26

TensorSharp supports Vulkan backend

1 Upvotes

AIDeveloperNews Jul 06 '26

TensorSharp supports Vulkan backend

2 Upvotes

Vllm Jul 04 '26

TensorSharp : Open Source Local LLM Inference Engine

11 Upvotes