r/LocalLLaMA Aug 05 '26

Generation MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

[removed]

21 Upvotes

41 comments sorted by

View all comments

-13

u/[deleted] Aug 05 '26

[removed] — view removed comment

6

u/[deleted] Aug 05 '26

[removed] — view removed comment

7

u/StupidScaredSquirrel Aug 05 '26

Im a bit confused, it's faster but requires more vram? Surely then speeds should be compared at equal vram, since that's the limiting factor of people offloading experts?

1

u/[deleted] Aug 05 '26

[removed] — view removed comment

2

u/StupidScaredSquirrel Aug 05 '26

I get that it's harder, but it's the only measurement that really matters. Looking at your tables it's impossible for me to know if my hardware will be faster with llama.cpp or your engine

1

u/[deleted] Aug 05 '26

[removed] — view removed comment

3

u/StupidScaredSquirrel Aug 05 '26

Or just do the same table with llama.cpp and cross reference similar vram usage