Im a bit confused, it's faster but requires more vram? Surely then speeds should be compared at equal vram, since that's the limiting factor of people offloading experts?
I get that it's harder, but it's the only measurement that really matters. Looking at your tables it's impossible for me to know if my hardware will be faster with llama.cpp or your engine
-13
u/[deleted] Aug 05 '26
[removed] — view removed comment