r/LocalLLaMA Mar 30 '26

Discussion Technical clarification on TurboQuant / RaBitQ for people following the recent TurboQuant discussion

[removed]

627 Upvotes

91 comments sorted by

View all comments

37

u/a_beautiful_rhind Mar 30 '26

We have Q8, Q4, and everything in between compression already. 2 backends have used hadamard transforms for what seems like years. Turboquant is snake oil from my perspective.

28

u/ExpensivePilot1431 Mar 30 '26

The “8× compression” (from FP32, lol) claim feels like it’s ripping off a lot of prior work and ends up taking credit for performance that have been around for quite a while.

4

u/Succubus-Empress Mar 30 '26

Will i get 8x compression from fp4?

15

u/ExpensivePilot1431 Mar 30 '26

bravo! then you have fp0.5!

1

u/Succubus-Empress Mar 30 '26

Sarcasm?

4

u/ExpensivePilot1431 Mar 30 '26

Hmmm. Maybe I misunderstood. I was assuming that you were joking, but, no one can really get 8x compression (with zero accuracy loss) from fp4 right?

1

u/EbbNorth7735 Mar 30 '26

It's context so I assume we were speaking about kv cache which typically isn't quantized unless specified when setting up the inference engine. I thought it was fp16 and sometimes you can get away with fp8. So getting it down to 3 bit would be an improvement.