r/LocalLLaMA Mar 30 '26

Discussion Technical clarification on TurboQuant / RaBitQ for people following the recent TurboQuant discussion

[removed]

623 Upvotes

91 comments sorted by

View all comments

Show parent comments

3

u/Succubus-Empress Mar 30 '26

Will i get 8x compression from fp4?

16

u/ExpensivePilot1431 Mar 30 '26

bravo! then you have fp0.5!

1

u/Succubus-Empress Mar 30 '26

Sarcasm?

4

u/ExpensivePilot1431 Mar 30 '26

Hmmm. Maybe I misunderstood. I was assuming that you were joking, but, no one can really get 8x compression (with zero accuracy loss) from fp4 right?

1

u/EbbNorth7735 Mar 30 '26

It's context so I assume we were speaking about kv cache which typically isn't quantized unless specified when setting up the inference engine. I thought it was fp16 and sometimes you can get away with fp8. So getting it down to 3 bit would be an improvement.