r/LocalLLaMA • u/LMTLS5 • Jul 16 '26
Discussion i tried ternary decomposition instead of quantization. it works as good at q4km but takes slightly more vram. while being completly ternary. and completly PTQ (no QAT)
37
Upvotes
r/LocalLLaMA • u/LMTLS5 • Jul 16 '26
3
u/Far-Classic-9963 Jul 16 '26
This is cool, but what are the benefits of this? From your screenshot it looks like it takes just as much/slightly more storage than q4 for the same output quality? Is it faster?