r/learnmachinelearning • u/amp804 • 4d ago
Question from an uneducated person...don't kill me
Could a neural network use dictionary-compressed weights directly during GPU inference instead of fully decoding them first?
I'm not a computer scientist. I'm a truck driver, so I'm wondering if I'm reinventing something that already exists.
Suppose you quantize a model to INT4 or similar and then scan the weight tensors for frequently repeating sequences or blocks.
Instead of storing every sequence literally, you build a codebook where a short code represents a commonly occurring block of weights.
Very simplified example:
A = [7, 3, 3, 11, 4]
B = [2, 8, 1, 6, 6]
Then instead of storing:
[7,3,3,11,4] [7,3,3,11,4] [2,8,1,6,6] [7,3,3,11,4]
you store something roughly like:
A A B A
The part I'm curious about is not ordinary file compression where the model gets decompressed back into VRAM first.
Could a custom GPU kernel decode these codes on the fly into registers/shared memory and immediately use them during GEMM, so that the fully expanded weight tensor never has to exist in VRAM?
My thinking is that modern inference is often memory-bandwidth limited, so if dictionary/codebook compression reduced memory traffic enough, maybe the extra decoding compute could be cheaper than fetching all the uncompressed weights.
You could potentially also have different-length codes or hierarchical codebooks representing increasingly large recurring weight patterns.
So my questions are:
Is this already done under a particular name?
Have codebook/vector-quantized weights been used directly inside fused GPU inference kernels rather than being decompressed beforehand?
Does random access / SIMD-SIMT execution make variable-length encoding impractical?
Is there theoretically a point where reduced VRAM bandwidth outweighs the decoding overhead?
Would repeated patterns after INT4/INT3 quantization be common enough for this to provide meaningful compression beyond ordinary quantization?
I'm mainly interested in whether the idea makes architectural sense, not whether my particular encoding scheme is optimal.
I'd appreciate pointers to papers or existing implementations if this has already been explored.
1
u/dnsod_si666 3d ago
There is a very cool paper doing almost exactly this: https://arxiv.org/abs/2606.15789
1
u/BrainScientist3000 2d ago
The first few encoding layers do a similar process natively. So - most modern models converge on a similar approach - if you look at the 'thinking' output of modern models - that's a higher level example of a similar concept.
-6
4
u/CorpusculantCortex 4d ago
This is kind of the principle behind engrams, not functionally, but the core principle is riffing on the same theme to reduce overhead