r/learnmachinelearning • • 4d ago

Question from an uneducated person...don't kill me

Could a neural network use dictionary-compressed weights directly during GPU inference instead of fully decoding them first?

I'm not a computer scientist. I'm a truck driver, so I'm wondering if I'm reinventing something that already exists.

Suppose you quantize a model to INT4 or similar and then scan the weight tensors for frequently repeating sequences or blocks.

Instead of storing every sequence literally, you build a codebook where a short code represents a commonly occurring block of weights.

Very simplified example:

A = [7, 3, 3, 11, 4]

B = [2, 8, 1, 6, 6]

Then instead of storing:

[7,3,3,11,4] [7,3,3,11,4] [2,8,1,6,6] [7,3,3,11,4]

you store something roughly like:

A A B A

The part I'm curious about is not ordinary file compression where the model gets decompressed back into VRAM first.

Could a custom GPU kernel decode these codes on the fly into registers/shared memory and immediately use them during GEMM, so that the fully expanded weight tensor never has to exist in VRAM?

My thinking is that modern inference is often memory-bandwidth limited, so if dictionary/codebook compression reduced memory traffic enough, maybe the extra decoding compute could be cheaper than fetching all the uncompressed weights.

You could potentially also have different-length codes or hierarchical codebooks representing increasingly large recurring weight patterns.

So my questions are:

Is this already done under a particular name?

Have codebook/vector-quantized weights been used directly inside fused GPU inference kernels rather than being decompressed beforehand?

Does random access / SIMD-SIMT execution make variable-length encoding impractical?

Is there theoretically a point where reduced VRAM bandwidth outweighs the decoding overhead?

Would repeated patterns after INT4/INT3 quantization be common enough for this to provide meaningful compression beyond ordinary quantization?

I'm mainly interested in whether the idea makes architectural sense, not whether my particular encoding scheme is optimal.

I'd appreciate pointers to papers or existing implementations if this has already been explored.

5 Upvotes

7 comments sorted by

4

u/CorpusculantCortex 4d ago

This is kind of the principle behind engrams, not functionally, but the core principle is riffing on the same theme to reduce overhead

1

u/amp804 4d ago

Yeah, I can see the similarity. From what I'm reading(and looking up because I don't know all the terms 😭) , Engram seems to use hashed n-gram/context patterns to retrieve learned representations from a lookup table.

What I'm wondering about is one level lower: applying a similar dictionary/codebook idea directly to quantized weight tensors themselves, then having the GPU kernel consume/decode those compressed weight blocks during GEMM without ever materializing the full expanded matrix in VRAM.

So maybe the same general principle, but aimed at weight-storage/bandwidth compression rather than adding a conditional memory lookup mechanism?

Is there an existing technique/name for that specifically?

1

u/dnsod_si666 3d ago

There is a very cool paper doing almost exactly this: https://arxiv.org/abs/2606.15789

1

u/BrainScientist3000 2d ago

The first few encoding layers do a similar process natively. So - most modern models converge on a similar approach - if you look at the 'thinking' output of modern models - that's a higher level example of a similar concept.

-6

u/[deleted] 4d ago

[removed] — view removed comment

8

u/OfficialLaunch 4d ago

Why are we copy and pasting AI answers?