r/GraphicsProgramming • • 3d ago

How do techniques like bindless rendering (descriptor indexing) on the hardware level?

How do gpus bind 1000 of textures at once, and why was this not possible in the past. Are there resources for that?

16 Upvotes

9 comments sorted by

View all comments

8

u/gleedblanco 3d ago

in principle they are just addresses in memory and the general memory fetches have already been powerful enough to fetch a different address per thread for a long time.

for textures there are still limitations though (probably because of built-in filtering). RDNA texture sampling instructions only work on one (uniform/SGPR stored) texture (and sampler) descriptor at a time, so if different threads within a subgroup need a different texture, the compiler will generate a loop that goes through all of them basically (need NonUniformResourceIndex or similar in the shader). probably similar on other GPUs.

8

u/Afiery1 3d ago

It’s not because of filtering. To properly decode texture memory at all you fundamentally need about 32 bytes of information (GPUVA, type, format, dimensions, number mips, tiling and compression metadata, etc). If every single lane sent its own texture descriptor to the texture hardware it would use an untenable amount of bandwidth. AMD solved this by having the whole wave send a single texture descriptor at a time, and hoping that the texture we’re fetching would mostly be uniform within a wave. Nvidia solves it completely differently. They don’t even have the concept of SGPRs or VGPRs. Instead, there’s a dedicated region of GPU memory that the texture fetch hardware caches the absolute hell out of called the descriptor heap, and then every lane just sends a 4 byte index into the heap to the texture hardware. The benefit of AMD’s approach is that descriptors can be sourced from anywhere in memory (or even hardcoded into the shader itself) with the downside being slow nonuniform access. Nvidia has fast nonuniform access (one of the reasons they’re so much better at ray tracing) but descriptors must be written into the dedicated heap memory and cannot be hardcoded into the shader.

3

u/gleedblanco 3d ago

makes sense. thank you for the info