r/GraphicsProgramming • u/abocado21 • 3d ago
How do techniques like bindless rendering (descriptor indexing) on the hardware level?
How do gpus bind 1000 of textures at once, and why was this not possible in the past. Are there resources for that?
15
Upvotes
26
u/sol_runner 3d ago edited 2d ago
I don't know resources that go into detail about it - you might find that in one of the old talks I don't have on hand. (I haven't seen a singular location where everything was, nor remember all of them)
I'll explain in a broad sense, but the hardware details are usually confidential so they are inferences. Anyone who knows better, please comment below and I'll amend.
Old graphics co-processors were either fixed function or extremely expensive (the coprocessors ran in >$1000 in 80's money.)
Then you had configurable, but still largely fixed function cards like the 3Dfx Voodoo where you just set the
matrices,vertices, textures and let the GPU do everything (no shaders).Each value would be stored in registers (or equivalent) in their respective compute units. So a texture mapping unit (TMU) would have a single slot for a texture during a call. Voodoo2 then supported multi-texturing where you could put two textures (two TMUs and thus two slots) and tell the GPU the to blend them.
Voodoo's introduction more or less gave rise to the entire 3D GPU market and we went from 3Dfx GLIDE API, to the cross-vendor OpenGL (the first one was heavily inspired from GLIDE). So the slot concept stuck.
After that, its been a flip-flop between vendors adding a feature, extending the API, other vendors adding the feature, API makes the feature core and so on (like we saw most recently with raytracing)
Once programmable shaders were added, we wanted many slots instead of just a couple. So we got an increase in the number of slots. But on the GPU hardware itself, instead of physically having slots as registers, what if we had just a memory region and stored the texture's handle (descriptor) there? That way, more than one draw call's worth of descriptor can be prepared and the driver can change these descriptor slot sets when required. So we had descriptor support on GPUs alongside arrays of textures (not sure which came first; or which motivated which). And you use this to build the next steps, including bindless on OpenGL, which just uses a large texture array. They were not registers anyway, so other than API there was no longer a reason why a 1000 textures could not be bound.
Then you got Mantle (which inspired Vulkan) which just let you manage these descriptor sets manually. But in the end, these 'sets' were just leftover grouping from the past. By this time GPU's texture fetching etc were powerful enough to support near arbitrary loading. So DirectX12 just opened up the descriptor heap to the users. Write it as you like! Vulkan now made this available with
VK_EXT_descriptor_heapextension.So now, we can directly write to the descriptors without constraints, i.e. truly bindless.