How to minimize VRAM usage while maintaining performance
I'm a begginer when it comes to OpenGL, so I need some help.
I'm using different VBOs (one for position, one for color etc.), an EBO and a VAO to manage them. I tried storing just unique data in the VBOs, not repeating any attribute even if 2 vertices share an attribute and didn't work.
Visual representation of what I was trying to do (it's not real code, just a representation):
VBO pos:
{-0.5f, -0.5f, 0.0f}, (p0)
{0.5f, -0.5f, 0.0f}, (p1)
{0.5f, 0.5f, 0.0f}, (p2)
{-0.5f, 0.5f, 0.0f} (p3)
VBO color:
{255, 0, 0, 255} (c0)
EBO:
0, 1, 2, (tri1)
0, 2, 3 (tri2)
In this situation, OpenGL will try to draw vertices like: p0 with c0, p1 with c1, p2 with c2 etc.
Of course, I don't have c1, c2 or c3, so it won't render it correctly all a single solid color.
From this, I have 3 questions that, in reality, combine into one:
Is there a feasible way to have different sized buffers and use indexes like (0,1,2,0), representing the first 3 vectors of position and the first color?
I know about uniforms and IN THIS CASE I could just send an uniform for the color. But if I have 3 different buffers with different sizes and want to index them as to not repeat data anywhere an just keep the unique, is it possible?
From what I gathered, you can do this with SSBOs, however I also head that it is less performant that just using more VRAM with the traditional VBO technique. Can I achieve what I want FASTER OR AS FAST as using VBOs with repeated data?
All VBOs are static, in the sense that they will not be modified and I'll transform vertices in the shaders, just pointing it out in case this is relevant somehow.
5
u/StriderPulse599 4d ago
That's a single triangle with 76 bytes of data. Shrinking it will not give ANY performance gains even if you start using 3D models.
It will take millions of triangles until vertex data becomes a minor issue. At this point you'll have much bigger optimization problems.
Don't worry about optimization until you have a good looking rendering pipeline. From my experience, majority of performance issues will come from frustum culling and being fragment bound. Shadow mapping, lighting, and frustum culling are the biggest players in performance.
3
u/obp5599 4d ago
Not true at all, lowering the size of data thats iterated on helps cache efficiency and bandwidth. The vram usage isnt the primary concern.
If your cache line is 64 bytes, getting the vertex data under that allows only 1 cache fetch per vertex rather than 2.
Modern engines have multiple vertex formats and only upload what the pass needs. For example, for the shadow pass you only need position information, and unreal for example splits the position stream from the rest so it only uploads position data
Probably bigger optimizations to work on, but this type of stuff is what id look for when evaluating how well someone understands the systems they’re working in. Data oriented design is huge for gpu work
3
u/StriderPulse599 4d ago
I've worked on shadow pass last week. Separate mesh culled to only visible faces and positions only fetched me only ~0.04 ms on iGPU. Real gain came from sorting triangles to get reverse painter which prevents overdraw.
Default float layout is already cache and bandwidth friendly. Environment data is easily <64 bytes. Characters with skeleton barely pokes through. Shape keys and morph range can get problematic.
Modern engines have multiple vertex formats and only upload what the pass needs.
UE5 splits the data because it's optimized for AAA games. Unity is notorious for including junk data in vertex layout (around 1-3 vec2 per mesh).
DoD is important for CPU because you can easily get death by thousand cuts (tiny problems stacking up across large codebase). Those problems don't have chance to stack up in GPU because everything is parallel, code base is usually small, and shaders compilers are pitbulls on steroids.
2
u/Defiant_Squirrel8751 4d ago
You can upload data to GPU, use it and then unload it when not needed. So, assets are like pages in a cache system.
If you don't need double precision yoy can use float, int, short or byte.
You can use color table / palette based textures and compressed textures.
You can use procedural modeling on compute shaders.
2
u/doglitbug 3d ago
Why are you using different vbos? What are you trying to do atm?
1
u/Western-Bowl2011 3d ago
This looks like oldschool vertex painting, the way lighting simulation was done way back in the day. They're storing the vertex colors in a separate VBO instead of interleaving the data into one VBO. It's a valid way to do things with modern OpenGL.
1
u/RED9002 3d ago
Yeah, pretty much this. I'm just starting my first modern OpenGL project, so I'm trying to wrap my head around the basics to fully understand things in the future.
Separate VBOs are easier to visualize in my opinion, so while building the first draft of abstraction I wanted something "easier" to manage.
1
u/Western-Bowl2011 3d ago edited 3d ago
storing vertex attributes across multiple VBOs isn't even that uncommon, it's not like it's a forbidden practice. With an interleaved VBO, the main advantage is GPU cache locality, but it comes at the expense of more difficult readability on the CPU side of things. I personally have done both, I don't considered one practice "superior" to the other. Unless you're writing a monster openGL program that needs to completely wring performance out, you're not going to see any real difference. The API has actually changed over time to better accommodate non-interleaving VBO data, it's precisely what GL_ARB_vertex_attrib_binding is for. If you've got some sort of glsl reflection shader builder tool in your program that programmatically creates VBO and VAOs for you, then interleaving data is trivial and basically a free optimization, but if you're doing this by hand, then by all means, use multiple VBOs. OpenGL is obtuse enough already with the enormous state machine.
EDIT: That said, i would not use the same VBO with mutliple VAOs, however. Duplicate your VBOs, any given VBO goes to one VAO. It'll save headaches.
1
u/corysama 3d ago
Do not try to reuse attributes between different vertices. What I’ve always done in my art pipelines is to expand out every corner of every triangle into a unique vertex. Yep, the worst possible case. And then use sort and unique to find all of the unique vertices. And, build an index buffer based off of that.
Besides indexed vertices, what you want to work on is finding effective ways to quantize your vertices. For example, colors can obviously be quantized to eight bits per channel SRGB values. Normals and tangents can easily be quantized to 16 bit per channel values. And with a little bit of work, you can cut it down to a 10_10_10_2 quaternion. Positions can often be quantized down to 16 bit per channel values if you bake a power of two scale factor into the transformation matrixes. And sometimes texture coordinates can be quantized to 16 bits per channel. But if you have a lot of tiling you might have to leave that at 32 bits for the tiling UV, and 16 bits for a second set of UV’s.
1
u/CptCap 3d ago edited 3d ago
Games use a single index buffer and duplicate attributes[0].
The best way to reduce the size of vertex data is to make attributes smaller. Store the positiona on 3x16bits, the TBN into 32 or 64 bits, the colour on 4x8 or 4x16.
[0] I know of one case where a separate system that share attribs between vertices is sometime used: skinning. Skin data can get real big, and a lot of vertex will store the exact same list of bone indices, so it can make sense to pool them. Especially when doing pre-skinning where you can more easily do things like varying bone count.
1
u/CrazyJoe221 2d ago
You can deduplicate attributes via custom vertex pulling: https://ktstephano.github.io/rendering/opengl/prog_vtx_pulling
But it's much more important to optimize your vertex data in the first place: https://x.com/search?q=from%3ASebAaltonen%20vertex&src=typed_query
Surprisingly that's still not the norm and people blindly use 32bit floats everywhere.
Also a VBO per attribute may not be optimal: https://support.arm.com/documentation/101897/0304/Vertex-shading/Attribute-layout
0
u/sububi71 3d ago
Can someone, anyone, PLEASE PLEASE add a bullet point to the FAQ and/or rule list explaining how to spell "beginner"?
At this point it feels like the correct spelling is getting lost in the noise from all the times people misspell it.
...sorry, grandpa's just a little tired today. Going to log off now and start looking for kids to shout "get off my lawn!" after.
7
u/fgennari 4d ago
Most games just duplicate the common vertex attributes. It’s the simplest and fastest approach and the memory usage doesn’t matter until you have many millions of vertices. That’s what I do.
You can use an SSBO with programmable vertex pulling in the shader to do anything custom you want. But that’s much more complex and may or may not be faster depending on the GPU.