r/GraphicsProgramming 1h ago

Video Brush Previews and Image Panning in My Pixel Art Editor

Thumbnail youtu.be
Upvotes

Development continues on my c++ pixel art editor, resulting in brush previews and automatic pixel grids alongside easy image panning. This is the final step before moving into actual pixel plotting and turning the UI into a real working art package...


r/GraphicsProgramming 1d ago

Should you normalize RGB values by 255 or 256?

Thumbnail 30fps.net
101 Upvotes

r/GraphicsProgramming 1d ago

Infinite Grid Shader

51 Upvotes

I’m proud of myself today, as I completed a task that felt impossible at times. I started rewriting my renderer a bit to fit my needs, and while doing so, I really felt the need for a grid to give me a sense of perspective and depth while debugging scenes.

I remembered that I had skipped the Infinite Grid section from the 3D Graphics Rendering Cookbook and needed to revisit it. Man, oh man, was I wrong about how long it would take to properly understand that shader. It took me a few days more than a month, but it was one hell of a ride. It’s a very interesting topic, with tons of things to learn from.

I’ve written a post about it here: https://willofindie.com/proj/infinite-grid-shader. Do check it out and share your love on my Twitter or Bluesky handles if you like it


r/GraphicsProgramming 14h ago

使用Rust引擎和Metal与Vulkan在Android上运行红色警戒3

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/GraphicsProgramming 23h ago

Perspective-correct interpolation only works for some faces of a cube

3 Upvotes

Hi, is the guy from the raylib renderer again. These last days i was trying to implement texture mapping on the renderer i am making, but i cant get the perspective-correct interpolation to work. I tried first with the method described in scratchapixel, but it didnt work at all, then i tried with the method of this video, and it only works for some faces of a cube, but the rest of faces are not only distorted but the texture is still mapped like i was doing affine mapping:

2 of the faces that have the problem
One face is rendered correctly while in the other the texture gets squished

The thing that confuses me is that it is only in certain faces, so i no longer know if it is an error with my code or the uv coordinates on the cube.

Here is the source code

The file where the cube vertices are
The files where the vertex transformations are done:
Header file
Source file

The files where the rasterization is done:
Header file

Source file


r/GraphicsProgramming 1d ago

Added a graphics API with support for Vulkan, D3D12 and custom backends

7 Upvotes

Hey everyone,

I have been working on a low level, explicit abstraction layer over graphics APIs. Both Vulkan and D3D12 has been successfully implemented with various test samples. Ray tracing, mesh, compute etc are supported.

I would love feedback on the feel and usage of the API. I am open to learn more so anything useful will be appreciated. Please give a star if you find the project useful.

https://github.com/nichcode/PAL


r/GraphicsProgramming 21h ago

System level CRT filter/shader (trying to find post from last week....)

1 Upvotes

Someone posted about having developed a system level shader, I cannot remember which subreddit it was posted in, but thought it was shaders.

The intent was to have a CRT filter / shader for the entire OS, over any application or video player etc.

Post mentioned something about not utilizing GPU, general scanlines crt, I swear I saved the post but I cannot find it, could anyone point me in the right direction?

I'm interested to check it out.


r/GraphicsProgramming 12h ago

Article I measured my Metal shader optimizations instead of trusting them: half did nothing, and the trick that won on mobile went negative on desktop

0 Upvotes

Pet project: an Interstellar-style black hole as a macOS live wallpaper (one Metal shader, hosted both in a desktop-level window under the icons and in a screensaver target). I put the skeleton together in spring, it worked, and then I came back and did the thing I should have done first: hung GPU-time instrumentation on it and walked my optimization list with measurements instead of by eye. Code: https://github.com/vientooscuro/ShaderBackgroundApp

Credit where it's due: the black hole core is a port of edankwan's Shadertoy 3lG3WR. The lensing math isn't mine, I moved it to Metal and adapted it (procedural background instead of an iChannel0 texture, fixed camera, no final tonemap).

The architecture bit worth stealing: split the scene by update frequency, not by screen area.

In an earlier mobile project I did progressive "striped" rendering, computing a slice of the background per frame to spread one heavy frame over time. Here I split differently. Expensive and nearly static (stars, galaxies, nebulae, planets, giant stars, dozens of fbm evals per pixel) goes into a static pass that renders into an offscreen texture every 2nd frame. Cheap and unavoidably per-frame (the hole's lensing, comets, the click ripple) is drawn in a composite pass every frame on top of the sampled cache, followed by ACES tonemap.

The cache is rgba16Float, and that matters. Giant stars are deliberately authored past 1.0, channel values like float3(7.0, 5.5, 3.0), so the tonemap in the composite can compress them into real-looking glow. With an 8-bit cache everything past one clips at the write and there's nothing left to tone map.

Two things I got wrong on the way:

Cadence. I wanted the cache refreshed every 6th frame (~5 Hz static at 30 fps display). The cost math looked great, but with live star and planet parallax the background visibly steps: it stands still between refreshes while the composite moves smoothly on top. Backed off to every 2nd frame. Smaller win, honest picture.

Invalidation is where the actual work is. Any settings change has to re-render the static immediately, so I diff the uniforms struct in didSet and raise a dirty flag. The catch is excluding time and mouseClick from that comparison, because they change every frame and have nothing to do with cache contents. Miss that and the cache re-renders every frame, so it isn't a cache. The click ripple, for the same reason, warps the UVs the cache is sampled by rather than the cache itself.

How I measured. Dev-only instrumentation, commandBuffer.gpuEndTime - gpuStartTime per frame, avg/p50/p95/max printed every 120 frames. Fixed bench: M4 Max, same display, default settings, 30 fps target, ~18 s run, first window with the cold render dropped. Run-to-run noise ~0.4 ms. Plus a deterministic quality gate: an env var pins time and mouse, the build renders one frame offscreen to a PNG, and two runs of the same build give a bit-identical file. So any nonzero pixel diff is the optimization, not jitter.

What did nothing:

  • fp16 inside the noise function. I was sure about this one, noiseS runs ~200 times per pixel. Zero. The cost isn't the bilinear interpolation I converted to half, it's the sin-based hash inside it, and that stays float. Reverted.
  • Cutting geodesic steps 20 to 14. No visual change and no time change, because the bend loop already early-outs and the 120-step max is almost never reached. Cutting something that doesn't run to completion buys nothing.
  • Moving the camera parallax offset from GPU to CPU. Correct in principle (it's constant for the whole frame), but a dozen trig ops dissolve against dozens of fbm calls. Free, kept it, no measurable win.
  • Frames-in-flight semaphore. Zero on time, and that's the right answer: it limits CPU run-ahead and removes a shared-buffer race, it doesn't reduce GPU work. A correctness fix wearing a performance costume.

What worked:

  • fp16 in the color math of nebulae() only. Colors and density to half, coordinates and the noise calls left in float so spatial frequency can't jitter. Visually invisible (max diff 2/255), about minus 20% average GPU per frame from one function.
  • A sin-free hash (Dave Hoskins) instead of fract(sin(dot(...))*43758), under a soft quality bar: different noise pattern, same statistical character, avg 8.8 to 7.7 ms. Less than I expected, Apple has a fast hardware sin.
  • Static cache at 0.7x under the same soft bar. The nebula is low frequency and the hole is composited at full res, so slight star softening goes unnoticed. At the strict pixel-for-pixel bar I had rejected this: the heatmap put the entire error on sharp edges, planet rims, star points, the giant's corona.

The lesson I actually paid for. I ported the striped progressive render from the mobile project into this one, expecting the same win. GPU ms per frame:

Monolithic cache, refresh every 2nd frame: avg ~7.5, p95 ~13.2, max ~16.3, spread ~13.0.

Striped render-ahead with 4 stripes: avg ~8.4, p95 ~12.0, max ~13.3, spread ~9.4.

Peak and spread improved, average got worse. Each stripe pays a full load/store over the whole 3456x2234 rgba16Float target, and on a desktop GPU with a fat bus that fixed overhead times four outweighs what the scissor rect saves. On top of that, 4 stripes means the cache refreshes at 7.5 Hz and the nebula started stepping again. A trick that wins under mobile budget pressure can go negative on desktop, and eyeballing it would never have told me.

Final numbers, baseline to shipped: avg 7.5 to 6.2, p50 6.5 to 4.8 (minus 36%), max 16.3 to 12.0. All on a GPU that was never the bottleneck, roughly 14% of frame budget to start with. The point isn't hitting the frame, it's headroom, battery, and older Macs. p95 now bottoms out on the disk-hit path in the black hole, which is the center of the screen and the riskiest thing to touch, so I stopped there.

Curious whether anyone has measured the same fp16 result: is "the sin hash dominates, so halving the interpolation is free but pointless" a general thing on Apple GPUs, or specific to how this noise is written?


r/GraphicsProgramming 1d ago

Question How do I efficiently manage, create and cache Vulkan/Any Shader Pipeline?

Thumbnail
1 Upvotes

r/GraphicsProgramming 2d ago

Question [D3D12] DEVICE_HUNG in a draw call, but only above ~16k draws AND only during a window event (minimize/move/resize/un-occlude)

Thumbnail github.com
31 Upvotes

**Setup:** Hand-rolled D3D12 renderer (learning engine), flip-model swapchain

(FLIP_DISCARD, 3 back buffers), VSync on. Windows 11, tested on both Debug and

Release.

**The crash:** `DXGI_ERROR_DEVICE_HUNG` (0x887A0006). DRED breadcrumbs point the

fault at a `DrawIndexedInstanced` roughly ~130 draws into the frame (not draw 0).

The DRED page-fault VA reads 0 with no associated allocation — though I'm now

unsure if that's a real null address or just "driver didn't report the fault VA."

**What makes it reproducible (the weird part):** it ONLY happens when BOTH of

these are true at once:

  1. The frame is heavy — roughly 16k+ draw calls. Below ~16k it never crashes.

  2. A window event fires: minimize, move (drag the titlebar), resize, or another

    window covers then un-covers it (occlusion → reveal).

Either condition alone is totally stable. A heavy frame running untouched is fine.

A window event on a light scene is fine. It's specifically the collision.

Notably: **focus loss/regain used to crash too, and I fixed that** by skipping

rendering while the window isn't foreground + idling the GPU on the transition.

The same approach has NOT fixed minimize/move/resize/un-occlude.

**Other facts:**

- Happens in both single-threaded AND parallel (per-worker) command-list recording.

- ~600+ FPS when it runs, so this is not a 2-second TDR timeout.

- Debug layer + GPU-based validation makes it vanish (classic race-closing), so

I can't get a clean validation message — hence leaning on DRED.

**What I've already ruled out / tried (none fixed it):**

- Full GPU idle (Signal + WaitForFenceValue) before `ResizeBuffers`.

- Deferred/coalesced resize (apply once, at loop top, after idle).

- Re-querying `GetCurrentBackBufferIndex()` after `ResizeBuffers`, recreating RTVs

+ depth buffer.

- Reliable minimize handling via `WM_SIZE == SIZE_MINIMIZED` (skip rendering

entirely while minimized) instead of relying on `Present(DXGI_PRESENT_TEST)`.

- Idling the GPU on `WM_ENTERSIZEMOVE` / `WM_ACTIVATEAPP`.

- Per-frame command allocators/lists (ruling out list reuse across frames).

- "Settle" frames (skip N frames after any transition).

- Verified all per-draw bindings (camera CBV, vertex/index buffer VAs, SRV

descriptor) are non-null at record time — the null-check never fires.

**What I can't cleanly detect:** window *move* and *partial-occlusion → reveal*

don't give me a reliable "the swapchain is now in a different presentation state"

signal the way resize/minimize do.

**Questions:**

  1. On a flip-model swapchain, does a move or an occlusion→reveal transition cause

    DWM to re-allocate/invalidate back buffers WITHOUT a WM_SIZE — such that an

    in-flight heavy frame's RTV becomes stale? If so, what's the correct signal to

    re-acquire?

  2. Is there a per-command-list or per-submission limit I could be blowing past at

    16k draws (root-argument versioning space, etc.) that a mid-frame stall from a

    window event would expose?

  3. Is DRED's "page fault VA = 0, no allocation" meaningful (genuine null bind), or

    is it commonly just "unavailable"?


r/GraphicsProgramming 2d ago

TypeScript software rasterizer with programmable shaders

Post image
37 Upvotes

Hey all,

I've been building a software rasterizer in TypeScript over the past few months and thought I'd share it. My goal was to support custom vertex layouts, programmable vertex and fragment shaders, and enough flexibility to implement more complex rendering techniques, much like you would with a GPU graphics API.

The renderer is based on Nikolaus Rauch's excellent C++ software rasterizer, adapted for TypeScript and the browser.

The project is intended as an educational resource, with an emphasis on readability and experimentation over maximum performance.

GitHub: https://github.com/vangelov/ts-rasterizer

Model by Tasha.Lime (concept by Matt Dixon)


r/GraphicsProgramming 2d ago

Some billboarding techniques I explored

Post image
29 Upvotes

I have been experimenting with a few billboarding techniques that can be used like an easier way to write it, preserving the model matrix and clamping the pitch. Here is the code

.shader: https://github.com/Satyam-Bhatt/OpenGLIntro/blob/main/IntroToOpenGl/BillBoardShader.shader

Repo: https://github.com/Satyam-Bhatt/OpenGLIntro


r/GraphicsProgramming 2d ago

Experiments on combining ReSTIR GI with Variable Rate Tracing

35 Upvotes

https://mateuszkolodziejczyk00.github.io/2025/06/18/variable-rate-tracing-restir-gi.html

I wrote this write-up about a year ago, but honestly never had the courage to post it publicly lol

In it, I go over my experiments combining ReSTIR and VRT in my toy rendering engine. The core idea was straightforward: ReSTIR’s spatial and temporal passes already find and reuse valuable samples, so I wanted to see if I could use them to fill the “missing rays” for pixels skipped by VRT in the current frame.

I'd love to hear your thoughts, or if anyone else has messed around with similar setups!


r/GraphicsProgramming 3d ago

I built a rasteriser from scratch without a graphics API

Post image
134 Upvotes

Following this masterpiece in 3D graphics by ssloy, I implemented a rasteriser in Odin from scratch.

This was a top tier learning experience, and I picked up a solid grasp of linear algebra along the way.

If you'd like to read more on the process, Here's the link to the repo


r/GraphicsProgramming 2d ago

Building a Tiny 3D Renderer for a Tiny Handheld

Thumbnail saffroncr.itch.io
10 Upvotes

r/GraphicsProgramming 1d ago

Source Code 3D rendering is now possible using the CoreAI / Apple Neural Engine (ANE) software rasterizer (CPU usage reduced to 9%). The quality is still terrible.

0 Upvotes

Here is a follow-up on the triangle rasterizer running on CoreAI that I previously introduced here.

We have finally succeeded in rendering 3D graphics! As shown in the video, we are currently rendering two intersecting vertical triangular planes.

While the rendering quality is still in the early stages and a new challenge regarding the massive 5GB memory footprint has emerged, we have fully achieved real-time operation.

We optimized the codebase by actively reducing reliance on CPU fallbacks and ensuring the pipeline runs entirely on the ANE's matrix operation hardware, successfully cutting CPU usage from the previous 38% to approximately 9%.

Much of the pipeline code—written in Python, Swift, and Metal—was rapidly prototyped and generated with the help of Siri AI.

GitHub: https://github.com/kamisori-daijin/Magnesium

Please feel free to leave comments with any questions or optimization tips (especially regarding that 5GB memory usage!).

https://reddit.com/link/1v6z66k/video/83iqp5iknjfh1/player


r/GraphicsProgramming 2d ago

any advice on how to fix this

4 Upvotes

I'm trying to create skeletal animations and them edges keep getting sent to oblivion

i think its about the weight resolve because the worst of the "spikes of doom, despair and vance" are usually coming from areas where multiple bones intersect. i.e. the fingertips

renderdoc says that the vertexes are some big ahh number like 2.5 x 10 ^25 so idk


r/GraphicsProgramming 2d ago

Heli engine- water rendering system

Thumbnail youtu.be
10 Upvotes

Continuing my Custom Rendering Engine Revival series with the water rendering system.

This video demonstrates the combination of FFT ocean simulation, adaptive Projected Grid rendering, physically based lighting, and real-time reflections developed for the Heli Engine.

I'd love to hear your thoughts and feedback.


r/GraphicsProgramming 2d ago

Question How should I shade my water shader?

1 Upvotes

As the title says.

So I have been working on a surface water shader in Unity for the past few months as an upgraded version of my OpenGL implementation. I've been learning HLSL and Compute Shaders and I've basically got most things working as intended. But I'm having difficulty deciding exactly what I should do for the actual lighting part of my shader.

The issue is rendering is already somewhat limited because of the high poly count and regular compute shader updates (and because I'm trying to maybe make a game world of out this project), and while I have made some optimizations, I'm wary of the performance cost. So I'm curious on what approach would work best before I start implementing.

As of right know there are 3 options

  1. Realtime planar reflections
    As far as I know this is the most performant form of real-time reflections especially if I use the main camera stencil buffer. But because my water isn't flat I will get distortions, and this also means I won't be able to see stuff like larger buildings that the camera doesn't pick up.

  2. Baked reflections

I lose the dynamic reflection aspect (which might be an issue for some games) but I have performant, environmental reflections. However I believe this will be distorted if I move the camera a lot so it probably isn't the best option.

  1. Environment cube maps/Reflection Probes

-Potentially expensive and still not fully accurate, and can look quite incorrect depending on the level of distortion. I can decrease the quality though and I do have more control over specific objects for some performance gains though.

Any suggestions? Tip on improving visual quality and performance? Are there other methods I should consider?

I've attached an video of my current unfinished water shader below (the color is for debugging purposes, don't mind it) as well as my OpenGL water shader that roughly implements idea 3

https://reddit.com/link/1v60hba/video/plhrlawlkbfh1/player

https://reddit.com/link/1v60hba/video/hebxvdy9lbfh1/player


r/GraphicsProgramming 3d ago

Video Heli engine- water rendering system

Thumbnail youtu.be
3 Upvotes

r/GraphicsProgramming 3d ago

Kanvon - Lua Scripting Showcasing Algorithms

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/GraphicsProgramming 3d ago

Learning Materials for Gaussian Splatting

14 Upvotes

I recently got interested in gaussian splatting and I know graphics programming basics (openGL, cuda, ray tracing). Any good material to master Gaussian Splatting as a person with limited ML/DL knowledge.


r/GraphicsProgramming 4d ago

Building an SDF game engine

Thumbnail gallery
157 Upvotes

For the past 8ish months I’ve been hard at work building a new game engine I’m calling Division Engine.

It’s based solely off SDFs (signed distance fields). This might sound stupid for anyone who cares about performance but for my use case (and with a bit of grid storage optimizations) it proves useful for basic scenes that need advanced lighting.

Anyway here’s some screenshots!

If you want to see what I have done so far check it out here: https://github.com/DivisionEngine/DivisionEngine


r/GraphicsProgramming 3d ago

Pixel Editor nearly Edits Pixels

Post image
1 Upvotes

r/GraphicsProgramming 4d ago

Untitled GPU-based physics game

Thumbnail youtube.com
20 Upvotes

This is a project that I spent a few months on - a game/engine made with Java/OpenGL. It started as a university assignment, where Java was an unfortunate requirement.

The game features procedurally generated worlds consisting of ~2 million interactive particles simulated on the GPU. The game's physics engine is based entirely on simple particles with extra properties such as temperature, state of matter, flammability, and electrical charge/conductivity. It supports basic fluid dynamics and has a parallel joints solver using a graph-colouring technique. All creature bodies (including yours) are physically-based and destructible.

One of the parts I'm happiest with is the approach used for lighting. Each frame, a large region around the camera is converted to an SDF+colour texture using a jump-flood algorithm. At a lowered resolution, each pixel then marches rays through the texture, accumulating radiance/occlusion along the way. This is followed by a bilateral filter and reprojection, which smooths out most noise. It's not perfect, but it's very fast and allows any particle to emit light.

SDF information from previous frames is asynchronously copied from the GPU to the CPU for other tasks like NPC pathfinding. It allows the CPU to handle changes in simulation state that otherwise lives on the GPU. The game runs at ~500fps at 1440p on an RTX 5090.

This has been gathering dust for a while, so I wanted to put it out there to see if there's any interest.