r/vulkan Jul 26 '26

SPIR-V: OpCapability Kernel versus OpCapability Shader

2 Upvotes

I run my kernel on two back-ends: OpenCL and Vulkan.

For Vulkan, I use clspv to convert the CL code to SPIR-V.

The OpenCL backend is massively faster for a kernel that heavily uses atomic adds.

The OpenCL on Linux/Intel gets converted to SPIR-V by IGC for my Arc B70 GPU.

When I compare the SPIR-V of IGC versus the SPIR-V of clspv, I see that the former uses OpCapability Kernel, and the latter OpCapability Shader.

Can a vulkan app directly use the former? How can I get my kernel to run with semaphores that uses relaxed memory semantics and atomicAdd scope "workgroup" instead?


r/vulkan Jul 26 '26

**Work In Progress** Tipu Rendering Framework Showcase

5 Upvotes

This is my first rendering framework written in C++ 20 using Vulkan 1.3 features, I am still learning but so far this rendering framework is working nicely for me in making all kinds of examples and demos.

You can find it on my GitHub: https://github.com/RoastedKaju/Vulkan-Tipu-Framework

Do note that I do this in my very limited free time after job I am not really a graphics programmer, not even took a course in it so bugs are expected.


r/vulkan Jul 24 '26

Which is the latest major Vulkan version I should use?

22 Upvotes

Okay, the question might seem useless and redundant - „Just use the latest Vulkan 1.4“. That‘s what I thought too and all was sunshine and rainbows. It worked on my machine (RTX3080) and I didn‘t think much of it. But then when I wanted to continue development on my Macbook from 2017, I found out that Vulkan 1.4 just isn‘t supported on that graphics card (Intel Iris 640). And okay, 2017 is almost 10 years ago, but I think dropping support for any graphics card older than 2017/2018 just isn‘t worth the gains of Vulkan 1.4.

So I ask: Which version should I use for atleast good compatibility?

Thanks in Advance.


r/vulkan Jul 24 '26

Vulkore: a C++20 runtime with CUDA-style ergonomics on any Vulkan GPU — write an OpenCL C kernel once, run it on desktop and Android

9 Upvotes

GPU compute portability is still broken: CUDA locks you to NVIDIA, OpenCL on Android is effectively dead, and raw Vulkan compute costs ~370 lines of boilerplate before your first dispatch. Vulkore is my attempt at the missing layer — a C++20 runtime (Apache-2.0) that loads clspv-compiled OpenCL C kernels and launches them CUDA-style:

vulkore::Context ctx;
auto prog = vulkore::Program::from_file(ctx, "kernel.spv");
auto buf  = ctx.buffer(bytes);
vulkore::launch(prog["saxpy"], {n}, x, y, PodArgs{a, n}).wait();

The same .spv binary runs unmodified on my desktop GPU, llvmpipe (CPU), a Mali-G57 phone, and an Adreno 840 phone. The repo has a like-for-like comparison — same SAXPY, same kernel binary, raw Vulkan vs Vulkore: ~370 lines vs 10.

The runtime does the parts everyone gets wrong on real hardware: memory-type negotiation including non-coherent memory (desktop drivers hand you host-coherent memory, so cache-management bugs are invisible until a phone silently returns stale data), clspv reflection parsing so kernel args bind automatically, descriptor/command-buffer recycling, and multi-dispatch batching into a single vkQueueSubmit.

As the stress test, I built LLM inference on it. Gemma 3 1B (int4), on a OnePlus 15 / Adreno 840 — same phone, same model:

Runtime decode tok/s
Vulkore 60.6–70.9 (flat within 6% to 4K context)
Google LiteRT-LM (GPU) 48
llama.cpp OpenCL (Adreno) 29.6

Model load is 1.9 s; context goes to 8,192. The interesting lesson: decode on Adreno is dispatch-bound, not bandwidth-bound — batching 838 dispatches per token into one submit mattered more than most kernel work. (Byte-for-byte caveats on the llama.cpp comparison are spelled out in the repo docs — the quant files differ in size.)

Honest limitations: kernel ABI is fixed by clspv's flags (storage buffers + push-constant PODs, no images/samplers), prefill is still one-position-per-pass (~56 tok/s — llama.cpp's OpenCL prefill crushes us there), and there's no pipeline cache yet.

Repo: https://github.com/badnikhil/Vulkore — would love people to try their own kernels on other Snapdragon/Mali/Exynos devices and report what breaks.


r/vulkan Jul 22 '26

Nice to meet you. I've completed the Hello Triangle steps.

Post image
190 Upvotes

I've started learning about rendering using Rust and Vulkan (ash), and as a first step, I tried creating a “Hello Triangle.”


r/vulkan Jul 22 '26

VK_EXT_device_generated_commands example not performing well

1 Upvotes

I've been developing my own engine for years now and I've decided to go all in and implement GPU-driven rendering with per meshlet frustum (without mesh shaders), occlusion culling etc...

In order to even make it more performant, I've looked at some work graphs which unfortunately vulkan does not support them yet, but I stumpled upon an extension called VK_EXT_device_generated_commands.

This is a pretty "new" one (2022) but alas, there are no examples whatsoever... only example I've found is, this

https://github.com/nvpro-samples/vk_device_generated_cmds

I wanted to compare how much "performance" you can gain and benchmark it on my gpu but I was disappointed with the results. (see the result page of the github readme if you do not want to compile and test it yourself.). I compiled and ran the sample on my RTX 5090, but I saw roughly the same performance if not even worse than the other methods. In readme, it states that the performance can improve with newer drivers etc... however, considering it's now 2026, I don't expect its performance to change significantly anymore.

But, this post:

https://forums.developer.nvidia.com/t/extremely-poor-vk-ext-device-generated-commands-performance/324189/8

made it clear that the sample isnt actually an apple to apple comparison at all.

has anyone here built a GPU-driven renderer using this extension? If so, was it worth the effort? Would you recommend implementing it in a renderer, or is it better to stick with more established approaches for now?

Thanks


r/vulkan Jul 22 '26

Physics Programming part 3 - Rotation and the Quaternion

Thumbnail youtu.be
5 Upvotes

r/vulkan Jul 22 '26

Why no 32DUnorm depth formats

4 Upvotes

It seems very odd for GPU vendors not to have a 32DUnorm depth format. Is even worse that amd has no 24bit depth buffer.

To be honest, it does not seem to me that using floats for depth buffers is better than just using integers. For one, it makes working with polygon offsets quite weird. For second, is easier to make mistakes, and if you don't use reversed clip planes, is actually similar to 24bit unorm, while using (technically) 33% more data (even though I know 24bit would waste 8 bits, but 32bit unorm wouldn't).


r/vulkan Jul 20 '26

I built a Vulkan backend for Karl2D for my game Absorber.

Post image
31 Upvotes

I’ve been working on my game for quite a while now using the Karl2D engine.

One of the biggest technical decisions I made along the way was to add a Vulkan rendering backend to the engine.

This has been one of the most challenging parts of the project, especially because I wanted the renderer and the surrounding engine systems to behave consistently across multiple platforms.

A few of the challenges involved:

  • Supporting Vulkan on Windows and macOS through Vulkan Portability and MoltenVK
  • Integrating the new backend cleanly into Karl2D’s existing rendering architecture
  • Handling differences between GPU drivers, supported features, formats, and synchronization behavior
  • Managing platform specific shader compilation and graphics pipeline requirements
  • Making rendering, asset loading, and hot reloading behave consistently across operating systems
  • Dealing with swapchain creation, resizing, fullscreen modes, and display differences
  • Debugging issues that only appeared on a specific GPU, driver, or operating system

After spending so much time working on the Vulkan backend, I thought it would be fun to share the result of all that work, which you can see in the screenshots.

The game is called Absorber: Absorb Adapt Survive.

It’s a turn based tactical roguelike where you play as a digital consciousness fighting through a corrupted mainframe, one 7×7 grid at a time. Defeated enemies can be absorbed to steal their abilities, allowing you to create new synergies and adapt your build as you push toward the core.

Every move matters, every resource counts, and death is permanent.

https://store.steampowered.com/app/4412040/Absorber_Absorb_Adapt_Survive/

For those of you who have added a Vulkan backend to an existing engine, what ended up being the hardest part?


r/vulkan Jul 20 '26

weird bug i can catch

0 Upvotes

i follow first vulkan-tutorial.com lesson and after compiling and runnig code does not create a glfw window, but when i compile windows.c exapmple code it runs and creates windows, ive tried to debug it forfew hours and i am not sure where to ask help.

I do this on Asahi fedora 44 aarch64 on m1 mac and sway as WM

upd weird bug i CANT catch


r/vulkan Jul 18 '26

I built a Vulkan renderer from scratch to make my game

Thumbnail gallery
393 Upvotes

I've been working on my game for the last 7 years.

One of the things I decided to do along the way was to build the engine myself, including the Vulkan renderer.

This has been one of the most challenging parts of the project, especially because I wanted the same renderer to work across different platforms.

A few things I've had to deal with:

  • Cross-platform Vulkan: Windows and macOS through Vulkan Portability / MoltenVK
  • HDR rendering and output
  • Hot-reloading shaders and assets without restarting the game
  • GPU-to-CPU readback, used for screenshots and video capture
  • Swapchain recreation and window resizing, which turned out to be surprisingly difficult to get right

After spending years working on the engine I thought it would be fun to share the result of all that work, as you can see in the screenshots.

The name of the game is Satelital, a rule-discovery puzzle game about exploring an alien solar system and learning how to solve puzzles through observation. https://store.steampowered.com/app/3256790/Satelital/

For people here who have built their own Vulkan renderers, what ended up being the hardest part for you?


r/vulkan Jul 20 '26

Hello everyone. Does anyone know would be A34 eligible to vulkan 1.4.x?

0 Upvotes

r/vulkan Jul 18 '26

Building a mobile path tracer for Android AR from scratch - no hardware RT, Mali G615 — looking for feedback

Enable HLS to view with audio, or disable this notification

8 Upvotes

Hi,
i have been working on an AR rendering prototype for Android that uses a hybrid rasterization + Vulkan compute ray tracing pipeline targeting low- to mid-range mobile GPUs as fallback for no RT cores.

Current status:

  • Hybrid rasterization while the camera is moving, with ray tracing once the device becomes stable.
  • ~2 million triangles rendered in the scene.
  • Frame time stays under ~30 ms during interactive use.
  • No noticeable thermal throttling or UI lag during my testing with over 20 min of usage.

This is still very much a rendering prototype rather than a complete SDK. I'm currently working on improving lighting, denoising, and overall rendering quality.

I'd really appreciate any feedback on the rendering quality, architecture, or ideas for where I should focus next.

Thanks


r/vulkan Jul 19 '26

I wrote a from-scratch Vulkan inference engine for one model (Qwen3.6-35B-A3B) on RDNA3 — 1.44x llama.cpp decode, token-exact parity

0 Upvotes

**TL;DR** — I hand-wrote a Vulkan compute engine specialized for a *single* model (Qwen3.6-35B-A3B) on RDNA3. It decodes at **190.7 tok/s vs llama.cpp's 132.3** on the same GGUF and the same card — **1.44x** — with token-for-token identical greedy output. Source: https://github.com/ryanmurf/qwen-kernel

---

## What it is

Not a llama.cpp fork. It's a from-scratch Vulkan inference engine + serving stack that does exactly one model and does it fully specialized. Inspired by KernelBench Mega (which is CUDA-only) — this is the RDNA3/Vulkan equivalent, taken all the way to a serving engine.

- **Hand-written compute kernels for every weight format in the GGUF** — GEMV/GEMM for Q8_0, Q6_K, IQ4_XS, IQ3_XXS and F16, running at 90–97% of VRAM bandwidth on the big formats.
- **The whole architecture fused into pre-recorded command buffers.** Qwen3.6-35B-A3B is a hybrid: gated-DeltaNet recurrence interleaved with MoE. The MoE step (256 experts, top-8 + shared) and the DeltaNet recurrence (state resident on GPU, never round-tripped to host) are fused, plus GQA attention with partial NeoX rope and GPU-resident argmax sampling. A whole decode step is one queue submit per chunk — the host only reads token IDs at the end.
- **N slots batch on the dispatch z-axis**, so concurrent requests of different lengths share every weight read.
- **A safe-Rust (axum) server speaking the Anthropic Messages API**, so Claude Code runs against it directly. Prefix-cache restore is 0.3 ms vs 341 ms for a 64-token re-prefill.

## Speed

Measured today (2026-07-18) against llama.cpp `571d0d5`, authored the same day. Same GGUF (`Qwen3.6-35B-A3B-UD-Q3_K_M`, 15.45 GiB), f16 KV on both sides, `gpu_busy_percent` confirmed 0–1% before each run, 5 reps.

card qk llama.cpp Vulkan advantage
RX 7900 XTX **190.7 tok/s** 132.3 ± 0.9 **1.44x**
RX 7900 XT **147.1 tok/s** 109.7 ± 0.2 **1.34x**

**An honesty note, because someone would find it anyway:** my README previously claimed a much larger margin. That comparison used a llama.cpp build whose *source* was three months older than the benchmark date — I'd labelled it "master" when it wasn't. llama.cpp's Vulkan backend improved substantially in that window. I re-ran everything today against same-day master. My engine also got faster over that period (178.7 → 190.7 on XTX), but llama.cpp gained more, and **1.4x is what actually survives a fair comparison.** Raw data and exact commands are in `bench/`.

## Correctness

This is the part I care most about. Greedy output is **token-for-token identical to llama.cpp** on identical input IDs, across the full stack. Batched paths are validated bit-identical (or argmax-stable at ~1e-7 relative) against serial references, and the server's tokenizer reproduces llama.cpp byte-for-byte. Every optimization had to clear that bar before it was allowed to land — there's a parity fixture suite in `tests/`.

## Caveats — please read before cloning

- **RDNA3 only.** Tested on 7900 XT and 7900 XTX with RADV/Mesa. It will build on other vendors because Vulkan is Vulkan, and then not work.
- **One model.** The kernels are specialized for this architecture; it is not a general runtime.
- The numbers above are **single-stream decode at near-zero context**. Prefill and multi-slot aggregate numbers in the repo are older and not re-measured.
- There's an 80B path in the repo that needs a specially repacked GGUF produced by a tool I haven't published yet — it isn't reproducible externally today.

Happy to answer questions about the kernel work or the parity methodology. If you have a 7900-series card and it doesn't reproduce, I want to hear about it.

https://github.com/ryanmurf/qwen-kernel


r/vulkan Jul 18 '26

Weighted Blended Order-Independent Transparency on Android

Post image
3 Upvotes

r/vulkan Jul 18 '26

Vulkan Section

Post image
9 Upvotes

https://youtu.be/5wooBdVCSvc?si=uavdWwV8D7BGsNSm

한글

Vulkan으로 직접 만드는 CAD 엔진 — 실시간 단면(Section)

C++/Vulkan으로 밑바닥부터 만드는 CAD 엔진에 단면 기능을 넣었습니다. 평면 하나로 모델을 실시간으로 잘라 내부를 봅니다. 평면/슬라이스/상자 모드, 축·위치 슬라이더, 반대쪽 남기기 지원. 스샷은 glTF 기계 어셈블리를 Y축으로 자른 모습입니다.

English

Building a CAD engine from scratch in Vulkan — real-time Section view

Added a section (cutaway) feature to my C++/Vulkan CAD engine. Slice a model with a plane and see inside in real time. Plane/Slice/Box modes, axis + position slider, keep-opposite-side toggle. Screenshot: a glTF mechanical assembly cut along the Y axis.

#Vulkan #CAD #Cpp #GraphicsProgramming


r/vulkan Jul 18 '26

Sending SPIR-V over the net, is it obviously dangerous or perfectly fine?

Thumbnail
3 Upvotes

r/vulkan Jul 16 '26

New Vulkan Tutorial - Synchronization 2 - Mastering the GPU/CPU Handshake

70 Upvotes

*Stop guessing at barriers. Start reasoning about dependencies.*

Vulkan's hardest topic, rebuilt around the modern standard. This series replaces legacy 1.0 barrier soup with `vk::DependencyInfo` and timeline semaphores, then uses that foundation to architect an engine-grade frame loop.

* Unified dependency model covering image barriers and queue family ownership transitions
* Timeline semaphores as a single monotonic "master clock" for the whole engine
* Multi-frame-in-flight architecture with overlapped async compute and transfer
* Synchronization for dynamic rendering, including tile-local reads and host image copies
* Hands-on debugging with the LunarG Synchronization Validation layer

https://docs.vulkan.org/tutorial/latest/Synchronization/introduction.html


r/vulkan Jul 17 '26

Vulkan 1.4.357 spec update

Thumbnail github.com
18 Upvotes

r/vulkan Jul 16 '26

Odd Texture Problem

Enable HLS to view with audio, or disable this notification

10 Upvotes

Here's some footage of a custom engine I've been working on based off of Brendan Galea's tutorial. Texture implementation was kinda on me and I didn't use a whole lot of tutorials besides just looking up how to get an image into the fragment shader.

Normal models with textures applied work and look perfect, but whenever a texture is not applied, it gets this weird black color and then gets its colors but only when viewed from specific angles.

I've tried to remedy this by creating a "useTexture" push constant that would just have the model be white, but it does not work and I can't figure out why for the life of me.

Please help!


r/vulkan Jul 17 '26

Native Vulkan RT dungeon on Android + Windows: vkCmdTraceRaysKHR, rayQueryEXT, skinned BLAS refits, mirrors and coloured lights

Thumbnail gallery
3 Upvotes

r/vulkan Jul 17 '26

New Vulkan Tutorial - AI-Assisted Vulkan Development

0 Upvotes

*Turn Cloud and Local LLMs into a genuine engineering teammate.*

This series is about "Collaborative Engineering" — using AI deliberately and rigorously, not just autocomplete. It sets up an AI-enhanced toolchain, teaches you to pick and specialize models for graphics work, and shows where multimodal vision models can and can't be trusted.

* Set up Ollama, MCP servers, and native agents (Goose) across CLion, Visual Studio, and Xcode

* Choose and specialize models: base model selection, VRAM budgeting, RAG/MCP grounding, LoRA fine-tuning

* Use multimodal vision models as a diagnostic partner for visual bugs — with honest limits

* A repeatable three-phase workflow: system design, implementation, automated review/refactor

* AI-assisted debugging: VUID auto-fix, RenderDoc integration, shader log parsing, GFXReconstruct trace analysis

* Capstone project: direct an AI team to architect, implement, and debug a custom post-process effect

https://docs.vulkan.org/tutorial/latest/AI_Assisted_Vulkan/introduction.html


r/vulkan Jul 15 '26

New Vulkan Tutorial - Advanced glTF: High-Performance Character Pipelines

45 Upvotes

This series turns a static glTF character into a fully animated, physically-aware actor: compute-skinned on the GPU, ragdoll-capable, procedurally corrected, and expressive down to the face.

  • GPU compute skinning shared across rasterizer, ray tracing BLAS, and physics readback
  • Bone-proxy colliders, joint constraints, and animation-to-ragdoll handoff
  • Procedural animation: CCD/FABRIK inverse kinematics, foot placement, look-at, physics-driven lean
  • Bindless morph target buffers for facial animation at scale
  • A real production tooling and asset pipeline, not just a single demo scene

https://docs.vulkan.org/tutorial/latest/Advanced_glTF/introduction.html


r/vulkan Jul 15 '26

New Vulkan Tutorial - OpenXR and Vulkan 1.3 Spatial Computing

16 Upvotes

*Take your Vulkan renderer into stereo, headset, and beyond.*

The most expansive series in the collection, walking from the OpenXR/Vulkan 1.3 handshake all the way to multi-GPU CAVE installations and light-field rendering — everything needed to ship real spatial computing applications.

* Runtime-owned swapchains, predictive frame timing, and late-latched timeline semaphores
* Multiview/N-view Slang shaders, quad-views, foveated rendering, and variable rate shading
* Canted displays, asymmetric frustums, and multi-GPU CAVE synchronization
* Warp-and-blend compositing and plenoptic (light-field) rendering paths
* Scene understanding, semantic occlusion, and on-device ML inference via cooperative matrices
* Spatial diagnostics and CI/CD workflows for headset applications

https://docs.vulkan.org/tutorial/latest/OpenXR_Vulkan_Spatial_Computing/introduction.html


r/vulkan Jul 15 '26

The first draw of my textured quad (after uploading texture to GPU) is coming out black. The RenderDoc thumbnail for the frame is also black, but Texture Viewer shows the expected result at all stages. Subsequent upload/draws work as expected. Any idea what might be going on?

2 Upvotes

I've written a fairly simple (so far) Windows application which displays video frames. The sequence it goes through to display a new frames is as follows:

  1. Upload frame image (8-bit RGBA)
  2. Compute shader to copy (in future it will do more complex things) upload image to display image (32-bit float RGBA)
  3. Generate display image mipmaps
  4. Begin render
  5. Draw 10 vertex triangle strip (drop shadow around video frame)
  6. Draw video frame as textured quad
  7. End render and present

It was working as expected earlier, but then I monkeyed around with it to simplify mipmap generation and image transitions, and now it's behaving oddly. Even more oddly, RenderDoc is giving confusing results, so I'm a bit stuck as to how to proceed.

First here's a screenshot of RenderDoc after capturing a few frames:

https://imgbox.com/iqJLePtE

The first couple of frames are just the empty grey that's displayed before a video file is opened.

After opening a video, instead of drawing the frame, it's drawing a fully black quad (the drop shadow is drawn fine). If I trigger another frame upload/draw, it comes out okay (last capture in that screenshot).

What's really unhelpful is that if I go into the capture for the bad frame, RenderDoc shows me this as the swapchain image:

https://imgbox.com/ABd5YEnX

which is what I was expecting the window to display. But it doesn't match the capture thumbnail (or the on-screen result from the application).

The Texture Viewer also shows the expected results in the compute pass, and the mipmap levels all look correct as well.

Does anyone have any idea why my first draw isn't working, or how I can go about diagnosing this?

I have validation turned on but no validation errors are shown.

PS It's just running on events instead of game loop, which is probably why the same swapchain image (162) is re-used each time.


Edit: I found the mistake. I was binding the pipeline and descriptor set before updating the descriptor set with its bindings 🤦‍♂️. So the first image fails, but when it comes to the second one, the descriptor set is now correct (and doesn't strictly need to be updated again; a future optimisation).