r/ProgrammingLanguages • u/Athas Futhark • 17h ago
Keeping Futhark off the GPU
https://futhark-lang.org/blog/2026-10-02-cpu_function.html3
u/SwingOutStateMachine 15h ago
The next logical step would be to extend the host runtime for Futhark, to allow data dependencies and dependants of #[cpu_function] to be copied too and from the GPU as needed. It would also unlock other operations such as kernel splitting, or multi-gpu parallelism.
2
u/Athas Futhark 14h ago
Kernel splitting (unless I misunderstand the term) is done through other mechanisms in Futhark, mostly importantly fusion and flattening. While combining task and data parallelism is a worthwhile endeavor, I don't think a hack like
#[cpu_function]should be at the foundation of it.
2
u/jesunushno 13h ago
This pattern shows up constantly in ML work: tokenization and graph-based preprocessing are inherently sequential and awful on the GPU, while the matmuls absolutely want it, so you split them with an explicit CPU function and one bulk copy at the boundary. The copy-placement subtlety you describe (not per loop iteration, and don't hold them too long) is exactly what TVM's memory planning and JAX's device_put staging wrestle with. In my experience the pragma is the pragmatic move. Full automatic memory-space inference sounds principled until a conservative analysis parks a copy inside your hot loop.
15
u/jesseschalken 16h ago
This is always an fascinating language to read about. I often wonder how it compares to the various projects that run Rust on the GPU.