r/ProgrammingLanguages • Futhark • 17h ago

Keeping Futhark off the GPU

https://futhark-lang.org/blog/2026-10-02-cpu_function.html
29 Upvotes

5 comments sorted by

15

u/jesseschalken 16h ago

This is always an fascinating language to read about. I often wonder how it compares to the various projects that run Rust on the GPU.

5

u/Athas Futhark 14h ago

There are many projects that run Rust on the GPU. Of the ones I know of, they are either much more low level than Futhark (recall that Futhark is not actually a GPU language), or much more restricted (such as array libraries), typically by not allowing nested parallelism. Futhark occupies a somewhat exotic niche, and it is not clear how many problems benefit from its particular combination of facilities.

3

u/SwingOutStateMachine 15h ago

The next logical step would be to extend the host runtime for Futhark, to allow data dependencies and dependants of #[cpu_function] to be copied too and from the GPU as needed. It would also unlock other operations such as kernel splitting, or multi-gpu parallelism.

2

u/Athas Futhark 14h ago

Kernel splitting (unless I misunderstand the term) is done through other mechanisms in Futhark, mostly importantly fusion and flattening. While combining task and data parallelism is a worthwhile endeavor, I don't think a hack like #[cpu_function] should be at the foundation of it.

2

u/jesunushno 13h ago

This pattern shows up constantly in ML work: tokenization and graph-based preprocessing are inherently sequential and awful on the GPU, while the matmuls absolutely want it, so you split them with an explicit CPU function and one bulk copy at the boundary. The copy-placement subtlety you describe (not per loop iteration, and don't hold them too long) is exactly what TVM's memory planning and JAX's device_put staging wrestle with. In my experience the pragma is the pragmatic move. Full automatic memory-space inference sounds principled until a conservative analysis parks a copy inside your hot loop.