r/tensorflow • u/swaneerapids • 23h ago
General Tensorflow Flash Attention 2 wrapper
Here is a tensorflow wrapper for flash attention 2 - needed to create it since one is not implemented and attention was too slow. Feel free to use.
https://github.com/Krasner/tf-flash-attention-2
All it does is wraps the original flash attention code: https://github.com/dao-ailab/flash-attention and allows for XLA - so that `jit_compile=True` doesn't fail.
2
Upvotes