Files
llama.cpp/ggml/src/ggml-vulkan/vulkan-shaders
Jeff Bolz 6efcd65945 vulkan: optimize flash attention split_k_reduce (#14554)
* vulkan: allow FA split_k with smaller KV values

* vulkan: spread split_k_reduce work across more threads

k_num can get rather large. Use the whole workgroup to reduce the M/L values.

Launch a thread for each element in the HSV dimension of the output. Helps a
lot for large HSV (like deepseek).
2025-07-08 20:11:42 +02:00
..
2025-05-02 20:54:30 +03:00
2025-07-01 10:14:21 +02:00
2025-02-28 07:52:51 +01:00