Files
llama.cpp/ggml/src/ggml-cuda
074944998d ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_argsort() to reduce temporary buffers memory usage (#24776)
* ggml : process data in smaller chunks in CUDA ggml_top_k() implementation to reduce temporary buffers memory usage

* ggml : allocate tmp_dst only only once before the loop

* chore : whitespaces

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* ggml : use chunked processing in both CUDA CUB top-k and argsort implementations

* chore : separate argsort_f32_i32_cuda_bitonic() call from return statement

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* chore : replace ternary operators with min/max

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
2026-07-09 20:07:12 +02:00
..
2025-08-20 10:17:37 +08:00
2025-08-05 22:10:36 +03:00
2026-04-10 10:24:09 +08:00
2026-06-18 22:23:01 +02:00
2026-06-18 22:23:01 +02:00
2025-06-20 09:50:24 +08:00
2025-06-20 09:50:24 +08:00
2025-08-28 20:33:03 +02:00
2025-12-09 20:28:57 +01:00
2025-12-09 20:28:57 +01:00
2025-12-08 21:10:12 +08:00
2026-07-09 17:56:32 +03:00
2025-06-22 12:39:54 +08:00
2026-01-29 11:10:53 +01:00
2025-07-29 14:45:18 +08:00
2025-07-29 14:45:18 +08:00
2025-11-13 08:50:01 +08:00
2025-07-29 14:22:03 +02:00
2025-03-31 18:05:13 +02:00
2025-06-22 12:39:54 +08:00
2026-04-23 10:28:56 +08:00
2025-11-30 21:57:31 +01:00