Files
llama.cpp/ggml/src
Radoslav Gerganov 0c58ba3365 rpc : reuse compute graph buffers (#21299)
Reuse the buffer for the ggml context which is used for creating the
compute graph on the server side. This partially addresses a memory leak
created by the CUDA backend due to using buffer addresses as cache
keys.

ref: #21265
ref: #20315
2026-04-03 10:28:09 +03:00
..
2026-03-25 19:57:40 +01:00
2026-03-25 12:53:16 +02:00