llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/counttokens) that bypass the task queue and access ctxserver.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when --sleep-idle-seconds is configured.
{
"binaries": [
{
"binary_version": "8681+dfsg-1",
"binary_name": "libllama0"
},
{
"binary_version": "8681+dfsg-1",
"binary_name": "llama.cpp"
},
{
"binary_version": "8681+dfsg-1",
"binary_name": "llama.cpp-examples"
},
{
"binary_version": "8681+dfsg-1",
"binary_name": "llama.cpp-tests"
},
{
"binary_version": "8681+dfsg-1",
"binary_name": "llama.cpp-tools"
},
{
"binary_version": "8681+dfsg-1",
"binary_name": "llama.cpp-tools-extra"
},
{
"binary_version": "8681+dfsg-1",
"binary_name": "python3-gguf"
}
]
}