llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and access ctx_server.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when --sleep-idle-seconds is configured.
{
"binaries": [
{
"binary_name": "libllama0",
"binary_version": "8681+dfsg-1"
},
{
"binary_name": "llama.cpp",
"binary_version": "8681+dfsg-1"
},
{
"binary_name": "llama.cpp-examples",
"binary_version": "8681+dfsg-1"
},
{
"binary_name": "llama.cpp-tests",
"binary_version": "8681+dfsg-1"
},
{
"binary_name": "llama.cpp-tools",
"binary_version": "8681+dfsg-1"
},
{
"binary_name": "llama.cpp-tools-extra",
"binary_version": "8681+dfsg-1"
},
{
"binary_name": "python3-gguf",
"binary_version": "8681+dfsg-1"
}
]
}