vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
{
"cna_assigner": "VulnCheck",
"cwe_ids": [
"CWE-401"
],
"osv_generated_from": "https://github.com/CVEProject/cvelistV5/tree/main/cves/2026/93xxx/CVE-2026-93436.json"
}{
"extracted_events": [
{
"introduced": "0"
},
{
"last_affected": "0.29.0"
},
{
"fixed": "0.29.0"
}
],
"source": [
"AFFECTED_FIELD",
"DESCRIPTION"
]
}