vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.
{
"cwe_ids": [
"CWE-200",
"CWE-681"
],
"osv_generated_from": "https://github.com/CVEProject/cvelistV5/tree/main/cves/2026/53xxx/CVE-2026-53923.json",
"cna_assigner": "GitHub_M"
}"2026-08-12T16:41:10Z"
[
{
"id": "CVE-2026-53923-f2bf7ca8",
"deprecated": false,
"signature_type": "Line",
"signature_version": "v1",
"digest": {
"threshold": 0.9,
"line_hashes": [
"118265656872246360285709774639688387291",
"61465877030539699673395513370793171388",
"272548032624709417653722042097387448651",
"319359192550244661342327622333718076226"
]
},
"source": "https://github.com/vllm-project/vllm/commit/f219788f91952827132fa4fdf916427cd20d225e",
"target": {
"file": "csrc/libtorch_stable/quantization/gguf/ggml-common.h"
}
}
]
"https://storage.googleapis.com/cve-osv-conversion/osv-output/CVE-2026-53923.json"