vllm752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.On late-interaction /score and /rerank deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the caller-controlled X-Request-Id header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring — so the victim's documents are scored against the attacker's query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's X-Request-Id. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.
This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
vllm/entrypoints/serve/engine/serving.py#L117-L124 — _base_request_id() copies the public X-Request-Id header directly.vllm/entrypoints/pooling/base/serving.py#L109 — the frontend request id is f"{self.request_id_prefix}-{self._base_request_id(raw_request)}".vllm/entrypoints/pooling/scoring/serving.py#L211 — flash_late_interaction() (at L191) derives worker cache keys directly from that id: query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)].vllm/v1/pool/late_interaction.py#L30-L36 — the data-parallel routing helper pins all requests sharing a query_key to the same engine via crc32(query_key), making collisions deterministic.vllm/v1/worker/gpu/pool/late_interaction_runner.py#L95 — the worker stores query embeddings in a process-local cache keyed only by that string: self._query_cache[query_key] = output.clone().The caller-controlled header enters as the request id, and the flash late-interaction path derives the worker cache key directly from it:
# vllm/entrypoints/serve/engine/serving.py Lines 116-126
@staticmethod
def _base_request_id(
raw_request: Request | None, default: str | None = None
) -> str | None:
"""Pulls the request id to use from a header, if provided"""
if raw_request is not None and (
(req_id := raw_request.headers.get("X-Request-Id")) is not None
):
return req_id
return random_uuid() if default is None else default
# vllm/entrypoints/pooling/scoring/serving.py Lines 207-212
n_queries = ctx.n_queries
n_docs = len(ctx.engine_inputs) - n_queries
query_engine_inputs = ctx.engine_inputs[:n_queries]
query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
query_uses = [n_docs if n_queries == 1 else 1] * n_queries
The worker then stores and reads the query embedding under that string with no check that the reader owns the entry — a colliding key returns another request's cached query, or (once the use counter is exhausted) raises a cache-miss error:
# vllm/v1/worker/gpu/pool/late_interaction_runner.py Lines 91-107
if mode == LATE_INTERACTION_MODE_CACHE_QUERY:
assert query_uses is not None
# `output` can be a view into the current step's hidden-states
# buffer, so clone it before storing across scheduling steps.
self._query_cache[query_key] = output.clone()
self._query_uses[query_key] = query_uses
outputs[i] = torch.zeros((), device=output.device, dtype=torch.float32)
continue
if mode == LATE_INTERACTION_MODE_SCORE_DOC:
query_output = self._query_cache.get(query_key)
if query_output is None:
raise ValueError(
"late-interaction query cache miss for key "
f"{query_key!r}. Ensure query requests are executed "
"before their paired document requests."
)
The bug is specific to the flash late-interaction path. Non-flash late-interaction scoring computes MaxSim directly from one request's in-memory outputs and does not create a cross-request worker cache key.
A network client of the standard scoring API can, on a flash late-interaction /score or /rerank deployment:
X-Request-Id, the attacker's query embedding overwrites the victim's cached entry, so the victim's documents are scored against the attacker's query (a cross-request integrity break).Both consequences follow deterministically from reusing the victim's header value, because same-key work is pinned to one engine. This is reachable through normal request handling and does not depend on any trusted inter-node network.
Derive the flash late-interaction query-cache key from a server-generated, unforgeable per-request identifier (a random_uuid() namespace) rather than the caller-supplied X-Request-Id, and thread that key through the PoolingServeContext to the doc-scoring pass so both passes reuse the same key and a caller cannot address another request's cache entry.
In vllm/entrypoints/pooling/scoring/serving.py, the encode-queries pass mints a fresh namespace and stashes the keys on the context:
- query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+ query_namespace = random_uuid()
+ query_keys = [
+ f"late-interaction-{query_namespace}-query-{i}" for i in range(n_queries)
+ ]
+ ctx.late_interaction_query_keys = query_keys
and the encode-docs pass reads those stored keys instead of re-deriving them from ctx.request_id:
- query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+ query_keys = ctx.late_interaction_query_keys
+ if query_keys is None:
+ raise RuntimeError("Late-interaction query keys were not initialized.")
This requires adding the late_interaction_query_keys: list[str] | None = None field to PoolingServeContext (vllm/entrypoints/pooling/typing.py). Because the namespace is a server-generated UUID, colliding X-Request-Id values no longer produce a shared cache key; a regression test asserting exactly that (colliding request ids yield distinct query-cache keys) accompanies the change.
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51445
{
"cwe_ids": [
"CWE-639"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-05T23:42:39Z",
"nvd_published_at": null,
"severity": "MODERATE"
}