vllm752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.On the GPT-OSS "Harmony" path (POST /v1/responses), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries request.cache_salt, but the tool-continuation re-submission rebuilds the engine input via tokens_input(token_ids) with no cache_salt. The continuation prefix is therefore cached in the global unsalted namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage — restoring the prompt-membership oracle that cache_salt is documented to prevent.
Silently dropping a preserved salt after the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.
This is distinct from GHSA-4qjh-9fv9-r85r (CVE-2025-46570): that advisory is the prefix-cache membership oracle for which cache_salt is the documented mitigation, and its PR-17045 fix does not close this site — the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact cached_tokens_per_turn counts from a different sink (the Responses serving continuation, not general TTFT timing).
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
vllm/entrypoints/openai/responses/serving.py#L712-L713 — token_ids = context.render_for_completion() then engine_input = tokens_input(token_ids), with no cache_salt.vllm/entrypoints/openai/responses/serving.py#L755 — tokens_input(prompt_token_ids, cache_salt=request.cache_salt).tokens_input stores the salt only if passed: vllm/inputs/engine.py#L51-L66 (if cache_salt is not None: inputs["cache_salt"] = cache_salt).vllm/v1/engine/input_processor.py#L380 (cache_salt=decoder_inputs.get("cache_salt") → None for the continuation).vllm/v1/core/kv_cache_utils.py#L560-L561 ([request.cache_salt] if (start_token_idx == 0 and request.cache_salt) else []).vllm/entrypoints/openai/responses/serving.py#L909 (cached_tokens_per_turn).vllm/entrypoints/openai/responses/protocol.py#L235 (cache_salt field).The tool-continuation re-submission rebuilds the engine input with no cache_salt:
# vllm/entrypoints/openai/responses/serving.py Lines 711-715
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
engine_input = tokens_input(token_ids)
sampling_params.max_tokens = max_model_len - len(token_ids)
Contrast with the correct turn-1 call, which does preserve the caller's salt:
# vllm/entrypoints/openai/responses/serving.py Lines 754-755
prompt_token_ids = render_for_completion(messages)
engine_input = tokens_input(prompt_token_ids, cache_salt=request.cache_salt)
tokens_input stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:
# vllm/inputs/engine.py Lines 51-66
def tokens_input(
prompt_token_ids: list[int],
*,
prompt: str | None = None,
cache_salt: str | None = None,
) -> TokensInput:
"""
Construct [`TokensInput`][vllm.inputs.engine.TokensInput]
from optional values.
"""
inputs = TokensInput(type="token", prompt_token_ids=prompt_token_ids)
if prompt is not None:
inputs["prompt"] = prompt
if cache_salt is not None:
inputs["cache_salt"] = cache_salt
An authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency — the exact prompt-membership oracle cache_salt is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.
Preconditions: a GPT-OSS Harmony model on /v1/responses; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets cache_salt and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The AC:H metric reflects that guessable-history precondition.
Propagate request.cache_salt into every Harmony (and Parsable) tool-continuation re-submission — at the continuation call site call tokens_input(token_ids, cache_salt=request.cache_salt), mirroring the correct turn-1 call. Carry the salt on the HarmonyContext (thread the originating request into the context) so no continuation path can omit it:
# vllm/entrypoints/openai/responses/serving.py
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
- engine_input = tokens_input(token_ids)
+ engine_input = tokens_input(
+ token_ids,
+ cache_salt=(
+ context.request.cache_salt
+ if context.request is not None
+ else None
+ ),
+ )
with HarmonyContext.__init__ gaining a request: ResponsesRequest | None = None parameter (stored as self.request) that _create_responses passes when constructing the context. The continuation prefix is then cached in the victim's salted namespace, mirroring turn 1.
Suggested regression test: assert cached_tokens_per_turn == 0 for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818
{
"cwe_ids": [
"CWE-200",
"CWE-524"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-06T00:02:02Z",
"nvd_published_at": null,
"severity": "LOW"
}