vllm752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches.vLLM's disaggregated scale-out transport splits a multimodal request into a trusted render step (POST /v1/chat/completions/render) and a separate generate step (POST /inference/v1/generate). The generate route decodes a caller-supplied features object — serialized encoder tensors (kwargs_data), multimodal hashes (mm_hashes), placeholder ranges (mm_placeholders), and the internal field-processor selection — and forwards it into the engine as if it had come from the trusted renderer, with no rebinding to (or validation against) the active model's renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the /inference prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field.
Depending on which field is forged, this produces:
ValueError, and a hard assert — three independent forged fields (sites 1, 2, 3);All five share one root cause and one fix shape: the reconstructed multimodal state on the scale-out path is trusted without being rebound to, and validated against, the active model's renderer output before it reaches the engine.
These sites are distinct from prior multimodal hardening. Site 1 survives GHSA-wv77-2vpf-vmmg (that fix validates full tensor shape in MultiModalDataParser/get_input_embeddings on the prompt-embeds path), because our request forges image_grid_thw metadata with the pixel bytes intact and reaches the Qwen2 vision RoPE/cu_seqlens and image_embeds.split sink, which the shape-check fix does not rebind. Site 4 is distinct from GHSA-c65p-x677-fgj6 (which folds metadata into MultiModalHasher.serialize_item to stop hash collisions), because the scale-out generate path trusts a caller-supplied mm_hash as the cache key with no origin binding, so that fix does not stop a caller from submitting a victim's hash or a kwargs_data=None cache read.
Links pinned to the confirmed commit 752a3a504485 (v0.25.1).
Shared entry point and control surface for all five sites:
POST /inference/v1/generate route: vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75.features schema (kwargs_data, mm_hashes, mm_placeholders): vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63.ServingTokens.serve_tokens() copies the decoded geometry and hashes into engine structures without rebinding to the renderer schema: vllm/entrypoints/scale_out/token_in_token_out/serving.py#L145-L172./inference is treated as an ordinary auth-guarded prefix: vllm/entrypoints/openai/api_server.py#L217-L219.Site 1 — forged Qwen grid geometry (engine-fatal DoS). The decoded image_grid_thw is never rebound to the rendered pixel-tensor element count.
serving.py#L147-L172.Qwen2VisionTransformer.prepare_encoder_metadata() derives RoPE tables, cu_seqlens, and the FlashAttention max_seqlen from the caller-supplied grid: qwen2_vl.py#L658-L712, consumed in forward() (#L713-L751)._process_image_input() computes image_embeds.split(sizes) from the same untrusted metadata: qwen2_vl.py#L1348-L1369; M-RoPE positions at qwen2_vl.py#L1223.The only shape check is assert grid_thw.ndim == 2; the split sizes and the vision-encoder call are then derived directly from the caller-supplied grid, with no cross-check against the pixel-tensor row count:
# vllm/model_executor/models/qwen2_vl.py Lines 1348-1369
def _process_image_input(
self, image_input: Qwen2VLImageInputs
) -> tuple[torch.Tensor, ...]:
grid_thw = image_input["image_grid_thw"]
assert grid_thw.ndim == 2
if image_input["type"] == "image_embeds":
image_embeds = image_input["image_embeds"]
else:
pixel_values = image_input["pixel_values"]
if self.use_data_parallel:
return run_dp_sharded_mrope_vision_model(
self.visual, pixel_values, grid_thw.tolist(), rope_type="rope_3d"
)
else:
image_embeds = self.visual(pixel_values, grid_thw=grid_thw)
# Split concatenated embeddings for each image item.
merge_size = self.visual.spatial_merge_size
sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()
return image_embeds.split(sizes)
Site 2 — wire-selected field-processor type confusion (engine-fatal DoS). MsgpackDecoder._decode_mm_field_elem() trusts a wire-selected field-factory name and constructs the internal field processor directly from caller data.
vllm/v1/serial_utils.py#L440-L454 reads factory_meth_name, factory_kw = obj["field"] and calls getattr(MultiModalFieldConfig, factory_meth_name).qwen2_vl.py#L763-L791 and Qwen2VLImagePixelInputs (#L119-L144); parse/validate at qwen2_vl.py#L1300.vllm/utils/tensor_schema.py#L155-L171, reached from TensorSchema.__init__ → validate() (#L63).Site 3 — non-positive placeholder length (engine-fatal DoS via reachable assert). PlaceholderRangeInfo{offset,length} is accepted as unconstrained integers and copied verbatim into the engine's PlaceholderRange.
protocol.py#L28-L35.PlaceholderRange: serving.py#L147-L153.num_embeds > mm_encoder_cache_size), with no non-positive check: vllm/v1/engine/input_processor.py#L459-L464; raw length returned by vllm/multimodal/inputs.py#L152-L154; window selection assumes non-empty ranges at vllm/multimodal/utils.py#L114-L134.assert start_idx < end_idx at vllm/v1/worker/gpu/mm/encoder_runner.py#L114 (duplicated at vllm/v1/worker/gpu_model_runner.py#L3192); turned into a fatal shutdown by EngineCore's uncaught-exception path at vllm/v1/engine/core.py#L1229-L1233.A length of 0 makes num_encoder_tokens == 0, so end_idx collapses to 0 and the bare assert fires inside the engine worker:
# vllm/v1/worker/gpu/mm/encoder_runner.py Lines 108-114
pos_info = mm_feature.mm_position
start_pos = pos_info.offset
num_encoder_tokens = pos_info.length
start_idx = max(cur_query_start - start_pos, 0)
end_idx = min(cur_query_end - start_pos, num_encoder_tokens)
assert start_idx < end_idx
Site 4 — cache hash not bound to payload (integrity / disclosure). features.mm_hashes (the cache key) and kwargs_data (the tensor) are independent fields with no origin or integrity binding.
None = resolve-from-cache semantics: protocol.py#L42-L63.mm_input(...): serving.py#L164-L170.vllm/v1/engine/input_processor.py#L165-L181.cache_key = feature.mm_hash or feature.identifier): vllm/multimodal/cache.py#L602-L607.The cache key is the caller-supplied hash with no verification against the tensor bytes, so a forged mm_hash both stores under and reads back another request's slot:
# vllm/multimodal/cache.py Lines 601-607
for feature in mm_features:
cache_key = feature.mm_hash or feature.identifier
self.touch_receiver_cache_item(cache_key, feature.data)
for feature in mm_features:
cache_key = feature.mm_hash or feature.identifier
feature.data = self.get_and_update_item(feature.data, cache_key)
return mm_features
Site 5 — dropped sparse placeholder mask (transport integrity loss). The render path serializes placeholders as only offset/length, so models relying on sparse is_embed masks lose the mask during render-to-generate replay.
ServingRender._extract_mm_features() builds each PlaceholderRangeInfo(offset=p.offset, length=p.length), discarding is_embed: vllm/entrypoints/scale_out/render/serving.py#L212-L229.protocol.py#L28.ServingTokens.serve_tokens() reconstructs a dense PlaceholderRange regardless of the original: serving.py#L148-L152.A single authenticated request to a scale-out multimodal deployment can:
/health → 503). Availability-only; no code execution or data disclosure demonstrated for these sites.is_embed masks.The forged multimodal payload is small; only the trust in its self-declared geometry/identity is the defect.
On the scale-out path, do not trust caller-supplied multimodal state as renderer-produced. After decoding features, rebind and validate the reconstructed MultiModalKwargsItem against the active model's renderer contract at the HTTP boundary:
prepare_encoder_metadata() (site 1).PlaceholderRangeInfo with length <= 0 (or out-of-range offset) with a request-scoped 4xx, and convert the encoder-runner invariant into a checked, request-scoped error rather than a process-fatal assert (site 3).kwargs_data before using it as a cache key, and namespace receiver-cache keys to a server-generated or principal scope, refusing cache-reads for hashes the caller did not legitimately produce (site 4).is_embed in PlaceholderRangeInfo, validate its length against the placeholder span, and reconstruct it on replay (site 5).Site 1 — validate grid geometry before the vision encoder. Replace the bare assert grid_thw.ndim == 2 in _process_image_input()/_process_video_input() with a shared helper that recomputes the split sizes and rejects a grid whose patch-row count does not match the pixel tensor (and rejects non-positive / non-merge-divisible dims), so the mismatch never reaches image_embeds.split():
# vllm/model_executor/models/qwen2_vl.py — _process_image_input()
- grid_thw = image_input["image_grid_thw"]
- assert grid_thw.ndim == 2
+ grid_thw = image_input["image_grid_thw"]
+ input_type = image_input["type"]
+ input_tensor = (
+ image_input["image_embeds"]
+ if input_type == "image_embeds"
+ else image_input["pixel_values"]
+ )
+ sizes = _validate_qwen2_vl_input_geometry(
+ modality="image",
+ input_type=input_type,
+ input_tensor=input_tensor,
+ grid_thw=grid_thw,
+ spatial_merge_size=self.visual.spatial_merge_size,
+ )
...
- # Split concatenated embeddings for each image item.
- merge_size = self.visual.spatial_merge_size
- sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()
return image_embeds.split(sizes)
where the helper raises before the encoder runs:
# vllm/model_executor/models/qwen2_vl.py — new _validate_qwen2_vl_input_geometry()
if t <= 0 or h <= 0 or w <= 0:
raise ValueError(f"{modality} grid_thw row {index} must be positive ...")
if h % spatial_merge_size != 0 or w % spatial_merge_size != 0:
raise ValueError(f"{modality} grid_thw row {index} must be divisible ...")
...
if actual_rows != expected_rows:
raise ValueError(
f"{modality} {row_kind} do not match grid_thw: "
f"expected {expected_rows}, got {actual_rows}."
)
Site 3 — constrain the placeholder schema. Make PlaceholderRangeInfo reject non-positive lengths and negative offsets at the Pydantic boundary (plus parallel-length, non-overlapping, and within-prompt validators), turning the process-fatal assert into a request-scoped 422:
# vllm/entrypoints/scale_out/token_in_token_out/protocol.py — PlaceholderRangeInfo
- offset: int
- length: int
+ offset: int = Field(ge=0)
+ length: int = Field(gt=0)
Site 4 — bind the cache key to the payload. Derive each cache key by hashing the submitted serialized tensor (ignoring the caller's mm_hashes) and refuse cache-only reads, so a forged hash can neither poison nor read a victim's slot:
# vllm/entrypoints/scale_out/token_in_token_out/serving.py — serve_tokens()
+ mm_hashes = _bind_mm_hashes_to_kwargs_data(
+ features.mm_hashes, features.kwargs_data,
+ )
engine_input = mm_input(
prompt_token_ids=request.token_ids,
mm_kwargs=MultiModalKwargsItems(mm_kwargs),
- mm_hashes=features.mm_hashes,
+ mm_hashes=mm_hashes,
mm_placeholders=mm_placeholders,
cache_salt=request.cache_salt,
)
where _bind_mm_hashes_to_kwargs_data() raises on kwargs_data is None (cache-only read) and derives sha256(modality || "\0" || serialized_item) per item. Site 2 applies the same rebind-to-declared-schema pattern in mm_serde.py (passing modality + mm_processor into decode_mm_kwargs_item), and site 5 adds an is_embed field to PlaceholderRangeInfo with from_placeholder_range/to_placeholder_range helpers so the sparse mask survives render-to-generate replay. Each site fix ships with a regression test. This packet groups the five sites because they share one entry point (/inference/v1/generate + /render) and one root cause; we are happy to split it into per-component advisories (for example, engine-fatal input-validation vs. cache-key binding vs. transport-schema integrity) if the vLLM team prefers.
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
These vulnerabilities were discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51898
{
"cwe_ids": [
"CWE-1284",
"CWE-20",
"CWE-617",
"CWE-639",
"CWE-668",
"CWE-704"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-05T23:42:50Z",
"nvd_published_at": null,
"severity": "MODERATE"
}