{"id":"CVE-2026-93436","title":"vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments","summary":"vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory witho…","severity":"high","cvss":7.5,"cvssVector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401","CWE-770"],"vendor":"vllm-project","product":"vllm","affected":["vllm <= 0.29.0"],"published":"2026-09-17","updated":"2026-09-22","sourceUpdated":"2026-09-22T20:25:55.870","source":"NVD","sourceUrl":"https://nvd.nist.gov/vuln/detail/CVE-2026-93436","references":[{"url":"https://github.com/vllm-project/vllm","label":"disclosure@vulncheck.com"},{"url":"https://github.com/vllm-project/vllm/blob/v0.29.0/vllm/distributed/kv_transfer/kv_connector/v1/nixl/push_worker.py#L162-L181","label":"disclosure@vulncheck.com"},{"url":"https://github.com/vllm-project/vllm/pull/55677","label":"disclosure@vulncheck.com"},{"url":"https://www.vulncheck.com/advisories/vllm-through-0.29.0-memory-exhaustion-via-rejected-requests","label":"disclosure@vulncheck.com"},{"url":"https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-93436.json"},{"url":"https://access.redhat.com/security/cve/CVE-2026-93436"},{"url":"https://bugzilla.redhat.com/show_bug.cgi?id=2538717"},{"url":"https://www.cve.org/CVERecord?id=CVE-2026-93436"},{"url":"https://nvd.nist.gov/vuln/detail/CVE-2026-93436"}],"tags":["nvd","cve.org","csaf","vex","red-hat"],"epss":0.00538,"epssPercentile":0.44214,"ssvc":{"exploitation":"none","automatable":"yes","technicalImpact":"partial","timestamp":"2026-09-22T01:58:51.466543Z"},"ingestedAt":"2026-09-17T22:30:21.417Z","slug":"CVE-2026-93436","body":"## Overview\n\nvLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.\n\n## Remediation\n\nRefer to the linked advisories for vendor-supplied fixes and affected version ranges.\n\n## Vendor advisories\n\n- **Red Hat VEX** · Important · affected: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · no fix planned: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · updated 2026-09-22 · [vex](https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-93436.json)","depth":"twilight","depthScore":41,"depthScoreParts":{"impact":41.3,"likelihood":0.1,"exploitation":0,"ransomware":0},"changes":[]}