{"id":"CVE-2026-54234","title":"vllm: vLLM: Denial of Service via malformed speculative decoding workload (CVE-2026-54234)","summary":"A flaw was found in vLLM, a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). A remote attacker can exploit this vulnerability by sending a specially crafted multi-request speculative decod…","severity":"high","cvss":7.5,"cvssVector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cvssSource":"vendor","cwe":["CWE-125","CWE-20"],"vendor":"Red Hat","product":"Red Hat AI Inference Server 3.4","affected":["ai_inference_server","enterprise_linux_ai_rhel_ai 3","openshift_ai_rhoai","ai_inference_server 3.2","ai_inference_server 3.3","ai_inference_server 3.4"],"patched":["ai_inference_server 3.2","ai_inference_server 3.3","ai_inference_server 3.4"],"published":"2026-07-06","updated":"2026-09-24","sourceUpdated":"2026-09-24T05:56:03+00:00","source":"CSAF","sourceUrl":"https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-54234.json","references":[{"url":"https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-54234.json"},{"url":"https://access.redhat.com/security/cve/CVE-2026-54234"},{"url":"https://bugzilla.redhat.com/show_bug.cgi?id=2497515"},{"url":"https://www.cve.org/CVERecord?id=CVE-2026-54234"},{"url":"https://nvd.nist.gov/vuln/detail/CVE-2026-54234"},{"url":"https://github.com/vllm-project/vllm/commit/8a5cf1ccd65e8ac7635c402c1ec0b08988bc26ca"},{"url":"https://github.com/vllm-project/vllm/pull/44744"},{"url":"https://github.com/vllm-project/vllm/security/advisories/GHSA-8wr5-jm2h-8r4f"},{"url":"https://access.redhat.com/errata/RHSA-2026:61627"},{"url":"https://access.redhat.com/errata/RHSA-2026:61629"},{"url":"https://access.redhat.com/errata/RHSA-2026:60363"},{"url":"https://access.redhat.com/errata/RHSA-2026:69466"},{"url":"https://access.redhat.com/errata/RHSA-2026:70965"},{"url":"https://access.redhat.com/errata/RHSA-2026:70979"},{"url":"https://access.redhat.com/errata/RHSA-2026:69467"},{"url":"https://access.redhat.com/errata/RHSA-2026:70995"},{"url":"https://access.redhat.com/errata/RHSA-2026:69469"},{"url":"https://access.redhat.com/errata/RHSA-2026:70969"},{"url":"https://access.redhat.com/errata/RHSA-2026:69464"},{"url":"https://github.com/vllm-project/vllm"},{"url":"https://github.com/advisories/GHSA-8wr5-jm2h-8r4f"}],"tags":["csaf","vex","red-hat","osv","pip","ghsa"],"epss":0.00616,"epssPercentile":0.47228,"aliases":["GHSA-8wr5-jm2h-8r4f","PYSEC-2026-3542"],"ecosystem":"pip","ingestedAt":"2026-07-17T17:13:27.259Z","slug":"CVE-2026-54234","body":"## Overview\n\nA flaw was found in vLLM, a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). A remote attacker can exploit this vulnerability by sending a specially crafted multi-request speculative decoding workload through public gRPC Generate and Abort endpoints. This malformed workload can cause the rejection sampler to produce an out-of-vocabulary token, which then crashes the engine worker. This leads to a service-wide Denial of Service (DoS) for all clients until the worker is restarted.\n\n## Vendor advisories\n\n- **RHSA-2026:61627** · Red Hat · fixed in: Red Hat AI Inference Server 3.2 · released 2026-08-31 · [advisory](https://access.redhat.com/errata/RHSA-2026:61627)\n- **RHSA-2026:61629** · Red Hat · fixed in: Red Hat AI Inference Server 3.2 · released 2026-08-31 · [advisory](https://access.redhat.com/errata/RHSA-2026:61629)\n- **RHSA-2026:60363** · Red Hat · fixed in: Red Hat AI Inference Server 3.3 · released 2026-08-26 · [advisory](https://access.redhat.com/errata/RHSA-2026:60363)\n- **RHSA-2026:69466** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-21 · [advisory](https://access.redhat.com/errata/RHSA-2026:69466)\n- **RHSA-2026:70965** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-23 · [advisory](https://access.redhat.com/errata/RHSA-2026:70965)\n- **RHSA-2026:70979** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-23 · [advisory](https://access.redhat.com/errata/RHSA-2026:70979)\n- **RHSA-2026:69467** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-21 · [advisory](https://access.redhat.com/errata/RHSA-2026:69467)\n- **RHSA-2026:70995** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-23 · [advisory](https://access.redhat.com/errata/RHSA-2026:70995)\n- **RHSA-2026:69469** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-21 · [advisory](https://access.redhat.com/errata/RHSA-2026:69469)\n- **RHSA-2026:70969** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-23 · [advisory](https://access.redhat.com/errata/RHSA-2026:70969)\n- **RHSA-2026:69464** · Red Hat · fixed in: Red Hat AI Inference Server 3.4 · released 2026-09-21 · [advisory](https://access.redhat.com/errata/RHSA-2026:69464)\n- **Red Hat VEX** · Important · affected: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · no fix planned: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · updated 2026-09-24 · [vex](https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-54234.json)\n\n**vllm: vLLM: Denial of Service via malformed speculative decoding workload** — rated Important by Red Hat. Released 2026-07-06, updated 2026-09-24.\n\nAffected:\n\n- Red Hat AI Inference Server\n- Red Hat Enterprise Linux AI (RHEL AI) 3\n- Red Hat OpenShift AI (RHOAI)\n\nFixed:\n\n- Red Hat AI Inference Server 3.2\n- Red Hat AI Inference Server 3.3\n- Red Hat AI Inference Server 3.4\n\nNo fix planned:\n\n- Red Hat AI Inference Server\n- Red Hat Enterprise Linux AI (RHEL AI) 3\n- Red Hat OpenShift AI (RHOAI)\n\nNot affected:\n\n- Red Hat OpenShift AI (RHOAI)\n\n## Remediation\n\nFor more information visit https://access.redhat.com/errata/RHSA-2026:61627 https://access.redhat.com/errata/RHSA-2026:61627\nFor more information visit https://access.redhat.com/errata/RHSA-2026:61629 https://access.redhat.com/errata/RHSA-2026:61629\nFor more information visit https://access.redhat.com/errata/RHSA-2026:60363 https://access.redhat.com/errata/RHSA-2026:60363\n\nWorkarounds / mitigations:\n\n- To mitigate this issue, restrict network access to the vLLM inference engine's gRPC Generate and Abort endpoints. Configure firewall rules to limit incoming connections to trusted clients or internal networks only. This will prevent remote, unauthenticated attackers from sending malformed workloads and triggering a denial of service. If the service is exposed via a proxy or load balancer, ensure that access controls are in place at that layer.\n\n## Package advisory (CVE-2026-54234)\n\nAffected packages:\n\n- `vllm >= 0.17.1, < 0.24.0`\n\nPatched in:\n\n- `vllm 0.24.0`\n\nSource: https://osv.dev/vulnerability/GHSA-8wr5-jm2h-8r4f","depth":"twilight","depthScore":41,"depthScoreParts":{"impact":41.3,"likelihood":0.1,"exploitation":0,"ransomware":0},"changes":[]}