---
id: CVE-2026-93436
title: >-
  vLLM through 0.29.0 fails to properly clean up decode-side metadata for
  rejected inference requests in prefill/decode disaggregated deployments
summary: >-
  vLLM through 0.29.0 fails to properly clean up decode-side metadata for
  rejected inference requests in prefill/decode disaggregated deployments.
  Remote attackers can submit requests with max_tokens=0 to exhaust
  decode-worker memory witho…
severity: high
cvss: 7.5
cvssVector: 'CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H'
cwe:
  - CWE-401
  - CWE-770
vendor: vllm-project
product: vllm
affected:
  - vllm <= 0.29.0
published: '2026-09-17'
updated: '2026-09-22'
sourceUpdated: '2026-09-22T20:25:55.870'
source: NVD
sourceUrl: 'https://nvd.nist.gov/vuln/detail/CVE-2026-93436'
references:
  - url: 'https://github.com/vllm-project/vllm'
    label: disclosure@vulncheck.com
  - url: >-
      https://github.com/vllm-project/vllm/blob/v0.29.0/vllm/distributed/kv_transfer/kv_connector/v1/nixl/push_worker.py#L162-L181
    label: disclosure@vulncheck.com
  - url: 'https://github.com/vllm-project/vllm/pull/55677'
    label: disclosure@vulncheck.com
  - url: >-
      https://www.vulncheck.com/advisories/vllm-through-0.29.0-memory-exhaustion-via-rejected-requests
    label: disclosure@vulncheck.com
  - url: >-
      https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-93436.json
  - url: 'https://access.redhat.com/security/cve/CVE-2026-93436'
  - url: 'https://bugzilla.redhat.com/show_bug.cgi?id=2538717'
  - url: 'https://www.cve.org/CVERecord?id=CVE-2026-93436'
  - url: 'https://nvd.nist.gov/vuln/detail/CVE-2026-93436'
tags:
  - nvd
  - cve.org
  - csaf
  - vex
  - red-hat
epss: 0.00759
epssPercentile: 0.53252
ssvc:
  exploitation: none
  automatable: 'yes'
  technicalImpact: partial
  timestamp: '2026-09-22T01:58:51.466543Z'
ingestedAt: '2026-09-17T22:30:21.417Z'
---

## Overview

vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.

## Remediation

Refer to the linked advisories for vendor-supplied fixes and affected version ranges.

## Vendor advisories

- **Red Hat VEX** · Important · affected: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · no fix planned: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · updated 2026-09-24 · [vex](https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-93436.json)
