---
id: CVE-2026-94627
title: >-
  vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache
  block ownership when concurrent child requests share a single transfer ID in
  prefill/decode disaggregated deployments
summary: >-
  vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache
  block ownership when concurrent child requests share a single transfer ID in
  prefill/decode disaggregated deployments. Attackers can trigger GPU memory
  exhausti…
severity: high
cvss: 7.5
cvssVector: 'CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H'
cwe:
  - CWE-401
  - CWE-772
vendor: vllm-project
product: vllm
affected:
  - vllm <= 0.29.0
published: '2026-09-21'
updated: '2026-09-22'
sourceUpdated: '2026-09-22T20:25:55.870'
source: NVD
sourceUrl: 'https://nvd.nist.gov/vuln/detail/CVE-2026-94627'
references:
  - url: 'https://github.com/vllm-project/vllm'
    label: disclosure@vulncheck.com
  - url: >-
      https://github.com/vllm-project/vllm/blob/v0.29.0/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/mooncake_connector.py#L1978-L1989
    label: disclosure@vulncheck.com
  - url: 'https://github.com/vllm-project/vllm/pull/49796'
    label: disclosure@vulncheck.com
  - url: >-
      https://www.vulncheck.com/advisories/vllm-through-0.29.0-gpu-kv-cache-leak-via-mooncake-transfer-id-collision
    label: disclosure@vulncheck.com
  - url: >-
      https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-94627.json
  - url: 'https://access.redhat.com/security/cve/CVE-2026-94627'
  - url: 'https://bugzilla.redhat.com/show_bug.cgi?id=2537919'
  - url: 'https://www.cve.org/CVERecord?id=CVE-2026-94627'
  - url: 'https://nvd.nist.gov/vuln/detail/CVE-2026-94627'
tags:
  - nvd
  - cve.org
  - csaf
  - vex
  - red-hat
epss: 0.0063
epssPercentile: 0.47983
ssvc:
  exploitation: none
  automatable: 'yes'
  technicalImpact: partial
  timestamp: '2026-09-22T12:50:58.697049Z'
ingestedAt: '2026-09-21T22:54:37.974Z'
---

## Overview

vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart and eventually preventing legitimate requests from executing.

## Remediation

Refer to the linked advisories for vendor-supplied fixes and affected version ranges.

## Vendor advisories

- **Red Hat VEX** · Important · affected: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · no fix planned: Red Hat AI Inference Server, Red Hat Enterprise Linux AI (RHEL AI) 3, Red Hat OpenShift AI (RHOAI) · updated 2026-09-25 · [vex](https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-94627.json)
