---
id: CVE-2026-105752
title: vLLM is an inference and serving engine for large language models
summary: >-
  vLLM is an inference and serving engine for large language models. Prior to
  0.30.0, Harmony tool continuations submitted through "POST /v1/responses"
  requests rebuild the next-turn engine input without preserving the cache_salt
  value, pl…
severity: low
cvss: 3.1
cvssVector: 'CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:N'
cwe:
  - CWE-200
  - CWE-524
vendor: vllm-project
product: vllm
affected:
  - vllm < 0.30.0
published: '2026-10-05'
updated: '2026-10-05'
sourceUpdated: '2026-10-05T23:17:01.710'
source: NVD
sourceUrl: 'https://nvd.nist.gov/vuln/detail/CVE-2026-105752'
references:
  - url: >-
      https://github.com/vllm-project/vllm/commit/6a2a2bb02b563b83f946012959fd3927984d072a
    label: security-advisories@github.com
  - url: 'https://github.com/vllm-project/vllm/pull/50195'
    label: security-advisories@github.com
  - url: 'https://github.com/vllm-project/vllm/pull/51818'
    label: security-advisories@github.com
  - url: 'https://github.com/vllm-project/vllm/releases/tag/v0.30.0'
    label: security-advisories@github.com
  - url: >-
      https://github.com/vllm-project/vllm/security/advisories/GHSA-935w-9g4m-p28p
    label: security-advisories@github.com
  - url: 'https://github.com/advisories/GHSA-935w-9g4m-p28p'
tags:
  - nvd
  - cve.org
  - ghsa
  - pip
ingestedAt: '2026-10-05T23:36:21.155Z'
aliases:
  - GHSA-935w-9g4m-p28p
ecosystem: pip
patched:
  - vllm 0.30.0
---

## Overview

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the continuation prefix in the global unsalted cache namespace even when the caller enabled salting. On deployments with prefix caching enabled, which is the default, an authenticated tenant who can reconstruct a victim's low-entropy post-tool history can submit the same continuation and use the cached_tokens_per_turn count to determine whether the prefix was previously processed, defeating the intended tenant isolation of salted prefix caching. This issue is fixed in version 0.30.0.

## Remediation

Refer to the linked advisories for vendor-supplied fixes and affected version ranges.

## Package advisory (CVE-2026-105752)

Affected packages:

- `vllm < 0.30.0`

Patched in:

- `vllm 0.30.0`

Source: https://github.com/advisories/GHSA-935w-9g4m-p28p
