LMCache is a distributed key-value cache that sits in front of LLM inference engines, vLLM chief among them. It exists to stop a model recomputing attention state it has already computed, which is why it ends up in production serving stacks rather than in experiments.

JFrog published an advisory on 7 October for CVE-2026-105192, scored 9.8. The bug is in its multiprocess transport, and it is the kind that gets taught.

The order of operations is the bug

The transport accepts connections without authentication. When a message arrives, it deserialises it with pickle before checking what type of message it is.

Python's pickle is not a data format. It is a small instruction set for reconstructing objects, and reconstructing an object can mean calling things. Unpickling attacker-controlled bytes is executing attacker-controlled code, which is why the module's own documentation warns against it in the first paragraph.

So the sequence is: receive bytes from anyone, run them, then decide whether they were a message you wanted. One crafted message is remote code execution as whatever user the LMCache process runs as.

Affected builds run from 0.3.9 onward, including the 0.5.5 stable release, the 0.5.6 release candidates and the development branch.

The score and the default disagree

The 9.8 assumes network reachability, no privileges and no user interaction. Two of those are right and the first deserves an asterisk.

By default LMCache binds to loopback, so a stock single-node install is not exposed to the internet. The exposure appears when somebody sets a routable host value — which is exactly what multi-node and Kubernetes deployments do, because that is the point of a distributed cache.

This is a 9.8 that is either unreachable or trivially reachable depending on one flag, and the flag is set by the deployments most likely to matter. Reading the number without the binding gets you the wrong answer in both directions.

We made a related point about a different vendor this week, where a researcher's speculation about grid impact ended up quoted inside the official CVE description. A severity score is a model with assumptions in it. It is not a measurement of your environment.

Root in the container

The container images ship with the process running as root.

That turns a compromise of the cache into a compromise of the host rather than something contained to a namespace. It is a packaging decision rather than a vulnerability, and it is the difference between the two outcomes here.

No patch, and a proof of concept the same day

As of the advisory there is no fixed release, no vendor security bulletin and no published timeline. Exploit code was posted publicly on the day of disclosure.

Two further identifiers have appeared against the same project: another 9.8 for remote Python execution through a script-submission route, and a 9.4 for an unauthenticated route that returns environment variables — which, in a serving deployment, is where the API keys are. Their patch status is unclear.

The mitigations available now are network ones. Keep the port off routable addresses, restrict it to local or trusted-cluster traffic, and treat any host allowed to connect as a host allowed to run code, because that is what it is. A durable fix means not unpickling untrusted input at all and adding cryptographic peer authentication — a redesign of the transport, not a patch to it.

Why this keeps happening to AI infrastructure

Exposed inference infrastructure is the second story of its kind here in three days. A botnet has been sweeping ports 3000 and 4000 for exposed LiteLLM and Gotenberg instances and turning what it finds into miners.

The pattern is not that AI software is written badly. It is that this layer is young, it is being deployed fast by teams measured on inference throughput, and a great deal of it was written when the assumption was a trusted cluster network. Pickle over an unauthenticated socket is a perfectly reasonable choice inside a single process boundary and an indefensible one across a cluster, and the code did not change when the deployment did.

What to do

  • Find out whether your LMCache instances bind to anything other than loopback. That single answer is most of your risk assessment.
  • Put the transport behind network policy now. A fix is not available to wait for.
  • Check whether your images run the process as root, and stop doing that regardless of this bug.
  • Rotate anything an environment-variable leak would have exposed if your deployment also exposes that route.

What is not established

  • When a fix will ship. No timeline has been published.
  • Whether this is being exploited. We have seen no report of it, and a public proof of concept on day one is not evidence of use.
  • The patch status of the two related identifiers.
  • Which downstream platforms actually run the vulnerable mode. Several large serving stacks are named in coverage as LMCache users, which is not the same as being configured this way.
  • Whether moving to a newer vLLM release addresses this. That claim appears in secondary coverage without confirmation.