A critical vulnerability in LMCache, used to speed up large language model (LLM) servers such as vLLM, permits unauthenticated remote code execution on the cache server. The flaw lies in LMCache’s multiprocess mode, where the cache operates as a standalone server accessed by LLM workers over ZeroMQ. A single network message can execute code with the privileges of the LMCache process.
The issue arises because the ZeroMQ socket has no authentication, and messages are unpacked with pickle during decoding, enabling arbitrary code execution before any type checks are performed. The LMCache process on official container images runs as root, according to JFrog. The vulnerability is triggered when the multiprocess server is configured to listen on a routable address rather than localhost, which is the default.
The CVE is CVE-2026-105192 and affects LMCache from version 0.3.9 (October 2025) through 0.5.5, with 0.5.6 release candidates and the development branch also affected; no fixed version is available. JFrog’s advisory notes this as a 9.8/10 critical issue and recommends operators keep the multiprocess server bound to localhost or a trusted cluster network, avoid routable addresses, and use a firewall to limit access.
Separately, a related vLLM flaw (pre-0.30.0) could cause a denial-of-service via a malformed cache_salt value (CVE-2026-105756), but it does not allow code execution. The article emphasises that the core error is handling unauthenticated network data via pickle. The source does not indicate a published LMCache security advisory or a patch.