
JFrog researchers disclosed CVE-2026-105192, a critical flaw in LMCache, open-source software that speeds up large language model servers such as vLLM. In LMCache’s multiprocess mode the cache runs as a standalone server that LLM workers reach over ZeroMQ. That socket has no authentication, and one message type is unpacked with Python’s pickle format, which can carry code and run it as it is decoded. Because the server unpacks the message arguments before checking the message type, a single crafted network message can run commands as the user the LMCache process runs as. According to JFrog, that process runs as root in the project’s official container images.
The Hacker News reports that the server is reachable from other machines only when an operator binds it to a routable address, since it listens on localhost by default; LMCache’s own example Kubernetes deployment, however, starts it listening on every network interface. JFrog rates the bug 9.8 out of 10 for that exposed configuration. No fixed release exists, and LMCache has not published a security advisory. Separately, a GitHub user filed six more unconfirmed LMCache security reports on October 6, which have no CVE, maintainer confirmation or fix.
How to check if you’re affected
Affected versions include LMCache 0.3.9 through 0.5.5, the latest stable release, and the 0.5.6 release candidates and development branch also contain the flaw. Only multiprocess-mode deployments are exposed: a copy of LMCache running inside a single vLLM process does not open the port. Check whether your multiprocess server is started with a routable address (for example, a deployment copied from LMCache’s Kubernetes example) and whether its port is reachable from other hosts.
Until a patch ships, JFrog advises not assigning the multiprocess server a routable address and keeping its port on the local machine or a trusted cluster network. A firewall lowers the risk but does not remove it, because any host that can still open a connection can run code. JFrog’s advisory gives no way to tell whether a server has already been attacked.
