Business AI Now Generates More Than 10% of TCS Revenue: Annualized Run Rate Hits $3.1 Billion, Up 19% in a Quarter AI Agents Stuut Raises $52.5 Million to Put AI Agents on Collections Calls: 81.7% of Outbound Chasing Now Runs Without a Human Security CVSS 9.8, No Patch: JFrog Finds Unauthenticated RCE in LMCache, the KV-Cache Layer Running Beside vLLM Inference Servers Business Nvidia Prepares to Back a Rival: Report Says a d-Matrix Investment Is in the Works, With the Startup's Next Inference Chip Already Slated for Nvidia's Own Racks ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data Business AI Now Generates More Than 10% of TCS Revenue: Annualized Run Rate Hits $3.1 Billion, Up 19% in a Quarter AI Agents Stuut Raises $52.5 Million to Put AI Agents on Collections Calls: 81.7% of Outbound Chasing Now Runs Without a Human Security CVSS 9.8, No Patch: JFrog Finds Unauthenticated RCE in LMCache, the KV-Cache Layer Running Beside vLLM Inference Servers Business Nvidia Prepares to Back a Rival: Report Says a d-Matrix Investment Is in the Works, With the Startup's Next Inference Chip Already Slated for Nvidia's Own Racks ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data

CVSS 9.8, No Patch: JFrog Finds Unauthenticated RCE in LMCache, the KV-Cache Layer Running Beside vLLM Inference Servers

CVSS 9.8, No Patch: JFrog Finds Unauthenticated RCE in LMCache, the KV-Cache Layer Running Beside vLLM Inference Servers

JFrog researchers disclosed CVE-2026-105192 on Oct 7: an unauthenticated remote-code-execution flaw in LMCache's multiprocess mode, rated 9.8 on the CVSS scale. Every release from 0.3.9 through the current 0.5.5 and the 0.5.6 candidates is affected and no fixed version exists. A single ZeroMQ message can execute code as the LMCache process user, which runs as root in the official containers.

A critical, still-unpatched vulnerability has been disclosed in LMCache, the open-source KV-cache layer that sits alongside LLM serving engines such as vLLM to make inference faster and cheaper. JFrog Security Research published the details on Oct 7 under CVE-2026-105192, rating it 9.8 out of 10 — critical — and, as of publication, no fixed release of LMCache exists.

The flaw lives in LMCache's multiprocess mode, where the cache runs as a standalone server that worker processes reach over the ZeroMQ messaging library. That server opens a ROUTER socket, port 5555 by default and bound to localhost, with no CURVE encryption, no ZAP authentication and no message verification. The bind only becomes dangerous when an operator points it at a routable address with the --host flag — exactly what multi-node deployments do. JFrog notes that LMCache's own example Kubernetes configuration starts the server listening on every network interface.

What turns exposure into code execution is the serialization. Messages arrive as msgpack, and extension code 1 is wired to a wrapper class whose payload is decoded with Python's pickle before any validation happens. Pickle can carry executable code, so a single crafted message is enough to run arbitrary Python as the LMCache process — and the official container images run that process as root. JFrog researcher Yuval Moravchick documented the path down to the source code, where the same pickle-based deserialization is visible in the current release.

The affected range is wide: version 0.3.9, released in October 2025, through 0.5.5, the current stable release from September 2026, plus the 0.5.6 release candidates and the development branch as of disclosure. LMCache has not issued a security advisory, and JFrog's write-up gives operators no way to tell whether a server has already been hit. There is no public proof-of-concept and no report of exploitation so far; NVD's triage data marks exploitation as none but automation as yes.

Until a patch ships, the mitigations are configuration, not code: keep the multiprocess server on localhost or a trusted cluster network, firewall port 5555, and assume any host that can reach the port can run code on it. JFrog's longer-term recommendations to the maintainers are to replace pickle for extension code 1 with a safe serializer, authenticate the ZeroMQ transport with CURVE or per-message HMAC, and refuse to bind to a routable address unless authentication is configured.

The disclosure lands alongside a second, related bug: CVE-2026-105756, a 6.5-rated denial-of-service flaw in vLLM deployments that use the LMCache multiprocess connector, where a single malformed cache_salt request could crash the engine for every user. That one is already fixed in vLLM 0.30.0, released Sept 22. The underlying mistake — feeding unauthenticated network data to pickle — is the same pattern researchers flagged across other AI inference frameworks in November 2025, a bug class dubbed ShadowMQ. GPU inference hosts are an attractive target precisely because of what they hold: model weights, API keys and cloud credentials, often on the same machine.

Comments (0)

Log in to join the discussion

Log In

No comments yet