Recommended Free Tools
LMCache’s documented AES-GCM option encrypts serialized cache payloads in the durable L2 tier; it does not encrypt data in L0 GPU memory or L1 host RAM. Separately, GitHub’s advisory lists LMCache versions through 0.4.6 as affected by CVE-2026-10813, a low-severity multimodal cache-key collision issue, but the records available as of October 7, 2026 do not identify a confirmed fixed release. Operators should verify the exact version boundary with current maintainer guidance rather than assume a later version is patched.
What is exposed in LMCache?
CVE-2026-10813 concerns multimodal cache-key collisions
The GitHub Advisory Database describes a weak 16-bit hash conversion in hex_hash_to_int16 in lmcache/integration/vllm/utils.py, used by the KV Cache Handler. The linked maintainer issue explains that different image identifiers can reduce to the same 16-bit value. In that case, a cache lookup could retrieve KV state generated for another image.
This is a cache-key collision concern, not an advisory describing general remote code execution or disclosure of cache contents. The advisory rates it low severity and gives it a CVSS v4 score of 1.1, with a local attack vector and high attack complexity. Those are the advisory’s assessments, not independent exploitability test results. The issue author notes that 16 bits allow 65,536 possible values and describes collisions appearing after a few hundred generated inputs; that is the author’s demonstration, not a separate benchmark.
Encryption covers only the L2 payload
| Cache tier | Where data resides | What the documented AES-GCM feature protects |
|---|---|---|
| L0 | GPU memory | Not encrypted by this feature; cache data is plaintext. |
| L1 | Host RAM | Not encrypted by this feature; cache data is plaintext. |
| L2 | Durable backend, such as filesystem or a supported remote backend | Serialized payload bytes can be encrypted at rest through the AES-GCM serde. |
The LMCache Team’s August 19, 2026 technical post characterizes the feature as at-rest protection for the durable tier, not end-to-end encryption. Someone able to access the running multiprocess server is outside the protection boundary of this feature.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Which versions are affected, and what should operators do?
The GitHub Advisory Database lists versions through 0.4.6 as affected and lists no patched version. The linked maintainer issue is closed as “not planned.” Those records do not establish whether a later release fixed the issue, whether a report was rejected, or whether another mitigation exists. They are not enough to declare all versions after 0.4.6 safe—or to call them affected.
- Identify the exact LMCache release in use. Record the package version and deployment artifact or image tag so you can compare the running build against current project guidance.
- Check current maintainer guidance for a version boundary. Look for a release note or a direct maintainer statement that explicitly addresses CVE-2026-10813. Do not infer a fix solely from a version number newer than 0.4.6.
- Until the boundary is clear, treat exposure according to your use case. If your deployment uses the affected multimodal KV Cache Handler path, assess whether untrusted or competing inputs can reach it and whether a cache collision could affect integrity or availability. The advisory does not describe confidentiality impact for the vulnerable system.
- Document and validate the decision. Pin the version you have verified, record the maintainer guidance relied on, and test the relevant multimodal cache behavior when changing releases or mitigations.
Do not treat L2 encryption as a fix for this collision issue: it protects stored payload bytes, not the generation of cache keys or the correctness of lookups.
Reporting a suspected vulnerability
LMCache’s official SECURITY.md asks people who believe they have found a vulnerability to email [email protected] and include useful details, such as examples or screenshots. The policy names no individual contact and does not promise a response time.
What does LMCache’s AES-GCM option protect?
The LMCache Team’s August 19, 2026 post documents an aesgcm serde for the L2 path. It can wrap an L2 adapter such as filesystem storage; the post also describes it as usable with S3, filesystem, RESP, and other adapters behind the serde wrapper. The documented default is AES-128-GCM, which provides confidentiality and integrity for stored payload bytes. Enabling it does not encrypt L0 or L1 data.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Configuration shape
The project’s example uses an L2 adapter configuration shaped like this:
{
"serde": {
"type": "aesgcm",
"key_provider": "hkdf",
"master_key_path": "/etc/lmcache/keys/master",
"aes_bits": 128
}
}
Adapt the serde to the selected backend and deployment. The example identifies a master-key file path; it is not a complete production secret-management policy. The post says the master key can be mounted as a Kubernetes Secret.
Key derivation is fleet-level, not independent tenant isolation
The documented default HkdfKeyProvider reads a master key from master_key_path and derives keys using cache_salt as a tenant selector. The salt is not itself key material. Because the derived tenant keys share one master, anyone who obtains that master can derive every tenant’s key.
The post describes KMS-backed per-tenant keys, per-tenant mounts, and tenant-to-node placement as future work, not shipped defaults. It also says rotation is manual: operators must use a new master key and invalidate and refill the cache.
Metadata remains visible
Encryption does not hide the L2 object name. The post says it includes cache_salt and a content-derived chunk_hash. A storage observer may therefore learn tenant identifiers and detect content overlap across tenants without decrypting payloads.
Integrity behavior and performance figures
The documented chunk format contains a version byte, a 12-byte random IV, ciphertext, and a 16-byte GCM authentication tag: 29 bytes of fixed framing overhead per chunk, according to the LMCache Team post. The IV must not repeat for a given key. If authentication fails because of a tag mismatch or wrong key, LMCache treats the chunk as a cache load miss, prompting a refetch or recomputation rather than silently restoring corrupted state.
The same post estimates AES-128-GCM throughput at approximately 4–8 GB/s per core on server hardware with AES-NI. This is the vendor post’s estimate, not an independently verified benchmark; actual throughput depends on the hardware and workload.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should LMCache be deployed safely?
Security depends on which tiers are used, who can reach the server and backend, how tenants share keys and infrastructure, and whether the runtime configuration is supported. The project’s deployment guide describes the following topology and operational considerations; they are not a blanket security guarantee.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Docker and IPC
The guide’s default multiprocess example uses shared IPC to support CUDA IPC transfers and documents Docker flags for networking, GPUs, and IPC. Isolated IPC can remove the shared /dev/shm dependency only when both LMCache and vLLM enable it. The guide limits that mode to the vLLM MP connector and notes memory-allocation constraints. Confirm that the exact connector and runtime combination supports the mode before relying on it.
Kubernetes and health monitoring
The guide describes running one LMCache server per node as a DaemonSet shared by vLLM pods. It recommends the HTTP server variant for liveness and readiness probes through /healthcheck, and documents logs and Prometheus metrics for monitoring. A shared per-node service is a topology choice: assess which workloads can reach it and what isolation exists between them.
Quick Recap
Deployment checks
- Map the data path: identify whether each workload uses L0, L1, L2, or a combination, and decide which threat applies to each tier.
- Restrict backend access: identify who can read the filesystem, object store, RESP service, backups, and snapshots. Payload encryption complements rather than replaces backend access controls.
- Protect the master key: limit access to the key file or mounted Secret, and account for the fact that one master can derive all documented salt-based tenant keys.
- Review metadata exposure: determine whether visible salts and content-derived chunk hashes are acceptable for your tenant and storage-observer threat model.
- Validate IPC and networking: confirm both-side settings, connector support, and memory constraints before enabling isolated IPC or changing shared IPC behavior.
- Check the complete compatibility matrix: validate the exact Python, PyTorch, accelerator ABI, connector, model, and feature recipe. The compatibility documentation says unlisted combinations are unverified until tested.
- Exercise recovery paths: verify the expected cache miss, refetch, or recomputation behavior during key changes and when encrypted data cannot be authenticated.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




