Layers an LMCache tiered persistent prefix-KV cache (MP/HMA connector) onto the dual-max fidelity profile (FP8 + int8-PTH KV + MTP n=3 @262K). Controlled A/B shows ZERO decode penalty (74 narr / 94 code TPS == without LMCache, MTP intact ~83% accept); reuses cached prefix KV across long / multi-session workloads (cold->warm TTFT ~7-8x) instead of re-prefilling. - New vllm-lmcache engine profile: lmcache/vllm-openai DIGEST-pinned (the tag is mutable — bundled vLLM 0.22.1-dev then 0.23.1-dev under the same tag). - preflight_lmcache_ram guard: rejects --l1-size-gb over-allocation even under --force (the l1=100-on-94GB host-OOM that forced a reboot, #133). - env-gated L2 disk tier hook (LMCACHE_L2_ADAPTER, off by default). - INTERNALS.md RAM/disk sizing section; registry entry + count bumps; full guard suite green. Incubating (hidden, --force to launch): runs LMCache's third-party image with a newer vLLM than our v0.22.0 pin; L2 latency unmeasured on-rig; 38 GB on-demand. Refs #133. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 <[email protected]>