The upstream rolling `ghcr.io/ggml-org/llama.cpp:server-cuda` tag regressed at b9282 — `llama-server` crash-loops with `libllama-common.so.0: cannot open shared object file` (broken lib packaging). Since our composes defaulted to the rolling tag, every fresh pull of llamacpp/default + llamacpp/mtp-vision hit the crash loop and the endpoint never came up. Reported in #187. Pin both single-card composes to the validated build `server-cuda-b9246` (2026-05-20 — the build all current BENCHMARKS rows were measured on). Override to follow a newer build via LLAMACPP_IMAGE once validated. README + engine profile notes updated to match. Closes #187. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>