gemma-4-12b: promote the two vLLM MTP composes to ⚠️ Production w/ caveats
vllm/gemma-12b-dual-bf16-mtp + vllm/gemma-12b-single-int8-mtp → status="caveats"
(🧪 → ⚠️). Both cleared the gates:
- dual-bf16-mtp: rebench-full (verify-full + bench + verify-stress + 8-pack
94/150 + soak PASS), 256K NIAH overlay-free.
- single-int8-mtp: bench + 256K NIAH + 8-pack 105/150 + soak PASS (fresh 20x5:
0/100 silent-empty, 0 MiB growth, 95.1% retention).
CAVEAT (both): the gemma4-unified image is an ephemeral arch-preview tag (0.1.dev)
— pin a digest; promotes to ✅ Production when gemma4_unified ships in a STABLE
vLLM release.
Withheld at 🧪 (deliberate): beellama/gemma-12b-single-q8kxl + llamacpp/...-q8kxl
— no MTP/spec-dec for Gemma-4 on llama.cpp yet (blocked on llama.cpp#23398).
Header Status+Caveats, registry status, and BENCHMARKS markers all flipped
together; test-compose-status-drift + full guard suite 41/41 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>