The gateway config still routed to retired pre-club-3090 endpoints (:8000 vLLM-patched, :8001/:8002 SGLang, :8003 longctx/ngram, :8004 luce-dflash) — none of which gpu-mode brings up anymore — and had no route to the shipped 35B-A3B dual. Now 3 live primaries only: - qwen3.6-27b-autoround -> :8010 (gpu-mode 27b) - qwen3.6-35b-a3b-autoround -> :8051 (launch.sh vllm/qwen-35b-a3b-dual; served-model-name matched) - gemma-4-31b-autoround -> :8030 (gpu-mode gemma) Gemma-4-26B-A4B left unrouted on purpose: its registry-default compose (vllm/gemma-a4b, autoround-int4-mixed) is SM86-blocked (Marlin K-dim); the working AWQ variant isn't a gpu-mode primary. Noted inline. Takes effect on next `docker restart litellm` / gpu-mode switch. YAML validated; no repo scripts/tests reference the removed model_names. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>