The LiteLLM gateway routes qwen3.6-27b-autoround → :8010, but the fp8 / awq / lmcache 27b scenes served scene-specific names (qwen3.6-27b-fp8, qwen3.6-27b-awq-bf16-int4). Bring one of those up as the :8010 primary (e.g. via gpu-mode PORT override) and the gateway 404s on a served-name mismatch (#482). Standardize every 27b serving scene's --served-model-name to the canonical qwen3.6-27b-autoround so the route matches whichever scene is on :8010. The quant still differs by compose path/port — only the served name is unified. Weights --model paths are untouched. Document the invariant in services/litellm/config.yaml. Full test suite green (58/58). Closes #482 Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 <[email protected]>