A symmetric 4-slug family for the Qwen vLLM path:
*-fast = AutoRound INT4 + fp8_e5m2 KV (peak TPS, the proven path)
*-max = official FP8 + int8-PTH KV (higher fidelity @ 262K)
Slugs:
vllm/qwen-27b-dual-fast alias of vllm/dual (AutoRound INT4, TP=2) — production
vllm/qwen-27b-dual-max FP8 + int8-PTH, TP=2 — 🧪 live-validated 2026-06-07
vllm/qwen-27b-multi-fast AutoRound INT4, TP=4 — 🧪 cross-rig (dual is the proxy)
vllm/qwen-27b-multi-max FP8 + int8-PTH, TP=4 — 🧪 cross-rig (dual-max is the proxy)
dual-max validated live on this 2-card rig: FP8 weights load via MarlinFP8
W8A16 on Ampere sm_86, int8_per_token_head KV keeps the full 262K (KV pool
295K tok / 1.13x concurrency), MTP n=3 active, coherent. The multi4 configs
are byte-identical to their dual sibling apart from TP + gpu-count, so they
ship Experimental until a real >=4x 3090 host validates them (this dev rig
has 2 cards). No multi4 DEFAULT added — the resolver degrades to vllm/dual
(honest; no model carries a multi default), and the slugs stay reachable by name.
Catalog wiring:
- models/qwen3.6-27b.yml: new fp8 weights_variant (official e4m3 release,
layer-split safetensors + embedded mtp.safetensors + vision tower)
- vllm-stable.yml: + fp8 weight format (Marlin W8A16 on Ampere) and
+ int8_per_token_head KV — NATIVE on stock v0.22.0 for uniform-head-dim
models (qwen3-next); features flag stays false (no overlay, unlike Gemma's
#40391-provided int8-PTH on vllm-gemma-stable)
- compose_registry.py: 4 entries; test-compose-registry-disk counts 38->42 / 40->43
Full gate green (42/42); compat C4/C5/C10/C14 pass for all 4 slugs.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>