Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.
* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
comfortable KV headroom. No drafter on this base smoke compose.
* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.
* COMPOSE_REGISTRY: dual variants become the canonical keys
(vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
single variants renamed to vllm/gemma-a4b-single and
vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).
Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>