Files
club-3090/scripts/lib
noonghunnaandClaude Opus 4.7 2d1b1dc347 feat(moe): add dual-card composes for Gemma 26B-A4B + Qwen 35B-A3B preview
Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.

* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
  TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
  comfortable KV headroom. No drafter on this base smoke compose.

* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
  TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
  port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
  ~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
  mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.

* COMPOSE_REGISTRY: dual variants become the canonical keys
  (vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
  single variants renamed to vllm/gemma-a4b-single and
  vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).

Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
..