Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).
Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text single/tools-text.yml fp8_e5m2
- vllm/minimal single/minimal.yml fp8_e5m2
- vllm/dual dual/docker-compose.yml fp8_e5m2
- vllm/dual-bf16 dual/bf16.yml bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp dual/qwopus-bf16mtp.yml fp8_e5m2
- vllm/dual-nvlink dual/nvlink.yml (extends) fp8_e5m2
- vllm/dual4 multi4/docker-compose.yml fp8_e5m2
TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.
Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
Genesis is required only for TQ3 KV format; that's enforced at the
compose level via Engine-profile selection.
Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
coverage.
Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
requires_genesis=true, so the Genesis enforcement happens at C15
(engine feature) level, not C6. Test now asserts positive (Qwen +
fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
non-Genesis engine fails C15).
All compat tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Adds the v0.7.3 critical-path plumbing for both MoE models:
* Compose files:
- models/gemma-4-26b-a4b/vllm/compose/single/docker-compose.yml
(vllm-nightly-clean, TP=1, port 8040, bf16 KV, no drafter)
- models/qwen3.6-35b-a3b/vllm/compose/single/preview.yml
(vllm-nightly-clean, TP=1, port 8050, fp8_e5m2 KV, preview-only)
* COMPOSE_REGISTRY entries 'vllm/gemma-a4b' and 'vllm/qwen-a3b-preview'.
* DrafterProfile: new gemma-26b-it-assistant for the 26B-A4B
(google/gemma-4-26B-A4B-it-assistant — separate weights from 31B);
qwen-mtp-builtin.model_compat extended with qwen3.6-35b-a3b.
* vllm-nightly-clean.supported_model_families adds qwen3-next-moe as a
preview/smoke route; notes clarify Cliff 2 / TQ3 unavailable here.
* Qwen 3.6 35B-A3B model: requires_genesis flipped false → upstream is
strictly bootable on post-#42521 nightlies; Genesis is RECOMMENDED for
production (TQ3 KV + Cliff 2 mitigations), not strictly required to boot.
* ModelProfile gains kv_calc_supported (defaults true). Both new MoE
models set false because tools/kv-calc.py only models qwen3.6-27b and
gemma-4-31b today (MoE + asymmetric KV + DeltaNet+attention hybrid all
require new MODEL_SPECS + activation coefficients — queued for Codex).
fits() now skips C12 for kv_calc_supported=false instead of failing.
All compat tests pass; both new composes fit on 1x-3090 (Gemma) and
1x-3090/2x-3090/1x-5090 (Qwen preview).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Add vllm-nightly-clean engine profile (latest nightly, no Genesis) for
model families that don't need DeltaNet stabilization patches.
Rule:
- required_genesis: true → SHA pinned to Genesis anchor, moves only on
Sander's pin-bump cycle (vllm-nightly-mtp/dflash/full)
- required_genesis: false → free to ride latest nightly (vllm-nightly-clean)
Routes:
- qwen3-next-hybrid + qwen3-next-moe → vllm-nightly-mtp (Genesis-locked)
- gemma4-swa-moe → vllm-nightly-clean (unconstrained)
- gemma4-swa-dense → either
- dense → either
Also adds qwen3-next-moe + gemma4-swa-moe to vllm-nightly-mtp's family
list (factual, independent of SHA), corrects misleading notes that
referenced an unmerged pin bump, and bumps engine-count assertion in
test-profiles-compat.sh from 6 to 7.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Onboards Gemma 4 26B-A4B (gemma4-swa-moe family) as the fourth
ModelProfile. Real values sourced from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json
(file downloaded but checkpoint still streaming in background as of
this commit).
Schema extension: new optional field `num_global_kv_heads`. Gemma 4
26B-A4B has asymmetric KV head counts (8 sliding / 2 global) — distinct
from Gemma 4 31B which uses 16 for both. Field is None for symmetric
models; existing Gemma 4 31B + Qwen models work unchanged.
Confirmed architectural facts (from config.json):
- 30 total hidden layers (smaller than estimated 40-48)
- 5 full_attention layers at indices [5, 11, 17, 23, 29]
(one global per 6-layer block; last layer always global)
- 25 sliding_attention layers (sliding_window=1024)
- num_kv_heads=8 (sliding), num_global_kv_heads=2 (global) — ASYMMETRIC
- head_dim=256 (sliding), global_head_dim=512 (global)
- attention_k_eq_v: TRUE (K=V tied)
- num_experts: 128, top_k_experts: 8 (active per token)
- moe_intermediate_size: 704
- Multimodal: vision_config + audio token IDs present
- requires_genesis: false (Gemma 4 doesn't need Genesis like Qwen)
KV math implication: per-token growing KV at bf16 is
5 × 2 × 512 × 1 (K=V tied) = 5,120 bytes
vs Gemma 4 31B's 10 × 16 × 512 × 1 = 81,920 bytes per token
=> Gemma 4 26B-A4B has ~16x smaller growing-KV per token than 31B,
making long context dramatically cheaper. KV_MATH.md MoE section
estimates need refresh.
Drafter compat note: references gemma-it-assistant (the existing 31B
drafter). TODO Codex: extend that DrafterProfile's model_compat to
include gemma-4-26b-a4b, OR split into a separate
gemma-26b-it-assistant DrafterProfile (google/gemma-4-26B-A4B-it-assistant
is downloading in background alongside the main model).
Tests pass: 41 tests, kv-calc 18/18 invariant held.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>