The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.
vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.
This commit:
* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
install.sh: anchor-based Python patcher that inserts the 4 branches
into the in-container gemma4.py at runtime. Idempotent (sentinel
comment), drift-resistant (anchors on existing branch line, not
line numbers). Smoke-tested against bf610c2f: sentinel count=1
after first install, no-op on second install, file remains valid
Python after patching.
README.md: full vendor context, usage pattern, drop trigger.
* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
/mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
Entrypoint runs install.sh before vllm serve. Routes through
vllm-nightly-clean (bf610c2f, no Genesis).
* ModelProfile gemma-4-26b-a4b.yml updates:
- autoround_int4_mixed: status flipped to "ampere-blocked" with
a one-line forensic note (was "production", incorrect on SM86)
- awq_compressed_tensors: new variant pointing at the cyankiwi
AWQ-4bit weights, marked production via PR #40886 overlay
- default_weight_variant: switched to awq_compressed_tensors
(Ampere users get the variant that boots by default)
* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
path at port 8042.
* vllm-nightly-clean engine: supported_weight_formats gains
"compressed-tensors" so fits() C14 accepts the AWQ variant.
All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.
Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.
Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
was issue #137 / #138 territory — here we deliver the alternative
that unblocks Gemma 26B-A4B for the v0.7.3 ship).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).
Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text single/tools-text.yml fp8_e5m2
- vllm/minimal single/minimal.yml fp8_e5m2
- vllm/dual dual/docker-compose.yml fp8_e5m2
- vllm/dual-bf16 dual/bf16.yml bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp dual/qwopus-bf16mtp.yml fp8_e5m2
- vllm/dual-nvlink dual/nvlink.yml (extends) fp8_e5m2
- vllm/dual4 multi4/docker-compose.yml fp8_e5m2
TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.
Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
Genesis is required only for TQ3 KV format; that's enforced at the
compose level via Engine-profile selection.
Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
coverage.
Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
requires_genesis=true, so the Genesis enforcement happens at C15
(engine feature) level, not C6. Test now asserts positive (Qwen +
fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
non-Genesis engine fails C15).
All compat tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Genesis is strictly required only for TQ3 KV on vLLM. The Gemma 4 31B
default composes don't use TQ3 — they don't need Genesis. They DO need
a post-PR-#41745 nightly for Gemma 4 MTP support (per compose header
breadcrumbs), which 01d4d1ad (the Genesis-anchored MTP pin) predates.
Affected composes (all bf16/fp8 KV, no TQ3):
- models/gemma-4-31b/vllm/compose/single/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/bf16.yml
All three flipped from `Engine-profile: vllm-nightly-mtp` to
`vllm-nightly-clean` (bf610c2f, 2026-05-15) which is post-everything
relevant: PR #41745 (Gemma 4 MTP), PR #42102 (DFlash + INT8 PTH), and
PR #42521 (qwen3_5_moe weight loading).
COMPOSE_REGISTRY entries vllm/gemma-mtp-tp1, vllm/gemma-mtp, and
vllm/gemma-bf16 updated to match.
The DFlash and INT8 paths (vllm-nightly-dflash on e47c98ef and
vllm-nightly-full on e47c98ef) are untouched — their pins remain
correctly anchored to the SHAs that produced all current BENCHMARKS
rows for those compose families. TQ3-using composes
(vllm/gemma-int8-tq3) likewise stay on vllm-nightly-full (Genesis).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.
* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
comfortable KV headroom. No drafter on this base smoke compose.
* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.
* COMPOSE_REGISTRY: dual variants become the canonical keys
(vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
single variants renamed to vllm/gemma-a4b-single and
vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).
Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Commit 40f1ef78 (2026-05-14) accidentally set vllm-nightly-mtp.spec to
nightly-1acd67a7, which is the post-Gemma4-merge nightly used by the
gemma-4-31b compose. The Genesis v7.72.2 PROD pin is nightly-01d4d1ad:
- BENCHMARKS rows from 2026-05-05 onwards reference nightly-01d4d1ad3
- calibration/qwen3.6-27b.yml: engine_pin: vllm-nightly-01d4d1ad
- gemma-4-31b/vllm/compose/single header: "Qwen3.6 composes stay on the
v7.72.2 PROD pin (01d4d1ad3) since Genesis allowlist anchors there"
Effect of the bug: any Qwen 27B launch via launch_compat.py since
2026-05-14 would have used the un-blessed Gemma-merge SHA (1acd67a7),
potentially causing Genesis patch apply failures or silent skips.
Existing benches that hardcoded the image or used the older mechanism
were unaffected.
Notes field updated with the corrected anchor + a PRIOR-BUG breadcrumb
so future readers don't re-introduce the bump.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Adds the v0.7.3 critical-path plumbing for both MoE models:
* Compose files:
- models/gemma-4-26b-a4b/vllm/compose/single/docker-compose.yml
(vllm-nightly-clean, TP=1, port 8040, bf16 KV, no drafter)
- models/qwen3.6-35b-a3b/vllm/compose/single/preview.yml
(vllm-nightly-clean, TP=1, port 8050, fp8_e5m2 KV, preview-only)
* COMPOSE_REGISTRY entries 'vllm/gemma-a4b' and 'vllm/qwen-a3b-preview'.
* DrafterProfile: new gemma-26b-it-assistant for the 26B-A4B
(google/gemma-4-26B-A4B-it-assistant — separate weights from 31B);
qwen-mtp-builtin.model_compat extended with qwen3.6-35b-a3b.
* vllm-nightly-clean.supported_model_families adds qwen3-next-moe as a
preview/smoke route; notes clarify Cliff 2 / TQ3 unavailable here.
* Qwen 3.6 35B-A3B model: requires_genesis flipped false → upstream is
strictly bootable on post-#42521 nightlies; Genesis is RECOMMENDED for
production (TQ3 KV + Cliff 2 mitigations), not strictly required to boot.
* ModelProfile gains kv_calc_supported (defaults true). Both new MoE
models set false because tools/kv-calc.py only models qwen3.6-27b and
gemma-4-31b today (MoE + asymmetric KV + DeltaNet+attention hybrid all
require new MODEL_SPECS + activation coefficients — queued for Codex).
fits() now skips C12 for kv_calc_supported=false instead of failing.
All compat tests pass; both new composes fit on 1x-3090 (Gemma) and
1x-3090/2x-3090/1x-5090 (Qwen preview).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Add vllm-nightly-clean engine profile (latest nightly, no Genesis) for
model families that don't need DeltaNet stabilization patches.
Rule:
- required_genesis: true → SHA pinned to Genesis anchor, moves only on
Sander's pin-bump cycle (vllm-nightly-mtp/dflash/full)
- required_genesis: false → free to ride latest nightly (vllm-nightly-clean)
Routes:
- qwen3-next-hybrid + qwen3-next-moe → vllm-nightly-mtp (Genesis-locked)
- gemma4-swa-moe → vllm-nightly-clean (unconstrained)
- gemma4-swa-dense → either
- dense → either
Also adds qwen3-next-moe + gemma4-swa-moe to vllm-nightly-mtp's family
list (factual, independent of SHA), corrects misleading notes that
referenced an unmerged pin bump, and bumps engine-count assertion in
test-profiles-compat.sh from 6 to 7.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Onboards Gemma 4 26B-A4B (gemma4-swa-moe family) as the fourth
ModelProfile. Real values sourced from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json
(file downloaded but checkpoint still streaming in background as of
this commit).
Schema extension: new optional field `num_global_kv_heads`. Gemma 4
26B-A4B has asymmetric KV head counts (8 sliding / 2 global) — distinct
from Gemma 4 31B which uses 16 for both. Field is None for symmetric
models; existing Gemma 4 31B + Qwen models work unchanged.
Confirmed architectural facts (from config.json):
- 30 total hidden layers (smaller than estimated 40-48)
- 5 full_attention layers at indices [5, 11, 17, 23, 29]
(one global per 6-layer block; last layer always global)
- 25 sliding_attention layers (sliding_window=1024)
- num_kv_heads=8 (sliding), num_global_kv_heads=2 (global) — ASYMMETRIC
- head_dim=256 (sliding), global_head_dim=512 (global)
- attention_k_eq_v: TRUE (K=V tied)
- num_experts: 128, top_k_experts: 8 (active per token)
- moe_intermediate_size: 704
- Multimodal: vision_config + audio token IDs present
- requires_genesis: false (Gemma 4 doesn't need Genesis like Qwen)
KV math implication: per-token growing KV at bf16 is
5 × 2 × 512 × 1 (K=V tied) = 5,120 bytes
vs Gemma 4 31B's 10 × 16 × 512 × 1 = 81,920 bytes per token
=> Gemma 4 26B-A4B has ~16x smaller growing-KV per token than 31B,
making long context dramatically cheaper. KV_MATH.md MoE section
estimates need refresh.
Drafter compat note: references gemma-it-assistant (the existing 31B
drafter). TODO Codex: extend that DrafterProfile's model_compat to
include gemma-4-26b-a4b, OR split into a separate
gemma-26b-it-assistant DrafterProfile (google/gemma-4-26B-A4B-it-assistant
is downloading in background alongside the main model).
Tests pass: 41 tests, kv-calc 18/18 invariant held.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>