The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.
vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.
This commit:
* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
install.sh: anchor-based Python patcher that inserts the 4 branches
into the in-container gemma4.py at runtime. Idempotent (sentinel
comment), drift-resistant (anchors on existing branch line, not
line numbers). Smoke-tested against bf610c2f: sentinel count=1
after first install, no-op on second install, file remains valid
Python after patching.
README.md: full vendor context, usage pattern, drop trigger.
* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
/mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
Entrypoint runs install.sh before vllm serve. Routes through
vllm-nightly-clean (bf610c2f, no Genesis).
* ModelProfile gemma-4-26b-a4b.yml updates:
- autoround_int4_mixed: status flipped to "ampere-blocked" with
a one-line forensic note (was "production", incorrect on SM86)
- awq_compressed_tensors: new variant pointing at the cyankiwi
AWQ-4bit weights, marked production via PR #40886 overlay
- default_weight_variant: switched to awq_compressed_tensors
(Ampere users get the variant that boots by default)
* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
path at port 8042.
* vllm-nightly-clean engine: supported_weight_formats gains
"compressed-tensors" so fits() C14 accepts the AWQ variant.
All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.
Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.
Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
was issue #137 / #138 territory — here we deliver the alternative
that unblocks Gemma 26B-A4B for the v0.7.3 ship).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>