Files
club-3090/models
noonghunnaandClaude Opus 4.7 0053444e84 feat(gemma-4-26b-a4b): AWQ path via vLLM PR #40886 overlay
The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.

vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.

This commit:

* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
  install.sh: anchor-based Python patcher that inserts the 4 branches
  into the in-container gemma4.py at runtime. Idempotent (sentinel
  comment), drift-resistant (anchors on existing branch line, not
  line numbers). Smoke-tested against bf610c2f: sentinel count=1
  after first install, no-op on second install, file remains valid
  Python after patching.
  README.md: full vendor context, usage pattern, drop trigger.

* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
  TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
  /mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
  Entrypoint runs install.sh before vllm serve. Routes through
  vllm-nightly-clean (bf610c2f, no Genesis).

* ModelProfile gemma-4-26b-a4b.yml updates:
  - autoround_int4_mixed: status flipped to "ampere-blocked" with
    a one-line forensic note (was "production", incorrect on SM86)
  - awq_compressed_tensors: new variant pointing at the cyankiwi
    AWQ-4bit weights, marked production via PR #40886 overlay
  - default_weight_variant: switched to awq_compressed_tensors
    (Ampere users get the variant that boots by default)

* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
  path at port 8042.

* vllm-nightly-clean engine: supported_weight_formats gains
  "compressed-tensors" so fits() C14 accepts the AWQ variant.

All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.

Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.

Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
   was issue #137 / #138 territory — here we deliver the alternative
   that unblocks Gemma 26B-A4B for the v0.7.3 ship).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
..