Commit Graph
24 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 0053444e84 feat(gemma-4-26b-a4b): AWQ path via vLLM PR #40886 overlay
The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.

vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.

This commit:

* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
  install.sh: anchor-based Python patcher that inserts the 4 branches
  into the in-container gemma4.py at runtime. Idempotent (sentinel
  comment), drift-resistant (anchors on existing branch line, not
  line numbers). Smoke-tested against bf610c2f: sentinel count=1
  after first install, no-op on second install, file remains valid
  Python after patching.
  README.md: full vendor context, usage pattern, drop trigger.

* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
  TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
  /mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
  Entrypoint runs install.sh before vllm serve. Routes through
  vllm-nightly-clean (bf610c2f, no Genesis).

* ModelProfile gemma-4-26b-a4b.yml updates:
  - autoround_int4_mixed: status flipped to "ampere-blocked" with
    a one-line forensic note (was "production", incorrect on SM86)
  - awq_compressed_tensors: new variant pointing at the cyankiwi
    AWQ-4bit weights, marked production via PR #40886 overlay
  - default_weight_variant: switched to awq_compressed_tensors
    (Ampere users get the variant that boots by default)

* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
  path at port 8042.

* vllm-nightly-clean engine: supported_weight_formats gains
  "compressed-tensors" so fits() C14 accepts the AWQ variant.

All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.

Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.

Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
   was issue #137 / #138 territory — here we deliver the alternative
   that unblocks Gemma 26B-A4B for the v0.7.3 ship).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunna 99328b4cda feat(estate): add parallel boot mode 2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 3b2d940d26 fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean
Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).

Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text         single/tools-text.yml      fp8_e5m2
- vllm/minimal            single/minimal.yml         fp8_e5m2
- vllm/dual               dual/docker-compose.yml    fp8_e5m2
- vllm/dual-bf16          dual/bf16.yml              bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp  dual/qwopus-bf16mtp.yml  fp8_e5m2
- vllm/dual-nvlink        dual/nvlink.yml (extends)  fp8_e5m2
- vllm/dual4              multi4/docker-compose.yml  fp8_e5m2

TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.

Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
  Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
  Genesis is required only for TQ3 KV format; that's enforced at the
  compose level via Engine-profile selection.

Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
  coverage.

Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
  for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
  requires_genesis=true, so the Genesis enforcement happens at C15
  (engine feature) level, not C6. Test now asserts positive (Qwen +
  fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
  non-Genesis engine fails C15).

All compat tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 cf0451a7ad fix(gemma-4-31b): route default/bf16 composes to vllm-nightly-clean
Genesis is strictly required only for TQ3 KV on vLLM. The Gemma 4 31B
default composes don't use TQ3 — they don't need Genesis. They DO need
a post-PR-#41745 nightly for Gemma 4 MTP support (per compose header
breadcrumbs), which 01d4d1ad (the Genesis-anchored MTP pin) predates.

Affected composes (all bf16/fp8 KV, no TQ3):
- models/gemma-4-31b/vllm/compose/single/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/bf16.yml

All three flipped from `Engine-profile: vllm-nightly-mtp` to
`vllm-nightly-clean` (bf610c2f, 2026-05-15) which is post-everything
relevant: PR #41745 (Gemma 4 MTP), PR #42102 (DFlash + INT8 PTH), and
PR #42521 (qwen3_5_moe weight loading).

COMPOSE_REGISTRY entries vllm/gemma-mtp-tp1, vllm/gemma-mtp, and
vllm/gemma-bf16 updated to match.

The DFlash and INT8 paths (vllm-nightly-dflash on e47c98ef and
vllm-nightly-full on e47c98ef) are untouched — their pins remain
correctly anchored to the SHAs that produced all current BENCHMARKS
rows for those compose families. TQ3-using composes
(vllm/gemma-int8-tq3) likewise stay on vllm-nightly-full (Genesis).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 2d1b1dc347 feat(moe): add dual-card composes for Gemma 26B-A4B + Qwen 35B-A3B preview
Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.

* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
  TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
  comfortable KV headroom. No drafter on this base smoke compose.

* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
  TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
  port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
  ~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
  mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.

* COMPOSE_REGISTRY: dual variants become the canonical keys
  (vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
  single variants renamed to vllm/gemma-a4b-single and
  vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).

Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 87f0a0c528 fix(engines): vllm-nightly-mtp anchors to 01d4d1ad (Sander v7.72.2 PROD pin)
Commit 40f1ef78 (2026-05-14) accidentally set vllm-nightly-mtp.spec to
nightly-1acd67a7, which is the post-Gemma4-merge nightly used by the
gemma-4-31b compose. The Genesis v7.72.2 PROD pin is nightly-01d4d1ad:

- BENCHMARKS rows from 2026-05-05 onwards reference nightly-01d4d1ad3
- calibration/qwen3.6-27b.yml: engine_pin: vllm-nightly-01d4d1ad
- gemma-4-31b/vllm/compose/single header: "Qwen3.6 composes stay on the
  v7.72.2 PROD pin (01d4d1ad3) since Genesis allowlist anchors there"

Effect of the bug: any Qwen 27B launch via launch_compat.py since
2026-05-14 would have used the un-blessed Gemma-merge SHA (1acd67a7),
potentially causing Genesis patch apply failures or silent skips.
Existing benches that hardcoded the image or used the older mechanism
were unaffected.

Notes field updated with the corrected anchor + a PRIOR-BUG breadcrumb
so future readers don't re-introduce the bump.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 f7f6f444b9 feat(moe): wire Gemma 4 26B-A4B + Qwen 3.6 35B-A3B composes through fits()
Adds the v0.7.3 critical-path plumbing for both MoE models:

* Compose files:
  - models/gemma-4-26b-a4b/vllm/compose/single/docker-compose.yml
    (vllm-nightly-clean, TP=1, port 8040, bf16 KV, no drafter)
  - models/qwen3.6-35b-a3b/vllm/compose/single/preview.yml
    (vllm-nightly-clean, TP=1, port 8050, fp8_e5m2 KV, preview-only)

* COMPOSE_REGISTRY entries 'vllm/gemma-a4b' and 'vllm/qwen-a3b-preview'.

* DrafterProfile: new gemma-26b-it-assistant for the 26B-A4B
  (google/gemma-4-26B-A4B-it-assistant — separate weights from 31B);
  qwen-mtp-builtin.model_compat extended with qwen3.6-35b-a3b.

* vllm-nightly-clean.supported_model_families adds qwen3-next-moe as a
  preview/smoke route; notes clarify Cliff 2 / TQ3 unavailable here.

* Qwen 3.6 35B-A3B model: requires_genesis flipped false → upstream is
  strictly bootable on post-#42521 nightlies; Genesis is RECOMMENDED for
  production (TQ3 KV + Cliff 2 mitigations), not strictly required to boot.

* ModelProfile gains kv_calc_supported (defaults true). Both new MoE
  models set false because tools/kv-calc.py only models qwen3.6-27b and
  gemma-4-31b today (MoE + asymmetric KV + DeltaNet+attention hybrid all
  require new MODEL_SPECS + activation coefficients — queued for Codex).
  fits() now skips C12 for kv_calc_supported=false instead of failing.

All compat tests pass; both new composes fit on 1x-3090 (Gemma) and
1x-3090/2x-3090/1x-5090 (Qwen preview).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 15eda8a823 feat(profiles): split engine-pin policy by Genesis dependency
Add vllm-nightly-clean engine profile (latest nightly, no Genesis) for
model families that don't need DeltaNet stabilization patches.

Rule:
- required_genesis: true  → SHA pinned to Genesis anchor, moves only on
  Sander's pin-bump cycle (vllm-nightly-mtp/dflash/full)
- required_genesis: false → free to ride latest nightly (vllm-nightly-clean)

Routes:
- qwen3-next-hybrid + qwen3-next-moe → vllm-nightly-mtp (Genesis-locked)
- gemma4-swa-moe                     → vllm-nightly-clean (unconstrained)
- gemma4-swa-dense                   → either
- dense                              → either

Also adds qwen3-next-moe + gemma4-swa-moe to vllm-nightly-mtp's family
list (factual, independent of SHA), corrects misleading notes that
referenced an unmerged pin bump, and bumps engine-count assertion in
test-profiles-compat.sh from 6 to 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 abf0e327f9 feat(profiles): add Gemma 4 26B-A4B ModelProfile + num_global_kv_heads field
Onboards Gemma 4 26B-A4B (gemma4-swa-moe family) as the fourth
ModelProfile. Real values sourced from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json
(file downloaded but checkpoint still streaming in background as of
this commit).

Schema extension: new optional field `num_global_kv_heads`. Gemma 4
26B-A4B has asymmetric KV head counts (8 sliding / 2 global) — distinct
from Gemma 4 31B which uses 16 for both. Field is None for symmetric
models; existing Gemma 4 31B + Qwen models work unchanged.

Confirmed architectural facts (from config.json):
- 30 total hidden layers (smaller than estimated 40-48)
- 5 full_attention layers at indices [5, 11, 17, 23, 29]
  (one global per 6-layer block; last layer always global)
- 25 sliding_attention layers (sliding_window=1024)
- num_kv_heads=8 (sliding), num_global_kv_heads=2 (global) — ASYMMETRIC
- head_dim=256 (sliding), global_head_dim=512 (global)
- attention_k_eq_v: TRUE (K=V tied)
- num_experts: 128, top_k_experts: 8 (active per token)
- moe_intermediate_size: 704
- Multimodal: vision_config + audio token IDs present
- requires_genesis: false (Gemma 4 doesn't need Genesis like Qwen)

KV math implication: per-token growing KV at bf16 is
5 × 2 × 512 × 1 (K=V tied) = 5,120 bytes
vs Gemma 4 31B's 10 × 16 × 512 × 1 = 81,920 bytes per token
=> Gemma 4 26B-A4B has ~16x smaller growing-KV per token than 31B,
   making long context dramatically cheaper. KV_MATH.md MoE section
   estimates need refresh.

Drafter compat note: references gemma-it-assistant (the existing 31B
drafter). TODO Codex: extend that DrafterProfile's model_compat to
include gemma-4-26b-a4b, OR split into a separate
gemma-26b-it-assistant DrafterProfile (google/gemma-4-26B-A4B-it-assistant
is downloading in background alongside the main model).

Tests pass: 41 tests, kv-calc 18/18 invariant held.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 937871492a feat(profiles): add Qwen 3.6 35B-A3B ModelProfile (MoE schema extensions)
Onboards Qwen 3.6 35B-A3B (qwen3-next-moe family) as the third
ModelProfile in the v0.7.0 profile catalog. Architectural values sourced
from the actual config.json at /mnt/models/huggingface/qwen3.6-35b-a3b-
autoround-int4/config.json, NOT estimates.

Schema extensions (additive, optional — existing dense models work
unchanged):
- num_experts, num_experts_per_tok
- moe_intermediate_size, shared_expert_intermediate_size
- active_params_b (documentation only)
- mtp_num_hidden_layers (built-in MTP indicator)
- attn_output_gate (gated attention bool)
- vision_capable (multimodal support flag)

Confirmed architectural facts (from config.json):
- 40 total hidden layers
- 10 full_attention layers at indices [3,7,11,15,19,23,27,31,35,39]
  (full_attention_interval=4)
- 30 linear_attention (Gated DeltaNet) layers
- num_kv_heads=2 caps valid TP at [1, 2]
- 256 experts, 8 active per token (corrects KV_MATH.md's 128 estimate)
- mtp_num_hidden_layers=1 — built-in MTP drafter
- vision_config present — multimodal

5 weight variants on disk:
- autoround_int4 (production, 20 GB)
- gptq_int4 (experimental, 22 GB)
- gguf (production, 90 GB)
- dflash + dflash_gguf (experimental)

What's NOT in this commit (still to do for v0.7.3):
- Build first vllm/dual compose for the new model
- Take down llama estate fixtures + boot the new compose
- Capture boot log + back-solve per_token_bytes against KV math
- 4+ calibration anchors → calibration YAML
- kv-calc activation coefficients for Qwen 3.6 35B-A3B
- BENCHMARKS rows
- learnings/qwen3.6-35b-a3b.md
- KV_MATH.md MoE section update (remove "calibration pending", fill in
  measured coefficients)

Tests:
- test-profiles-compat.sh: 41 tests pass (was 34 pre-v0.7.2 topology
  + 7 topology + bumped model count 2 → 3)
- kv-calc --calibration: 18/18 = 100% (held)
- All 7 other test suites pass

Branch reasoning: feature branch for Codex review before merge to
master. Live boot + calibration discipline benefits from Codex's
focused-session bench rigor.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunna d116ba9ba3 feat(launch): add hardware topology advisor
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m22s
Release / release (push) Failing after 48s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-15 09:54:04 +00:00
noonghunna 2a148d702b feat(bench): surface prompt processing throughput
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m23s
Release / release (push) Failing after 45s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-15 09:03:10 +00:00
noonghunna c2adb3970d feat(scripts): add diagnose-profile triage
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 22m22s
Release / release (push) Failing after 46s
Build vLLM Club3090 Image / Detect self-hosted GPU runner (push) Canceled after 0s
Build vLLM Club3090 Image / GPU smoke and alias promotion (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-14 22:07:00 +00:00
noonghunna c3063838bf feat(launch): export profile vllm pins 2026-05-14 21:38:55 +00:00
noonghunna 40f1ef78f8 feat(profiles): resolve vllm nightly pins 2026-05-14 21:38:47 +00:00
noonghunna 52e43470c6 fix(launch): persist estate source of truth 2026-05-14 20:18:12 +00:00
noonghunna c9b153f91f feat(launch): add estate planner orchestration 2026-05-14 17:04:38 +00:00
noonghunna a142b1ce10 feat(launch): validate single-model profiles 2026-05-14 16:38:50 +00:00
noonghunna 6581ccca9e feat(compat): add profile validator and estate self-test 2026-05-14 15:25:28 +00:00
noonghunna 69825d7dae feat(profiles): ship v0.7.0 data layer 2026-05-14 14:46:42 +00:00
noonghunna 5882bbef6f feat(launch): add hardware-aware launcher 2026-05-13 23:43:32 +00:00
noonghunna 12a33fbdd1 feat(scripts): add hardware-aware setup picker 2026-05-13 20:44:35 +00:00
noonghunna ef770322f4 feat(scripts): add submit-bench flow
Release / release (push) Failing after 51s
2026-05-13 18:24:14 +00:00
noonghunna 26985527f7 Add hardware-aware compose preflight
Release / release (push) Failing after 50s
2026-05-13 18:03:37 +00:00