41 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 421114b9dd fix(vllm-pr35936): make overlay tolerate bf610c2f upstream drift (closes #144)
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m28s
Release / release (push) Failing after 50s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
The PR #35936 required-tool-fallback overlay was captured against vLLM
source HEAD a2696294 / image nightly-1acd67a7. The v0.7.3 engine-pin
policy split (3b2d940) routed 8 non-TQ3 Qwen 3.6-27B composes to
vllm-nightly-clean (bf610c2f), where three upstream API drifts broke the
overlay. All three are now defensively patched and validated end-to-end
on bf610c2f (boot + /v1/chat/completions + /v1/completions all green,
zero errors, system_fingerprint vllm-...gbf610c2f5):

1. engine/serving.py — `speech_to_text` relocated from
   entrypoints/openai/speech_to_text → entrypoints/speech_to_text and
   split into transcription/ + translation/ subpackages. try/except
   imports old path first (matches captured SHA), falls back to new.

2. chat_completion/serving.py — `RequestOutput.prompt_routed_experts`
   removed on the dense path in bf610c2f. getattr(..., None) so the
   overlay tolerates both layouts; dense path produces None as before.

3. engine/serving.py — bf610c2f's chat_completion / completion /
   responses serving call `self._with_kv_transfer_rejection_cleanup(...)`,
   a method absent from the captured base. Added a passthrough stub
   (club-3090 doesn't run disaggregated KV-transfer / remote-prefill).

Drift surface was bounded to these three points — verified by live boot
on bf610c2f, not just import-time. Engine-pin policy split (3b2d940)
stays intact; reverts d32e168/e7bca8e were undone.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 20:58:14 +00:00
noonghunna d7804107c9 Reapply "fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean"
This reverts commit d32e168a89.
2026-05-15 20:52:21 +00:00
noonghunna 0c4260f53d Revert "test(launch): re-align engine pin expectations after revert"
This reverts commit e7bca8e035.
2026-05-15 20:52:21 +00:00
noonghunna e7bca8e035 test(launch): re-align engine pin expectations after revert
Companion to the revert of 3b2d940. The test-launch-compat.sh
expectations were updated in 127f4f6 to expect CLEAN_SHA for vllm/dual;
since the routing went back to vllm-nightly-mtp (MTP_SHA = 01d4d1ad),
flip back. Removes the now-redundant CLEAN_SHA assertion in the
estate_cli compose_env block (both vllm/dual and vllm/dual-tq3-mtp now
resolve to MTP_SHA, so one assertion covers both).

All 8 test suites green.
2026-05-15 18:45:23 +00:00
noonghunna d32e168a89 Revert "fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean"
This reverts commit 3b2d940d26.
2026-05-15 18:40:12 +00:00
noonghunna 77802a3d8e docs(BENCHMARKS): @OVDEN13 dual.yml — PCIe Gen 4 x4+x8 asymmetric (#142)
First asymmetric PCIe-lane data point on the matrix. 75.80 narr / 99.04
code wall TPS, soak PASS (20×5 fresh, 0 errors, 100.7% retention,
0 MiB growth). −15% narr / −14% code vs @danbedford's same-compose
Gen 4 x16 PCIe-only A/B (#77) — gap consistent with the Gen 4 x4 card
being the bottleneck for cross-card NCCL allreduce. MTP per-position
acceptance also lower-floor (0.81/0.56/0.35) vs reference stable
(0.94/0.84/0.72), suggesting draft-forward sync stalls on the slow lane.
2026-05-15 17:30:44 +00:00
github-actions[bot] c3d1c5d67d chore(changelog): regenerate for v0.7.3 [skip ci] 2026-05-15 17:19:12 +00:00
noonghunna 27ca1a40bd feat(report): surface kv-calc calibration verdict
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m20s
Release / release (push) Failing after 52s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
When a user files a VRAM-OOM or context-ceiling bug, the maintainer's
first triage question is "does kv-calc still agree with measured
reality?" — a calibration failure means the projection model has drifted
from actual VRAM cost, so any "predicted PASS" verdict in compat or
acceptance projections becomes untrustworthy.

Run `tools/kv-calc.py --calibration` from report.sh and surface:
- The Overall: X/Y verdict line
- Any FAIL rows inline with an explicit "treat projections as suspect"
  warning
- Full calibration output in a <details> block

Skipped gracefully if python3 / tools/kv-calc.py are unavailable.
Output flows through the existing redact pipeline.
2026-05-15 22:18:09 +05:00
noonghunna 9a039d8c92 docs(soak-test): clarify PASS verdict semantics — closes #140
Soak-test PASS only verifies "no failure signal on this sample at this
depth," not "patches in the compose's overlay set are load-bearing for
this workload." On TP=2 / llama.cpp configs the topology itself takes
Cliff 2 off the table, so PASS on those composes can't attribute work
to any specific patch.

- soak-test.sh header docstring: new "PASS verdict semantics" block
- soak-test.sh --help: matching "PASS VERDICT" section
- soak-helper.py: PASS verdicts now print a one-line caveat pointing
  to scripts/soak-test.sh --help and docs/CLIFFS.md
- docs/CLIFFS.md: callout in "Why TP=2 escapes" explaining what a clean
  dual.yml soak does and does not validate

Patch-attribution path (rerun with overlays stripped) referenced in all
three surfaces.
2026-05-15 22:18:09 +05:00
noonghunna 39e18733aa feat(kv-calc): model v0.7.3 MoE architectures 2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 6dc9a0dce1 feat(gemma-4-26b-a4b): AWQ + MTP n=4 — +12% narr / +49% code over no-MTP baseline
New compose models/gemma-4-26b-a4b/vllm/compose/dual/awq-mtp.yml adds
external MTP drafter wiring via google/gemma-4-26B-A4B-it-assistant
(832 MB BF16, already downloaded). COMPOSE_REGISTRY entry
vllm/gemma-a4b-awq-mtp at port 8043.

Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:

  awq.yml        (no MTP) → 138.88 / 138.67 wall TPS (139 / 139 decode)
  awq-mtp.yml    (n=4)    → 155.05 / 207.02 wall TPS (157 / 211 decode)
                            +12% narr / +49% code

MTP metrics:
  AL: 3.04 narrative, 3.79 code
  Per-pos accept (narr): 0.77 / 0.55 / 0.40 / 0.29 (50.9% avg)
  Per-pos accept (code): 73.5% avg
  Both GPUs at 98% util / 357 W and 305 W (symmetric)

Cross-MoE finding (opposite direction from Qwen 35B-A3B MTP, see
preview-mtp.yml row in BENCHMARKS): Gemma's external assistant
drafter is a small dense model (~0.5 B params, Gemma4AssistantForCausalLM)
that BYPASSES the MoE expert routing entirely on the draft pass.
The Qwen built-in MTP head runs through the model's MoE forward, so
each draft step pays MoE routing + inter-GPU sync overhead — that's
why Qwen MTP n=3 was 50% SLOWER but Gemma MTP n=4 is 49% FASTER.

Practical rule for MoE models on Ampere: prefer external drafters
over built-in MTP heads where both are available. Documented in
learnings/gemma-4-26b-a4b.md + cross-ref in qwen3.6-35b-a3b.md.

Recommendation: awq-mtp.yml becomes the recommended default for
Gemma 26B-A4B going forward; awq.yml stays as the no-drafter A/B
reference.

PR #40886 overlay still applied via the same patches/install.sh
bind-mount pattern as awq.yml.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 e1d44bd732 feat(qwen-35b-a3b): preview-MTP compose + bench row — MTP measured SLOWER on MoE
New compose models/qwen3.6-35b-a3b/vllm/compose/dual/preview-mtp.yml
adds MTP n=3 via Qwen 35B-A3B's built-in head (mtp_num_hidden_layers=1).
COMPOSE_REGISTRY entry vllm/qwen-a3b-preview-mtp at port 8052.

Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:

  preview.yml         (no MTP) → 182.68 / 177.45 wall TPS
  preview-mtp.yml     (n=3)    →  90.36 / 115.33 wall TPS
                                  −51% narr / −35% code

Per-position MTP acceptance is healthy (0.927 / 0.810 / 0.698, AL 3.44,
81.2% avg) — the draft head is working correctly. The bottleneck is
inter-GPU sync overhead on the MoE forward path: asymmetric GPU util
(GPU 0 at 39% / 233 W vs GPU 1 at 79% / 191 W; non-MTP run had both
cards at 85-99%) shows one card stalling on draft-target communication.

vLLM warns at boot: "max_num_scheduled_tokens=4096 may lead to
suboptimal performance ... consider increasing max_num_batched_tokens
to accommodate the additional draft token slots."

Hypothesis: Cliff 2 mitigations (Genesis PN12 / PN25 / PN34) carry
the scheduler-side optimizations that make MTP profitable for
Qwen3-Next family on Ampere. Without them (preview path is no-Genesis
by definition), the MoE draft pass cost dominates the acceptance gain.

Re-test triggers (each separately worth trying):
  (1) Genesis v7.73.x re-anchors on post-#42521 nightly
  (2) bump max_num_batched_tokens to ~12K
  (3) MTP n=2 to see if smaller spec depth changes calculus

For v0.7.3 ship: preview.yml (no MTP) stays the recommended default
for Qwen 35B-A3B users. preview-mtp.yml ships as documented A/B
reference + calibration anchor #2.

Compose, BENCHMARKS row, and the qwen3.6-35b-a3b learnings file all
flag the finding honestly so users can A/B on their own hardware.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 273c017646 docs(UPSTREAM): add PR #41800 truncate_prompt_tokens row
Records the vendored overlay landed in 1872880 in the canonical
pin-bump tracker:

- Merged upstream 2026-05-06 at commit d5b31c95
- Wired into 18 affected composes (vllm-nightly-mtp / -dflash / -full
  pin targets)
- vllm-nightly-clean (bf610c2f) is post-fix and doesn't need the
  overlay (install.sh has upstream-fix detection that no-ops on it)
- Drop trigger per engine spelled out:
  * vllm-nightly-mtp: requires Genesis v7.73.x re-anchor
  * vllm-nightly-dflash: requires PR #41703 overlay re-validation on
    newer base
  * vllm-nightly-full: requires PR #42102 overlay re-validation on
    newer base
- Tracking issue #139 + triggered-by #138 (SEVENID) referenced

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 1d7aad112c feat(vllm-pr41800): vendor truncate_prompt_tokens overlay across all pre-fix engines (closes #139)
Resolves the HTTP 400 reported by opencode (and other agentic clients
sending `truncate_prompt_tokens` on chat-completion requests) on every
compose currently routing through `vllm-nightly-mtp` (01d4d1ad),
`vllm-nightly-dflash` (e47c98ef), or `vllm-nightly-full` (e47c98ef).
All three SHAs predate vLLM PR #41800 (merged 2026-05-06 at d5b31c95)
which adds `truncate_prompt_tokens` to `get_max_tokens()`'s signature.

This commit:

* `models/qwen3.6-27b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/`
  - install.sh: Python anchor-based in-place patcher for
    `vllm/entrypoints/utils.py`. Two anchored edits:
      1. signature: insert `truncate_prompt_tokens: int | None = None,`
      2. body: insert truncation-aware `input_length = min(...)` block
         before the existing `if max_model_len < input_length:` check
    Idempotent via a sentinel comment. Post-patch AST-validated.
    Includes upstream-fix detection: if the function signature already
    accepts `truncate_prompt_tokens` (e.g. on a post-#41800 nightly),
    install.sh no-ops cleanly so composes routing through
    vllm-nightly-clean (bf610c2f, post-fix) still boot.
  - README.md: full context, wiring pattern, smoke-test command,
    drop-trigger criteria.

* Wired into 18 affected composes (every compose with
  `Engine-profile: vllm-nightly-(mtp|dflash|full)` header) plus
  `qwen3.6-27b/dual/docker-compose.yml` (the latter is on
  vllm-nightly-clean post-engine-pin-split commit a67af0f and the
  overlay no-ops there cleanly — harmless safety wiring):

  Qwen 3.6 27B:
    dual: docker-compose.yml, dflash.yml, dflash-noviz.yml, int8.yml,
          tq3-mtp.yml, tq3-mtp-genesis.yml, tq3-nomtp.yml, turbo.yml
    single: docker-compose.yml, bounded-thinking.yml, long-text.yml,
            long-text-no-mtp.yml, long-vision.yml
    multi4: dflash.yml

  Gemma 4 31B:
    dual: awq.yml, dflash.yml, dflash-int8.yml, int8.yml, int8-tq3.yml

* NVLink-variant stubs (nvlink.yml, nvlink-turbo.yml,
  nvlink-dflash.yml, nvlink-dflash-noviz.yml) inherit the overlay
  automatically via `extends:` from the base compose — no separate
  patch needed.

Smoke-tested on `vllm/vllm-openai:nightly-01d4d1ad...` (Sander's
Genesis PROD pin): get_max_tokens signature lacks truncate_prompt_tokens
before install, has it after. Idempotency confirmed (second install =
no-op). Post-fix no-op confirmed on `bf610c2f` (vllm-nightly-clean).
All 18 wired composes parse as valid YAML; test-profiles-compat.sh ok.

For Pattern-B composes (Gemma 4 31B set, plus tq3-mtp-genesis.yml)
that have a `if/then/else` exec-vllm-serve branching for NVLink mode,
the install line is placed BEFORE the `if` block at the script's base
indent so it runs regardless of the NVLink path taken.

Tracking: #139 (root cause + plan); triggered by #138 (SEVENID's
opencode failure on dual-dflash-noviz). Drop when each affected engine
pin bumps past commit d5b31c95.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 e49c939748 docs(README): add v0.7.3 MoE models to Supported Models table
Two new rows in the Supported Models table:

* **Qwen 3.6 35B-A3B** ⭐ NEW v0.7.3 — preview track (production-track
  blocked on Genesis v7.73.x). MoE 256 experts × 8 active (~3 B active
  params), upstream native loader via vLLM PR #42521. Preview dual:
  182/177 wall TPS at 16K — ~2× the Qwen 3.6-27B dense baseline.

* **Gemma 4 26B-A4B** ⭐ NEW v0.7.3 — production via AWQ path. MoE
  128 experts × 8 active (~4 B active params). Intel AutoRound INT4
  variants are Ampere-blocked (Marlin K-dim alignment); the cyankiwi
  AWQ-4bit weights work on Ampere via vendored vLLM PR #40886
  (compressed-tensors MoE key remapping). AWQ dual: 139/139 wall TPS
  at 32K, CV 0.2% / 0.0%.

Both rows surface the key trade-offs (Genesis-pending for Qwen,
AutoRound-blocked for Gemma on Ampere) so users picking a model see
the constraints up front.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 92b69bd220 docs(BENCHMARKS): Gemma 4 26B-A4B AWQ first row + AutoRound row demoted
Replace the AutoRound boot-fail row with two rows that capture both
the working path and the documented blocker:

* **`gemma-4-26b-a4b/dual/awq.yml` (TP=2)** ⭐ — first measured row
  on the AWQ path. 138.88 / 138.67 wall TPS, decode 139.92 / 139.98,
  CV 0.2% / 0.0%, 23.45 GB/card, TTFT 53 ms.
  Engine vllm-nightly-clean + PR #40886 overlay (compressed-tensors
  AWQ MoE key remapping). Compose at
  models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml. Weights:
  cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit (17 GB).

* `dual/docker-compose.yml` (Intel AutoRound INT4) — still documented
  as Ampere-blocked (Marlin K-dim alignment), but now framed as the
  SM90+ path with AWQ as the Ampere alternative just above.

Key observation: identical narr/code wall TPS (139 / 139) matches the
MoE memory-bandwidth-bound signature seen on Qwen 3.6 35B-A3B preview
(187 / 187 decode) — both confirm that per-token weight reads dominate
for these models on Ampere, prompt content distribution doesn't matter.
Gemma's 138 vs Qwen's 182 (both narr wall) is consistent with 4 B vs
3 B active params per forward.

CV 0.0% on code is the tightest measurement we've recorded on the
matrix — characteristic of a deterministic MoE decode path where
expert routing repeats predictably on similar prompts.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 0053444e84 feat(gemma-4-26b-a4b): AWQ path via vLLM PR #40886 overlay
The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.

vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.

This commit:

* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
  install.sh: anchor-based Python patcher that inserts the 4 branches
  into the in-container gemma4.py at runtime. Idempotent (sentinel
  comment), drift-resistant (anchors on existing branch line, not
  line numbers). Smoke-tested against bf610c2f: sentinel count=1
  after first install, no-op on second install, file remains valid
  Python after patching.
  README.md: full vendor context, usage pattern, drop trigger.

* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
  TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
  /mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
  Entrypoint runs install.sh before vllm serve. Routes through
  vllm-nightly-clean (bf610c2f, no Genesis).

* ModelProfile gemma-4-26b-a4b.yml updates:
  - autoround_int4_mixed: status flipped to "ampere-blocked" with
    a one-line forensic note (was "production", incorrect on SM86)
  - awq_compressed_tensors: new variant pointing at the cyankiwi
    AWQ-4bit weights, marked production via PR #40886 overlay
  - default_weight_variant: switched to awq_compressed_tensors
    (Ampere users get the variant that boots by default)

* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
  path at port 8042.

* vllm-nightly-clean engine: supported_weight_formats gains
  "compressed-tensors" so fits() C14 accepts the AWQ variant.

All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.

Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.

Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
   was issue #137 / #138 territory — here we deliver the alternative
   that unblocks Gemma 26B-A4B for the v0.7.3 ship).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunna 99328b4cda feat(estate): add parallel boot mode 2026-05-15 22:18:09 +05:00
noonghunna 127f4f6d8f test(launch): align engine pin expectations 2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 bdfb939edd docs(BENCHMARKS): add v0.7.3 MoE preview section
New section "MoE models (v0.7.3 — preview track)" between Gemma 4 31B
and the Aider Polyglot 30 quality benches. Captures the first two MoE
rows for v0.7.3 onboarding:

* **Qwen 3.6 35B-A3B preview dual** (TP=2, fp8_e5m2 KV, 16K ctx) ⭐
  - 182.68 / 177.45 wall TPS, decode 186.98 / 186.90, CV 0.3-0.9%
  - 21.94 GB/card, TTFT 126 ms
  - ~2x the Qwen 3.6 27B dense `dual.yml` baseline on the same hardware
  - Engine `vllm-nightly-clean` (nightly `bf610c2f`, post-PR-#42521)
  - No drafter, no TQ3, no Genesis — production track lands in v0.7.4
    after Genesis v7.73.x re-anchors on a post-#42521 nightly

* **Gemma 4 26B-A4B `dual/docker-compose.yml`** — Ampere-blocked
  - `moe_intermediate_size=704` is not a multiple of Marlin's
    `group_size=128` (5.5x). SM86 has no WNA16 kernel for unaligned K
    on Intel AutoRound INT4; only SM90+ Cutlass W4A8 / Machete handle
    arbitrary shapes.
  - Same failure on TP=1 and TP=2; both Intel quant variants
    (int4-mixed-AutoRound and int4-AutoRound) hit the same error
  - Documented workaround: route through cyankiwi AWQ-4bit + vendor
    vLLM PR #40886 (compressed-tensors AWQ MoE keys remap, tested by
    PR author on RTX 3090 24 GB)
  - Existing AutoRound compose preserved for SM90+ rigs

Both rows follow the v0.7.3 row convention: VRAM captured, engine pin
called out, related upstream PRs linked.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 9cd854dbcb fix(gpu-mode): mode_off tears down estate-managed instances
Before this fix, gpu-mode off only stopped services it knew about
(stop_all_27b vLLM variants, stop_all_gemma vLLM variants, stop_comfyui,
and the SERVICES array). Instances booted via launch.sh --estate or
--estate-file persist via Docker `restart: unless-stopped` and so kept
running after gpu-mode off — appearing to "auto-restart". The user had
to know about launch.sh --down-estate to clean them up.

New stop_estate function:
- Reads ~/.club3090/estate.yml (constant ESTATE_YAML)
- No-ops if file missing, launch.sh unavailable, or estate list empty
- Otherwise calls launch.sh --down-estate <yaml> which cleanly removes
  each estate-managed container + its docker network

Called from mode_off only — keeps blast radius small. Other mode
switchers (mode_27b, mode_gemma, etc.) aren't touched on purpose:
estate instances and gpu-mode classic services can in principle share
a rig at TP=1 + small contexts, so the user should choose explicitly
whether to keep them up when switching gpu-mode targets. mode_off is
the "stop everything" intent — clearly should include estate.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 8cf38b0590 docs(HARDWARE): add note on PCIe Gen 3 + older CPU TP=2 headwind
New "Note for older host platforms" section between the SM86 long-ctx
note and the WSL2 section. Documents the 30-40% TP=2 throughput hit
that PCIe Gen 3 + pre-Zen3 / pre-2018 Xeon rigs see vs the reference
Gen 4 / Ryzen 5950X / EPYC rigs in BENCHMARKS.md.

Covers:
- Two compounding causes (PCIe bandwidth halved, weaker CPU per-core).
- Symptom (asymmetric GPU util during decode — communication-starved
  TP=2).
- Three concrete mitigations: enable persistence mode, prefer single-
  card paths over dual.yml, raise host RAM allocation on VMs.
- Worked example linking to issue #137 (Xeon Gold 6138 + Gen 3) and
  @lolren disc #18 (Ryzen 5950X + Gen 4) for the side-by-side.

Follows from comment on #137 that promised this addition.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 1a233cd9fd docs(KERNEL_MATRIX): add KV Cache Impact subsection
Adds a new subsection between the Core Modern Kernels and Engine Support
Matrix tables. Classifies each kernel/system by whether it impacts KV
cache size and how:

* FA2/FA3 + FlashInfer + TensorRT-LLM kernels: no KV size change (pure
  attention-compute optimizations / fusion).
* PagedAttention (vLLM): effective KV size win via fragmentation
  reduction.
* RadixAttention (SGLang): very strong KV win via prefix-tree sharing
  across requests.
* Block Diffusion (DFlash/Zaya): indirect but large — fewer KV updates
  per output token.

Each row includes "Best For" and "Notes for 3090 / Consumer" so the
table cross-references KERNEL_MATRIX's existing 3090-routing guidance.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 97195fe443 docs: add KERNEL_MATRIX.md (attention backend + engine support matrix)
New companion to DTYPE_MATRIX, KV_MATH, and INFERENCE_ENGINES. Covers:

* Core modern kernels (FlashAttention-2/3, FlashInfer, RadixAttention,
  PagedAttention, Triton custom, TensorRT-LLM kernels, MLA/FlashMLA)
  with Ampere (3090) maturity ratings.

* Engine support matrix (vLLM / SGLang / TensorRT-LLM / llama.cpp)
  across attention, prefix-cache, paged KV, FA3, DFlash, MTP/spec-decode,
  TQ3 / advanced KV quant, MoE / hybrid, structured output, and Ampere
  optimization — with 3090-specific notes per row.

* Routing guidance for the club-3090 stack:
  - Interactive/agents/RAG → SGLang (RadixAttention + DFlash)
  - Max raw throughput on Ada/Blackwell → TensorRT-LLM
  - Broad compatibility + stability on 3090 → vLLM + FlashInfer
  - Low VRAM / single-card → llama.cpp
  - Hybrid MoE (Qwen 3.6 35B-A3B, Gemma 4 26B) → SGLang or vLLM
    with FlashInfer

Cross-links to DTYPE_MATRIX (quantization), KV_MATH (cache math), and
INFERENCE_ENGINES (high-level engine comparison).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 3b2d940d26 fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean
Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).

Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text         single/tools-text.yml      fp8_e5m2
- vllm/minimal            single/minimal.yml         fp8_e5m2
- vllm/dual               dual/docker-compose.yml    fp8_e5m2
- vllm/dual-bf16          dual/bf16.yml              bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp  dual/qwopus-bf16mtp.yml  fp8_e5m2
- vllm/dual-nvlink        dual/nvlink.yml (extends)  fp8_e5m2
- vllm/dual4              multi4/docker-compose.yml  fp8_e5m2

TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.

Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
  Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
  Genesis is required only for TQ3 KV format; that's enforced at the
  compose level via Engine-profile selection.

Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
  coverage.

Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
  for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
  requires_genesis=true, so the Genesis enforcement happens at C15
  (engine feature) level, not C6. Test now asserts positive (Qwen +
  fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
  non-Genesis engine fails C15).

All compat tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 cf0451a7ad fix(gemma-4-31b): route default/bf16 composes to vllm-nightly-clean
Genesis is strictly required only for TQ3 KV on vLLM. The Gemma 4 31B
default composes don't use TQ3 — they don't need Genesis. They DO need
a post-PR-#41745 nightly for Gemma 4 MTP support (per compose header
breadcrumbs), which 01d4d1ad (the Genesis-anchored MTP pin) predates.

Affected composes (all bf16/fp8 KV, no TQ3):
- models/gemma-4-31b/vllm/compose/single/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/bf16.yml

All three flipped from `Engine-profile: vllm-nightly-mtp` to
`vllm-nightly-clean` (bf610c2f, 2026-05-15) which is post-everything
relevant: PR #41745 (Gemma 4 MTP), PR #42102 (DFlash + INT8 PTH), and
PR #42521 (qwen3_5_moe weight loading).

COMPOSE_REGISTRY entries vllm/gemma-mtp-tp1, vllm/gemma-mtp, and
vllm/gemma-bf16 updated to match.

The DFlash and INT8 paths (vllm-nightly-dflash on e47c98ef and
vllm-nightly-full on e47c98ef) are untouched — their pins remain
correctly anchored to the SHAs that produced all current BENCHMARKS
rows for those compose families. TQ3-using composes
(vllm/gemma-int8-tq3) likewise stay on vllm-nightly-full (Genesis).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 2d1b1dc347 feat(moe): add dual-card composes for Gemma 26B-A4B + Qwen 35B-A3B preview
Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.

* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
  TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
  comfortable KV headroom. No drafter on this base smoke compose.

* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
  TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
  port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
  ~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
  mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.

* COMPOSE_REGISTRY: dual variants become the canonical keys
  (vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
  single variants renamed to vllm/gemma-a4b-single and
  vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).

Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 87f0a0c528 fix(engines): vllm-nightly-mtp anchors to 01d4d1ad (Sander v7.72.2 PROD pin)
Commit 40f1ef78 (2026-05-14) accidentally set vllm-nightly-mtp.spec to
nightly-1acd67a7, which is the post-Gemma4-merge nightly used by the
gemma-4-31b compose. The Genesis v7.72.2 PROD pin is nightly-01d4d1ad:

- BENCHMARKS rows from 2026-05-05 onwards reference nightly-01d4d1ad3
- calibration/qwen3.6-27b.yml: engine_pin: vllm-nightly-01d4d1ad
- gemma-4-31b/vllm/compose/single header: "Qwen3.6 composes stay on the
  v7.72.2 PROD pin (01d4d1ad3) since Genesis allowlist anchors there"

Effect of the bug: any Qwen 27B launch via launch_compat.py since
2026-05-14 would have used the un-blessed Gemma-merge SHA (1acd67a7),
potentially causing Genesis patch apply failures or silent skips.
Existing benches that hardcoded the image or used the older mechanism
were unaffected.

Notes field updated with the corrected anchor + a PRIOR-BUG breadcrumb
so future readers don't re-introduce the bump.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 f7f6f444b9 feat(moe): wire Gemma 4 26B-A4B + Qwen 3.6 35B-A3B composes through fits()
Adds the v0.7.3 critical-path plumbing for both MoE models:

* Compose files:
  - models/gemma-4-26b-a4b/vllm/compose/single/docker-compose.yml
    (vllm-nightly-clean, TP=1, port 8040, bf16 KV, no drafter)
  - models/qwen3.6-35b-a3b/vllm/compose/single/preview.yml
    (vllm-nightly-clean, TP=1, port 8050, fp8_e5m2 KV, preview-only)

* COMPOSE_REGISTRY entries 'vllm/gemma-a4b' and 'vllm/qwen-a3b-preview'.

* DrafterProfile: new gemma-26b-it-assistant for the 26B-A4B
  (google/gemma-4-26B-A4B-it-assistant — separate weights from 31B);
  qwen-mtp-builtin.model_compat extended with qwen3.6-35b-a3b.

* vllm-nightly-clean.supported_model_families adds qwen3-next-moe as a
  preview/smoke route; notes clarify Cliff 2 / TQ3 unavailable here.

* Qwen 3.6 35B-A3B model: requires_genesis flipped false → upstream is
  strictly bootable on post-#42521 nightlies; Genesis is RECOMMENDED for
  production (TQ3 KV + Cliff 2 mitigations), not strictly required to boot.

* ModelProfile gains kv_calc_supported (defaults true). Both new MoE
  models set false because tools/kv-calc.py only models qwen3.6-27b and
  gemma-4-31b today (MoE + asymmetric KV + DeltaNet+attention hybrid all
  require new MODEL_SPECS + activation coefficients — queued for Codex).
  fits() now skips C12 for kv_calc_supported=false instead of failing.

All compat tests pass; both new composes fit on 1x-3090 (Gemma) and
1x-3090/2x-3090/1x-5090 (Qwen preview).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 15eda8a823 feat(profiles): split engine-pin policy by Genesis dependency
Add vllm-nightly-clean engine profile (latest nightly, no Genesis) for
model families that don't need DeltaNet stabilization patches.

Rule:
- required_genesis: true  → SHA pinned to Genesis anchor, moves only on
  Sander's pin-bump cycle (vllm-nightly-mtp/dflash/full)
- required_genesis: false → free to ride latest nightly (vllm-nightly-clean)

Routes:
- qwen3-next-hybrid + qwen3-next-moe → vllm-nightly-mtp (Genesis-locked)
- gemma4-swa-moe                     → vllm-nightly-clean (unconstrained)
- gemma4-swa-dense                   → either
- dense                              → either

Also adds qwen3-next-moe + gemma4-swa-moe to vllm-nightly-mtp's family
list (factual, independent of SHA), corrects misleading notes that
referenced an unmerged pin bump, and bumps engine-count assertion in
test-profiles-compat.sh from 6 to 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 abf0e327f9 feat(profiles): add Gemma 4 26B-A4B ModelProfile + num_global_kv_heads field
Onboards Gemma 4 26B-A4B (gemma4-swa-moe family) as the fourth
ModelProfile. Real values sourced from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json
(file downloaded but checkpoint still streaming in background as of
this commit).

Schema extension: new optional field `num_global_kv_heads`. Gemma 4
26B-A4B has asymmetric KV head counts (8 sliding / 2 global) — distinct
from Gemma 4 31B which uses 16 for both. Field is None for symmetric
models; existing Gemma 4 31B + Qwen models work unchanged.

Confirmed architectural facts (from config.json):
- 30 total hidden layers (smaller than estimated 40-48)
- 5 full_attention layers at indices [5, 11, 17, 23, 29]
  (one global per 6-layer block; last layer always global)
- 25 sliding_attention layers (sliding_window=1024)
- num_kv_heads=8 (sliding), num_global_kv_heads=2 (global) — ASYMMETRIC
- head_dim=256 (sliding), global_head_dim=512 (global)
- attention_k_eq_v: TRUE (K=V tied)
- num_experts: 128, top_k_experts: 8 (active per token)
- moe_intermediate_size: 704
- Multimodal: vision_config + audio token IDs present
- requires_genesis: false (Gemma 4 doesn't need Genesis like Qwen)

KV math implication: per-token growing KV at bf16 is
5 × 2 × 512 × 1 (K=V tied) = 5,120 bytes
vs Gemma 4 31B's 10 × 16 × 512 × 1 = 81,920 bytes per token
=> Gemma 4 26B-A4B has ~16x smaller growing-KV per token than 31B,
   making long context dramatically cheaper. KV_MATH.md MoE section
   estimates need refresh.

Drafter compat note: references gemma-it-assistant (the existing 31B
drafter). TODO Codex: extend that DrafterProfile's model_compat to
include gemma-4-26b-a4b, OR split into a separate
gemma-26b-it-assistant DrafterProfile (google/gemma-4-26B-A4B-it-assistant
is downloading in background alongside the main model).

Tests pass: 41 tests, kv-calc 18/18 invariant held.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 937871492a feat(profiles): add Qwen 3.6 35B-A3B ModelProfile (MoE schema extensions)
Onboards Qwen 3.6 35B-A3B (qwen3-next-moe family) as the third
ModelProfile in the v0.7.0 profile catalog. Architectural values sourced
from the actual config.json at /mnt/models/huggingface/qwen3.6-35b-a3b-
autoround-int4/config.json, NOT estimates.

Schema extensions (additive, optional — existing dense models work
unchanged):
- num_experts, num_experts_per_tok
- moe_intermediate_size, shared_expert_intermediate_size
- active_params_b (documentation only)
- mtp_num_hidden_layers (built-in MTP indicator)
- attn_output_gate (gated attention bool)
- vision_capable (multimodal support flag)

Confirmed architectural facts (from config.json):
- 40 total hidden layers
- 10 full_attention layers at indices [3,7,11,15,19,23,27,31,35,39]
  (full_attention_interval=4)
- 30 linear_attention (Gated DeltaNet) layers
- num_kv_heads=2 caps valid TP at [1, 2]
- 256 experts, 8 active per token (corrects KV_MATH.md's 128 estimate)
- mtp_num_hidden_layers=1 — built-in MTP drafter
- vision_config present — multimodal

5 weight variants on disk:
- autoround_int4 (production, 20 GB)
- gptq_int4 (experimental, 22 GB)
- gguf (production, 90 GB)
- dflash + dflash_gguf (experimental)

What's NOT in this commit (still to do for v0.7.3):
- Build first vllm/dual compose for the new model
- Take down llama estate fixtures + boot the new compose
- Capture boot log + back-solve per_token_bytes against KV math
- 4+ calibration anchors → calibration YAML
- kv-calc activation coefficients for Qwen 3.6 35B-A3B
- BENCHMARKS rows
- learnings/qwen3.6-35b-a3b.md
- KV_MATH.md MoE section update (remove "calibration pending", fill in
  measured coefficients)

Tests:
- test-profiles-compat.sh: 41 tests pass (was 34 pre-v0.7.2 topology
  + 7 topology + bumped model count 2 → 3)
- kv-calc --calibration: 18/18 = 100% (held)
- All 7 other test suites pass

Branch reasoning: feature branch for Codex review before merge to
master. Live boot + calibration discipline benefits from Codex's
focused-session bench rigor.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 b1c68b4fe3 docs(kv-math): extend k_v_tensors=N notation to sliding-KV formulas
Earlier passes harmonized the 4 GROWING-KV formulas to the
'× k_v_tensors=N × bpe' form, but the 2 SLIDING-KV formulas (Gemma 4
31B + Gemma 4 26B-A4B sliding-window portions) still used the old
'× 1 (K=V tied)' parenthetical.

All 6 formulas (4 growing + 2 sliding) now use consistent notation.

(Note: Grok's third-pass review claimed a leftover at Qwen 27B line
~211 — that's actually a Python comment in the config.json extraction
section, not a formula. Grok was hallucinating against a cached older
state. But the sliding-KV inconsistency was real and worth fixing.)

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 11:01:55 +00:00
noonghunnaandClaude Opus 4.7 54d8bf0b3e docs(kv-math): tighten k_v_tensors notation across all 4 formulas
Per Grok third-pass nit: parenthetical (N, K=V tie note) was redundant
with the architecture summary above each section. Switched to compact
'k_v_tensors=N' form across all 4 formulas:

- Qwen 27B: × k_v_tensors=2
- Qwen 35B-A3B: × k_v_tensors=2
- Gemma 31B: × k_v_tensors=1
- Gemma 26B-A4B: × k_v_tensors=1

K=V tie context still present in each section's architecture summary;
no information lost.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:56:49 +00:00
noonghunnaandClaude Opus 4.7 67de3eca3e docs(kv-math): third-pass Grok polish
1. Quick reference table — new "Vs Qwen 27B" column anchored at fp8_e5m2
   TP=2 (common production config). Makes the family hierarchy
   self-evident in the table without needing to read the commentary:
   - Qwen 3.6 27B: 1.00× (baseline)
   - Qwen 3.6 35B-A3B: 0.31× (~3.2× lighter)
   - Gemma 4 31B: 2.50× (~2.5× heavier)
   - Gemma 4 26B-A4B: 0.16× (~6.4× lighter)
   Footnote explains the anchor choice + that ratios shift slightly
   under other formats but family hierarchy is stable.

2. DeltaNet general formula — added worked example for Qwen 3.6 35B-A3B
   showing plug-and-chug of the formula with real config.json values
   (30 layers, 16 k_heads × 128, 32 v_heads × 128, conv_kernel=4, fp32).
   Result: ~3.5 MB at max_num_seqs=1, ~14 MB at seqs=4. Makes the formula
   actionable for future DeltaNet onboardings.

Skipped Grok #1 (Qwen 27B notation parenthetical tightening) — Grok
explicitly says "already good" and didn't suggest a concrete form.
Aesthetic-perfection rabbit hole; not worth a further commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:55:56 +00:00
noonghunnaandClaude Opus 4.7 78f94cacf0 docs(kv-math): second-pass Grok polish
1. Formula notation — flip from '× 2 (k_v_tensors — no K=V tie) × bpe'
   to '× k_v_tensors (2, no K=V tie) × bpe'. Variable name leads the
   factor; value is parenthetical. Applied across all 4 model formulas
   for consistency with the general formula's reading order.

2. Quick reference table — new section after the per-card budget
   composition. At-a-glance per-token growing-KV bytes for all 4 models
   across bf16/fp8/TQ3-or-INT8 at TP=1/2. Includes 'what jumps out'
   commentary calling out:
   - Gemma 26B-A4B ~16x smaller per token vs 31B
   - Qwen 35B-A3B ~3.2x smaller per token vs 27B
   - Sliding-window KV is fixed (not per-token) for Gemma family
   - TQ3 vs INT8 PTH family applicability

3. DeltaNet general formula — new subsection inside the general KV
   formula block. Single-line formula covering K state + V state +
   conv1d kernel state across num_gdn_layers × max_num_seqs streams,
   using the linear_* config fields. Future DeltaNet model onboardings
   can compute their state size from this formula directly without
   re-deriving from scratch.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:53:38 +00:00
noonghunnaandClaude Opus 4.7 3114399983 docs(kv-math): address Grok review feedback
1. Formula consistency — all per-token formulas now annotate the k_v_tensors
   factor explicitly:
   - Qwen 27B: '× 2 (k_v_tensors — no K=V tie)'
   - Qwen 35B-A3B: '× 2 (k_v_tensors — no K=V tie)'
   - Gemma 31B: '× 1 (k_v_tensors — K=V tied)'
   - Gemma 26B-A4B: '× 1 (k_v_tensors — K=V tied)'
   Ties the per-model math back to the general formula's named variable.

2. MoE column format — summary table now uses compact 'Yes (256×8)' /
   'Yes (128×8)' notation (experts × active-per-token). Note added
   explaining the convention.

3. Qwen 35B-A3B router workspace math — corrected for real 256 experts
   (was using estimate-era 128). Now shows ~1 MB per router × 40 layers
   ≈ 40 MB total.

4. DeltaNet recurrent state — new dedicated subsection in both Qwen 3.6
   27B and Qwen 3.6 35B-A3B sections. Concrete sizes (~7.5 MB for 27B,
   ~3.5 MB for 35B-A3B per stream at seqs=1) computed from
   linear_num_*_heads + linear_*_head_dim + linear_conv_kernel_dim.
   Distinguishes the recurrent state (small, persistent) from the
   block-wise intermediate during forward (large, captured in activation
   peak §). Grok suggested ~50-150 MB total but that overstates by ~10×;
   I used computed values from the real config.json fields instead.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:50:22 +00:00
noonghunnaandClaude Opus 4.7 6ec6a6761f docs(kv-math): config-verify Qwen 35B-A3B + Gemma 26B-A4B MoE sections
Both MoE model sections rewritten from "math-ready estimates" to
"config-verified, calibration pending". Architectural facts now sourced
from on-disk config.json + layer_types arrays rather than guesses.

Qwen 3.6 35B-A3B corrections (real values from
/mnt/models/huggingface/qwen3.6-35b-a3b-autoround-int4/config.json):
- num_experts: 256 (was estimated 128 — 2x correction)
- Confirmed full_attention indices [3,7,11,15,19,23,27,31,35,39]
  via full_attention_interval=4 and layer_types array
- Built-in MTP (mtp_num_hidden_layers=1) confirmed
- attn_output_gate=True (gated attention) confirmed
- 5 quant variants on disk (autoround_int4 production, plus gptq_int4,
  gguf, dflash, dflash-gguf)

Gemma 4 26B-A4B major rewrite (real values from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json):
- 30 total layers (was estimated 40-48 — smaller than expected)
- 5 full_attention layers at indices [5,11,17,23,29]
- 25 sliding_attention layers
- *** ASYMMETRIC KV HEADS *** (the architectural surprise):
  - num_key_value_heads: 8 for sliding layers
  - num_global_key_value_heads: 2 for global layers (distinct field)
- Per-token growing KV at bf16 = 5 × 2 × 512 × 1 = 5,120 bytes
  vs Gemma 4 31B's 81,920 bytes — *** ~16x smaller per token ***
- 128 experts (was estimated 64), 8 active per token
- moe_intermediate_size: 704
- Multimodal: vision + audio token IDs
- requires_genesis: false (Gemma 4 has no DeltaNet quirks)
- Per-card budget projection: ~10-12 GB at 200K + fp8 KV →
  massive headroom on 24 GB; single-card serving may be viable at 262K

Top intro table + architecture summary table both updated with
verified values; status changed from "Math-ready" to "Config-verified".

Calibration remains pending: empirical activation-peak coefficients
need ≥4 measured BENCHMARKS rows per model before locking in.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:42:43 +00:00
github-actions[bot] 5cb993be99 chore(changelog): regenerate for v0.7.2 [skip ci] 2026-05-15 09:58:10 +00:00
noonghunna d116ba9ba3 feat(launch): add hardware topology advisor
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m22s
Release / release (push) Failing after 48s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-15 09:54:04 +00:00
github-actions[bot] 6775d10919 chore(changelog): regenerate for v0.7.1 [skip ci] 2026-05-15 09:25:34 +00:00
73 changed files with 3146 additions and 134 deletions
+16
View File
@@ -115,6 +115,7 @@ Primary serving model. Hybrid Qwen3-Next architecture (DeltaNet GDN + standard a
| `bounded-thinking.yml` | @lolren (2× 3090 PCIe + Ryzen 9 5950X, 250W cap, **MTP-disabled-suspected**) | TQ3 | 180K | 64.86 / 64.96 (CV **0.1%**) | — | ~22.3 GB | 2026-05-05 | **Anomaly:** lolren reports "no spec-decode" on this run despite `bounded-thinking.yml` shipping `--speculative-config mtp n=3` by default. Near-identical narr=code TPS + extreme CV stability (0.1%) suggests MTP was inactive — likely because his image was older `nightly-7a1eb8ac2` (pre-v7.72.2 + pre-PN35). Re-test on `nightly-01d4d1ad3` should restore MTP path → expect ~50/66 narr/code with normal CV. Tracked. [Disc #18](https://github.com/noonghunna/club-3090/discussions/18#discussioncomment-16820303). |
| `dual.yml` | @JDWarner (**Mixed RTX A5000 + RTX 3090**, both **Razer Core X eGPU enclosures over Thunderbolt 3**, Intel NUC11TNH i5-1135G7, **16 GB RAM**, headless, A5000=230W cap / 3090=290W cap, PCIe **x4 Gen 3** per card) | fp8 | 262K | **56.83 / 72.47** (soak p50 93.09) | — | ~23.6 GB/card | 2026-05-09 | **Soak: ✓ PASS** (5×5, 0 errors, 0 silent-empty, 100% TPS retention, 0 MiB growth). **First TB3 dual-eGPU + mixed-arch cross-rig data**. The setup that "shouldn't work": each card on a separate TB3 controller → ~3.94 GB/s effective per card vs ~32 GB/s on PCIe x16 Gen 4 (~8× cut), mixed Ampere SKUs (workstation A5000 + consumer 3090 with different mem bandwidth + clocks), 16 GB system RAM total. **Result: matches `dual.yml` PCIe x16 baseline within run-to-run noise** — confirms decode on Qwen3.6-27B is per-card-bandwidth bound, cross-card NCCL allreduce is small enough that even an 8× link cut doesn't dominate. Extends @aaronlockhartdev's #91/#95 finding (patched-P2P only +2%/+9% on `dual.yml`) in the opposite direction: even with 8× *less* cross-card bandwidth, decode holds. MTP AL 3.39-3.52, per-pos accept 0.93/0.83/0.70 (89% avg). verify-full + verify-stress all PASS. Genesis pin `7b9fd319` (v7.72.2). [Issue #107](https://github.com/noonghunna/club-3090/issues/107). |
| `dual/docker-compose.yml` (default) | @ygafarov (**3090 via USB4 eGPU dock + 5070 Ti via OCuLink** — heterogeneous Ampere + Blackwell consumer dual-eGPU, AMD Ryzen AI MAX+ 395 / Strix Halo miniPC, CachyOS, 123 GB RAM, 290 W cap both cards, PCIe **x4** per card — USB4 ≈ 3.94 GB/s, OCuLink ≈ 7.88 GB/s) | fp8 | 200K | **65.10 / 85.81** | — | 17.1 / 15.7 GB | 2026-05-12 | **First heterogeneous Ampere + Blackwell consumer dual-eGPU on the matrix.** TP=2 bound by the slower USB4 link in allreduce + sm_86 kernels (5070 Ti spends back-half of step waiting — 91% util but only 125 W out of 290 W cap). KV pool 200K @ 1.00× concurrency — VRAM cap from the 5070 Ti's 16 GiB (model takes 13.8 GiB/card → only ~2.2 GiB left for KV on the smaller card). verify-stress 7/7 incl. **91K needle recall** (Cliff 2 clean). Soak ⚠ borderline (360 MiB > 200 MiB threshold — same eGPU-bus accretion as ygafarov's own #113 single-card row above at 240 MiB; 100% TPS retention + 0 silent-empty + 0 errors so not a leak). MTP AL 3.50, per-pos accept 0.94/0.86/0.70. CV 4.5%/1.8%. **Slower than ygafarov's own single-3090 #113 row** (68.86/91.70 at 48K) — on this rig the single-card path is recommended; the 5070 Ti adds VRAM cap pain without TPS gain. Driver 595.71.05, vLLM `nightly-1acd67a79`, no Genesis (Blackwell consumer not on allowlist). [Issue #120](https://github.com/noonghunna/club-3090/issues/120). |
| `dual.yml` | @OVDEN13 (2× 3090 + Ryzen 7 5700X, **PCIe Gen 4 x4 + x8 asymmetric**, 300 W cap, no NVLink, Ubuntu 26.04 / driver 595.58.03 / CUDA 13.2, 92 GB RAM, bare metal) | fp8 | 262K | **75.80 / 99.04** | — | 22.3 GB/card | 2026-05-15 | **First asymmetric PCIe-lane data point on the matrix** (Gen 4 x4 to one card, x8 to the other — half the cross-card bandwidth of @danbedford's Gen 4 x16 baseline). **−15% narr / −14% code vs @danbedford's same-compose PCIe-only A/B** (89.24 / 114.57, [#77](https://github.com/noonghunna/club-3090/issues/77)) — gap consistent with the Gen 4 x4 card being the bottleneck for cross-card NCCL allreduce. MTP AL 2.71–3.51, per-pos accept ranges 0.81/0.56/0.35 → 0.95/0.85/0.72; the wide variance and lower floor vs @lolren's stable 0.94/0.84/0.72 ([disc #18](https://github.com/noonghunna/club-3090/discussions/18#discussioncomment-16820303)) suggest draft-forward sync stalls on the slow lane. Both GPUs pegged at 297 W power cap (80% / 87% util — mildly asymmetric, consistent with x4+x8 imbalance). CV 1.2% narr / 4.2% code. **Soak: ✓ PASS** (20×5 fresh mode, p50 101.41 TPS, p95 TTFT 1645 ms, 0 errors, 0 silent-empty, 100.7% retention, 0 MiB growth). Stack: master @ `b1c68b4` (v0.7.2-7-g, pre-v0.7.3 tag), vLLM `nightly-1acd67a795eb`, Genesis pin `7b9fd319`. [Issue #142](https://github.com/noonghunna/club-3090/issues/142). |
### Quad-card (4× RTX 3090, TP=4)
@@ -256,6 +257,21 @@ None close the **-13% narr / -11% code gap to 3dluvr's anchor**. Remaining gap l
| `dual-dflash.yml`-shape forced TP=1 (mem-util 0.96, max-model-len 12000) | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | **12K** | **150.40 / 261.06** (decode 151.16 / 264.62) | 28.8 GB | 2026-05-07 | **First single-5090 Gemma 4 DFlash data point.** Trade vs MTP row above: ~6% narr loss, **+21% code lift** (215→261). 1st-warmup TTFT outlier (73 s) suggests cudagraph warmup taking longer on first request; subsequent warmups stable at <40 ms. CV 3.6%/2.8%, peak 440 W. **Required mem-util 0.96 + max-model-len 12K** to fit BF16 weights + DFlash N=5 drafter on 32 GB — DFlash drafter footprint pushes out ctx ceiling vs MTP's 32K. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
---
## MoE models (v0.7.3 — preview track)
First-pass numbers for the two MoE models onboarded in v0.7.3. **Preview** track = `vllm-nightly-clean` engine (no Genesis patches, no TQ3 KV, no MTP yet) — exercises the upstream MoE loader [PR #42521](https://github.com/vllm-project/vllm/pull/42521) (qwen3_5_moe weight loading, merged 2026-05-14) without other overlays. Production track (Genesis-anchored, MTP, longer context) lands in v0.7.4 after Genesis v7.73.x re-anchors on a post-#42521 nightly.
| Compose | Rig | KV | Max ctx | Narr / Code TPS | PP tok/s | AL | Per-pos accept | Peak VRAM | Date | Notes |
| --- | --- | --- | ---: | ---: | ---: | ---: | --- | ---: | --- | --- |
| **`qwen3.6-35b-a3b/dual/preview.yml` (TP=2)** ⭐ | @noonghunna (2× 3090 PCIe, 230 W cap) | fp8_e5m2 | 16K | **182.68 / 177.45** (decode 186.98 / 186.90) | — | n/a (no drafter) | n/a | **21.94 GB/card** | 2026-05-15 | **First v0.7.3 MoE preview row on the matrix.** Engine `vllm-nightly-clean` (nightly `bf610c2f`, post-#42521). No Genesis, no MTP — exercises the upstream qwen3_5_moe loader cleanly. CV **0.3% / 0.9%** (very stable). TTFT 126 ms. GPU 0 at 99% util / 292 W, GPU 1 at 85% / 244 W. Decode TPS basically identical narr vs code (187 / 187) — characteristic of MoE memory-bandwidth-bound decode (3 B active params per forward, weights fit cache). **~2× the Qwen 3.6 27B dense `dual.yml` baseline** (89/118 wall) on the same hardware — MoE's active-params advantage on Ampere. Next steps: TQ3 KV + MTP (built-in head) after Genesis v7.73.x lands; longer ctx after upstream MoE expert dispatch overhead is measured. Compose: `models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml`. |
| `qwen3.6-35b-a3b/dual/preview-mtp.yml` (TP=2, MTP n=3, built-in head) ⚠️ | @noonghunna (2× 3090 PCIe, 230 W cap) | fp8_e5m2 | 16K | **90.36 / 115.33** (decode 91.48 / 119.65) | — | **3.44** (narr) | 0.927 / 0.810 / 0.698 | 22.72 GB/card | 2026-05-15 | **MTP MAKES THINGS SLOWER on Qwen MoE preview path.** A/B vs preview.yml above (same config, same nightly, same hardware): **−51% narr / −35% code wall TPS** despite 81.2% avg draft acceptance and AL 3.44. Per-position accept (0.927 / 0.810 / 0.698) is healthy — the draft head works correctly. Surfacing cause: asymmetric GPU util (GPU 0 at **39%** / 233 W vs GPU 1 at 79% / 191 W; non-MTP run had both at 85-99%) indicates the draft forward pass on MoE imposes inter-GPU sync overhead the acceptance gain can't amortize. Vendor warning at boot: `max_num_scheduled_tokens=4096 ... suboptimal performance ... increase max_num_batched_tokens to accommodate the additional draft token slots`. **CV 5.4% / 2.0%** (less stable than no-MTP). VRAM +0.78 GB/card vs no-MTP (MTP head workspace). **Practical implication**: for v0.7.3, route Qwen 35B-A3B users to `preview.yml` (no MTP) as the default — MTP gives worse latency despite high acceptance. Re-test after (a) Genesis v7.73.x re-anchors on a post-#42521 nightly (Cliff 2 mitigations may unblock the underlying scheduler overhead), (b) max_num_batched_tokens raised to ~12K to satisfy the vLLM warning, or (c) MTP n=2 to see if smaller spec depth changes the calculus. Compose: `models/qwen3.6-35b-a3b/vllm/compose/dual/preview-mtp.yml`. |
| **`gemma-4-26b-a4b/dual/awq.yml` (TP=2)** ⭐ | @noonghunna (2× 3090 PCIe, 230 W cap) | bf16 | 32K | **138.88 / 138.67** (decode 139.92 / 139.98) | — | n/a (no drafter) | n/a | **23.45 GB/card** | 2026-05-15 | **First v0.7.3 Gemma MoE production-track row.** Engine `vllm-nightly-clean` (nightly `bf610c2f`) + vendored [vLLM PR #40886](https://github.com/vllm-project/vllm/pull/40886) overlay (compressed-tensors AWQ MoE key remapping — applied at boot via anchor-based Python patcher in `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/install.sh`). Weights: `cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit` (17 GB). CV **0.2% / 0.0%** — extraordinarily stable, characteristic of MoE memory-bandwidth-bound decode. TTFT 53 ms. GPU 0 at 98% util / 360 W, GPU 1 at 59% / 299 W. **Identical narr/code wall TPS (139 / 139)** — same MoE signature as Qwen 3.6 35B-A3B preview above (per-token weight reads dominate, prompt content distribution doesn't matter). ~76% of Qwen 35B-A3B preview's narr TPS — consistent with 4 B vs 3 B active params per forward. Compose: `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml`. |
| **`gemma-4-26b-a4b/dual/awq-mtp.yml` (TP=2, MTP n=4)** ⭐⭐ | @noonghunna (2× 3090 PCIe, 230 W cap) | bf16 | 32K | **155.05 / 207.02** (decode 156.68 / 210.61) | — | **3.04 narr / 3.79 code** | 0.77 / 0.55 / 0.40 / 0.29 (narr); 50.9% avg (narr), 73.5% avg (code) | 23.50 GB/card | 2026-05-15 | **MTP boosts Gemma MoE: +12% narr / +49% code over awq.yml**. Engine `vllm-nightly-clean` + PR #40886 overlay + external `google/gemma-4-26B-A4B-it-assistant` drafter (832 MB BF16). Both GPUs symmetric at 98% util — **the external-drafter path doesn't pay the inter-GPU sync penalty that the Qwen 35B-A3B built-in MTP head pays** (see `preview-mtp.yml` row above where MTP made things 50% slower). Mechanism: external drafter is a small dense model (`Gemma4AssistantForCausalLM`, ~0.5 B params) — the draft forward bypasses MoE expert routing entirely. The 32-49% code lift confirms MTP is structurally compatible with Gemma's hybrid-SWA attention on the AWQ compressed-tensors path. **Recommended default for Gemma 26B-A4B going forward; awq.yml stays as the no-drafter A/B reference.** Compose: `models/gemma-4-26b-a4b/vllm/compose/dual/awq-mtp.yml`. |
| `gemma-4-26b-a4b/dual/docker-compose.yml` (TP=2, Intel AutoRound INT4) | @noonghunna (2× 3090 PCIe) | — | — | **boot fail (SM86)** | — | — | — | — | 2026-05-15 | **Ampere-blocked on Intel AutoRound INT4.** `moe_intermediate_size=704` is not a multiple of `group_size=128` (5.5×); Marlin K-dim alignment fails. SM86 has no WNA16 kernel for unaligned K-dim — only SM90+ Cutlass W4A8 / Machete handle arbitrary shapes. Same failure on TP=1 (no split) and TP=2 (split to 352 per rank). Both Intel quant variants (`int4-mixed-AutoRound` and `int4-AutoRound`) hit the same error. **AWQ is the Ampere path** (row above). AutoRound compose preserved here as a documented-blocker for SM90+ rigs (RTX 5090 / Pro 6000 should boot it). |
---
## Quality benches — Aider Polyglot 30
+83
View File
@@ -16,6 +16,89 @@ history; SemVer takes over from `v0.3.0` onward.
---
## v0.7.3 — 2026-05-15
### ✨ Features
- feat(report): surface kv-calc calibration verdict ([#143](https://github.com/noonghunna/club-3090/pull/143) by @noonghunna)
- feat(kv-calc): model v0.7.3 MoE architectures ([39e1873](https://github.com/noonghunna/club-3090/commit/39e18733aa8f14c83a1f76d5bed23156fda61568))
- feat(gemma-4-26b-a4b): AWQ + MTP n=4 — +12% narr / +49% code over no-MTP baseline ([6dc9a0d](https://github.com/noonghunna/club-3090/commit/6dc9a0dce1b4ecef787e50808a8a15439cabe209))
- feat(qwen-35b-a3b): preview-MTP compose + bench row — MTP measured SLOWER on MoE ([e1d44bd](https://github.com/noonghunna/club-3090/commit/e1d44bd732cef0757b8d3f0f872b5b0a1a4fe5cc))
- feat(vllm-pr41800): vendor truncate_prompt_tokens overlay across all pre-fix engines (closes #139) ([1d7aad1](https://github.com/noonghunna/club-3090/commit/1d7aad112c097709d954750ab0e725a785ad062a))
- feat(gemma-4-26b-a4b): AWQ path via vLLM PR #40886 overlay ([0053444](https://github.com/noonghunna/club-3090/commit/0053444e84f80ce9f3fb910435c8d627591f59ed))
- feat(estate): add parallel boot mode ([99328b4](https://github.com/noonghunna/club-3090/commit/99328b4cda7ec9efc723a001295525a6d1f86e2a))
- feat(moe): add dual-card composes for Gemma 26B-A4B + Qwen 35B-A3B preview ([2d1b1dc](https://github.com/noonghunna/club-3090/commit/2d1b1dc347c3897f8b902f7ab60c8bd9c97abbeb))
- feat(moe): wire Gemma 4 26B-A4B + Qwen 3.6 35B-A3B composes through fits() ([f7f6f44](https://github.com/noonghunna/club-3090/commit/f7f6f444b9e937396c3de70df7dd6a5b957be80b))
- feat(profiles): split engine-pin policy by Genesis dependency ([15eda8a](https://github.com/noonghunna/club-3090/commit/15eda8a823dd273ee9b238369cba365021b61a29))
- feat(profiles): add Gemma 4 26B-A4B ModelProfile + num_global_kv_heads field ([abf0e32](https://github.com/noonghunna/club-3090/commit/abf0e327f9b877f989fe16c88bda1a457c144e0a))
- feat(profiles): add Qwen 3.6 35B-A3B ModelProfile (MoE schema extensions) ([9378714](https://github.com/noonghunna/club-3090/commit/937871492ae08b58b041a1fded41ab793f0c8549))
### 🐛 Bug fixes
- fix(gpu-mode): mode_off tears down estate-managed instances ([9cd854d](https://github.com/noonghunna/club-3090/commit/9cd854dbcb9b8b215876e0f88443c269072b93b9))
- fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean ([3b2d940](https://github.com/noonghunna/club-3090/commit/3b2d940d26ec2ffe6daf631e77986898c1d2849d))
- fix(gemma-4-31b): route default/bf16 composes to vllm-nightly-clean ([cf0451a](https://github.com/noonghunna/club-3090/commit/cf0451a7ad5c56bd7b63327aac7178df70cba1fb))
- fix(engines): vllm-nightly-mtp anchors to 01d4d1ad (Sander v7.72.2 PROD pin) ([87f0a0c](https://github.com/noonghunna/club-3090/commit/87f0a0c528473a225f06997bdcb9b79fefa13b04))
### 📝 Documentation
- docs(soak-test): clarify PASS verdict semantics — closes #140 ([9a039d8](https://github.com/noonghunna/club-3090/commit/9a039d8c922d3141102e8244d5468e8de46c6674))
- docs(UPSTREAM): add PR #41800 truncate_prompt_tokens row ([273c017](https://github.com/noonghunna/club-3090/commit/273c017646087fe508626be7ecea5e2106be54be))
- docs(README): add v0.7.3 MoE models to Supported Models table ([e49c939](https://github.com/noonghunna/club-3090/commit/e49c9397481b5fba226323b59c4fa37bdc2aeeab))
- docs(BENCHMARKS): Gemma 4 26B-A4B AWQ first row + AutoRound row demoted ([92b69bd](https://github.com/noonghunna/club-3090/commit/92b69bd220ce8360d2dfd5fdcf89d26ee2264791))
- docs(BENCHMARKS): add v0.7.3 MoE preview section ([bdfb939](https://github.com/noonghunna/club-3090/commit/bdfb939edd98c3872439a7a05dcde13cce7ccaa2))
- docs(HARDWARE): add note on PCIe Gen 3 + older CPU TP=2 headwind ([8cf38b0](https://github.com/noonghunna/club-3090/commit/8cf38b05906e3954405af3db09a822ac87a25ae5))
- docs(KERNEL_MATRIX): add KV Cache Impact subsection ([1a233cd](https://github.com/noonghunna/club-3090/commit/1a233cd9fd05a6ec51adfadae79813bc80cea4b0))
- docs: add KERNEL_MATRIX.md (attention backend + engine support matrix) ([97195fe](https://github.com/noonghunna/club-3090/commit/97195fe443770bc8adc883755fd3baf9933df874))
- docs(kv-math): extend k_v_tensors=N notation to sliding-KV formulas ([b1c68b4](https://github.com/noonghunna/club-3090/commit/b1c68b4fe3211af3442cd3bb7e9fd011e30e5554))
- docs(kv-math): tighten k_v_tensors notation across all 4 formulas ([54d8bf0](https://github.com/noonghunna/club-3090/commit/54d8bf0b3e2b322dbd67e6939ff9f9995c642437))
- docs(kv-math): third-pass Grok polish ([67de3ec](https://github.com/noonghunna/club-3090/commit/67de3eca3eb4b75b43fd0f83d2ed304976b3f1b5))
- docs(kv-math): second-pass Grok polish ([78f94ca](https://github.com/noonghunna/club-3090/commit/78f94cacf006064f68d945122907d30be8723eec))
- docs(kv-math): address Grok review feedback ([3114399](https://github.com/noonghunna/club-3090/commit/31143999837639b5a7d492d2f2b64e164787bcfe))
- docs(kv-math): config-verify Qwen 35B-A3B + Gemma 26B-A4B MoE sections ([6ec6a67](https://github.com/noonghunna/club-3090/commit/6ec6a6761f36cca44ea819434be7205bcb6fdbc6))
### 🧹 Maintenance
- test(launch): align engine pin expectations ([127f4f6](https://github.com/noonghunna/club-3090/commit/127f4f6d8fe104a954b7865a4d7550017a1c629b))
[Pin: `git checkout v0.7.3`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.7.2...v0.7.3)
## v0.7.2 — 2026-05-15
### ✨ Features
- feat(launch): add hardware topology advisor ([d116ba9](https://github.com/noonghunna/club-3090/commit/d116ba9ba3442482e0005bab9067d304b40e5e24))
[Pin: `git checkout v0.7.2`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.7.1...v0.7.2)
## v0.7.1 — 2026-05-15
### ✨ Features
- feat(bench): surface prompt processing throughput ([2a148d7](https://github.com/noonghunna/club-3090/commit/2a148d702b9415129d4c4ec9d3e7d30765927aa4))
- feat(llamacpp): expose batch tuning knobs ([02249ab](https://github.com/noonghunna/club-3090/commit/02249ab1939f354ac062d343efefe32677203174))
### 🐛 Bug fixes
- fix(ci): simplify vllm image workflow, drop smoke-gate (#135) ([ce2617e](https://github.com/noonghunna/club-3090/commit/ce2617e0bc0f56d42caf64e96847d966380be80c))
### 📝 Documentation
- docs(upstream): PR #42102 closed-as-slop; local overlay permanent ([57eb269](https://github.com/noonghunna/club-3090/commit/57eb269cd70935fc3069b85e46ead8f0f0af13dc))
[Pin: `git checkout v0.7.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.7.0...v0.7.1)
## v0.7.0 — 2026-05-14
+2
View File
@@ -44,6 +44,8 @@ Each hardware page lists every supported model with the working composes for tha
|---|---|---|---|---|
| **[Qwen3.6-27B](models/qwen3.6-27b/)** | Production-ready ⭐ | 1× / 2× 3090 | vLLM ✅ · llama.cpp ✅ · SGLang ❌ blocked | Vision · tools · MTP n=3 · up to 262K ctx · vLLM dual = 89/127 TPS · llama.cpp single = full 262K, no prefill cliffs |
| **[Gemma 4 31B](models/gemma-4-31b/)** | Production-ready (dual-card only on Ampere 24 GB) | 2× 3090 only ¹ | vLLM ✅ · llama.cpp ❌ · SGLang ❌ | Vision · tools · MTP n=3 (Google official drafter) **OR** DFlash n=7 (z-lab drafter) · up to 262K ctx via INT8 PTH KV (PR [#40391](https://github.com/vllm-project/vllm/pull/40391) vendored) · MTP dual = 106/141 TPS at 32K, 95/126 at 262K · DFlash dual = 105/177 TPS at 32K (code-optimal) |
| **[Qwen3.6 35B-A3B](models/qwen3.6-35b-a3b/)** ⭐ NEW v0.7.3 | Preview (production-track blocked on Genesis v7.73.x) | 2× 3090 | vLLM ✅ (preview) · SGLang ❌ · llama.cpp ❌ | **MoE (256 experts × 8 active, ~3 B active params)** · vision · tools · upstream native loader via [vLLM PR #42521](https://github.com/vllm-project/vllm/pull/42521) · preview dual = **182/177 TPS at 16K** (no MTP, no TQ3, no Genesis) · ~2× the Qwen 3.6-27B dense baseline on the same hardware · production path (Genesis + TQ3 + MTP) pending upstream Genesis re-anchor |
| **[Gemma 4 26B-A4B](models/gemma-4-26b-a4b/)** ⭐ NEW v0.7.3 | Production via AWQ (Intel AutoRound INT4 blocked on Ampere) | 2× 3090 | vLLM ✅ (AWQ overlay) · SGLang ❌ · llama.cpp ❌ | **MoE (128 experts × 8 active, ~4 B active params)** · vision · tools · cyankiwi AWQ-4bit weights via vendored [vLLM PR #40886](https://github.com/vllm-project/vllm/pull/40886) (compressed-tensors MoE key remapping) · AWQ dual = **139/139 TPS at 32K**, CV 0.2% / 0.0% · Intel AutoRound variants Ampere-blocked (Marlin K-dim alignment — moe_intermediate_size=704 not aligned to group_size=128) — AutoRound works on SM90+ |
¹ Single-card boot OOMs on Ampere 24 GB regardless of KV format (weights + drafter + profiling at 8K ctx leaves no KV pool). Single-card Gemma 4 is feasible on 32 GB+ GPUs (validated on RTX 5090 32 GB by [@apnar](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042)).
+2
View File
@@ -45,6 +45,8 @@ Plus per-card model weights drop from ~14 GB to ~7 GB (sharded), KV cache from f
Cost paid for this: NCCL allreduce per layer between cards (~30-50µs/token on PCIe), ~10-20% TPS overhead vs single-card if single-card actually worked.
> **Reading soak-test results for TP=2 / llama.cpp configs.** A clean `verdict PASS` on `dual.yml` / `dual-turbo.yml` / `llamacpp/default` does NOT mean the Cliff 2 mitigation patches in the compose's overlay set (PN-* sidecars, FLA chunked-prefill stabilizers, etc.) are doing the work — the topology alone takes that failure mode off the table. PASS on TP=2 reflects "the configuration is stable end-to-end at this depth," not "patches X/Y are load-bearing here." For per-patch attribution, run the same soak with overlays stripped and compare. See `scripts/soak-test.sh --help` ("PASS VERDICT" block) and [#140](https://github.com/noonghunna/club-3090/issues/140).
## Why llama.cpp escapes Cliff 2b on a single card
llama.cpp uses **different kernels and a different memory allocator** than vLLM. Three concrete differences:
+19
View File
@@ -390,6 +390,25 @@ If you're on **SM89+ hardware (RTX 4090 / 5090, A6000 Ada / Blackwell)**, the pe
---
## Note for older host platforms (PCIe Gen 3 + older CPUs)
If your rig is on **PCIe Gen 3** (rather than Gen 4) **and/or paired with a pre-Zen3 / pre-2018 CPU** (e.g. Xeon Gold 61xx Skylake, Xeon E5 v4 Broadwell), TP=2 paths take a 30-40% throughput hit vs the Gen 4 / Ryzen 5950X / EPYC rigs in `BENCHMARKS.md`. Two compounding causes:
1. **PCIe Gen 3 x16 ≈ 15.75 GB/s** per direction vs Gen 4 x16 ≈ 31.5 GB/s. TP=2 all-reduce on the residual stream every layer is GB/s-class traffic — halving interconnect bandwidth roughly halves the all-reduce wall time, and decode-TPS is sensitive to that.
2. **Older Xeon / Broadwell CPUs** have lower per-core clock and IPC than current Ryzen / EPYC parts. Affects prefill throughput, TTFT, and host-side coordination between the two GPUs.
**Symptom**: GPU utilization asymmetry during decode (e.g. `GPU 0: 28% util / 174W` vs `GPU 1: 85% util / 254W`) — communication-starved TP=2, where one card finishes its half-step and stalls waiting on all-reduce.
**Mitigation on Gen 3 rigs**:
- **Enable persistence mode** (`sudo nvidia-smi -pm 1`) — common to find this off on KVM/VM hosts; with it disabled the driver tears down between idle periods and adds per-request init latency.
- **Prefer single-card paths**: with interconnect being the bottleneck, `vllm/minimal` (single-card fp8 KV, no MTP) or `vllm/long-text-no-mtp` (single-card TQ3 KV) often beats `dual.yml` on these rigs. You give up max context ceiling but get back the decode TPS the interconnect was eating.
- **More host RAM** if VM-passthrough: 32+ GB recommended; vLLM uses host RAM for tokenizer staging, paged weight loading, and IPC buffers — VMs with 15 GB total tend to thrash.
See [issue #137](https://github.com/noonghunna/club-3090/issues/137) for a worked example: Xeon Gold 6138 + PCIe Gen 3 x16 + 2× 3090 (KVM passthrough) → 32 / 41 TPS on `dual.yml`, vs Ryzen 5950X + Gen 4 + same KV config → 89 / 117 TPS ([@lolren disc #18](https://github.com/noonghunna/club-3090/discussions/18#discussioncomment-16820303)).
---
## Note for WSL2 / Windows users
### GPU memory budget on WSL2
+71
View File
@@ -0,0 +1,71 @@
# Kernel & Attention Backend Matrix (Mid-2026)
**Last updated:** 2026-05-15
**Focus:** Consumer / Prosumer GPUs (RTX 3090 → 5090) + hybrid/MoE models (Qwen3.6, Gemma4)
This matrix helps decide which inference engine + kernel combination to route composes to.
## Core Modern Kernels
| Kernel / Backend | Primary Purpose | Best Hardware | Key Strengths | Maturity on Ampere (3090) |
|-------------------------------|----------------------------------------|----------------------------|--------------------------------------------|---------------------------|
| **FlashAttention-2** | Standard attention | Ampere+ | Excellent balance, wide compatibility | Very High |
| **FlashAttention-3** | Hopper/Blackwell optimized | H100/B200+ | FP8, asynchrony, TMA/WGMMA | Medium (falls back) |
| **FlashInfer** | Flexible paged / custom attention | All NVIDIA | High performance, Triton-based, very tunable | Very High |
| **RadixAttention** | Prefix sharing (radix tree) | All | Best for chat, RAG, agents | High |
| **PagedAttention** (vLLM) | Memory-efficient KV management | All | Low fragmentation, high concurrency | Very High |
| **Triton Custom Kernels** | Rapid prototyping / specialized ops | All | Easy to write, high flexibility | High |
| **TensorRT-LLM Kernels** | Deep fusion + hardware-specific | NVIDIA (best on Ada+) | Highest raw speed on supported hardware | High |
| **MLA / FlashMLA** | Multi-head Latent Attention (DeepSeek/Qwen) | All | Optimized for compressed KV | High |
## KV Cache Impact
Different kernels affect KV cache size/efficiency differently. Some only speed up attention computation (no KV size change); others fundamentally change how the cache is laid out, paged, or shared across requests.
| Kernel / System | Impacts KV Cache Size? | Main Impact on KV Cache | Best For | Notes for 3090 / Consumer |
|-----------------------------------|------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------|---------------------------------------|------------------------------------------|
| **FlashAttention-2 / -3** | No | Reduces HBM traffic during attention computation (IO-aware tiling). Makes attention faster without materializing full attention matrix. | Speed (especially prefill) | FA2 is excellent on Ampere. FA3 falls back. |
| **FlashInfer** | No | Highly optimized paged + custom attention kernels. Excellent at handling non-contiguous KV blocks with minimal overhead. | Flexibility + speed on all NVIDIA | Often the fastest backend on 3090. |
| **PagedAttention** (vLLM) | Yes (effective) | Breaks KV cache into small pages/blocks. Dramatically reduces fragmentation → much higher real-world memory utilization. | High concurrency + long context | Core reason vLLM can serve many users. |
| **RadixAttention** (SGLang) | Yes (very strong) | Stores KV cache in a prefix tree (radix trie). Shares common prefixes across requests → massive memory savings in chat/RAG/agent workloads. | Prefix-heavy workloads | Biggest win for multi-turn / agents. |
| **TensorRT-LLM kernels** | No | Highly fused, hardware-specific kernels. Excellent with FP8/TQ3 but same base KV size. | Raw speed on new NVIDIA cards | Less flexible on 3090. |
| **Block Diffusion** (DFlash/Zaya) | Indirect but big | Generates multiple tokens per forward pass → fewer total KV updates per output token. KV cache grows slower in practice. | Throughput on bandwidth-limited cards| Very promising for 3090. |
## Engine Support Matrix
| Feature / Kernel | **vLLM** | **SGLang** | **TensorRT-LLM** | **llama.cpp** | Notes for 3090 / Consumer |
|-----------------------------------|-----------------------------------|-------------------------------------|------------------------------------|--------------------------------|---------------------------|
| **Default Attention** | FlashAttention / FlashInfer | FlashInfer (preferred) | Custom TRT kernels + FA3 | Basic FA2 / custom | FlashInfer best on 3090 |
| **RadixAttention / Prefix Cache**| APC (Automatic Prefix Caching) | **Native RadixAttention** (best) | Limited | Basic | SGLang wins for agents/RAG |
| **Paged KV Cache** | **Native PagedAttention** | Yes | Yes | Basic | vLLM strongest |
| **FlashAttention-3** | Good (Hopper+) | Good | **Best** | No | Falls back on Ampere |
| **DFlash (Block Diffusion)** | Good | **Excellent** (early + deep) | Emerging | Limited | SGLang currently strongest |
| **MTP / Speculative Decoding** | Strong (but TQ3 issues) | Strong | Excellent | Good | vLLM + MTP works well |
| **TQ3 / Advanced KV Quant** | Yes (but MTP broken) | Partial / WIP | Excellent | Good (GGUF) | vLLM best but buggy |
| **MoE / Hybrid Models** | Very Good | Excellent (esp. Qwen/DeepSeek) | Excellent | Decent | SGLang & TRT-LLM shine |
| **Structured Output** | Good | **Best-in-class** | Good | Basic | SGLang wins |
| **3090 / Ampere Optimization** | Solid | **Strong** | Good | Very Good | SGLang + FlashInfer often fastest |
## Recommendations for club-3090 / Composes
### Primary Routing Guidance
- **Interactive / Chat / Agents / RAG** → **SGLang** (RadixAttention + DFlash)
- **Max raw throughput on new cards (Ada/Blackwell)** → **TensorRT-LLM**
- **Broad compatibility + stability on 3090** → **vLLM** (with FlashInfer backend)
- **Low VRAM / single card / maximum compression** → **llama.cpp** (accept GGUF tradeoffs)
- **Hybrid MoE models (Qwen3.6 35B-A3B, Gemma4 26B)** → **SGLang** or **vLLM** with FlashInfer
### 3090-Specific Tips
- Use **FlashInfer** backend wherever possible (`--attention-backend flashinfer` in vLLM/SGLang).
- TQ3 + MTP is currently unstable in vLLM → prefer FP8/INT8 KV or DFlash.
- SGLang + RadixAttention + DFlash often gives the best real-world experience on consumer hardware.
---
**Cross-references**:
- See [DTYPE_MATRIX.md](DTYPE_MATRIX.md) for quantization + hardware accelerators
- See [KV_MATH.md](KV_MATH.md) for cache calculations per model
- See [INFERENCE_ENGINES.md](INFERENCE_ENGINES.md) for high-level engine comparison & setup
Contributions & updates welcome — this space moves fast!
+189 -72
View File
@@ -8,10 +8,10 @@ Four model families are documented:
|---|---|---|
| **Qwen 3.6 27B** (dense) | Calibrated 11/11 on this stack | Qwen3-Next hybrid: 16 full-attention + 48 GDN (Gated DeltaNet) layers |
| **Gemma 4 31B** (dense) | Calibrated 7/7 on this stack | Sliding-window + dense MLP: 50 SWA + 10 full-attention layers |
| **Qwen 3.6 35B-A3B** (MoE) | Math-ready, calibration pending | Qwen3-Next hybrid + MoE: 30 GDN + 10 gated-attention layers |
| **Gemma 4 26B-A4B** (MoE) | Math-ready, calibration pending | Sliding-window + dense MoE: SWA + occasional global layers |
| **Qwen 3.6 35B-A3B** (MoE) | **Config-verified, calibration pending** | Qwen3-Next hybrid + MoE: 30 GDN + 10 gated-attention layers. Confirmed from `config.json` 2026-05-15 — see [Qwen section](#qwen-36-35b-a3b-moe--per-card-budget-components). |
| **Gemma 4 26B-A4B** (MoE) | **Config-verified, calibration pending** | Sliding-window + dense MoE: 25 SWA + 5 full-attention layers. **Asymmetric KV heads** (8 sliding / 2 global). Confirmed from `config.json` 2026-05-15 — see [Gemma section](#gemma-4-26b-a4b-moe--per-card-budget-components). |
For models marked **calibration pending**, the architectural math derivations below are anchored to published config + model card material; absolute numbers ship as estimates until we measure them on the stack. See [Sources of Error & Accuracy](#sources-of-error--accuracy) at the end.
For models marked **config-verified, calibration pending**: architectural facts (layer counts, head dims, K=V tying, MoE expert counts, layer-type pattern) are sourced directly from the on-disk `config.json` and `layer_types` arrays — not estimates. What remains pending is the **empirical activation-peak coefficient** for each (model, KV-format) pair, which needs ≥4 measured BENCHMARKS rows per model. See [Sources of Error & Accuracy](#sources-of-error--accuracy) at the end.
## TL;DR
@@ -84,16 +84,70 @@ peak ≈ weights/TP ← exact, from checkpoint
+ drafter_overhead/TP ← speculative-decoding drafter weights, if any
```
### DeltaNet recurrent state (DeltaNet-family models only)
Hybrid DeltaNet models (Qwen3-Next family) maintain a fixed-size recurrent state between tokens, separate from the per-token growing KV. This state is tiny but worth noting for completeness:
```
delta_state_bytes ≈ num_gdn_layers
× (linear_num_k_heads × linear_k_head_dim
+ linear_num_v_heads × linear_v_head_dim
+ linear_conv_kernel_dim × (linear_num_k_heads × linear_k_head_dim
+ linear_num_v_heads × linear_v_head_dim))
× 4 (fp32, mamba_ssm_dtype)
× max_num_seqs
```
Three components per layer: K state + V state + conv1d kernel state. All four `linear_*` fields are in `config.json → text_config`. Concrete sizes for our models in per-model §"DeltaNet recurrent state" subsections — typically single-digit MB total, negligible vs activation peak.
**Worked example — Qwen 3.6 35B-A3B at `max_num_seqs=1`:**
```
linear_num_k_heads = 16, linear_k_head_dim = 128 → 2,048 elements per layer
linear_num_v_heads = 32, linear_v_head_dim = 128 → 4,096 elements per layer
linear_conv_kernel_dim = 4 → conv state = 4 × (2,048 + 4,096) = 24,576 elements
num_gdn_layers = 30, fp32 (4 bytes), max_num_seqs = 1
delta_state_bytes = 30 × (2,048 + 4,096 + 24,576) × 4 × 1
= 30 × 30,720 × 4
= 3,686,400 bytes
≈ 3.5 MB
```
Plug `max_num_seqs = 4` → ~14 MB. Both well below activation-peak scale.
Per-model sections below derive each term concretely.
## Quick reference: per-token growing-KV bytes
Headline numbers for the four shipping models, computed from each per-model formula in the deep sections. Useful for at-a-glance capacity planning.
| Model | bf16 (TP=1 / TP=2) | fp8_e5m2 (TP=1 / TP=2) | INT8 PTH or TQ3 (TP=1 / TP=2) | Vs Qwen 27B† |
|---|---:|---:|---:|---:|
| Qwen 3.6 27B | 65,536 B / 32,768 B | 32,768 B / 16,384 B | **13,927 B / 6,963 B** (TQ3) | **1.00×** (baseline) |
| Qwen 3.6 35B-A3B (MoE) | 20,480 B / 10,240 B | 10,240 B / 5,120 B | **4,352 B / 2,176 B** (TQ3) | **0.31×** (~3.2× lighter) |
| Gemma 4 31B | 163,840 B / 81,920 B | 81,920 B / 40,960 B | ~82,700 B / ~41,400 B (INT8 PTH) | **2.50×** (~2.5× heavier) |
| Gemma 4 26B-A4B (MoE) | **10,240 B / 5,120 B** | 5,120 B / 2,560 B | ~5,170 B / ~2,585 B (INT8 PTH) | **0.16×** (~6.4× lighter) |
† Ratio at fp8_e5m2 TP=2 — pick this as the comparison anchor because it's a common production config. Ratios shift slightly under other formats but the family hierarchy is stable.
**What jumps out:**
- **Gemma 4 26B-A4B vs 31B**: ~16× smaller per-token growing KV thanks to asymmetric KV head counts (2 global vs 16). At 200K context + fp8 + TP=2, growing KV per card is ~512 MB for the MoE vs ~8 GB for the 31B. Long-context serving on 24 GB Ampere is dramatically cheaper.
- **Qwen 3.6 35B-A3B vs 27B**: ~3.2× smaller per token (10 growing layers × 2 KV heads vs 16 × 4). The MoE shifts the bottleneck from KV to weights + activation.
- **Sliding-window KV** for Gemma models is **fixed** (not per-token): ~50 MB total (26B-A4B) / ~200 MB total (31B) at bf16. Excluded from per-token math but included in the per-model deep sections.
- **TQ3 (Genesis) only applies to Qwen-family** (DeltaNet kernel dependency); **INT8 PTH (PR #40391/#42102) is the long-context unlock for Gemma family on Ampere**.
## Model architecture summary
| Model | Total layers | Growing layers | Sliding / fixed | KV heads | Head dim | K=V tied | MoE | Special notes |
|---|---:|---:|---:|---:|---:|:---:|:---:|---|
| **Qwen 3.6 27B** | 64 | 16 (full-attention) | 48 (GDN recurrent) | 4 | 256 | No (×2) | No | DeltaNet block-wise activation peak (Cliff 2). `linear_attn` in-proj stays fp16 even under INT4 quant. |
| **Qwen 3.6 35B-A3B** | 40 | 10 (gated attention) | 30 (Gated DeltaNet) | 2 | 256 | No (×2) | Yes | Pattern: `10 × (3× GDN → MoE → 1× Gated Attn → MoE)`. Active params ~3B, total 35B. MoE experts gate decode FLOPs but not KV size. |
| **Qwen 3.6 35B-A3B** | 40 | **10** (gated attention at idx 3,7,11,15,19,23,27,31,35,39) | 30 (Gated DeltaNet) | **2** | 256 | No (×2) | **Yes (256×8)** | `full_attention_interval=4`: every 4th layer is attention. Built-in MTP (`mtp_num_hidden_layers=1`). `attn_output_gate=True` (gated attention). Vision-capable. Active params ~3B, total 35B. |
| **Gemma 4 31B** | 60 | 10 (full-attention) | 50 (SWA, window=1024) | 16 | 256 sliding / **512 global** | Yes (×1) | No | Global layers use 2× head_dim of sliding layers. K=V tying confirmed empirically against boot-log KV cache reports. |
| **Gemma 4 26B-A4B** | TBD (likely ~40-48) | TBD (likely sparse global) | Majority SWA (window=1024) | TBD | TBD | Likely yes (×1) | Yes | A4B = 4B active params from 26B total. Layer pattern from model card README. **Numbers pending config.json + first boot.** |
| **Gemma 4 26B-A4B** | **30** | **5** (full-attention at idx 5,11,17,23,29) | 25 (SWA, window=1024) | **8 sliding / 2 global** (asymmetric) | 256 sliding / **512 global** | **Yes (×1)** | **Yes (128×8)** | Asymmetric KV-head split per layer type. Every 6th layer is global, last layer always global. Per-token growing KV is **~16× smaller** than Gemma 4 31B (see [Gemma section](#gemma-4-26b-a4b-moe--per-card-budget-components)). Vision + audio support. **No Genesis required.** |
> **MoE column format**: `N×K` = `num_experts × num_experts_per_tok` (e.g. "256×8" = 256 experts, 8 active per token).
**Hybrid quirks to internalize:**
@@ -208,7 +262,7 @@ In the Qwen3-Next hybrid architecture, **only the 16 full_attention layers contr
Applying the general formula:
```
per_token_bytes = 16 (growing layers) × 4 (kv_heads) × 256 (head_dim) × 2 (no K=V tie) × bpe
per_token_bytes = 16 (growing layers) × 4 (kv_heads) × 256 (head_dim) × k_v_tensors=2 × bpe
= 32,768 × bpe bytes
```
@@ -268,36 +322,54 @@ overhead = 0.5 + 1.0 × mem_util + 0.3 × (TP - 1) # GB
This is rough — actual overhead depends on how many graphs vLLM captures, which depends on `max_num_seqs`, `compile_sizes`, and other internals.
### 5. DFlash draft model
### 5. DeltaNet recurrent state (per-stream, constant)
The 48 GDN layers maintain a fixed-size recurrent state between tokens (separate from the block-wise intermediate during forward — that's the activation peak in §3). Concrete size for Qwen 3.6 27B:
- K state: `16 × 128 × fp32 = 8 KB` per layer
- V state: `48 × 128 × fp32 = 24 KB` per layer
- Conv state: `4 × (16×128 + 48×128) × fp32 = ~128 KB` per layer
- **Total per layer: ~160 KB** × 48 layers × `max_num_seqs` streams
At `max_num_seqs=1`: ~7.5 MB total per card. At `max_num_seqs=4`: ~30 MB. Negligible vs activation peak (GB-scale) and KV pool (sub-GB). Listed for completeness; don't model in budget projections.
### 6. DFlash draft model
Only present on `dual-dflash*.yml` composes. `z-lab/Qwen3.6-27B-DFlash` is a ~1.75 GB draft model (per card, FP16). With TP > 1, the draft itself is sharded.
## Qwen 3.6 35B-A3B (MoE) — per-card budget components
**Status**: math-ready, **calibration pending** (not yet served on this stack). Numerical values below are derived from the architecture pattern in the model card; expect re-calibration once we measure boot peaks.
**Status**: **config-verified** (architecture confirmed from on-disk `config.json` 2026-05-15), **calibration pending** (not yet served on this stack — activation coefficients TBD). All architectural numbers below are sourced from the model checkpoint, not estimates.
### Architecture summary
Qwen 3.6 35B-A3B is a Qwen3-Next hybrid MoE:
Qwen 3.6 35B-A3B is a Qwen3-Next hybrid MoE (`model_type: qwen3_5_moe`, `architectures: Qwen3_5MoeForConditionalGeneration`):
- 40 transformer layers
- Pattern: `10 × (3× Gated DeltaNet → MoE → 1× Gated Attention → MoE)`
- **10 growing-attention layers** (gated attention with KV cache)
- **30 GDN layers** (recurrent state, fixed size)
- MoE: 128 experts, 8 active per token (typical Qwen3-Next MoE config — verify from `config.json`)
- **40 transformer layers**
- `full_attention_interval: 4` → every 4th layer is full attention; the other 3 are Gated DeltaNet
- `layer_types` array confirms **10 full_attention layers at indices [3, 7, 11, 15, 19, 23, 27, 31, 35, 39]** + **30 linear_attention (GDN) layers**
- **2 KV heads** (`num_key_value_heads: 2`) — caps `valid_tp` at `[1, 2]`
- **16 attention heads**, **head_dim: 256**
- **MoE: 256 experts, 8 active per token** (was estimated as 128 — real config has 2× more experts)
- `moe_intermediate_size: 512`, `shared_expert_intermediate_size: 512`
- Built-in MTP drafter (`mtp_num_hidden_layers: 1`) — same pattern as Qwen 3.6 27B
- `attn_output_gate: True` — gated attention
- Vision-capable (`vision_config` + image/video token IDs present)
- Active params: ~3B; total params: 35B
### 1. Model weights
MoE weights are dominated by the expert FFNs. Per-card budget under TP:
MoE weights are dominated by the expert FFNs. **5 quant variants on disk** as of 2026-05-15:
| Quant (planned) | On-disk estimate | Per-card at TP=2 |
|---|---:|---:|
| AutoRound INT4 (when available) | ~22-25 GB | 11-12 GB |
| AWQ-4bit (community) | ~22-25 GB | 11-12 GB |
| BF16 (unquantized) | ~70 GB | 35 GB (does not fit on 24 GB) |
| Quant | On-disk | Per-card at TP=2 | Notes |
|---|---:|---:|---|
| AutoRound INT4 (`qwen3.6-35b-a3b-autoround-int4`) | 20 GB | 10 GB | Production; matches our Qwen 3.6 27B AutoRound pipeline |
| GPTQ INT4 (`qwen3.6-35b-a3b-gptq-int4`) | 22 GB | 11 GB | Experimental |
| GGUF (`qwen3.6-35b-a3b-gguf`) | 90 GB | n/a (llama.cpp single-card path) | Multi-bit-depth |
| DFlash variants (`*-dflash`, `*-dflash-gguf`) | variable | n/a | Experimental (z-lab) |
| BF16 unquantized | ~70 GB | 35 GB | Does not fit on 24 GB |
Like the dense Qwen 3.6 27B, DeltaNet `linear_attn` in-projection layers will likely stay at fp16 even under INT4 quantization. The byte count will be included in the total checkpoint size.
Like the dense Qwen 3.6 27B, DeltaNet `linear_attn` in-projection layers stay at fp16 even under INT4 quantization. The byte count is included in the total checkpoint size.
**Note**: MoE expert weights all live in VRAM (they're sparse-activated at FLOPs level, not at memory level). Don't confuse "active params" with "loaded params" — the budget is for the full 35B.
@@ -306,7 +378,7 @@ Like the dense Qwen 3.6 27B, DeltaNet `linear_attn` in-projection layers will li
Applying the general formula:
```
per_token_bytes = 10 (growing layers) × 2 (kv_heads) × 256 (head_dim) × 2 (no K=V tie) × bpe
per_token_bytes = 10 (growing layers) × 2 (kv_heads) × 256 (head_dim) × k_v_tensors=2 × bpe
= 10,240 × bpe bytes
```
@@ -338,11 +410,22 @@ The activation peak should be **~60-70% of dense Qwen 3.6 27B's** (30/48 layers
MoE introduces a few new accounting items:
- **Router workspace**: small (`hidden_size × num_experts` weights, ~100-200 MB). One-time cost, not per-token.
- **Expert dispatch buffers**: vLLM allocates buffers for top-k expert routing. Empirical ~200-400 MB per card.
- **Router workspace**: `hidden_size × num_experts × bf16_bytes = 2048 × 256 × 2 = ~1 MB` per router. Across 40 layers ≈ 40 MB. Tiny one-time cost.
- **Expert dispatch buffers**: vLLM allocates buffers for top-k expert routing across all 256 experts. Empirical ~200-400 MB per card.
- **No KV-side impact**: MoE only gates FFN compute. The KV cache for the gated-attention layers is unaffected.
### 5. Cudagraph + workspace overhead
### 5. DeltaNet recurrent state (per-stream, constant)
The 30 GDN layers maintain a fixed-size recurrent state between tokens (separate from the block-wise intermediate during forward, which is the activation peak). Concrete size:
- K state: `linear_num_k_heads × linear_k_head_dim × fp32 = 16 × 128 × 4 = 8 KB` per layer
- V state: `linear_num_v_heads × linear_v_head_dim × fp32 = 32 × 128 × 4 = 16 KB` per layer
- Conv state: `linear_conv_kernel_dim × (16×128 + 32×128) × fp32 = ~96 KB` per layer
- **Total per layer: ~120 KB** × 30 layers × `max_num_seqs` streams
At `max_num_seqs=1`: ~3.5 MB total per card. At `max_num_seqs=4`: ~14 MB. **Negligible** vs activation peak (which is GB-scale) and KV pool (sub-GB). Listed here for completeness; don't bother modelling in budget projections.
### 6. Cudagraph + workspace overhead
Same form as dense models:
@@ -394,7 +477,7 @@ Two shipped quants on this stack: AutoRound INT4 (default) and AWQ-4bit (Tier 2
Each stores K and V at `global_head_dim=512`, with K==V tying meaning a single store per element:
```
per_token_bytes_growing = 10 (growing layers) × 16 (kv_heads) × 512 (global_head_dim) × 1 (K=V tied) × bpe
per_token_bytes_growing = 10 (growing layers) × 16 (kv_heads) × 512 (global_head_dim) × k_v_tensors=1 × bpe
= 81,920 × bpe bytes
```
@@ -419,7 +502,7 @@ Total growing-KV pool per card = `per_token_bytes_growing / TP × max_ctx × max
The 50 sliding-attention layers maintain a fixed-size KV window (`sliding_window=1024`). K==V tying applies here too:
```
sliding_kv_bytes_total = 50 (sliding layers) × 16 (kv_heads) × 256 (head_dim) × 1 (K=V tied) × bpe × 1024 (window)
sliding_kv_bytes_total = 50 (sliding layers) × 16 (kv_heads) × 256 (head_dim) × k_v_tensors=1 × bpe × 1024 (window)
= 209,715,200 × bpe bytes
≈ 200 MB × bpe
```
@@ -463,84 +546,118 @@ At TP > 1, drafter weights shard across cards (`drafter_gb / TP`).
### Architecture summary
Gemma 4 26B-A4B is a Gemma 4 MoE with sliding-window attention:
Gemma 4 26B-A4B is a Gemma 4 MoE (`model_type: gemma4`, `architectures: Gemma4ForConditionalGeneration`):
- Likely ~40-48 transformer layers (smaller than Gemma 4 31B's 60)
- Sliding-window + occasional global layers (Gemma 4 family pattern); final layer typically global
- Sliding window 1024 (same as 31B)
- K=V tying expected (Gemma 4 family convention)
- MoE: number of experts + active-per-token TBD from `config.json`
- **30 transformer layers** (notably smaller than Gemma 4 31B's 60)
- `layer_types` array confirms **5 full_attention layers at indices [5, 11, 17, 23, 29]** + **25 sliding_attention layers**
- Pattern: every 6th layer is global; **last layer is always global** (per Gemma 4 family convention)
- `sliding_window: 1024`
- **`attention_k_eq_v: True`** — K and V share storage (×1)
- **Asymmetric KV head counts** (the big architectural surprise vs Gemma 4 31B):
- `num_key_value_heads: 8` — for sliding-attention layers
- `num_global_key_value_heads: 2` — for full-attention layers
- `head_dim: 256` (sliding), `global_head_dim: 512` (global)
- **MoE: 128 experts, 8 active per token** (`top_k_experts: 8` in config)
- `moe_intermediate_size: 704`
- Multimodal: `vision_config` + `audio_config` token IDs + image/video token IDs present
- **Does NOT require Genesis** (Gemma 4 family has no DeltaNet quirks)
- Active params: ~4B; total params: 26B
**Exact layer counts pending the model's actual config.json on disk.** The placeholders below show the math shape — substitute real values once measured.
### 1. Model weights
| Quant (planned) | On-disk estimate | Per-card at TP=2 |
|---|---:|---:|
| AutoRound INT4 (when available) | ~16-18 GB | 8-9 GB |
| BF16 (unquantized) | ~52 GB | 26 GB (does not fit on 24 GB) |
| Quant | On-disk | Per-card at TP=2 | Notes |
|---|---:|---:|---|
| **Intel AutoRound INT4 mixed** (`gemma-4-26b-a4b-autoround-int4-mixed`) | ~14-15 GB | 7-8 GB | Production target. Mixed precision protects routing-critical layers; matches our AutoRound pipeline. |
| Intel AutoRound INT4 (pure) | ~13 GB | 6.5 GB | Alternative; slightly worse routing quality than mixed. |
| Community AWQ-4bit (cyankiwi) | ~13-14 GB | 6.5-7 GB | Different quant pipeline → activation coefficients don't transfer from our AutoRound calibration. |
| BF16 (unquantized) | ~52 GB | 26 GB | Does not fit on 24 GB. |
MoE expert weights all live in VRAM (sparse-activation at FLOPs, not at memory). Active-params count (4B) doesn't reduce the loaded budget.
MoE expert weights all live in VRAM (sparse-activation at FLOPs level, not at memory). Active-params count (4B) doesn't reduce the loaded budget.
### 2. KV pool — growing portion (global layers only)
### 2. KV pool — growing portion (5 full_attention layers)
Estimated `N_global` global layers (TBD from README; Gemma 4 family typically uses 1 global per 5 sliding):
The asymmetric KV head count dramatically reduces per-token growing KV vs Gemma 4 31B:
```
per_token_bytes_growing = N_global × num_kv_heads × global_head_dim × 1 (K=V tied) × bpe
per_token_bytes_growing = num_full_attn_layers × num_global_kv_heads × global_head_dim × k_v_tensors=1 × bpe
= 5 × 2 × 512 × 1 × bpe
= 5,120 × bpe bytes
```
If `N_global = 8` and shape matches 31B family (`num_kv_heads=16`, `global_head_dim=512`):
**Compare to Gemma 4 31B's growing KV** = `10 × 16 × 512 × 1 × bpe = 81,920 × bpe bytes` per token. The 26B-A4B is **~16× lighter per token**:
- Fewer full-attention layers: 5 vs 10
- Fewer KV heads on global layers: 2 vs 16
- Same head_dim and K=V tying
| KV format | bpe | per-token growing KV (TP=1) | per-token (TP=2) |
|---|---:|---:|---:|
| `bf16` / `fp16` | 2.0 | 10,240 B (~10 KB) | 5,120 B |
| `fp8_e5m2` / `fp8_e4m3` | 1.0 | 5,120 B (~5 KB) | 2,560 B |
| `int8_per_token_head` | ~1.01 | ~5,170 B | ~2,585 B |
| `q4_0` | ~0.56 | ~2,867 B | ~1,434 B |
**Implication**: at 200K context, growing KV pool per card at TP=2 + fp8 = `2,560 × 200,000 = ~512 MB`. **The 26B-A4B is extremely KV-light** — even at full 262K context, growing KV per card is under 700 MB at fp8. The constraint shifts decisively to weights + activation peak, NOT to KV.
This means BF16 KV becomes viable at 262K on Ampere consumer cards (~1.3 GB growing KV per card) — a contrast to Gemma 4 31B where INT8 PTH was the unlock for long context.
### 3. KV pool — fixed sliding portion (25 sliding_attention layers)
The 25 SWA layers maintain a fixed-size KV window (`sliding_window: 1024`):
```
per_token_bytes_growing = 8 × 16 × 512 × 1 × bpe = 65,536 × bpe bytes
sliding_kv_bytes_total = num_sliding_layers × num_kv_heads × head_dim × k_v_tensors=1 × bpe × sliding_window
= 25 × 8 × 256 × 1 × bpe × 1024
= 52,428,800 × bpe bytes
≈ 50 MB × bpe
```
That's ~80% of Gemma 4 31B's growing KV per token. INT8 KV remains the right format for 24 GB Ampere at long contexts.
**Constant** — doesn't scale with `max_ctx` or `max_num_seqs`. At fp8 KV: ~50 MB per card (TP=1) or ~25 MB at TP=2. Negligible.
### 3. KV pool — fixed sliding portion
Note: this is dramatically smaller than Gemma 4 31B's sliding portion (`50 × 16 × 256 × 1 × bpe × 1024 ≈ 200 MB × bpe`) due to fewer sliding layers (25 vs 50) and fewer KV heads (8 vs 16).
```
sliding_kv_bytes_total = N_sliding × num_kv_heads × head_dim × 1 (K=V tied) × bpe × 1024
```
### 4. Activation peak (SWA prefill + dense MoE intermediate buffer)
If `N_sliding = 32` and shape matches 31B (`head_dim=256`):
Same mechanism as Gemma 4 31B (SWA prefill + dense MoE intermediate buffer). MoE adds small per-expert routing overhead but **shouldn't dominate**.
```
sliding_kv_bytes_total = 32 × 16 × 256 × 1 × bpe × 1024 = 134,217,728 × bpe bytes ≈ 128 MB × bpe
```
Projected coefficient (calibration pending; expect ≥4 BENCHMARKS rows before locking in):
Constant; doesn't scale with `max_ctx`.
| KV format | Projected bytes/layer/token | Reasoning |
|---|---:|---|
| `bf16` / `fp16` | ~1.0-1.5 KB | Smaller than Gemma 4 31B due to fewer total layers (30 vs 60) and smaller `hidden_size` (2816 vs 5376) |
| `fp8_e5m2` / `int8_per_token_head` | ~1.0-1.5 KB | Similar to BF16; minimal dequant overhead |
### 4. Activation peak
Expected activation peak: ~1-2 GB at TP=2 dual-card configs, but **calibration TBD**.
Same mechanism as Gemma 4 31B (SWA prefill + dense MoE intermediate buffer). MoE may add a small per-expert routing overhead but **shouldn't dominate**. Empirical coefficient TBD; expected ~1-2 GB at TP=2.
### 5. MoE-specific considerations
### 5. MoE considerations
Same accounting as Qwen 3.6 35B-A3B:
Same as Qwen 3.6 35B-A3B MoE:
- Router workspace (~100-200 MB, one-time)
- Expert dispatch buffers (~200-400 MB per card)
- No KV-side impact from MoE
- **Router workspace**: `hidden_size × num_experts` = `2816 × 128` ≈ 360 K weights. Tiny (~700 KB at BF16). One-time cost.
- **Expert dispatch buffers**: vLLM allocates buffers for top-k expert routing. Empirical ~200-400 MB per card.
- **No KV-side impact**: MoE only gates FFN compute. KV cache for full-attention layers is unaffected.
### 6. Cudagraph + workspace overhead + drafter
Same empirical form as 31B; drafter family TBD (Google may release a Gemma 4 26B MTP assistant similar to the 31B-it-assistant).
Same empirical form as Gemma 4 31B; standard `0.5 + 1.0 × mem_util + 0.3 × (TP - 1) GB`.
**Drafter family**:
- `google/gemma-4-26B-A4B-it-assistant` released as MTP drafter (~0.5-1 GB, FP16). Same pattern as our existing `gemma-4-31b-it-assistant` drafter.
- `z-lab/gemma-4-26B-A4B-it-DFlash` released as DFlash drafter (community).
### Estimated per-card budget at TP=2, 24 GB VRAM
| Term | Value (INT8 KV, 100K ctx, seqs=1) | Notes |
| Term | Value (fp8 KV, 200K ctx, seqs=1) | Notes |
|---|---:|---|
| Weights / 2 | ~8-9 GB | INT4 quant |
| KV pool growing | ~3-4 GB | 100K × 32 KB/tok = ~3.2 GB |
| KV pool sliding | ~0.13 GB | Constant |
| Activation peak | ~1-2 GB | Smaller than 31B if fewer total layers |
| Cudagraph + overhead | ~1.2 GB | |
| **Predicted peak** | **~13-17 GB** | Comfortable headroom on 24 GB, fits 20 GB at lower ctx |
| Weights / 2 | ~7-8 GB | AutoRound INT4 mixed (~14-15 GB on-disk) |
| KV pool growing | ~0.5 GB | Asymmetric KV heads + few global layers |
| KV pool sliding | ~0.05 GB | Constant; trivially small |
| Activation peak | ~1-2 GB | Smaller than Gemma 4 31B |
| Cudagraph + overhead | ~1.2 GB | Empirical fit |
| MoE expert dispatch buffers | ~0.3 GB | Per-card |
| **Predicted peak** | **~10-12 GB** | Massive headroom on 24 GB; could likely run at higher mem_util or push to BF16 KV at full 262K |
**Calibration pending**. Numbers will shift once real `config.json` values replace estimates.
**Calibration pending**. The headline finding to verify on first boot: Gemma 4 26B-A4B at full 262K context should fit on a single 3090 with INT4 weights — single-card serving may be the right default for this model.
## Best practices for building a KV calculator
+51
View File
@@ -42,6 +42,57 @@ isn't your topology.
---
## Topology classification
The launcher classifies your selected hardware and emits strategy guidance
when the cards are not matched. You can run the classifier without booting
anything:
```bash
bash scripts/launch.sh --topology
```
Use `--gpus 0,1` or `--cards 2` with `--topology` if you only want advice
for a subset.
| Class | What it means | Example | Recommended |
|---|---|---|---|
| `single_card` | 1 GPU detected | 1x RTX 3090 | Use the largest single-card compose that fits (`vllm/default`, `vllm/long-text`, `llamacpp/default`). |
| `homogeneous` | All cards have matched VRAM and matched SM | 2x RTX 3090 | TP=N is the optimal default; use the shipped `vllm/dual*` or `vllm/dual4*` composes. |
| `vram_matched_compute_mismatched` | Same VRAM, different compute tier | RTX 3090 + RTX 4090 | TP=N works correctly, but faster cards wait at NCCL allreduce. Estate planner is better for multi-model workloads. |
| `vram_mismatched` | Different VRAM sizes | RTX 3060 12 GB + RTX 3090 24 GB | Prefer llama.cpp `--tensor-split`, manual PP=N experiments, or estate planner. Avoid TP=N across the full mismatched set. |
| `heterogeneous_mixed` | Multiple VRAM and compute tiers | RTX 3060 + RTX 3090 + RTX 4090 | Manual selection. Run one model on the largest matched subset or use estate planner for separate endpoints. |
### Why TP=N is poor on VRAM-mismatched cards
Tensor parallelism splits weights evenly across cards. If one card has 24 GB
and another has 12 GB, TP=2 still puts roughly half the model on each card.
The smaller card becomes the hard ceiling for weights, KV cache, activations,
and fragmentation. For Qwen 3.6 27B INT4, that usually leaves too little KV
headroom to be useful.
For mismatched VRAM, the practical paths are:
- llama.cpp `--tensor-split` for weighted layer placement.
- PP=N as a manual vLLM flag flip (`--pipeline-parallel-size N`) when you are
deliberately experimenting. club-3090 does not ship a PP compose today.
- Estate planner: `bash scripts/launch.sh --estate` runs different models on
different card subsets without forcing one model across uneven VRAM.
### When compute-mismatched TP is fine
Matched VRAM with different SM, such as RTX 3090 + RTX 4090, is a different
trade-off. TP=2 works because both cards have enough memory for the same model
shard and KV budget. The cost is throughput: the faster card waits at NCCL
allreduce barriers, so effective pair speed caps near the slower card. You
preserve per-card VRAM capacity, but waste some compute on the faster card.
That is acceptable for one-model serving. If your goal is maximum aggregate
throughput from two different cards, estate planner usually wins because each
card runs its own model at full speed.
---
## Valid TP values for Qwen3.6-27B
vLLM's tensor parallelism splits attention heads across cards. The TP
+1
View File
@@ -83,6 +83,7 @@ recommended pinned target; `:latest` follows the most-recent dated nightly.
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [#35936](https://github.com/vllm-project/vllm/pull/35936) — `tool_choice="required"` falls back to configured tool parser | 🟡 Open / **local overlay active** | Qwen3-Coder with `--tool-call-parser qwen3_coder` emits XML-style tool calls. On pinned nightly `1acd67a79`, non-streaming `tool_choice="required"` validates JSON only, bypasses the configured parser, and returns `tool_calls=[]`. MLS-Bench hits this when `thinking.enabled=false`. | Vendored overlay: [`models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md). Drop when #35936 or equivalent lands in our pinned image. |
| **[#41800](https://github.com/vllm-project/vllm/pull/41800)** — `truncate_prompt_tokens` kwarg on `get_max_tokens()` | ✅ **Merged upstream 2026-05-06 at `d5b31c95`** / **local overlay active on pre-fix engine pins** | opencode (and other agentic clients sending `truncate_prompt_tokens`) fail with HTTP 400 `get_max_tokens() got an unexpected keyword argument` on engines pinned to `01d4d1ad` (Genesis MTP), `e47c98ef` (DFlash, full). All three SHAs predate `d5b31c95`. `vllm-nightly-clean` (`bf610c2f`, post-fix) doesn't need the overlay. Tracking issue: club-3090 #139. Triggered by club-3090 #138 (SEVENID's opencode failure). | Vendored overlay: [`models/qwen3.6-27b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/README.md). Wired into 18 affected composes (every compose routing through `vllm-nightly-(mtp\|dflash\|full)`). Install script has upstream-fix detection: no-ops cleanly when run against a post-`d5b31c95` nightly. **Drop trigger per engine**: bump each affected engine's pin past `d5b31c95`. For `vllm-nightly-mtp` that requires Genesis v7.73.x; for `vllm-nightly-dflash` and `vllm-nightly-full` it requires re-validating PR #41703 and PR #42102 overlays on a newer base. |
| [#40361](https://github.com/vllm-project/vllm/pull/40361) — Marlin pad-sub-tile-n | 🟡 Open, mergeable, **stale 13d** (last update 2026-04-20) | All 4 dual-card composes + `dual-nvlink.yml` + `dual-nvlink-turbo.yml` mount the patched files vendored in-repo at `models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/`. Drops out as a setup dependency when this merges + propagates. | Vendored mount: see [`models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md). Queued for rebase + ping next week (see "Active follow-ups" table above). |
| [#40807](https://github.com/vllm-project/vllm/issues/40807) — `.tolist()` cudagraph crash on continuation-prefill | ✅ **Retired locally** (2026-05-05 Genesis v7.72.2 bump) — Genesis ships [P78 `TOLIST_CAPTURE_GUARD`](../models/qwen3.6-27b/vllm/compose/dual/tq3-mtp-genesis.yml) as the equivalent fix. Currently disabled (`=0`) on `tq3-mtp-genesis.yml` after rebench-full leg 6 (2026-05-11) passed clean with it off — apparent root cause is now covered by Genesis PN34 (workspace-lock relax) + post-#41434 attention rework. Non-Genesis composes on `1acd67a79` pin run without any guard for this bug; unvalidated at long-context TurboQuant chunked-prefill (worth testing per cferra's vllm#41403 validation pass — see vllm#40798 row below). | None active. Drop the Genesis env var permanently if a future v7.73.x rebench leaves it OFF without regression. |
| [#40798](https://github.com/vllm-project/vllm/pull/40798) + [#42215](https://github.com/vllm-project/vllm/pull/42215) — share decode scratch workspace pre-CUDA-graph + decode-kernel warmup | 🟡 Open, validated cross-rig | Pair closes the `AssertionError: Workspace is locked but allocation requires NMB` crash that fires at `turboquant_attn.py:_continuation_prefill` for ≥48K-token chunked-prefill with TurboQuant KV. Independently validated on 2× 3090 sm_86 by cferra (vllm#41403 [comment](https://github.com/vllm-project/vllm/issues/41403#issuecomment-4435164709), 2026-05-12). | Genesis [PN34 `WORKSPACE_LOCK_RELAX`](../models/qwen3.6-27b/vllm/compose/dual/tq3-mtp-genesis.yml) addresses the same symptom via a different mechanism (relax-lock vs reserve-before-capture). On non-Genesis composes (`1acd67a79` pin) we currently have no guard — re-validate against this PR pair once they propagate to a nightly we pin to, then A/B PN34 vs upstream. |
@@ -0,0 +1,110 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 26B-A4B MoE (cyankiwi AWQ-4bit, compressed-tensors)
# Topology: Dual 3090 (TP=2)
# Drafter: MTP n=4 — google/gemma-4-26B-A4B-it-assistant (external)
# KV: bfloat16 (sidesteps Ampere fp8 dispatch issues)
# Vision: off (limit-mm-per-prompt image=0 audio=0)
# Max ctx: 32K (matches awq.yml for direct A/B)
# Genesis: none (Gemma 4 family doesn't need Genesis)
# Status: 🔵 v0.7.3 PRIMARY — AWQ + MTP, ships alongside awq.yml
# ---------------------------------------------------------------------------
# Same shape as awq.yml plus the external MTP assistant drafter wired via
# `--speculative-config`. n=4 per the gemma-26b-it-assistant DrafterProfile.
#
# Direct A/B target: awq.yml gives 138.88 / 138.67 wall TPS (2026-05-15, no
# spec-decode). This compose measures MTP's contribution. Per the Qwen
# 35B-A3B preview-MTP finding (slower despite high acceptance), it's worth
# being skeptical — measure both before declaring a default.
#
# Why a separate drafter path (not built-in MTP head):
# Gemma 4 26B-A4B doesn't ship a built-in MTP head — `mtp_num_hidden_layers`
# is null in the model YAML. Google released an external assistant model
# (`google/gemma-4-26B-A4B-it-assistant`, ~0.97 GB BF16) for this purpose.
# Same family pattern as the 31B assistant we already use for `gemma-mtp`.
#
# PR #40886 overlay still applied (compressed-tensors AWQ MoE key remap).
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 7.5+
services:
vllm-gemma-4-26b-a4b-awq-mtp-tp2:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-gemma-4-26b-a4b-awq-mtp-tp2}"
restart: "no"
ports:
- "${ESTATE_PORT:-${PORT:-8043}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# vLLM PR #40886 overlay — AWQ compressed-tensors MoE key remapping.
# See ../../patches/vllm-pr40886-awq-moe-keys/README.md.
- ../../patches/vllm-pr40886-awq-moe-keys/install.sh:/etc/club3090/install-pr40886.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- OMP_NUM_THREADS=1
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:512
- VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
- TRITON_CACHE_DIR=/root/.triton/cache
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
bash /etc/club3090/install-pr40886.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/gemma-4-26b-a4b-awq-4bit
- --served-model-name
- gemma-4-26b-a4b-awq
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-32768}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "256"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- auto
- --trust-remote-code
- --enable-auto-tool-choice
- --tool-call-parser
- gemma4
- --chat-template
- /vllm-workspace/examples/tool_chat_template_gemma4.jinja
- --enable-prefix-caching
- --enable-chunked-prefill
# External MTP assistant drafter — n=4 per gemma-26b-it-assistant DrafterProfile.
- --speculative-config
- '{"model":"/root/.cache/huggingface/gemma-4-26b-a4b-it-assistant","num_speculative_tokens":4}'
@@ -0,0 +1,117 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 26B-A4B MoE (cyankiwi AWQ-4bit, compressed-tensors)
# Topology: Dual 3090 (TP=2)
# Drafter: none (base smoke compose — MTP via gemma-26b-it-assistant
# can be added in a follow-up compose)
# KV: bfloat16 (sidesteps Ampere fp8 dispatch issues — Triton
# fp8e4nv not supported on sm_86)
# Vision: off (limit-mm-per-prompt image=0 audio=0) for first boot
# Max ctx: 32K (first-boot smoke; can extend after validation)
# Genesis: none (Gemma 4 family doesn't need Genesis)
# Status: 🔵 v0.7.3 PRIMARY — AWQ path with PR #40886 overlay
# Active params: ~4B (128 experts × 8 active = ~4B routed)
# ---------------------------------------------------------------------------
# Gemma 4 26B-A4B-it (cyankiwi AWQ-4bit, compressed-tensors format) — the
# v0.7.3 production Ampere path for this model, replacing the Intel
# AutoRound INT4 attempt which is structurally blocked on SM86 by Marlin
# K-dim alignment (moe_intermediate_size=704 not aligned to group_size=128).
#
# Why this works on Ampere where AutoRound INT4 doesn't:
# - AWQ via compressed-tensors routes through a different vLLM kernel
# path that handles arbitrary K shapes (unlike Marlin)
# - PR #40886 overlay (mounted below) fixes the KeyError on the
# `_packed` / `_scale` suffixed MoE expert weights — without this,
# vLLM's `gemma4.py::_weight_iterator` doesn't know how to remap
# the compressed-tensors per-expert keys into the FusedMoE format
#
# Overlay drops when: vLLM PR #40886 merges upstream AND vllm-nightly-clean
# bumps past the merge commit. Track in docs/UPSTREAM.md.
#
# Vendor: cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit (~17 GB on disk, 4 shards).
# Format: compressed-tensors pack-quantized, group-quantized scales.
# PR #40886 head: tajwali/vllm @ 652819dad0bf9bbb0436d6660822e7aff30c3ff0.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 7.5+
services:
vllm-gemma-4-26b-a4b-awq-tp2:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-gemma-4-26b-a4b-awq-tp2}"
restart: "no"
ports:
- "${ESTATE_PORT:-${PORT:-8042}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# vLLM PR #40886 overlay — AWQ compressed-tensors MoE key remapping.
# See ../../patches/vllm-pr40886-awq-moe-keys/README.md.
- ../../patches/vllm-pr40886-awq-moe-keys/install.sh:/etc/club3090/install-pr40886.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- OMP_NUM_THREADS=1
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:512
- VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
- TRITON_CACHE_DIR=/root/.triton/cache
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
# PR #40886 overlay must run BEFORE `vllm serve` imports the model
# module — installs the AWQ compressed-tensors key remapping into
# gemma4.py via an anchor-based Python patcher (idempotent).
bash /etc/club3090/install-pr40886.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/gemma-4-26b-a4b-awq-4bit
- --served-model-name
- gemma-4-26b-a4b-awq
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-32768}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "256"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- auto
- --trust-remote-code
- --enable-auto-tool-choice
- --tool-call-parser
- gemma4
- --chat-template
- /vllm-workspace/examples/tool_chat_template_gemma4.jinja
- --enable-prefix-caching
- --enable-chunked-prefill
@@ -0,0 +1,110 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 26B-A4B MoE (Intel AutoRound INT4 mixed)
# Topology: Dual 3090 (TP=2)
# Drafter: none (base smoke compose — MTP via gemma-26b-it-assistant
# can be added once base boot is validated)
# KV: bfloat16 (sidesteps Ampere fp8 dispatch issues — same logic
# as gemma-4-31b/dual composes)
# Vision: off (limit-mm-per-prompt image=0 audio=0) for first boot
# Max ctx: 32K (first-boot smoke; can extend after validation)
# Genesis: none (Gemma 4 family doesn't need Genesis — engine
# vllm-nightly-clean rides latest nightly)
# Status: 🔵 v0.7.3 ONBOARDING — primary bench target
# Active params: ~4B (128 experts × 8 active = ~4B routed)
# ---------------------------------------------------------------------------
# Gemma 4 26B-A4B-it (Intel AutoRound INT4 mixed) — dual-card production
# bench target for v0.7.3 MoE onboarding alongside Qwen 3.6 35B-A3B.
#
# Why TP=2: weights ~16 GB → ~8 GB/card, leaving ~14 GB/card for KV pool
# and activations. Mirrors the production posture used across the matrix.
# Single-card variant exists at single/docker-compose.yml (24 GB-only).
#
# Why vllm-nightly-clean (not vllm-nightly-mtp):
# Gemma 4 family doesn't need Genesis patches (no DeltaNet quirks).
# By routing through the unconstrained nightly we get latest upstream
# features (incl. MoE-loader improvements) without waiting on Genesis
# re-anchor cycles. Genesis-anchored MTP path is still vllm-nightly-mtp,
# pinned to nightly-01d4d1ad (Sander's v7.72.2 PROD pin).
#
# Why no drafter on this compose:
# The 26B-A4B has an external MTP assistant (google/gemma-4-26B-A4B-it-
# assistant, ~0.97 GB) — separate weights from the 31B assistant.
# Validating base boot first, then layering MTP on top in a follow-up
# compose.
#
# Models:
# target: Intel/gemma-4-26B-A4B-it-int4-mixed-AutoRound
# (~16 GB on disk — quant mix of MoE expert layers)
# draft : none on this compose (gemma-26b-it-assistant compose TBD)
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 7.5+
services:
vllm-gemma-4-26b-a4b-tp2:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-gemma-4-26b-a4b-tp2}"
restart: "no"
ports:
- "${ESTATE_PORT:-${PORT:-8041}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- OMP_NUM_THREADS=1
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:512
- VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
- TRITON_CACHE_DIR=/root/.triton/cache
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/gemma-4-26b-a4b-autoround-int4-mixed
- --served-model-name
- gemma-4-26b-a4b-autoround
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-32768}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "256"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- auto
- --trust-remote-code
- --enable-auto-tool-choice
- --tool-call-parser
- gemma4
- --chat-template
- /vllm-workspace/examples/tool_chat_template_gemma4.jinja
- --enable-prefix-caching
- --enable-chunked-prefill
@@ -0,0 +1,108 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 26B-A4B MoE (Intel AutoRound INT4 mixed)
# Topology: Single 3090 (TP=1)
# Drafter: none (base smoke compose — MTP via gemma-26b-it-assistant
# can be added once base boot is validated)
# KV: bfloat16 (sidesteps Ampere fp8 dispatch issues — same logic
# as gemma-4-31b/single/docker-compose.yml)
# Vision: off (limit-mm-per-prompt image=0 audio=0) for first boot
# Max ctx: 8K (conservative — first-boot smoke; bump after validation)
# Genesis: none (Gemma 4 family doesn't need Genesis — engine
# vllm-nightly-clean rides latest nightly)
# Status: 🔵 v0.7.3 ONBOARDING — first boot pending
# Active params: ~4B (128 experts × 8 active = ~4B routed)
# ---------------------------------------------------------------------------
# Gemma 4 26B-A4B-it (Intel AutoRound INT4 mixed) — first MoE model added to
# club-3090 alongside Qwen 3.6 35B-A3B.
#
# Why vllm-nightly-clean (not vllm-nightly-mtp):
# Gemma 4 family doesn't need Genesis patches (no DeltaNet quirks).
# By routing through the unconstrained nightly we get latest upstream
# features (incl. continued MoE-loader improvements) without waiting on
# Genesis re-anchor cycles.
#
# Why no drafter on this compose:
# The 26B-A4B has an external MTP assistant (google/gemma-4-26B-A4B-it-
# assistant, ~0.97 GB) — separate weights from the 31B assistant.
# Validating base boot first, then layering MTP on top in a follow-up.
#
# Models:
# target: Intel/gemma-4-26B-A4B-it-int4-mixed-AutoRound
# (~14 GB — quant mix of MoE expert layers)
# draft : none on this compose (gemma-26b-it-assistant compose TBD)
#
# KV format pinned to bf16:
# - fp8_e5m2 → blocked by gemma4_mm.py assert (vLLM allowlist)
# - fp8_e4m3 → Triton "fp8e4nv not supported" on sm_86 Ampere
# - default bf16 → sidesteps both — same pattern as 31B single
# Smaller KV pool than fp8 but at 8K initial ctx not the bottleneck.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
# Requires-sm: 7.5+
services:
vllm-gemma-4-26b-a4b-tp1:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-gemma-4-26b-a4b-tp1}"
restart: "no"
ports:
- "${ESTATE_PORT:-${PORT:-8040}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- OMP_NUM_THREADS=1
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:512
- VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
- TRITON_CACHE_DIR=/root/.triton/cache
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/gemma-4-26b-a4b-autoround-int4-mixed
- --served-model-name
- gemma-4-26b-a4b-autoround
- --tensor-parallel-size
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-8192}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "256"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- auto
- --trust-remote-code
- --enable-auto-tool-choice
- --tool-call-parser
- gemma4
- --chat-template
- /vllm-workspace/examples/tool_chat_template_gemma4.jinja
@@ -0,0 +1,79 @@
# vLLM PR #40886 overlay — AWQ compressed-tensors MoE key remapping
## What this fixes
`cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit` ships in `compressed-tensors` pack-quantized format, which stores per-expert weights with `_packed` (int32) and `_scale` (bfloat16) suffixes. The vLLM `gemma4.py::_weight_iterator` (as of bf610c2f, 2026-05-15) only handles the float-checkpoint key pattern and does not remap the `_packed` / `_scale` suffixed variants on MoE expert weights. Without this overlay, model load fails with a `KeyError` on the first packed expert key.
The fix is from [vLLM PR #40886](https://github.com/vllm-project/vllm/pull/40886) by @tajwali (open as of 2026-05-15, last updated 2026-04-25). PR author tested on **RTX 3090 24 GB** (same SKU as ours) with vLLM 0.19.1. The patch is +23 / -0 — pure insertion of 4 conditional branches in `_weight_iterator` that intercept the four `_packed` / `_scale` MoE key patterns and yield per-expert float-shaped keys that the existing FusedMoE loader path expects.
## Why an anchor-based Python patcher (not full-file replacement)
The PR is a small insertion. Full-file replacement risks shadowing unrelated upstream changes in `gemma4.py` that we DO want (PR #41745 Gemma 4 MTP support, etc.). The anchor-based patcher in `install.sh` locates the existing `if "moe.gate_up_proj" in name and weight.dim() == 3:` line inside `_weight_iterator` and inserts the patch immediately above it. Idempotent: a sentinel comment is included in the inserted block, and re-running on an already-patched file is a no-op.
## How to use
### From a compose
Bind-mount `install.sh` into the container at a known path, then invoke it from the entrypoint **before** `vllm serve` runs:
```yaml
services:
vllm-gemma-4-26b-a4b-awq-tp2:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
volumes:
- ../../patches/vllm-pr40886-awq-moe-keys/install.sh:/etc/club3090/install-pr40886.sh:ro
entrypoint:
- /bin/bash
- -c
- |
bash /etc/club3090/install-pr40886.sh
exec vllm serve "$@"
- --
command:
- --model
- /root/.cache/huggingface/gemma-4-26b-a4b-awq-4bit
# ... other flags ...
```
Same sidecar pattern as `vllm-pr35936-required-fallback/install.sh` — runs once per container start, leaves the container's RW layer in the patched state.
### Override env vars
If vLLM moves the install path of `gemma4.py`:
```bash
CLUB3090_PR40886_TARGET=/some/other/path/gemma4.py bash install.sh
```
## When to drop this overlay
When **both** of these are true:
1. PR #40886 has merged upstream
2. The engine's pinned nightly SHA is past the merge commit
Track in `docs/UPSTREAM.md`.
## Smoke test
Manual:
```bash
# Spin up a transient container, install the patch, verify the sentinel.
docker run --rm \
-v $(pwd)/install.sh:/install.sh:ro \
vllm/vllm-openai:nightly-bf610c2f56764e1b30bc6065f4ceace3d6e59036 \
bash -c 'bash /install.sh && grep -c "club3090/pr40886" /usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/gemma4.py'
```
Expected: prints `1` (sentinel present after install).
## Source PR
- PR head: `tajwali/vllm` @ `652819dad0bf9bbb0436d6660822e7aff30c3ff0` (branch `fix/gemma4-compressed-tensors-moe-key-remapping`)
- Vendored as of 2026-05-15
- Patch summary: 4 `if` branches added to `_weight_iterator` in `vllm/model_executor/models/gemma4.py`
- `moe.gate_up_proj_packed [E, 2I, H/8]` → split into per-expert `gate_proj.weight_packed` + `up_proj.weight_packed`
- `moe.gate_up_proj_scale [E, 2I, G]` → same split for scales
- `moe.down_proj_packed [E, H, I/8]` → yield per-expert `down_proj.weight_packed`
- `moe.down_proj_scale [E, H, G]` → yield per-expert `down_proj.weight_scale`
@@ -0,0 +1,104 @@
#!/usr/bin/env bash
# Install vLLM PR #40886 — compressed-tensors AWQ MoE key remapping for
# Gemma 4 26B-A4B. Without this patch, `cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit`
# fails to load with KeyError on `moe.gate_up_proj_packed` (or similar)
# because vLLM's `gemma4.py::_weight_iterator` doesn't handle the
# `_packed`/`_scale` suffix on MoE expert weights.
#
# PR: https://github.com/vllm-project/vllm/pull/40886
# Head commit at time of vendor: 652819dad0bf9bbb0436d6660822e7aff30c3ff0
# Author tested on: RTX 3090 24 GB (same SKU as ours), vLLM 0.19.1
#
# WHY a Python anchor-based patcher instead of full-file replacement:
# The PR's diff is +23 / -0 — pure insertion before an existing branch
# in `_weight_iterator`. Anchor-based insertion is robust to upstream
# drift around the function (vLLM nightly may add unrelated logic
# elsewhere in gemma4.py without breaking our patch).
#
# Idempotent: the patcher checks for a sentinel comment before inserting,
# so re-running it on an already-patched file is a no-op.
#
# Drop when: vLLM PR #40886 merges upstream AND our engine pin bumps
# past the merge commit.
set -euo pipefail
# Container's vLLM install path. Override via env if vLLM moves.
GEMMA4_PY="${CLUB3090_PR40886_TARGET:-/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/gemma4.py}"
if [ ! -f "$GEMMA4_PY" ]; then
echo "[club3090/pr40886] ERROR: $GEMMA4_PY not found; aborting overlay install" >&2
exit 1
fi
python3 - <<'PY'
import os
import re
import sys
target = os.environ.get(
"CLUB3090_PR40886_TARGET",
"/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/gemma4.py",
)
SENTINEL = "# PATCH: AWQ compressed-tensors key remapping (club3090/pr40886)"
PATCH_BLOCK = ''' # PATCH: AWQ compressed-tensors key remapping (club3090/pr40886)
if "moe.gate_up_proj_packed" in name and weight.dim() == 3:
mid = weight.size(1) // 2
for e in range(weight.size(0)):
base = name.replace("moe.", f"moe.experts.{e}.")
yield base.replace("gate_up_proj_packed", "gate_proj.weight_packed"), weight[e, :mid]
yield base.replace("gate_up_proj_packed", "up_proj.weight_packed"), weight[e, mid:]
continue
if "moe.gate_up_proj_scale" in name and weight.dim() == 3:
mid = weight.size(1) // 2
for e in range(weight.size(0)):
base = name.replace("moe.", f"moe.experts.{e}.")
yield base.replace("gate_up_proj_scale", "gate_proj.weight_scale"), weight[e, :mid]
yield base.replace("gate_up_proj_scale", "up_proj.weight_scale"), weight[e, mid:]
continue
if "moe.down_proj_packed" in name and weight.dim() == 3:
for e in range(weight.size(0)):
yield name.replace("moe.", f"moe.experts.{e}.").replace("down_proj_packed", "down_proj.weight_packed"), weight[e]
continue
if "moe.down_proj_scale" in name and weight.dim() == 3:
for e in range(weight.size(0)):
yield name.replace("moe.", f"moe.experts.{e}.").replace("down_proj_scale", "down_proj.weight_scale"), weight[e]
continue
'''
# Insert immediately above the existing `if "moe.gate_up_proj" in name and weight.dim() == 3:`
# branch inside `_weight_iterator`. This is the line the upstream PR puts the patch above.
ANCHOR_RE = re.compile(
r'^(?P<indent>[ \t]+)(?P<line>if "moe\.gate_up_proj" in name and weight\.dim\(\) == 3:)',
re.MULTILINE,
)
with open(target, "r", encoding="utf-8") as f:
src = f.read()
if SENTINEL in src:
print(f"[club3090/pr40886] {target}: sentinel present, patch already applied; no-op", file=sys.stderr)
sys.exit(0)
m = ANCHOR_RE.search(src)
if not m:
print(
f"[club3090/pr40886] ERROR: anchor 'if \"moe.gate_up_proj\" in name and weight.dim() == 3:' "
f"not found in {target}. vLLM nightly may have changed gemma4.py — overlay needs re-anchoring.",
file=sys.stderr,
)
sys.exit(1)
# Insertion point: start of the matched line
insert_at = m.start()
patched = src[:insert_at] + PATCH_BLOCK + src[insert_at:]
with open(target, "w", encoding="utf-8") as f:
f.write(patched)
print(f"[club3090/pr40886] {target}: PR #40886 (AWQ MoE key remapping) applied", file=sys.stderr)
PY
echo "[club3090/pr40886] install complete" >&2
@@ -78,6 +78,10 @@ services:
- ../../cache/triton_awq:/root/.triton/cache
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -112,6 +116,7 @@ services:
pip install --quiet --upgrade transformers==5.8.0
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
bash /etc/club3090/install-pr41800.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
@@ -91,7 +91,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -145,6 +145,10 @@ services:
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -177,6 +181,7 @@ services:
pip install --quiet --upgrade transformers==5.8.0
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
bash /etc/club3090/install-pr41800.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
@@ -110,6 +110,10 @@ services:
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -142,6 +146,7 @@ services:
pip install --quiet --upgrade transformers==5.8.0
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
bash /etc/club3090/install-pr41800.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
@@ -50,7 +50,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -220,6 +220,10 @@ services:
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -253,6 +257,7 @@ services:
echo "[club3090-tq3] Launching vllm serve..." >&2
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
bash /etc/club3090/install-pr41800.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
@@ -124,6 +124,10 @@ services:
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -151,6 +155,7 @@ services:
- |
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
bash /etc/club3090/install-pr41800.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
@@ -50,7 +50,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 32
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
# Requires-sm: 9.0+
@@ -22,7 +22,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -50,7 +50,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -75,6 +75,9 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode
# HTTP 400 on get_max_tokens(). See patches/vllm-pr41800.../README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -114,6 +117,7 @@ services:
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
@@ -99,6 +99,9 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode
# HTTP 400 on get_max_tokens(). See patches/vllm-pr41800.../README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -138,6 +141,7 @@ services:
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
@@ -51,7 +51,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -87,6 +87,9 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode
# HTTP 400 on get_max_tokens(). See patches/vllm-pr41800.../README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -128,6 +131,7 @@ services:
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
@@ -57,6 +57,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -88,6 +92,7 @@ services:
- |
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
@@ -35,7 +35,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -105,6 +105,10 @@ services:
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -246,6 +250,7 @@ services:
# v7.72.2 (P78 + PN34). Mounts and invocations dropped 2026-05-05.
# VLLM_ENFORCE_EAGER=1 in compose/.env disables CUDA graphs — use on
# hardware where Cliff 2 GDN activation spikes occur at runtime.
bash /etc/club3090/install-pr41800.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
@@ -104,6 +104,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -136,6 +140,7 @@ services:
- |
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
@@ -88,6 +88,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -119,6 +123,7 @@ services:
- |
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
@@ -84,6 +84,9 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode
# HTTP 400 on get_max_tokens(). See patches/vllm-pr41800.../README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -222,6 +225,7 @@ services:
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
@@ -92,6 +92,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -128,6 +132,7 @@ services:
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
@@ -59,7 +59,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 4
# Tensor-parallel: 4
services:
@@ -161,6 +161,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -290,6 +294,7 @@ services:
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
@@ -132,6 +132,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -221,6 +225,7 @@ services:
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
@@ -149,6 +149,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -315,6 +319,7 @@ services:
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
@@ -159,6 +159,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -332,6 +336,7 @@ services:
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
@@ -127,6 +127,10 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# vLLM PR #41800 (truncate_prompt_tokens kwarg) overlay — fixes opencode-style
# HTTP 400 on get_max_tokens(). No-op on post-fix nightlies.
# See ../../patches/vllm-pr41800-truncate-prompt-tokens/README.md.
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
@@ -256,6 +260,7 @@ services:
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
bash /etc/club3090/install-pr41800.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
@@ -33,7 +33,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 20
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -35,7 +35,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -1358,9 +1358,13 @@ class OpenAIServingChat(OpenAIServing):
request_metadata.final_usage_info = usage
# club-3090 patch: `prompt_routed_experts` was removed from
# `RequestOutput` in vLLM bf610c2f. Use getattr so the overlay tolerates
# both pre-bf610c2f and post-bf610c2f layouts. Dense path produces None.
_prompt_routed_experts_raw = getattr(final_res, "prompt_routed_experts", None)
prompt_routed_experts = None
if final_res.prompt_routed_experts is not None:
prompt_routed_experts = final_res.prompt_routed_experts.tolist()
if _prompt_routed_experts_raw is not None:
prompt_routed_experts = _prompt_routed_experts_raw.tolist()
response = ChatCompletionResponse(
id=request_id,
@@ -38,11 +38,24 @@ from vllm.entrypoints.openai.engine.protocol import (
)
from vllm.entrypoints.openai.models.serving import OpenAIServingModels
from vllm.entrypoints.openai.responses.protocol import ResponsesRequest
from vllm.entrypoints.openai.speech_to_text.protocol import (
TranscriptionRequest,
TranscriptionResponse,
TranslationRequest,
)
# club-3090 patch: speech_to_text was relocated from
# `vllm.entrypoints.openai.speech_to_text` (pre-bf610c2f) to
# `vllm.entrypoints.speech_to_text` and split into transcription/ + translation/
# subpackages. Try old path first (matches captured SHA), fall back to new.
try:
from vllm.entrypoints.openai.speech_to_text.protocol import (
TranscriptionRequest,
TranscriptionResponse,
TranslationRequest,
)
except ImportError:
from vllm.entrypoints.speech_to_text.transcription.protocol import (
TranscriptionRequest,
TranscriptionResponse,
)
from vllm.entrypoints.speech_to_text.translation.protocol import (
TranslationRequest,
)
from vllm.entrypoints.serve.disagg.protocol import GenerateRequest, GenerateResponse
from vllm.entrypoints.serve.tokenize.protocol import (
DetokenizeRequest,
@@ -615,6 +628,19 @@ class OpenAIServing:
except ValueError:
return None
# club-3090 patch: bf610c2f's `chat_completion/serving.py`, `completion/serving.py`,
# and `responses/serving.py` call `self._with_kv_transfer_rejection_cleanup(...)`.
# The pre-bf610c2f base this overlay was captured against didn't define it.
# club-3090 doesn't run disaggregated KV-transfer / remote-prefill workflows,
# so this stub just delegates to the awaitable — the cleanup branch never fires.
async def _with_kv_transfer_rejection_cleanup(
self,
awaitable,
request,
raw_request,
):
return await awaitable
@staticmethod
def _parse_tool_calls_from_content(
request: ResponsesRequest | ChatCompletionRequest,
@@ -0,0 +1,94 @@
# vLLM PR #41800 overlay — `truncate_prompt_tokens` kwarg on `get_max_tokens`
## What this fixes
Agentic clients (opencode, codex-cli, and similar IDE/agent runtimes) send `truncate_prompt_tokens` on chat-completion requests. Pre-[vLLM PR #41800](https://github.com/vllm-project/vllm/pull/41800), `vllm.entrypoints.utils.get_max_tokens()` doesn't accept that kwarg — and the kwarg propagates from the request handler down into the function call — so requests fail with:
```
HTTP 400: {"error":{"message":"get_max_tokens() got an unexpected keyword argument 'truncate_prompt_tokens'",...}}
```
The fix is upstream PR #41800 (merged 2026-05-06 at commit `d5b31c95`). It adds the kwarg to the function signature and a small body block that clamps `input_length` to `min(input_length, truncate_prompt_tokens or max_model_len)` before the existing length check.
## When this overlay is needed
This overlay is needed on engines pinned to vLLM SHAs that **predate `d5b31c95`**:
| Engine | Pinned SHA | Pre-fix? |
|---|---|---|
| `vllm-nightly-mtp` | `01d4d1ad` (2026-05-04) | ✅ needs overlay |
| `vllm-nightly-dflash` | `e47c98ef` (~2026-05-05) | ✅ needs overlay (20 commits behind d5b31c95) |
| `vllm-nightly-full` | `e47c98ef` | ✅ needs overlay |
| `vllm-nightly-clean` | `bf610c2f` (2026-05-15) | ❌ already includes fix |
If a compose routes through `vllm-nightly-clean`, the overlay is unnecessary — the function signature already accepts the kwarg upstream.
## How the overlay works
`install.sh` is a Python anchor-based in-place patcher. It does two surgical edits to the in-container `/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/utils.py`:
1. **Signature**: adds `truncate_prompt_tokens: int | None = None,` to `get_max_tokens`'s signature, anchored to the existing `override_max_tokens: int | None = None,` line.
2. **Body**: inserts a 6-line truncation-aware `input_length` adjustment block before the existing `if max_model_len < input_length:` check, anchored to that line.
Each insertion carries a sentinel comment (`# PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)`) so re-running the install on an already-patched file is a no-op. Post-patch the file is AST-validated before write.
Why anchor-based and not full-file replacement: the PR diff is +14 / -0 across a 200-line file — replacing the full file would shadow other upstream changes in `utils.py`. Anchor-based insertion is drift-resistant to unrelated upstream movement.
## Composes that wire this overlay in (as of v0.7.3 ship)
* `models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml` (gpu-mode `27b`)
* `models/qwen3.6-27b/vllm/compose/dual/turbo.yml` (gpu-mode `27b-turbo`)
* `models/qwen3.6-27b/vllm/compose/dual/dflash.yml` (gpu-mode `27b-dflash`)
* `models/qwen3.6-27b/vllm/compose/dual/dflash-noviz.yml` (gpu-mode `27b-dflash-noviz`, the compose from issue #138)
## How to add this overlay to another affected compose
In any compose that routes through `vllm-nightly-mtp` / `vllm-nightly-dflash` / `vllm-nightly-full`, add:
1. **Volume mount** in the `volumes:` block:
```yaml
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
```
2. **Install line** in the `entrypoint:` bash script, before `exec vllm serve`:
```bash
bash /etc/club3090/install-pr41800.sh
```
Run `bash install.sh` (the file in this directory) standalone to test against a transient vLLM container before wiring into a compose. See the smoke test in the next section.
## Smoke test
```bash
docker run --rm --entrypoint /bin/bash \
-v $(pwd)/install.sh:/install.sh:ro \
vllm/vllm-openai:nightly-01d4d1ad375dc5854779c593eee093bcebb0cada \
-c '
python3 -c "from vllm.entrypoints.utils import get_max_tokens; import inspect; print(inspect.signature(get_max_tokens))"
bash /install.sh
python3 -c "from vllm.entrypoints.utils import get_max_tokens; import inspect; print(inspect.signature(get_max_tokens))"
'
```
Expected: signature lacks `truncate_prompt_tokens` BEFORE install, has it AFTER. Verified on `01d4d1ad` (2026-05-15).
## When to drop this overlay
When **both** are true:
1. PR #41800 has merged upstream (it has — 2026-05-06 at `d5b31c95`)
2. The engine's pinned nightly SHA bumps past `d5b31c95`
For the Genesis-anchored engines, the bump happens with Sander's next Genesis release cycle (v7.73.x). For `vllm-nightly-dflash` and `vllm-nightly-full`, the bump happens when their respective overlays (PR #41703 DFlash, PR #42102 INT8 PTH KV) are re-validated against a newer nightly.
Track in `docs/UPSTREAM.md`.
## Source
- vLLM PR #41800: https://github.com/vllm-project/vllm/pull/41800
- Merged commit: `d5b31c95`
- Tracking issue: noonghunna/club-3090#139
- Triggered by: noonghunna/club-3090#138 (SEVENID's opencode boot failure)
- Patch summary: +7 lines in `vllm/entrypoints/utils.py` (the actual fix) + 5 call-site forward-compat additions in other files (we skip those — the signature fix alone unblocks all known TypeError reports)
@@ -0,0 +1,141 @@
#!/usr/bin/env bash
# Install vLLM PR #41800 — `truncate_prompt_tokens` kwarg on get_max_tokens.
#
# WHY THIS OVERLAY EXISTS:
# opencode (and other agentic clients like codex-cli) send `truncate_prompt_tokens`
# on chat-completion requests. Pre-#41800, vLLM's `get_max_tokens()` doesn't
# accept that kwarg — and somewhere upstream of the function the kwarg gets
# unpacked into the call — so requests fail with:
# HTTP 400: get_max_tokens() got an unexpected keyword argument 'truncate_prompt_tokens'
#
# PR: https://github.com/vllm-project/vllm/pull/41800
# Merged: 2026-05-06 at commit d5b31c95
# Affected pins on master:
# - vllm-nightly-mtp (01d4d1ad, 2026-05-04) — pre-fix
# - vllm-nightly-dflash (e47c98ef) — pre-fix
# - vllm-nightly-full (e47c98ef) — pre-fix
# (vllm-nightly-clean at bf610c2f is POST-fix; doesn't need the overlay)
#
# Tracking issue: #139 (noonghunna/club-3090)
# Triggered by: #138 — SEVENID's opencode boot failure on dual-dflash-noviz.
#
# WHY A PYTHON ANCHOR-BASED PATCHER:
# The PR is +7 lines in `vllm/entrypoints/utils.py` (the actual fix) plus a
# handful of forward-compat call-site additions in 5 other files. The
# function-signature change in utils.py is the ONLY thing required to fix
# the TypeError — once `get_max_tokens` accepts the kwarg, requests stop
# crashing. The call-site changes are nice-to-have semantic completeness
# (actually applying the truncation), so we patch those too via anchors.
#
# Idempotent: each anchor checks for a sentinel marker before inserting.
set -euo pipefail
# Container's vLLM install path. Override via env if vLLM moves.
SITE_PACKAGES="${CLUB3090_PR41800_SITE_PACKAGES:-/usr/local/lib/python3.12/dist-packages}"
UTILS_PY="$SITE_PACKAGES/vllm/entrypoints/utils.py"
if [ ! -f "$UTILS_PY" ]; then
echo "[club3090/pr41800] ERROR: $UTILS_PY not found; aborting overlay install" >&2
exit 1
fi
python3 - <<'PY'
import os
import re
import sys
site_packages = os.environ.get(
"CLUB3090_PR41800_SITE_PACKAGES",
"/usr/local/lib/python3.12/dist-packages",
)
utils_py = f"{site_packages}/vllm/entrypoints/utils.py"
SENTINEL = "# PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)"
# Function signature update: add `truncate_prompt_tokens: int | None = None,`
# as a kwarg on `get_max_tokens`. Anchor on the existing line that closes
# the signature (`override_max_tokens: int | None = None,` line right before `) -> int:`).
SIGNATURE_ANCHOR_RE = re.compile(
r'^(?P<indent>[ \t]+)override_max_tokens: int \| None = None,\n(?P<close>[ \t]*\) -> int:)',
re.MULTILINE,
)
SIGNATURE_INSERT = ''' override_max_tokens: int | None = None,
truncate_prompt_tokens: int | None = None, # PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)
) -> int:'''
# Body update: insert truncation-aware input_length adjustment BEFORE the
# `if max_model_len < input_length:` check. Anchor on that line.
BODY_ANCHOR_RE = re.compile(
r'^(?P<indent>[ \t]+)if max_model_len < input_length:',
re.MULTILINE,
)
BODY_INSERT = ''' # PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)
if truncate_prompt_tokens is not None:
limit = truncate_prompt_tokens
input_length = min(
input_length,
max_model_len if limit == -1 else limit,
)
if max_model_len < input_length:'''
with open(utils_py, "r", encoding="utf-8") as f:
src = f.read()
if SENTINEL in src:
print(f"[club3090/pr41800] {utils_py}: sentinel present, patch already applied; no-op", file=sys.stderr)
sys.exit(0)
# Upstream-fix detection: if the function signature already accepts the kwarg
# (i.e. the engine pinned a post-#41800 nightly), the overlay is unnecessary
# and should no-op gracefully so composes that mount it on a post-fix image
# (e.g. via vllm-nightly-clean) still boot cleanly.
UPSTREAM_RE = re.compile(
r'def get_max_tokens\([^)]*truncate_prompt_tokens\b',
re.DOTALL,
)
if UPSTREAM_RE.search(src):
print(f"[club3090/pr41800] {utils_py}: upstream get_max_tokens() already accepts truncate_prompt_tokens; no-op", file=sys.stderr)
sys.exit(0)
# Apply signature patch first (so the function accepts the kwarg)
m = SIGNATURE_ANCHOR_RE.search(src)
if not m:
print(
f"[club3090/pr41800] ERROR: signature anchor "
f"'override_max_tokens: int | None = None, ... ) -> int:' not found in {utils_py}. "
f"vLLM nightly may have changed entrypoints/utils.py — overlay needs re-anchoring.",
file=sys.stderr,
)
sys.exit(1)
src = SIGNATURE_ANCHOR_RE.sub(SIGNATURE_INSERT, src, count=1)
# Apply body patch
m = BODY_ANCHOR_RE.search(src)
if not m:
print(
f"[club3090/pr41800] ERROR: body anchor 'if max_model_len < input_length:' not found in {utils_py} "
f"after signature patch. vLLM nightly diverged unexpectedly — overlay needs re-anchoring.",
file=sys.stderr,
)
sys.exit(1)
src = BODY_ANCHOR_RE.sub(BODY_INSERT, src, count=1)
with open(utils_py, "w", encoding="utf-8") as f:
f.write(src)
# Quick validity check
import ast
try:
ast.parse(src)
except SyntaxError as e:
print(f"[club3090/pr41800] ERROR: post-patch utils.py is not valid Python: {e}", file=sys.stderr)
sys.exit(1)
print(f"[club3090/pr41800] {utils_py}: signature + body patches applied (truncate_prompt_tokens kwarg)", file=sys.stderr)
PY
echo "[club3090/pr41800] install complete" >&2
@@ -0,0 +1,103 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen 3.6 35B-A3B MoE (AutoRound INT4, ~20 GB weights)
# Topology: Dual 3090 (TP=2)
# Drafter: MTP n=3 (built-in head — qwen-mtp-builtin DrafterProfile)
# KV: fp8_e5m2 (no TQ3 — Genesis-only, not available here)
# Vision: off (limit-mm-per-prompt image=0 audio=0)
# Max ctx: 16K (conservative — matches preview.yml for direct A/B)
# Genesis: none — preview-MTP path on vllm-nightly-clean
# Status: 🔵 PREVIEW — measures MTP n=3 uplift on the Qwen MoE
# ---------------------------------------------------------------------------
# Same shape as preview.yml plus the built-in MTP head wired via
# `--speculative-config`. Qwen 3.6 35B-A3B's `mtp_num_hidden_layers: 1`
# (per the model config) provides the head natively — no separate drafter
# download needed. n=3 mirrors the production Qwen 27B convention.
#
# Direct A/B target: preview.yml gives 182.68 / 177.45 wall TPS (2026-05-15,
# no spec-decode). This compose measures MTP's contribution on the same
# rig + config so calibration data has a tight baseline pair.
#
# Cliff 2 mitigations still UNAVAILABLE without Genesis — keep max_ctx
# below ~21K accumulated until Genesis v7.73.x re-anchors on a post-#42521
# nightly. Same caveat as preview.yml.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 7.5+
services:
vllm-qwen36-35b-a3b-preview-mtp-tp2:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-qwen36-35b-a3b-preview-mtp-tp2}"
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${ESTATE_PORT:-${PORT:-8052}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- PYTORCH_CUDA_ALLOC_CONF=${PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True}
- OMP_NUM_THREADS=1
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/qwen3.6-35b-a3b-autoround-int4
- --served-model-name
- qwen3.6-35b-a3b-autoround
- --quantization
- auto_round
- --dtype
- float16
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-16384}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "1"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- "${KV_CACHE_DTYPE:-fp8_e5m2}"
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
# Built-in MTP head — n=3 matches the production Qwen 27B convention.
# `model` points at the same model dir; vLLM extracts the MTP head
# from `mtp_num_hidden_layers: 1` in the model config.
- --speculative-config
- '{"method":"mtp","num_speculative_tokens":3}'
@@ -0,0 +1,124 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen 3.6 35B-A3B MoE (AutoRound INT4, ~20 GB weights)
# Topology: Dual 3090 (TP=2)
# Drafter: none (smoke compose — built-in MTP head added in a follow-up
# once base boot is validated)
# KV: fp8_e5m2 (no TQ3 — Genesis-only, not available here)
# Vision: off (limit-mm-per-prompt image=0 audio=0) for first boot
# Max ctx: 16K (conservative for first boot; can extend after validation)
# Genesis: none — preview path on vllm-nightly-clean
# Status: 🔵 PREVIEW — v0.7.3 MoE onboarding primary bench target
# Caveats: Cliff 2 mitigations UNAVAILABLE without Genesis. Do NOT use
# past ~21-26K accumulated context — risk of OOM / instability.
# Production path (Genesis-anchored, TQ3 KV, MTP, higher ctx)
# is parked until Genesis v7.73.x re-anchors on a post-#42521
# nightly. This compose exists to exercise PR #42521 (qwen3_5_moe
# weight loading) and produce a first benchmark row.
# ---------------------------------------------------------------------------
# Qwen 3.6 35B-A3B (AutoRound INT4) — dual-card preview alongside Gemma 4
# 26B-A4B for v0.7.3 MoE onboarding.
#
# Why TP=2 (not TP=1): weights ~20 GB. Single-card 24 GB leaves only ~2 GB
# for KV pool + activations + cudagraph capture — boot OOMs are highly
# likely. TP=2 splits weights to ~10 GB/card, leaving ~12 GB/card for KV.
# num_kv_heads=2 caps valid_tp at [1, 2] — TP=2 is the maximum.
#
# Why vllm-nightly-clean (not vllm-nightly-mtp):
# vllm-nightly-clean rides nightly-bf610c2f (2026-05-15), which INCLUDES
# PR #42521 (qwen3_5_moe weight loading, merged 2026-05-14). The Genesis-
# anchored vllm-nightly-mtp is pinned to nightly-01d4d1ad (Sander's
# v7.72.2 PROD pin), which PRE-DATES that fix. Trade-off: no Genesis
# patches → Cliff 2 stabilization missing, no TQ3 KV. Acceptable for
# low-ctx smoke + first bench row; not acceptable for production
# long-ctx workloads.
#
# Re-test trigger to graduate this compose to production-anchored variant:
# Genesis v7.73.x is released → bump vllm-nightly-mtp.spec to a
# post-#42521 nightly SHA → port this compose to engine vllm-nightly-mtp
# → drop the "preview" naming → add TQ3 KV + MTP variant.
#
# Models:
# target: Qwen/Qwen3-MoE-A3B-Instruct-AutoRound-Int4-mixed (path:
# qwen3.6-35b-a3b-autoround-int4, arch:
# Qwen3_5MoeForConditionalGeneration, ~20 GB on disk)
# draft : none on this compose
#
# KV format:
# - fp8_e5m2 chosen over fp8_e4m3 (Triton fp8e4nv unsupported on sm_86).
# - TQ3 not available without Genesis.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 7.5+
services:
vllm-qwen36-35b-a3b-preview-tp2:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-qwen36-35b-a3b-preview-tp2}"
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${ESTATE_PORT:-${PORT:-8051}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- PYTORCH_CUDA_ALLOC_CONF=${PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True}
- OMP_NUM_THREADS=1
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/qwen3.6-35b-a3b-autoround-int4
- --served-model-name
- qwen3.6-35b-a3b-autoround
- --quantization
- auto_round
- --dtype
- float16
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-16384}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "1"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- "${KV_CACHE_DTYPE:-fp8_e5m2}"
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
@@ -0,0 +1,118 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen 3.6 35B-A3B MoE (AutoRound INT4, ~20 GB weights)
# Topology: Single 3090 (TP=1)
# Drafter: none (smoke compose — built-in MTP head added in a follow-up
# once base boot is validated)
# KV: fp8_e5m2 (no TQ3 — Genesis-only, not available here)
# Vision: off (limit-mm-per-prompt image=0 audio=0) for first boot
# Max ctx: 8K (conservative — KV pool tight after 20 GB weights on 24 GB card)
# Genesis: none — preview path on vllm-nightly-clean
# Status: 🔵 PREVIEW — v0.7.3 MoE onboarding smoke
# Caveats: Cliff 2 mitigations UNAVAILABLE without Genesis. Do NOT use
# past ~21-26K accumulated context — risk of OOM / instability.
# Production path (Genesis-anchored, TQ3 KV, MTP, higher ctx)
# is parked until Genesis v7.73.x re-anchors on a post-#42521
# nightly. This compose exists to exercise PR #42521 (qwen3_5_moe
# weight loading) and produce a first benchmark row.
# ---------------------------------------------------------------------------
# Qwen 3.6 35B-A3B (AutoRound INT4) — first MoE model from the Qwen3-Next
# family on club-3090, alongside Gemma 4 26B-A4B.
#
# Why vllm-nightly-clean (not vllm-nightly-mtp):
# vllm-nightly-clean rides nightly-bf610c2f (2026-05-15), which INCLUDES
# PR #42521 (qwen3_5_moe weight loading, merged 2026-05-14). The Genesis-
# anchored vllm-nightly-mtp is pinned to nightly-1acd67a7 (2026-05-08),
# which PRE-DATES that fix. Trade-off: no Genesis patches → Cliff 2
# stabilization missing, no TQ3 KV. Acceptable for low-ctx smoke + first
# bench row; not acceptable for production long-ctx workloads.
#
# Re-test trigger to graduate this compose to production-anchored variant:
# Genesis v7.73.x is released → bump vllm-nightly-mtp.spec to a post-#42521
# nightly SHA → port this compose to engine vllm-nightly-mtp → drop the
# "preview" naming → add a Cliff-2-mitigated dual.yml.
#
# Models:
# target: Qwen/Qwen3-MoE-A3B-Instruct-AutoRound-Int4-mixed (path:
# qwen3.6-35b-a3b-autoround-int4, arch:
# Qwen3_5MoeForConditionalGeneration, ~20 GB on disk)
# draft : none on this compose
#
# KV format:
# - fp8_e5m2 chosen over fp8_e4m3 (Triton fp8e4nv unsupported on sm_86).
# - TQ3 not available without Genesis.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-clean
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
# Requires-sm: 7.5+
services:
vllm-qwen36-35b-a3b-preview:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
container_name: "${ESTATE_CONTAINER:-vllm-qwen36-35b-a3b-preview}"
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${ESTATE_PORT:-${PORT:-8050}}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
environment:
- NVIDIA_VISIBLE_DEVICES=${ESTATE_GPUS:-${NVIDIA_VISIBLE_DEVICES:-all}}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- PYTORCH_CUDA_ALLOC_CONF=${PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True}
- OMP_NUM_THREADS=1
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
command:
- --host
- 0.0.0.0
- --port
- "8000"
- --model
- /root/.cache/huggingface/qwen3.6-35b-a3b-autoround-int4
- --served-model-name
- qwen3.6-35b-a3b-autoround
- --quantization
- auto_round
- --dtype
- float16
- --tensor-parallel-size
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-8192}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "1"
- --max-num-batched-tokens
- "4096"
- --limit-mm-per-prompt
- '{"image":0,"audio":0}'
- --kv-cache-dtype
- "${KV_CACHE_DTYPE:-fp8_e5m2}"
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
+29
View File
@@ -12,6 +12,13 @@ COMPOSE_BASE="$CLUB3090_DIR/services"
DUAL_27B_DIR="$CLUB3090_DIR/models/qwen3.6-27b/vllm/compose/dual"
GEMMA_DUAL_DIR="$CLUB3090_DIR/models/gemma-4-31b/vllm/compose/dual"
# Estate planner state file (v0.7.0+). Instances booted via launch.sh --estate
# or --estate-file are tracked here and persist via Docker `restart:
# unless-stopped`, so they DO survive a plain mode-switch unless explicitly
# torn down via launch.sh --down-estate. mode_off uses this path to clean
# them up alongside the older vLLM/Gemma/ComfyUI services.
ESTATE_YAML="${HOME}/.club3090/estate.yml"
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
RED='\033[0;31m'
@@ -559,11 +566,33 @@ mode_bigmodel() {
echo -e " --n-gpu-layers 99 --ctx-size 32768 --host 0.0.0.0 --port 8001"
}
stop_estate() {
# Tear down any estate-managed instances (launch.sh --estate-file or --estate
# bookings persist via Docker `restart: unless-stopped`). No-op if no estate
# plan exists or launch.sh is unavailable.
if [[ ! -f "$ESTATE_YAML" ]]; then
return 0
fi
if ! command -v bash >/dev/null 2>&1 || [[ ! -x "$CLUB3090_DIR/scripts/launch.sh" ]]; then
return 0
fi
if ! python3 -c "import yaml; d=yaml.safe_load(open('$ESTATE_YAML')); raise SystemExit(0 if d and d.get('estate') else 1)" 2>/dev/null; then
return 0 # empty/missing estate list
fi
printf " ${RED}▼${NC} Stopping estate-managed instances..."
if bash "$CLUB3090_DIR/scripts/launch.sh" --down-estate "$ESTATE_YAML" >/dev/null 2>&1; then
echo "done"
else
echo "skipped (no instances or already down)"
fi
}
mode_off() {
echo -e "${CYAN}═══ Stopping ALL services ═══${NC}"
stop_all_27b
stop_all_gemma
stop_comfyui
stop_estate
for svc in "${SERVICES[@]}"; do
stop_service "$svc"
done
+89
View File
@@ -11,8 +11,10 @@
# bash scripts/launch.sh --variant <name> # skip wizard, boot directly
# bash scripts/launch.sh --estate # multi-model estate wizard
# bash scripts/launch.sh --estate-file <path> # boot an existing estate plan
# bash scripts/launch.sh --estate-file <path> --parallel --parallel-jobs 3 --parallel-stagger 30
# bash scripts/launch.sh --validate-estate <path> # validate estate.yml, no boot
# bash scripts/launch.sh --down-estate <path> # stop estate instances
# bash scripts/launch.sh --topology # print GPU topology advisory, no boot
# bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1
# bash scripts/launch.sh --engine vllm --cards 1 # deprecated; prefer --gpus
# bash scripts/launch.sh --workload long-ctx-single # profile-aware filter
@@ -66,6 +68,10 @@ ESTATE_FILE=""
VALIDATE_ESTATE=""
DOWN_ESTATE=""
ONLY_NAMES=""
TOPOLOGY_ONLY=0
PARALLEL_BOOT=0
PARALLEL_JOBS=""
PARALLEL_STAGGER=""
CARDS=""
VARIANT=""
MODEL_NAME=""
@@ -87,6 +93,10 @@ while [[ $# -gt 0 ]]; do
--validate-estate) VALIDATE_ESTATE="$2"; shift 2 ;;
--down-estate) DOWN_ESTATE="$2"; shift 2 ;;
--only) ONLY_NAMES="$2"; shift 2 ;;
--topology) TOPOLOGY_ONLY=1; SKIP_PREFLIGHT=1; shift ;;
--parallel) PARALLEL_BOOT=1; shift ;;
--parallel-jobs) PARALLEL_BOOT=1; PARALLEL_JOBS="$2"; shift 2 ;;
--parallel-stagger) PARALLEL_BOOT=1; PARALLEL_STAGGER="$2"; shift 2 ;;
--engine) ENGINE="$2"; shift 2 ;;
--workload) WORKLOAD_ID="$2"; shift 2 ;;
--drafter) DRAFTER_ID="$2"; shift 2 ;;
@@ -602,6 +612,77 @@ selected_gpu_profile_spec() {
printf '%s' "$joined"
}
select_topology_gpus() {
GPU_LINES="$(compose_hw_detect_gpus 2>/dev/null || true)"
[[ -n "$GPU_LINES" ]] || return 1
CARD_INDICES=()
CARD_NAMES=()
CARD_MEM_MIB=()
CARD_SM=()
if [[ -n "$GPU_ARG" && "$GPU_ARG" != "all" ]]; then
IFS=',' read -ra _launch_topology_tokens <<< "$GPU_ARG"
local idx
for idx in "${_launch_topology_tokens[@]}"; do
idx="$(_compose_meta_trim "$idx")"
[[ -z "$idx" ]] && continue
gpu_exists "$idx" || { echo "[launch] ERROR: requested GPU ${idx}, but it was not detected." >&2; exit 1; }
append_selected_gpu "$idx"
done
elif [[ -n "$CARDS" ]]; then
[[ "$CARDS" =~ ^[0-9]+$ && "$CARDS" -ge 1 ]] || { echo "[launch] ERROR: --cards expects a positive integer." >&2; exit 1; }
local idx name mem_mib sm selected=0
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
append_selected_gpu "$idx"
selected=$((selected + 1))
(( selected >= CARDS )) && break
done <<< "$GPU_LINES"
(( selected == CARDS )) || { echo "[launch] ERROR: --cards ${CARDS} requested, but only ${selected} GPU(s) were detected." >&2; exit 1; }
else
local idx name mem_mib sm
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
append_selected_gpu "$idx"
done <<< "$GPU_LINES"
fi
[[ "${#CARD_INDICES[@]}" -gt 0 ]] || return 1
SELECTED_GPU_CSV="$(IFS=','; echo "${CARD_INDICES[*]}")"
summarize_selected_vram >/dev/null
return 0
}
print_topology_advisory() {
local output
output="$(python3 "$LAUNCH_PROFILE" topology --gpu-spec "$(selected_gpu_profile_spec)" --format wizard 2>&1)" || {
echo "$output" >&2
exit 2
}
if [[ -n "$output" ]]; then
echo "$output" >&2
fi
}
print_topology_and_exit() {
local output
if ! select_topology_gpus; then
echo "Detected hardware:"
echo " no NVIDIA GPUs detected"
echo ""
echo "Topology class: unavailable"
echo ""
echo "For details, see docs/MULTI_CARD.md."
exit 0
fi
output="$(python3 "$LAUNCH_PROFILE" topology --gpu-spec "$(selected_gpu_profile_spec)" --format standalone 2>&1)" || {
echo "$output" >&2
exit 0
}
echo "$output"
exit 0
}
launch_nvlink_active() {
if [[ "${#CARD_INDICES[@]}" -ne 2 ]]; then
printf '0'
@@ -927,11 +1008,18 @@ if [[ "$ESTATE_MODE" -eq 1 || -n "$ESTATE_FILE" ]]; then
else
_estate_cmd=(python3 "$ESTATE_HELPER" boot --file "$ESTATE_FILE")
[[ -n "$ONLY_NAMES" ]] && _estate_cmd+=(--only "$ONLY_NAMES")
[[ "$PARALLEL_BOOT" -eq 1 ]] && _estate_cmd+=(--parallel)
[[ -n "$PARALLEL_JOBS" ]] && _estate_cmd+=(--parallel-jobs "$PARALLEL_JOBS")
[[ -n "$PARALLEL_STAGGER" ]] && _estate_cmd+=(--parallel-stagger "$PARALLEL_STAGGER")
fi
"${_estate_cmd[@]}"
exit $?
fi
if [[ "$TOPOLOGY_ONLY" -eq 1 ]]; then
print_topology_and_exit
fi
# --- wizard ---
if [[ -z "$VARIANT" ]]; then
echo "" >&2
@@ -939,6 +1027,7 @@ if [[ -z "$VARIANT" ]]; then
echo "(Use --variant <name> next time to skip the wizard.)" >&2
choose_model
choose_gpus
print_topology_advisory
pick_parallelism
if [[ "$MODEL_NAME" == "gemma-4-31b" && "${#CARD_INDICES[@]}" -eq 1 && "$MIN_VRAM_GB" -lt 32 ]]; then
gemma_single_24gb_guidance
@@ -0,0 +1,22 @@
schema_version: 1
model: gemma-4-26b-a4b
# Low-anchor calibration: both rows use the same bf16 / 32K / seqs=256 / TP=2
# envelope and only vary the external MTP assistant. Add lower-concurrency and
# longer-context rows before tightening the MoE activation coefficient.
rows:
- compose: vllm/gemma-a4b-awq
vram_gb: 24
measured_peak_gb: 23.45
ctx_override: null
status: low-anchor-calibration
engine_pin: vllm-nightly-bf610c2f
genesis_pin: null
source: "BENCHMARKS.md#MoE models gemma-4-26b-a4b/dual/awq.yml @noonghunna 2026-05-15"
- compose: vllm/gemma-a4b-awq-mtp
vram_gb: 24
measured_peak_gb: 23.50
ctx_override: null
status: low-anchor-calibration
engine_pin: vllm-nightly-bf610c2f
genesis_pin: null
source: "BENCHMARKS.md#MoE models gemma-4-26b-a4b/dual/awq-mtp.yml @noonghunna 2026-05-15"
@@ -0,0 +1,22 @@
schema_version: 1
model: qwen3.6-35b-a3b
# Low-anchor calibration: both rows use the same fp8_e5m2 / 16K / seqs=1 /
# TP=2 envelope and only vary built-in MTP. Add max_ctx / max_num_seqs A/B
# rows before tightening the MoE activation coefficient.
rows:
- compose: vllm/qwen-a3b-preview
vram_gb: 24
measured_peak_gb: 21.94
ctx_override: null
status: low-anchor-calibration
engine_pin: vllm-nightly-bf610c2f
genesis_pin: null
source: "BENCHMARKS.md#MoE models qwen3.6-35b-a3b/dual/preview.yml @noonghunna 2026-05-15"
- compose: vllm/qwen-a3b-preview-mtp
vram_gb: 24
measured_peak_gb: 22.72
ctx_override: null
status: low-anchor-calibration
engine_pin: vllm-nightly-bf610c2f
genesis_pin: null
source: "BENCHMARKS.md#MoE models qwen3.6-35b-a3b/dual/preview-mtp.yml @noonghunna 2026-05-15"
+108 -1
View File
@@ -12,6 +12,7 @@ import os
import subprocess
import time
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import Any, Optional
@@ -26,7 +27,7 @@ from .compose_registry import COMPOSE_REGISTRY
SUPPORTED_SCHEMA_VERSIONS = {1}
PROFILE_ROOT = Path(__file__).resolve().parent
REPO_ROOT = Path(__file__).resolve().parents[3]
CONSTRAINT_IDS = [f"C{i}" for i in range(1, 16)]
CONSTRAINT_IDS = [f"C{i}" for i in range(1, 17)]
ESTATE_CONSTRAINT_IDS = [f"E{i}" for i in range(1, 5)]
@@ -42,6 +43,37 @@ class CrossReferenceError(ProfileError):
"""Raised when a profile references a missing profile id."""
class TopologyClass(str, Enum):
SINGLE_CARD = "single_card"
HOMOGENEOUS = "homogeneous"
VRAM_MATCHED_COMPUTE_MISMATCHED = "vram_matched_compute_mismatched"
VRAM_MISMATCHED = "vram_mismatched"
HETEROGENEOUS_MIXED = "heterogeneous_mixed"
TOPOLOGY_ADVISORY = {
TopologyClass.SINGLE_CARD: None,
TopologyClass.HOMOGENEOUS: None,
TopologyClass.VRAM_MATCHED_COMPUTE_MISMATCHED: (
"Compute mismatch detected (VRAM matched). TP=N works fine but the faster card "
"waits at every NCCL allreduce — effective throughput caps at slower card's speed "
"(~30% of faster card idle at allreduce). Full per-card VRAM capacity preserved. "
"Alternative: estate planner (--estate) to run different models per card at full speed."
),
TopologyClass.VRAM_MISMATCHED: (
"VRAM mismatch detected. TP=N would cap to smaller card's usable model size. "
"Recommended paths: (a) llama.cpp `--tensor-split` for weighted layer split, "
"(b) PP=N (manual flag flip — `--pipeline-parallel-size N` on a vllm/dual compose; "
"no shipping PP compose), (c) estate planner (--estate) to run different models per card."
),
TopologyClass.HETEROGENEOUS_MIXED: (
"Heterogeneous hardware detected (multiple VRAM and compute tiers). Manual selection "
"recommended. Consider the estate planner (--estate) to put different models on "
"different card subsets, or run a single model on the largest matched subset."
),
}
def _logger() -> logging.Logger:
logger = logging.getLogger("compat")
if not logger.handlers:
@@ -92,6 +124,30 @@ class HardwareProfile:
notes: Optional[str] = None
def classify_hardware_topology(hardware: list[HardwareProfile]) -> TopologyClass:
"""Classify selected GPUs for TP-vs-PP/estate advisory output."""
if not hardware:
raise ProfileError("classify_hardware_topology requires at least one HardwareProfile")
if len(hardware) == 1:
return TopologyClass.SINGLE_CARD
vrams = sorted(hw.vram_gb for hw in hardware)
sms = {hw.sm for hw in hardware}
vram_clusters = 1
for i in range(1, len(vrams)):
if vrams[i] - vrams[i - 1] > 1.0:
vram_clusters += 1
if vram_clusters == 1 and len(sms) == 1:
return TopologyClass.HOMOGENEOUS
if vram_clusters == 1 and len(sms) > 1:
return TopologyClass.VRAM_MATCHED_COMPUTE_MISMATCHED
if vram_clusters > 1:
return TopologyClass.VRAM_MISMATCHED
return TopologyClass.HETEROGENEOUS_MIXED
@dataclass(frozen=True)
class ModelProfile:
schema_version: int
@@ -123,6 +179,25 @@ class ModelProfile:
head_dim_sliding: Optional[int] = None
global_head_dim: Optional[int] = None
sliding_window: Optional[int] = None
# Asymmetric KV head counts for SWA-hybrid models where global layers
# have a different KV head count than sliding layers (e.g. Gemma 4
# 26B-A4B: 8 sliding, 2 global). Leave None for symmetric models.
num_global_kv_heads: Optional[int] = None
# MoE fields (None for dense models; set for MoE variants)
num_experts: Optional[int] = None
num_experts_per_tok: Optional[int] = None
moe_intermediate_size: Optional[int] = None
shared_expert_intermediate_size: Optional[int] = None
active_params_b: Optional[float] = None
# Optional architectural metadata
mtp_num_hidden_layers: Optional[int] = None
attn_output_gate: Optional[bool] = None
vision_capable: Optional[bool] = None
# C12 (KV projection via tools/kv-calc.py) only supports models whose
# architecture has been added to MODEL_SPECS in kv-calc.py. New MoE /
# hybrid models can set kv_calc_supported=false to skip C12 until
# kv-calc gains MoE-aware activation/KV formulas.
kv_calc_supported: bool = True
@dataclass(frozen=True)
@@ -203,6 +278,7 @@ class FitsResult:
world_size: Optional[int] = None
bottleneck_vram_gb: Optional[float] = None
homogeneous: Optional[bool] = None
topology_class: Optional[TopologyClass] = None
kv_projection: Optional[dict[str, Any]] = None
compose_name: Optional[str] = None
weights_variant: Optional[str] = None
@@ -306,6 +382,15 @@ def _model(data: dict[str, Any]) -> ModelProfile:
head_dim_sliding=data.get("head_dim_sliding"),
global_head_dim=data.get("global_head_dim"),
sliding_window=data.get("sliding_window"),
num_global_kv_heads=data.get("num_global_kv_heads"),
num_experts=data.get("num_experts"),
num_experts_per_tok=data.get("num_experts_per_tok"),
moe_intermediate_size=data.get("moe_intermediate_size"),
shared_expert_intermediate_size=data.get("shared_expert_intermediate_size"),
active_params_b=data.get("active_params_b"),
mtp_num_hidden_layers=data.get("mtp_num_hidden_layers"),
attn_output_gate=data.get("attn_output_gate"),
vision_capable=data.get("vision_capable"),
max_ctx_supported=int(data["max_ctx_supported"]),
attention_k_eq_v=bool(data["attention_k_eq_v"]),
weights=_dict(data.get("weights")),
@@ -313,6 +398,7 @@ def _model(data: dict[str, Any]) -> ModelProfile:
compatible_drafters=_tuple(data.get("compatible_drafters")),
valid_tp=tuple(int(x) for x in _tuple(data.get("valid_tp"))),
requires_genesis=bool(data.get("requires_genesis", False)),
kv_calc_supported=bool(data.get("kv_calc_supported", True)),
)
@@ -576,6 +662,14 @@ def _kv_calc_weights_variant(model: ModelProfile, variant: str) -> str:
if variant == "bf16":
return "bf16"
return "int4"
if model.family == "gemma4-swa-moe":
if variant == "awq_compressed_tensors":
return "awq"
return "int4"
if model.family == "qwen3-next-moe":
if variant == "gptq_int4":
return "gptq"
return "default"
return "default"
@@ -686,6 +780,7 @@ def fits(
effective_max_num_seqs = max_num_seqs if max_num_seqs is not None else int(workload.defaults.get("max_num_seqs", 1))
effective_weights = resolve_weights_variant(model, engine, weights_variant)
homogeneous = len({hw.id for hw in hardware}) <= 1
topology_class = classify_hardware_topology(hardware) if hardware else None
bottleneck = min((hw.vram_gb for hw in hardware), default=None)
effective_cudagraph = _cudagraph_mode(hardware)
@@ -793,6 +888,9 @@ def fits(
elif engine.type != "vllm":
skipped.append("C12")
notes.append("KV projection not available for non-vLLM engines")
elif not model.kv_calc_supported:
skipped.append("C12")
notes.append(f"KV projection skipped: {model.id} not yet wired into tools/kv-calc.py")
else:
kv_calc_invoked = True
if effective_mem_util is None or bottleneck is None:
@@ -830,6 +928,14 @@ def fits(
else:
ok("C12")
if topology_class is None:
skip("C16", "Topology advisory not run; no hardware profiles provided.")
else:
ok("C16")
advisory = TOPOLOGY_ADVISORY.get(topology_class)
if advisory:
notes.append(f"C16 topology={topology_class.value}: {advisory}")
diagnostics = {
"constraints_evaluated": list(CONSTRAINT_IDS),
"constraints_passed": passed,
@@ -850,6 +956,7 @@ def fits(
world_size=world_size,
bottleneck_vram_gb=bottleneck,
homogeneous=homogeneous,
topology_class=topology_class,
kv_projection=kv_projection,
weights_variant=effective_weights,
diagnostics=diagnostics,
+66 -11
View File
@@ -88,14 +88,14 @@ COMPOSE_REGISTRY = {
),
"vllm/tools-text": _entry(
model="qwen3.6-27b", weights_variant="autoround_int4", workload="tool-heavy",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=1, max_ctx=75000, max_num_seqs=1, mem_util=0.97,
compose_path="models/qwen3.6-27b/vllm/compose/single/tools-text.yml",
default_port=8020,
),
"vllm/minimal": _entry(
model="qwen3.6-27b", weights_variant="autoround_int4", workload="fast-chat",
engine="vllm-nightly-mtp", drafter=None, kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter=None, kv_format="fp8_e5m2",
tp=1, max_ctx=32768, max_num_seqs=1, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/single/minimal.yml",
default_port=8020,
@@ -104,7 +104,7 @@ COMPOSE_REGISTRY = {
# Qwen 3.6 27B, vLLM dual/multi-card.
"vllm/dual": _entry(
model="qwen3.6-27b", weights_variant="autoround_int4", workload="long-ctx-single",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=2, max_ctx=262144, max_num_seqs=2, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml",
default_port=8010, recommended_engine_features=["marlin_pad_sub_tile_n"],
@@ -133,7 +133,7 @@ COMPOSE_REGISTRY = {
),
"vllm/dual-bf16": _entry(
model="qwen3.6-27b", weights_variant="autoround_int4", workload="long-ctx-single",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="bf16",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="bf16",
tp=2, max_ctx=200000, max_num_seqs=1, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/dual/bf16.yml",
default_port=8012,
@@ -168,21 +168,21 @@ COMPOSE_REGISTRY = {
),
"vllm/dual-carnice-bf16mtp": _entry(
model="qwen3.6-27b", weights_variant="carnice_bf16mtp", workload="long-ctx-single",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=2, max_ctx=262144, max_num_seqs=2, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/dual/carnice-bf16mtp.yml",
default_port=8070,
),
"vllm/dual-qwopus-bf16mtp": _entry(
model="qwen3.6-27b", weights_variant="qwopus_bf16mtp", workload="long-ctx-single",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=2, max_ctx=262144, max_num_seqs=2, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/dual/qwopus-bf16mtp.yml",
default_port=8071,
),
"vllm/dual-nvlink": _entry(
model="qwen3.6-27b", weights_variant="autoround_int4", workload="long-ctx-single",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=2, max_ctx=262144, max_num_seqs=2, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/dual/nvlink.yml",
default_port=8014, requires_nvlink=True, recommended_engine_features=["marlin_pad_sub_tile_n"],
@@ -211,7 +211,7 @@ COMPOSE_REGISTRY = {
),
"vllm/dual4": _entry(
model="qwen3.6-27b", weights_variant="autoround_int4", workload="multi-stream-tenant",
engine="vllm-nightly-mtp", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=4, max_ctx=262144, max_num_seqs=4, mem_util=0.92,
compose_path="models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml",
default_port=8015,
@@ -243,14 +243,14 @@ COMPOSE_REGISTRY = {
# Gemma 4 31B, vLLM.
"vllm/gemma-mtp-tp1": _entry(
model="gemma-4-31b", weights_variant="autoround_int4", workload="fast-chat",
engine="vllm-nightly-mtp", drafter="gemma-it-assistant", kv_format="fp8_e4m3",
engine="vllm-nightly-clean", drafter="gemma-it-assistant", kv_format="fp8_e4m3",
tp=1, max_ctx=8192, max_num_seqs=256, mem_util=0.95,
compose_path="models/gemma-4-31b/vllm/compose/single/docker-compose.yml",
default_port=8031, required_sm=9.0,
),
"vllm/gemma-mtp": _entry(
model="gemma-4-31b", weights_variant="autoround_int4", workload="fast-chat",
engine="vllm-nightly-mtp", drafter="gemma-it-assistant", kv_format="bf16",
engine="vllm-nightly-clean", drafter="gemma-it-assistant", kv_format="bf16",
tp=2, max_ctx=32768, max_num_seqs=4, mem_util=0.92,
compose_path="models/gemma-4-31b/vllm/compose/dual/docker-compose.yml",
default_port=8030,
@@ -292,7 +292,7 @@ COMPOSE_REGISTRY = {
),
"vllm/gemma-bf16": _entry(
model="gemma-4-31b", weights_variant="autoround_int4", workload="long-ctx-single",
engine="vllm-nightly-mtp", drafter="gemma-it-assistant", kv_format="bf16",
engine="vllm-nightly-clean", drafter="gemma-it-assistant", kv_format="bf16",
tp=2, max_ctx=200000, max_num_seqs=1, mem_util=0.95,
compose_path="models/gemma-4-31b/vllm/compose/dual/bf16.yml",
default_port=8033,
@@ -304,5 +304,60 @@ COMPOSE_REGISTRY = {
compose_path="models/gemma-4-31b/vllm/compose/dual/awq.yml",
default_port=8033,
),
# v0.7.3 MoE onboarding — Gemma 4 26B-A4B + Qwen 3.6 35B-A3B.
# Both target the unconstrained-nightly engine (vllm-nightly-clean) which
# rides nightly-bf610c2f (2026-05-15, post-PR-#42521). Gemma is the
# shippable path; Qwen 35B-A3B is preview-only until Genesis v7.73.x
# re-anchors on a post-#42521 nightly.
"vllm/gemma-a4b-single": _entry(
model="gemma-4-26b-a4b", weights_variant="autoround_int4_mixed", workload="fast-chat",
engine="vllm-nightly-clean", drafter=None, kv_format="bf16",
tp=1, max_ctx=8192, max_num_seqs=256, mem_util=0.92,
compose_path="models/gemma-4-26b-a4b/vllm/compose/single/docker-compose.yml",
default_port=8040,
),
"vllm/gemma-a4b": _entry(
model="gemma-4-26b-a4b", weights_variant="autoround_int4_mixed", workload="fast-chat",
engine="vllm-nightly-clean", drafter=None, kv_format="bf16",
tp=2, max_ctx=32768, max_num_seqs=256, mem_util=0.92,
compose_path="models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml",
default_port=8041,
),
"vllm/gemma-a4b-awq": _entry(
model="gemma-4-26b-a4b", weights_variant="awq_compressed_tensors", workload="fast-chat",
engine="vllm-nightly-clean", drafter=None, kv_format="bf16",
tp=2, max_ctx=32768, max_num_seqs=256, mem_util=0.92,
compose_path="models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml",
default_port=8042,
),
"vllm/gemma-a4b-awq-mtp": _entry(
model="gemma-4-26b-a4b", weights_variant="awq_compressed_tensors", workload="fast-chat",
engine="vllm-nightly-clean", drafter="gemma-26b-it-assistant", kv_format="bf16",
tp=2, max_ctx=32768, max_num_seqs=256, mem_util=0.92,
compose_path="models/gemma-4-26b-a4b/vllm/compose/dual/awq-mtp.yml",
default_port=8043,
),
"vllm/qwen-a3b-preview-single": _entry(
model="qwen3.6-35b-a3b", weights_variant="autoround_int4", workload="fast-chat",
engine="vllm-nightly-clean", drafter=None, kv_format="fp8_e5m2",
tp=1, max_ctx=8192, max_num_seqs=1, mem_util=0.92,
compose_path="models/qwen3.6-35b-a3b/vllm/compose/single/preview.yml",
default_port=8050,
),
"vllm/qwen-a3b-preview": _entry(
model="qwen3.6-35b-a3b", weights_variant="autoround_int4", workload="fast-chat",
engine="vllm-nightly-clean", drafter=None, kv_format="fp8_e5m2",
tp=2, max_ctx=16384, max_num_seqs=1, mem_util=0.92,
compose_path="models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml",
default_port=8051,
),
"vllm/qwen-a3b-preview-mtp": _entry(
model="qwen3.6-35b-a3b", weights_variant="autoround_int4", workload="fast-chat",
engine="vllm-nightly-clean", drafter="qwen-mtp-builtin", kv_format="fp8_e5m2",
tp=2, max_ctx=16384, max_num_seqs=1, mem_util=0.92,
compose_path="models/qwen3.6-35b-a3b/vllm/compose/dual/preview-mtp.yml",
default_port=8052,
),
}
@@ -0,0 +1,14 @@
schema_version: 1
id: gemma-26b-it-assistant
display_name: Google Gemma 4 26B-A4B MTP assistant
spec_method: mtp_assistant
model_compat:
- gemma-4-26b-a4b
n_default: 4
n_max: 4
download:
hf_repo: google/gemma-4-26B-A4B-it-assistant
size_gb: 0.97
format: fp16
vram_footprint_gb: 0.97
status: production
@@ -1,9 +1,10 @@
schema_version: 1
id: qwen-mtp-builtin
display_name: Qwen 3.6 27B built-in MTP head
display_name: Qwen 3.6 built-in MTP head
spec_method: mtp
model_compat:
- qwen3.6-27b
- qwen3.6-35b-a3b
n_default: 3
n_max: 3
download: false
@@ -0,0 +1,42 @@
schema_version: 1
id: vllm-nightly-clean
display_name: vLLM nightly (no Genesis)
type: vllm
stability: nightly
install:
method: docker_image
spec: vllm/vllm-openai:nightly-bf610c2f56764e1b30bc6065f4ceace3d6e59036
min_sm: 7.5
supported_model_families:
- dense
- gemma4-swa-dense
- gemma4-swa-moe
- qwen3-next-hybrid
- qwen3-next-moe
features:
int8_per_token_head: false
turboquant_3bit_nc: false
marlin_pad_sub_tile_n: false
qwen3_coder_tool_parser: false
supported_kv_formats:
- bf16
- fp16
- fp8_e5m2
- fp8_e4m3
- q4_0
- k8v4
supported_drafters:
- eagle3
- mtp
- mtp_assistant
supported_weight_formats:
- bf16
- fp16
- autoround
- awq
- gptq
- compressed-tensors
required_overlays: []
vendored_overlays: []
required_genesis: false
notes: "Unconstrained-nightly path: free to bump to whatever's current on Docker Hub, since no Genesis patches are anchored here. Per the TQ3-only Genesis policy (Genesis is strictly required only for turboquant_3bit_nc KV), this engine accepts every supported model family — non-TQ3 composes for Qwen 3.6-27B (qwen3-next-hybrid) and Gemma 4-31B (gemma4-swa-dense) route here, alongside the new MoE additions (qwen3-next-moe, gemma4-swa-moe). For Qwen3-Next workloads at long context (>21-26K), Genesis-anchored vllm-nightly-mtp is *recommended* for Cliff 2 mitigations but not strictly required to boot. Excludes turboquant_3bit_nc KV (Genesis-only feature) — TQ3 composes must route to vllm-nightly-mtp. Launch exports VLLM_NIGHTLY_SHA from this spec; VLLM_IMAGE override still works."
@@ -5,12 +5,14 @@ type: vllm
stability: nightly
install:
method: docker_image
spec: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
spec: vllm/vllm-openai:nightly-01d4d1ad375dc5854779c593eee093bcebb0cada
min_sm: 7.5
supported_model_families:
- dense
- qwen3-next-hybrid
- qwen3-next-moe
- gemma4-swa-dense
- gemma4-swa-moe
features:
int8_per_token_head: false
turboquant_3bit_nc: true
@@ -40,4 +42,4 @@ required_overlays: []
vendored_overlays: []
required_genesis: true
genesis_pin: v7.72.2
notes: "Production nightly path for Qwen MTP and Gemma MTP without DFlash or INT8 PTH overlays. Launch exports this SHA as VLLM_NIGHTLY_SHA; VLLM_IMAGE can override the full image ref."
notes: "Production nightly path for Qwen MTP. Pin held at nightly-01d4d1ad (2026-05-04, Sander's v7.72.2 PROD pin) — this is the SHA Genesis v7.72.2 patches were anchored against and all v7.72.2 BENCHMARKS rows reference. Bumping requires either a Genesis pin bump cycle (Sander v7.73.x) or local patch re-anchor + Qwen 27B/Gemma 31B rebench gate. supported_model_families lists qwen3-next-moe + gemma4-swa-moe so the schema is ready; the actual MoE production boot needs a Genesis re-anchor on a post-#42521 nightly. Until then qwen3-next-moe routes to the preview path on vllm-nightly-clean. Gemma 4 composes that don't need Genesis should route to vllm-nightly-clean directly. Launch exports this SHA as VLLM_NIGHTLY_SHA; VLLM_IMAGE can override the full image ref. PRIOR-BUG: commit 40f1ef78 (2026-05-14) accidentally set this to nightly-1acd67a7 (the post-Gemma4-merge Gemma SHA); fixed back to 01d4d1ad."
+145 -2
View File
@@ -8,6 +8,7 @@ existing compose registry and validate_estate() profile checks.
from __future__ import annotations
import argparse
import concurrent.futures
import os
import re
import shlex
@@ -43,6 +44,8 @@ from scripts.lib.profiles.launch_compat import _hardware_id_from_gpu, resolve_en
SUPPORTED_ESTATE_SCHEMA_VERSIONS = {1}
DEFAULT_ESTATE_PATH = Path("~/.club3090/estate.yml").expanduser()
DEFAULT_BOOT_LOG_DIR = Path("/tmp/club3090-estate-boot")
BOOT_LOG_KEEP = 5
class EstateCliError(Exception):
@@ -58,6 +61,17 @@ class GpuInfo:
hardware_id: str
@dataclass(frozen=True)
class BootOutcome:
index: int
total: int
inst: InstanceSpec
ok: bool
elapsed_s: int
log_path: Path
error: str = ""
def utc_now() -> str:
return datetime.now(timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z")
@@ -353,11 +367,54 @@ def compose_cmd() -> list[str]:
return shlex.split(os.environ.get("COMPOSE_BIN", "docker compose"))
def run_compose(inst: InstanceSpec, action: str) -> None:
def boot_log_dir() -> Path:
return Path(os.environ.get("CLUB3090_ESTATE_BOOT_LOG_DIR", str(DEFAULT_BOOT_LOG_DIR))).expanduser()
def instance_log_path(inst: InstanceSpec) -> Path:
return boot_log_dir() / f"{safe_name(inst.name)}.log"
def rotate_log(path: Path, keep: int = BOOT_LOG_KEEP) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
for i in range(keep - 1, 0, -1):
src = path.with_name(f"{path.name}.{i}")
dst = path.with_name(f"{path.name}.{i + 1}")
if src.exists():
src.replace(dst)
if path.exists():
path.replace(path.with_name(f"{path.name}.1"))
def prepare_instance_log(inst: InstanceSpec) -> Path:
path = instance_log_path(inst)
rotate_log(path)
path.touch(mode=0o600, exist_ok=True)
return path
def summarize_error(error: str, limit: int = 220) -> str:
line = " ".join(part.strip() for part in str(error).splitlines() if part.strip())
if not line:
line = "unknown error"
return line if len(line) <= limit else line[: limit - 1] + "…"
def append_log(path: Path, message: str) -> None:
with path.open("a", encoding="utf-8") as fh:
fh.write(message.rstrip() + "\n")
def run_compose(inst: InstanceSpec, action: str, log_path: Path | None = None) -> None:
cmd = compose_cmd() + ["-p", project_name(inst.name), "-f", str(compose_abs_path(inst.compose_name)), action]
if action == "up":
cmd.append("-d")
proc = subprocess.run(cmd, cwd=REPO_ROOT, env=compose_env(inst), text=True)
if log_path is not None:
append_log(log_path, f"$ {' '.join(cmd)}")
with log_path.open("a", encoding="utf-8") as fh:
proc = subprocess.run(cmd, cwd=REPO_ROOT, env=compose_env(inst), text=True, stdout=fh, stderr=subprocess.STDOUT)
else:
proc = subprocess.run(cmd, cwd=REPO_ROOT, env=compose_env(inst), text=True)
if proc.returncode != 0:
raise EstateCliError(f"`{' '.join(cmd)}` failed with exit {proc.returncode}")
@@ -406,6 +463,21 @@ def wait_ready(inst: InstanceSpec, timeout: int) -> None:
time.sleep(4)
def wait_ready_quiet(inst: InstanceSpec, timeout: int) -> int:
start = time.monotonic()
poll_interval = max(float(os.environ.get("CLUB3090_ESTATE_POLL_INTERVAL", "4")), 0.1)
while True:
if endpoint_ready(inst.port):
return int(time.monotonic() - start)
if not container_running(inst.name):
logs = docker_logs_tail(inst.name)
raise EstateCliError(f"container {container_name(inst.name)} stopped during boot\n{logs}")
elapsed = time.monotonic() - start
if elapsed >= timeout:
raise EstateCliError(f"timeout waiting for {inst.name} after {timeout}s; logs: docker logs {container_name(inst.name)}")
time.sleep(min(poll_interval, max(timeout - elapsed, 0.1)))
def select_instances(instances: list[InstanceSpec], only: set[str] | None) -> list[InstanceSpec]:
if not only:
return instances
@@ -430,6 +502,67 @@ def command_validate(args: argparse.Namespace) -> int:
return 0 if result.valid else 1
def effective_parallel_jobs(requested: int | None, total: int) -> int:
if requested is None:
return min(total, 4)
if requested < 1:
raise EstateCliError("--parallel-jobs must be >= 1")
return min(requested, total, 4)
def boot_instance_parallel(index: int, total: int, inst: InstanceSpec, timeout: int, log_path: Path) -> BootOutcome:
start = time.monotonic()
try:
append_log(log_path, f"[estate] booting {inst.name}: {inst.compose_name} GPUs={list(inst.gpu_indices)} port={inst.port}")
run_compose(inst, "up", log_path=log_path)
elapsed_s = wait_ready_quiet(inst, timeout)
append_log(log_path, f"[estate] healthy after {elapsed_s}s")
return BootOutcome(index=index, total=total, inst=inst, ok=True, elapsed_s=elapsed_s, log_path=log_path)
except Exception as exc:
elapsed_s = int(time.monotonic() - start)
summary = summarize_error(str(exc))
append_log(log_path, f"[estate] ERROR after {elapsed_s}s: {summary}")
return BootOutcome(index=index, total=total, inst=inst, ok=False, elapsed_s=elapsed_s, log_path=log_path, error=summary)
def boot_instances_parallel(selected: list[InstanceSpec], timeout: int, requested_jobs: int | None, stagger_s: float) -> int:
if stagger_s < 0:
raise EstateCliError("--parallel-stagger must be >= 0")
total = len(selected)
jobs = effective_parallel_jobs(requested_jobs, total)
print(f"[estate] parallel boot: {total} instance(s), jobs={jobs}, stagger={stagger_s:g}s")
futures: dict[concurrent.futures.Future[BootOutcome], tuple[int, InstanceSpec, Path]] = {}
with concurrent.futures.ThreadPoolExecutor(max_workers=jobs) as pool:
for i, inst in enumerate(selected, start=1):
log_path = prepare_instance_log(inst)
print(f"[estate] [{i}/{total}] booting {inst.name} on GPUs {','.join(str(g) for g in inst.gpu_indices)} port {inst.port}... (started)")
future = pool.submit(boot_instance_parallel, i, total, inst, timeout, log_path)
futures[future] = (i, inst, log_path)
if i < total and stagger_s > 0:
time.sleep(stagger_s)
outcomes = [future.result() for future in concurrent.futures.as_completed(futures)]
outcomes.sort(key=lambda outcome: outcome.index)
healthy = 0
failed: list[BootOutcome] = []
for outcome in outcomes:
if outcome.ok:
healthy += 1
print(f"[estate] [{outcome.index}/{outcome.total}] {outcome.inst.name} ✓ healthy after {outcome.elapsed_s}s")
else:
failed.append(outcome)
print(
f"[estate] [{outcome.index}/{outcome.total}] {outcome.inst.name} ✗ failed after {outcome.elapsed_s}s: {outcome.error}"
)
print(f"[estate] Summary: {healthy}/{total} healthy, {len(failed)} failed.")
for outcome in failed:
print(f"[estate] Failed instance: {outcome.inst.name}. See {outcome.log_path}")
return 0 if not failed else 1
def command_boot(args: argparse.Namespace) -> int:
path = estate_path(args.file)
try:
@@ -439,6 +572,13 @@ def command_boot(args: argparse.Namespace) -> int:
return 1
selected = select_instances(instances, parse_only(args.only))
persist_default_estate_source(path, data, instances, gpus, nvlink_active)
if getattr(args, "parallel", False) and len(selected) > 1:
return boot_instances_parallel(
selected,
args.timeout,
getattr(args, "parallel_jobs", None),
getattr(args, "parallel_stagger", 15.0),
)
total = len(selected)
for i, inst in enumerate(selected, start=1):
print(f"[estate] [{i}/{total}] booting {inst.name}: {inst.compose_name} GPUs={list(inst.gpu_indices)} port={inst.port}")
@@ -699,6 +839,9 @@ def build_parser() -> argparse.ArgumentParser:
boot.add_argument("--file", default=str(DEFAULT_ESTATE_PATH))
boot.add_argument("--only", default="")
boot.add_argument("--timeout", type=int, default=int(os.environ.get("READY_TIMEOUT", "600")))
boot.add_argument("--parallel", action="store_true")
boot.add_argument("--parallel-jobs", type=int, default=None)
boot.add_argument("--parallel-stagger", type=float, default=15.0)
boot.set_defaults(func=command_boot)
down = sub.add_parser("down")
+119 -1
View File
@@ -19,7 +19,16 @@ if str(REPO_ROOT) not in sys.path:
os.environ.setdefault("CLUB3090_LOG_LEVEL", "ERROR")
from scripts.lib.profiles.compat import FitsResult, ProfileError, fits, load_profiles, to_compose_name # noqa: E402
from scripts.lib.profiles.compat import ( # noqa: E402
TOPOLOGY_ADVISORY,
FitsResult,
ProfileError,
TopologyClass,
classify_hardware_topology,
fits,
load_profiles,
to_compose_name,
)
from scripts.lib.profiles.compose_registry import COMPOSE_REGISTRY # noqa: E402
@@ -103,6 +112,26 @@ def _parse_gpu_specs(value: str, profiles) -> list:
return hardware
def _parse_gpu_specs_with_indices(value: str, profiles) -> list[tuple[str, object]]:
hardware = []
for raw in value.split(";"):
raw = raw.strip()
if not raw:
continue
try:
idx, name, mem_mib, sm = raw.split("|", 3)
except ValueError as exc:
raise LaunchCompatError(f"invalid --gpu-spec entry `{raw}`") from exc
hardware_id = _hardware_id_from_gpu(name, int(mem_mib), float(sm))
try:
hardware.append((idx, profiles.hardware[hardware_id]))
except KeyError as exc:
raise LaunchCompatError(f"hardware profile `{hardware_id}` is not installed") from exc
if not hardware:
raise LaunchCompatError("no GPU specs were provided for topology classification")
return hardware
def _engine_family(engine_type: str) -> str:
return "llamacpp" if engine_type == "llama.cpp" else engine_type
@@ -364,6 +393,90 @@ def command_resolve_variant_pin(args: argparse.Namespace) -> int:
return 0
def _hardware_line(index: str, hardware) -> str:
return f" GPU {index}: {hardware.display_name} ({hardware.vram_gb:g} GB, sm {hardware.sm:g})"
def _standalone_recommendation(topology: TopologyClass, count: int) -> list[str]:
if topology == TopologyClass.SINGLE_CARD:
return [
"Recommended:",
" 1. Use the largest single-card compose your model fits.",
" 2. Add another matched card for TP=2 when long-context concurrency matters.",
]
if topology == TopologyClass.HOMOGENEOUS:
return [
"Recommended:",
f" 1. TP={count} is the default path for matched cards; use the shipped vllm/dual* or multi-card composes.",
" 2. Estate planner remains useful when you want separate models/endpoints instead of one larger TP instance.",
]
if topology == TopologyClass.VRAM_MATCHED_COMPUTE_MISMATCHED:
return [
"Recommended:",
f" 1. TP={count} works as-is. Compute mismatch means the faster card waits at every NCCL allreduce; effective throughput caps at the slower card's speed (~30% of faster card idle). Full per-card VRAM capacity preserved.",
" 2. Estate planner — `bash scripts/launch.sh --estate` runs different models per card, each at full speed.",
"",
"Not recommended:",
" - PP=N: possible as a manual flag flip (`--pipeline-parallel-size N`) on a vllm/dual compose, but no PP compose ships today.",
]
if topology == TopologyClass.VRAM_MISMATCHED:
return [
"Recommended:",
" 1. llama.cpp `--tensor-split` for weighted layer split on mismatched VRAM.",
" 2. PP=N as a manual vLLM flag flip (`--pipeline-parallel-size N`) if you are deliberately experimenting.",
" 3. Estate planner — run different models per card or use the largest matched subset.",
"",
"Not recommended:",
" - TP=N on the full mismatched set: the smaller card caps usable model size and KV headroom.",
]
return [
"Recommended:",
" 1. Manual selection. Use the largest matched subset for one model.",
" 2. Estate planner — put different models on different card subsets.",
]
def command_topology(args: argparse.Namespace) -> int:
_quiet_compat_logger()
profiles = load_profiles()
indexed_hardware = _parse_gpu_specs_with_indices(args.gpu_spec, profiles)
hardware = [item[1] for item in indexed_hardware]
topology = classify_hardware_topology(hardware)
advisory = TOPOLOGY_ADVISORY.get(topology)
if args.format == "wizard":
if topology in (TopologyClass.SINGLE_CARD, TopologyClass.HOMOGENEOUS):
return 0
detected = " + ".join(
f"1x {hw.display_name} ({hw.vram_gb:g} GB, sm {hw.sm:g})"
for _idx, hw in indexed_hardware
)
print(f"Detected: {detected}")
print("")
print(f"Topology: {topology.value}")
if advisory:
print(f" {advisory}")
print("")
print("Continue with the selected parallelism if that trade-off is acceptable.")
return 0
print("Detected hardware:")
for idx, hw in indexed_hardware:
print(_hardware_line(idx, hw))
print("")
print(f"Topology class: {topology.value}")
print("")
for line in _standalone_recommendation(topology, len(hardware)):
print(line)
print("")
if advisory:
print("Advisory:")
print(f" {advisory}")
print("")
print("For details, see docs/MULTI_CARD.md.")
return 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description="Profile bridge for scripts/launch.sh")
sub = parser.add_subparsers(dest="command", required=True)
@@ -404,6 +517,11 @@ def build_parser() -> argparse.ArgumentParser:
variant_pin.add_argument("--format", choices=("shell", "json", "value"), default="shell")
variant_pin.set_defaults(func=command_resolve_variant_pin)
topology = sub.add_parser("topology")
topology.add_argument("--gpu-spec", required=True)
topology.add_argument("--format", choices=("standalone", "wizard"), default="standalone")
topology.set_defaults(func=command_topology)
return parser
@@ -0,0 +1,54 @@
schema_version: 1
id: gemma-4-26b-a4b
display_name: Gemma 4 26B-A4B (MoE)
family: gemma4-swa-moe
hidden_size: 2816
intermediate_size: 2112
num_hidden_layers: 30
# Hybrid SWA pattern: 5 full_attention layers at indices [5, 11, 17, 23, 29]
# (every 6th layer, with last layer always global per Gemma 4 family convention)
num_full_attn_layers: 5
num_sliding_attn_layers: 25
num_attn_heads: 16
num_kv_heads: 8 # sliding-layer KV head count
num_global_kv_heads: 2 # NEW: global-layer KV head count (asymmetric — distinct from sliding)
head_dim_sliding: 256
global_head_dim: 512
sliding_window: 1024
max_ctx_supported: 262144
attention_k_eq_v: true # Gemma 4 family default — K and V share storage
# MoE (from config.json text_config)
num_experts: 128
num_experts_per_tok: 8 # config field: top_k_experts
moe_intermediate_size: 704
active_params_b: 4.0
mtp_num_hidden_layers: null # MTP drafter is external (gemma-4-26B-A4B-it-assistant), not built-in
vision_capable: true # multimodal — vision_config + audio_config tokens present
weights:
autoround_int4_mixed:
path: gemma-4-26b-a4b-autoround-int4-mixed
size_gb: 16.0 # measured post-download 2026-05-15
format: autoround
status: ampere-blocked # moe_intermediate_size=704 % group_size=128 = 5.5
# → Marlin K-dim alignment fails on SM86.
# Boots fine on SM90+ (Cutlass W4A8 / Machete).
awq_compressed_tensors:
path: gemma-4-26b-a4b-awq-4bit
size_gb: 17.0 # measured post-download 2026-05-15
format: compressed-tensors # AWQ pack-quantized via vLLM compressed-tensors loader
status: production # via vLLM PR #40886 overlay (Ampere-bootable)
default_weight_variant: awq_compressed_tensors
compatible_drafters:
- gemma-26b-it-assistant # dedicated 26B-A4B MTP assistant (separate weights from 31B)
# Sliding layers cap TP at num_kv_heads=8 → [1,2,4,8]
# Global layers cap TP at num_global_kv_heads=2 → [1,2]
# Effective valid_tp = intersection = [1, 2]
valid_tp:
- 1
- 2
requires_genesis: false # Gemma 4 family doesn't need Genesis (no DeltaNet quirks)
# kv-calc.py models MoE + asymmetric KV heads (8 sliding / 2 global) as of
# v0.7.3. Calibration is low-anchor (AWQ no-MTP + AWQ-MTP rows only); add
# longer context / lower max_num_seqs rows before treating projections as
# production-grade.
kv_calc_supported: true
+8 -1
View File
@@ -51,5 +51,12 @@ valid_tp:
- 1
- 2
- 4
requires_genesis: true
# Strictly bootable on any vLLM nightly that supports the qwen3-next-hybrid
# family. Genesis is REQUIRED for TQ3 KV (turboquant_3bit_nc — Genesis-only
# feature). For non-TQ3 KV formats (fp8, bf16, q4_0, k8v4) Genesis is
# recommended but not strictly needed: it adds Cliff 2 mitigations
# (PN12/PN25/PN34) for long-context stability past ~21-26K. Composes encode
# the engine choice via their `Engine-profile:` header — non-TQ3 composes
# route to vllm-nightly-clean, TQ3 composes route to vllm-nightly-mtp.
requires_genesis: false
@@ -0,0 +1,73 @@
schema_version: 1
id: qwen3.6-35b-a3b
display_name: Qwen 3.6 35B-A3B (MoE)
family: qwen3-next-moe
hidden_size: 2048
num_hidden_layers: 40
# Hybrid: 10 full_attention layers at indices 3,7,11,15,19,23,27,31,35,39
# (one per 4-layer block, matching full_attention_interval=4 in config.json)
num_gdn_layers: 30
num_attn_layers: 10
num_attn_heads: 16
num_kv_heads: 2
head_dim_attn: 256
linear_num_v_heads: 32
linear_num_k_heads: 16
linear_v_head_dim: 128
linear_k_head_dim: 128
linear_conv_kernel_dim: 4
max_ctx_supported: 262144
attention_k_eq_v: false
# MoE (from config.json text_config)
num_experts: 256
num_experts_per_tok: 8
moe_intermediate_size: 512
shared_expert_intermediate_size: 512
active_params_b: 3.0
# Built-in MTP drafter (1 dedicated MTP head)
mtp_num_hidden_layers: 1
attn_output_gate: true
vision_capable: true
weights:
autoround_int4:
path: qwen3.6-35b-a3b-autoround-int4
size_gb: 20.0
format: autoround
status: production
gptq_int4:
path: qwen3.6-35b-a3b-gptq-int4
size_gb: 22.0
format: gptq
status: experimental
gguf:
path: qwen3.6-35b-a3b-gguf
size_gb: 90.0
format: gguf
status: production
dflash:
path: qwen3.6-35b-a3b-dflash
size_gb: variable
format: autoround
status: experimental
dflash_gguf:
path: qwen3.6-35b-a3b-dflash-gguf
size_gb: variable
format: gguf
status: experimental
default_weight_variant: autoround_int4
compatible_drafters:
- qwen-mtp-builtin
# num_kv_heads=2 caps TP at 2 (each rank needs at least 1 KV head)
valid_tp:
- 1
- 2
# Strictly bootable on upstream vLLM (PR #42521 onwards). Genesis is
# RECOMMENDED for production: it provides TQ3 KV (Genesis-only) plus the
# Cliff 2 / DeltaNet stabilization patches that matter past ~21-26K ctx.
# Preview/smoke composes target vllm-nightly-clean (no Genesis); production
# composes (TBD — gated on Genesis v7.73.x re-anchor) target vllm-nightly-mtp.
requires_genesis: false
# kv-calc.py models the MoE + DeltaNet+attention hybrid as of v0.7.3.
# Calibration is low-anchor (preview + preview-MTP rows only); add longer
# context / max_num_seqs rows before treating projections as production-grade.
kv_calc_supported: true
+34
View File
@@ -384,6 +384,40 @@ if [[ -x scripts/lib/profiles/estate_cli.py || -f scripts/lib/profiles/estate_cl
python3 scripts/lib/profiles/estate_cli.py report-state 2>&1 | redact || true
fi
# ---------------------------------------------------------------------------
# KV math calibration
# ---------------------------------------------------------------------------
# When a user files a VRAM-OOM or context-ceiling bug, the maintainer's first
# question is "does kv-calc still agree with measured reality?" — a calibration
# failure means the projection model has drifted from the actual VRAM cost of a
# compose, so any "predicted PASS" verdict can't be trusted. Surface the
# verdict line + any FAIL rows here so a triage reply can immediately see
# whether to trust kv-calc projections for this user's config.
if have python3 && [[ -f tools/kv-calc.py ]]; then
section "KV math calibration"
calib_output=$(python3 tools/kv-calc.py --calibration 2>&1 || true)
overall=$(echo "$calib_output" | grep -E '^Overall:' | head -1)
fail_rows=$(echo "$calib_output" | grep -E '\bFAIL\b' || true)
{
if [[ -n "$overall" ]]; then
echo "- ${overall}"
else
echo "- _kv-calc --calibration produced no Overall line; see output below._"
fi
if [[ -n "$fail_rows" ]]; then
echo "- ⚠ Failing rows:"
echo '```'
echo "$fail_rows"
echo '```'
echo "- Math model is mis-calibrated against measured reality for the rows above. Any kv-calc projection on this checkout should be treated as suspect until the calibration anchors / formulas are reconciled."
else
echo "- No FAIL rows. kv-calc projections should agree with measured VRAM within the ±1.5 GB error band."
fi
} | redact
echo "$calib_output" | redact | details "Full kv-calc --calibration output"
fi
# ---------------------------------------------------------------------------
# Active container
# ---------------------------------------------------------------------------
+5
View File
@@ -828,6 +828,11 @@ def cmd_summary(turn_log, summary_path, boot_vram, growth_limit, timed_out, expe
print(f"[soak] {label}:")
for item in items:
print(f"[soak] - {item}")
if verdict == "PASS":
print("[soak] note PASS = no failure signal on this sample;")
print("[soak] not patch validation (topology alone can")
print("[soak] sidestep what overlays target). See")
print("[soak] scripts/soak-test.sh --help and docs/CLIFFS.md.")
sys.exit(exit_code)
+35 -3
View File
@@ -12,6 +12,27 @@
# retention across sessions.
# - Read-only against the running deployment.
#
# PASS verdict semantics:
# PASS = no failure signal fired on the test sample. Specifically:
# - silent_empty turns: 0 (no HTTP 200 + 0 completion tokens)
# - max VRAM growth: under SOAK_MAX_GROWTH_MIB (default 200 MiB)
# - TPS retention: first-5 vs last-5 median >= 98%
# - request errors / stream interruptions: 0
# PASS does NOT mean:
# - "Patches in this compose's overlay set are doing useful work."
# PASS-on-patched is consistent with patches working OR with patches
# not being load-bearing for this workload + topology. Cliff 2 / 2b
# mitigations target single-card 24 GB pressure; TP=2 (dual.yml)
# structurally escapes Cliff 2 regardless of which patches load.
# - "Deeper-context workloads will also pass." Continuous mode ramps
# to ~22-25K accumulated tokens by turn 5; it does not push to
# model max_ctx. Longer-context regimes can still fail.
# - "The configuration is optimally tuned." Soak detects failures,
# not whether perf is on the table.
# For patch attribution, run the same soak on the same compose with
# the overlay bind-mounts stripped (or on a baseline image) and compare
# metrics. See https://github.com/noonghunna/club-3090/issues/140.
#
# Time budget:
# Default SOAK_SESSIONS=20 x SOAK_TURNS=5, capped by SOAK_TIMEOUT_S=1800.
# Expect 10-30 minutes depending on config.
@@ -101,9 +122,20 @@ EXAMPLES
CONTAINER=none ENDPOINT=http://localhost:8030 bash scripts/soak-test.sh
NOTES
Soak-continuous is the only test that catches Cliff 2b. If you're filing a
bench contribution, run with --continuous and paste the [soak] summary
alongside your bench numbers. See docs/CLIFFS.md for context.
Soak-continuous is the only test that surfaces Cliff 2b under
multi-turn accumulating-context traffic on single-card configs.
If you're filing a bench contribution, run with --continuous and
paste the [soak] summary alongside your bench numbers.
See docs/CLIFFS.md for context.
PASS VERDICT — WHAT IT DOES AND DOES NOT MEAN
PASS = no failure signal on the test sample (silent_empty=0, VRAM
growth under threshold, TPS retention >= 98%, zero errors).
PASS does NOT validate that patches in the compose's overlay set
are load-bearing for the workload — topology alone (e.g. TP=2)
can sidestep the failure mode patches target. For patch attribution,
re-run the same soak with overlays stripped and compare.
Full discussion: docs/CLIFFS.md and issue #140.
EOF
}
+13 -6
View File
@@ -4,7 +4,8 @@ set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
HELPER="${ROOT_DIR}/scripts/lib/profiles/launch_compat.py"
GPU_3090='0|RTX_3090|24576|8.6'
MTP_SHA="1acd67a795ebccdf9b9db7697ae9082058301657"
MTP_SHA="01d4d1ad375dc5854779c593eee093bcebb0cada"
CLEAN_SHA="bf610c2f56764e1b30bc6065f4ceace3d6e59036"
DFLASH_SHA="e47c98ef7a38792996e452ef53914e21e41928e9"
assert_contains() {
@@ -85,16 +86,19 @@ fi
assert_contains "$out" "install.spec is not a docker nightly image"
out="$(python3 "$HELPER" resolve-variant-pin --variant vllm/dual --format shell)"
assert_contains "$out" "VLLM_NIGHTLY_SHA=${CLEAN_SHA}"
out="$(python3 "$HELPER" resolve-variant-pin --variant vllm/dual-tq3-mtp --format shell)"
assert_contains "$out" "VLLM_NIGHTLY_SHA=${MTP_SHA}"
out="$(python3 "$HELPER" resolve-variant-pin --variant vllm/gemma-dflash --format shell)"
assert_contains "$out" "VLLM_NIGHTLY_SHA=${DFLASH_SHA}"
if command -v docker >/dev/null 2>&1 && docker compose version >/dev/null 2>&1; then
out="$(VLLM_NIGHTLY_SHA="$MTP_SHA" docker compose -f "$ROOT_DIR/models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml" config 2>/dev/null)"
assert_contains "$out" "image: vllm/vllm-openai:nightly-${MTP_SHA}"
out="$(VLLM_NIGHTLY_SHA="$CLEAN_SHA" docker compose -f "$ROOT_DIR/models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml" config 2>/dev/null)"
assert_contains "$out" "image: vllm/vllm-openai:nightly-${CLEAN_SHA}"
out="$(VLLM_NIGHTLY_SHA="$MTP_SHA" VLLM_IMAGE=ghcr.io/noonghunna/vllm-club3090:latest docker compose -f "$ROOT_DIR/models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml" config 2>/dev/null)"
out="$(VLLM_NIGHTLY_SHA="$CLEAN_SHA" VLLM_IMAGE=ghcr.io/noonghunna/vllm-club3090:latest docker compose -f "$ROOT_DIR/models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml" config 2>/dev/null)"
assert_contains "$out" "image: ghcr.io/noonghunna/vllm-club3090:latest"
fi
@@ -102,12 +106,15 @@ out="$(python3 - <<'PY'
from scripts.lib.profiles.compat import InstanceSpec
from scripts.lib.profiles.estate_cli import compose_env
mtp = compose_env(InstanceSpec(name="qwen", compose_name="vllm/dual", gpu_indices=(0, 1), port=8010))
clean = compose_env(InstanceSpec(name="qwen", compose_name="vllm/dual", gpu_indices=(0, 1), port=8010))
tq3 = compose_env(InstanceSpec(name="qwen-tq3", compose_name="vllm/dual-tq3-mtp", gpu_indices=(0, 1), port=8010))
dflash = compose_env(InstanceSpec(name="gemma", compose_name="vllm/gemma-dflash", gpu_indices=(0, 1), port=8032))
print(mtp["VLLM_NIGHTLY_SHA"])
print(clean["VLLM_NIGHTLY_SHA"])
print(tq3["VLLM_NIGHTLY_SHA"])
print(dflash["VLLM_NIGHTLY_SHA"])
PY
)"
assert_contains "$out" "$CLEAN_SHA"
assert_contains "$out" "$MTP_SHA"
assert_contains "$out" "$DFLASH_SHA"
+240
View File
@@ -0,0 +1,240 @@
#!/usr/bin/env bash
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
TMP_DIR="$(mktemp -d)"
trap 'rm -rf "$TMP_DIR"' EXIT
assert_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" != *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
FAKE_GPUS_4X='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6'
ESTATE4="${TMP_DIR}/estate4.yml"
cat > "$ESTATE4" <<'YAML'
schema_version: 1
created: 2026-05-15T00:00:00Z
rig:
hardware_id: rtx-3090
gpu_count: 4
nvlink_active: false
estate:
- name: one
compose: llamacpp/default
gpus: [0]
port: 8110
- name: two
compose: llamacpp/default
gpus: [1]
port: 8120
- name: three
compose: llamacpp/default
gpus: [2]
port: 8130
- name: four
compose: llamacpp/default
gpus: [3]
port: 8140
YAML
FAKE_ESTATE_HELPER="${TMP_DIR}/estate-helper.py"
cat > "$FAKE_ESTATE_HELPER" <<'PY'
#!/usr/bin/env python3
import sys
print("ARGS " + " ".join(sys.argv[1:]))
PY
chmod +x "$FAKE_ESTATE_HELPER"
out="$(ESTATE_HELPER="$FAKE_ESTATE_HELPER" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight \
--estate-file "$ESTATE4" \
--only one \
--parallel \
--parallel-jobs 3 \
--parallel-stagger 0 2>&1)"
assert_contains "$out" "ARGS boot --file $ESTATE4 --only one --parallel --parallel-jobs 3 --parallel-stagger 0"
out="$(
cd "$ROOT_DIR"
HOME="${TMP_DIR}/home-success" \
CLUB3090_FAKE_GPUS="$FAKE_GPUS_4X" \
CLUB3090_ESTATE_BOOT_LOG_DIR="${TMP_DIR}/logs-success" \
python3 - "$ESTATE4" <<'PY'
import argparse
import sys
import threading
import time
from scripts.lib.profiles import estate_cli as ec
state = {"active": 0, "max_active": 0}
ready = []
lock = threading.Lock()
def fake_run_compose(inst, action, log_path=None):
if log_path is not None:
ec.append_log(log_path, f"compose {action} {inst.name}")
with lock:
state["active"] += 1
state["max_active"] = max(state["max_active"], state["active"])
time.sleep(0.05)
with lock:
state["active"] -= 1
def fake_wait_ready_quiet(inst, timeout):
ready.append(inst.name)
return 1
ec.run_compose = fake_run_compose
ec.wait_ready_quiet = fake_wait_ready_quiet
rc = ec.command_boot(
argparse.Namespace(
file=sys.argv[1],
only="",
timeout=5,
parallel=True,
parallel_jobs=2,
parallel_stagger=0.0,
)
)
print(f"rc={rc}")
print(f"max_active={state['max_active']}")
print("ready=" + ",".join(sorted(ready)))
for name in ("one", "two", "three", "four"):
print(f"log:{name}={ec.instance_log_path(ec.InstanceSpec(name=name, compose_name='llamacpp/default', gpu_indices=(0,), port=8000)).exists()}")
print(f"hard_cap={ec.effective_parallel_jobs(9, 8)}")
PY
)"
assert_contains "$out" "[estate] parallel boot: 4 instance(s), jobs=2, stagger=0s"
assert_contains "$out" "[estate] Summary: 4/4 healthy, 0 failed."
assert_contains "$out" "rc=0"
assert_contains "$out" "max_active=2"
assert_contains "$out" "ready=four,one,three,two"
assert_contains "$out" "log:one=True"
assert_contains "$out" "log:four=True"
assert_contains "$out" "hard_cap=4"
ESTATE2="${TMP_DIR}/estate2.yml"
cat > "$ESTATE2" <<'YAML'
schema_version: 1
created: 2026-05-15T00:00:00Z
rig:
hardware_id: rtx-3090
gpu_count: 2
nvlink_active: false
estate:
- name: good
compose: llamacpp/default
gpus: [0]
port: 8110
- name: bad
compose: llamacpp/default
gpus: [1]
port: 8120
YAML
out="$(
cd "$ROOT_DIR"
HOME="${TMP_DIR}/home-fail" \
CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6' \
CLUB3090_ESTATE_BOOT_LOG_DIR="${TMP_DIR}/logs-fail" \
python3 - "$ESTATE2" <<'PY'
import argparse
import sys
from scripts.lib.profiles import estate_cli as ec
def fake_run_compose(inst, action, log_path=None):
if log_path is not None:
ec.append_log(log_path, f"compose {action} {inst.name}")
if inst.name == "bad":
raise ec.EstateCliError("mock compose failed")
def fake_wait_ready_quiet(inst, timeout):
return 2
ec.run_compose = fake_run_compose
ec.wait_ready_quiet = fake_wait_ready_quiet
rc = ec.command_boot(
argparse.Namespace(
file=sys.argv[1],
only="",
timeout=5,
parallel=True,
parallel_jobs=2,
parallel_stagger=0.0,
)
)
print(f"rc={rc}")
PY
)"
assert_contains "$out" "[estate] Summary: 1/2 healthy, 1 failed."
assert_contains "$out" "bad ✗ failed"
assert_contains "$out" "Failed instance: bad. See"
assert_contains "$out" "rc=1"
out="$(
cd "$ROOT_DIR"
HOME="${TMP_DIR}/home-single" \
CLUB3090_FAKE_GPUS="$FAKE_GPUS_4X" \
CLUB3090_ESTATE_BOOT_LOG_DIR="${TMP_DIR}/logs-single" \
python3 - "$ESTATE4" <<'PY'
import argparse
import sys
from scripts.lib.profiles import estate_cli as ec
events = []
def fake_run_compose(inst, action):
events.append(f"{action}:{inst.name}")
def fake_wait_ready(inst, timeout):
events.append(f"ready:{inst.name}")
def fake_wait_ready_quiet(inst, timeout):
raise AssertionError("single-instance --parallel should use sequential boot")
ec.run_compose = fake_run_compose
ec.wait_ready = fake_wait_ready
ec.wait_ready_quiet = fake_wait_ready_quiet
rc = ec.command_boot(
argparse.Namespace(
file=sys.argv[1],
only="one",
timeout=5,
parallel=True,
parallel_jobs=2,
parallel_stagger=0.0,
)
)
print(f"rc={rc}")
print("events=" + ",".join(events))
PY
)"
assert_contains "$out" "[estate] all selected instances are healthy"
assert_contains "$out" "rc=0"
assert_contains "$out" "events=up:one,ready:one"
echo "test-parallel-boot: ok"
+89 -8
View File
@@ -24,11 +24,11 @@ run_test "load_profiles parses all profile groups" <<'PY'
from scripts.lib.profiles.compat import load_profiles
p = load_profiles()
assert len(p.hardware) == 9
assert len(p.models) == 2
assert len(p.models) == 4
assert len(p.workloads) == 5
assert len(p.engines) == 6
assert len(p.drafters) == 5
assert len(p.calibration) == 2
assert len(p.engines) == 7
assert len(p.drafters) == 6
assert len(p.calibration) == 4
PY
run_test "fits() happy path: Qwen dual on 2x3090" <<'PY'
@@ -49,6 +49,77 @@ assert r.recommended_kv_format == "turboquant_3bit_nc"
assert r.diagnostics["constraints_skipped"] == ["C12"]
PY
run_test "topology: single card classified" <<'PY'
from scripts.lib.profiles.compat import load_profiles, classify_hardware_topology, TopologyClass
p = load_profiles()
r = classify_hardware_topology([p.hardware["rtx-3090"]])
assert r == TopologyClass.SINGLE_CARD
PY
run_test "topology: 2x3090 classified homogeneous" <<'PY'
from scripts.lib.profiles.compat import load_profiles, classify_hardware_topology, TopologyClass
p = load_profiles()
r = classify_hardware_topology([p.hardware["rtx-3090"], p.hardware["rtx-3090"]])
assert r == TopologyClass.HOMOGENEOUS
PY
run_test "topology: 3090+4090 classified compute-mismatched" <<'PY'
from scripts.lib.profiles.compat import load_profiles, classify_hardware_topology, TopologyClass
p = load_profiles()
r = classify_hardware_topology([p.hardware["rtx-3090"], p.hardware["rtx-4090"]])
assert r == TopologyClass.VRAM_MATCHED_COMPUTE_MISMATCHED
PY
run_test "topology: 3090+3060 classified VRAM-mismatched" <<'PY'
from scripts.lib.profiles.compat import load_profiles, classify_hardware_topology, TopologyClass
p = load_profiles()
r = classify_hardware_topology([p.hardware["rtx-3090"], p.hardware["rtx-3060-12gb"]])
assert r == TopologyClass.VRAM_MISMATCHED
PY
run_test "topology: VRAM cluster wins over mixed compute" <<'PY'
from scripts.lib.profiles.compat import load_profiles, classify_hardware_topology, TopologyClass
p = load_profiles()
r = classify_hardware_topology([p.hardware["rtx-3090"], p.hardware["rtx-3060-12gb"], p.hardware["rtx-4090"]])
assert r == TopologyClass.VRAM_MISMATCHED
PY
run_test "C16 topology advisory emits note for compute mismatch" <<'PY'
from scripts.lib.profiles.compat import load_profiles, fits, TopologyClass
p = load_profiles()
r = fits(
hardware=[p.hardware["rtx-3090"], p.hardware["rtx-4090"]],
model=p.models["qwen3.6-27b"],
workload=p.workloads["long-ctx-single"],
engine=p.engines["vllm-nightly-mtp"],
drafter=p.drafters["qwen-mtp-builtin"],
tp=2,
pp=1,
project_vram=False,
)
assert r.topology_class == TopologyClass.VRAM_MATCHED_COMPUTE_MISMATCHED
assert "C16" in r.diagnostics["constraints_passed"]
assert any("C16" in n and "vram_matched_compute_mismatched" in n for n in r.notes), r.notes
PY
run_test "C16 topology advisory is silent for homogeneous GPUs" <<'PY'
from scripts.lib.profiles.compat import load_profiles, fits, TopologyClass
p = load_profiles()
r = fits(
hardware=[p.hardware["rtx-3090"], p.hardware["rtx-3090"]],
model=p.models["qwen3.6-27b"],
workload=p.workloads["long-ctx-single"],
engine=p.engines["vllm-nightly-mtp"],
drafter=p.drafters["qwen-mtp-builtin"],
tp=2,
pp=1,
project_vram=False,
)
assert r.topology_class == TopologyClass.HOMOGENEOUS
assert "C16" in r.diagnostics["constraints_passed"]
assert not any("C16" in n for n in r.notes), r.notes
PY
run_test "C1 card count: world size mismatch rejected" <<'PY'
from scripts.lib.profiles.compat import load_profiles, fits
p = load_profiles()
@@ -92,12 +163,22 @@ assert not r.valid
assert any(reason.startswith("C5:") for reason in r.reasons), r.reasons
PY
run_test "C6 Genesis one-way implication: Qwen on non-Genesis vLLM rejected" <<'PY'
run_test "C6 Genesis one-way: TQ3 KV on non-Genesis engine rejected (via C15)" <<'PY'
# Under the TQ3-only Genesis policy, no model declares requires_genesis=true
# (so C6 has no current model-level trigger). Genesis is enforced at the
# *feature* level via C15: requesting turboquant_3bit_nc on an engine that
# doesn't expose it (e.g. vllm-stable-next) fails C15. The previous
# C6 assertion (Qwen on non-Genesis vLLM rejected) no longer holds —
# Qwen 27B with fp8 is valid on non-Genesis engines.
from scripts.lib.profiles.compat import load_profiles, fits
p = load_profiles()
# Positive: Qwen 27B + fp8 on non-Genesis engine is now valid.
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["vllm-stable-next"], kv_format="fp8_e5m2", tp=1, project_vram=False)
assert r.valid, r.reasons
# Negative: Qwen 27B + TQ3 on non-Genesis engine fails C15.
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["vllm-stable-next"], kv_format="turboquant_3bit_nc", tp=1, project_vram=False, required_engine_features=["turboquant_3bit_nc"])
assert not r.valid
assert any(reason.startswith("C6:") for reason in r.reasons), r.reasons
assert any(reason.startswith("C15:") for reason in r.reasons), r.reasons
PY
run_test "C7 drafter method: DFlash on MTP-only engine rejected" <<'PY'
@@ -218,7 +299,7 @@ from scripts.lib.profiles.compat import load_profiles, to_compose_name
p = load_profiles()
name = to_compose_name(
p.models["qwen3.6-27b"],
p.engines["vllm-nightly-mtp"],
p.engines["vllm-nightly-clean"],
p.drafters["qwen-mtp-builtin"],
"fp8_e5m2",
2,
@@ -236,7 +317,7 @@ from scripts.lib.profiles.compat import load_profiles, fits
p = load_profiles()
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["vllm-nightly-mtp"], tp=1, project_vram=False)
d = r.diagnostics
assert d["constraints_evaluated"] == [f"C{i}" for i in range(1, 16)]
assert d["constraints_evaluated"] == [f"C{i}" for i in range(1, 17)]
assert "constraints_passed" in d and "constraints_failed" in d and "constraints_skipped" in d
assert isinstance(d["elapsed_ms"], float)
PY
+19
View File
@@ -143,6 +143,7 @@ out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1 --no-projection 2>&1)"
assert_contains "$out" "[launch] Tensor parallel TP=2"
assert_not_contains "$out" "Topology:"
assert_contains "$out" "SWITCHED vllm/dual CUDA=0,1 NVD=0,1 TP=2 PP=1"
selected_count="$(grep -c "\[launch\] selected variant:" <<< "$out" || true)"
if [[ "$selected_count" != "1" ]]; then
@@ -179,6 +180,24 @@ if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6
fi
assert_contains "$out" "Gemma 4 31B does not fit on a single 24 GB card today"
out="$(CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6' \
bash "${ROOT_DIR}/scripts/launch.sh" --topology 2>&1)"
assert_contains "$out" "Topology class: homogeneous"
assert_not_contains "$out" "Compute mismatch detected"
out="$(CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_4090:24576:8.9' \
bash "${ROOT_DIR}/scripts/launch.sh" --topology 2>&1)"
assert_contains "$out" "Topology class: vram_matched_compute_mismatched"
assert_contains "$out" "Compute mismatch detected"
assert_contains "$out" "Estate planner"
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_4090:24576:8.9' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1 --no-projection 2>&1)"
assert_contains "$out" "Topology: vram_matched_compute_mismatched"
assert_contains "$out" "Compute mismatch detected"
assert_contains "$out" "SWITCHED vllm/dual CUDA=0,1 NVD=0,1 TP=2 PP=1"
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6,4:RTX_3090:24576:8.6,5:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1,2,3,4,5 --tp 6 --no-projection 2>&1)"; then
+134 -9
View File
@@ -15,9 +15,11 @@ Predicts (per card, after TP split):
- Total vs available VRAM
- Verdict: PASS / TIGHT / FAIL
Two models modelled:
Four models modelled:
- Qwen 3.6 27B (DeltaNet hybrid: 16 full_attention + 48 GDN)
- Qwen 3.6 35B-A3B (MoE + DeltaNet hybrid: 10 attention + 30 GDN)
- Gemma 4 31B (SWA + dense MLP: 10 full_attention + 50 sliding_attention)
- Gemma 4 26B-A4B (MoE + SWA: 5 full_attention + 25 sliding_attention)
vLLM rate-limits KV pool to fit available budget; this predictor models that
capping behavior. When the requested KV pool exceeds what fits, the verdict
@@ -39,7 +41,7 @@ Usage:
bash tools/kv-calc.py --compose dual-turbo --vram 24 # Qwen (default model)
bash tools/kv-calc.py --model gemma-4-31b --compose gemma-dual-int8 --vram 24
bash tools/kv-calc.py --model gemma-4-31b --solve-max-ctx --kv-format int8_per_token_head --tp 2 --vram 24
bash tools/kv-calc.py --calibration # both models, grouped per-model
bash tools/kv-calc.py --calibration # all calibrated models, grouped per-model
"""
import argparse
@@ -90,16 +92,28 @@ def _weight_size(model, variant):
def _load_model_specs_from_yaml(profiles):
qwen, gemma = profiles.models["qwen3.6-27b"], profiles.models["gemma-4-31b"]
qwen_moe, gemma_moe = profiles.models["qwen3.6-35b-a3b"], profiles.models["gemma-4-26b-a4b"]
q_fields = ("hidden_size", "num_hidden_layers", "num_gdn_layers", "num_attn_layers", "num_attn_heads", "num_kv_heads", "head_dim_attn", "linear_num_v_heads", "linear_num_k_heads", "linear_v_head_dim", "linear_k_head_dim", "linear_conv_kernel_dim", "max_ctx_supported", "attention_k_eq_v")
g_fields = ("hidden_size", "intermediate_size", "num_hidden_layers", "num_full_attn_layers", "num_sliding_attn_layers", "num_attn_heads", "num_kv_heads", "head_dim_sliding", "global_head_dim", "sliding_window", "max_ctx_supported", "attention_k_eq_v")
qspec = {"model_id": qwen.id, "model_family": qwen.family, **{k: getattr(qwen, k) for k in q_fields}, "valid_tp": list(qwen.valid_tp), "weights_total_gb": _weight_size(qwen, qwen.default_weight_variant), "mamba_state_bytes": 4, "chunk_size": 256}
gspec = {"model_id": gemma.id, "model_family": gemma.family, **{k: getattr(gemma, k) for k in g_fields}, "valid_tp": list(gemma.valid_tp), "weights_int4_gb": _weight_size(gemma, "autoround_int4"), "weights_awq_gb": _weight_size(gemma, "awq"), "weights_bf16_gb": _weight_size(gemma, "bf16"), "drafter_mtp_gb": float(profiles.drafters["gemma-it-assistant"].vram_footprint_gb), "drafter_dflash_gb": float(profiles.drafters["gemma-dflash"].vram_footprint_gb)}
return {"qwen3.6-27b": qspec, "gemma-4-31b": gspec}
gm_fields = (*g_fields, "num_global_kv_heads", "num_experts", "num_experts_per_tok", "moe_intermediate_size", "active_params_b", "mtp_num_hidden_layers")
qm_fields = (*q_fields, "num_experts", "num_experts_per_tok", "moe_intermediate_size", "shared_expert_intermediate_size", "active_params_b", "mtp_num_hidden_layers")
qspec = {"model_id": qwen.id, "model_family": qwen.family, **{k: getattr(qwen, k) for k in q_fields}, "valid_tp": list(qwen.valid_tp), "weights_total_gb": _weight_size(qwen, qwen.default_weight_variant), "mamba_state_bytes": 4, "chunk_size": 256, "mtp_n_default": profiles.drafters["qwen-mtp-builtin"].n_default}
qmspec = {"model_id": qwen_moe.id, "model_family": qwen_moe.family, **{k: getattr(qwen_moe, k) for k in qm_fields}, "valid_tp": list(qwen_moe.valid_tp), "weights_total_gb": _weight_size(qwen_moe, qwen_moe.default_weight_variant), "weights_gptq_gb": _weight_size(qwen_moe, "gptq_int4"), "mamba_state_bytes": 4, "chunk_size": 256, "mtp_n_default": profiles.drafters["qwen-mtp-builtin"].n_default}
gspec = {"model_id": gemma.id, "model_family": gemma.family, **{k: getattr(gemma, k) for k in g_fields}, "valid_tp": list(gemma.valid_tp), "weights_int4_gb": _weight_size(gemma, "autoround_int4"), "weights_awq_gb": _weight_size(gemma, "awq"), "weights_bf16_gb": _weight_size(gemma, "bf16"), "drafter_mtp_gb": float(profiles.drafters["gemma-it-assistant"].vram_footprint_gb), "drafter_dflash_gb": float(profiles.drafters["gemma-dflash"].vram_footprint_gb), "mtp_n_default": profiles.drafters["gemma-it-assistant"].n_default}
gmspec = {"model_id": gemma_moe.id, "model_family": gemma_moe.family, **{k: getattr(gemma_moe, k) for k in gm_fields}, "valid_tp": list(gemma_moe.valid_tp), "weights_int4_gb": _weight_size(gemma_moe, "autoround_int4_mixed"), "weights_awq_gb": _weight_size(gemma_moe, "awq_compressed_tensors"), "drafter_mtp_gb": float(profiles.drafters["gemma-26b-it-assistant"].vram_footprint_gb), "mtp_n_default": profiles.drafters["gemma-26b-it-assistant"].n_default}
return {
"qwen3.6-27b": qspec,
"qwen3.6-35b-a3b": qmspec,
"gemma-4-31b": gspec,
"gemma-4-26b-a4b": gmspec,
}
MODEL_SPECS = _load_model_specs_from_yaml(PROFILES)
QWEN36_27B = MODEL_SPECS["qwen3.6-27b"]
QWEN36_35B_A3B = MODEL_SPECS["qwen3.6-35b-a3b"]
GEMMA4_31B = MODEL_SPECS["gemma-4-31b"]
GEMMA4_26B_A4B = MODEL_SPECS["gemma-4-26b-a4b"]
# =============================================================================
@@ -141,6 +155,25 @@ QWEN_GDN_ACTIVATION_COEF = {
"turboquant_3bit_nc": 165,
}
# ---- Qwen MoE activation + built-in MTP workspace ----
# Path-B low-anchor fit from the two v0.7.3 preview rows. The per-token GDN
# coefficient follows the dense-Qwen shape; the small constant captures MoE
# expert dispatch/router buffers. Current vLLM TP preview effectively keeps
# the quantized MoE weights resident per card, so weights are not divided by TP
# for qwen3-next-moe in _weights_per_card_gb().
QWEN_MOE_ACTIVATION_COEF = {
"fp16": 110,
"bf16": 110,
"fp8_e5m2": 105,
"fp8_e4m3": 105,
"int8_per_token_head": 105,
"q4_0": 130,
"k8v4": 130,
"turboquant_3bit_nc": 140,
}
QWEN_MOE_EXPERT_DISPATCH_GB = 0.20
QWEN_MOE_BUILTIN_MTP_WORKSPACE_GB = 0.10
# ---- Gemma activation peak (mostly constant in ctx) ----
# Unlike Qwen GDN, Gemma's activation peak comes from dense MLP forward +
# SWA windowed-attention prefill, both bounded by chunked-prefill chunk_size.
@@ -149,13 +182,22 @@ QWEN_GDN_ACTIVATION_COEF = {
GEMMA_ACTIVATION_CONST_GB = 1.5 # per card at TP=1 — calibrated, ~scales as 1/TP
GEMMA_ACTIVATION_PER_TOKEN_BYTES = 8 # tiny ctx scaling term to keep solver well-behaved
# ---- Gemma MoE activation peak ----
# Low-anchor fit from awq.yml and awq-mtp.yml. MoE dispatch is folded into the
# constant term; external assistant weights are modelled separately via
# drafter_gb.
GEMMA_MOE_ACTIVATION_CONST_GB = 1.8
GEMMA_MOE_ACTIVATION_PER_TOKEN_BYTES = 12
# =============================================================================
# Compose presets (per-model)
# =============================================================================
COMPOSE_ALIAS_TEXT = {
"qwen3.6-27b": "minimal=vllm/minimal long-text=vllm/long-text long-text-no-mtp=vllm/long-text-no-mtp long-vision=vllm/long-vision bounded-thinking=vllm/bounded-thinking tools-text=vllm/tools-text dual=vllm/dual dual-turbo=vllm/dual-turbo dual-dflash=vllm/dual-dflash dual-dflash-noviz=vllm/dual-dflash-noviz dual4=vllm/dual4 dual4-dflash=vllm/dual4-dflash",
"qwen3.6-35b-a3b": "qwen-a3b-preview-single=vllm/qwen-a3b-preview-single qwen-a3b-preview=vllm/qwen-a3b-preview qwen-a3b-preview-mtp=vllm/qwen-a3b-preview-mtp",
"gemma-4-31b": "gemma-dual=vllm/gemma-mtp gemma-dual-int8=vllm/gemma-int8 gemma-dual-int8-262k=vllm/gemma-int8-262k gemma-dual-bf16=vllm/gemma-bf16 gemma-dual-int8-tq3=vllm/gemma-int8-tq3 gemma-dual-dflash=vllm/gemma-dflash gemma-dual-dflash-int8=vllm/gemma-dflash-int8 gemma-dual-awq=vllm/gemma-awq gemma-single=vllm/gemma-mtp-tp1",
"gemma-4-26b-a4b": "gemma-a4b-single=vllm/gemma-a4b-single gemma-a4b=vllm/gemma-a4b gemma-a4b-awq=vllm/gemma-a4b-awq gemma-a4b-awq-mtp=vllm/gemma-a4b-awq-mtp",
}
COMPOSE_ALIASES = {model: tuple(part.split("=", 1) for part in text.split()) for model, text in COMPOSE_ALIAS_TEXT.items()}
@@ -182,10 +224,18 @@ def _compose_cfg_from_registry(profiles, model_id, legacy_name, registry_name):
cfg["mtp"] = drafter is not None and drafter.spec_method in ("mtp", "mtp_assistant")
if drafter is not None and drafter.spec_method == "dflash":
cfg.update({"mtp": False, "dflash_draft_gb": float(drafter.vram_footprint_gb)})
if drafter is not None:
cfg["mtp_n"] = int(drafter.n_default)
if model_id == "gemma-4-31b" and drafter is not None:
cfg["drafter_gb"] = float(drafter.vram_footprint_gb)
if model_id == "gemma-4-31b":
cfg["weights_variant"] = {"awq": "awq", "bf16": "bf16"}.get(entry["weights_variant"], "int4")
if model_id == "gemma-4-26b-a4b":
cfg["weights_variant"] = "awq" if entry["weights_variant"] == "awq_compressed_tensors" else "int4"
if drafter is not None:
cfg["drafter_gb"] = float(drafter.vram_footprint_gb)
if model_id == "qwen3.6-35b-a3b":
cfg["weights_variant"] = "gptq" if entry["weights_variant"] == "gptq_int4" else "default"
cfg.update(COMPOSE_COMPAT_OVERRIDES.get((model_id, legacy_name), {}))
return cfg
@@ -235,6 +285,12 @@ def _weights_per_card_gb(spec, tp, weights_variant="default"):
"""Return per-card weights footprint in GB after TP split."""
if spec["model_family"] == "qwen3-next-hybrid":
return spec["weights_total_gb"] / tp
elif spec["model_family"] == "qwen3-next-moe":
# Current vLLM MoE preview keeps expert weights effectively resident
# per TP rank; live 16K rows calibrate to full quant weight per card.
if weights_variant == "gptq":
return spec["weights_gptq_gb"]
return spec["weights_total_gb"]
elif spec["model_family"] == "gemma4-swa-dense":
if weights_variant == "awq":
return spec["weights_awq_gb"] / tp
@@ -242,6 +298,12 @@ def _weights_per_card_gb(spec, tp, weights_variant="default"):
return spec["weights_bf16_gb"] / tp
else: # int4 default
return spec["weights_int4_gb"] / tp
elif spec["model_family"] == "gemma4-swa-moe":
# Same MoE-residency assumption as Qwen A3B; validated by the
# v0.7.3 AWQ rows where TP=2 still peaks near a full AWQ shard/card.
if weights_variant == "int4":
return spec["weights_int4_gb"]
return spec["weights_awq_gb"]
raise ValueError(f"Unknown model_family: {spec['model_family']}")
@@ -278,6 +340,30 @@ def kv_pool_per_card_bytes(spec, kv_format, max_ctx, max_num_seqs, tp, mtp_n=0):
growing = (per_token / tp) * effective_ctx * max_num_seqs
return growing, 0.0
elif spec["model_family"] == "qwen3-next-moe":
# K and V stored independently. GDN recurrent state is fixed-size and
# per-stream, not context-linear.
per_token = (
spec["num_attn_layers"]
* spec["num_kv_heads"]
* spec["head_dim_attn"]
* 2
* bpe
)
effective_ctx = max_ctx + mtp_n * 32
growing = (per_token / tp) * effective_ctx * max_num_seqs
recurrent_per_stream = (
spec["num_gdn_layers"]
* (
spec["linear_num_v_heads"] * spec["linear_v_head_dim"]
+ spec["linear_num_k_heads"] * spec["linear_k_head_dim"]
+ spec["linear_conv_kernel_dim"] * spec["hidden_size"]
)
* 2 # recurrent state kept in bf16 on this stack
)
recurrent_fixed = recurrent_per_stream * max_num_seqs
return growing, recurrent_fixed
elif spec["model_family"] == "gemma4-swa-dense":
# K==V tied → ×1 storage
per_token_growing = (
@@ -302,6 +388,28 @@ def kv_pool_per_card_bytes(spec, kv_format, max_ctx, max_num_seqs, tp, mtp_n=0):
sliding_per_card = sliding_fixed_total / tp
return growing, sliding_per_card
elif spec["model_family"] == "gemma4-swa-moe":
# K==V tied. Global layers use their own KV-head count; sliding
# layers keep the windowed KV head count.
per_token_growing = (
spec["num_full_attn_layers"]
* spec["num_global_kv_heads"]
* spec["global_head_dim"]
* 1
* bpe
)
growing = (per_token_growing / tp) * max_ctx * max_num_seqs
sliding_fixed_total = (
spec["num_sliding_attn_layers"]
* spec["num_kv_heads"]
* spec["head_dim_sliding"]
* 1
* bpe
* spec["sliding_window"]
)
sliding_per_card = sliding_fixed_total / tp
return growing, sliding_per_card
raise ValueError(f"Unknown model_family: {spec['model_family']}")
@@ -320,11 +428,20 @@ def activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp):
coef = QWEN_GDN_ACTIVATION_COEF[kv_format]
return (coef * spec["num_gdn_layers"] * max_ctx) / tp
elif spec["model_family"] == "qwen3-next-moe":
coef = QWEN_MOE_ACTIVATION_COEF[kv_format]
return (coef * spec["num_gdn_layers"] * max_ctx) / tp + QWEN_MOE_EXPERT_DISPATCH_GB * 1e9
elif spec["model_family"] == "gemma4-swa-dense":
const_bytes = GEMMA_ACTIVATION_CONST_GB * 1e9
per_token = GEMMA_ACTIVATION_PER_TOKEN_BYTES * max_ctx
return (const_bytes + per_token) / tp
elif spec["model_family"] == "gemma4-swa-moe":
const_bytes = GEMMA_MOE_ACTIVATION_CONST_GB * 1e9
per_token = GEMMA_MOE_ACTIVATION_PER_TOKEN_BYTES * max_ctx
return (const_bytes + per_token) / tp
raise ValueError(f"Unknown model_family: {spec['model_family']}")
@@ -375,9 +492,10 @@ def predict(
weights_gb = _weights_per_card_gb(spec, tp, weights_variant)
mtp_n = int(spec.get("mtp_n_default", 3)) if mtp else 0
growing_b, sliding_b = kv_pool_per_card_bytes(
spec, kv_format, max_ctx, max_num_seqs, tp,
mtp_n=3 if mtp else 0,
mtp_n=mtp_n,
)
kv_pool_requested_gb = growing_b / 1e9
kv_pool_sliding_fixed_gb = sliding_b / 1e9
@@ -387,6 +505,8 @@ def predict(
# Drafter: prefer drafter_gb; fall back to legacy dflash_draft_gb.
drafter_total = drafter_gb if drafter_gb > 0 else dflash_draft_gb
if mtp and spec["model_family"] == "qwen3-next-moe":
drafter_total += QWEN_MOE_BUILTIN_MTP_WORKSPACE_GB
drafter_per_card = drafter_total / tp if tp > 1 else drafter_total
fixed_gb = weights_gb + activation_gb + overhead_gb + drafter_per_card + kv_pool_sliding_fixed_gb
@@ -407,7 +527,10 @@ def predict(
# concurrency reduced (BOOT OK, but `--max-num-seqs` may not be
# honored at full max_ctx).
# - PASS: requested KV fits with room to spare.
MIN_KV_GB = 1.0 # vLLM needs at least ~1 GB for paged-attention blocks
# The Qwen A3B preview is KV-light enough that the live 16K rows boot with
# <0.1 GB requested growing KV. Keep the older 1 GB guard for dense/long-KV
# models, but avoid false FAILs on this MoE family.
MIN_KV_GB = 0.05 if spec["model_family"] == "qwen3-next-moe" else 1.0
if available_for_kv < MIN_KV_GB:
verdict = "FAIL"
notes.append(
@@ -434,6 +557,8 @@ def predict(
notes.append("⚠ fp8_e4m3 on Ampere (sm_86): Triton `fp8e4nv` kernel unsupported; use int8_per_token_head instead (PR #40391 via #42102)")
if spec["model_family"] == "gemma4-swa-dense" and tp == 1 and vram_gb < 32:
notes.append("⚠ Gemma 4 31B TP=1 needs ≥32 GB VRAM; 24 GB Ampere boot-OOMs (model weights + drafter + min KV)")
if spec["model_family"] in {"qwen3-next-moe", "gemma4-swa-moe"}:
notes.append("MoE projection uses low-anchor calibration; add max_ctx/max_num_seqs A/B rows before treating this as production-grade.")
if tp > 4:
notes.append("TP > 4 predictions are extrapolated; report deltas via scripts/report.sh --bench")
@@ -554,7 +679,7 @@ def run_calibration():
print()
total_c, total_n = 0, 0
for model_key in ("qwen3.6-27b", "gemma-4-31b"):
for model_key in MODEL_SPECS:
c, n = _calibration_block(model_key)
total_c += c
total_n += n
@@ -642,7 +767,7 @@ def main():
help="(deprecated alias for --drafter-gb)")
p.add_argument("--weights-variant", choices=["default", "int4", "awq", "bf16"], default=None,
help="Gemma 4 only: which weight quant variant. Default: from --compose, or int4.")
p.add_argument("--calibration", action="store_true", help="Print predicted vs measured for both models.")
p.add_argument("--calibration", action="store_true", help="Print predicted vs measured for all calibrated models.")
p.add_argument("--solve-max-ctx", action="store_true", help="Binary-search for the largest max_ctx that fits.")
p.add_argument("--json", action="store_true", help="Output prediction as JSON.")
args = p.parse_args()