Follow-up on c3e7c7e: pinning to a specific build (b9246) was cargo-culted from
our vLLM pattern, but the vLLM pinning serves a real purpose (Genesis-patch
anchor, Docker Hub purge resistance) that doesn't apply here. llama.cpp on the
club-3090 stack is stock upstream — no patches, no Genesis equivalent — and
GHCR tag retention is more reliable than Docker Hub.
Switch to rolling `:server-cuda` tag so users automatically get MTP improvements,
EAGLE3 fixes, kernel updates from upstream without us being a bottleneck.
Override path preserved via `LLAMACPP_IMAGE` env if a future upstream build
regresses and a user needs to pin reactively.
Bench numbers in #170 footnoted as "measured on b9246"; expect ±5% drift on
newer builds.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
v0.8.3 shipped composes (llamacpp/default, llamacpp/mtp, llamacpp/mtp-vision) all
reference `image: llama-cpp:local`, a custom image that exists ONLY on the
maintainer's rig. There is no Dockerfile, no build script, and no setup.sh hook
to produce it for users. Anyone running `bash scripts/switch.sh llamacpp/mtp`
on a fresh clone hits "image llama-cpp:local not found" and dies at boot.
The custom image was a v0.8.3-dev artifact from when MTP PR #22673 was bleeding
edge. The official upstream `ghcr.io/ggml-org/llama.cpp:server-cuda` now has it
merged (build b9246 = commit 871b0b70f, 2026-05-20) — pinning to b9246 reproduces
the v0.8.3 numbers (50.25 narr / 58.04 code on single 3090, vs shipped 51.28/59.72).
Surfaced by @zemaphore in discussion #170. README.md was also lying: claimed
"both use the official ghcr.io image, no custom build needed" while composes
referenced llama-cpp:local.
Override the pin via `LLAMACPP_IMAGE=ghcr.io/.../server-cuda-bXXXX` env if you
want to follow upstream master.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
_render_recommendation keyed solely on res.ok, so a confirm→proceed /
override-accepted terminal (raw_verdict=fits-clean, res.ok=False because
the run needs an explicit --yes/--force-download) fell into the generic
"DOES NOT FIT / BLOCKED" branch — dishonest by imprecision (the model
fits; only acceptance is pending; CONTRACT-4 is "honest recommendation —
fits?"). Add a presentation-only needs-acceptance classification: a
fits-clean acceptance terminal now renders "FITS (estimated) — NOT YET
ACCEPTED" with the acceptance-gate guidance (not the failure on-ramp);
genuine hard-blocks still render DOES NOT FIT / BLOCKED. Pure
presentation — still derived only from res, no decision logic. Caught by
the V5 on-rig gate (microsoft/phi-2 --dry-run). test-pull.sh rec(2a)
updated to assert the honest rendering; suite 25/25, kv-calc N/N.
CONTRACT-4: add `--recommend` — an honest aggregated recommendation that
is PURE presentation/aggregation over the SHIPPED run_pull verdict. Every
line is read straight off the real PullResult (ok/confidence/raw_verdict/
terminal/stratum/abort_reason/notices/emitted); it introduces no decision
logic and does not change the exit code. Carries the §7 boot-fit≠runtime
caveat + soak-continuous pointer ONLY when the gate itself marked the run
boot-fit-satisfied (echoed from res.notices, never re-derived), states
which gate decided, is vLLM-only by construction, and never implies a
non-emitted artifact (the compose line appears only when res.emitted).
CONTRACT-1 user doc: docs/PULL.md gains a "Report a failed pull" section
documenting the SHIPPED V1/V2 on-ramp (capture-on-hard-block → surfaced
pointer → scripts/pull.sh --submit-last / --submit <dir>, consent prompt,
gh + gh-less). Every documented command/flag/output string was verified
verbatim against the live shipped CLI on this branch (docs-fidelity RED-
LINE). Leak-clean: only repo-relative .pull-captures/<slug>/<ts> forms,
no absolute paths.
§9-reconciliation: the release headline AND the readiness ledger in
docs/PULL.md now state explicitly that GGUF is deferred to a §9 cross-
engine design-unlock proposal, and that v0.8.2's scope is the failure
on-ramp + registry-expansion + whichllm-hw-detect + recommend — same
location/pattern v0.8.0 used for its §9-headline reconciliation.
Zero decision-logic change: gates.py / deriver.py / capture.py /
loop_input.py / classifier.py / dedup.py / submit_pull.py / hwdetect.py /
failure_fingerprints.yml / arch_patches.yml all byte-unchanged. pull.py
is a pure addition (zero removed lines): a new _render_recommendation()
function + a --recommend flag + one presentation-only call site.
test-pull.sh adds the CONTRACT-4 V5 section asserting the recommendation
TRACKS a real differing verdict — four genuinely-different real outcomes
(fit+emitted / confirm→proceed-blocked / estimated-lower-bound-fit /
hard-block) render four pairwise-different blocks, each matching its own
real res; rig-independent leak assertion (str(root) absent), not a
substring allowlist. Full shipped suite 25/25 green in the CI condition;
kv-calc --calibration 11/11 unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
CONTRACT-3 §8: an OPTIONAL, bounded subprocess that augments hardware
ENUMERATION for the eval path where nvidia-smi does not apply (AMD ROCm /
Apple / other-vendor). Strictly detect-only; never feeds kv-calc (kv-calc
stays the sole fit authority); no new hard dependency.
New isolated leaf module scripts/lib/profiles/hwdetect.py:
- detect_non_nvidia_hw()/detect_non_nvidia_sm(): bounded `whichllm list
--json` subprocess, defensively parsed into a structured HwDetectResult;
maps a recognised non-NVIDIA device class to an SM-equivalent for the
[C0] SM gate ONLY.
- Every non-delivery path (tool absent / failed / timeout / unparseable /
NVIDIA-only / unrecognised) degrades to None and NEVER raises out.
Additive consume-point wiring in run_pull (the eval path): a new optional
`hwdetect_fn` kwarg, consulted ONLY inside the existing
`if hardware_sm is None:` stratum-3 block — i.e. only when nvidia-smi
already returned nothing. The NVIDIA majority never enters the seam, so
that path is byte-identical whether the augment is absent OR
present-but-degrading. On a recognised non-NVIDIA device the eval path
gets an SM-equivalent (the [C0] SM gate runs instead of the blind-refuse
`hardware-sm-undetermined` terminal) plus an additive notice/diagnostic;
no shipped decision field is mutated and kv-calc is not consulted.
BOTH RED-LINE halves covered + proven by scripts/tests/test-hwdetect.sh:
(a) safety — optional/no-hard-dep, graceful degrade, NVIDIA-path
byte-identity, never feeds kv-calc; (b) delivery — a simulated
(explicit, deterministic) non-nvidia env yields a structured enumeration
the eval path observably consumes (outcome moves OFF the degrade
terminal). Rig-independent leak assertion (str(abs_dir) not in shared).
All 9 shipped v0.8.0/V1/V2/V3 decision modules byte-unchanged
(dedup.py:262 FInput.dedup_hash(_EffProxy()) idiom fenced/untouched).
Full scripts/tests/test-*.sh suite green in the CI condition (25/25);
kv-calc --calibration unchanged (11/11).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The first expansion marked ALL added arches requires_trust_remote_code:
unverified, so a registry-recognised model only moved no-arch-row ->
needs-trust-remote-code-ack — a lateral relabel, NOT the "materially more
models pass [C0] engine-supported" CONTRACT-2 requires (net newly-passing:
zero; caught on-rig via microsoft/phi-2).
Two-class TRC posture: long-standing native vLLM built-in classes (no
remote code — a documented upstream constraint) carry
requires_trust_remote_code:"false" with a documented-constraint evidence
anchor; arch families with genuine remote-code lineage
(Phi3SmallForCausalLM, InternLM2ForCausalLM) stay unverified/fail-closed.
Zero-false-pass preserved: gates.py's has_auto_map is an INDEPENDENT
OR-term, so a repo shipping auto_map still hard-blocks needs-trc-ack
regardless of the row flag — "false" removes only the arch-row-level
over-refusal, never the per-repo trust boundary.
On-rig (2026-05-18): microsoft/phi-2 (no auto_map) -> engine-supported
clean; PhiForCausalLM+auto_map -> needs-trc-ack; absent arch ->
no-arch-row; Phi3Small/InternLM2 -> needs-trc-ack. Suite 24/24,
kv-calc 22/22. test-pullgate-gates updated to the two-class invariant.
CONTRACT-2 (§10-R4) arch-family registry expansion: +13 safetensors arch
rows in arch_patches.yml (PhiForCausalLM — the microsoft/phi-2 STEP V1
on-rig no-arch-row anchor — Phi3Small, Gemma/Gemma3/Gemma3-CG, Starcoder2,
Cohere, InternLM2, Mixtral/Qwen2Moe/Qwen3Moe MoE, Qwen2-VL). Additive data
only, zero [C0]/decision-logic change. Zero false-pass by construction:
each follows the established estimated-lower-bound/unverified-TRC precedent
so [C0] still resolves needs-trust-remote-code-ack (fail-closed, bypassable
ONLY by --trust-remote-code) — the expansion drops only the
--experimental-arch requirement, never auto-passes; an arch still absent
still hard-blocks no-arch-row. test-pullgate-gates.sh proves both, plus the
#146-shape worked acceptance case (a hand-added awq_bf16_int4 weights
variant the expanded flag schema/parity machinery absorbs cleanly).
CONTRACT-2b-i chat-template attribution + behavioral drift_guard: new
`chat_template` delivery class (VALID_DELIVERY_MECHANISM); froggeric (22
composes — 18 direct + 4 nvlink* via REAL Docker Compose extends: merge)
and carnice (mount-only) brought under load_bearing_when + a behavioral
drift_guard whose check encodes the self-contained symmetric restart+settle
protocol (identical docker restart both arms, /v1/models healthy, 60s
settle, >=3 bench runs/arm, grand-mean same-segment compare, flag only a
3/3 deterministic regression). Effective coverage uses REAL merge
semantics: docker compose config (preferred) or a deterministic offline
extends: merge applying the same rules (additive sequence merge; `!reset`
removal) — never the unsound single-base text concat. .jinja artifact
discovery catches an orphan vendored template. test-patch-attribution.sh
adds the class checks + an H4 fixture asserting a `!reset` child AND a
stopped-extending child both lose coverage (the false-negative is the
dangerous direction). Generator emit kept in lock-step with reaches().
Documented as PATCH_POLICY.md §3.1. Rig-independent leak assertions added
(str(abs_dir) not in shared; repo-relative-only — never a /opt|/home
substring allowlist).
RED-LINE: gates.py/pull.py/deriver.py/capture.py/loop_input.py/
classifier.py/dedup.py/submit_pull.py/kv-calc.py/failure_fingerprints.yml
byte-unchanged; no shipped compose changed; patch_attribution.py c0_state/
is_artifact/compose_text/service_body byte-identical (additive only). Full
test-*.sh suite green in the CI condition; kv-calc --calibration N/N.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The gh-less paste fallback embedded the absolute bundle dir into the
PUBLIC issue body ("full redacted bundle at `/abs/.../.pull-captures/...`"),
violating the acceptance that nothing the on-ramp tells a user to share
contains an unredacted absolute path. Render a repo-relative
`.pull-captures/<slug>/<ts>` pointer instead. Strengthen the gh-less
leak assertion to a rig-independent check (the absolute bundle dir must
not appear; only the relative pointer may) — the prior /opt|/home check
passed under a tmp sandbox dir and missed this.
CONTRACT-1.2: pull prints the honest one-line on-ramp pointer whenever a
gate bundle was emitted for the run, keyed on the V1-recorded capture dir
— explicitly NOT gated on the exit code (the bypassable no-arch-row C0
advisory path exits 0 yet emits the #1 §10-R9 bundle). Gate path stays
I/O-free: a single stdout line, no network/prompt/auto-send. It does not
classify (suppression is loop-side at submit).
CONTRACT-1.3: scripts/pull.sh --submit-last / --submit <dir> is a distinct
top-level verb parsed before the slug/--profile-like requirement.
--submit-last re-reads the V1 shared .last marker at submit (the race
defense — surfaces the CURRENT bundle, never a silent wrong-bundle).
Re-shows bundle identity + the exact already-redacted payload, requires an
explicit y before any network, then reuses the shipped F5 dedup.submit
(effective_dedup_hash, bounded loop:dedup-<hash> labels, +1-or-open,
collision-safe verify, suppression/review-queue) — not reimplemented.
gh-less fallback runs post-F2 classification, gated on should_file:
should_file=True -> a prefilled public issues/new URL with the
loop:dedup-<hash> label and the deterministic title template; review-
queued (unknown / correct-refusal) -> the local _review-queue spool path
and the no-public-issue line, with NO public issues/new URL. Never raises;
degrades to the local spool + printed paste-path. Console is never a
submission source — only the redacted artifact is emitted.
New scripts/tests/test-submit-pull.sh (mocked gh, zero network): the
.last-marker race re-read, the bundle-emitted-but-exit-0 surfacing,
F5-reuse, the gh-less should_file branch with no public URL for review-
queued, gate-path I/O-free, and leak-hygiene. Full shipped suite green in
the CI condition; kv-calc --calibration unchanged at 22/22; safetensors
decision path byte-unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
CONTRACT-1.1 capture-on-hard-block (additive only; zero v0.8.0 decision-logic
change — the safetensors/GGUF paths are byte-unchanged):
- capture.py: new SEPARATE emit_gate_capture() (the emit_override_capture
byte-preserving precedent — NOT invoked by emit_capture()) writing a
pt1-gate.json + schema:2 manifest.json (outcome:hard-block, exact shipped
abort_reason, failure_class:null) per the per-abort-stratum key table
(model/arch/quant null pre-deriver; topology best-effort/nullable, capture-
only resolve; post-C0 always null). New shared write_last_marker() helper
(atomic tmp+os.replace) called from BOTH emit_capture() and the gate
emitter (centralization mandate — gate-only is the commonest failure).
- pull.py: pass-through capture on the 7 terminal hard-block return paths
(deriver / profile-like / hardware-sm-undetermined / C0 / C2a / no-fit-model
/ C1) — emits a bundle before the existing `return res`; the decision is
byte-unchanged; injectable gate_capture_fn; never raises.
- loop_input.py: BaseCaptureBundle typing.Protocol (Optional[dict] pt2-5);
FInput satisfies it by construction (verified: no isinstance(finput,FInput)
anywhere in F2/F5 — pure static retype, schema==1 byte-identical incl.
dedup_hash); new FInputGate + read_gate_bundle() (schema==2; validates ONLY
the always-present row + outcome==hard-block + failure_class is None — does
NOT reuse the 22-key validator); FInputGate.dedup_tuple() uses .get(k,None)
(behaviour-neutral schema-1, crash-safe schema-2, deterministic null-topo).
- classifier.py / dedup.py: F2+F5 parameter annotations retyped FInput ->
BaseCaptureBundle. The dedup.py FInput.dedup_hash(_EffProxy()) unbound-class
idiom is FENCED (unchanged — not tidied). Additive gate_abort_reason
_match_condition kind (reads pt1_gate.abort_reason; bool like sibling kinds;
no enum/routing change).
- failure_fingerprints.yml: seeded gate_abort_reason rules keyed on the
verified shipped abort_reason strings — only engine-support-unknown/
no-arch-row -> kernel-unsupported (public-filed); runtime-incompatible /
disk-short / hard-block / catch-alls -> unknown (review-queued, not filed).
- tests: extended test-{pull,pullemit-capture,loop-input,classifier,dedup}.sh
with the V1 RED-LINE proofs (emit_capture() still writes ONLY pt1-4; a
schema==1 bundle yields byte-identical FInput / ClassificationResult /
dedup_hash / effective_dedup_hash pre/post the protocol lift; the dedup
fence holds) + gate-emitter / read_gate_bundle / gate_abort_reason routing
/ shared .last marker coverage; all 22 shipped test-*.sh green in the CI
condition (gitignored .pull-captures absent), kv-calc --calibration 11/11.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
deriver.py:343 and :743 told users GGUF/.bin is "not supported until
v0.8.1". v0.8.1 has shipped (the fix/docs-fidelity stack) and GGUF was
deliberately de-scoped from the v0.8.2 feature work too (cross-engine
serving = a deferred §2/§9 design-unlock, not a near-term version). The
strings actively mislead users on master ("wait for v0.8.1" — which
exists and won't add it). Re-anchored both to accurate, version-free
wording: GGUF/.bin not supported — this path is vLLM + safetensors only.
String-only; zero decision-logic change. Surfaced by the v0.8.2 brief
r1 review (Major finding). Same docs-fidelity class as the v0.8.1 stack.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
PR #154 vendored models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/
to fix#153, but did not add its patch-attribution entry. test-patch-attribution
flags the install.sh as an orphan artifact (lacks patches.yml entry) → rc=1.
The repo has no PR-CI so the merge didn't catch it; master is red on this test.
Caught by the v0.8.1 pre-tag gate (full suite in CI condition) — exactly the
v0.8.0-lesson failure class that per-step verification misses.
Fix: add `gemma-vllm-pr41800-truncate-prompt-tokens` to patches.yml mirroring
the canonical `qwen-vllm-pr41800-truncate-prompt-tokens` entry (model=gemma-4-31b,
the 6 gemma dual compose registry ids that wire it, same delivery_spec/drift_guard/
upstream block), and list it in the Gemma4ForConditionalGeneration arch
required_patches for modeling consistency with the qwen arch. Engine-level overlay,
no behavior change. Full scripts/tests suite + kv-calc calibration 13/13 GREEN.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
v0.8.0 docs-fidelity finding #1. `pull.py` defined `_EXIT_USAGE=64` and
docs/pull.sh-header promised "64 = usage", but argparse's default
`error()` hard-exits `2` — colliding with `_EXIT_ABORT` (honest gate
hard-stop). A typo and a legitimate gate-block were indistinguishable to
callers/automation (both `2`).
Fix: a contained `_UsageExit64Parser(argparse.ArgumentParser)` overriding
`error()` to exit `_EXIT_USAGE` (64). `--help` is unaffected (goes through
`exit()`, still 0). Verified: no-args / missing-required / unknown-flag
-> 64; --help -> 0; honest hard-stop -> 2 (distinct again); full v0.8.0
suite + kv-calc 22/22 zero regression. Regression-locked by a new
CLI-contract block in test-pull.sh (the pure truth-table can't cover the
argv/exit boundary). docs/PULL.md exit-code table updated to the fixed
contract, with a note that the v0.8.0 *tag* still exits 2 (this lands on
master post-v0.8.0, ships with the next release — not a separate patch
tag, per the maintainer call).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Pre-tag full-branch review caught two classes the per-STEP verification
missed (each gated by transient dev-rig state, not visible per-commit):
1. Internal-path leaks in COMMITTED source (8 files): F-series sub-agent
docstrings/comments cited the internal locked-brief / design / on-rig
F8-log absolute paths. Non-functional but ship internal paths in a
public repo. Scrubbed to non-leaky grounding (which CONTRACT / §;
point to in-repo docs/LOOP.md). The test-pullemit-capture.sh
redaction-canary `/opt/ai` strings are deliberately LEFT (they test
that redaction strips them).
2. CI-robustness: test-pullemit-capture.sh (F6 G1 gate) and test-dedup.sh
(F5 real-data block) HARD-asserted `>=2 real .pull-captures/ bundles`.
`.pull-captures/` is gitignored runtime state — populated only after a
real on-rig pull, ALWAYS absent on a fresh clone / in CI. These passed
on the dev rig only because on-rig E5/F8 left captures behind; they
would RED the v0.8.0 tag's CI. Converted to skip-when-absent /
verify-when-present (the real invariant is the serialization-format /
round-trip of any captures present, not that the corpus exists).
Verified: full 14-suite run with .pull-captures ABSENT (the exact CI
condition) all RC=0; kv-calc --calibration 22/22; v0.8.0-changeset leak
sweep clean (only legit redaction canaries remain). Comment/docstring +
test-skip logic only — zero production decision-logic change.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
STEP 3: scripts/generate-compose.sh (+ scripts/lib/generate_compose.py).
Implements the brief's steps 1-10: scope gates first (type!=vllm /
genesis_equipped -> clean refuse), engine-pin loads:true validation
(image NEVER rewritten), arch via model_slugs/arch_model_xref, tp/kv
validation, trc {true,unverified} security refusal, compose-keyed patch
selection, delivery-gap-before-drift-guard, graded drift-guard
(capability-scoped -> OMIT+DEGRADED+--accept-degraded, foundational ->
hard-refuse, never repair, never wire a failed patch). Emits from the
captured compose_service_template: param-slots/constants verbatim, image
expression passed through verbatim, --trust-remote-code never emitted
in-scope (governed slot, locked §88), wiring re-derived only at the two
named insertion points, synthesizes nothing else. 3-category provenance
header above services: so STEP-2 service_body() discards it.
STEP 4: scripts/tests/test-generate-compose.sh. 5 golden triples
(vllm/minimal, vllm/dual, vllm/gemma-mtp, vllm/gemma-int8 [full,
multi-file overlay], vllm/gemma-dflash [dflash]) — all verified
genesis_equipped:false. Per triple: semantic diff vs shipped confined to
the two insertion points (image + constants verbatim), selected+wired
subset-of-shipped, wired pass reaches() on the GENERATED compose,
selected-but-undelivered NOT reachable, 3-category header, no
--trust-remote-code emitted. Plus the refusal/degraded matrix and
kv_arg() unit table. Imports (does not reimplement) patch_attribution.
test-patch-attribution.sh stays byte-identical (61 patch / 11 arch / 18
calibration, same 15 known-gap lines); all other test-*.sh remain RC=0.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Factor the embedded patch-attribution logic out of
scripts/tests/test-patch-attribution.sh into a reusable module
scripts/lib/profiles/patch_attribution.py (load, compose_text,
gap_declared, reaches, c0_state, schema/coverage helpers + key-sets).
The test now imports and calls the module; output is byte-identical
(same PASS summary + known-delivery-gaps list, RC=0).
reaches() is now sound (brief v9 correction #4): it parses the
comment-stripped service body only (ignoring the file-header banner and
the generator's own header WARNING block) and validates the patch's
actual delivery_spec wiring (declared mount target / entrypoint invoke
at the wired_at insertion points) instead of a bare patch["id"] in text
substring. Accepts a COMPOSE_REGISTRY profile name OR an absolute path.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
New compose models/gemma-4-26b-a4b/vllm/compose/dual/awq-mtp.yml adds
external MTP drafter wiring via google/gemma-4-26B-A4B-it-assistant
(832 MB BF16, already downloaded). COMPOSE_REGISTRY entry
vllm/gemma-a4b-awq-mtp at port 8043.
Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:
awq.yml (no MTP) → 138.88 / 138.67 wall TPS (139 / 139 decode)
awq-mtp.yml (n=4) → 155.05 / 207.02 wall TPS (157 / 211 decode)
+12% narr / +49% code
MTP metrics:
AL: 3.04 narrative, 3.79 code
Per-pos accept (narr): 0.77 / 0.55 / 0.40 / 0.29 (50.9% avg)
Per-pos accept (code): 73.5% avg
Both GPUs at 98% util / 357 W and 305 W (symmetric)
Cross-MoE finding (opposite direction from Qwen 35B-A3B MTP, see
preview-mtp.yml row in BENCHMARKS): Gemma's external assistant
drafter is a small dense model (~0.5 B params, Gemma4AssistantForCausalLM)
that BYPASSES the MoE expert routing entirely on the draft pass.
The Qwen built-in MTP head runs through the model's MoE forward, so
each draft step pays MoE routing + inter-GPU sync overhead — that's
why Qwen MTP n=3 was 50% SLOWER but Gemma MTP n=4 is 49% FASTER.
Practical rule for MoE models on Ampere: prefer external drafters
over built-in MTP heads where both are available. Documented in
learnings/gemma-4-26b-a4b.md + cross-ref in qwen3.6-35b-a3b.md.
Recommendation: awq-mtp.yml becomes the recommended default for
Gemma 26B-A4B going forward; awq.yml stays as the no-drafter A/B
reference.
PR #40886 overlay still applied via the same patches/install.sh
bind-mount pattern as awq.yml.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
New compose models/qwen3.6-35b-a3b/vllm/compose/dual/preview-mtp.yml
adds MTP n=3 via Qwen 35B-A3B's built-in head (mtp_num_hidden_layers=1).
COMPOSE_REGISTRY entry vllm/qwen-a3b-preview-mtp at port 8052.
Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:
preview.yml (no MTP) → 182.68 / 177.45 wall TPS
preview-mtp.yml (n=3) → 90.36 / 115.33 wall TPS
−51% narr / −35% code
Per-position MTP acceptance is healthy (0.927 / 0.810 / 0.698, AL 3.44,
81.2% avg) — the draft head is working correctly. The bottleneck is
inter-GPU sync overhead on the MoE forward path: asymmetric GPU util
(GPU 0 at 39% / 233 W vs GPU 1 at 79% / 191 W; non-MTP run had both
cards at 85-99%) shows one card stalling on draft-target communication.
vLLM warns at boot: "max_num_scheduled_tokens=4096 may lead to
suboptimal performance ... consider increasing max_num_batched_tokens
to accommodate the additional draft token slots."
Hypothesis: Cliff 2 mitigations (Genesis PN12 / PN25 / PN34) carry
the scheduler-side optimizations that make MTP profitable for
Qwen3-Next family on Ampere. Without them (preview path is no-Genesis
by definition), the MoE draft pass cost dominates the acceptance gain.
Re-test triggers (each separately worth trying):
(1) Genesis v7.73.x re-anchors on post-#42521 nightly
(2) bump max_num_batched_tokens to ~12K
(3) MTP n=2 to see if smaller spec depth changes calculus
For v0.7.3 ship: preview.yml (no MTP) stays the recommended default
for Qwen 35B-A3B users. preview-mtp.yml ships as documented A/B
reference + calibration anchor #2.
Compose, BENCHMARKS row, and the qwen3.6-35b-a3b learnings file all
flag the finding honestly so users can A/B on their own hardware.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.
vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.
This commit:
* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
install.sh: anchor-based Python patcher that inserts the 4 branches
into the in-container gemma4.py at runtime. Idempotent (sentinel
comment), drift-resistant (anchors on existing branch line, not
line numbers). Smoke-tested against bf610c2f: sentinel count=1
after first install, no-op on second install, file remains valid
Python after patching.
README.md: full vendor context, usage pattern, drop trigger.
* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
/mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
Entrypoint runs install.sh before vllm serve. Routes through
vllm-nightly-clean (bf610c2f, no Genesis).
* ModelProfile gemma-4-26b-a4b.yml updates:
- autoround_int4_mixed: status flipped to "ampere-blocked" with
a one-line forensic note (was "production", incorrect on SM86)
- awq_compressed_tensors: new variant pointing at the cyankiwi
AWQ-4bit weights, marked production via PR #40886 overlay
- default_weight_variant: switched to awq_compressed_tensors
(Ampere users get the variant that boots by default)
* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
path at port 8042.
* vllm-nightly-clean engine: supported_weight_formats gains
"compressed-tensors" so fits() C14 accepts the AWQ variant.
All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.
Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.
Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
was issue #137 / #138 territory — here we deliver the alternative
that unblocks Gemma 26B-A4B for the v0.7.3 ship).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).
Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text single/tools-text.yml fp8_e5m2
- vllm/minimal single/minimal.yml fp8_e5m2
- vllm/dual dual/docker-compose.yml fp8_e5m2
- vllm/dual-bf16 dual/bf16.yml bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp dual/qwopus-bf16mtp.yml fp8_e5m2
- vllm/dual-nvlink dual/nvlink.yml (extends) fp8_e5m2
- vllm/dual4 multi4/docker-compose.yml fp8_e5m2
TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.
Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
Genesis is required only for TQ3 KV format; that's enforced at the
compose level via Engine-profile selection.
Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
coverage.
Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
requires_genesis=true, so the Genesis enforcement happens at C15
(engine feature) level, not C6. Test now asserts positive (Qwen +
fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
non-Genesis engine fails C15).
All compat tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Genesis is strictly required only for TQ3 KV on vLLM. The Gemma 4 31B
default composes don't use TQ3 — they don't need Genesis. They DO need
a post-PR-#41745 nightly for Gemma 4 MTP support (per compose header
breadcrumbs), which 01d4d1ad (the Genesis-anchored MTP pin) predates.
Affected composes (all bf16/fp8 KV, no TQ3):
- models/gemma-4-31b/vllm/compose/single/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/bf16.yml
All three flipped from `Engine-profile: vllm-nightly-mtp` to
`vllm-nightly-clean` (bf610c2f, 2026-05-15) which is post-everything
relevant: PR #41745 (Gemma 4 MTP), PR #42102 (DFlash + INT8 PTH), and
PR #42521 (qwen3_5_moe weight loading).
COMPOSE_REGISTRY entries vllm/gemma-mtp-tp1, vllm/gemma-mtp, and
vllm/gemma-bf16 updated to match.
The DFlash and INT8 paths (vllm-nightly-dflash on e47c98ef and
vllm-nightly-full on e47c98ef) are untouched — their pins remain
correctly anchored to the SHAs that produced all current BENCHMARKS
rows for those compose families. TQ3-using composes
(vllm/gemma-int8-tq3) likewise stay on vllm-nightly-full (Genesis).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.
* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
comfortable KV headroom. No drafter on this base smoke compose.
* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.
* COMPOSE_REGISTRY: dual variants become the canonical keys
(vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
single variants renamed to vllm/gemma-a4b-single and
vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).
Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Commit 40f1ef78 (2026-05-14) accidentally set vllm-nightly-mtp.spec to
nightly-1acd67a7, which is the post-Gemma4-merge nightly used by the
gemma-4-31b compose. The Genesis v7.72.2 PROD pin is nightly-01d4d1ad:
- BENCHMARKS rows from 2026-05-05 onwards reference nightly-01d4d1ad3
- calibration/qwen3.6-27b.yml: engine_pin: vllm-nightly-01d4d1ad
- gemma-4-31b/vllm/compose/single header: "Qwen3.6 composes stay on the
v7.72.2 PROD pin (01d4d1ad3) since Genesis allowlist anchors there"
Effect of the bug: any Qwen 27B launch via launch_compat.py since
2026-05-14 would have used the un-blessed Gemma-merge SHA (1acd67a7),
potentially causing Genesis patch apply failures or silent skips.
Existing benches that hardcoded the image or used the older mechanism
were unaffected.
Notes field updated with the corrected anchor + a PRIOR-BUG breadcrumb
so future readers don't re-introduce the bump.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>