Commit Graph
68 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 4a53edab43 llama-cpp: switch to rolling :server-cuda tag (no patches → no pin needed)
Follow-up on c3e7c7e: pinning to a specific build (b9246) was cargo-culted from
our vLLM pattern, but the vLLM pinning serves a real purpose (Genesis-patch
anchor, Docker Hub purge resistance) that doesn't apply here. llama.cpp on the
club-3090 stack is stock upstream — no patches, no Genesis equivalent — and
GHCR tag retention is more reliable than Docker Hub.

Switch to rolling `:server-cuda` tag so users automatically get MTP improvements,
EAGLE3 fixes, kernel updates from upstream without us being a bottleneck.
Override path preserved via `LLAMACPP_IMAGE` env if a future upstream build
regresses and a user needs to pin reactively.

Bench numbers in #170 footnoted as "measured on b9246"; expect ±5% drift on
newer builds.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-20 12:57:58 +00:00
noonghunnaandClaude Opus 4.7 c3e7c7ed80 llama-cpp: replace orphan llama-cpp:local with upstream pinned image (#170)
v0.8.3 shipped composes (llamacpp/default, llamacpp/mtp, llamacpp/mtp-vision) all
reference `image: llama-cpp:local`, a custom image that exists ONLY on the
maintainer's rig. There is no Dockerfile, no build script, and no setup.sh hook
to produce it for users. Anyone running `bash scripts/switch.sh llamacpp/mtp`
on a fresh clone hits "image llama-cpp:local not found" and dies at boot.

The custom image was a v0.8.3-dev artifact from when MTP PR #22673 was bleeding
edge. The official upstream `ghcr.io/ggml-org/llama.cpp:server-cuda` now has it
merged (build b9246 = commit 871b0b70f, 2026-05-20) — pinning to b9246 reproduces
the v0.8.3 numbers (50.25 narr / 58.04 code on single 3090, vs shipped 51.28/59.72).

Surfaced by @zemaphore in discussion #170. README.md was also lying: claimed
"both use the official ghcr.io image, no custom build needed" while composes
referenced llama-cpp:local.

Override the pin via `LLAMACPP_IMAGE=ghcr.io/.../server-cuda-bXXXX` env if you
want to follow upstream master.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-20 12:37:31 +00:00
noonghunnaandClaude Opus 4.7 ed1507122c llama.cpp single: thinking-off policy alignment + MTP profile family
Closes the dataextract 0/15 quality regression caused by mounting the
vLLM-only froggeric Jinja template on the llama.cpp engine, which
silently suppressed --reasoning off. Native template + --reasoning off
restores it. Stack-wide Qwen3.6 thinking-off-by-default policy now
shipped on the llama.cpp path (matches 22/24 vLLM composes).

Adds three named single-card profiles:
- llamacpp/default (262K vanilla, unchanged) — cliff-immune fallback
- llamacpp/mtp (NEW: 131K + MTP n=2) — single-card workhorse, ~60 TPS code
- llamacpp/mtp-vision (NEW: 49K + MTP + vision) — first stack profile
  combining MTP + vision on build 9235 (the older "strip mmproj when
  MTP" rule was obsolete; sweep-verified MTP + vision coexist)

Retires single/concurrent.yml — single-card concurrency is anti-value
(per-slot ctx ~48K, worse than every other profile). Concurrency
belongs on dual.

Fixes a latent user-facing bug: engines/llama-cpp-mainline.yml had
supported_drafters: [draft-mtp] (the llama.cpp CLI flag value) while
the compat-layer C7 gate compares against drafter spec_method (mtp).
Net effect pre-fix: llamacpp/mtp would have been silently filtered
out of launch.sh candidate lists. Fixed [draft-mtp] -> [mtp].

Lowers -ub default 2048 -> 1024 on the llama.cpp single composes. The
per-pass activation peak halves; verify-stress goes 5/7 -> 7/7
including the 60K + 91K needle rungs previously treated as
architectural Cliff 2 territory. Cliff 2 single-prompt at 50-60K was
config-driven on llama.cpp, not architectural; CLIFFS.md note added
(vLLM Cliff 2 narrative unchanged — different kernel-level failure).

Migrates 6 stale call sites for the engine-id rename
(llama-cpp-mainline -> llama-cpp-local) that earlier work lagged.

Measured impact, Config A (llamacpp/mtp, 131K, MTP n=2, ub=1024):
- bench: 51.28 narr / 59.72 code decode TPS (n=3, CV 1.9%/0.5%)
- verify-stress: 7/7 PASS (incl. 60K + 91K needle recall)
- quality 8-pack: 102/150 (68%) — beats every Qwen vLLM dual config in
  discussion #119 by 6-16 pp
- aider-polyglot-30: 17/30 (56.7%) — matches Qwen vLLM bf16 dual
  exactly (17/30) on half the hardware
- per-GPU code TPS (59.72) ~equal to vLLM dual configs (60-63),
  confirming the engine-side per-card rate is identical and vLLM
  dual's aggregate advantage is purely from the second card

Config B (llamacpp/mtp-vision, 49K, MTP + vision):
- multimodal vision probe passed
- bench: 56.52 narr / 66.17 code decode TPS (n=5)
- verify-stress: 7/7 PASS

Tests green: test-launch-compat, test-profiles-compat,
test-switch-registry-parity, test-pullgate-gates, test-patch-attribution.
Leak-grep clean. YAML lint clean. Performance charts regenerated to
include the two new entries.

Doc alignment: SINGLE_CARD, CLIFFS, FAQ, EXAMPLES, README, llama-cpp
README, COMPOSE_GENERATOR, PULL_GATE all reflect the new family.
CHANGELOG.md + UPSTREAM.md left intact per append-only history rule.

Followups (queued, not blocking): #400 MTP retune A/B (n=4 + p-min=0
on the fixed config), #403 soak-test container glob (unlocks Cliff 2b
validation on llama.cpp), #405 benchlocal-cli port-offset for parallel
rebench, #401 + #404 ik_llama + Tom's TurboQuant trackers.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-20 01:01:26 +00:00
noonghunna 26949d7fa8 fix(pull): v0.8.2 STEP V5 — recommend must not label a fits-clean model "DOES NOT FIT"
_render_recommendation keyed solely on res.ok, so a confirm→proceed /
override-accepted terminal (raw_verdict=fits-clean, res.ok=False because
the run needs an explicit --yes/--force-download) fell into the generic
"DOES NOT FIT / BLOCKED" branch — dishonest by imprecision (the model
fits; only acceptance is pending; CONTRACT-4 is "honest recommendation —
fits?"). Add a presentation-only needs-acceptance classification: a
fits-clean acceptance terminal now renders "FITS (estimated) — NOT YET
ACCEPTED" with the acceptance-gate guidance (not the failure on-ramp);
genuine hard-blocks still render DOES NOT FIT / BLOCKED. Pure
presentation — still derived only from res, no decision logic. Caught by
the V5 on-rig gate (microsoft/phi-2 --dry-run). test-pull.sh rec(2a)
updated to assert the honest rendering; suite 25/25, kv-calc N/N.
2026-05-18 21:05:12 +00:00
noonghunnaandClaude Opus 4.7 c5b5e9b27e feat(pull): v0.8.2 STEP V5 — recommend UX + report-a-failed-pull doc + §9-reconciliation
CONTRACT-4: add `--recommend` — an honest aggregated recommendation that
is PURE presentation/aggregation over the SHIPPED run_pull verdict. Every
line is read straight off the real PullResult (ok/confidence/raw_verdict/
terminal/stratum/abort_reason/notices/emitted); it introduces no decision
logic and does not change the exit code. Carries the §7 boot-fit≠runtime
caveat + soak-continuous pointer ONLY when the gate itself marked the run
boot-fit-satisfied (echoed from res.notices, never re-derived), states
which gate decided, is vLLM-only by construction, and never implies a
non-emitted artifact (the compose line appears only when res.emitted).

CONTRACT-1 user doc: docs/PULL.md gains a "Report a failed pull" section
documenting the SHIPPED V1/V2 on-ramp (capture-on-hard-block → surfaced
pointer → scripts/pull.sh --submit-last / --submit <dir>, consent prompt,
gh + gh-less). Every documented command/flag/output string was verified
verbatim against the live shipped CLI on this branch (docs-fidelity RED-
LINE). Leak-clean: only repo-relative .pull-captures/<slug>/<ts> forms,
no absolute paths.

§9-reconciliation: the release headline AND the readiness ledger in
docs/PULL.md now state explicitly that GGUF is deferred to a §9 cross-
engine design-unlock proposal, and that v0.8.2's scope is the failure
on-ramp + registry-expansion + whichllm-hw-detect + recommend — same
location/pattern v0.8.0 used for its §9-headline reconciliation.

Zero decision-logic change: gates.py / deriver.py / capture.py /
loop_input.py / classifier.py / dedup.py / submit_pull.py / hwdetect.py /
failure_fingerprints.yml / arch_patches.yml all byte-unchanged. pull.py
is a pure addition (zero removed lines): a new _render_recommendation()
function + a --recommend flag + one presentation-only call site.

test-pull.sh adds the CONTRACT-4 V5 section asserting the recommendation
TRACKS a real differing verdict — four genuinely-different real outcomes
(fit+emitted / confirm→proceed-blocked / estimated-lower-bound-fit /
hard-block) render four pairwise-different blocks, each matching its own
real res; rig-independent leak assertion (str(root) absent), not a
substring allowlist. Full shipped suite 25/25 green in the CI condition;
kv-calc --calibration 11/11 unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-18 20:11:22 +00:00
noonghunnaandClaude Opus 4.7 39177282b7 feat(pull): v0.8.2 STEP V4 — optional whichllm hw-detect subprocess (CONTRACT-3, hw-detect-only)
CONTRACT-3 §8: an OPTIONAL, bounded subprocess that augments hardware
ENUMERATION for the eval path where nvidia-smi does not apply (AMD ROCm /
Apple / other-vendor). Strictly detect-only; never feeds kv-calc (kv-calc
stays the sole fit authority); no new hard dependency.

New isolated leaf module scripts/lib/profiles/hwdetect.py:
- detect_non_nvidia_hw()/detect_non_nvidia_sm(): bounded `whichllm list
  --json` subprocess, defensively parsed into a structured HwDetectResult;
  maps a recognised non-NVIDIA device class to an SM-equivalent for the
  [C0] SM gate ONLY.
- Every non-delivery path (tool absent / failed / timeout / unparseable /
  NVIDIA-only / unrecognised) degrades to None and NEVER raises out.

Additive consume-point wiring in run_pull (the eval path): a new optional
`hwdetect_fn` kwarg, consulted ONLY inside the existing
`if hardware_sm is None:` stratum-3 block — i.e. only when nvidia-smi
already returned nothing. The NVIDIA majority never enters the seam, so
that path is byte-identical whether the augment is absent OR
present-but-degrading. On a recognised non-NVIDIA device the eval path
gets an SM-equivalent (the [C0] SM gate runs instead of the blind-refuse
`hardware-sm-undetermined` terminal) plus an additive notice/diagnostic;
no shipped decision field is mutated and kv-calc is not consulted.

BOTH RED-LINE halves covered + proven by scripts/tests/test-hwdetect.sh:
(a) safety — optional/no-hard-dep, graceful degrade, NVIDIA-path
byte-identity, never feeds kv-calc; (b) delivery — a simulated
(explicit, deterministic) non-nvidia env yields a structured enumeration
the eval path observably consumes (outcome moves OFF the degrade
terminal). Rig-independent leak assertion (str(abs_dir) not in shared).

All 9 shipped v0.8.0/V1/V2/V3 decision modules byte-unchanged
(dedup.py:262 FInput.dedup_hash(_EffProxy()) idiom fenced/untouched).
Full scripts/tests/test-*.sh suite green in the CI condition (25/25);
kv-calc --calibration unchanged (11/11).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-18 19:29:43 +00:00
noonghunna d78b9a9496 fix(pull): v0.8.2 STEP V3 — deliver CONTRACT-2's engine-supported broadening (TRC two-class)
The first expansion marked ALL added arches requires_trust_remote_code:
unverified, so a registry-recognised model only moved no-arch-row ->
needs-trust-remote-code-ack — a lateral relabel, NOT the "materially more
models pass [C0] engine-supported" CONTRACT-2 requires (net newly-passing:
zero; caught on-rig via microsoft/phi-2).

Two-class TRC posture: long-standing native vLLM built-in classes (no
remote code — a documented upstream constraint) carry
requires_trust_remote_code:"false" with a documented-constraint evidence
anchor; arch families with genuine remote-code lineage
(Phi3SmallForCausalLM, InternLM2ForCausalLM) stay unverified/fail-closed.
Zero-false-pass preserved: gates.py's has_auto_map is an INDEPENDENT
OR-term, so a repo shipping auto_map still hard-blocks needs-trc-ack
regardless of the row flag — "false" removes only the arch-row-level
over-refusal, never the per-repo trust boundary.

On-rig (2026-05-18): microsoft/phi-2 (no auto_map) -> engine-supported
clean; PhiForCausalLM+auto_map -> needs-trc-ack; absent arch ->
no-arch-row; Phi3Small/InternLM2 -> needs-trc-ack. Suite 24/24,
kv-calc 22/22. test-pullgate-gates updated to the two-class invariant.
2026-05-18 18:51:22 +00:00
noonghunnaandClaude Opus 4.7 999c93fe8c feat(pull): v0.8.2 STEP V3 — arch-registry expansion + chat-template attribution/drift_guard
CONTRACT-2 (§10-R4) arch-family registry expansion: +13 safetensors arch
rows in arch_patches.yml (PhiForCausalLM — the microsoft/phi-2 STEP V1
on-rig no-arch-row anchor — Phi3Small, Gemma/Gemma3/Gemma3-CG, Starcoder2,
Cohere, InternLM2, Mixtral/Qwen2Moe/Qwen3Moe MoE, Qwen2-VL). Additive data
only, zero [C0]/decision-logic change. Zero false-pass by construction:
each follows the established estimated-lower-bound/unverified-TRC precedent
so [C0] still resolves needs-trust-remote-code-ack (fail-closed, bypassable
ONLY by --trust-remote-code) — the expansion drops only the
--experimental-arch requirement, never auto-passes; an arch still absent
still hard-blocks no-arch-row. test-pullgate-gates.sh proves both, plus the
#146-shape worked acceptance case (a hand-added awq_bf16_int4 weights
variant the expanded flag schema/parity machinery absorbs cleanly).

CONTRACT-2b-i chat-template attribution + behavioral drift_guard: new
`chat_template` delivery class (VALID_DELIVERY_MECHANISM); froggeric (22
composes — 18 direct + 4 nvlink* via REAL Docker Compose extends: merge)
and carnice (mount-only) brought under load_bearing_when + a behavioral
drift_guard whose check encodes the self-contained symmetric restart+settle
protocol (identical docker restart both arms, /v1/models healthy, 60s
settle, >=3 bench runs/arm, grand-mean same-segment compare, flag only a
3/3 deterministic regression). Effective coverage uses REAL merge
semantics: docker compose config (preferred) or a deterministic offline
extends: merge applying the same rules (additive sequence merge; `!reset`
removal) — never the unsound single-base text concat. .jinja artifact
discovery catches an orphan vendored template. test-patch-attribution.sh
adds the class checks + an H4 fixture asserting a `!reset` child AND a
stopped-extending child both lose coverage (the false-negative is the
dangerous direction). Generator emit kept in lock-step with reaches().
Documented as PATCH_POLICY.md §3.1. Rig-independent leak assertions added
(str(abs_dir) not in shared; repo-relative-only — never a /opt|/home
substring allowlist).

RED-LINE: gates.py/pull.py/deriver.py/capture.py/loop_input.py/
classifier.py/dedup.py/submit_pull.py/kv-calc.py/failure_fingerprints.yml
byte-unchanged; no shipped compose changed; patch_attribution.py c0_state/
is_artifact/compose_text/service_body byte-identical (additive only). Full
test-*.sh suite green in the CI condition; kv-calc --calibration N/N.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-18 18:29:58 +00:00
noonghunna 52451ca0b0 fix(pull): v0.8.2 STEP V2 — gh-less issue body must not carry the absolute capture path
The gh-less paste fallback embedded the absolute bundle dir into the
PUBLIC issue body ("full redacted bundle at `/abs/.../.pull-captures/...`"),
violating the acceptance that nothing the on-ramp tells a user to share
contains an unredacted absolute path. Render a repo-relative
`.pull-captures/<slug>/<ts>` pointer instead. Strengthen the gh-less
leak assertion to a rig-independent check (the absolute bundle dir must
not appear; only the relative pointer may) — the prior /opt|/home check
passed under a tmp sandbox dir and missed this.
2026-05-18 17:49:00 +00:00
noonghunnaandClaude Opus 4.7 e1cdcb53c7 feat(pull): v0.8.2 STEP V2 — surface pointer + --submit-last/--submit (gh + gh-less, consented, F5 reuse)
CONTRACT-1.2: pull prints the honest one-line on-ramp pointer whenever a
gate bundle was emitted for the run, keyed on the V1-recorded capture dir
— explicitly NOT gated on the exit code (the bypassable no-arch-row C0
advisory path exits 0 yet emits the #1 §10-R9 bundle). Gate path stays
I/O-free: a single stdout line, no network/prompt/auto-send. It does not
classify (suppression is loop-side at submit).

CONTRACT-1.3: scripts/pull.sh --submit-last / --submit <dir> is a distinct
top-level verb parsed before the slug/--profile-like requirement.
--submit-last re-reads the V1 shared .last marker at submit (the race
defense — surfaces the CURRENT bundle, never a silent wrong-bundle).
Re-shows bundle identity + the exact already-redacted payload, requires an
explicit y before any network, then reuses the shipped F5 dedup.submit
(effective_dedup_hash, bounded loop:dedup-<hash> labels, +1-or-open,
collision-safe verify, suppression/review-queue) — not reimplemented.

gh-less fallback runs post-F2 classification, gated on should_file:
should_file=True -> a prefilled public issues/new URL with the
loop:dedup-<hash> label and the deterministic title template; review-
queued (unknown / correct-refusal) -> the local _review-queue spool path
and the no-public-issue line, with NO public issues/new URL. Never raises;
degrades to the local spool + printed paste-path. Console is never a
submission source — only the redacted artifact is emitted.

New scripts/tests/test-submit-pull.sh (mocked gh, zero network): the
.last-marker race re-read, the bundle-emitted-but-exit-0 surfacing,
F5-reuse, the gh-less should_file branch with no public URL for review-
queued, gate-path I/O-free, and leak-hygiene. Full shipped suite green in
the CI condition; kv-calc --calibration unchanged at 22/22; safetensors
decision path byte-unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-18 17:37:41 +00:00
noonghunnaandClaude Opus 4.7 20f1557d29 feat(pull): v0.8.2 STEP V1 — capture-on-hard-block pt1-gate emitter + BaseCaptureBundle protocol lift
CONTRACT-1.1 capture-on-hard-block (additive only; zero v0.8.0 decision-logic
change — the safetensors/GGUF paths are byte-unchanged):

- capture.py: new SEPARATE emit_gate_capture() (the emit_override_capture
  byte-preserving precedent — NOT invoked by emit_capture()) writing a
  pt1-gate.json + schema:2 manifest.json (outcome:hard-block, exact shipped
  abort_reason, failure_class:null) per the per-abort-stratum key table
  (model/arch/quant null pre-deriver; topology best-effort/nullable, capture-
  only resolve; post-C0 always null). New shared write_last_marker() helper
  (atomic tmp+os.replace) called from BOTH emit_capture() and the gate
  emitter (centralization mandate — gate-only is the commonest failure).
- pull.py: pass-through capture on the 7 terminal hard-block return paths
  (deriver / profile-like / hardware-sm-undetermined / C0 / C2a / no-fit-model
  / C1) — emits a bundle before the existing `return res`; the decision is
  byte-unchanged; injectable gate_capture_fn; never raises.
- loop_input.py: BaseCaptureBundle typing.Protocol (Optional[dict] pt2-5);
  FInput satisfies it by construction (verified: no isinstance(finput,FInput)
  anywhere in F2/F5 — pure static retype, schema==1 byte-identical incl.
  dedup_hash); new FInputGate + read_gate_bundle() (schema==2; validates ONLY
  the always-present row + outcome==hard-block + failure_class is None — does
  NOT reuse the 22-key validator); FInputGate.dedup_tuple() uses .get(k,None)
  (behaviour-neutral schema-1, crash-safe schema-2, deterministic null-topo).
- classifier.py / dedup.py: F2+F5 parameter annotations retyped FInput ->
  BaseCaptureBundle. The dedup.py FInput.dedup_hash(_EffProxy()) unbound-class
  idiom is FENCED (unchanged — not tidied). Additive gate_abort_reason
  _match_condition kind (reads pt1_gate.abort_reason; bool like sibling kinds;
  no enum/routing change).
- failure_fingerprints.yml: seeded gate_abort_reason rules keyed on the
  verified shipped abort_reason strings — only engine-support-unknown/
  no-arch-row -> kernel-unsupported (public-filed); runtime-incompatible /
  disk-short / hard-block / catch-alls -> unknown (review-queued, not filed).
- tests: extended test-{pull,pullemit-capture,loop-input,classifier,dedup}.sh
  with the V1 RED-LINE proofs (emit_capture() still writes ONLY pt1-4; a
  schema==1 bundle yields byte-identical FInput / ClassificationResult /
  dedup_hash / effective_dedup_hash pre/post the protocol lift; the dedup
  fence holds) + gate-emitter / read_gate_bundle / gate_abort_reason routing
  / shared .last marker coverage; all 22 shipped test-*.sh green in the CI
  condition (gitignored .pull-captures absent), kv-calc --calibration 11/11.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-18 16:43:28 +00:00
noonghunnaandClaude Opus 4.7 344ab87dd3 fix(deriver): correct stale "GGUF not supported until v0.8.1" message — now misleading post-v0.8.1-ship
deriver.py:343 and :743 told users GGUF/.bin is "not supported until
v0.8.1". v0.8.1 has shipped (the fix/docs-fidelity stack) and GGUF was
deliberately de-scoped from the v0.8.2 feature work too (cross-engine
serving = a deferred §2/§9 design-unlock, not a near-term version). The
strings actively mislead users on master ("wait for v0.8.1" — which
exists and won't add it). Re-anchored both to accurate, version-free
wording: GGUF/.bin not supported — this path is vLLM + safetensors only.
String-only; zero decision-logic change. Surfaced by the v0.8.2 brief
r1 review (Major finding). Same docs-fidelity class as the v0.8.1 stack.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-18 11:38:12 +00:00
noonghunnaandClaude Opus 4.7 b4b20ff7b6 fix(patch-attribution): register vendored gemma-4-31b pr41800 overlay (follow-up to #153/#154)
PR #154 vendored models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/
to fix #153, but did not add its patch-attribution entry. test-patch-attribution
flags the install.sh as an orphan artifact (lacks patches.yml entry) → rc=1.
The repo has no PR-CI so the merge didn't catch it; master is red on this test.

Caught by the v0.8.1 pre-tag gate (full suite in CI condition) — exactly the
v0.8.0-lesson failure class that per-step verification misses.

Fix: add `gemma-vllm-pr41800-truncate-prompt-tokens` to patches.yml mirroring
the canonical `qwen-vllm-pr41800-truncate-prompt-tokens` entry (model=gemma-4-31b,
the 6 gemma dual compose registry ids that wire it, same delivery_spec/drift_guard/
upstream block), and list it in the Gemma4ForConditionalGeneration arch
required_patches for modeling consistency with the qwen arch. Engine-level overlay,
no behavior change. Full scripts/tests suite + kv-calc calibration 13/13 GREEN.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 22:03:05 +00:00
noonghunnaandClaude Opus 4.7 820eb3845c fix(pull): argparse usage errors exit 64, not 2 — distinguishable from honest hard-stop (#370)
v0.8.0 docs-fidelity finding #1. `pull.py` defined `_EXIT_USAGE=64` and
docs/pull.sh-header promised "64 = usage", but argparse's default
`error()` hard-exits `2` — colliding with `_EXIT_ABORT` (honest gate
hard-stop). A typo and a legitimate gate-block were indistinguishable to
callers/automation (both `2`).

Fix: a contained `_UsageExit64Parser(argparse.ArgumentParser)` overriding
`error()` to exit `_EXIT_USAGE` (64). `--help` is unaffected (goes through
`exit()`, still 0). Verified: no-args / missing-required / unknown-flag
-> 64; --help -> 0; honest hard-stop -> 2 (distinct again); full v0.8.0
suite + kv-calc 22/22 zero regression. Regression-locked by a new
CLI-contract block in test-pull.sh (the pure truth-table can't cover the
argv/exit boundary). docs/PULL.md exit-code table updated to the fixed
contract, with a note that the v0.8.0 *tag* still exits 2 (this lands on
master post-v0.8.0, ships with the next release — not a separate patch
tag, per the maintainer call).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 20:58:10 +00:00
noonghunnaandClaude Opus 4.7 49d9bb4313 v0.8.0 [review] pre-tag fixes: scrub internal-path leaks from shipped source + make .pull-captures-corpus tests CI-safe (skip-when-absent)
Pre-tag full-branch review caught two classes the per-STEP verification
missed (each gated by transient dev-rig state, not visible per-commit):

1. Internal-path leaks in COMMITTED source (8 files): F-series sub-agent
   docstrings/comments cited the internal locked-brief / design / on-rig
   F8-log absolute paths. Non-functional but ship internal paths in a
   public repo. Scrubbed to non-leaky grounding (which CONTRACT / §;
   point to in-repo docs/LOOP.md). The test-pullemit-capture.sh
   redaction-canary `/opt/ai` strings are deliberately LEFT (they test
   that redaction strips them).

2. CI-robustness: test-pullemit-capture.sh (F6 G1 gate) and test-dedup.sh
   (F5 real-data block) HARD-asserted `>=2 real .pull-captures/ bundles`.
   `.pull-captures/` is gitignored runtime state — populated only after a
   real on-rig pull, ALWAYS absent on a fresh clone / in CI. These passed
   on the dev rig only because on-rig E5/F8 left captures behind; they
   would RED the v0.8.0 tag's CI. Converted to skip-when-absent /
   verify-when-present (the real invariant is the serialization-format /
   round-trip of any captures present, not that the corpus exists).

Verified: full 14-suite run with .pull-captures ABSENT (the exact CI
condition) all RC=0; kv-calc --calibration 22/22; v0.8.0-changeset leak
sweep clean (only legit redaction canaries remain). Comment/docstring +
test-skip logic only — zero production decision-logic change.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 12:01:25 +00:00
noonghunnaandClaude Opus 4.7 f92624d9a7 v0.8.0 [F] F8-fix: widen §6.1 Tier-1 OOM signature + pt3.actual regexes to real vLLM v0.21.0+ KV-cache-too-large phrasing — on-rig F8 caught classic-torch-only regexes miss the common KV-prediction failure
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 05:19:51 +00:00
noonghunnaandClaude Opus 4.7 1ac048189d v0.8.0 [F] F6: CONTRACT-5 mandatory content-hash kv_calc_version (G2) + G1 topo-verify + L2 fixture sync
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 04:50:50 +00:00
noonghunnaandClaude Opus 4.7 5de7224a73 v0.8.0 [F] F5: §6.3 canonical-tuple-hash dedup + bounded label scheme + collision-safe submit path (CONTRACT-4)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 04:43:55 +00:00
noonghunnaandClaude Opus 4.7 d758f08fce v0.8.0 [F] F4: §6.2 inbound-trust pipeline raw→candidate→validated→Tier-1 + CONTRACT-3a derived-deferral (CONTRACT-3)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 04:35:17 +00:00
noonghunnaandClaude Opus 4.7 b1009793f2 v0.8.0 [F] F3: G6-A 3-part additive [E] touch (pt1.predicted_b_breakdown, pt3.failure_log_excerpt+actual, container-log capture) + §6.1 Tier-1 (CONTRACT-2)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 04:19:53 +00:00
noonghunnaandClaude Opus 4.7 9f80d29fbe v0.8.0 [F] F2: §6.1 Tier-2 semantic-fingerprint classifier + Appendix A seed DB (CONTRACT-2 Tier-2)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 04:08:02 +00:00
noonghunnaandClaude Opus 4.7 1491cbc7af v0.8.0 [F] F1: FInput capture-bundle reader + schema-1 validation + key-normalization (CONTRACT-1)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 04:00:54 +00:00
noonghunnaandClaude Opus 4.7 71148d6054 v0.8.0 [E] E-outcome-fix: honest 3-state manifest outcome (partial-success != failed) — §6.2 partial is a capability-scoped success
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 02:34:44 +00:00
noonghunnaandClaude Opus 4.7 f7c405a06d v0.8.0 [E] E3/E4-fix: boot lifecycle as context manager (server stays up for smoke+capture, teardown on ctx-exit) — on-rig E5 caught teardown-in-finally-before-smoke
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 02:10:40 +00:00
noonghunnaandClaude Opus 4.7 16a1e4d944 v0.8.0 [E] E3-fix: smoke probes the real served-model-name (not literal "derived") + capture failure detail — on-rig E5 caught red-smoke-on-healthy-boot
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 01:47:16 +00:00
noonghunnaandClaude Opus 4.7 3ae74bfdcf v0.8.0 [E] E2-fix-2: verify *.safetensors against HF API lfs.sha256 (not Xet-redirect-fragile HEAD x-linked-etag) — on-rig E5 caught false no-etag
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 01:28:22 +00:00
noonghunnaandClaude Opus 4.7 806a298522 v0.8.0 [E] E2-fix: download via hf CLI subprocess (not huggingface_hub lib-import) — on-rig E5 caught ModuleNotFoundError
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 01:03:26 +00:00
noonghunnaandClaude Opus 4.7 2ed18aad3f v0.8.0 [E] E4: post-[C1] derived-[E] orchestration + trigger semantics + override force-capture (pt5)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 00:50:13 +00:00
noonghunnaandClaude Opus 4.7 f327887c39 v0.8.0 [E] E3: derived boot (HF_HOME mount) + 4 §6 capture emitters + manifest + derived smoke floor
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 00:34:39 +00:00
noonghunnaandClaude Opus 4.7 7a2ec86640 v0.8.0 [E] E2: HF download stage (download_set allowlist + x-linked-etag SHA, no-etag fail-closed, atomic staging)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 00:24:36 +00:00
noonghunnaandClaude Opus 4.7 411c84fd8a v0.8.0 [E] E1: generate_from_profile + derived-vllm template + EInput + CONTRACT-5 gate
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 00:12:26 +00:00
noonghunnaandClaude Opus 4.7 087a8ea692 v0.8.0 Pull-Gate P4-fix: price Tier-1 curated via curated-exact kv-calc spec, not generic-dense (+ non-mocked regression test)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 22:03:09 +00:00
noonghunnaandClaude Opus 4.7 adf7a3bf13 v0.8.0 Pull-Gate P4: stratum-5 + [C1] §4.1 total fn + stratum-6 [D] dry-run + pull orchestrator + exhaustive test-pull.sh
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 21:31:53 +00:00
noonghunnaandClaude Opus 4.7 4a1d3857f2 v0.8.0 Pull-Gate P3: stratum-2 precondition + [C0] engine-support/runtime/hardware gate + [C2a] disk
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 21:07:36 +00:00
noonghunnaandClaude Opus 4.7 818b79ccb5 v0.8.0 Pull-Gate P2: transformers deriver + ModelProfile/confidence + variant-scoped hf_repos schema
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 20:56:19 +00:00
noonghunnaandClaude Opus 4.7 6d7a043d90 v0.8.0 STEP 3+4: compose generator + 5-triple golden-parity test (#141)
STEP 3: scripts/generate-compose.sh (+ scripts/lib/generate_compose.py).
Implements the brief's steps 1-10: scope gates first (type!=vllm /
genesis_equipped -> clean refuse), engine-pin loads:true validation
(image NEVER rewritten), arch via model_slugs/arch_model_xref, tp/kv
validation, trc {true,unverified} security refusal, compose-keyed patch
selection, delivery-gap-before-drift-guard, graded drift-guard
(capability-scoped -> OMIT+DEGRADED+--accept-degraded, foundational ->
hard-refuse, never repair, never wire a failed patch). Emits from the
captured compose_service_template: param-slots/constants verbatim, image
expression passed through verbatim, --trust-remote-code never emitted
in-scope (governed slot, locked §88), wiring re-derived only at the two
named insertion points, synthesizes nothing else. 3-category provenance
header above services: so STEP-2 service_body() discards it.

STEP 4: scripts/tests/test-generate-compose.sh. 5 golden triples
(vllm/minimal, vllm/dual, vllm/gemma-mtp, vllm/gemma-int8 [full,
multi-file overlay], vllm/gemma-dflash [dflash]) — all verified
genesis_equipped:false. Per triple: semantic diff vs shipped confined to
the two insertion points (image + constants verbatim), selected+wired
subset-of-shipped, wired pass reaches() on the GENERATED compose,
selected-but-undelivered NOT reachable, 3-category header, no
--trust-remote-code emitted. Plus the refusal/degraded matrix and
kv_arg() unit table. Imports (does not reimplement) patch_attribution.

test-patch-attribution.sh stays byte-identical (61 patch / 11 arch / 18
calibration, same 15 known-gap lines); all other test-*.sh remain RC=0.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 15:49:37 +00:00
noonghunnaandClaude Opus 4.7 60f3983283 v0.8.0 STEP 2: extract patch_attribution.py (sound body-only reaches(), test imports it)
Factor the embedded patch-attribution logic out of
scripts/tests/test-patch-attribution.sh into a reusable module
scripts/lib/profiles/patch_attribution.py (load, compose_text,
gap_declared, reaches, c0_state, schema/coverage helpers + key-sets).
The test now imports and calls the module; output is byte-identical
(same PASS summary + known-delivery-gaps list, RC=0).

reaches() is now sound (brief v9 correction #4): it parses the
comment-stripped service body only (ignoring the file-header banner and
the generator's own header WARNING block) and validates the patch's
actual delivery_spec wiring (declared mount target / entrypoint invoke
at the wired_at insertion points) instead of a bare patch["id"] in text
substring. Accepts a COMPOSE_REGISTRY profile name OR an absolute path.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 15:29:10 +00:00
noonghunnaandClaude Opus 4.7 9f23736f01 v0.8.0 Phase A-prime: enrich patch/profile data for #141 generator (compose_service_template, genesis_equipped, delivery metadata, drift_guards, drafter/model_slug/trc fold-ins)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-16 15:20:35 +00:00
noonghunna 91a9622619 Add v0.8 Phase A patch attribution data 2026-05-16 05:13:00 +00:00
noonghunna d7804107c9 Reapply "fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean"
This reverts commit d32e168a89.
2026-05-15 20:52:21 +00:00
noonghunna d32e168a89 Revert "fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean"
This reverts commit 3b2d940d26.
2026-05-15 18:40:12 +00:00
noonghunna 39e18733aa feat(kv-calc): model v0.7.3 MoE architectures 2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 6dc9a0dce1 feat(gemma-4-26b-a4b): AWQ + MTP n=4 — +12% narr / +49% code over no-MTP baseline
New compose models/gemma-4-26b-a4b/vllm/compose/dual/awq-mtp.yml adds
external MTP drafter wiring via google/gemma-4-26B-A4B-it-assistant
(832 MB BF16, already downloaded). COMPOSE_REGISTRY entry
vllm/gemma-a4b-awq-mtp at port 8043.

Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:

  awq.yml        (no MTP) → 138.88 / 138.67 wall TPS (139 / 139 decode)
  awq-mtp.yml    (n=4)    → 155.05 / 207.02 wall TPS (157 / 211 decode)
                            +12% narr / +49% code

MTP metrics:
  AL: 3.04 narrative, 3.79 code
  Per-pos accept (narr): 0.77 / 0.55 / 0.40 / 0.29 (50.9% avg)
  Per-pos accept (code): 73.5% avg
  Both GPUs at 98% util / 357 W and 305 W (symmetric)

Cross-MoE finding (opposite direction from Qwen 35B-A3B MTP, see
preview-mtp.yml row in BENCHMARKS): Gemma's external assistant
drafter is a small dense model (~0.5 B params, Gemma4AssistantForCausalLM)
that BYPASSES the MoE expert routing entirely on the draft pass.
The Qwen built-in MTP head runs through the model's MoE forward, so
each draft step pays MoE routing + inter-GPU sync overhead — that's
why Qwen MTP n=3 was 50% SLOWER but Gemma MTP n=4 is 49% FASTER.

Practical rule for MoE models on Ampere: prefer external drafters
over built-in MTP heads where both are available. Documented in
learnings/gemma-4-26b-a4b.md + cross-ref in qwen3.6-35b-a3b.md.

Recommendation: awq-mtp.yml becomes the recommended default for
Gemma 26B-A4B going forward; awq.yml stays as the no-drafter A/B
reference.

PR #40886 overlay still applied via the same patches/install.sh
bind-mount pattern as awq.yml.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 e1d44bd732 feat(qwen-35b-a3b): preview-MTP compose + bench row — MTP measured SLOWER on MoE
New compose models/qwen3.6-35b-a3b/vllm/compose/dual/preview-mtp.yml
adds MTP n=3 via Qwen 35B-A3B's built-in head (mtp_num_hidden_layers=1).
COMPOSE_REGISTRY entry vllm/qwen-a3b-preview-mtp at port 8052.

Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:

  preview.yml         (no MTP) → 182.68 / 177.45 wall TPS
  preview-mtp.yml     (n=3)    →  90.36 / 115.33 wall TPS
                                  −51% narr / −35% code

Per-position MTP acceptance is healthy (0.927 / 0.810 / 0.698, AL 3.44,
81.2% avg) — the draft head is working correctly. The bottleneck is
inter-GPU sync overhead on the MoE forward path: asymmetric GPU util
(GPU 0 at 39% / 233 W vs GPU 1 at 79% / 191 W; non-MTP run had both
cards at 85-99%) shows one card stalling on draft-target communication.

vLLM warns at boot: "max_num_scheduled_tokens=4096 may lead to
suboptimal performance ... consider increasing max_num_batched_tokens
to accommodate the additional draft token slots."

Hypothesis: Cliff 2 mitigations (Genesis PN12 / PN25 / PN34) carry
the scheduler-side optimizations that make MTP profitable for
Qwen3-Next family on Ampere. Without them (preview path is no-Genesis
by definition), the MoE draft pass cost dominates the acceptance gain.

Re-test triggers (each separately worth trying):
  (1) Genesis v7.73.x re-anchors on post-#42521 nightly
  (2) bump max_num_batched_tokens to ~12K
  (3) MTP n=2 to see if smaller spec depth changes calculus

For v0.7.3 ship: preview.yml (no MTP) stays the recommended default
for Qwen 35B-A3B users. preview-mtp.yml ships as documented A/B
reference + calibration anchor #2.

Compose, BENCHMARKS row, and the qwen3.6-35b-a3b learnings file all
flag the finding honestly so users can A/B on their own hardware.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 0053444e84 feat(gemma-4-26b-a4b): AWQ path via vLLM PR #40886 overlay
The Intel AutoRound INT4 quants for Gemma 4 26B-A4B are structurally
blocked on SM86 (moe_intermediate_size=704 not aligned to Marlin's
group_size=128 = 5.5x; no SM86 WNA16 kernel handles unaligned K-dim).
The AWQ-4bit variant ships in compressed-tensors format, which routes
through a different vLLM kernel path that DOES handle arbitrary K
shapes — but vLLM's gemma4.py::_weight_iterator (as of bf610c2f) lacks
the _packed/_scale key remapping for compressed-tensors MoE experts.

vLLM PR #40886 (open, last updated 2026-04-25, author tested on RTX
3090 24 GB) adds the 4 remapping branches that yield per-expert
weight_packed / weight_scale keys for FusedMoE's loader. +23 / -0
in vllm/model_executor/models/gemma4.py.

This commit:

* `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/`
  install.sh: anchor-based Python patcher that inserts the 4 branches
  into the in-container gemma4.py at runtime. Idempotent (sentinel
  comment), drift-resistant (anchors on existing branch line, not
  line numbers). Smoke-tested against bf610c2f: sentinel count=1
  after first install, no-op on second install, file remains valid
  Python after patching.
  README.md: full vendor context, usage pattern, drop trigger.

* `models/gemma-4-26b-a4b/vllm/compose/dual/awq.yml` (new):
  TP=2, port 8042, bf16 KV, max_ctx 32K, no drafter. Targets
  /mnt/models/huggingface/gemma-4-26b-a4b-awq-4bit (17 GB on disk).
  Entrypoint runs install.sh before vllm serve. Routes through
  vllm-nightly-clean (bf610c2f, no Genesis).

* ModelProfile gemma-4-26b-a4b.yml updates:
  - autoround_int4_mixed: status flipped to "ampere-blocked" with
    a one-line forensic note (was "production", incorrect on SM86)
  - awq_compressed_tensors: new variant pointing at the cyankiwi
    AWQ-4bit weights, marked production via PR #40886 overlay
  - default_weight_variant: switched to awq_compressed_tensors
    (Ampere users get the variant that boots by default)

* COMPOSE_REGISTRY: new "vllm/gemma-a4b-awq" entry for the dual AWQ
  path at port 8042.

* vllm-nightly-clean engine: supported_weight_formats gains
  "compressed-tensors" so fits() C14 accepts the AWQ variant.

All 7 test suites still pass; test-profiles-compat now validates
40 COMPOSE_REGISTRY entries.

Track in docs/UPSTREAM.md: drop overlay when PR #40886 merges upstream
AND vllm-nightly-clean pin bumps past the merge commit.

Refs: #138 (the user-facing trigger that surfaced the AWQ-on-Ampere path
   was issue #137 / #138 territory — here we deliver the alternative
   that unblocks Gemma 26B-A4B for the v0.7.3 ship).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunna 99328b4cda feat(estate): add parallel boot mode 2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 3b2d940d26 fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean
Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).

Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text         single/tools-text.yml      fp8_e5m2
- vllm/minimal            single/minimal.yml         fp8_e5m2
- vllm/dual               dual/docker-compose.yml    fp8_e5m2
- vllm/dual-bf16          dual/bf16.yml              bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp  dual/qwopus-bf16mtp.yml  fp8_e5m2
- vllm/dual-nvlink        dual/nvlink.yml (extends)  fp8_e5m2
- vllm/dual4              multi4/docker-compose.yml  fp8_e5m2

TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.

Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
  Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
  Genesis is required only for TQ3 KV format; that's enforced at the
  compose level via Engine-profile selection.

Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
  coverage.

Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
  for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
  requires_genesis=true, so the Genesis enforcement happens at C15
  (engine feature) level, not C6. Test now asserts positive (Qwen +
  fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
  non-Genesis engine fails C15).

All compat tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 cf0451a7ad fix(gemma-4-31b): route default/bf16 composes to vllm-nightly-clean
Genesis is strictly required only for TQ3 KV on vLLM. The Gemma 4 31B
default composes don't use TQ3 — they don't need Genesis. They DO need
a post-PR-#41745 nightly for Gemma 4 MTP support (per compose header
breadcrumbs), which 01d4d1ad (the Genesis-anchored MTP pin) predates.

Affected composes (all bf16/fp8 KV, no TQ3):
- models/gemma-4-31b/vllm/compose/single/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/docker-compose.yml
- models/gemma-4-31b/vllm/compose/dual/bf16.yml

All three flipped from `Engine-profile: vllm-nightly-mtp` to
`vllm-nightly-clean` (bf610c2f, 2026-05-15) which is post-everything
relevant: PR #41745 (Gemma 4 MTP), PR #42102 (DFlash + INT8 PTH), and
PR #42521 (qwen3_5_moe weight loading).

COMPOSE_REGISTRY entries vllm/gemma-mtp-tp1, vllm/gemma-mtp, and
vllm/gemma-bf16 updated to match.

The DFlash and INT8 paths (vllm-nightly-dflash on e47c98ef and
vllm-nightly-full on e47c98ef) are untouched — their pins remain
correctly anchored to the SHAs that produced all current BENCHMARKS
rows for those compose families. TQ3-using composes
(vllm/gemma-int8-tq3) likewise stay on vllm-nightly-full (Genesis).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 2d1b1dc347 feat(moe): add dual-card composes for Gemma 26B-A4B + Qwen 35B-A3B preview
Both new MoE models now have dual-card (TP=2) variants as their primary
bench targets, mirroring the production posture used across the matrix.

* models/gemma-4-26b-a4b/vllm/compose/dual/docker-compose.yml
  TP=2, port 8041, bf16 KV, max_ctx 32K. ~8 GB/card weights leaves
  comfortable KV headroom. No drafter on this base smoke compose.

* models/qwen3.6-35b-a3b/vllm/compose/dual/preview.yml
  TP=2 (which is the max — num_kv_heads=2 caps valid_tp at [1, 2]),
  port 8051, fp8_e5m2 KV, max_ctx 16K. ~10 GB/card weights leaves
  ~12 GB for KV. Preview-only path on vllm-nightly-clean — Cliff 2
  mitigations and TQ3 KV unavailable until Genesis v7.73.x re-anchors.

* COMPOSE_REGISTRY: dual variants become the canonical keys
  (vllm/gemma-a4b → dual TP=2, vllm/qwen-a3b-preview → dual TP=2);
  single variants renamed to vllm/gemma-a4b-single and
  vllm/qwen-a3b-preview-single (for 1-GPU users / debugging).

Single Qwen 35B-A3B is technically risky on 24 GB (20 GB weights leaves
only ~2 GB for KV+activations), but preserved as a debugging option.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 87f0a0c528 fix(engines): vllm-nightly-mtp anchors to 01d4d1ad (Sander v7.72.2 PROD pin)
Commit 40f1ef78 (2026-05-14) accidentally set vllm-nightly-mtp.spec to
nightly-1acd67a7, which is the post-Gemma4-merge nightly used by the
gemma-4-31b compose. The Genesis v7.72.2 PROD pin is nightly-01d4d1ad:

- BENCHMARKS rows from 2026-05-05 onwards reference nightly-01d4d1ad3
- calibration/qwen3.6-27b.yml: engine_pin: vllm-nightly-01d4d1ad
- gemma-4-31b/vllm/compose/single header: "Qwen3.6 composes stay on the
  v7.72.2 PROD pin (01d4d1ad3) since Genesis allowlist anchors there"

Effect of the bug: any Qwen 27B launch via launch_compat.py since
2026-05-14 would have used the un-blessed Gemma-merge SHA (1acd67a7),
potentially causing Genesis patch apply failures or silent skips.
Existing benches that hardcoded the image or used the older mechanism
were unaffected.

Notes field updated with the corrected anchor + a PRIOR-BUG breadcrumb
so future readers don't re-introduce the bump.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00