beellama.cpp composes (ghcr.io/{anbeeld/beellama.cpp,noonghunna/beellama-cpp})
were not matched by the is_llamacpp image regex, so they fell into the vLLM
HF-cache branch and skipped the model-presence check entirely. And even on the
llama.cpp branch, only `-m`/`--model` (the target) was checked — never the
`--spec-draft-model` drafter. Result: a missing target OR drafter GGUF passed
preflight and surfaced as a cryptic in-container "failed to open GGUF file"
crash instead of preflight's "download this:" hint (#288 eddie + George).
- Add `beellama` to the is_llamacpp image regex.
- Collect drafter paths from --spec-draft-model/--model-draft/-md into their own
presence check, with a DRAFT_FILE env override mirroring GGUF_FILE/MMPROJ_FILE.
- New test case: beellama target-present + drafter-missing must refuse with the
drafter path + its `hf download Anbeeld/...` hint; both present must pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Rename vllm/gemma-int8→gemma-int8-mtp + vllm/gemma-mtp→gemma-bf16-mtp (both carry MTP n=4; only KV format differs); kv-calc alias fix. Ladder gemma-bf16-mtp default 32K→131K (measured BF16 KV pool 196,527 tok @ 0.95; verify-stress 8/8, NIAH→120K, bench 118.8/154.3, zero TPS regression). Codify the dual-card max-context-priority rule in DUAL_CARD.md. Suite 41/41.
Single beellama/gemma-dflash 128K (caveats; 140K OOMs under load, CTX_SIZE=102400 agent-safe). New dual beellama/gemma-dflash-dual parked upstream-gated (262K works but DFlash multi-GPU broken in 07ac3ce; fixes on v0.3.0 dev branch, no release). switch.sh --list hides upstream-gated (parked) like deprecated with its own note. UPSTREAM row tracks Anbeeld #39. Suite 38/38.
Adds the max context size to switch.sh --list. registry-emit threads a ctx label through the variant TSV (before status_note); switch.sh renders it rightmost. Production -> bare ctx; caveats/NA -> folded into the health paren (comma-separated). Rounded to nearest K (163840->164K; 32768 reads 33K). Shows registry vs compose default as a single value when they match, 'validated/compose' (e.g. 164K/200K) when they drift -- the only 3 drifts are experimental lanes; all production/caveats match. Suite 38/38.
Move the two default qwen3.6-27b vLLM composes off the purge-prone nightly
(#167) onto immutable v0.22.0 — the engine the 35B-A3B already runs on:
- vllm/minimal: nightly -> v0.22.0; drop the PR-35936 required-tool-fallback
overlay + its install (its serving.py does `import vllm.beam_search`, removed
in newer vLLM -> crashes on v0.22.0; the bug it patched doesn't trip on
current stable). Keep the froggeric chat template.
- vllm/dual: v0.21.0 -> v0.22.0 (already overlay-free since cf1f14f).
Validated live on 2x3090, stock v0.22.0 (no Genesis, no source overlays):
- vllm/minimal TP=1: engine v0.22.0, 5/5 coherence probes clean, 21.1 GB.
- vllm/dual TP=2 + MTP n=3: engine v0.22.0, 5/5 clean (built-in MTP works on
stock v0.22.0 without Genesis), 22.3 GB. No first-call warmup garble.
Confirms qwen3.6-27b AutoRound INT4 + fp8 KV (+ MTP) runs clean on stable
v0.22.0 — the path off the nightly treadmill (#167). Also updates the
test-launch-compat hardcoded v0.21.0 expectation for fp8-mtp.yml -> v0.22.0.
Suite 38/38.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* Deprecate Genesis vLLM composes; repoint vLLM single default to vllm/minimal
Mark all 8 Genesis-loading qwen3.6-27b vLLM composes deprecated (registry
status + compose-header Status): vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/tools-text, vllm/dual-turbo,
vllm/dual-tq3-mtp-genesis. Genesis is on hold pending Sandermage's next stable
release; the stack is moving to stable vLLM + beellama single-card.
Repoint DEFAULTS[(qwen3.6-27b, vllm, single)]: vllm/default -> vllm/minimal (the
Genesis-free fp8-KV config). The `vllm/default` token now resolves single ->
vllm/minimal, dual -> vllm/dual. The model single-card default stays beellama/dflash.
Resolver guard tests updated (vllm/default single -> vllm/minimal); SINGLE_CARD.md
example commands redirected off the deprecated slugs. Suite 38/38.
Follow-ups: (1) bump vllm/minimal + vllm/dual to v0.22.0 (Genesis-free -> stable);
(2) SINGLE_CARD.md narrative reframe (single-card vLLM deprecated, beellama default).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* switch.sh --list: hide deprecated by default, reveal with --all
With 9 Genesis composes now deprecated, --list was cluttered. Hide deprecated
variants from the default --list (tally + display loops); --all reveals them
alongside other-topology variants. Footer shows "(+N deprecated hidden --all)"
so they stay discoverable, never silently dropped.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Gemma single-card had NO functional default (vLLM single is upstream-gated at
head_dim=512; ik-llama walls ~24K; stock llama.cpp ~12 TPS). beellama is the
only viable fast single-card Gemma-4 path, so promote it to fill the gap.
Registry: beellama/gemma-dflash experimental -> caveats + new
DEFAULTS[(gemma-4-31b, beellama, single)]. Resolver now returns it for Gemma
single (beellama is #1 in ENGINE_PREFERENCE[single]); Gemma dual unchanged.
Validated #441: 131-200K ctx, 47/88 TPS, 8-pack 109/150 think-off /
114/150 think-on (reasoning net-positive on Gemma). Reuses the multi-arch
image from #271 (sm_86/89/120; sm_89/120 compiled-not-validated).
Docs: compose header (caveats + Quality + ctx ceiling), README Gemma row,
BENCHMARKS row marker + intro, UPSTREAM.md row. Resolver guard test flipped
from a degradation assertion to beellama/gemma-dflash; full suite 38/38.
Re-test trigger: re-point the default to the no-fork mainline path when
llama.cpp#23398 (Gemma-4 MTP) merges (docs/UPSTREAM.md).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Make `beellama/dflash` the single-GPU default for Qwen3.6-27B, served via
an unofficial multi-arch image, and demote ik-llama to the balanced alt.
Decision basis (single-3090, 370 W; bench.sh n=5 + quality-test --full):
- Code TPS ~100 vs ik 69 — fastest single-card 3090 code path.
- 8-pack 107/150 (71%) think-off vs ik 99/150 (66%); 113/150 think-on
(same-session). DFlash is output-lossless, so quality == the Q5_K_S target.
- Context ceiling ladder (measured): 130K comfortable, 160K usable
(115K-tok prefill OK, 0.8 GB headroom), 200K boots but OOMs on prefill.
Ships 102K default; raise to 160K via CTX_SIZE.
- ik still wins narrative TPS (63), VRAM, and shipped vision — kept as the
balanced alt (`--variant ik-llama/iq4ks-mtp`).
Registry: `beellama/dflash` experimental -> caveats + a new
DEFAULTS[(qwen3.6-27b, beellama, single)] row. beellama is already #1 in
ENGINE_PREFERENCE[single], so the resolver now returns it for Qwen single.
Image: unofficial multi-arch ghcr build
`beellama-cpp:multiarch-b9459-07ac3ce` (sm_86/89/120 = RTX 3090/4090/5090),
built from Anbeeld/beellama.cpp .devops/cuda.Dockerfile with
CUDA_DOCKER_ARCH="86;89;120" + GGML_CUDA_FA_ALL_QUANTS=ON. Engine reports
`ARCHS = 860,890,1200`; cuobjdump confirms all 3 SASS. CAVEAT: sm_89/sm_120
are compiled but UNVALIDATED — only sm_86/3090 is verified on-rig.
Also bumps the gemma beellama compose to the same multi-arch image (stays
experimental — not promoted). Docs: README single-card framing,
INFERENCE_ENGINES.md, UPSTREAM.md row, BENCHMARKS.md Qwen beellama row.
Resolver guard tests flipped to the new default; full suite 38/38.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Every vLLM compose + engine-pin now defaults to stock vllm/vllm-openai
(nightly-SHA / v0.21.0 / v0.22.0); nothing builds or pulls the baked
vllm-club3090 image. Repoint the test-preflight-compose-deps fixture off the
retired club image to a stock tag (the image is incidental — the test asserts
on missing model weights). Document the vLLM delivery model in AGENTS.md:
patches are volume-mounted into the pinned stock image, not baked; the
vllm-club3090 GHCR package is retired-by-disuse (kept as historical release
artifacts, not deleted). Leaves the legacy dockerfile_bake delivery block +
its patch_attribution handler (test-covered, marked read-only) untouched.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Onboards Anbeeld's llama.cpp fork (beellama.cpp) as a registry engine
with two single-card DFlash composes, both experimental pending a
published Docker image:
- beellama/dflash Qwen3.6-27B Q5_K_S + Anbeeld DFlash-IQ4_XS,
q5_0/q4_1 KV, 102K ctx, port 8060
- beellama/gemma-dflash Gemma-4-31B Q4_K_S + Anbeeld DFlash-IQ4_XS,
q5_0/q4_1 KV, 102K ctx, port 8061
Composes reference the locally-built ${BEELLAMA_IMAGE:-beellama-cpp:local}
(entrypoint /app/llama-server) — upstream ships no pullable image, so
they ship as Status: Experimental and surface as (NA: experimental) in
--list (launch is --force-gated). Configs follow the fork's quickstart
docs (Qwen "Precision combo", Gemma "Balanced combo"), deviating only on
the stack-wide --reasoning off default.
Catalog wiring: engine profile beellama-local.yml; two DFlash GGUF
drafters (anbeeld-qwen-dflash, anbeeld-gemma-dflash); compose_registry.py
entries with kvcalc_key="SKIP" (llama.cpp-family, like ik-llama);
weights-map entries (target + draft GGUFs) on both models; q4_1 added to
rtx-3090 supported KV; INFERENCE_ENGINES + UPSTREAM notes; test counts
bumped (engines 8->9, drafters 6->8, registry 45->47, disk 46->48).
Default-resolver invariant preserved: beellama is #1 in
ENGINE_PREFERENCE[single], but its (NA) status + no DEFAULTS row means
the resolver skips it, so qwen3.6-27b single still resolves to
ik-llama/iq4ks-mtp (asserted by test-model-default-resolver.sh). It
auto-promotes only when a published image lands and the composes flip to
production. Full catalog suite green (only the pre-existing
test-submit-bench.sh fixture failure remains).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
`switch.sh --list` showed every slug regardless of topology, so a
single-GPU box listed dual/multi configs it can't actually run — noise.
Filter `--list` to the topologies the detected GPU count supports:
1 GPU -> single only, 2 -> single+dual, 4+ -> everything. Detection
reuses switch_topology_from_gpus (CUDA/NVIDIA_VISIBLE_DEVICES, then
nvidia-smi) mapped to a topology rank.
- `--all` (and the `--list-all` alias) bypass the filter for
discoverability; --list is deferred until args are parsed so
`--list --all` works in either order.
- Don't silently hide: print a one-line note naming the detected GPU
count and exactly which topologies were hidden, plus a `(+N … hidden
— --all)` tally in the header. No note under --all / when nothing is
hidden.
- Fail-open: with no detection signal (no selector, no nvidia-smi) show
ALL — never hide off a guessed "single" fallback.
- Counts (header + per-model) reflect the visible set; PR-A health
markers/grouping and PR-B Defaults view unchanged.
New test scripts/tests/test-list-topology-filter.sh covers 1/2-GPU
filtering, --all, --list-all, order independence, the --all-without-list
guard, and fail-open. Docs: --help usage, FAQ model-switch entry, README.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* feat(switch): model-default resolver + user-pinnable defaults
Add a two-layer default scheme on top of the existing <engine>/default map.
<engine>/default stays the maintainer's recommendation (read-only to users);
the new <model>/default is the user's preference — their .env pin if set,
else a curated pick for the detected topology.
compose_registry.py: two maintainer knobs next to DEFAULTS —
RECOMMENDED_DEFAULT_MODELS (short opt-in shortlist, not an exhaustive
ranking; new models are NOT auto-added) and ENGINE_PREFERENCE (per-topology
engine order). Plus resolver helpers (curated_default_target,
community_default_target stub, model_default_pin_key, engine_set/model_set,
slug_topology, model_of_slug).
registry-emit.sh: shared resolver model_default_target(root, model, topology)
— the single injection point for both launchers. Precedence ladder: --variant
(caller) -> .env pin -> community seam (None today) -> curated ENGINE_PREFERENCE
walk (skips non-functional (NA) slugs) -> degradation (notice + nearest-lower
topology, else a clear "pick explicitly" message; never crashes). Plus
x_default_dispatch: X/default with X in engine-set -> engine rec; X in
model-set -> model default; else error (engines + model-ids are disjoint).
switch.sh: <model>/default token; --set-default <slug> / --clear-default
<model> (round-trip the .env pin CLUB3090_DEFAULT_<MODELID>); a Defaults view
appended to --list (also standalone via --defaults) marking user-pin vs
curated. PR-A's grouping/markers/counts and the --force gate preserved.
launch.sh: bare invocation -> first installed shortlist model -> its
<model>/default (no full wizard); a pinned fast-path ("Launch your default
<slug>? [Y/n]"); a post-boot offer ("Make <slug> your default? [y/N]"). Any
narrowing flag keeps the explicit wizard path. Also load CLUB3090_DEFAULT_*
pin keys from .env even when MODEL_DIR is exported in the shell (the existing
.env loader is MODEL_DIR-gated, which would otherwise hide the pin).
Pin validation is warn + fall back, never blocking: unknown slug / wrong
model / topology-mismatch / (NA) status -> notice + curated default.
Tests: new test-model-default-resolver.sh (curated walk, (NA) skip,
degradation, X/default dispatch, pin override + all validation paths,
community seam skipped, .env round-trip); test-default-resolver.sh extended
with <model>/default launch.sh dispatch. Full suite green (only the
pre-existing test-submit-bench.sh fixture failure remains); 45 entries
unchanged; kv-calc calibration 17/17.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* docs: document the model-default resolver + user pins
Ship the user/contributor docs alongside the resolver code (repo convention —
update existing docs, no new top-level doc).
- README: single-card realign (ik-llama = fastest blessed single default,
llama.cpp = cliff-immune alternative — matches ENGINE_PREFERENCE single
order); add "pin your default" + <model>/default to Quick start.
- CLAUDE.md (= AGENTS.md): document RECOMMENDED_DEFAULT_MODELS +
ENGINE_PREFERENCE + the shared resolver as maintainer knobs next to DEFAULTS.
- FAQ: extend "switch to a different model" with <model>/default; new "How do
I set my own default config?" Q (two-layer model, --set-default/--clear-
default/--defaults, .env key, warn+fallback validation).
- SINGLE_CARD / DUAL_CARD: the resolver + the per-topology engine order; pin
hint.
- GETTING_STARTED: first-run uses <model>/default + the in-flow "set as
default?" prompt.
- ADDING_MODELS: a new model resolves via ENGINE_PREFERENCE; add a DEFAULTS row
per engine; RECOMMENDED_DEFAULT_MODELS is not auto-grown.
- UPSTREAM: beellama Docker-image row (gates beellama onboarding; the resolver
skips it today and it auto-promotes to single default on catalog).
- models/qwen3.6-27b/CHANGELOG: dated PR-B entry.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* feat(registry): add slug health/availability flag
Add a lifecycle `status` to every registry slug so `switch.sh --list`,
launch, and switch are no longer blind to a compose's health. Previously
status lived only in compose-header comments, which drifted: the Genesis
dual compose declared "Working (with Genesis)" while its pin is parked and
it won't boot clean — a user could boot a broken slug unknowingly.
- compose_registry.py: `_entry()` gains keyword-only `status`
(default "production") + `status_note`, validated against the enum
(production/caveats/experimental/preview/upstream-gated/deprecated).
Add `compose_header_status()` mapping a compose's profile-schema
`Status:` emoji to that enum.
- Sweep every compose `Status:` header to a canonical enum value and
re-flag the non-functional slugs: all *genesis* + gemma-4-31b single
fp8 -> upstream-gated; carnice -> caveats; qwopus + qwen-a3b-preview
-> preview; bf16/int8 A/Bs, llamacpp bounded-thinking, PRISM/APEX eval
lanes, gemma-4-26b-a4b onboarding -> experimental; tq3-mtp -> deprecated.
- registry-emit.sh emits `status` + `status_note` as the last two VARIANT
fields; both loaders + the parity tests read the extended field list.
- switch.sh --list: status marker (caveats -> "(caveats)", the NA set ->
"(NA: <word>)"); model/topology grouping preserved. Launch/switch gate:
production launches, caveats launches with a notice, NA warns + requires
--force. launch.sh surfaces the flag before delegating to switch.sh.
- New drift-guard test test-compose-status-drift.sh: registry status in
enum, compose header maps to enum, and the two agree.
45 entries unchanged; kv-calc calibration 17/17; full suite green (only
the pre-existing test-submit-bench fixture failure remains).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* feat(switch): add model/variant counts to --list
Header line shows supported-model count + total variants with the health
split (N production · N caveats · N NA); each model group shows its variant
count. Widen the marker column so the (NA: …)/(caveats) markers align.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
The four dual/autoround-int4/nvlink-*.yml composes (nvlink-fp8-mtp,
nvlink-turbo, nvlink-dflash, nvlink-dflash-noviz) and their
vllm/dual-nvlink* launch slugs are removed. They were thin Docker Compose
`extends:` stubs whose ONLY override was NVLINK_MODE=force_on — fully
redundant since every base dual compose now auto-detects NVLink at boot
via detect_nvlink.sh (NVLINK_MODE=auto): it probes nvidia-smi and flips on
NCCL_P2P_LEVEL=NVL + custom-all-reduce when a bridge is present, else
NCCL_P2P_DISABLE=1 + --disable-custom-all-reduce. NVLink rigs get the
identical fast path from the base dual compose with no separate slug;
force it explicitly with `NVLINK_MODE=force_on scripts/switch.sh vllm/dual`
if auto-detect ever misses.
Scope (prune only; no engine-image change — the v0.22.0 stable-engine
consolidation is a separate follow-up):
- delete 4 nvlink-*.yml composes
- compose_registry.py: drop the 4 _entry blocks (50 -> 46 entries)
- test-compose-registry-disk.sh: count guards 50/51 -> 46/47
- profile_runtime.yml + patches.yml: drop the 4 nvlink blocks/slug refs
- test-profiles-compat.sh: retire the C13/E3 NVLink-required-compose
scenarios (the estate NVLink-gating code is now dormant; cleanup tracked
as a follow-up, code retained)
- docs: switch.sh help, AGENTS.md, HARDWARE.md, UPSTREAM.md, patches/README.md,
fp8-mtp/dflash/dflash-noviz compose headers, per-model CHANGELOG removal
entry; add a docs/README.md pointer to `switch.sh --list` as the
authoritative registry-derived compose x slug matrix
Historical NVLink bench rows (JusefPol #31, danbedford #74/#92/#96) are
preserved in BENCHMARKS.md. Full test suite green (the lone
test-submit-bench failure is a worktree-isolation artifact — it needs
gitignored results/rebench/ fixtures absent from a fresh worktree; passes
on the working tree).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
fp8-mtp.yml was only ever an A/B reference -- MTP is net-negative on this MoE
at vLLM TP=2 (n2 ~= -45% / n3 ~= -51% wall TPS; the built-in-MTP draft forward
pays inter-GPU sync the acceptance can't amortize). It served no production
purpose, so remove it and its wiring.
Removed:
- models/qwen3.6-35b-a3b/vllm/compose/dual/autoround-int4/fp8-mtp.yml
- compose_registry.py entry `vllm/qwen-a3b-preview-mtp` (port 8052)
- its profile_runtime.yml capture block + calibration anchor (-> 17 anchors)
- the kv-calc COMPOSE_ALIAS_TEXT token
Test/doc updates for the new counts:
- test-kv-generic-dense + test-submit-pull: calibration contract 18/18 -> 17/17
- test-compose-registry-disk: registry 50->49, disk 51->50 composes
- test-report-calib: container ref -> live vllm-qwen36-35b-a3b-dual
- BENCHMARKS.md + KV_MATH.md notes
The MTP-net-negative finding stays recorded in the BENCHMARKS production row
and learnings/qwen3.6-35b-a3b.md. Full tests/ suite green (sole red is the
pre-existing fixture-dependent test-submit-bench, identical on clean master).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Promote the 35B-A3B MoE dual compose out of preview after full live validation
on 2× RTX 3090 (2026-05-30): max-ctx probe (262K fits, 7.81× concurrency, 2.05M-
token KV pool), bench (178/174 wall · 182 decode TPS, CV <0.5%), verify-stress
(NIAH-clean to 240K), soak-continuous (0 growth / 0 err / 0 silent / 100%
retention), quality --full (det 79/90 = 88%, aider 13/30 — ≥ ik-llama ref), and a
live vision smoke (read a test image correctly @ 262K).
- compose: preview.yml → fp8.yml (serving-stack filename per layout convention),
preview-mtp.yml → fp8-mtp.yml; fp8.yml = 262K + vision-on + no-MTP, vLLM v0.22.0
stable (no overlays), ✅ Production header.
- registry: slug vllm/qwen-a3b-preview → vllm/qwen-35b-a3b-dual, max_ctx
16384 → 262144, compose_path, DEFAULTS, kvcalc alias (kv-calc.py).
- profile_runtime capture re-synced to fp8.yml; test-pull slug updated.
- BENCHMARKS row + ADDING_MODELS pointer.
- MTP intentionally OFF: built-in MTP is net-negative on this MoE under vLLM TP=2
at BOTH n=2 (≈−45%) and n=3 (≈−51%) — the draft forward pays inter-GPU sync the
acceptance can't amortize (ik-llama avoids this single-card). fp8-mtp.yml kept as
the A/B reference only.
GATED: the registry-wide test-diagnose-profile fails for this entry because
kv-calc's qwen-MoE branch does NOT divide weights by TP (a deliberate hack:
inflated weights ≈ the vLLM-filled peak, which coincidentally passes at 16K but
starves the 262K KV in the fit-check). The config is empirically validated; this
is a kv-calc MoE-VRAM-model limitation. Lands AFTER the kv-calc fix PR (separate)
that distinguishes diagnose fit-need from the calibration filled-peak for cheap-KV
MoE. Calibration anchor left untouched here (the re-baseline belongs to that PR).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Every gemma-4-31b vLLM compose pinned a purged Docker Hub nightly
(e47c98ef / bf610c2f), un-bootable for fresh users (#250, #167 class).
Prune the 9-variant set to 3 and repoint survivors onto the immutable
stable release tag vllm/vllm-openai:v0.21.0 (never :latest).
Survivors: vllm/gemma-mtp (bf16 dual, default), vllm/gemma-int8 (int8-PTH
dual), vllm/gemma-mtp-tp1 (fp8 single). Dropped: gemma-dflash,
gemma-dflash-int8, gemma-int8-tq3, gemma-bf16, gemma-awq; the 262K path
folds into gemma-int8 via a CTX env override. Gemma is decoupled from the
shared Qwen nightly profiles via a new vllm-gemma-stable engine profile;
Qwen nightly profiles are untouched.
Diagnose decoupling: the pruned gemma dflash patch dirs doubled as the
diagnose tool's cross-model disk-source proxies for two Qwen overlays
(vllm-pr41703-dflash, vllm-pr42102-dflash-kv-quant). Those capabilities
are image-baked in the pinned nightly, not mountable files, so mark them
image_baked and have diagnose skip the disk-source check for image-baked
overlays. Runtime-neutral (no Qwen compose mounts them).
switch.sh launcher: handle the new VLLM_IMAGE engine-pin export (the twin
loop in launch.sh had it; switch.sh did not -> "unexpected engine pin
export"), and guard the VLLM_NIGHTLY_SHA echo for the image-only stable
profile (was unbound under set -u).
Validated on 2x RTX 3090 (Ampere, v0.21.0): full scripts/tests/*.sh suite
35/0; live boots of gemma-mtp (bf16) and gemma-int8 serve coherent output,
qwen vllm/dual regression-clean. gemma-mtp-tp1 (fp8_e4m3) is correctly
SM-gated (required_sm=9.0) -- confirmed Ampere-incompatible (Triton
"fp8e4nv not supported on sm_86"), a documented 32 GB+/Hopper variant;
beellama remains the single-card gemma path on Ampere.
Refs #451, #250, #167.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The producer could write a `boot-fit-measured` record with null decode
TPS when bench.sh output drifted (summary block absent/unparseable) or a
metric was missing. A measured record with null TPS is worse than no
record for optimizer calibration: it looks like real data.
For a measured result_class:
- raise MeasuredRecordError if no parseable bench summary block was found
(decode_TPS mean= absent => output drift), or if the parse produced no
decode TPS. This matches the module's existing fail-loud posture
(KeyError on an unknown registry tag).
- a malformed/absent `=== GPU state ===` line is a SOFT gap (VRAM is a
fingerprint extension, not the core measured TPS): surface it in a new
top-level `parse_warnings` list instead of raising, so the gap is
explicit and never a silent null.
Non-measured classes (predicted/derived) impose no decode-TPS
requirement; genuinely-optional optimizer fields stay null. The CLI
catches MeasuredRecordError (exit 2, clean message) and echoes warnings
to stderr. Extends the test with: measured + no decode summary => fail
loud; measured + bad/absent GPU line => parse_warnings; non-measured +
no decode => no raise; and a happy-path regression asserting empty
parse_warnings.
Addresses the Codex review Medium finding on #249.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Pure additive capture: parse a completed scripts/bench.sh stdout (no GPU,
no live model) + a compose_registry tag into ONE measurement-record JSON,
written to a per-rig gitignored corpus (results/measurement-records/).
Conforms to the optimizer design's FROZEN measurement-record field names
verbatim; optimizer-only fields (objective/confidence_tier/margin_applied)
emitted null, never fabricated. Producer-proposed additions (a context-depth
TPS ladder + power_cap_w fingerprint) namespaced under measured_extensions
and flagged as Lock-criteria #6 candidates. No consumer, no lookup, no
decision logic — that half is gated behind a separate design unlock.
bench.sh left untouched (load-bearing); emitter runs standalone on saved
bench output. Test runs GPU-free via bench.sh BENCH_MOCK output.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* Add ik-llama apex-fit-q8q5 variant for Qwen3.6-35B-A3B (#242)
Captures @laurimyllari's `--fit` + asymmetric q8_0(K)/q5_0(V) KV config
from discussion #241 as a single-card ik-llama variant on the APEX
I-Compact GGUF (registry tag `ik-llama/apex-fit-q8q5`, port 8057).
Measured 1× 3090 + 370 W, n=5:
q4/q4 mtp.yml baseline (first real run — APEX weights weren't on disk
before today): 96.47 / 144.01 wall TPS narr/code (CV 5.4% / 4.5%)
20.46 GB VRAM
q8/q5 fit-mtp.yml (this variant): 103.25 / 149.12 wall TPS
(CV 3.0% / 1.5%) — +7% narr / +4% code at tighter CV, +0.6 GB VRAM
The Anbeeld K-high/V-low asymmetric-KV pattern materialised here.
Gates:
verify-full 8/8 PASS
verify-stress 8/8 PASS incl. 180K NIAH (91% of n_ctx 196608)
bench above
soak-continuous PASS (0 errors, 0/25 silent_empty, 0 VRAM growth,
100% TPS retention, p50 decode 223 TPS)
deterministic q 76/90 = 84% on PR #38 verifiers
toolcall 14/15 · instructfollow 15/15 ·
structoutput 14/15 · dataextract 11/15 ·
reasonmath 12/15 · bugfind 10/15
MoE × MTP sub-question (from #242 body) — answered: ik-llama built-in
MTP on the 35B-A3B MoE does NOT pay the vLLM −51% / −35% penalty
(cf. `qwen3.6-35b-a3b/dual/preview-mtp.yml` BENCHMARKS row). MTP context
ready at n_ctx=196608, decode bursts 270+ TPS in soak. The MoE×MTP
penalty is vLLM-scheduler-specific, not architectural.
Status set to ⚠️ Production w/ caveats because sandbox-pack quality
(hermesagent-20 0/20, aider-polyglot-30 0/30, cli-40 timeouts) hit
benchlocal-cli sandbox infrastructure issues 2026-05-28 — hermes
`tool_events=0` after the agent runs (suspected PR #38 thinking-on
sampler interaction); aider fails at git checkout
`fatal: path 'aider/__init__.py' does not exist in 'f46766c'` before
any LLM call. Neither is attributable to the model or the compose.
Pre-PR-#38 cross-rig reference from @laurimyllari's 4090 in #241:
hermes 9-12/20, aider 14-18/30 — the model class is capable; the
sandbox state on this rig needs separate work.
Catalog gates: test-compose-registry-disk count bumped 55→56 per the
documented model-add workflow. test-profiles-compat, test-model-
weights-registry, test-switch-registry-parity, test-launch-registry-
parity all PASS. Two inherited reds (test-compose-mounts-resolve and
test-patch-attribution) unchanged from master baseline — they affect
qwen3.6-27b vLLM composes, not this addition.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
* fit-mtp.yml: tighten --fit/--no-mmap/--cache-ram rationale (PR #243 review)
@laurimyllari clarified on PR #243 that for the I-Compact GGUF
(~17 GB on 24 GB card) both `--fit` and `--no-mmap` are largely inert
since the model fits in VRAM — the reason to keep them as defaults
is forward-compat: swapping GGUF_FILE for a bigger quant
(UD-Q8_K_XL, APEX Quality, etc.) "just works" with reasonable
partial-offload performance, without re-tuning the compose.
Replaced the two "Why" blocks in the compose docstring to lead with
his framing. Also flagged `--cache-ram 4096` (lower than ik-llama's
8192 default) with the same suspected forward-compat rationale,
pending his confirmation on the PR.
No flag changes; docstring-only.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
* Promote fit-mtp.yml ✅ Production + update sandbox-pack quality
Updates the compose Status + Quality line + BENCHMARKS row after live
validation of the sandbox-pack fixes (benchlocal-cli #42/#43/#44 +
club-3090 #245, all merged today):
hermesagent-20: 0/20 → 11/20 (55%) via #42 deterministic sampler
aider-polyglot-30: 0/30 → 12/30 (40%) via #44 git-checkout from AIDER_DIR
cli-40: 11/40 → 12/40 (30%, ±1 noise) via #43 budget fix
(helps aider/hermes wall-clock; cli-40 "timeouts"
turn out to be sandbox-internal agent-give-up, not
wall-clock budget — see follow-up benchlocal-cli issue)
All gates clean: verify-full 8/8, verify-stress 8/8 (incl. 180K NIAH),
bench n=5 (+7% narr / +4% code vs q4/q4 mtp.yml), soak-continuous PASS
(0 errors, 0 silent_empty, 0 VRAM growth, 100% retention). Status:
🧪/⚠️ → ✅ Production. Caveats trimmed to the 3090 power-sensitivity
note.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
PR-B of the compose-quant-hierarchy refactor (follows PR-A #231). launch.sh no
longer hardcodes its variant/port/container/kvcalc maps — it derives them from
compose_registry.py via a shared emitter, the same single source of truth
switch.sh already uses. Adds a topology-aware <engine>/default resolver.
- scripts/lib/registry-emit.sh (new): shared emitter — emits switch/launch
tables from COMPOSE_REGISTRY, parses container_name from each compose, and
exposes registry_default_target() for <engine>[/<topology>]/default.
- launch.sh: drop the hardcoded `declare -A LAUNCH_*` maps; populate from the
emitter. PRIMARY_MODEL constant (default qwen3.6-27b) backs bare <engine>/default.
- switch.sh: drop inline derivation; resolve vllm/default, vllm/dual/default,
vllm/multi4/default via the shared emitter.
- compose_registry.py: add kvcalc_key to entries; tools/kv-calc.py aliases.
- bench-row-formatter.sh: DEFAULTS-aware compose_display (drops the now-dead
docker-compose.yml branch).
- tests: test-launch-registry-parity.sh + test-default-resolver.sh (new);
test-switch-registry-parity refreshed onto the shared emitter.
Default resolution is an alias layer over existing registry keys — no keys
renamed, no compose content changed. Side effect: the launcher<->registry drift
that left vllm/gemma-mtp pointing at a non-existent dual/.../fp8-mtp.yml is gone
(now derives the registry's bf16-mtp.yml).
Implemented via Codex per the PR-B brief; independently re-validated before
commit: 7/7 gates PASS (switch+launch parity, default-resolver, launch-compat,
registry-disk, mounts-resolve, kv-calc 22/22); live vllm/default->dual +
vllm/dual/default->dual both verify-full 8/8; leak-clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
`report.sh --full` always dumped kv-calc's full catalog-wide calibration
matrix (4 models × every compose), regardless of what the reporter runs, and
the vLLM-only skip didn't cover ik_llama. Three fixes:
(a) Scope to the running model. Resolve the active container -> kv-calc model
id and filter `--calibration` output to that model's `== <id> ==` section.
A Qwen single-card reporter no longer gets Gemma/MoE rows.
(b) Fix the ggml-engine skip. The skip only matched `llama-cpp-*`; `ik-llama-*`
fell through to "unknown" and ran the (inapplicable) calibration anyway.
Both ggml engines now emit the skip note.
(c) Opt-in full matrix. `--full-calibration` (or REPORT_FULL_CALIBRATION=1)
restores the catalog-wide matrix for maintainer triage. Unknown/unresolved
model also falls back to the full matrix.
Logic is factored into a pure, side-effect-free lib (scripts/lib/report_calib.sh:
calib_engine_for_container / calib_model_for_container / calib_filter_model_section)
so it's unit-testable. New test scripts/tests/test-report-calib.sh covers the
engine map (incl. the ik_llama regression), the model map (all 4 models), and
the section filter (keeps banner + target section + Overall, drops others;
empty id = passthrough).
Validated: bash -n; test-report-calib ok; live `kv-calc --calibration` scoped to
qwen3.6-27b keeps only that section + the Overall line.
Closes#168.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
We don't distribute a club-3090 vLLM image to users, so the
build-vllm-image.yml workflow + its surrounding docs are dead weight. The
workflow had also been failing on every tag push and the weekly cron because
its base nightly (nightly-1acd67a7) was purged from Docker Hub (#167/#407) —
that red CI on the v0.8.4 tag is what surfaced this.
Removed:
- .github/workflows/build-vllm-image.yml (the GHCR image builder)
- docker/vllm-club3090/Dockerfile (its build recipe)
- docs/CI_RUNNER_SETUP.md (build/distribution doc)
- README + UPSTREAM references to ghcr.io/noonghunna/vllm-club3090
- docs/README index link to the deleted CI doc
Kept: the generic VLLM_IMAGE override (now documented against
vllm/vllm-openai:latest — also the #167 workaround). The launch-compat test
still exercises that override, just with an upstream image.
Recoverable from history if we ever want to ship an image again.
NOT touched: patch_attribution.py still references the dockerfile_bake
delivery mode — left for a separate decision (internal, harmless).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Bug 1 — get_vram_free_mb inflated on multi-GPU hosts:
Summed ALL GPUs → reported 24507 MB when GPU0 truly had 381 free
(added idle GPU1's 24 GB). Margin gate was defeated on any host with
more GPUs than the model uses.
FIX: reads Docker HostConfig.DeviceRequests[0].DeviceIDs to identify
the model's GPU(s), passes nvidia-smi -i <those>. Falls back to
CUDA_VISIBLE_DEVICES / NVIDIA_VISIBLE_DEVICES env, then all-GPUs with
a warning. Live-validated: 353 MB (GPU0 only) vs old 24479 MB.
Bug 2 — filler→token ratio overshoots ~18x:
The rung used scale = target_tokens / 3.5, treating scale as chars
when it's actually block-repetition count. The 95K rung produced
1.77M tokens (674% of n_ctx) → HTTP 400 on a healthy engine.
FIX: calibration probe before the ladder sends scale=100, reads back
prompt_tokens, computes the real tok/scale_unit ratio. Live-validated:
95K rung now produces 94788 tokens (0.2% error, 36% of n_ctx).
Bug 3 — no-op ladder reported PASS:
A ladder that tested nothing (all rungs HTTP 400) said 'All stress
checks passed.' A rung with target < n_ctx returning 400 is a sizing
error, not a clean skip.
FIX: distinguishes sizing errors (target < n_ctx + HTTP 400) from
legitimate engine rejections (target > n_ctx + HTTP 400). Sizing
errors FAIL the probe. 'All rungs skipped' now FAILs with a clear
diagnostic instead of silently passing.
Tests: 37 pass (9 new — GPU-scoped VRAM query for Bug 1, dual/single/
scoped GPU variants, streaming helper timing extraction with real
llama.cpp response shape).
All three bugs were caught by running the ladder against the live
262K compose on the dual-3090 dev rig. The unit tests alone didn't
catch them — mocks fed the wrong shapes (top-level n_ctx, sum-all-GPUs,
clean-skip semantics). Live validation is now part of the workflow.
Add prefill speed measurement to all NIAH rungs (probes 1, 7, and 8)
so the ladder shows whether a context depth is *usable* (fast enough),
not just whether it *fits*. Prefill is ~O(N²) over context, so t/s
collapses as depth climbs — that latency cliff hits agent workloads
(which re-prefill a growing context every turn) before the VRAM cliff.
Implementation:
- send_streaming_niah() helper: sends NIAH requests with stream:true,
measures wall-clock TTFT (time to first token), extracts prefill
throughput from the response.
- Primary (llama.cpp): timings.prompt_per_second + timings.prompt_ms
from the final streaming chunk (confirmed against live endpoint).
- Fallback (cross-engine): prompt_tokens / TTFT_seconds when timings
is absent (vLLM, SGLang).
Output: each rung line now includes prefill data:
✓ 14500 tokens: recalled 'crimson otter 42' (got: ...) prefill=890 t/s (16s)
✓ rung 2/6: target=125K actual=124K tok (47%) recalled '...' prefill=512 t/s (242s) VRAM_free=15800MB
Purely additive/informational — no change to rung pass/fail logic or
the VRAM-margin gate. A slow-but-fits rung still passes.
Tests: 36 pass (8 new — streaming helper timing extraction with mocked
llama.cpp timings shape, vLLM cross-engine fallback, HTTP 500 handling).
verify-stress.sh's NIAH ladder used **fixed rungs** (~10K / 30K / 60K / 90K)
that **do not scale to the compose's CTX_SIZE**. Any compose with CTX_SIZE >
~115K was never exercised above ~90K — its top window was untested.
#197 surfaced the consequence: a hermes agent accumulated 127K tokens and
OOMed on a compose that had 'passed' at 90K max.
Fix: probe 8 is now a staggered ladder that climbs from ~95K to ~92% of
n_ctx in ~30K increments. Each rung:
- sends a NIAH request at that depth
- captures VRAM before/after
- stops at the first SYSTEM failure (OOM, crash, timeout) or recall miss
Recall miss = quality ceiling found:
- HTTP 200 + wrong answer → log △, break out of the ladder, pass the probe
- HTTP 500 / timeout / 000 → system wall, break + fail the probe
This gives you the exact fillable ceiling and the VRAM slope, not just
pass/fail at one arbitrary depth.
Engine health check + auto-restart:
After probes 7 and 8 (both crash-prone at deep contexts), verify-stress
checks if the engine is still alive. If it crashed (OOM kill, inductor
ICE), it attempts `docker restart` and polls for recovery (up to 120s).
This prevents cascade failures in rebench-full.sh's subsequent steps
(quality-test, soak-test, aider) that would otherwise run against a dead
engine. No-op for CONTAINER=none (endpoint-first mode).
Example output for a 262K compose:
✓ rung 1/6: target=95K actual=94K (36%) recalled 'crimson otter 42' VRAM=18200MB
✓ rung 2/6: target=125K actual=124K (47%) recalled 'amber falcon 17' VRAM=15800MB
△ rung 3/6: target=155K actual=153K (58%) recall MISS — quality ceiling reached VRAM=12100MB
✓ ceiling ladder: quality ceiling at 153000 tok (58% of n_ctx=262144) — recall miss, passed up to 124000 tok
Configuration:
CEILING_START_TOKENS First rung target (default: 95000)
CEILING_STEP_TOKENS Increment between rungs (default: 30000)
CEILING_FRACTION Top rung as fraction of n_ctx (default: 0.92)
VRAM_MARGIN_MB Warn if free VRAM drops below this (default: 1024)
SKIP_CEILING=1 Skip the entire ladder
Also adds:
- get_n_ctx() helper: reads /props (llama.cpp) → /v1/models (vLLM) →
docker inspect fallback
- get_vram_free_mb() helper: sums nvidia-smi memory.free across GPUs
- ensure_engine_alive(): health check + auto-restart after crash-prone probes
- Probes 1/7 recall miss also downgraded to informational (yellow △)
- Probe numbering updated from [N/7] to [N/8] throughout
- test-verify-stress-ceiling.sh: 26 tests covering helpers + ladder math
Closes#195.
Add ENABLE_THINKING support to bench.sh and pass --enable-thinking / --thinking-max-tokens through quality-test.sh. Propagate the same env through rebench-full.sh, warn on likely reasoning-on server/request-off mismatches, and document the workflow for reasoning-model evals.
Co-authored-by: noonghunna <[email protected]>
_render_recommendation keyed solely on res.ok, so a confirm→proceed /
override-accepted terminal (raw_verdict=fits-clean, res.ok=False because
the run needs an explicit --yes/--force-download) fell into the generic
"DOES NOT FIT / BLOCKED" branch — dishonest by imprecision (the model
fits; only acceptance is pending; CONTRACT-4 is "honest recommendation —
fits?"). Add a presentation-only needs-acceptance classification: a
fits-clean acceptance terminal now renders "FITS (estimated) — NOT YET
ACCEPTED" with the acceptance-gate guidance (not the failure on-ramp);
genuine hard-blocks still render DOES NOT FIT / BLOCKED. Pure
presentation — still derived only from res, no decision logic. Caught by
the V5 on-rig gate (microsoft/phi-2 --dry-run). test-pull.sh rec(2a)
updated to assert the honest rendering; suite 25/25, kv-calc N/N.
CONTRACT-4: add `--recommend` — an honest aggregated recommendation that
is PURE presentation/aggregation over the SHIPPED run_pull verdict. Every
line is read straight off the real PullResult (ok/confidence/raw_verdict/
terminal/stratum/abort_reason/notices/emitted); it introduces no decision
logic and does not change the exit code. Carries the §7 boot-fit≠runtime
caveat + soak-continuous pointer ONLY when the gate itself marked the run
boot-fit-satisfied (echoed from res.notices, never re-derived), states
which gate decided, is vLLM-only by construction, and never implies a
non-emitted artifact (the compose line appears only when res.emitted).
CONTRACT-1 user doc: docs/PULL.md gains a "Report a failed pull" section
documenting the SHIPPED V1/V2 on-ramp (capture-on-hard-block → surfaced
pointer → scripts/pull.sh --submit-last / --submit <dir>, consent prompt,
gh + gh-less). Every documented command/flag/output string was verified
verbatim against the live shipped CLI on this branch (docs-fidelity RED-
LINE). Leak-clean: only repo-relative .pull-captures/<slug>/<ts> forms,
no absolute paths.
§9-reconciliation: the release headline AND the readiness ledger in
docs/PULL.md now state explicitly that GGUF is deferred to a §9 cross-
engine design-unlock proposal, and that v0.8.2's scope is the failure
on-ramp + registry-expansion + whichllm-hw-detect + recommend — same
location/pattern v0.8.0 used for its §9-headline reconciliation.
Zero decision-logic change: gates.py / deriver.py / capture.py /
loop_input.py / classifier.py / dedup.py / submit_pull.py / hwdetect.py /
failure_fingerprints.yml / arch_patches.yml all byte-unchanged. pull.py
is a pure addition (zero removed lines): a new _render_recommendation()
function + a --recommend flag + one presentation-only call site.
test-pull.sh adds the CONTRACT-4 V5 section asserting the recommendation
TRACKS a real differing verdict — four genuinely-different real outcomes
(fit+emitted / confirm→proceed-blocked / estimated-lower-bound-fit /
hard-block) render four pairwise-different blocks, each matching its own
real res; rig-independent leak assertion (str(root) absent), not a
substring allowlist. Full shipped suite 25/25 green in the CI condition;
kv-calc --calibration 11/11 unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
CONTRACT-3 §8: an OPTIONAL, bounded subprocess that augments hardware
ENUMERATION for the eval path where nvidia-smi does not apply (AMD ROCm /
Apple / other-vendor). Strictly detect-only; never feeds kv-calc (kv-calc
stays the sole fit authority); no new hard dependency.
New isolated leaf module scripts/lib/profiles/hwdetect.py:
- detect_non_nvidia_hw()/detect_non_nvidia_sm(): bounded `whichllm list
--json` subprocess, defensively parsed into a structured HwDetectResult;
maps a recognised non-NVIDIA device class to an SM-equivalent for the
[C0] SM gate ONLY.
- Every non-delivery path (tool absent / failed / timeout / unparseable /
NVIDIA-only / unrecognised) degrades to None and NEVER raises out.
Additive consume-point wiring in run_pull (the eval path): a new optional
`hwdetect_fn` kwarg, consulted ONLY inside the existing
`if hardware_sm is None:` stratum-3 block — i.e. only when nvidia-smi
already returned nothing. The NVIDIA majority never enters the seam, so
that path is byte-identical whether the augment is absent OR
present-but-degrading. On a recognised non-NVIDIA device the eval path
gets an SM-equivalent (the [C0] SM gate runs instead of the blind-refuse
`hardware-sm-undetermined` terminal) plus an additive notice/diagnostic;
no shipped decision field is mutated and kv-calc is not consulted.
BOTH RED-LINE halves covered + proven by scripts/tests/test-hwdetect.sh:
(a) safety — optional/no-hard-dep, graceful degrade, NVIDIA-path
byte-identity, never feeds kv-calc; (b) delivery — a simulated
(explicit, deterministic) non-nvidia env yields a structured enumeration
the eval path observably consumes (outcome moves OFF the degrade
terminal). Rig-independent leak assertion (str(abs_dir) not in shared).
All 9 shipped v0.8.0/V1/V2/V3 decision modules byte-unchanged
(dedup.py:262 FInput.dedup_hash(_EffProxy()) idiom fenced/untouched).
Full scripts/tests/test-*.sh suite green in the CI condition (25/25);
kv-calc --calibration unchanged (11/11).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The first expansion marked ALL added arches requires_trust_remote_code:
unverified, so a registry-recognised model only moved no-arch-row ->
needs-trust-remote-code-ack — a lateral relabel, NOT the "materially more
models pass [C0] engine-supported" CONTRACT-2 requires (net newly-passing:
zero; caught on-rig via microsoft/phi-2).
Two-class TRC posture: long-standing native vLLM built-in classes (no
remote code — a documented upstream constraint) carry
requires_trust_remote_code:"false" with a documented-constraint evidence
anchor; arch families with genuine remote-code lineage
(Phi3SmallForCausalLM, InternLM2ForCausalLM) stay unverified/fail-closed.
Zero-false-pass preserved: gates.py's has_auto_map is an INDEPENDENT
OR-term, so a repo shipping auto_map still hard-blocks needs-trc-ack
regardless of the row flag — "false" removes only the arch-row-level
over-refusal, never the per-repo trust boundary.
On-rig (2026-05-18): microsoft/phi-2 (no auto_map) -> engine-supported
clean; PhiForCausalLM+auto_map -> needs-trc-ack; absent arch ->
no-arch-row; Phi3Small/InternLM2 -> needs-trc-ack. Suite 24/24,
kv-calc 22/22. test-pullgate-gates updated to the two-class invariant.
switch.sh now DERIVES its VARIANTS + VARIANT_DEFAULT_PORT tables from
compose_registry.py (the single source of truth) instead of a hardcoded
`declare -A` map that had drifted: 20 registered composes (incl.
vllm/dual-int8 — shipped as dual/int8.yml but unlaunchable, which cost a
real A/B a config pivot) were not launchable. All 42 registered composes
are now launchable; zero launcher-only ghosts.
New deterministic test-switch-registry-parity.sh (no docker/GPU/network)
fails CI on ANY registry↔launcher mismatch in EITHER direction: registry ⊆
launcher (zero registered-but-unlaunchable), launcher ⊆ registry (zero
ghosts) — driven through the FULL shipped `switch.sh --list` path so a
manual post-derivation ghost is caught too — plus spec parity, port parity,
and every resolved compose file exists on disk. Negative-case verified: a
synthetic registry/launcher mismatch makes the test exit 1.
Additive: no [C0]/decision-logic change; no shipped compose touched.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
CONTRACT-2 (§10-R4) arch-family registry expansion: +13 safetensors arch
rows in arch_patches.yml (PhiForCausalLM — the microsoft/phi-2 STEP V1
on-rig no-arch-row anchor — Phi3Small, Gemma/Gemma3/Gemma3-CG, Starcoder2,
Cohere, InternLM2, Mixtral/Qwen2Moe/Qwen3Moe MoE, Qwen2-VL). Additive data
only, zero [C0]/decision-logic change. Zero false-pass by construction:
each follows the established estimated-lower-bound/unverified-TRC precedent
so [C0] still resolves needs-trust-remote-code-ack (fail-closed, bypassable
ONLY by --trust-remote-code) — the expansion drops only the
--experimental-arch requirement, never auto-passes; an arch still absent
still hard-blocks no-arch-row. test-pullgate-gates.sh proves both, plus the
#146-shape worked acceptance case (a hand-added awq_bf16_int4 weights
variant the expanded flag schema/parity machinery absorbs cleanly).
CONTRACT-2b-i chat-template attribution + behavioral drift_guard: new
`chat_template` delivery class (VALID_DELIVERY_MECHANISM); froggeric (22
composes — 18 direct + 4 nvlink* via REAL Docker Compose extends: merge)
and carnice (mount-only) brought under load_bearing_when + a behavioral
drift_guard whose check encodes the self-contained symmetric restart+settle
protocol (identical docker restart both arms, /v1/models healthy, 60s
settle, >=3 bench runs/arm, grand-mean same-segment compare, flag only a
3/3 deterministic regression). Effective coverage uses REAL merge
semantics: docker compose config (preferred) or a deterministic offline
extends: merge applying the same rules (additive sequence merge; `!reset`
removal) — never the unsound single-base text concat. .jinja artifact
discovery catches an orphan vendored template. test-patch-attribution.sh
adds the class checks + an H4 fixture asserting a `!reset` child AND a
stopped-extending child both lose coverage (the false-negative is the
dangerous direction). Generator emit kept in lock-step with reaches().
Documented as PATCH_POLICY.md §3.1. Rig-independent leak assertions added
(str(abs_dir) not in shared; repo-relative-only — never a /opt|/home
substring allowlist).
RED-LINE: gates.py/pull.py/deriver.py/capture.py/loop_input.py/
classifier.py/dedup.py/submit_pull.py/kv-calc.py/failure_fingerprints.yml
byte-unchanged; no shipped compose changed; patch_attribution.py c0_state/
is_artifact/compose_text/service_body byte-identical (additive only). Full
test-*.sh suite green in the CI condition; kv-calc --calibration N/N.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The gh-less paste fallback embedded the absolute bundle dir into the
PUBLIC issue body ("full redacted bundle at `/abs/.../.pull-captures/...`"),
violating the acceptance that nothing the on-ramp tells a user to share
contains an unredacted absolute path. Render a repo-relative
`.pull-captures/<slug>/<ts>` pointer instead. Strengthen the gh-less
leak assertion to a rig-independent check (the absolute bundle dir must
not appear; only the relative pointer may) — the prior /opt|/home check
passed under a tmp sandbox dir and missed this.
CONTRACT-1.2: pull prints the honest one-line on-ramp pointer whenever a
gate bundle was emitted for the run, keyed on the V1-recorded capture dir
— explicitly NOT gated on the exit code (the bypassable no-arch-row C0
advisory path exits 0 yet emits the #1 §10-R9 bundle). Gate path stays
I/O-free: a single stdout line, no network/prompt/auto-send. It does not
classify (suppression is loop-side at submit).
CONTRACT-1.3: scripts/pull.sh --submit-last / --submit <dir> is a distinct
top-level verb parsed before the slug/--profile-like requirement.
--submit-last re-reads the V1 shared .last marker at submit (the race
defense — surfaces the CURRENT bundle, never a silent wrong-bundle).
Re-shows bundle identity + the exact already-redacted payload, requires an
explicit y before any network, then reuses the shipped F5 dedup.submit
(effective_dedup_hash, bounded loop:dedup-<hash> labels, +1-or-open,
collision-safe verify, suppression/review-queue) — not reimplemented.
gh-less fallback runs post-F2 classification, gated on should_file:
should_file=True -> a prefilled public issues/new URL with the
loop:dedup-<hash> label and the deterministic title template; review-
queued (unknown / correct-refusal) -> the local _review-queue spool path
and the no-public-issue line, with NO public issues/new URL. Never raises;
degrades to the local spool + printed paste-path. Console is never a
submission source — only the redacted artifact is emitted.
New scripts/tests/test-submit-pull.sh (mocked gh, zero network): the
.last-marker race re-read, the bundle-emitted-but-exit-0 surfacing,
F5-reuse, the gh-less should_file branch with no public URL for review-
queued, gate-path I/O-free, and leak-hygiene. Full shipped suite green in
the CI condition; kv-calc --calibration unchanged at 22/22; safetensors
decision path byte-unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
CONTRACT-1.1 capture-on-hard-block (additive only; zero v0.8.0 decision-logic
change — the safetensors/GGUF paths are byte-unchanged):
- capture.py: new SEPARATE emit_gate_capture() (the emit_override_capture
byte-preserving precedent — NOT invoked by emit_capture()) writing a
pt1-gate.json + schema:2 manifest.json (outcome:hard-block, exact shipped
abort_reason, failure_class:null) per the per-abort-stratum key table
(model/arch/quant null pre-deriver; topology best-effort/nullable, capture-
only resolve; post-C0 always null). New shared write_last_marker() helper
(atomic tmp+os.replace) called from BOTH emit_capture() and the gate
emitter (centralization mandate — gate-only is the commonest failure).
- pull.py: pass-through capture on the 7 terminal hard-block return paths
(deriver / profile-like / hardware-sm-undetermined / C0 / C2a / no-fit-model
/ C1) — emits a bundle before the existing `return res`; the decision is
byte-unchanged; injectable gate_capture_fn; never raises.
- loop_input.py: BaseCaptureBundle typing.Protocol (Optional[dict] pt2-5);
FInput satisfies it by construction (verified: no isinstance(finput,FInput)
anywhere in F2/F5 — pure static retype, schema==1 byte-identical incl.
dedup_hash); new FInputGate + read_gate_bundle() (schema==2; validates ONLY
the always-present row + outcome==hard-block + failure_class is None — does
NOT reuse the 22-key validator); FInputGate.dedup_tuple() uses .get(k,None)
(behaviour-neutral schema-1, crash-safe schema-2, deterministic null-topo).
- classifier.py / dedup.py: F2+F5 parameter annotations retyped FInput ->
BaseCaptureBundle. The dedup.py FInput.dedup_hash(_EffProxy()) unbound-class
idiom is FENCED (unchanged — not tidied). Additive gate_abort_reason
_match_condition kind (reads pt1_gate.abort_reason; bool like sibling kinds;
no enum/routing change).
- failure_fingerprints.yml: seeded gate_abort_reason rules keyed on the
verified shipped abort_reason strings — only engine-support-unknown/
no-arch-row -> kernel-unsupported (public-filed); runtime-incompatible /
disk-short / hard-block / catch-alls -> unknown (review-queued, not filed).
- tests: extended test-{pull,pullemit-capture,loop-input,classifier,dedup}.sh
with the V1 RED-LINE proofs (emit_capture() still writes ONLY pt1-4; a
schema==1 bundle yields byte-identical FInput / ClassificationResult /
dedup_hash / effective_dedup_hash pre/post the protocol lift; the dedup
fence holds) + gate-emitter / read_gate_bundle / gate_abort_reason routing
/ shared .last marker coverage; all 22 shipped test-*.sh green in the CI
condition (gitignored .pull-captures absent), kv-calc --calibration 11/11.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
v0.8.0 docs-fidelity finding #1. `pull.py` defined `_EXIT_USAGE=64` and
docs/pull.sh-header promised "64 = usage", but argparse's default
`error()` hard-exits `2` — colliding with `_EXIT_ABORT` (honest gate
hard-stop). A typo and a legitimate gate-block were indistinguishable to
callers/automation (both `2`).
Fix: a contained `_UsageExit64Parser(argparse.ArgumentParser)` overriding
`error()` to exit `_EXIT_USAGE` (64). `--help` is unaffected (goes through
`exit()`, still 0). Verified: no-args / missing-required / unknown-flag
-> 64; --help -> 0; honest hard-stop -> 2 (distinct again); full v0.8.0
suite + kv-calc 22/22 zero regression. Regression-locked by a new
CLI-contract block in test-pull.sh (the pure truth-table can't cover the
argv/exit boundary). docs/PULL.md exit-code table updated to the fixed
contract, with a note that the v0.8.0 *tag* still exits 2 (this lands on
master post-v0.8.0, ships with the next release — not a separate patch
tag, per the maintainer call).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Pre-tag full-branch review caught two classes the per-STEP verification
missed (each gated by transient dev-rig state, not visible per-commit):
1. Internal-path leaks in COMMITTED source (8 files): F-series sub-agent
docstrings/comments cited the internal locked-brief / design / on-rig
F8-log absolute paths. Non-functional but ship internal paths in a
public repo. Scrubbed to non-leaky grounding (which CONTRACT / §;
point to in-repo docs/LOOP.md). The test-pullemit-capture.sh
redaction-canary `/opt/ai` strings are deliberately LEFT (they test
that redaction strips them).
2. CI-robustness: test-pullemit-capture.sh (F6 G1 gate) and test-dedup.sh
(F5 real-data block) HARD-asserted `>=2 real .pull-captures/ bundles`.
`.pull-captures/` is gitignored runtime state — populated only after a
real on-rig pull, ALWAYS absent on a fresh clone / in CI. These passed
on the dev rig only because on-rig E5/F8 left captures behind; they
would RED the v0.8.0 tag's CI. Converted to skip-when-absent /
verify-when-present (the real invariant is the serialization-format /
round-trip of any captures present, not that the corpus exists).
Verified: full 14-suite run with .pull-captures ABSENT (the exact CI
condition) all RC=0; kv-calc --calibration 22/22; v0.8.0-changeset leak
sweep clean (only legit redaction canaries remain). Comment/docstring +
test-skip logic only — zero production decision-logic change.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>