The FP8 max-accuracy tier had bench (83.1/108.2) + quality (107/150) on the
2x3090 reference rig but no soak-continuous run — the one missing gate item.
Ran the full operational gate fresh on the v0.24.0 pin:
- verify-full 9/9
- verify-stress: all 6 rungs, fillable to 240,636 tok clean (91%)
- soak-continuous PASS: 0 err / 0 growth / 100% retention, p50 decode 85
That clears the production bar.
- registry + compose header: status experimental -> production (drift guard
green). status_note also de-staled: dropped the '~56 TPS' probe / 'slowest of
the three' framing (corrected to decode 83/108, the slow axis is PREFILL/TTFT
from MarlinFP8 W8A16 on Ampere, not decode) + the full-gate results.
- baselines.yml: row comment notes the soak PASS completing the gate.
- No DEFAULTS[(qwen,vllm,dual)] change -> vllm/dual (fast) stays the dual
default; dual-max just joins the actionable list as the max-fidelity tier.
Bonus for the 5090 crowd: dual-max is now a non-experimental config, so the
launch command drops --force (bash scripts/switch.sh vllm/qwen-27b-dual-max).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
#588 fixed the fp8-weights DeepGEMM crash but the _deepgemm_env gate (and my
guard) matched weights_variant == 'fp8' EXACTLY — missing agents-a1, whose
weights are 'fp8-dynamic' (compressed-tensors FP8). Those are genuine FP8
weights that route FP8 GEMM, so a 5090 running vllm/agents-a1-dual would still
hit the 'recipe not found' crash the whole framework is meant to prevent.
- _deepgemm_env gate: weights_variant == 'fp8' -> startswith('fp8'), catching
both 'fp8' and 'fp8-dynamic' (compressed-tensors). INT4/AWQ/W8A8/bf16 never
route FP8 GEMM so they stay correctly excluded.
- agents-a1 compose: add the - VLLM_USE_DEEP_GEMM pass-through.
- test-deepgemm-fp8-parity: broadened to startswith('fp8') so it now covers
agents-a1 (5 fp8-family composes) + any future fp8-* variant.
- test-launch-compat: new case — fp8-dynamic on 5090 injects VLLM_USE_DEEP_GEMM=0.
Now ALL fp8-family vLLM slugs are 5090-safe via the launcher (dual-max,
multi-max, dual-lmcache, diffusiongemma-dual, agents-a1). Convention: fp8
weight variants must be named fp8-* for the gate/guard to catch them.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
#580 added the consumer-Blackwell DeepGEMM auto-disable pass-through to
dual-max, but the other fp8-weights vLLM composes — hand-maintained
parallel copies, no extends/include — had silently drifted without it:
multi-max, dual-lmcache, diffusiongemma-dual. On a 5090 / PRO 6000 / GB10
(sm_120/121) the launcher's _deepgemm_env injects VLLM_USE_DEEP_GEMM=0, but
with no receiving line docker-compose drops it → the fp8-GEMM 'recipe not
found' boot crash #580 fixed for dual-max (disc #571).
- Add the pass-through (+ the shared comment) to all
three. All 4 fp8-weights vLLM composes now carry it.
- NEW test-deepgemm-fp8-parity: asserts every fp8-weights vLLM compose has
the pass-through — REDs on the exact drift class that caused this (verified
it fails when the line is removed, passes when restored). Scope matches the
_deepgemm_env gate (weights_variant==fp8 + vllm engine); INT4/AWQ/bf16 never
invoke DeepGEMM so they're correctly excluded.
Removes ONE 5090 blocker (fp8 crash) — NOT a blanket 'runs on 5090': beellama
composes stay broken on sm_120 (Anbeeld#85, CUDA 12.4), llama.cpp/INT4 on
sm_120 remain unvalidated (no 5090 on the reference rig).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
@ryanmpelletier's flawless 4x 3090 report (#584) is the cross-rig
validation the compose header required ('until a real >=4x 3090 host
validates it'): v0.24.0, no power caps, all x16, verify-full 9/9,
verify-stress needles clean to 240K, soak-continuous PASS, bench n=5.
Third concordant 4-card validation (after @alanspires 6xVFIO + @Whamp),
first on v0.24.0.
- registry + compose header: status experimental -> production (both, drift
guard green). No DEFAULTS[(qwen,vllm,multi4)] entry -> no resolver cascade;
promotion just moves it onto the actionable list. Quality is TP-invariant,
carried from the vllm/dual proxy (109/150); 4-card 8-pack is the open
confirmation (report.sh runs none).
- baselines.yml: vllm/qwen-27b-multi-fast gains its 4x3090-pcie submission
(tier: submitted; decode 74.76/90.83, prefill 1288->1175, KV pool
1.77M/6.77x, NIAH 240K). Submission-only by construction - the 2-card
reference rig can never produce a local 4-card row.
- BENCHMARKS.md: ryan's row (all-x16 74.8/90.8 beats @Whamp's mixed-lane
59/74 on the same fp8-KV config -> PCIe lane width matters for TP=4).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The dual-max FP8 2x3090 decode bench (83.1/108.2, v0.24.0, 2026-06-30) lived
only in BENCHMARKS.md:122 + disc #515 (c17490565) — never inducted into
baselines.yml, so c3 / the registry-emit join showed only guybrush01's
2x5090 submission and no local bar. Surfaced while diagnosing #585.
- baselines.yml: vllm/qwen-27b-dual-max gains its primary tier: local row
(83.1/108.2 · TTFT 158 · prefill 1364->875 · 8-pack 107/150 · NIAH 240K ·
v0.24.0 = current pin -> FRESH); guybrush's 2x5090 submission preserved.
- compose header + DUAL_CARD.md: the stale '~56 TPS' probe (and the now-false
'slowest of the three' framing) -> real decode 83/108; the genuine
tradeoffs (smallest KV pool 295K/1.13x, slowest prefill/TTFT 158ms from
FP8's compute-heavy Marlin W8A16 dequant) kept. BENCHMARKS.md already
corrected; this closes the two spots that still read ~56.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Live-dogfood round 1 (Qwythos-9B-Claude-Mythos-5-1M-GGUF): the repo ships
base + MTP builds per quant (…-Q4_K_M.gguf / …-MTP-Q4_K_M.gguf); token-
keyed grouping merged them into ONE "2-part" variant with a summed, wrong
size (10.7 GiB shown for a 5.2 GiB pick) and no way to select just one.
Group by STEM instead (basename minus the -NNNNN-of-NNNNN part suffix) —
true multi-part shards share a stem so 'parts' still counts them; distinct
artifacts don't. Label = stem minus the repo-wide common prefix (standard
repos keep their plain quant token; multi-artifact repos keep the
distinguishing part: Q4_K_M vs MTP-Q4_K_M). Guard case added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
artifact_inventory(api) enumerates a repo's servable artifacts WITHOUT
gating on format — a GGUF-only repo is a first-class bring (design §2b-1/2;
select_weight_files stays the vLLM/safetensors gate). GGUF variants are
enumerated at any depth, grouped by quant token with multi-part files
summed, mmproj projectors split out (never a variant); safetensors sets
reuse the adapter-excluding filter; cardData.base_model rides along
(friction #11 — lineage for ⑤'s taxonomy + credits). inspect_repo() wraps
the API fetch with structured errors; CLI: deriver.py --inventory <repo>
--json (the c3 Bring pane's INSPECT subprocess). Offline guard test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
catalog-baseline.sh extracts runner_version from the benchlocal results
JSON and stamps quality_env: { harness: "benchlocal-cli X.Y.Z" } next to
quality_8pk — the half-deployed-sandbox lesson (§2.1.4): a quality number
without its harness fingerprint is unreproducible. sandbox_digest rides
along once benchlocal emits it (upstream candidate). Arms on DIFFERENT
harness versions → loud warn + field omitted (the arm delta isn't
comparable). Guard validates the shape; fixtures updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Bundle mode inducts a volunteer's rebench bundle into the slug's
submissions: map with provenance FROM THE BUNDLE — rig/power from rig.txt,
engine pin from container-config.json Config.Image — never from this rig
(nvidia-smi / resolve_variant_pin would stamp our fingerprint onto foreign
numbers). --source and --submitted-by are required, no $USER default.
- splice safety both directions: a primary re-induction preserves an
existing submissions map; bundle mode never touches the primary row
- one row per rig_class (newest replaces; history stays in git)
- multi-tag bundles refuse without --from-tag selection
- test-catalog-baseline: synthetic-bundle fixture covering refusals,
bundle-derived provenance, add/replace, submission-only entries,
splice preservation
Dogfood: first external row — guybrush01's #571 fp8w bundle lands as
vllm/qwen-27b-dual-max submissions[2x5090-pcie] (134.52/165.11 decode,
NIAH-clean 240,635 tok, v0.24.0, tier: submitted, TPS-only pending his
quality run).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Add the slice-3 trust-boundary schema to baselines.yml: every primary row
gains tier: local; an optional submissions: map (keyed by rig_class, e.g.
2x5090-pcie) carries cross-rig rows with required source + tier
submitted|reproduced. A slug may be submission-only (hardware we don't have).
- test-baselines: shared field validator across both row shapes; rig_class
key format + rig-field parity; tier enums; submission-only entries legal
- registry-emit _baseline_for: submissions ride the join with per-submission
staleness (pin comparison is rig-independent)
- c3: a submission-only baseline is NOT the bar (TPS column stays em-dash);
detail panel renders rig-labeled, tier-badged cross-rig lines, never merged
into the local bar
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Two fixes from guybrush01's fp8-weights 5090 run (disc #571), which needed both
to boot the fp8w arm on consumer Blackwell.
1) VLLM_USE_DEEP_GEMM: vLLM's DeepGEMM fp8-GEMM path is built for Hopper (sm_90)
+ datacenter Blackwell (sm_100/103). Consumer Blackwell (5090 / PRO 6000 /
GB10, sm_120/121) has no recipe and hard-fails "recipe not found" at boot.
Add _deepgemm_env: for fp8-weights slugs, inject VLLM_USE_DEEP_GEMM=0 on the
consumer SMs (sm_120/121 confirmed-broken; sm_89 Ada added proactively — it
routes fp8 via Marlin/CUTLASS so disabling is a harmless no-op that pre-empts
the same wall for 4090 owners). Hopper/datacenter untouched. dual/fp8/mtp.yml
gains a pass-through env; both launchers whitelist the export.
2) arch-ab.sh fp8w arm: add --force. vllm/qwen-27b-dual-max is status=experimental
so switch.sh gates it without --force.
test-launch-compat locks 5090/Ada-down / Hopper-keep / non-fp8-skip / user-pin.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The concurrency envelope spends the KV pool, but nothing sized the pool per
card: the composes default to --gpu-memory-utilization 0.92 and the launcher
never adjusted it. For discrete cards that's fine-to-conservative (their
mem_util_safe is 0.95-0.96, above the default). For DGX Spark it's a real
safety hole: its 128 GB is unified LPDDR5X shared with the Grace CPU/OS, so
mem_util_safe is 0.85 — booting at 0.92 would take ~118 GB and starve the OS.
The seeded Spark envelope row already assumes 0.85; the launch would wrongly
use 0.92.
Add _mem_util_env (same seam as _envelope_env): inject GPU_MEMORY_UTILIZATION
DOWNWARD only — when a detected card's mem_util_safe is below the compose's
registry mem_util. Heterogeneous rigs clamp to the lowest ceiling (one GMU
applies across all ranks; a unified-memory card forces the rig down). Today
this fires for exactly one card: Spark -> 0.85.
It deliberately NEVER raises above the tested default. A 3090/5090 could give
0.95-0.96, but that changes the validated Cliff-2b margin and risks boot-OOM,
so the upward move stays a validated opt-in on the soak protocol, not an
automatic bump (and the big-card concurrency ceiling is bandwidth-bound anyway,
so the extra pool mostly buys context headroom).
Both launchers whitelist the new export; test-launch-compat locks Spark-down /
discrete-no-raise / het-min / user-pin-wins.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Enhances concurrency-probe.sh into the tool that upgrades a `computed` envelope
row to `validated` (see /opt/ai/docs/phase2-soak-validation-protocol.md). All
additive — the plain fit-check stays the default behaviour.
The design's core is two row classes, two bars, and this probe serves both:
• pool-ceiling rows (5090): VALIDATE=1 fills each stream to the served
--max-model-len (or TARGET_CTX), runs 6 rounds, gates fit + >=98% TPS
retention. The value IS the kv-calc ceiling, so this is a fit+stability test.
• bandwidth-cap rows (PRO 6000 / Spark): SWEEP="4 8 12" SLUG=... TPS_FLOOR=15
reboots per N (vLLM can't hot-change max-num-seqs), probes decode-dominated,
and prints the throughput KNEE — the largest clean N whose per-stream decode
TPS clears the floor. A fit test is useless here (N=8 trivially fits 96 GB).
Key addition: streamed per-stream DECODE tok/s (TTFT-separated) — the only
honest throughput number at deep context, where prefill dominates wall time.
Retention drops the round-1 cudagraph warmup once >=4 rounds so it isn't
flattered. Machine-readable RESULT line drives the sweep's knee-finding.
SWEEP needs SLUG (refuses with exit 2 otherwise); SWEEP_DRY=1 plans the reboots
without booting. New test-concurrency-probe.sh guards syntax + refusal + dry-plan
(offline / CI-safe). Live-validated on the dev-rig Qwen: streaming TPS, floor
gate (FAIL on high floor), and VALIDATE fill all behaved.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Follow-up to #576. Adds the two >24 GB rows deferred there and fixes the GPU
detector wrinkle that made them impossible to reach.
Detector (launch_compat.py): the `sm >= 12 -> rtx-5090` catch-all collapsed
every Blackwell to a 32 GB 5090 — a 96 GB PRO 6000 and a 128 GB GB10 both
mis-detected. Split it by SM then VRAM: sm_121 -> dgx-spark, sm_120 + >=64 GB
-> rtx-6000-pro-blackwell, else rtx-5090. Also fixed the PRO 6000 name alias
(the shipped "6000 pro blackwell" never matched the real "RTX PRO 6000
Blackwell" word order; now "pro 6000", which does NOT catch the sm_89 RTX 6000
Ada). Locked with 5 detector regression cases.
New hardware profile: dgx-spark.yml (GB10, sm_12.1, 128 GB unified LPDDR5X,
conservative mem_util for CPU-shared memory, no nvfp4 KV — same FMHA gap as
consumer Blackwell).
Rows (all `computed`, live-verified injecting):
vllm/dual @ rtx-6000-pro-blackwell -> 8 (2x96 GB, NVLink)
vllm/minimal @ rtx-6000-pro-blackwell -> 16 (1x96 GB)
vllm/minimal @ dgx-spark -> 8 (1x128 GB)
These are OPERATIONAL CAPS, not raw pool ceilings: kv-calc gives >56 (PRO 6000
minimal) and >80 (Spark minimal), but decode concurrency is BANDWIDTH-bound
there, not pool-bound. PRO 6000 (~1.8 TB/s GDDR7 ≈ a 5090) lifts modestly above
the 5090; Spark (~273 GB/s, ≈1/7) does not — its huge pool buys context
headroom, not streams. Each row's basis cites both numbers. Spark-dual is
unseeded (2-unit clustering is ConnectX/RDMA, network-bound for TP). A soak on
real hardware refines each and upgrades it to validated.
test-profiles-compat: hardware count 9 -> 10 (dgx-spark).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Previously a mixed-card rig no-op'd the concurrency injection ("no single
card-class row applies"). But it does have a right answer: vLLM already sizes
the KV pool to min(free blocks) across TP ranks (the cache is symmetric-
sharded), so the smallest card dictates the pool. Clamping the envelope lookup
to the smallest-VRAM card therefore MIRRORS the engine — it's the real ceiling,
not a conservative guess — and stays safe for TP=1 too (the ceiling fits
whichever single card vLLM lands on, all >= the smallest).
Behaviour:
5090 + 3090 -> smallest 3090 has no row -> compose default (unchanged)
5090 + H100 -> smallest 5090 is seeded -> inject its ceiling (was: no-op)
2x 5090 -> homogeneous -> unchanged
The common 5090+3090 case still lands on the compose default (now for the
principled reason: the 24 GB card caps the pool), so nothing regresses; the
gain is heterogeneous rigs whose SMALLEST card is itself a seeded >24 GB class.
Fixtures lock both the new inject-on-smallest and the still-no-op paths.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Populate envelopes.yml with the first memory-envelope rows, computed (not
guessed) from the kv-calc pool ceiling at model-max context:
vllm/dual @ rtx-5090 (2x32 GB): 2 -> 4 concurrent full-262K sessions
vllm/minimal @ rtx-5090 (1x32 GB): 1 -> 9 concurrent full-32K sessions
Both are the no-preemption ceiling: N *full-context* sequences whose KV
blocks fit the pool. Concurrency is a capacity question kv-calc answers
deterministically (arithmetic, arch-independent), and max_num_seqs is a cap
with graceful preemption (never OOM), so an over-optimistic ceiling costs an
occasional preempt, not a crash. kv-calc is sm_120-calibrated (disc #571
paulp83's real-5090 verify-stress PASS), so a 32 GB projection is trusted
arithmetic. A concurrency-soak upgrades a `computed` row to `validated`.
Guard: test-envelopes now accepts a `computed` (kv-calc basis) block as
provenance alongside `validated` (soak) — a computed row must name its basis
(invocation + PASS/cap boundary). Added a fixture proving injection is
provenance-agnostic (the launcher reads max_num_seqs; only the guard cares
about provenance).
Multi-GPU: the seam already fires on homogeneous >2-card rigs (collapses N
identical cards to one class; verified vllm/dual injects on a 4x5090 spec) and
kv-calc computes TP=4 pools; the experimental TP=4 slugs are documented as
computable-but-deferred rather than seeded. Heterogeneous rigs no-op by design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Found testing vibethinker-3b (a reasoning model): with a small max_tokens
it spends the budget mid-<think> and returns HTTP 200 with 256 tokens but
EMPTY content (never emits the final answer). The old content-based check
wrongly flagged those streams as silent-empty failures. For a KV-pool
stress test the stream DID run (generated tokens, held KV) -> ok =
completion_tokens > 0; a true silent-empty is HTTP 200 with ZERO tokens
(what soak-test means). After the fix vibethinker sustains N=8 clean on a
3090 (8x the compose default, 0 post-warm growth).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The #246 Phase 2 first pass — spend a bigger card's KV pool on
concurrency, weights-invariant. Ships the machinery + measurement; the
32 GB envelope rows land from volunteer probe runs (baselines-style
plumbing-first).
Why concurrency-only (agreed 2026-07-05, design doc): most headline
configs are already at 262K model-max on 24 GB (GGUF single, dual vLLM),
so the extra 32 GB pool serves MORE STREAMS, not more context. The
context lever (bigger max_model_len) only helps single-card vLLM and is
Cliff-2b-capped anyway — deferred.
- Composes: `--max-num-seqs` parametrized to `${MAX_NUM_SEQS:-N}` on the
two pilot slugs (dual=2, minimal=1). Weights-invariant, default kept.
- envelopes.yml (empty): measured per-(slug, card-class) `max_num_seqs`,
born-from-measurement (a `validated` block is required — no guesses).
- launch_compat.py `_envelope_env`: injects MAX_NUM_SEQS from a row when
it EXCEEDS the compose default; 24 GB / no-row / heterogeneous /
user-env-set -> no-op. Both launchers whitelist MAX_NUM_SEQS.
- test-envelopes: schema + injection contract (inject / no-row /
heterogeneous / user-env-wins / no-gain).
- concurrency-probe.sh: the measurement — N concurrent long streams x R
rounds, separating expected pool-fill from a real leak (post-warm
growth). Live-validated N=2 on the dual-3090 (0 MB post-warm growth,
all rounds clean — the shipped default is sound).
Reuses the Phase 1 injection seam; independent of the KV verdict.
Full scripts gate 67/67.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Two 5090 owners hit `--kv-cache-dtype nvfp4 requires sm100f` crashing
mid-boot on the #246 A/B (disc #571). Root-caused from vLLM #43562 /
TRT-LLM #10241: nvfp4 KV forces the trtllm-gen FP4 FMHA, built ONLY for
datacenter Blackwell sm_100/sm_103. Consumer Blackwell (sm_120/121 —
RTX 5090 / PRO 6000 Blackwell) is a HIGHER cc number but a different
family with no FMHA build. NVFP4 *weights* work there; only the KV path
doesn't. Our #246 gate had used a numeric ">=10.0" floor that wrongly
passed sm_120 — a floor can't express "sm_100/103 but not the
numerically-higher sm_120".
- gates.py: new `_ARCH_KERNEL_SM_FAMILY` allowlist ({nvfp4: sm_100/103});
dropped nvfp4 from the numeric `_ARCH_KERNEL_SM` floor; family-membership
reject with the FMHA reason + fp8_e4m3 fallback.
- arch-ab.sh: nvfp4 arm now refuses on consumer Blackwell (not just
<sm_10), naming the FMHA gap + the fp8_e4m3 path; dropped nvfp4 from the
recommended arms in help.
- hardware profiles: removed nvfp4 from rtx-5090 / rtx-6000-pro-blackwell
KV lists (both sm_120) + added a why-not note.
- docs (DTYPE_MATRIX / HARDWARE / KV_MATH / QUANTIZATION): corrected the
"Blackwell sm >= 10.0" framing to "datacenter sm_100/103 only".
- UPSTREAM.md: #43562 / TRT-LLM #10241 row + re-test trigger.
- test-arch-ab: nvfp4 refuses on sm_86 AND sm_120, allowed on sm_100;
the dual-5090 all-arms test drops nvfp4.
Full scripts gate 66/66.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
report.sh already printed both raw inputs (host capability sections +
the [nvlink] engagement trail added after #446/#488); what was missing
was the CROSS-REFERENCE. New scripts/lib/p2p-state.sh implements the
verdict matrix once, consumed by report.sh (full verdict line) and
preflight.sh (capability one-liner):
- <2 GPUs / no capability -> SILENT (stock-PCIe owners never nagged;
that's ~95% of dual-3090 rigs and us)
- capability + engaged -> one OK line
- NVLink bridge + P2P off -> WARN (bridge idle, ~15% decode on the
table per the #77 controlled A/B; names
the fix: launcher boot / force_on)
- P2P-capable driver + off -> INFO (launcher auto-engages since #291;
residual case = direct docker compose)
The lib is the read-only AUDITOR; scripts/detect_nvlink.sh stays the
boot-time DECIDER. Their capability probes mirror each other by design,
and test-p2p-state runs BOTH against shared faked-nvidia-smi fixtures
and asserts agreement, so they cannot drift apart silently. Engagement
classification is pure (stdin text: boot trail beats env fallback).
Live-validated on this rig (stock-PCIe dual): preflight and report both
correctly silent; verdict matrix + classifier + probes covered
hermetically. Full scripts gate 66/66.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Everything the #246 runner dogfood on this rig surfaced, in one PR:
- verify-stress STRESS_FAST=1 (opt-in; gates keep full mode): skips
probe 7's large fresh needles (near-duplicate the ladder's depths)
and caps the ceiling ladder at 2 rungs (~2/3 anchor + ceiling).
MEASURED on the identical vllm/dual boot: 861s -> 361s (-58%).
Depth coverage stays 4-anchor (bench probe 10K/90K + mid + ceiling).
- VRAM margin attribution: compose `count: all` (DeviceRequests
Count:-1) and VISIBLE_DEVICES=all now resolve as a DETERMINED
"all host GPUs" -- kills the spurious per-rung WARN on dual rigs.
- VRAM margin aggregation: MIN per-card free instead of sum. TP OOMs
on the first card to run out; the sum overstated dual margins ~2x
(906 MB reported where the honest figure was 453 MB). SEMANTIC
TIGHTENING: multi-GPU margin advisories now fire against the real
per-card floor. Ceiling test updated to the new contract.
- arch-ab: lean tier exports STRESS_FAST=1 (per-arm ~30 -> ~20 min);
margin advisory surfaced in the summary table + a mid-run "ladder
CLEAN, not a failure" note so first-time runners don't abort on the
rc=1 advisory; bundle print points at the #246 test thread.
Full scripts gate 65/65. The e4m3-below-sm_8.9 refusal is the prior
commit on this branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
vLLM hard-rejects fp8_e4m3 KV below SM 8.9 at boot -- without the guard
a bare run on an Ampere rig (default arms include e4m3) sits in
switch.sh's ready-wait until the 10-minute timeout instead of getting a
clear refusal. The message says why: on Ampere there is no arch delta
to measure. --arms e5m2 (control-only) stays possible.
test-arch-ab: bare-run-on-3090 refusal + control-only-on-3090 pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
report.sh --full-calibration instead of the bare default: seconds, no
test re-runs, and the predicted-vs-actual VRAM matrix on the volunteer
card class is direct Phase 2 input (32 GB envelope calibration).
Deliberately NOT --full -- that re-runs the whole battery (~43 min)
against only the last arm's container; the arms already carry that data.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
report.sh default mode (~2 s, path/host/user redaction ON) rides in the
tarball — PCIe lane width / driver / power caps / WSL-vs-bare-metal is
the context that makes cross-rig A/B variance interpretable (the same
mechanism that auto-flagged the x4-lane slot in #158).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
One command per volunteer rig: runs the pilot variant once per KV-dtype
arm (fresh symmetric boot each), verify-full + bench n=5 + NIAH ladder
per arm -- soak and quality packs deliberately skipped (quality
consolidates at promotion time per the #246 test plan; ~25-30 min/arm).
Ends with a per-arm comparison table and ONE attachable tarball.
Arms: e5m2 (control) / e4m3 (native FP8 KV, sm_89+) / nvfp4 (Blackwell
opt-in, refused below sm 10.0) / fp8w (vllm/qwen-27b-dual-max STOCK --
FP8-weights checkpoints reject fp8 KV, so this arm measures the
native-FP8-WEIGHTS lift instead; dual-rig only).
Guard rails: variant auto-pick by GPU count (2+ -> vllm/dual, 1 ->
vllm/minimal); KV arms are explicit KV_CACHE_DTYPE pins (unambiguous on
any launcher version); the running container's --kv-cache-dtype is
asserted live before any bench time is spent; --dry-run / --resume /
--full escape hatch.
test-arch-ab: hermetic plan/refusal contract via CLUB3090_FAKE_GPUS.
Full scripts gate 65/65.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
launch.sh/switch.sh now detect GPU arch and export KV_CACHE_DTYPE=
fp8_e4m3 (native FP8 compute) on sm_89+ cards for the pilot slugs
(vllm/dual, vllm/minimal). The injected value is the hardware
profiles' dormant kv_format_default.balanced -- one source of truth
shared with the pull gates; the Ampere no-op is data equality
(3090-class balanced = fp8_e5m2 = compose default -> nothing emitted),
not a code branch.
Injection guards (all load-bearing):
- pilot allowlist only; expansion gated on the #246 cross-rig A/B
- per-variant: registry kv_format == fp8_e5m2 only (int8-PTH/TQ/bf16
slugs never touched -- compressed-tensors weights reject fp8 KV)
- vLLM-family variants only; explicit user KV_CACHE_DTYPE= wins;
unmapped cards / heterogeneous rigs / no nvidia-smi -> no injection
- direct `docker compose up` keeps Ampere-safe compose defaults
Rides the existing resolve-variant-pin export seam (new optional
--gpu-spec); VLLM_ATTENTION_BACKEND is whitelisted but ships no value
(vLLM auto-detect stays the default until measured). Preflight banner
names the detected arch class.
Consistency fixes the injection exposed:
- gates.py _ARCH_KERNEL_SM: fp8_e4m3 9.0 -> 8.9 (vLLM's real floor is
SM89+; we'd otherwise inject e4m3 on 4090s our own pull-gate calls
unloadable) + new nvfp4: 10.0 row (v0.24.0 literal had NO gate --
a 3090 pull of an nvfp4 config wouldn't have been rejected)
- Blackwell hardware profiles: nvfp4 declared as CANDIDATE capability
(engine list unchanged until validated -- gates take the intersection)
- kv-calc: projected nvfp4 rows (bytes/elem + activation coefs),
calibration unchanged
Docs: HARDWARE.md new section, KV_MATH/QUANTIZATION/DTYPE_MATRIX rows
(incl. retiring the Genesis-era "e4m3 undertuned per #51" advisory in
favor of the A/B).
Validated: 8-case injection matrix in test-launch-compat; live no-op
on the real 2x3090 (spec built, nothing injected); faked 4090 through
the real bash detection path emits e4m3; compose interpolation both
ways; kv-calc --calibration green; full scripts gate 64/64.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The single-card Gemma-4 default was the catalog's only default with no
baseline row (wave-2 PRIORITY gap). Full gate on Anbeeld's OFFICIAL
v0.3.2-preview digest (the current beellama-local pin), rebench tag
gemma-dflash-regate:
- decode 44.91 narr / 79.76 code (n=5, CV 2.1%/5.6%), TTFT 115 ms
- NIAH ladder clean to 117,513 tok (91% of 128K); ceiling margin 929 MB
(< 1024 MB bar -> caveat stays, but up from ~177 MB pre-v0.3.2)
- 8-pack current harness: 108/150 off / 113/150 on (+5 in-band wash,
think-off default re-confirmed)
- soak PASS (p50 59.8, 100.6% retention, 0 growth, 0/100 silent-empty)
Row seeded via catalog-baseline.sh (first non-vLLM engine through the
producer flow); BENCHMARKS row + compose header (Quality + margin
numbers) updated to match.
Tooling fallout the gate exposed:
- rebench-full.sh hardcoded RUNS:-3/WARMUPS:-1, silently under-running
the bench protocol on every orchestrated gate. The A1 "n=3 env leak"
attribution was wrong -- it was this line (baselines comment
corrected). Fixed to protocol 5/3; n=3 had flattered gemma-dflash
code TPS by +8%.
- catalog-baseline.sh counted prefill-probe run-lines into its bench-n
gate (read n=8 for an n=5 log); now reads the summary headers.
Full scripts gate 64/64.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Reviewed dispositions of the wave-1 gap list (verdict sheet approved
2026-07-04):
- vllm/dual + vllm/qwen-27b-dual-fast (alias mirror): FRESH from the
2026-06-30 v0.24.0 pin-bump gate row (decode 70.7/93.5, NIAH->240K,
soak PASS) — measured on the current vllm-stable pin, so no
born-stale archaeology needed after all.
- ik-llama/iq4ks-mtp-vision: carried decode (60.39/72.4) from iq4ks-mtp
per the 2026-05-25 vision re-tune row; carried qualifies because the
vision delta doesn't touch the decode path.
- ik-llama/apex-fit-q8q5: 2026-05-28 row (decode 105.63/156.80, n=5),
NIAH clean@180K.
- vllm/gemma-26ba4b-single: 2026-06-06 row (decode 169.2/219.9) + full
8-pack (98 off / 109 think-on), NIAH clean@161K — FRESH (gemma-stable
still pins v0.22.0).
- vllm/agents-a1-dual: decode pair order corrected from the tag
artifact (narr 154.0 / code 153.8 — wave-1 had them swapped).
- Footer: SEED WAVE 2 list -> KNOWN GAPS dispositions (vllm/minimal,
gemma-12b pair, beellama/gemma-dflash PRIORITY re-gate, deckard).
catalog-baseline.sh: insertion anchor updated for the retitled footer
(accepts both titles); test-catalog-baseline.sh now guarantees the ADD
path (strips the seeded vllm/dual row from its copy) and asserts added
rows land before the footer — the marker-drift class this bump caught.
Guards: test-baselines 15 rows joined, same 4 wave-1 stale WARNs, all
5 new rows fresh. Full scripts gate 64/64.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Closes the rescore-materialization gap the induction tool's first live
run surfaced (#562): the benchlocal-cli 'rescore' subcommand ALREADY
supports write-back (--in-place / --output) — the gap was practice, not
a missing upstream feature. The T2 rescore ran to stdout only, so the
published A1 thinking-on 110/150 diverged from the tag artifact (108).
- A1's rescore MATERIALIZED into the tag JSONs (rescore --in-place; OFF
105 unchanged, ON 108→110) — induction now extracts 110 and the
regenerated corpus record carries 110; artifact, corpus, baseline row
and publication all agree. Pre-rescore JSON backed up outside the tag.
- docs/QUALITY_TEST.md: 'Rescoring saved results — MATERIALIZE, don't
just read' — the rule (a rescore that changes a published number must
be written back in the same session), the command, what rescore can't
re-run (sandbox packs), and the A1 case as the cautionary example.
- catalog-baseline.sh header: reads-artifacts-as-truth warning + the
header's gate text updated to the actual n>=3/warn-under-5 behavior.
- baselines.yml A1 comment: gap note -> materialized + doc pointer.
Guards green (test-catalog-baseline, test-baselines).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Decode-only TPS can't express real trade-offs anymore (the W8A8-vs-FP8
result was a prefill-corner-vs-decode-corner split) — this adds the
CANONICAL prefill/TTFT measurement to bench.sh per the design's sourcing
rule, with the protocol gotchas productized:
- bench.sh PREFILL PROBE (default-on; PREFILL_PROBE=0 / PREFILL_DEPTHS /
PREFILL_RUNS): warm + n measured per depth (10K + 90K anchors; 90K is
inside the DeltaNet degradation regime and pairs with the NIAH ladder's
~94K rung). CACHE-BUSTED: fresh salted haystack per request — composes
serve enable_prefix_caching, an identical prompt re-measures the CACHE
HIT (vLLM's prefix cache is block-chained; unique first line breaks the
chain). SELF-CALIBRATING: word-count heuristics overshoot tokens ~1.3x;
the warmup's reported prompt_toks scales the measured runs (target^2/
actual) to within ~4% of the requested depth. Depths exceeding the
served ctx SKIP with a note. DUAL METRIC, labeled: prompt_tokens/TTFT =
client-observed (user-truth: incl tokenization+transfer+scheduling) AND
the vLLM stats-log windowed rate = engine-internal (compute-truth) —
at 93K on A1 they differ by ~7s of non-prefill overhead (5.5K vs ~10K
t/s); never cross-compare kinds (stack LEARNINGS row added).
- measurement_record parser: per-block pass -> prefill_tps_by_ctx +
ttft_ms_by_ctx extensions; the canonical short-prompt ttft_s is
PROTECTED from the probe blocks (the old last-occurrence rule would
have swallowed the 90K block's 17s TTFT).
- catalog-baseline.sh: rows gain prefill_tps {10k: N, 90k: M} (parsed
via THE record parser, no second grammar) + ANCHOR CALIBRATION at
induction: the probe's deep anchor vs the NIAH ladder's nearest rung —
agreement (0.7-1.3) certifies the ladder's whole depth curve; A1 live:
probe 5459-5584 t/s @93K vs ladder 7403 @94K = ratio 0.74-0.75, OK.
Divergence warns with an investigate message (design: a finding).
- test-baselines schema: prefill_tps = dict of numeric depth points.
test-catalog-baseline fixture: probe blocks + TTFT-pollution guard +
anchor-OK assertion.
Live-validated 3x against the serving A1 (262K): 10K = 7977 t/s CV 1.1%
TTFT 1.25s; 93K = 5584 t/s CV 0.5% TTFT 16.1s; engine-log ~10K t/s.
Full scripts gate green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The consumer half of the producer wiring: THIS RIG's corpus records now
render next to the shipped bar in the catalog preview — the design's
'c3 overlays the bar with your numbers, badged'.
- LocalMeasured projection + CatalogEntry.local_measurement: the newest
#249 corpus record per variant slug (decode point, both 8pk arms, pin,
date). One canonical-short decode point per record — the narr/code
pair split is a parser note for 2c.
- services.local_measurements(): pure-fs corpus scan; newest-per-slug by
the record's _recorded_at stamp (NEW — rebench-full now stamps it at
write; mtime+line-order fallback covers pre-stamp records); malformed
lines skipped. Joins in the same enrichment pass — still zero
subprocess on the catalog path.
- Preview rework: the measured line is the BAR with provenance
(date · rig · submitted_by from the joined baseline); a stale bar gets
an explicit detail line ('† bar measured on <pin> — current <pin>;
re-bench owed'); a slug this rig has gated adds
'yours ~<decode> · 8pk <P/150> (<date> · this rig)'.
Live-verified: the corpus's first record (A1, from 2a's live run) joins
and renders 'yours ~154 decode · 8pk 105/150'. Tests: overlay semantics
(newest-wins by stamp, malformed-skip, entry join, no-record → None) +
the three-state preview (fresh bar + provenance + yours · stale pin
detail · no-yours). c3 suite 766/766; scripts gate green serial.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The producer wiring: a validated gate run now flows into the corpus
automatically and into the shipped bar via one reviewed command.
- scripts/catalog-baseline.sh (§2.3, one mechanism / three entry points):
gate-validates the tag dir (verify-full pass HARD · bench n>=3 HARD,
n<5 WARN vs the canonical target · quality present unless --tps-only),
extracts decode TPS/TTFT/8pk both arms/ctx_validated (NIAH-ladder
verdict incl. token count)/provenance (pin via the same resolution the
emit join uses; rig+power via nvidia-smi, all overridable), UPSERTS the
baselines.yml row textually (comments preserved), prints a unified
diff. Never commits — rows ship via PR. --baselines-file for tests.
- rebench-full.sh: on completion, EXACT-container-match the served
engine to a registry slug (identity semantics — never port/substring,
the F9 rule) and append a fingerprint-complete #249 record (engine_pin
via resolve_variant_pin/compose fallback · hardware · power-cap ·
quality_8pk extensions · soak status) to the gitignored corpus, then
print the catalog-baseline.sh induction prompt. BYO/swap serves skip
with a note (no registry identity to record against).
- measurement_record.py: --quality-8pk / --quality-8pk-think-on ->
measured_extensions (the designed extension namespace; frozen schema
untouched).
- test-catalog-baseline.sh: hermetic (synthetic tag dir + a COPY of
baselines.yml) — extraction, dry-run no-write, add->replace upsert,
refusals (missing verify/quality/unknown slug), real-file checksum
guard.
TRUTH CORRECTIONS the tool's first live run surfaced (run against the
real agents-a1-fp8-dual tag): the A1 gate's bench was n=3 (a RUNS=3 env
leak into the overnight gate — 'n=5' was misstated in publications), the
BENCHMARKS TPS pair was order-garbled (artifact: narrative 154.0 / code
153.8), and the #81-rescored think-on total (110) was never materialized
back into the tag JSON (reads 108) — materialization gap documented in
the row comment. baselines.yml + BENCHMARKS.md corrected.
Live-validated: rebench record block ran against the serving A1 → first
corpus record written (vllm-agents-a1-dual__e9d28e22.jsonl: pin v0.24.0,
rtx-3090, 370W, quality extensions); induction dry-run reproduces the
row from the tag. Deferred to 2b/2c: c3 local-overlay + richer badges;
the bench 10K/90K prefill/TTFT probe. Gates: test-catalog-baseline +
test-measurement-record + test-baselines green; full scripts gate green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The catalog's measured columns (TPS / 8pk) now come from ONE productized
source: scripts/lib/profiles/baselines.yml — the shipped, PR-reviewed
'bar' (accepted display projection of a validated gate run, with pin/rig/
power provenance) — joined per-slug into registry-emit --json with an
emit-computed staleness verdict. Consumers never read baselines.yml or
BENCHMARKS.md directly (BENCHMARKS stays the public human/cross-rig
ledger). Design: the catalog-baselines note (2026-07-02, §2/§5 slice 1).
- baselines.yml SEED WAVE 1: 10 rows with airtight traceability only
(rebench tags on disk: agents-a1 golden specimen + gemma-31b-dual;
unambiguous decode-class BENCHMARKS rows for the rest). 4 rows are
HONESTLY born-stale with documented reasons (llamacpp rolling-tag-era
x2, beellama pre-#296 image, 35B v0.22.0) — the guard's demo cases.
9 slugs still owed rows are listed as wave-2 gaps, no guessed numbers.
- registry-emit join: per-variant 'baseline' field; current pin resolved
the way launchers actually resolve it (engine-profile install.spec,
compose-image-default fallback for ik/llama.cpp) → 'stale' =
measured-pin != current-pin, null when undeterminable.
- test-baselines.sh: schema + slug-membership + ctx-parity (compose ctx
default == registry max_ctx, functional slugs) RED; pin-staleness
WARN-only (pin bumps must not block on immediate re-bench — the debt
stays visible). test-registry-json gains the 'baseline' contract key.
- c3: enrich_measurements = pure in-memory map off the joined field —
deletes BOTH the per-slug --explain fan-out (~4s/slug, the #439
option-3 leg) and the BENCHMARKS.md scrape from the catalog path
(Explain modal + cross-rig explorer keep their readers). Stale rows
render a † on the TPS cell + a status-line legend; full badge/overlay
treatment is slice 2. Measurement gains source='baseline' + stale.
- results/baselines/README: the three-store relationship (regression
corpus vs measurement records vs display bar) + the no-drift rule the
slice-2 induction tool enforces.
Verified: guard test green (10 rows, 4 stale-warned); live join emits
57 variants / 10 baselines; live c3 catalog shows all 10 with daggers on
exactly the stale 4 (A1 serving during the check: 154/154 - 105/150);
c3 suite 764/764; full scripts gate green serial. (First gate pass
tripped classifier/dedup by running CONCURRENTLY with the c3 suite —
.pull-captures pollution class; serial rerun clean.)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Investigation of a Discord cross-rig report (WSL, 3060Ti+4090: verify
fails on the pinned v0.3.2-preview digest, passes on the old noonghunna
snapshot; NOT reproducible on our 3090s — verify-full all-pass) surfaced
three doc-level truths worth recording:
- The 50-series gap is a TOOLCHAIN ceiling: Anbeeld's CI builds on the
Dockerfile-default CUDA 12.4, whose nvcc cannot target sm_120 (no
cubin, max PTX compute_90) — every official tag lacks Blackwell.
Upstream ask filed: Anbeeld#85 (CUDA_VERSION 12.8.1 + arch list incl
120); when it lands we retire the noonghunna snapshot entirely.
- Three registry status_notes claimed launchers inject 'server-cuda-
v0.3.0' — stale since the 2026-06-12 pin bump; they inject the
v0.3.2-preview digest from engines/beellama-local.yml install.spec.
Fixed all three (qwen dflash, gemma-12b, gemma dflash) + honest
labeling of the noonghunna snapshot as v0.3.0-feature-level and
unmaintained (predates KVarN + v0.3.1 fixes).
- The engine-notes self-build guidance was outdated: FA_ALL_QUANTS is
hardcoded in Anbeeld's cuda.Dockerfile since our PR Anbeeld#48, so a
self-build needs only CUDA_DOCKER_ARCH (+ CUDA_VERSION=12.8.1 for
sm_120). Recipe verified against his master Dockerfile 2026-07-04.
UPSTREAM.md beellama row updated (dated entry + next-triggers; the 'no
official image' claim struck through as historical). Gates: YAML +
registry import clean; status-drift / profiles-compat / switch-parity
green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The A1 compose shipped (#546) without the 'Hardware metadata' block every
sibling carries, so preflight_compose_gpu_fit logged 'WARN: compose has
no hardware metadata; allowing boot' and skipped enforcement — visible in
guybrush01's #548 report. Add the block (min-vram 24 / gpu-count 2 /
TP 2 / Requires-sm 8.0+ — Marlin FP8 weight-only needs sm_80, unlike the
sibling's INT4 7.5+); preflight now parses + enforces (live rc=0).
Also propagate the #548 finding as Caveats (4): Blackwell sm_120 hits an
upstream vLLM v0.24.0 kernel-selection bug at boot ('QKVParallelLinear'
has no attribute 'workspace' — Cutlass W8A8 selected, a Marlin-path
attribute expected). Workaround shipped as a commented env line
(VLLM_TEST_FORCE_FP8_MARLIN=1 — forces the exact weight-only path this
compose's gate validated; verified present in the v0.24.0 image) +
mirrored into the registry status_note so switch.sh's NOTE carries it.
Gates: compose_meta_get parses all 4 keys; docker compose config valid;
test-compose-status-drift / registry-disk / preflight-gpu-fit green;
full scripts gate green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
F2 — the Catalog preview strip clipped a WRAPPING caveat line: border
eats 2 rows of max-height 6 → 4 content lines; the dual-fast caveat
wrapped past that and lost its tail. max-height 6→8 + overflow-y auto
(height stays auto — short previews don't grow). Regression test pins a
wrapped-caveat entry to its full height with the tail visible.
F5 — the #544 deliberate deferrals:
- wait_gpu_vram_settle wired into ALL model/studio scene handlers
(27b, 35b-a3b, gemma-12b, deckard, ai-studio) — was gemma-int8 only;
every scene switch now lets the torn-down scene's VRAM release before
the next boot (#535 class).
- mode_off gains an engine-prefix CATCH-ALL: the enumerated stop_*
lists cover gpu-mode scenes, but a catalog-launched engine
(switch.sh <slug>) survived 'off' — caught LIVE during validation
when off left vllm-qwen36-27b-minimal serving and the 27b TP=2 scene
booted straight into its residue (the exact #535 failure). Any
remaining vllm-/llama-cpp-/ik-llama-/sglang-/beellama- container is
now stopped, with a named notice.
- c3 preflight-error visibility (the third residual): verified
already-plumbed — switch.sh's #544 refusal exits fast, its
[preflight] ERROR lines stream into the serve pane, serve_failed
stamps ✗ + [!] capture (pinned by existing tests). No change needed.
Validation: bash -n + full scripts gate green (by exit code); c3 suite
763/763; LIVE scene cycle off → 27b → off on the rig — boot ready
(qwen3.6-27b, 21.5 GiB/card), fixed off left ZERO engine containers
(verified with a non-enumerated probe container that the catch-all
stopped), both GPUs at 1 MiB after.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
estate.yml is a desired-state PLAN, but report-state reported its
instances as 'active' without checking docker — so every consumer
treated a leftover ~/.club3090/estate.yml as live GPU claims. The T1.1
audit hit this on an EMPTY rig: the c3 serve confirm warned '⚠ Starting
this will STOP estate llama-gpu0 (GPU [0]), estate llama-gpu1 (GPU [1])'
with nothing running. Scary + wrong for a fresh user.
estate_cli (the layer that owns estate liveness — fix here, not a c3
filter):
- report-state --json: each instance gains 'container' (the
club3090-<name> it would own) + 'running' (docker probe: true / false
/ null when docker itself is unavailable), payload gains
'running_count'. Additive — existing keys unchanged.
- report-state text view: per instance '— running' / '— down (plan
only)' / '— liveness unknown'.
c3 reconcile gate (consumer): an estate instance is a claim unless
running == false. true / null / MISSING (older CLI output) all stay
claims — the dual-writer gate fails CLOSED on unknown liveness, and a
mid-boot instance (running container, VRAM not yet allocated) still
conflicts.
Tests: estate JSON contract test extended (running/container/
running_count; running asserted 'in (False, None)' so the test passes
with or without docker); 3 new c3 gate tests (stale plan on empty rig →
SAFE; running:true → claim even with idle GPUs; running:null → fail
closed; the legacy no-key fixtures double as the missing-key case).
Scripts gate all green; c3 suite 758/758. Live repro on this rig (which
has the exact stale estate.yml): running_count 0, reconcile now reports
estate_claims=[] while keeping the TRUE conflict (the actually-serving
vllm container).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The T1 audit's top adoption gap: even while a model SERVES, the best
endpoint info anywhere was a bare port — no scheme/host/path, no auth
note, no copy. "Point your agent at this URL" is the whole reason the
stack exists, and the URL was underivable from the UI.
One shared derivation (layer rule; #512 precedence env/.env ->
c3_lan_ip subshell -> localhost), four surfaces:
- switch.sh: ready-line on successful boot —
"▶ API: http://<lan>:<port>/v1 (model: <served-id> · OpenAI-compatible
· no auth)" — CLI parity; served id probed from /v1/models.
- services.lan_ip(): c3 consumes the SAME derivation, session-cached.
- Orchestration serving card: bare ":8020" -> full URL + "no auth ·
[u] copy".
- Estate rail: compact "api :8020/v1 · [u] copy" (rail is ~30 cols; the
full URL lives on the card).
- NEW [u] copy-API-URL key on every Run & Operate tab (kept separate
from [Y] so row-copy semantics are untouched); honest notify no-op
when nothing serves.
F3b (readiness lag): api_booting only cleared when the HEAVY docker+
health batch re-ran, so "⏳ booting" outlived actual readiness by a poll
cycle (audit: >=25s stale across three surfaces). While booting, the
fast GPU tick now piggybacks a 1.5s /v1/models probe; on 200 it pulls
the next heavy poll forward (burst regime) — every booting surface
flips within a tick or two of the API answering. Bounded: fires only in
the booting state.
Verified: c3 pytest suite 516 passed; live serve via switch.sh prints
the ready-line with the real LAN IP + the neutral served name
(http://192.168.86.33:8020/v1 · qwen3.6-27b · no auth).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Adds a stdlib HTTP control plane that wraps scripts/switch.sh so a harness
can POST /switch and block until the new model is serving. Introduces no new
orchestration logic — switch.sh stays the single source of truth (registry
lookup, down/up, readiness).
- tools/model-switch/server.py: GET /healthz|/status|/models, POST /switch
({slug}|{model}); registry-validated; /health readiness (works with or
without VLLM_API_KEY); single-flight lock; refuses to start unauthenticated
on a non-loopback bind.
- scripts/systemd/club3090-model-switch.service: host daemon unit.
- scripts/tests/test-model-switch.sh: hermetic HTTP/auth/validation contract.
- docs/EXAMPLES.md, .env.example: usage + config.
Mirrors the existing stdlib HTTP style (services/studio/*); zero new deps.
Experimental/opt-in per the repo's staging convention.
InternScience's 35B agentic MoE (Qwen3-Next MoE arch; the card declares its
OWN base -> own model entry, not a qwen fine-tune slug), served from the
official FP8-dynamic compressed-tensors checkpoint on stock vLLM v0.24.0,
dual 3090 TP=2, full 262K. The agentic thinking-ON specialist.
First model onboarded END-TO-END through the Bring & Validate lane (T2
producer-zero): pull.sh route-C sibling swap (post deriver fix) ->
generate-compose -> documented swap -> full gate -> ④ vs-bar -> ⑤ scaffold.
Full gate PASS (rebench tag agents-a1-fp8-dual, 2026-07-03):
- bench 153.9/154.0 decode TPS (n=5, CV 0.1%, TTFT ~130ms), ~22.0 GB/card
- verify-stress 8/8 — staggered NIAH exact-recall to 240K (91%), VRAM Δ0
- soak PASS (0 growth, 0/100 silent-empty, 99.8% retention)
- 8-pack --full OFF 105/150 · ON 110/150 (post benchlocal #79+#81 harness):
toolcall 15/15 OFF · IF 15/15 ON · cli-40 thinking-ON 23/40 (the stack's
highest; base 17/40) · hermes 12/20 OFF, REGRESSES to 9/20 thinking-ON
(verified model behavior — disclosed as caveat 1)
vs qwen3.6-35b-a3b: general capability TIES (ON 110=110), decode −13% —
NOT a general upgrade; reach for it on tool/CLI-agent work thinking-ON.
Catalog wiring: compose dual/fp8-dynamic/fp8.yml (non-default KV named per
convention; OWN port 8072 — sibling-shared ports masquerade the slug in
estate detection); model profile (geometry verified == 35B-A3B from
config.json; mtp_num_hidden_layers=0 — safetensors header shows 0 mtp
tensors, config's 1 is an inherited default); registry entry status=caveats;
kv-calc agents-a1 spec + alias (fit-all prices, calibration 100%);
weights.py aliases (hf auto-fetch); LiteLLM route :8072; BENCHMARKS section;
count bumps (registry 57, disk 58, models 11).
Full catalog suite 60/60 (1 = known worktree-fixture submit-bench).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
pull.sh (and c3's ① Bring pane) rejected EVERY llm-compressor checkpoint
(FP8-dynamic, INT8 W8A8, INT4 pack-quantized) with `quant-dtype-unknown
(bits undeterminable)`: _quant_bpw only read top-level bits keys + method-
name heuristics, and "compressed-tensors" as a method matches none — but
the bits live nested at config_groups.<g>.weights.num_bits. Parse them
(widest group wins — mixed-precision groups exist; the widest dominates
the VRAM footprint the fit-check prices).
Found by the T2 producer-zero dogfood (Agents-A1-FP8-dynamic): step ①
hard-stopped with arch=null + swap_path=null; post-fix the lane resolves
the arch and emits the honest route-C verdict (sibling qwen3.6-35b-a3b,
BRING_YOUR_OWN swap pointer) — validated live, and the swapped compose
boots + serves on stock v0.24.0 (Marlin weight-only FP8 MoE on sm_86).
+3 deriver test cases: FP8-dynamic (the A1 shape), INT4 pack-quantized,
mixed groups -> widest.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Switching gpu-mode ai-studio → gemma on a desktop rig (GNOME running on the GPUs,
~1.3 GiB/card) leaves < the 0.95 util budget free, so gemma failed vLLM's boot
free-memory check (`free < util × total`); with restart:unless-stopped it restart-
looped and switch.sh waited the full 600s READY_TIMEOUT for an endpoint that never
came. The existing gpu_preflight only enforces a coarse 80%-free floor (the reporter
had ~90% free), so it sailed through. qwen (0.92) fit → that path worked, which is
why *only* gemma broke.
- preflight.sh: new preflight_compose_gpu_fit — parses the effective
GPU_MEMORY_UTILIZATION (env override or compose default) + TP card-count, compares
per-card free VRAM to `util × total × 0.98` (~ vLLM's own mem_get_info total), and
HARD-fails (unless --force) with an actionable message (free VRAM / lower util). A
~10s settle-retry covers teardown lag (docker `down` returns before CUDA frees).
- switch.sh: call it in up_variant (vllm-only, after down_running so the retry also
covers the just-torn-down container). launch.sh delegates to switch.sh → covered.
- gpu-mode.sh: wait_gpu_vram_settle after scene teardown in the gemma handler, so the
incoming TP=2 model boots into freed VRAM instead of ai-studio's residue.
- test-preflight-gpu-fit.sh: mocks nvidia-smi + the #535 numbers — fit / short+message
/ --force / env-override / single-card.
Full shell gate 60/60 (1 = known worktree-fixture). Validated end-to-end against the
real gemma compose: idle rig fits; simulated 22060 MiB free → instant "GPU 1 has 21.5
GiB free, needs ~22.3" + lower-util hint (matches vLLM's 22.38 threshold).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The shared served-model-name `qwen3.6-27b-autoround` (from #490) mislabels the
non-autoround 27b scenes (fp8 dual-max, lmcache): /v1/models advertises
"autoround" while `root` points at qwen3.6-27b-fp8. Same class on gemma-4-31b —
after the v0.24.0 consolidation the default is cyankiwi qat-AWQ-INT4 (bf16 KV),
yet the LiteLLM route still targeted `gemma-4-31b-autoround` (a latent #482 drift).
Fix WITHOUT breaking anything, via vLLM multi-served-name:
- Every 27b scene now serves `qwen3.6-27b <its-quant-name>`; every gemma-31b
scene serves `gemma-4-31b <its-quant-name>`. The neutral name is PRIMARY
(honest /v1/models id); the quant-specific name is retained as a live ALIAS.
- LiteLLM: add `qwen3.6-27b` / `gemma-4-31b` canonical public routes; keep the
`-autoround` routes as back-compat aliases (same upstream). Repairs the gemma drift.
- Migrate our own MODEL= defaults + docs (bench/verify/quality/launch/setup, c3,
tui-core, EXAMPLES, ...) to the neutral name. Weights slugs (`-autoround-int4`)
untouched; CHANGELOG + results/ history left as-is.
Retiring the `-autoround` alias entirely is a deliberate later step once nothing
still asks for it.
Live-verified on-rig (single/minimal, stock v0.24.0): /v1/models lists BOTH names
(root=...-autoround-int4); chat to `qwen3.6-27b` AND `qwen3.6-27b-autoround` both
return 200; `qwen3.6-27b-fp8` correctly 404s. Full shell gate 59/59 (1 = known
worktree-fixture); c3 pytest 41 passed; served-name arg-order + YAML validated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Our vllm#40361 sub-tile-n Marlin pad was closed-superseded by mgoin's
vllm#45295 (consolidated marlin_padded_nk across all dense Marlin paths),
native in vLLM v0.24.0. `vllm-stable` now pins v0.24.0 and the AutoRound
INT4 TP=2 path (vllm/dual) boots clean without the overlay (validated in
#533 Phase 0b). No live compose mounts the patch (archive-only), so this
is a tracking-only de-registration — the cleanup deferred from #533.
- patches.yml: qwen-vllm-marlin-pad -> deprecated (upstream.status
open->merged, load_bearing_when [], delivery none, drift_guard null),
mirroring the gemma-vllm-pr41800 merged-and-dropped precedent. Kept as
history (foundational false; entry not deleted).
- arch_patches.yml: correct the stale kernel_constraints note (#40361 ->
#45295 native in v0.24.0). required_patches / marlin_alignment_required
unchanged: the alignment is a real arch property (now satisfied stock),
and deprecated patches stay listed per the pr41800 precedent.
- UPSTREAM.md: mark the #40361 / #40354 / v0.24.0-bump marlin rows DONE
(native in v0.24.0, patch de-registered, no live mount) and correct the
stale "composes still mount it" line (all mounts are under _archive/).
Full shell gate green (59/59; the 1 = known worktree-fixture-absent
test-submit-bench). test-patch-attribution (reads both registry files) passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm