Commit Graph

1156 Commits

Author SHA1 Message Date
noonghunna
055986a5d0 Merge pull request #572 from noonghunna/fix/a1-untrack-compile-cache
Un-track A1's compiled torch cache — it crashed Blackwell boots (#548 root cause)
2026-07-05 03:59:14 +05:00
noonghunna
717cb431a9 Un-track A1's compiled torch cache — it crashed Blackwell boots
Root cause of club-3090 #548, pinned from the stack trace: A1's
promotion accidentally committed the whole torch_compile cache (3,667
files, ~20 MB) compiled on THIS sm_86 rig, and the compose warm-start
mount fed it to every fresh pull. On sm_120 the AOT graph (baked
against Marlin-processed FP8 layers -> reads layer.workspace) loads
onto Cutlass-processed layers -> AttributeError -> restart loop.
vLLM's AOT cache key doesn't include arch/kernel selection, so the
cross-arch hit is silent (upstream issue to follow; UPSTREAM.md row
updated with the pinned mechanism).

Fix = the house pattern every other model already has: cache contents
gitignored (cache/.gitignore + README), directory kept for the mount,
local files untouched (our warm-start intact). Fresh users pay one
~60-90s compile on first boot and warm-start locally thereafter --
against THEIR OWN silicon's kernel selection.

Also retires the recurring dirty-tree noise from best_config files
updating during our own runs.

Full scripts gate 65/65.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 22:54:50 +00:00
noonghunna
6b7001aae6 Docs: first Blackwell A1 row (#567) + sumo quality (#552) + sm_120 FP8 tracking
- BENCHMARKS: guybrush01 2x5090 A1 row (decode 220.21/220.36 n=5 CV
  0.0%, +43% over the 3090 gate) with the forced-Marlin workaround
  caveat labeled explicitly (NOT native FP8 GEMMs -- headroom pending
  the upstream fix); sumo-dandan row gains his --medium quality
  (69/75, current harness).
- UPSTREAM: row for the v0.24.0 sm_120 FP8 kernel-selection
  AttributeError (to-file status; workaround validated cross-rig via
  #548 -> #567).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 22:41:57 +00:00
noonghunna
c8b840f406 Merge pull request #570 from noonghunna/fix/arch-ab-e4m3-sm-floor
arch-ab dogfood findings: e4m3 SM-floor refusal, STRESS_FAST (−58%), honest dual VRAM margins
2026-07-05 03:28:14 +05:00
noonghunna
2790a7adcb STRESS_FAST mode + honest dual-rig VRAM margins (dogfood findings)
Everything the #246 runner dogfood on this rig surfaced, in one PR:

- verify-stress STRESS_FAST=1 (opt-in; gates keep full mode): skips
  probe 7's large fresh needles (near-duplicate the ladder's depths)
  and caps the ceiling ladder at 2 rungs (~2/3 anchor + ceiling).
  MEASURED on the identical vllm/dual boot: 861s -> 361s (-58%).
  Depth coverage stays 4-anchor (bench probe 10K/90K + mid + ceiling).
- VRAM margin attribution: compose `count: all` (DeviceRequests
  Count:-1) and VISIBLE_DEVICES=all now resolve as a DETERMINED
  "all host GPUs" -- kills the spurious per-rung WARN on dual rigs.
- VRAM margin aggregation: MIN per-card free instead of sum. TP OOMs
  on the first card to run out; the sum overstated dual margins ~2x
  (906 MB reported where the honest figure was 453 MB). SEMANTIC
  TIGHTENING: multi-GPU margin advisories now fire against the real
  per-card floor. Ceiling test updated to the new contract.
- arch-ab: lean tier exports STRESS_FAST=1 (per-arm ~30 -> ~20 min);
  margin advisory surfaced in the summary table + a mid-run "ladder
  CLEAN, not a failure" note so first-time runners don't abort on the
  rc=1 advisory; bundle print points at the #246 test thread.

Full scripts gate 65/65. The e4m3-below-sm_8.9 refusal is the prior
commit on this branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 22:22:54 +00:00
noonghunna
2477f53e2f arch-ab: refuse the e4m3 arm below sm_8.9 (found dogfooding on sm_86)
vLLM hard-rejects fp8_e4m3 KV below SM 8.9 at boot -- without the guard
a bare run on an Ampere rig (default arms include e4m3) sits in
switch.sh's ready-wait until the 10-minute timeout instead of getting a
clear refusal. The message says why: on Ampere there is no arch delta
to measure. --arms e5m2 (control-only) stays possible.

test-arch-ab: bare-run-on-3090 refusal + control-only-on-3090 pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:25:53 +00:00
noonghunna
3c86246844 Merge pull request #569 from noonghunna/feat/246-arch-ab-runner
Add arch-ab.sh: cross-rig KV-dtype A/B runner for #246 (lean tier)
2026-07-05 02:23:55 +05:00
noonghunna
8b5e6cee60 Merge pull request #568 from noonghunna/feat/246-arch-aware-kv-injection
Launcher arch-aware KV dtype injection for pilot slugs (#246 Phase 1)
2026-07-05 02:23:52 +05:00
noonghunna
47a9dacc1c arch-ab: rig report gains the kv-calc calibration matrix
report.sh --full-calibration instead of the bare default: seconds, no
test re-runs, and the predicted-vs-actual VRAM matrix on the volunteer
card class is direct Phase 2 input (32 GB envelope calibration).
Deliberately NOT --full -- that re-runs the whole battery (~43 min)
against only the last arm's container; the arms already carry that data.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:14:47 +00:00
noonghunna
a98a2468bf arch-ab: include the redacted rig triage report in the bundle
report.sh default mode (~2 s, path/host/user redaction ON) rides in the
tarball — PCIe lane width / driver / power caps / WSL-vs-bare-metal is
the context that makes cross-rig A/B variance interpretable (the same
mechanism that auto-flagged the x4-lane slot in #158).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:13:26 +00:00
noonghunna
b569b62d89 Add arch-ab.sh: cross-rig KV-dtype A/B runner for #246 (lean tier)
One command per volunteer rig: runs the pilot variant once per KV-dtype
arm (fresh symmetric boot each), verify-full + bench n=5 + NIAH ladder
per arm -- soak and quality packs deliberately skipped (quality
consolidates at promotion time per the #246 test plan; ~25-30 min/arm).
Ends with a per-arm comparison table and ONE attachable tarball.

Arms: e5m2 (control) / e4m3 (native FP8 KV, sm_89+) / nvfp4 (Blackwell
opt-in, refused below sm 10.0) / fp8w (vllm/qwen-27b-dual-max STOCK --
FP8-weights checkpoints reject fp8 KV, so this arm measures the
native-FP8-WEIGHTS lift instead; dual-rig only).

Guard rails: variant auto-pick by GPU count (2+ -> vllm/dual, 1 ->
vllm/minimal); KV arms are explicit KV_CACHE_DTYPE pins (unambiguous on
any launcher version); the running container's --kv-cache-dtype is
asserted live before any bench time is spent; --dry-run / --resume /
--full escape hatch.

test-arch-ab: hermetic plan/refusal contract via CLUB3090_FAKE_GPUS.
Full scripts gate 65/65.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:02:00 +00:00
noonghunna
e8bcfd8da1 Launcher arch-aware KV dtype injection for pilot slugs (#246 Phase 1)
launch.sh/switch.sh now detect GPU arch and export KV_CACHE_DTYPE=
fp8_e4m3 (native FP8 compute) on sm_89+ cards for the pilot slugs
(vllm/dual, vllm/minimal). The injected value is the hardware
profiles' dormant kv_format_default.balanced -- one source of truth
shared with the pull gates; the Ampere no-op is data equality
(3090-class balanced = fp8_e5m2 = compose default -> nothing emitted),
not a code branch.

Injection guards (all load-bearing):
- pilot allowlist only; expansion gated on the #246 cross-rig A/B
- per-variant: registry kv_format == fp8_e5m2 only (int8-PTH/TQ/bf16
  slugs never touched -- compressed-tensors weights reject fp8 KV)
- vLLM-family variants only; explicit user KV_CACHE_DTYPE= wins;
  unmapped cards / heterogeneous rigs / no nvidia-smi -> no injection
- direct `docker compose up` keeps Ampere-safe compose defaults

Rides the existing resolve-variant-pin export seam (new optional
--gpu-spec); VLLM_ATTENTION_BACKEND is whitelisted but ships no value
(vLLM auto-detect stays the default until measured). Preflight banner
names the detected arch class.

Consistency fixes the injection exposed:
- gates.py _ARCH_KERNEL_SM: fp8_e4m3 9.0 -> 8.9 (vLLM's real floor is
  SM89+; we'd otherwise inject e4m3 on 4090s our own pull-gate calls
  unloadable) + new nvfp4: 10.0 row (v0.24.0 literal had NO gate --
  a 3090 pull of an nvfp4 config wouldn't have been rejected)
- Blackwell hardware profiles: nvfp4 declared as CANDIDATE capability
  (engine list unchanged until validated -- gates take the intersection)
- kv-calc: projected nvfp4 rows (bytes/elem + activation coefs),
  calibration unchanged

Docs: HARDWARE.md new section, KV_MATH/QUANTIZATION/DTYPE_MATRIX rows
(incl. retiring the Genesis-era "e4m3 undertuned per #51" advisory in
favor of the A/B).

Validated: 8-case injection matrix in test-launch-compat; live no-op
on the real 2x3090 (spec built, nothing injected); faked 4090 through
the real bash detection path emits e4m3; compose interpolation both
ways; kv-calc --calibration green; full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 20:45:28 +00:00
noonghunna
7dd7ec6332 Merge pull request #566 from noonghunna/feat/gemma-dflash-regate
Re-gate beellama/gemma-dflash on the official pin; fix rebench-full n=3
2026-07-05 00:38:36 +05:00
noonghunna
5ea64bc539 Re-gate beellama/gemma-dflash on the official pin; fix rebench-full n=3
The single-card Gemma-4 default was the catalog's only default with no
baseline row (wave-2 PRIORITY gap). Full gate on Anbeeld's OFFICIAL
v0.3.2-preview digest (the current beellama-local pin), rebench tag
gemma-dflash-regate:

- decode 44.91 narr / 79.76 code (n=5, CV 2.1%/5.6%), TTFT 115 ms
- NIAH ladder clean to 117,513 tok (91% of 128K); ceiling margin 929 MB
  (< 1024 MB bar -> caveat stays, but up from ~177 MB pre-v0.3.2)
- 8-pack current harness: 108/150 off / 113/150 on (+5 in-band wash,
  think-off default re-confirmed)
- soak PASS (p50 59.8, 100.6% retention, 0 growth, 0/100 silent-empty)

Row seeded via catalog-baseline.sh (first non-vLLM engine through the
producer flow); BENCHMARKS row + compose header (Quality + margin
numbers) updated to match.

Tooling fallout the gate exposed:
- rebench-full.sh hardcoded RUNS:-3/WARMUPS:-1, silently under-running
  the bench protocol on every orchestrated gate. The A1 "n=3 env leak"
  attribution was wrong -- it was this line (baselines comment
  corrected). Fixed to protocol 5/3; n=3 had flattered gemma-dflash
  code TPS by +8%.
- catalog-baseline.sh counted prefill-probe run-lines into its bench-n
  gate (read n=8 for an n=5 log); now reads the summary headers.

Full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 19:28:03 +00:00
noonghunna
0588507960 Merge pull request #565 from noonghunna/feat/baselines-wave-2
Seed baselines wave-2: 5 rows + A1 decode-pair fix + gap dispositions
2026-07-04 21:58:45 +05:00
noonghunna
a627efa470 Seed baselines wave-2: 5 rows + A1 decode-pair fix + gap dispositions
Reviewed dispositions of the wave-1 gap list (verdict sheet approved
2026-07-04):

- vllm/dual + vllm/qwen-27b-dual-fast (alias mirror): FRESH from the
  2026-06-30 v0.24.0 pin-bump gate row (decode 70.7/93.5, NIAH->240K,
  soak PASS) — measured on the current vllm-stable pin, so no
  born-stale archaeology needed after all.
- ik-llama/iq4ks-mtp-vision: carried decode (60.39/72.4) from iq4ks-mtp
  per the 2026-05-25 vision re-tune row; carried qualifies because the
  vision delta doesn't touch the decode path.
- ik-llama/apex-fit-q8q5: 2026-05-28 row (decode 105.63/156.80, n=5),
  NIAH clean@180K.
- vllm/gemma-26ba4b-single: 2026-06-06 row (decode 169.2/219.9) + full
  8-pack (98 off / 109 think-on), NIAH clean@161K — FRESH (gemma-stable
  still pins v0.22.0).
- vllm/agents-a1-dual: decode pair order corrected from the tag
  artifact (narr 154.0 / code 153.8 — wave-1 had them swapped).
- Footer: SEED WAVE 2 list -> KNOWN GAPS dispositions (vllm/minimal,
  gemma-12b pair, beellama/gemma-dflash PRIORITY re-gate, deckard).

catalog-baseline.sh: insertion anchor updated for the retitled footer
(accepts both titles); test-catalog-baseline.sh now guarantees the ADD
path (strips the seeded vllm/dual row from its copy) and asserts added
rows land before the footer — the marker-drift class this bump caught.

Guards: test-baselines 15 rows joined, same 4 wave-1 stale WARNs, all
5 new rows fresh. Full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:56:11 +00:00
noonghunna
d19d00b9d3 quality/baselines: rescore-materialization practice — artifacts carry the accepted truth
Closes the rescore-materialization gap the induction tool's first live
run surfaced (#562): the benchlocal-cli 'rescore' subcommand ALREADY
supports write-back (--in-place / --output) — the gap was practice, not
a missing upstream feature. The T2 rescore ran to stdout only, so the
published A1 thinking-on 110/150 diverged from the tag artifact (108).

- A1's rescore MATERIALIZED into the tag JSONs (rescore --in-place; OFF
  105 unchanged, ON 108→110) — induction now extracts 110 and the
  regenerated corpus record carries 110; artifact, corpus, baseline row
  and publication all agree. Pre-rescore JSON backed up outside the tag.
- docs/QUALITY_TEST.md: 'Rescoring saved results — MATERIALIZE, don't
  just read' — the rule (a rescore that changes a published number must
  be written back in the same session), the command, what rescore can't
  re-run (sandbox packs), and the A1 case as the cautionary example.
- catalog-baseline.sh header: reads-artifacts-as-truth warning + the
  header's gate text updated to the actual n>=3/warn-under-5 behavior.
- baselines.yml A1 comment: gap note -> materialized + doc pointer.

Guards green (test-catalog-baseline, test-baselines).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:30:53 +00:00
noonghunna
71e04fb214 Merge pull request #564 from noonghunna/feat/baselines-slice-2c
catalog-baselines slice 2c: canonical two-depth prefill/TTFT probe + anchor calibration
2026-07-04 21:27:33 +05:00
noonghunna
34b31c0aa4 catalog-baselines slice 2c: canonical two-depth prefill/TTFT probe + anchor calibration
Decode-only TPS can't express real trade-offs anymore (the W8A8-vs-FP8
result was a prefill-corner-vs-decode-corner split) — this adds the
CANONICAL prefill/TTFT measurement to bench.sh per the design's sourcing
rule, with the protocol gotchas productized:

- bench.sh PREFILL PROBE (default-on; PREFILL_PROBE=0 / PREFILL_DEPTHS /
  PREFILL_RUNS): warm + n measured per depth (10K + 90K anchors; 90K is
  inside the DeltaNet degradation regime and pairs with the NIAH ladder's
  ~94K rung). CACHE-BUSTED: fresh salted haystack per request — composes
  serve enable_prefix_caching, an identical prompt re-measures the CACHE
  HIT (vLLM's prefix cache is block-chained; unique first line breaks the
  chain). SELF-CALIBRATING: word-count heuristics overshoot tokens ~1.3x;
  the warmup's reported prompt_toks scales the measured runs (target^2/
  actual) to within ~4% of the requested depth. Depths exceeding the
  served ctx SKIP with a note. DUAL METRIC, labeled: prompt_tokens/TTFT =
  client-observed (user-truth: incl tokenization+transfer+scheduling) AND
  the vLLM stats-log windowed rate = engine-internal (compute-truth) —
  at 93K on A1 they differ by ~7s of non-prefill overhead (5.5K vs ~10K
  t/s); never cross-compare kinds (stack LEARNINGS row added).

- measurement_record parser: per-block pass -> prefill_tps_by_ctx +
  ttft_ms_by_ctx extensions; the canonical short-prompt ttft_s is
  PROTECTED from the probe blocks (the old last-occurrence rule would
  have swallowed the 90K block's 17s TTFT).

- catalog-baseline.sh: rows gain prefill_tps {10k: N, 90k: M} (parsed
  via THE record parser, no second grammar) + ANCHOR CALIBRATION at
  induction: the probe's deep anchor vs the NIAH ladder's nearest rung —
  agreement (0.7-1.3) certifies the ladder's whole depth curve; A1 live:
  probe 5459-5584 t/s @93K vs ladder 7403 @94K = ratio 0.74-0.75, OK.
  Divergence warns with an investigate message (design: a finding).

- test-baselines schema: prefill_tps = dict of numeric depth points.
  test-catalog-baseline fixture: probe blocks + TTFT-pollution guard +
  anchor-OK assertion.

Live-validated 3x against the serving A1 (262K): 10K = 7977 t/s CV 1.1%
TTFT 1.25s; 93K = 5584 t/s CV 0.5% TTFT 16.1s; engine-log ~10K t/s.
Full scripts gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:22:48 +00:00
noonghunna
f6a1caf6da Merge pull request #563 from noonghunna/feat/baselines-slice-2b
catalog-baselines slice 2b: c3 'yours vs the bar' overlay + bar provenance/staleness detail
2026-07-04 21:07:20 +05:00
noonghunna
c6f56e0c88 catalog-baselines slice 2b: c3 'yours vs the bar' overlay + bar provenance/staleness detail
The consumer half of the producer wiring: THIS RIG's corpus records now
render next to the shipped bar in the catalog preview — the design's
'c3 overlays the bar with your numbers, badged'.

- LocalMeasured projection + CatalogEntry.local_measurement: the newest
  #249 corpus record per variant slug (decode point, both 8pk arms, pin,
  date). One canonical-short decode point per record — the narr/code
  pair split is a parser note for 2c.
- services.local_measurements(): pure-fs corpus scan; newest-per-slug by
  the record's _recorded_at stamp (NEW — rebench-full now stamps it at
  write; mtime+line-order fallback covers pre-stamp records); malformed
  lines skipped. Joins in the same enrichment pass — still zero
  subprocess on the catalog path.
- Preview rework: the measured line is the BAR with provenance
  (date · rig · submitted_by from the joined baseline); a stale bar gets
  an explicit detail line ('† bar measured on <pin> — current <pin>;
  re-bench owed'); a slug this rig has gated adds
  'yours ~<decode> · 8pk <P/150> (<date> · this rig)'.

Live-verified: the corpus's first record (A1, from 2a's live run) joins
and renders 'yours ~154 decode · 8pk 105/150'. Tests: overlay semantics
(newest-wins by stamp, malformed-skip, entry join, no-record → None) +
the three-state preview (fresh bar + provenance + yours · stale pin
detail · no-yours). c3 suite 766/766; scripts gate green serial.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:05:51 +00:00
noonghunna
746752e63a Merge pull request #562 from noonghunna/feat/baselines-slice-2a
catalog-baselines slice 2a: induction tool + rebench auto-record + completion prompt
2026-07-04 20:44:38 +05:00
noonghunna
1d974459c9 catalog-baselines slice 2a: induction tool + rebench auto-record + completion prompt
The producer wiring: a validated gate run now flows into the corpus
automatically and into the shipped bar via one reviewed command.

- scripts/catalog-baseline.sh (§2.3, one mechanism / three entry points):
  gate-validates the tag dir (verify-full pass HARD · bench n>=3 HARD,
  n<5 WARN vs the canonical target · quality present unless --tps-only),
  extracts decode TPS/TTFT/8pk both arms/ctx_validated (NIAH-ladder
  verdict incl. token count)/provenance (pin via the same resolution the
  emit join uses; rig+power via nvidia-smi, all overridable), UPSERTS the
  baselines.yml row textually (comments preserved), prints a unified
  diff. Never commits — rows ship via PR. --baselines-file for tests.

- rebench-full.sh: on completion, EXACT-container-match the served
  engine to a registry slug (identity semantics — never port/substring,
  the F9 rule) and append a fingerprint-complete #249 record (engine_pin
  via resolve_variant_pin/compose fallback · hardware · power-cap ·
  quality_8pk extensions · soak status) to the gitignored corpus, then
  print the catalog-baseline.sh induction prompt. BYO/swap serves skip
  with a note (no registry identity to record against).

- measurement_record.py: --quality-8pk / --quality-8pk-think-on ->
  measured_extensions (the designed extension namespace; frozen schema
  untouched).

- test-catalog-baseline.sh: hermetic (synthetic tag dir + a COPY of
  baselines.yml) — extraction, dry-run no-write, add->replace upsert,
  refusals (missing verify/quality/unknown slug), real-file checksum
  guard.

TRUTH CORRECTIONS the tool's first live run surfaced (run against the
real agents-a1-fp8-dual tag): the A1 gate's bench was n=3 (a RUNS=3 env
leak into the overnight gate — 'n=5' was misstated in publications), the
BENCHMARKS TPS pair was order-garbled (artifact: narrative 154.0 / code
153.8), and the #81-rescored think-on total (110) was never materialized
back into the tag JSON (reads 108) — materialization gap documented in
the row comment. baselines.yml + BENCHMARKS.md corrected.

Live-validated: rebench record block ran against the serving A1 → first
corpus record written (vllm-agents-a1-dual__e9d28e22.jsonl: pin v0.24.0,
rtx-3090, 370W, quality extensions); induction dry-run reproduces the
row from the tag. Deferred to 2b/2c: c3 local-overlay + richer badges;
the bench 10K/90K prefill/TTFT probe. Gates: test-catalog-baseline +
test-measurement-record + test-baselines green; full scripts gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 15:40:54 +00:00
noonghunna
68f52d5c5e Merge pull request #561 from noonghunna/feat/baselines-slice-1
catalog-baselines slice 1: baselines.yml + registry-emit join + guards; catalog drops the BENCHMARKS scrape
2026-07-04 20:27:42 +05:00
noonghunna
75b29556c1 catalog-baselines slice 1: baselines.yml + registry-emit join + guards; catalog drops the BENCHMARKS scrape
The catalog's measured columns (TPS / 8pk) now come from ONE productized
source: scripts/lib/profiles/baselines.yml — the shipped, PR-reviewed
'bar' (accepted display projection of a validated gate run, with pin/rig/
power provenance) — joined per-slug into registry-emit --json with an
emit-computed staleness verdict. Consumers never read baselines.yml or
BENCHMARKS.md directly (BENCHMARKS stays the public human/cross-rig
ledger). Design: the catalog-baselines note (2026-07-02, §2/§5 slice 1).

- baselines.yml SEED WAVE 1: 10 rows with airtight traceability only
  (rebench tags on disk: agents-a1 golden specimen + gemma-31b-dual;
  unambiguous decode-class BENCHMARKS rows for the rest). 4 rows are
  HONESTLY born-stale with documented reasons (llamacpp rolling-tag-era
  x2, beellama pre-#296 image, 35B v0.22.0) — the guard's demo cases.
  9 slugs still owed rows are listed as wave-2 gaps, no guessed numbers.
- registry-emit join: per-variant 'baseline' field; current pin resolved
  the way launchers actually resolve it (engine-profile install.spec,
  compose-image-default fallback for ik/llama.cpp) → 'stale' =
  measured-pin != current-pin, null when undeterminable.
- test-baselines.sh: schema + slug-membership + ctx-parity (compose ctx
  default == registry max_ctx, functional slugs) RED; pin-staleness
  WARN-only (pin bumps must not block on immediate re-bench — the debt
  stays visible). test-registry-json gains the 'baseline' contract key.
- c3: enrich_measurements = pure in-memory map off the joined field —
  deletes BOTH the per-slug --explain fan-out (~4s/slug, the #439
  option-3 leg) and the BENCHMARKS.md scrape from the catalog path
  (Explain modal + cross-rig explorer keep their readers). Stale rows
  render a † on the TPS cell + a status-line legend; full badge/overlay
  treatment is slice 2. Measurement gains source='baseline' + stale.
- results/baselines/README: the three-store relationship (regression
  corpus vs measurement records vs display bar) + the no-drift rule the
  slice-2 induction tool enforces.

Verified: guard test green (10 rows, 4 stale-warned); live join emits
57 variants / 10 baselines; live c3 catalog shows all 10 with daggers on
exactly the stale 4 (A1 serving during the check: 154/154 - 105/150);
c3 suite 764/764; full scripts gate green serial. (First gate pass
tripped classifier/dedup by running CONCURRENTLY with the c3 suite —
.pull-captures pollution class; serial rerun clean.)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 15:25:36 +00:00
noonghunna
c2c06e47a7 beellama docs: sm_120 root cause + Anbeeld#85 ask + verified self-build recipe
Investigation of a Discord cross-rig report (WSL, 3060Ti+4090: verify
fails on the pinned v0.3.2-preview digest, passes on the old noonghunna
snapshot; NOT reproducible on our 3090s — verify-full all-pass) surfaced
three doc-level truths worth recording:

- The 50-series gap is a TOOLCHAIN ceiling: Anbeeld's CI builds on the
  Dockerfile-default CUDA 12.4, whose nvcc cannot target sm_120 (no
  cubin, max PTX compute_90) — every official tag lacks Blackwell.
  Upstream ask filed: Anbeeld#85 (CUDA_VERSION 12.8.1 + arch list incl
  120); when it lands we retire the noonghunna snapshot entirely.
- Three registry status_notes claimed launchers inject 'server-cuda-
  v0.3.0' — stale since the 2026-06-12 pin bump; they inject the
  v0.3.2-preview digest from engines/beellama-local.yml install.spec.
  Fixed all three (qwen dflash, gemma-12b, gemma dflash) + honest
  labeling of the noonghunna snapshot as v0.3.0-feature-level and
  unmaintained (predates KVarN + v0.3.1 fixes).
- The engine-notes self-build guidance was outdated: FA_ALL_QUANTS is
  hardcoded in Anbeeld's cuda.Dockerfile since our PR Anbeeld#48, so a
  self-build needs only CUDA_DOCKER_ARCH (+ CUDA_VERSION=12.8.1 for
  sm_120). Recipe verified against his master Dockerfile 2026-07-04.

UPSTREAM.md beellama row updated (dated entry + next-triggers; the 'no
official image' claim struck through as historical). Gates: YAML +
registry import clean; status-drift / profiles-compat / switch-parity
green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 14:17:55 +00:00
noonghunna
1e2419e320 Merge pull request #549 from titan550/feat/model-switch-service
tools: add HTTP model-switch service (thin wrapper over switch.sh)
2026-07-04 16:55:45 +05:00
noonghunna
1c0ea8fcb6 BENCHMARKS: Agents-A1 cross-rig row — @sumo-dandan #552 (x4 OCuLink eGPU, 250W)
First cross-rig confirmation: decode 152.2/152.3 (CV 0.0%) within ~1% of
the gate rig despite one card on an OCuLink eGPU dock at PCIe 4.0 x4 and
both capped 250W — single-stream TP=2 decode is VRAM-bandwidth-bound.
Full battery mirrors the gate (NIAH 6/6 to 240K, soak 100% retention,
agentic TTFT sub-linear); boot ran custom all-reduce via the PCIe-P2P
auto-detect, clean throughout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 11:44:12 +00:00
noonghunna
5f3567f4e7 Merge pull request #560 from noonghunna/fix/548-a1-hardware-metadata
agents-a1: hardware-metadata header + Blackwell caveat (#548 follow-up)
2026-07-04 16:36:49 +05:00
noonghunna
ff96f89506 agents-a1: hardware-metadata header + Blackwell caveat (#548 follow-up)
The A1 compose shipped (#546) without the 'Hardware metadata' block every
sibling carries, so preflight_compose_gpu_fit logged 'WARN: compose has
no hardware metadata; allowing boot' and skipped enforcement — visible in
guybrush01's #548 report. Add the block (min-vram 24 / gpu-count 2 /
TP 2 / Requires-sm 8.0+ — Marlin FP8 weight-only needs sm_80, unlike the
sibling's INT4 7.5+); preflight now parses + enforces (live rc=0).

Also propagate the #548 finding as Caveats (4): Blackwell sm_120 hits an
upstream vLLM v0.24.0 kernel-selection bug at boot ('QKVParallelLinear'
has no attribute 'workspace' — Cutlass W8A8 selected, a Marlin-path
attribute expected). Workaround shipped as a commented env line
(VLLM_TEST_FORCE_FP8_MARLIN=1 — forces the exact weight-only path this
compose's gate validated; verified present in the v0.24.0 image) +
mirrored into the registry status_note so switch.sh's NOTE carries it.

Gates: compose_meta_get parses all 4 keys; docker compose config valid;
test-compose-status-drift / registry-disk / preflight-gpu-fit green;
full scripts gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 11:34:53 +00:00
noonghunna
011467682b Merge pull request #559 from noonghunna/fix/f2-f5-t1lite-tail
c3/gpu-mode: preview-clip fix + scene-boundary hardening (F2+F5, T1-lite tail)
2026-07-04 16:12:54 +05:00
noonghunna
b00b5a0955 c3/gpu-mode: preview-clip fix + scene-boundary hardening (F2+F5, T1-lite tail)
F2 — the Catalog preview strip clipped a WRAPPING caveat line: border
eats 2 rows of max-height 6 → 4 content lines; the dual-fast caveat
wrapped past that and lost its tail. max-height 6→8 + overflow-y auto
(height stays auto — short previews don't grow). Regression test pins a
wrapped-caveat entry to its full height with the tail visible.

F5 — the #544 deliberate deferrals:
- wait_gpu_vram_settle wired into ALL model/studio scene handlers
  (27b, 35b-a3b, gemma-12b, deckard, ai-studio) — was gemma-int8 only;
  every scene switch now lets the torn-down scene's VRAM release before
  the next boot (#535 class).
- mode_off gains an engine-prefix CATCH-ALL: the enumerated stop_*
  lists cover gpu-mode scenes, but a catalog-launched engine
  (switch.sh <slug>) survived 'off' — caught LIVE during validation
  when off left vllm-qwen36-27b-minimal serving and the 27b TP=2 scene
  booted straight into its residue (the exact #535 failure). Any
  remaining vllm-/llama-cpp-/ik-llama-/sglang-/beellama- container is
  now stopped, with a named notice.
- c3 preflight-error visibility (the third residual): verified
  already-plumbed — switch.sh's #544 refusal exits fast, its
  [preflight] ERROR lines stream into the serve pane, serve_failed
  stamps ✗ + [!] capture (pinned by existing tests). No change needed.

Validation: bash -n + full scripts gate green (by exit code); c3 suite
763/763; LIVE scene cycle off → 27b → off on the rig — boot ready
(qwen3.6-27b, 21.5 GiB/card), fixed off left ZERO engine containers
(verified with a non-enumerated probe container that the catch-all
stopped), both GPUs at 1 MiB after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 11:12:32 +00:00
noonghunna
9a6bed3389 Merge pull request #558 from noonghunna/fix/f4-f8-livepane-copy-placeholder
c3/tui-core: [Y]-copy the visible live-log tail + pane-specific placeholders (F4+F8)
2026-07-04 15:58:26 +05:00
noonghunna
da11b4cf56 c3/tui-core: [Y]-copy the visible live-log tail + pane-specific placeholders (F4+F8)
F4 — the Logs drill had no copy affordance: LivePane is RichLog-family,
which captures the mouse and implements no text selection (Top/Config
are selectable Statics). LivePane (tui-core, shared) now keeps a
plain-text tail buffer (markup stripped, capped 2000 lines) exposed via
tail_text(); c3 routes [Y] to the visible pane's tail on the Containers
Logs drill and the ③ Gate run output. Placeholders and display-only
notes are NOT buffered (append_line buffer=False), so [Y] on an idle
drill falls through to the highlighted container name as before. The
transient #serve-live pane keeps the established slug-copy semantics.

F8 — the idle placeholder was hardcoded test-runner wording ('Ready.
Select a test and press Enter to run.') that leaked into the docker-logs
drill. The placeholder is now constructor-owned (neutral 'Ready.'
default); c3 mounts pass pane-specific copy (logs drill / gate output /
'' for the hidden serve pane). c3t is unaffected (own vendored pane).

Tests: 6 tui-core unit tests (tail buffer, unbuffered notes, cap,
placeholder contract; 76/76) + 4 headless (placeholders, tail copy wins
over row-primary on the Logs drill, idle fall-through to container name,
③ Gate tail); c3t 83/83; c3 suite 762/762.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 10:58:11 +00:00
noonghunna
9b9ef63de3 Merge pull request #557 from noonghunna/fix/f7-stale-estate-claims
estate_cli: probe per-instance liveness so a leftover plan can't fake conflicts (F7)
2026-07-04 15:39:05 +05:00
noonghunna
a939fb1831 estate_cli: probe per-instance liveness so a leftover plan can't fake conflicts (F7)
estate.yml is a desired-state PLAN, but report-state reported its
instances as 'active' without checking docker — so every consumer
treated a leftover ~/.club3090/estate.yml as live GPU claims. The T1.1
audit hit this on an EMPTY rig: the c3 serve confirm warned '⚠ Starting
this will STOP estate llama-gpu0 (GPU [0]), estate llama-gpu1 (GPU [1])'
with nothing running. Scary + wrong for a fresh user.

estate_cli (the layer that owns estate liveness — fix here, not a c3
filter):
- report-state --json: each instance gains 'container' (the
  club3090-<name> it would own) + 'running' (docker probe: true / false
  / null when docker itself is unavailable), payload gains
  'running_count'. Additive — existing keys unchanged.
- report-state text view: per instance '— running' / '— down (plan
  only)' / '— liveness unknown'.

c3 reconcile gate (consumer): an estate instance is a claim unless
running == false. true / null / MISSING (older CLI output) all stay
claims — the dual-writer gate fails CLOSED on unknown liveness, and a
mid-boot instance (running container, VRAM not yet allocated) still
conflicts.

Tests: estate JSON contract test extended (running/container/
running_count; running asserted 'in (False, None)' so the test passes
with or without docker); 3 new c3 gate tests (stale plan on empty rig →
SAFE; running:true → claim even with idle GPUs; running:null → fail
closed; the legacy no-key fixtures double as the missing-key case).
Scripts gate all green; c3 suite 758/758. Live repro on this rig (which
has the exact stale estate.yml): running_count 0, reconcile now reports
estate_claims=[] while keeping the TRUE conflict (the actually-serving
vllm container).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 10:35:00 +00:00
noonghunna
746baa28fc Merge pull request #556 from noonghunna/fix/f6-catalog-column-budget
c3: keep the catalog's money columns on-screen at 120-140 cols (F6)
2026-07-04 15:19:18 +05:00
noonghunna
f163dca673 c3: keep the catalog's money columns on-screen at 120-140 cols (F6)
At 140 cols the columns a user actually picks by — TPS (our rig) /
8pk (our rig) / status — sat past the fold: the two wide '(our rig)'
headers alone cost 8 cols (right edge 141 > 140), and topology+engine
sat BEFORE the numbers despite being largely redundant with the slug
(which encodes both). A fresh user may never discover the money columns
exist (T1.1 audit, F6).

- Reorder: model · slug · ctx · TPS (rig) · 8pk (rig) · status · topo ·
  engine — a narrow terminal now folds the slug-redundant tail, not the
  numbers.
- Headers: '(our rig)' → '(rig)' keeps the this-rig provenance at 4
  chars; 'topology' → 'topo'.

Verified against the REAL registry (53 rows, headless Pilot at 140
cols): total width 133 ≤ 140 (was 141+), money right-edge 103 — visible
down to a 110-col terminal. Column test updated to pin the fold order
(every money column LEFT of topo/engine). Suite 755/755.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 10:16:33 +00:00
noonghunna
8c00cad86d Merge pull request #555 from noonghunna/feat/f10-gate-run-observer
c3: Evidence pane doubles as the live gate-run observer (F10 MVP)
2026-07-04 15:02:55 +05:00
noonghunna
a6ff86feba c3: Evidence pane doubles as the live gate-run observer (F10 MVP)
rebench-full runs (incl. nohup/CLI-launched) were invisible in c3 — a
producer facing a silent 3-hr gate assumes a hang (T2 friction #7). The
Evidence pane now OBSERVES runs in flight from the script-owned artifacts
alone: run_step tees each step to <step>.log and records completed steps
in timings.json, and REPORT.md is synthesized LAST — so a report-less dir
is a run in flight (fresh writes) or an aborted one (gone quiet).

- services.evidence_list(): report-less dirs gain live/stale grade, the
  done-steps ladder (timings.json, torn-read tolerant), the active step
  (newest step log absent from timings), and an ANSI/CR-cleaned tail of
  its log. Pure filesystem READ — observer, never executor.
- Evidence pane: live rows badge ▶ "running — <step> · N/6 steps done";
  aborted rows ⚠ incomplete; the preview renders the gate ladder
  (✓ done+duration · ▶ active · · pending) + the live tail; a 4s
  self-refresh ticks ONLY while a run is live; cursor preserved.
- Guards: ⏎ report and [s] submit refuse on live/incomplete tags
  (rebench-report WRITES REPORT.md into the dir — generating mid-run
  would corrupt the live/complete signal; incomplete runs are not
  submittable evidence).
- GATE_STEPS render constant pinned to rebench-full.sh's run_step call
  sites by a drift-guard test.

Tests: 4 service (live/stale/torn-timings/drift-guard) + 3 headless
(badge+ladder+tail, mid-run guards, stale rendering); suite 755/755.
Live-validated against the rig's real results/rebench/ (45 tags): 3
historical aborted runs surfaced honestly (ornith-35b died mid
quality-full, gemma-12b-coder-kvarn4 in verify-stress), no false LIVE.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 10:02:22 +00:00
noonghunna
ff733583d6 Merge pull request #554 from noonghunna/fix/f9-slug-masquerade
c3/detect: grade slug matches — shape guesses never masquerade as identity (F9)
2026-07-04 14:42:20 +05:00
noonghunna
4390340040 c3/detect: grade slug matches so shape guesses never masquerade as identity (F9)
A brought model serving on a curated sibling's port was PRESENTED as the
sibling: match_target_to_registry falls back exact-container -> port ->
substring, and the UI treated every match as the running model's identity.
Observed live during the Agents-A1 bring (2026-07-03): c3 showed
vllm/qwen-35b-a3b-dual serving while the rig ran agents-a1.

Detect layer (tools/tui-core, shared by c3 / c3t / estate tooling):
- ServingTarget.match_confidence records the match GRADE:
  "identity" (exact container) | "shape" (port/substring fallback) | "".

c3 renders the grade on every masquerade surface:
- estate rail: leads with the probed served id + a 👤 badge
  ("model agents-a1 👤" / dim "shape vllm/qwen-35b-a3b-dual")
- Orchestration serving card: "agents-a1 · 👤 on <slug> shape"
- Run catalog: the row badge downgrades "● serving" -> "👤 <slug> port in
  use" — a guess never claims "serving"
- Containers table + config drill: 👤-badged slug for shape matches
  (ContainerInfo.match_confidence threaded through)

Tests: 4 new tui-core matcher-grade tests; 4 new headless tests (serving
card shape/identity, rail, catalog badge); the N3 catalog fixture gains
the registry container so its intent (identity badge) is preserved.
tui-core 70/70, test-console 83/83, serve-cockpit 748/748. Live controls
on the rig: serving vllm/minimal grades identity (no visual change);
replaying the A1 port-8051 squat against the real registry grades shape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 09:37:01 +00:00
noonghunna
e7c061f3cd Merge pull request #553 from noonghunna/fix/f3-endpoint-readiness
c3+switch: show/copy the serving API URL; fast-flip booting→ready (F3/F3b)
2026-07-04 14:12:28 +05:00
noonghunna
d3836ecd87 c3+switch: show/copy the serving API URL; fast-flip booting→ready (F3/F3b)
The T1 audit's top adoption gap: even while a model SERVES, the best
endpoint info anywhere was a bare port — no scheme/host/path, no auth
note, no copy. "Point your agent at this URL" is the whole reason the
stack exists, and the URL was underivable from the UI.

One shared derivation (layer rule; #512 precedence env/.env ->
c3_lan_ip subshell -> localhost), four surfaces:
- switch.sh: ready-line on successful boot —
  "▶ API: http://<lan>:<port>/v1 (model: <served-id> · OpenAI-compatible
  · no auth)" — CLI parity; served id probed from /v1/models.
- services.lan_ip(): c3 consumes the SAME derivation, session-cached.
- Orchestration serving card: bare ":8020" -> full URL + "no auth ·
  [u] copy".
- Estate rail: compact "api :8020/v1 · [u] copy" (rail is ~30 cols; the
  full URL lives on the card).
- NEW [u] copy-API-URL key on every Run & Operate tab (kept separate
  from [Y] so row-copy semantics are untouched); honest notify no-op
  when nothing serves.

F3b (readiness lag): api_booting only cleared when the HEAVY docker+
health batch re-ran, so " booting" outlived actual readiness by a poll
cycle (audit: >=25s stale across three surfaces). While booting, the
fast GPU tick now piggybacks a 1.5s /v1/models probe; on 200 it pulls
the next heavy poll forward (burst regime) — every booting surface
flips within a tick or two of the API answering. Bounded: fires only in
the booting state.

Verified: c3 pytest suite 516 passed; live serve via switch.sh prints
the ready-line with the real LAN IP + the neutral served name
(http://192.168.86.33:8020/v1 · qwen3.6-27b · no auth).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 09:09:12 +00:00
John Shojaei
c22a9d2d84 tools: add HTTP model-switch service (thin wrapper over switch.sh)
Adds a stdlib HTTP control plane that wraps scripts/switch.sh so a harness
can POST /switch and block until the new model is serving. Introduces no new
orchestration logic — switch.sh stays the single source of truth (registry
lookup, down/up, readiness).

- tools/model-switch/server.py: GET /healthz|/status|/models, POST /switch
  ({slug}|{model}); registry-validated; /health readiness (works with or
  without VLLM_API_KEY); single-flight lock; refuses to start unauthenticated
  on a non-loopback bind.
- scripts/systemd/club3090-model-switch.service: host daemon unit.
- scripts/tests/test-model-switch.sh: hermetic HTTP/auth/validation contract.
- docs/EXAMPLES.md, .env.example: usage + config.

Mirrors the existing stdlib HTTP style (services/studio/*); zero new deps.
Experimental/opt-in per the repo's staging convention.
2026-07-03 17:44:17 -07:00
noonghunna
e4d33e09be Merge pull request #546 from noonghunna/feat/agents-a1-promotion
Promote Agents-A1 to the catalog: vllm/agents-a1-dual (⚠️ production w/ caveats)
2026-07-03 08:53:11 +05:00
noonghunna
79fb331201 Merge pull request #545 from noonghunna/t2/deriver-compressed-tensors
deriver: resolve bit-width from compressed-tensors config_groups
2026-07-03 08:53:08 +05:00
noonghunna
6926dfde48 Promote Agents-A1 to the catalog: vllm/agents-a1-dual (production w/ caveats)
InternScience's 35B agentic MoE (Qwen3-Next MoE arch; the card declares its
OWN base -> own model entry, not a qwen fine-tune slug), served from the
official FP8-dynamic compressed-tensors checkpoint on stock vLLM v0.24.0,
dual 3090 TP=2, full 262K. The agentic thinking-ON specialist.

First model onboarded END-TO-END through the Bring & Validate lane (T2
producer-zero): pull.sh route-C sibling swap (post deriver fix) ->
generate-compose -> documented swap -> full gate -> ④ vs-bar -> ⑤ scaffold.

Full gate PASS (rebench tag agents-a1-fp8-dual, 2026-07-03):
- bench 153.9/154.0 decode TPS (n=5, CV 0.1%, TTFT ~130ms), ~22.0 GB/card
- verify-stress 8/8 — staggered NIAH exact-recall to 240K (91%), VRAM Δ0
- soak PASS (0 growth, 0/100 silent-empty, 99.8% retention)
- 8-pack --full OFF 105/150 · ON 110/150 (post benchlocal #79+#81 harness):
  toolcall 15/15 OFF · IF 15/15 ON · cli-40 thinking-ON 23/40 (the stack's
  highest; base 17/40) · hermes 12/20 OFF, REGRESSES to 9/20 thinking-ON
  (verified model behavior — disclosed as caveat 1)

vs qwen3.6-35b-a3b: general capability TIES (ON 110=110), decode −13% —
NOT a general upgrade; reach for it on tool/CLI-agent work thinking-ON.

Catalog wiring: compose dual/fp8-dynamic/fp8.yml (non-default KV named per
convention; OWN port 8072 — sibling-shared ports masquerade the slug in
estate detection); model profile (geometry verified == 35B-A3B from
config.json; mtp_num_hidden_layers=0 — safetensors header shows 0 mtp
tensors, config's 1 is an inherited default); registry entry status=caveats;
kv-calc agents-a1 spec + alias (fit-all prices, calibration 100%);
weights.py aliases (hf auto-fetch); LiteLLM route :8072; BENCHMARKS section;
count bumps (registry 57, disk 58, models 11).

Full catalog suite 60/60 (1 = known worktree-fixture submit-bench).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-03 03:50:14 +00:00
noonghunna
15a4e75180 deriver: resolve bit-width from compressed-tensors config_groups
pull.sh (and c3's ① Bring pane) rejected EVERY llm-compressor checkpoint
(FP8-dynamic, INT8 W8A8, INT4 pack-quantized) with `quant-dtype-unknown
(bits undeterminable)`: _quant_bpw only read top-level bits keys + method-
name heuristics, and "compressed-tensors" as a method matches none — but
the bits live nested at config_groups.<g>.weights.num_bits. Parse them
(widest group wins — mixed-precision groups exist; the widest dominates
the VRAM footprint the fit-check prices).

Found by the T2 producer-zero dogfood (Agents-A1-FP8-dynamic): step ①
hard-stopped with arch=null + swap_path=null; post-fix the lane resolves
the arch and emits the honest route-C verdict (sibling qwen3.6-35b-a3b,
BRING_YOUR_OWN swap pointer) — validated live, and the swapped compose
boots + serves on stock v0.24.0 (Marlin weight-only FP8 MoE on sm_86).

+3 deriver test cases: FP8-dynamic (the A1 shape), INT4 pack-quantized,
mixed groups -> widest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-02 23:24:28 +00:00
noonghunna
fc7fff717c Merge pull request #544 from noonghunna/fix/535-gpu-vram-preflight
fix(#535): fail fast + actionable when gpu_memory_utilization doesn't fit free VRAM
2026-07-02 17:52:06 +05:00