Commit Graph

477 Commits

Author SHA1 Message Date
noonghunna
7174ad91c4 Promote vllm/qwen-27b-dual-max to Production (soak completes the gate)
The FP8 max-accuracy tier had bench (83.1/108.2) + quality (107/150) on the
2x3090 reference rig but no soak-continuous run — the one missing gate item.
Ran the full operational gate fresh on the v0.24.0 pin:
  - verify-full 9/9
  - verify-stress: all 6 rungs, fillable to 240,636 tok clean (91%)
  - soak-continuous PASS: 0 err / 0 growth / 100% retention, p50 decode 85
That clears the production bar.

- registry + compose header: status experimental -> production (drift guard
  green). status_note also de-staled: dropped the '~56 TPS' probe / 'slowest of
  the three' framing (corrected to decode 83/108, the slow axis is PREFILL/TTFT
  from MarlinFP8 W8A16 on Ampere, not decode) + the full-gate results.
- baselines.yml: row comment notes the soak PASS completing the gate.
- No DEFAULTS[(qwen,vllm,dual)] change -> vllm/dual (fast) stays the dual
  default; dual-max just joins the actionable list as the max-fidelity tier.

Bonus for the 5090 crowd: dual-max is now a non-experimental config, so the
launch command drops --force (bash scripts/switch.sh vllm/qwen-27b-dual-max).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-06 01:12:42 +00:00
noonghunna
7e474f4c2b DeepGEMM: cover fp8-dynamic weights too (agents-a1) — all fp8 slugs 5090-safe
#588 fixed the fp8-weights DeepGEMM crash but the _deepgemm_env gate (and my
guard) matched weights_variant == 'fp8' EXACTLY — missing agents-a1, whose
weights are 'fp8-dynamic' (compressed-tensors FP8). Those are genuine FP8
weights that route FP8 GEMM, so a 5090 running vllm/agents-a1-dual would still
hit the 'recipe not found' crash the whole framework is meant to prevent.

- _deepgemm_env gate: weights_variant == 'fp8' -> startswith('fp8'), catching
  both 'fp8' and 'fp8-dynamic' (compressed-tensors). INT4/AWQ/W8A8/bf16 never
  route FP8 GEMM so they stay correctly excluded.
- agents-a1 compose: add the - VLLM_USE_DEEP_GEMM pass-through.
- test-deepgemm-fp8-parity: broadened to startswith('fp8') so it now covers
  agents-a1 (5 fp8-family composes) + any future fp8-* variant.
- test-launch-compat: new case — fp8-dynamic on 5090 injects VLLM_USE_DEEP_GEMM=0.

Now ALL fp8-family vLLM slugs are 5090-safe via the launcher (dual-max,
multi-max, dual-lmcache, diffusiongemma-dual, agents-a1). Convention: fp8
weight variants must be named fp8-* for the gate/guard to catch them.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 23:20:06 +00:00
noonghunna
f3b55a03af fp8 composes: add VLLM_USE_DEEP_GEMM pass-through parity + drift guard
#580 added the consumer-Blackwell DeepGEMM auto-disable pass-through to
dual-max, but the other fp8-weights vLLM composes — hand-maintained
parallel copies, no extends/include — had silently drifted without it:
multi-max, dual-lmcache, diffusiongemma-dual. On a 5090 / PRO 6000 / GB10
(sm_120/121) the launcher's _deepgemm_env injects VLLM_USE_DEEP_GEMM=0, but
with no receiving line docker-compose drops it → the fp8-GEMM 'recipe not
found' boot crash #580 fixed for dual-max (disc #571).

- Add the  pass-through (+ the shared comment) to all
  three. All 4 fp8-weights vLLM composes now carry it.
- NEW test-deepgemm-fp8-parity: asserts every fp8-weights vLLM compose has
  the pass-through — REDs on the exact drift class that caused this (verified
  it fails when the line is removed, passes when restored). Scope matches the
  _deepgemm_env gate (weights_variant==fp8 + vllm engine); INT4/AWQ/bf16 never
  invoke DeepGEMM so they're correctly excluded.

Removes ONE 5090 blocker (fp8 crash) — NOT a blanket 'runs on 5090': beellama
composes stay broken on sm_120 (Anbeeld#85, CUDA 12.4), llama.cpp/INT4 on
sm_120 remain unvalidated (no 5090 on the reference rig).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 22:46:20 +00:00
noonghunna
b8f2d2b927 Promote vllm/qwen-27b-multi-fast to Production + induct 4x3090 baseline (#584)
@ryanmpelletier's flawless 4x 3090 report (#584) is the cross-rig
validation the compose header required ('until a real >=4x 3090 host
validates it'): v0.24.0, no power caps, all x16, verify-full 9/9,
verify-stress needles clean to 240K, soak-continuous PASS, bench n=5.
Third concordant 4-card validation (after @alanspires 6xVFIO + @Whamp),
first on v0.24.0.

- registry + compose header: status experimental -> production (both, drift
  guard green). No DEFAULTS[(qwen,vllm,multi4)] entry -> no resolver cascade;
  promotion just moves it onto the actionable list. Quality is TP-invariant,
  carried from the vllm/dual proxy (109/150); 4-card 8-pack is the open
  confirmation (report.sh runs none).
- baselines.yml: vllm/qwen-27b-multi-fast gains its 4x3090-pcie submission
  (tier: submitted; decode 74.76/90.83, prefill 1288->1175, KV pool
  1.77M/6.77x, NIAH 240K). Submission-only by construction - the 2-card
  reference rig can never produce a local 4-card row.
- BENCHMARKS.md: ryan's row (all-x16 74.8/90.8 beats @Whamp's mixed-lane
  59/74 on the same fp8-KV config -> PCIe lane width matters for TP=4).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 22:33:26 +00:00
noonghunna
fd56fe7ad5 dual-max: induct 2x3090 baseline row + correct the stale ~56 TPS probe
The dual-max FP8 2x3090 decode bench (83.1/108.2, v0.24.0, 2026-06-30) lived
only in BENCHMARKS.md:122 + disc #515 (c17490565) — never inducted into
baselines.yml, so c3 / the registry-emit join showed only guybrush01's
2x5090 submission and no local bar. Surfaced while diagnosing #585.

- baselines.yml: vllm/qwen-27b-dual-max gains its primary tier: local row
  (83.1/108.2 · TTFT 158 · prefill 1364->875 · 8-pack 107/150 · NIAH 240K ·
  v0.24.0 = current pin -> FRESH); guybrush's 2x5090 submission preserved.
- compose header + DUAL_CARD.md: the stale '~56 TPS' probe (and the now-false
  'slowest of the three' framing) -> real decode 83/108; the genuine
  tradeoffs (smallest KV pool 295K/1.13x, slowest prefill/TTFT 158ms from
  FP8's compute-heavy Marlin W8A16 dequant) kept. BENCHMARKS.md already
  corrected; this closes the two spots that still read ~56.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 22:07:39 +00:00
noonghunna
81beaec132 deriver inventory: distinct gguf artifacts sharing a quant token stay separate
Live-dogfood round 1 (Qwythos-9B-Claude-Mythos-5-1M-GGUF): the repo ships
base + MTP builds per quant (…-Q4_K_M.gguf / …-MTP-Q4_K_M.gguf); token-
keyed grouping merged them into ONE "2-part" variant with a summed, wrong
size (10.7 GiB shown for a 5.2 GiB pick) and no way to select just one.

Group by STEM instead (basename minus the -NNNNN-of-NNNNN part suffix) —
true multi-part shards share a stem so 'parts' still counts them; distinct
artifacts don't. Label = stem minus the repo-wide common prefix (standard
repos keep their plain quant token; multi-artifact repos keep the
distinguishing part: Q4_K_M vs MTP-Q4_K_M). Guard case added.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 17:12:10 +00:00
noonghunna
a2f616521d deriver: artifact inventory — bring-funnel stage-1 INSPECT (F1-1)
artifact_inventory(api) enumerates a repo's servable artifacts WITHOUT
gating on format — a GGUF-only repo is a first-class bring (design §2b-1/2;
select_weight_files stays the vLLM/safetensors gate). GGUF variants are
enumerated at any depth, grouped by quant token with multi-part files
summed, mmproj projectors split out (never a variant); safetensors sets
reuse the adapter-excluding filter; cardData.base_model rides along
(friction #11 — lineage for ⑤'s taxonomy + credits). inspect_repo() wraps
the API fetch with structured errors; CLI: deriver.py --inventory <repo>
--json (the c3 Bring pane's INSPECT subprocess). Offline guard test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 10:01:13 +00:00
noonghunna
48238e4cbb baselines: quality_env harness provenance on quality rows (friction #8)
catalog-baseline.sh extracts runner_version from the benchlocal results
JSON and stamps quality_env: { harness: "benchlocal-cli X.Y.Z" } next to
quality_8pk — the half-deployed-sandbox lesson (§2.1.4): a quality number
without its harness fingerprint is unreproducible. sandbox_digest rides
along once benchlocal emits it (upstream candidate). Arms on DIFFERENT
harness versions → loud warn + field omitted (the arm delta isn't
comparable). Guard validates the shape; fixtures updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 09:55:54 +00:00
noonghunna
3970549ff7 catalog-baseline --from-bundle: ingest volunteer bundles (slice 3b)
Bundle mode inducts a volunteer's rebench bundle into the slug's
submissions: map with provenance FROM THE BUNDLE — rig/power from rig.txt,
engine pin from container-config.json Config.Image — never from this rig
(nvidia-smi / resolve_variant_pin would stamp our fingerprint onto foreign
numbers). --source and --submitted-by are required, no $USER default.

- splice safety both directions: a primary re-induction preserves an
  existing submissions map; bundle mode never touches the primary row
- one row per rig_class (newest replaces; history stays in git)
- multi-tag bundles refuse without --from-tag selection
- test-catalog-baseline: synthetic-bundle fixture covering refusals,
  bundle-derived provenance, add/replace, submission-only entries,
  splice preservation

Dogfood: first external row — guybrush01's #571 fp8w bundle lands as
vllm/qwen-27b-dual-max submissions[2x5090-pcie] (134.52/165.11 decode,
NIAH-clean 240,635 tok, v0.24.0, tier: submitted, TPS-only pending his
quality run).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 09:35:19 +00:00
noonghunna
3841582728 Baselines slice 3a: cross-rig submissions schema (slug × rig-class) + tier
Add the slice-3 trust-boundary schema to baselines.yml: every primary row
gains tier: local; an optional submissions: map (keyed by rig_class, e.g.
2x5090-pcie) carries cross-rig rows with required source + tier
submitted|reproduced. A slug may be submission-only (hardware we don't have).

- test-baselines: shared field validator across both row shapes; rig_class
  key format + rig-field parity; tier enums; submission-only entries legal
- registry-emit _baseline_for: submissions ride the join with per-submission
  staleness (pin comparison is rig-independent)
- c3: a submission-only baseline is NOT the bar (TPS column stays em-dash);
  detail panel renders rig-labeled, tier-badged cross-rig lines, never merged
  into the local bar

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 09:35:01 +00:00
noonghunna
deb58a5f55 fp8w on Blackwell: auto-disable DeepGEMM + --force the arch-ab arm
Two fixes from guybrush01's fp8-weights 5090 run (disc #571), which needed both
to boot the fp8w arm on consumer Blackwell.

1) VLLM_USE_DEEP_GEMM: vLLM's DeepGEMM fp8-GEMM path is built for Hopper (sm_90)
   + datacenter Blackwell (sm_100/103). Consumer Blackwell (5090 / PRO 6000 /
   GB10, sm_120/121) has no recipe and hard-fails "recipe not found" at boot.
   Add _deepgemm_env: for fp8-weights slugs, inject VLLM_USE_DEEP_GEMM=0 on the
   consumer SMs (sm_120/121 confirmed-broken; sm_89 Ada added proactively — it
   routes fp8 via Marlin/CUTLASS so disabling is a harmless no-op that pre-empts
   the same wall for 4090 owners). Hopper/datacenter untouched. dual/fp8/mtp.yml
   gains a pass-through env; both launchers whitelist the export.

2) arch-ab.sh fp8w arm: add --force. vllm/qwen-27b-dual-max is status=experimental
   so switch.sh gates it without --force.

test-launch-compat locks 5090/Ada-down / Hopper-keep / non-fp8-skip / user-pin.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 07:21:34 +00:00
noonghunna
4243a19634 Merge pull request #579 from noonghunna/feat/phase2-mem-util-floor
Phase 2: GPU_MEMORY_UTILIZATION floor for unified-memory cards (DGX Spark)
2026-07-05 11:57:45 +05:00
noonghunna
47b43f9a3a Phase 2: inject GPU_MEMORY_UTILIZATION floor for unified-memory cards
The concurrency envelope spends the KV pool, but nothing sized the pool per
card: the composes default to --gpu-memory-utilization 0.92 and the launcher
never adjusted it. For discrete cards that's fine-to-conservative (their
mem_util_safe is 0.95-0.96, above the default). For DGX Spark it's a real
safety hole: its 128 GB is unified LPDDR5X shared with the Grace CPU/OS, so
mem_util_safe is 0.85 — booting at 0.92 would take ~118 GB and starve the OS.
The seeded Spark envelope row already assumes 0.85; the launch would wrongly
use 0.92.

Add _mem_util_env (same seam as _envelope_env): inject GPU_MEMORY_UTILIZATION
DOWNWARD only — when a detected card's mem_util_safe is below the compose's
registry mem_util. Heterogeneous rigs clamp to the lowest ceiling (one GMU
applies across all ranks; a unified-memory card forces the rig down). Today
this fires for exactly one card: Spark -> 0.85.

It deliberately NEVER raises above the tested default. A 3090/5090 could give
0.95-0.96, but that changes the validated Cliff-2b margin and risks boot-OOM,
so the upward move stays a validated opt-in on the soak protocol, not an
automatic bump (and the big-card concurrency ceiling is bandwidth-bound anyway,
so the extra pool mostly buys context headroom).

Both launchers whitelist the new export; test-launch-compat locks Spark-down /
discrete-no-raise / het-min / user-pin-wins.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 06:53:27 +00:00
noonghunna
4667330425 Phase 2: validation-grade concurrency-probe (per-stream TPS, VALIDATE, SWEEP)
Enhances concurrency-probe.sh into the tool that upgrades a `computed` envelope
row to `validated` (see /opt/ai/docs/phase2-soak-validation-protocol.md). All
additive — the plain fit-check stays the default behaviour.

The design's core is two row classes, two bars, and this probe serves both:
  • pool-ceiling rows (5090): VALIDATE=1 fills each stream to the served
    --max-model-len (or TARGET_CTX), runs 6 rounds, gates fit + >=98% TPS
    retention. The value IS the kv-calc ceiling, so this is a fit+stability test.
  • bandwidth-cap rows (PRO 6000 / Spark): SWEEP="4 8 12" SLUG=... TPS_FLOOR=15
    reboots per N (vLLM can't hot-change max-num-seqs), probes decode-dominated,
    and prints the throughput KNEE — the largest clean N whose per-stream decode
    TPS clears the floor. A fit test is useless here (N=8 trivially fits 96 GB).

Key addition: streamed per-stream DECODE tok/s (TTFT-separated) — the only
honest throughput number at deep context, where prefill dominates wall time.
Retention drops the round-1 cudagraph warmup once >=4 rounds so it isn't
flattered. Machine-readable RESULT line drives the sweep's knee-finding.

SWEEP needs SLUG (refuses with exit 2 otherwise); SWEEP_DRY=1 plans the reboots
without booting. New test-concurrency-probe.sh guards syntax + refusal + dry-plan
(offline / CI-safe). Live-validated on the dev-rig Qwen: streaming TPS, floor
gate (FAIL on high floor), and VALIDATE fill all behaved.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 06:41:07 +00:00
noonghunna
e4b15fa67b Phase 2: PRO 6000 + DGX Spark envelopes, fix Blackwell detector
Follow-up to #576. Adds the two >24 GB rows deferred there and fixes the GPU
detector wrinkle that made them impossible to reach.

Detector (launch_compat.py): the `sm >= 12 -> rtx-5090` catch-all collapsed
every Blackwell to a 32 GB 5090 — a 96 GB PRO 6000 and a 128 GB GB10 both
mis-detected. Split it by SM then VRAM: sm_121 -> dgx-spark, sm_120 + >=64 GB
-> rtx-6000-pro-blackwell, else rtx-5090. Also fixed the PRO 6000 name alias
(the shipped "6000 pro blackwell" never matched the real "RTX PRO 6000
Blackwell" word order; now "pro 6000", which does NOT catch the sm_89 RTX 6000
Ada). Locked with 5 detector regression cases.

New hardware profile: dgx-spark.yml (GB10, sm_12.1, 128 GB unified LPDDR5X,
conservative mem_util for CPU-shared memory, no nvfp4 KV — same FMHA gap as
consumer Blackwell).

Rows (all `computed`, live-verified injecting):
  vllm/dual    @ rtx-6000-pro-blackwell -> 8   (2x96 GB, NVLink)
  vllm/minimal @ rtx-6000-pro-blackwell -> 16  (1x96 GB)
  vllm/minimal @ dgx-spark              -> 8   (1x128 GB)

These are OPERATIONAL CAPS, not raw pool ceilings: kv-calc gives >56 (PRO 6000
minimal) and >80 (Spark minimal), but decode concurrency is BANDWIDTH-bound
there, not pool-bound. PRO 6000 (~1.8 TB/s GDDR7 ≈ a 5090) lifts modestly above
the 5090; Spark (~273 GB/s, ≈1/7) does not — its huge pool buys context
headroom, not streams. Each row's basis cites both numbers. Spark-dual is
unseeded (2-unit clustering is ConnectX/RDMA, network-bound for TP). A soak on
real hardware refines each and upgrades it to validated.

test-profiles-compat: hardware count 9 -> 10 (dgx-spark).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 06:16:59 +00:00
noonghunna
19df532441 Phase 2: heterogeneous rigs clamp the envelope to the smallest-VRAM card
Previously a mixed-card rig no-op'd the concurrency injection ("no single
card-class row applies"). But it does have a right answer: vLLM already sizes
the KV pool to min(free blocks) across TP ranks (the cache is symmetric-
sharded), so the smallest card dictates the pool. Clamping the envelope lookup
to the smallest-VRAM card therefore MIRRORS the engine — it's the real ceiling,
not a conservative guess — and stays safe for TP=1 too (the ceiling fits
whichever single card vLLM lands on, all >= the smallest).

Behaviour:
  5090 + 3090  -> smallest 3090 has no row -> compose default (unchanged)
  5090 + H100  -> smallest 5090 is seeded  -> inject its ceiling (was: no-op)
  2x 5090      -> homogeneous               -> unchanged

The common 5090+3090 case still lands on the compose default (now for the
principled reason: the 24 GB card caps the pool), so nothing regresses; the
gain is heterogeneous rigs whose SMALLEST card is itself a seeded >24 GB class.
Fixtures lock both the new inject-on-smallest and the still-no-op paths.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 05:56:27 +00:00
noonghunna
babf1aa69c Phase 2: seed computed 5090 concurrency envelopes from kv-calc
Populate envelopes.yml with the first memory-envelope rows, computed (not
guessed) from the kv-calc pool ceiling at model-max context:

  vllm/dual    @ rtx-5090 (2x32 GB): 2 -> 4 concurrent full-262K sessions
  vllm/minimal @ rtx-5090 (1x32 GB): 1 -> 9 concurrent full-32K sessions

Both are the no-preemption ceiling: N *full-context* sequences whose KV
blocks fit the pool. Concurrency is a capacity question kv-calc answers
deterministically (arithmetic, arch-independent), and max_num_seqs is a cap
with graceful preemption (never OOM), so an over-optimistic ceiling costs an
occasional preempt, not a crash. kv-calc is sm_120-calibrated (disc #571
paulp83's real-5090 verify-stress PASS), so a 32 GB projection is trusted
arithmetic. A concurrency-soak upgrades a `computed` row to `validated`.

Guard: test-envelopes now accepts a `computed` (kv-calc basis) block as
provenance alongside `validated` (soak) — a computed row must name its basis
(invocation + PASS/cap boundary). Added a fixture proving injection is
provenance-agnostic (the launcher reads max_num_seqs; only the guard cares
about provenance).

Multi-GPU: the seam already fires on homogeneous >2-card rigs (collapses N
identical cards to one class; verified vllm/dual injects on a 4x5090 spec) and
kv-calc computes TP=4 pools; the experimental TP=4 slugs are documented as
computable-but-deferred rather than seeded. Heterogeneous rigs no-op by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 05:48:01 +00:00
noonghunna
dcf10b6cdc concurrency-probe: success = tokens generated, not non-empty content
Found testing vibethinker-3b (a reasoning model): with a small max_tokens
it spends the budget mid-<think> and returns HTTP 200 with 256 tokens but
EMPTY content (never emits the final answer). The old content-based check
wrongly flagged those streams as silent-empty failures. For a KV-pool
stress test the stream DID run (generated tokens, held KV) -> ok =
completion_tokens > 0; a true silent-empty is HTTP 200 with ZERO tokens
(what soak-test means). After the fix vibethinker sustains N=8 clean on a
3090 (8x the compose default, 0 post-warm growth).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 05:11:09 +00:00
noonghunna
65c150d562 Phase 2 (concurrency-only): memory-envelope MAX_NUM_SEQS injection + probe
The #246 Phase 2 first pass — spend a bigger card's KV pool on
concurrency, weights-invariant. Ships the machinery + measurement; the
32 GB envelope rows land from volunteer probe runs (baselines-style
plumbing-first).

Why concurrency-only (agreed 2026-07-05, design doc): most headline
configs are already at 262K model-max on 24 GB (GGUF single, dual vLLM),
so the extra 32 GB pool serves MORE STREAMS, not more context. The
context lever (bigger max_model_len) only helps single-card vLLM and is
Cliff-2b-capped anyway — deferred.

- Composes: `--max-num-seqs` parametrized to `${MAX_NUM_SEQS:-N}` on the
  two pilot slugs (dual=2, minimal=1). Weights-invariant, default kept.
- envelopes.yml (empty): measured per-(slug, card-class) `max_num_seqs`,
  born-from-measurement (a `validated` block is required — no guesses).
- launch_compat.py `_envelope_env`: injects MAX_NUM_SEQS from a row when
  it EXCEEDS the compose default; 24 GB / no-row / heterogeneous /
  user-env-set -> no-op. Both launchers whitelist MAX_NUM_SEQS.
- test-envelopes: schema + injection contract (inject / no-row /
  heterogeneous / user-env-wins / no-gain).
- concurrency-probe.sh: the measurement — N concurrent long streams x R
  rounds, separating expected pool-fill from a real leak (post-warm
  growth). Live-validated N=2 on the dual-3090 (0 MB post-warm growth,
  all rounds clean — the shipped default is sound).

Reuses the Phase 1 injection seam; independent of the KV verdict.
Full scripts gate 67/67.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 05:01:23 +00:00
noonghunna
6f674fa2ce Gate nvfp4 KV to datacenter Blackwell only (sm_100/103) — found on #571
Two 5090 owners hit `--kv-cache-dtype nvfp4 requires sm100f` crashing
mid-boot on the #246 A/B (disc #571). Root-caused from vLLM #43562 /
TRT-LLM #10241: nvfp4 KV forces the trtllm-gen FP4 FMHA, built ONLY for
datacenter Blackwell sm_100/sm_103. Consumer Blackwell (sm_120/121 —
RTX 5090 / PRO 6000 Blackwell) is a HIGHER cc number but a different
family with no FMHA build. NVFP4 *weights* work there; only the KV path
doesn't. Our #246 gate had used a numeric ">=10.0" floor that wrongly
passed sm_120 — a floor can't express "sm_100/103 but not the
numerically-higher sm_120".

- gates.py: new `_ARCH_KERNEL_SM_FAMILY` allowlist ({nvfp4: sm_100/103});
  dropped nvfp4 from the numeric `_ARCH_KERNEL_SM` floor; family-membership
  reject with the FMHA reason + fp8_e4m3 fallback.
- arch-ab.sh: nvfp4 arm now refuses on consumer Blackwell (not just
  <sm_10), naming the FMHA gap + the fp8_e4m3 path; dropped nvfp4 from the
  recommended arms in help.
- hardware profiles: removed nvfp4 from rtx-5090 / rtx-6000-pro-blackwell
  KV lists (both sm_120) + added a why-not note.
- docs (DTYPE_MATRIX / HARDWARE / KV_MATH / QUANTIZATION): corrected the
  "Blackwell sm >= 10.0" framing to "datacenter sm_100/103 only".
- UPSTREAM.md: #43562 / TRT-LLM #10241 row + re-test trigger.
- test-arch-ab: nvfp4 refuses on sm_86 AND sm_120, allowed on sm_100;
  the dual-5090 all-arms test drops nvfp4.

Full scripts gate 66/66.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 03:23:49 +00:00
noonghunna
b0c8933800 p2p verdict: point WARN/INFO at docs/PCIE_P2P.md; doc catches up with the automated verdict
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 00:10:57 +00:00
noonghunna
218a0a9a92 Interconnect verdict: warn when P2P hardware sits idle (#488 matrix)
report.sh already printed both raw inputs (host capability sections +
the [nvlink] engagement trail added after #446/#488); what was missing
was the CROSS-REFERENCE. New scripts/lib/p2p-state.sh implements the
verdict matrix once, consumed by report.sh (full verdict line) and
preflight.sh (capability one-liner):

- <2 GPUs / no capability      -> SILENT (stock-PCIe owners never nagged;
                                  that's ~95% of dual-3090 rigs and us)
- capability + engaged         -> one OK line
- NVLink bridge + P2P off      -> WARN (bridge idle, ~15% decode on the
                                  table per the #77 controlled A/B; names
                                  the fix: launcher boot / force_on)
- P2P-capable driver + off     -> INFO (launcher auto-engages since #291;
                                  residual case = direct docker compose)

The lib is the read-only AUDITOR; scripts/detect_nvlink.sh stays the
boot-time DECIDER. Their capability probes mirror each other by design,
and test-p2p-state runs BOTH against shared faked-nvidia-smi fixtures
and asserts agreement, so they cannot drift apart silently. Engagement
classification is pure (stdin text: boot trail beats env fallback).

Live-validated on this rig (stock-PCIe dual): preflight and report both
correctly silent; verdict matrix + classifier + probes covered
hermetically. Full scripts gate 66/66.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-05 00:08:03 +00:00
noonghunna
2790a7adcb STRESS_FAST mode + honest dual-rig VRAM margins (dogfood findings)
Everything the #246 runner dogfood on this rig surfaced, in one PR:

- verify-stress STRESS_FAST=1 (opt-in; gates keep full mode): skips
  probe 7's large fresh needles (near-duplicate the ladder's depths)
  and caps the ceiling ladder at 2 rungs (~2/3 anchor + ceiling).
  MEASURED on the identical vllm/dual boot: 861s -> 361s (-58%).
  Depth coverage stays 4-anchor (bench probe 10K/90K + mid + ceiling).
- VRAM margin attribution: compose `count: all` (DeviceRequests
  Count:-1) and VISIBLE_DEVICES=all now resolve as a DETERMINED
  "all host GPUs" -- kills the spurious per-rung WARN on dual rigs.
- VRAM margin aggregation: MIN per-card free instead of sum. TP OOMs
  on the first card to run out; the sum overstated dual margins ~2x
  (906 MB reported where the honest figure was 453 MB). SEMANTIC
  TIGHTENING: multi-GPU margin advisories now fire against the real
  per-card floor. Ceiling test updated to the new contract.
- arch-ab: lean tier exports STRESS_FAST=1 (per-arm ~30 -> ~20 min);
  margin advisory surfaced in the summary table + a mid-run "ladder
  CLEAN, not a failure" note so first-time runners don't abort on the
  rc=1 advisory; bundle print points at the #246 test thread.

Full scripts gate 65/65. The e4m3-below-sm_8.9 refusal is the prior
commit on this branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 22:22:54 +00:00
noonghunna
2477f53e2f arch-ab: refuse the e4m3 arm below sm_8.9 (found dogfooding on sm_86)
vLLM hard-rejects fp8_e4m3 KV below SM 8.9 at boot -- without the guard
a bare run on an Ampere rig (default arms include e4m3) sits in
switch.sh's ready-wait until the 10-minute timeout instead of getting a
clear refusal. The message says why: on Ampere there is no arch delta
to measure. --arms e5m2 (control-only) stays possible.

test-arch-ab: bare-run-on-3090 refusal + control-only-on-3090 pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:25:53 +00:00
noonghunna
3c86246844 Merge pull request #569 from noonghunna/feat/246-arch-ab-runner
Add arch-ab.sh: cross-rig KV-dtype A/B runner for #246 (lean tier)
2026-07-05 02:23:55 +05:00
noonghunna
47a9dacc1c arch-ab: rig report gains the kv-calc calibration matrix
report.sh --full-calibration instead of the bare default: seconds, no
test re-runs, and the predicted-vs-actual VRAM matrix on the volunteer
card class is direct Phase 2 input (32 GB envelope calibration).
Deliberately NOT --full -- that re-runs the whole battery (~43 min)
against only the last arm's container; the arms already carry that data.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:14:47 +00:00
noonghunna
a98a2468bf arch-ab: include the redacted rig triage report in the bundle
report.sh default mode (~2 s, path/host/user redaction ON) rides in the
tarball — PCIe lane width / driver / power caps / WSL-vs-bare-metal is
the context that makes cross-rig A/B variance interpretable (the same
mechanism that auto-flagged the x4-lane slot in #158).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:13:26 +00:00
noonghunna
b569b62d89 Add arch-ab.sh: cross-rig KV-dtype A/B runner for #246 (lean tier)
One command per volunteer rig: runs the pilot variant once per KV-dtype
arm (fresh symmetric boot each), verify-full + bench n=5 + NIAH ladder
per arm -- soak and quality packs deliberately skipped (quality
consolidates at promotion time per the #246 test plan; ~25-30 min/arm).
Ends with a per-arm comparison table and ONE attachable tarball.

Arms: e5m2 (control) / e4m3 (native FP8 KV, sm_89+) / nvfp4 (Blackwell
opt-in, refused below sm 10.0) / fp8w (vllm/qwen-27b-dual-max STOCK --
FP8-weights checkpoints reject fp8 KV, so this arm measures the
native-FP8-WEIGHTS lift instead; dual-rig only).

Guard rails: variant auto-pick by GPU count (2+ -> vllm/dual, 1 ->
vllm/minimal); KV arms are explicit KV_CACHE_DTYPE pins (unambiguous on
any launcher version); the running container's --kv-cache-dtype is
asserted live before any bench time is spent; --dry-run / --resume /
--full escape hatch.

test-arch-ab: hermetic plan/refusal contract via CLUB3090_FAKE_GPUS.
Full scripts gate 65/65.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 21:02:00 +00:00
noonghunna
e8bcfd8da1 Launcher arch-aware KV dtype injection for pilot slugs (#246 Phase 1)
launch.sh/switch.sh now detect GPU arch and export KV_CACHE_DTYPE=
fp8_e4m3 (native FP8 compute) on sm_89+ cards for the pilot slugs
(vllm/dual, vllm/minimal). The injected value is the hardware
profiles' dormant kv_format_default.balanced -- one source of truth
shared with the pull gates; the Ampere no-op is data equality
(3090-class balanced = fp8_e5m2 = compose default -> nothing emitted),
not a code branch.

Injection guards (all load-bearing):
- pilot allowlist only; expansion gated on the #246 cross-rig A/B
- per-variant: registry kv_format == fp8_e5m2 only (int8-PTH/TQ/bf16
  slugs never touched -- compressed-tensors weights reject fp8 KV)
- vLLM-family variants only; explicit user KV_CACHE_DTYPE= wins;
  unmapped cards / heterogeneous rigs / no nvidia-smi -> no injection
- direct `docker compose up` keeps Ampere-safe compose defaults

Rides the existing resolve-variant-pin export seam (new optional
--gpu-spec); VLLM_ATTENTION_BACKEND is whitelisted but ships no value
(vLLM auto-detect stays the default until measured). Preflight banner
names the detected arch class.

Consistency fixes the injection exposed:
- gates.py _ARCH_KERNEL_SM: fp8_e4m3 9.0 -> 8.9 (vLLM's real floor is
  SM89+; we'd otherwise inject e4m3 on 4090s our own pull-gate calls
  unloadable) + new nvfp4: 10.0 row (v0.24.0 literal had NO gate --
  a 3090 pull of an nvfp4 config wouldn't have been rejected)
- Blackwell hardware profiles: nvfp4 declared as CANDIDATE capability
  (engine list unchanged until validated -- gates take the intersection)
- kv-calc: projected nvfp4 rows (bytes/elem + activation coefs),
  calibration unchanged

Docs: HARDWARE.md new section, KV_MATH/QUANTIZATION/DTYPE_MATRIX rows
(incl. retiring the Genesis-era "e4m3 undertuned per #51" advisory in
favor of the A/B).

Validated: 8-case injection matrix in test-launch-compat; live no-op
on the real 2x3090 (spec built, nothing injected); faked 4090 through
the real bash detection path emits e4m3; compose interpolation both
ways; kv-calc --calibration green; full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 20:45:28 +00:00
noonghunna
5ea64bc539 Re-gate beellama/gemma-dflash on the official pin; fix rebench-full n=3
The single-card Gemma-4 default was the catalog's only default with no
baseline row (wave-2 PRIORITY gap). Full gate on Anbeeld's OFFICIAL
v0.3.2-preview digest (the current beellama-local pin), rebench tag
gemma-dflash-regate:

- decode 44.91 narr / 79.76 code (n=5, CV 2.1%/5.6%), TTFT 115 ms
- NIAH ladder clean to 117,513 tok (91% of 128K); ceiling margin 929 MB
  (< 1024 MB bar -> caveat stays, but up from ~177 MB pre-v0.3.2)
- 8-pack current harness: 108/150 off / 113/150 on (+5 in-band wash,
  think-off default re-confirmed)
- soak PASS (p50 59.8, 100.6% retention, 0 growth, 0/100 silent-empty)

Row seeded via catalog-baseline.sh (first non-vLLM engine through the
producer flow); BENCHMARKS row + compose header (Quality + margin
numbers) updated to match.

Tooling fallout the gate exposed:
- rebench-full.sh hardcoded RUNS:-3/WARMUPS:-1, silently under-running
  the bench protocol on every orchestrated gate. The A1 "n=3 env leak"
  attribution was wrong -- it was this line (baselines comment
  corrected). Fixed to protocol 5/3; n=3 had flattered gemma-dflash
  code TPS by +8%.
- catalog-baseline.sh counted prefill-probe run-lines into its bench-n
  gate (read n=8 for an n=5 log); now reads the summary headers.

Full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 19:28:03 +00:00
noonghunna
a627efa470 Seed baselines wave-2: 5 rows + A1 decode-pair fix + gap dispositions
Reviewed dispositions of the wave-1 gap list (verdict sheet approved
2026-07-04):

- vllm/dual + vllm/qwen-27b-dual-fast (alias mirror): FRESH from the
  2026-06-30 v0.24.0 pin-bump gate row (decode 70.7/93.5, NIAH->240K,
  soak PASS) — measured on the current vllm-stable pin, so no
  born-stale archaeology needed after all.
- ik-llama/iq4ks-mtp-vision: carried decode (60.39/72.4) from iq4ks-mtp
  per the 2026-05-25 vision re-tune row; carried qualifies because the
  vision delta doesn't touch the decode path.
- ik-llama/apex-fit-q8q5: 2026-05-28 row (decode 105.63/156.80, n=5),
  NIAH clean@180K.
- vllm/gemma-26ba4b-single: 2026-06-06 row (decode 169.2/219.9) + full
  8-pack (98 off / 109 think-on), NIAH clean@161K — FRESH (gemma-stable
  still pins v0.22.0).
- vllm/agents-a1-dual: decode pair order corrected from the tag
  artifact (narr 154.0 / code 153.8 — wave-1 had them swapped).
- Footer: SEED WAVE 2 list -> KNOWN GAPS dispositions (vllm/minimal,
  gemma-12b pair, beellama/gemma-dflash PRIORITY re-gate, deckard).

catalog-baseline.sh: insertion anchor updated for the retitled footer
(accepts both titles); test-catalog-baseline.sh now guarantees the ADD
path (strips the seeded vllm/dual row from its copy) and asserts added
rows land before the footer — the marker-drift class this bump caught.

Guards: test-baselines 15 rows joined, same 4 wave-1 stale WARNs, all
5 new rows fresh. Full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:56:11 +00:00
noonghunna
d19d00b9d3 quality/baselines: rescore-materialization practice — artifacts carry the accepted truth
Closes the rescore-materialization gap the induction tool's first live
run surfaced (#562): the benchlocal-cli 'rescore' subcommand ALREADY
supports write-back (--in-place / --output) — the gap was practice, not
a missing upstream feature. The T2 rescore ran to stdout only, so the
published A1 thinking-on 110/150 diverged from the tag artifact (108).

- A1's rescore MATERIALIZED into the tag JSONs (rescore --in-place; OFF
  105 unchanged, ON 108→110) — induction now extracts 110 and the
  regenerated corpus record carries 110; artifact, corpus, baseline row
  and publication all agree. Pre-rescore JSON backed up outside the tag.
- docs/QUALITY_TEST.md: 'Rescoring saved results — MATERIALIZE, don't
  just read' — the rule (a rescore that changes a published number must
  be written back in the same session), the command, what rescore can't
  re-run (sandbox packs), and the A1 case as the cautionary example.
- catalog-baseline.sh header: reads-artifacts-as-truth warning + the
  header's gate text updated to the actual n>=3/warn-under-5 behavior.
- baselines.yml A1 comment: gap note -> materialized + doc pointer.

Guards green (test-catalog-baseline, test-baselines).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:30:53 +00:00
noonghunna
34b31c0aa4 catalog-baselines slice 2c: canonical two-depth prefill/TTFT probe + anchor calibration
Decode-only TPS can't express real trade-offs anymore (the W8A8-vs-FP8
result was a prefill-corner-vs-decode-corner split) — this adds the
CANONICAL prefill/TTFT measurement to bench.sh per the design's sourcing
rule, with the protocol gotchas productized:

- bench.sh PREFILL PROBE (default-on; PREFILL_PROBE=0 / PREFILL_DEPTHS /
  PREFILL_RUNS): warm + n measured per depth (10K + 90K anchors; 90K is
  inside the DeltaNet degradation regime and pairs with the NIAH ladder's
  ~94K rung). CACHE-BUSTED: fresh salted haystack per request — composes
  serve enable_prefix_caching, an identical prompt re-measures the CACHE
  HIT (vLLM's prefix cache is block-chained; unique first line breaks the
  chain). SELF-CALIBRATING: word-count heuristics overshoot tokens ~1.3x;
  the warmup's reported prompt_toks scales the measured runs (target^2/
  actual) to within ~4% of the requested depth. Depths exceeding the
  served ctx SKIP with a note. DUAL METRIC, labeled: prompt_tokens/TTFT =
  client-observed (user-truth: incl tokenization+transfer+scheduling) AND
  the vLLM stats-log windowed rate = engine-internal (compute-truth) —
  at 93K on A1 they differ by ~7s of non-prefill overhead (5.5K vs ~10K
  t/s); never cross-compare kinds (stack LEARNINGS row added).

- measurement_record parser: per-block pass -> prefill_tps_by_ctx +
  ttft_ms_by_ctx extensions; the canonical short-prompt ttft_s is
  PROTECTED from the probe blocks (the old last-occurrence rule would
  have swallowed the 90K block's 17s TTFT).

- catalog-baseline.sh: rows gain prefill_tps {10k: N, 90k: M} (parsed
  via THE record parser, no second grammar) + ANCHOR CALIBRATION at
  induction: the probe's deep anchor vs the NIAH ladder's nearest rung —
  agreement (0.7-1.3) certifies the ladder's whole depth curve; A1 live:
  probe 5459-5584 t/s @93K vs ladder 7403 @94K = ratio 0.74-0.75, OK.
  Divergence warns with an investigate message (design: a finding).

- test-baselines schema: prefill_tps = dict of numeric depth points.
  test-catalog-baseline fixture: probe blocks + TTFT-pollution guard +
  anchor-OK assertion.

Live-validated 3x against the serving A1 (262K): 10K = 7977 t/s CV 1.1%
TTFT 1.25s; 93K = 5584 t/s CV 0.5% TTFT 16.1s; engine-log ~10K t/s.
Full scripts gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:22:48 +00:00
noonghunna
c6f56e0c88 catalog-baselines slice 2b: c3 'yours vs the bar' overlay + bar provenance/staleness detail
The consumer half of the producer wiring: THIS RIG's corpus records now
render next to the shipped bar in the catalog preview — the design's
'c3 overlays the bar with your numbers, badged'.

- LocalMeasured projection + CatalogEntry.local_measurement: the newest
  #249 corpus record per variant slug (decode point, both 8pk arms, pin,
  date). One canonical-short decode point per record — the narr/code
  pair split is a parser note for 2c.
- services.local_measurements(): pure-fs corpus scan; newest-per-slug by
  the record's _recorded_at stamp (NEW — rebench-full now stamps it at
  write; mtime+line-order fallback covers pre-stamp records); malformed
  lines skipped. Joins in the same enrichment pass — still zero
  subprocess on the catalog path.
- Preview rework: the measured line is the BAR with provenance
  (date · rig · submitted_by from the joined baseline); a stale bar gets
  an explicit detail line ('† bar measured on <pin> — current <pin>;
  re-bench owed'); a slug this rig has gated adds
  'yours ~<decode> · 8pk <P/150> (<date> · this rig)'.

Live-verified: the corpus's first record (A1, from 2a's live run) joins
and renders 'yours ~154 decode · 8pk 105/150'. Tests: overlay semantics
(newest-wins by stamp, malformed-skip, entry join, no-record → None) +
the three-state preview (fresh bar + provenance + yours · stale pin
detail · no-yours). c3 suite 766/766; scripts gate green serial.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 16:05:51 +00:00
noonghunna
1d974459c9 catalog-baselines slice 2a: induction tool + rebench auto-record + completion prompt
The producer wiring: a validated gate run now flows into the corpus
automatically and into the shipped bar via one reviewed command.

- scripts/catalog-baseline.sh (§2.3, one mechanism / three entry points):
  gate-validates the tag dir (verify-full pass HARD · bench n>=3 HARD,
  n<5 WARN vs the canonical target · quality present unless --tps-only),
  extracts decode TPS/TTFT/8pk both arms/ctx_validated (NIAH-ladder
  verdict incl. token count)/provenance (pin via the same resolution the
  emit join uses; rig+power via nvidia-smi, all overridable), UPSERTS the
  baselines.yml row textually (comments preserved), prints a unified
  diff. Never commits — rows ship via PR. --baselines-file for tests.

- rebench-full.sh: on completion, EXACT-container-match the served
  engine to a registry slug (identity semantics — never port/substring,
  the F9 rule) and append a fingerprint-complete #249 record (engine_pin
  via resolve_variant_pin/compose fallback · hardware · power-cap ·
  quality_8pk extensions · soak status) to the gitignored corpus, then
  print the catalog-baseline.sh induction prompt. BYO/swap serves skip
  with a note (no registry identity to record against).

- measurement_record.py: --quality-8pk / --quality-8pk-think-on ->
  measured_extensions (the designed extension namespace; frozen schema
  untouched).

- test-catalog-baseline.sh: hermetic (synthetic tag dir + a COPY of
  baselines.yml) — extraction, dry-run no-write, add->replace upsert,
  refusals (missing verify/quality/unknown slug), real-file checksum
  guard.

TRUTH CORRECTIONS the tool's first live run surfaced (run against the
real agents-a1-fp8-dual tag): the A1 gate's bench was n=3 (a RUNS=3 env
leak into the overnight gate — 'n=5' was misstated in publications), the
BENCHMARKS TPS pair was order-garbled (artifact: narrative 154.0 / code
153.8), and the #81-rescored think-on total (110) was never materialized
back into the tag JSON (reads 108) — materialization gap documented in
the row comment. baselines.yml + BENCHMARKS.md corrected.

Live-validated: rebench record block ran against the serving A1 → first
corpus record written (vllm-agents-a1-dual__e9d28e22.jsonl: pin v0.24.0,
rtx-3090, 370W, quality extensions); induction dry-run reproduces the
row from the tag. Deferred to 2b/2c: c3 local-overlay + richer badges;
the bench 10K/90K prefill/TTFT probe. Gates: test-catalog-baseline +
test-measurement-record + test-baselines green; full scripts gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 15:40:54 +00:00
noonghunna
75b29556c1 catalog-baselines slice 1: baselines.yml + registry-emit join + guards; catalog drops the BENCHMARKS scrape
The catalog's measured columns (TPS / 8pk) now come from ONE productized
source: scripts/lib/profiles/baselines.yml — the shipped, PR-reviewed
'bar' (accepted display projection of a validated gate run, with pin/rig/
power provenance) — joined per-slug into registry-emit --json with an
emit-computed staleness verdict. Consumers never read baselines.yml or
BENCHMARKS.md directly (BENCHMARKS stays the public human/cross-rig
ledger). Design: the catalog-baselines note (2026-07-02, §2/§5 slice 1).

- baselines.yml SEED WAVE 1: 10 rows with airtight traceability only
  (rebench tags on disk: agents-a1 golden specimen + gemma-31b-dual;
  unambiguous decode-class BENCHMARKS rows for the rest). 4 rows are
  HONESTLY born-stale with documented reasons (llamacpp rolling-tag-era
  x2, beellama pre-#296 image, 35B v0.22.0) — the guard's demo cases.
  9 slugs still owed rows are listed as wave-2 gaps, no guessed numbers.
- registry-emit join: per-variant 'baseline' field; current pin resolved
  the way launchers actually resolve it (engine-profile install.spec,
  compose-image-default fallback for ik/llama.cpp) → 'stale' =
  measured-pin != current-pin, null when undeterminable.
- test-baselines.sh: schema + slug-membership + ctx-parity (compose ctx
  default == registry max_ctx, functional slugs) RED; pin-staleness
  WARN-only (pin bumps must not block on immediate re-bench — the debt
  stays visible). test-registry-json gains the 'baseline' contract key.
- c3: enrich_measurements = pure in-memory map off the joined field —
  deletes BOTH the per-slug --explain fan-out (~4s/slug, the #439
  option-3 leg) and the BENCHMARKS.md scrape from the catalog path
  (Explain modal + cross-rig explorer keep their readers). Stale rows
  render a † on the TPS cell + a status-line legend; full badge/overlay
  treatment is slice 2. Measurement gains source='baseline' + stale.
- results/baselines/README: the three-store relationship (regression
  corpus vs measurement records vs display bar) + the no-drift rule the
  slice-2 induction tool enforces.

Verified: guard test green (10 rows, 4 stale-warned); live join emits
57 variants / 10 baselines; live c3 catalog shows all 10 with daggers on
exactly the stale 4 (A1 serving during the check: 154/154 - 105/150);
c3 suite 764/764; full scripts gate green serial. (First gate pass
tripped classifier/dedup by running CONCURRENTLY with the c3 suite —
.pull-captures pollution class; serial rerun clean.)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 15:25:36 +00:00
noonghunna
c2c06e47a7 beellama docs: sm_120 root cause + Anbeeld#85 ask + verified self-build recipe
Investigation of a Discord cross-rig report (WSL, 3060Ti+4090: verify
fails on the pinned v0.3.2-preview digest, passes on the old noonghunna
snapshot; NOT reproducible on our 3090s — verify-full all-pass) surfaced
three doc-level truths worth recording:

- The 50-series gap is a TOOLCHAIN ceiling: Anbeeld's CI builds on the
  Dockerfile-default CUDA 12.4, whose nvcc cannot target sm_120 (no
  cubin, max PTX compute_90) — every official tag lacks Blackwell.
  Upstream ask filed: Anbeeld#85 (CUDA_VERSION 12.8.1 + arch list incl
  120); when it lands we retire the noonghunna snapshot entirely.
- Three registry status_notes claimed launchers inject 'server-cuda-
  v0.3.0' — stale since the 2026-06-12 pin bump; they inject the
  v0.3.2-preview digest from engines/beellama-local.yml install.spec.
  Fixed all three (qwen dflash, gemma-12b, gemma dflash) + honest
  labeling of the noonghunna snapshot as v0.3.0-feature-level and
  unmaintained (predates KVarN + v0.3.1 fixes).
- The engine-notes self-build guidance was outdated: FA_ALL_QUANTS is
  hardcoded in Anbeeld's cuda.Dockerfile since our PR Anbeeld#48, so a
  self-build needs only CUDA_DOCKER_ARCH (+ CUDA_VERSION=12.8.1 for
  sm_120). Recipe verified against his master Dockerfile 2026-07-04.

UPSTREAM.md beellama row updated (dated entry + next-triggers; the 'no
official image' claim struck through as historical). Gates: YAML +
registry import clean; status-drift / profiles-compat / switch-parity
green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 14:17:55 +00:00
noonghunna
1e2419e320 Merge pull request #549 from titan550/feat/model-switch-service
tools: add HTTP model-switch service (thin wrapper over switch.sh)
2026-07-04 16:55:45 +05:00
noonghunna
ff96f89506 agents-a1: hardware-metadata header + Blackwell caveat (#548 follow-up)
The A1 compose shipped (#546) without the 'Hardware metadata' block every
sibling carries, so preflight_compose_gpu_fit logged 'WARN: compose has
no hardware metadata; allowing boot' and skipped enforcement — visible in
guybrush01's #548 report. Add the block (min-vram 24 / gpu-count 2 /
TP 2 / Requires-sm 8.0+ — Marlin FP8 weight-only needs sm_80, unlike the
sibling's INT4 7.5+); preflight now parses + enforces (live rc=0).

Also propagate the #548 finding as Caveats (4): Blackwell sm_120 hits an
upstream vLLM v0.24.0 kernel-selection bug at boot ('QKVParallelLinear'
has no attribute 'workspace' — Cutlass W8A8 selected, a Marlin-path
attribute expected). Workaround shipped as a commented env line
(VLLM_TEST_FORCE_FP8_MARLIN=1 — forces the exact weight-only path this
compose's gate validated; verified present in the v0.24.0 image) +
mirrored into the registry status_note so switch.sh's NOTE carries it.

Gates: compose_meta_get parses all 4 keys; docker compose config valid;
test-compose-status-drift / registry-disk / preflight-gpu-fit green;
full scripts gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 11:34:53 +00:00
noonghunna
b00b5a0955 c3/gpu-mode: preview-clip fix + scene-boundary hardening (F2+F5, T1-lite tail)
F2 — the Catalog preview strip clipped a WRAPPING caveat line: border
eats 2 rows of max-height 6 → 4 content lines; the dual-fast caveat
wrapped past that and lost its tail. max-height 6→8 + overflow-y auto
(height stays auto — short previews don't grow). Regression test pins a
wrapped-caveat entry to its full height with the tail visible.

F5 — the #544 deliberate deferrals:
- wait_gpu_vram_settle wired into ALL model/studio scene handlers
  (27b, 35b-a3b, gemma-12b, deckard, ai-studio) — was gemma-int8 only;
  every scene switch now lets the torn-down scene's VRAM release before
  the next boot (#535 class).
- mode_off gains an engine-prefix CATCH-ALL: the enumerated stop_*
  lists cover gpu-mode scenes, but a catalog-launched engine
  (switch.sh <slug>) survived 'off' — caught LIVE during validation
  when off left vllm-qwen36-27b-minimal serving and the 27b TP=2 scene
  booted straight into its residue (the exact #535 failure). Any
  remaining vllm-/llama-cpp-/ik-llama-/sglang-/beellama- container is
  now stopped, with a named notice.
- c3 preflight-error visibility (the third residual): verified
  already-plumbed — switch.sh's #544 refusal exits fast, its
  [preflight] ERROR lines stream into the serve pane, serve_failed
  stamps ✗ + [!] capture (pinned by existing tests). No change needed.

Validation: bash -n + full scripts gate green (by exit code); c3 suite
763/763; LIVE scene cycle off → 27b → off on the rig — boot ready
(qwen3.6-27b, 21.5 GiB/card), fixed off left ZERO engine containers
(verified with a non-enumerated probe container that the catch-all
stopped), both GPUs at 1 MiB after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 11:12:32 +00:00
noonghunna
a939fb1831 estate_cli: probe per-instance liveness so a leftover plan can't fake conflicts (F7)
estate.yml is a desired-state PLAN, but report-state reported its
instances as 'active' without checking docker — so every consumer
treated a leftover ~/.club3090/estate.yml as live GPU claims. The T1.1
audit hit this on an EMPTY rig: the c3 serve confirm warned '⚠ Starting
this will STOP estate llama-gpu0 (GPU [0]), estate llama-gpu1 (GPU [1])'
with nothing running. Scary + wrong for a fresh user.

estate_cli (the layer that owns estate liveness — fix here, not a c3
filter):
- report-state --json: each instance gains 'container' (the
  club3090-<name> it would own) + 'running' (docker probe: true / false
  / null when docker itself is unavailable), payload gains
  'running_count'. Additive — existing keys unchanged.
- report-state text view: per instance '— running' / '— down (plan
  only)' / '— liveness unknown'.

c3 reconcile gate (consumer): an estate instance is a claim unless
running == false. true / null / MISSING (older CLI output) all stay
claims — the dual-writer gate fails CLOSED on unknown liveness, and a
mid-boot instance (running container, VRAM not yet allocated) still
conflicts.

Tests: estate JSON contract test extended (running/container/
running_count; running asserted 'in (False, None)' so the test passes
with or without docker); 3 new c3 gate tests (stale plan on empty rig →
SAFE; running:true → claim even with idle GPUs; running:null → fail
closed; the legacy no-key fixtures double as the missing-key case).
Scripts gate all green; c3 suite 758/758. Live repro on this rig (which
has the exact stale estate.yml): running_count 0, reconcile now reports
estate_claims=[] while keeping the TRUE conflict (the actually-serving
vllm container).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 10:35:00 +00:00
noonghunna
d3836ecd87 c3+switch: show/copy the serving API URL; fast-flip booting→ready (F3/F3b)
The T1 audit's top adoption gap: even while a model SERVES, the best
endpoint info anywhere was a bare port — no scheme/host/path, no auth
note, no copy. "Point your agent at this URL" is the whole reason the
stack exists, and the URL was underivable from the UI.

One shared derivation (layer rule; #512 precedence env/.env ->
c3_lan_ip subshell -> localhost), four surfaces:
- switch.sh: ready-line on successful boot —
  "▶ API: http://<lan>:<port>/v1 (model: <served-id> · OpenAI-compatible
  · no auth)" — CLI parity; served id probed from /v1/models.
- services.lan_ip(): c3 consumes the SAME derivation, session-cached.
- Orchestration serving card: bare ":8020" -> full URL + "no auth ·
  [u] copy".
- Estate rail: compact "api :8020/v1 · [u] copy" (rail is ~30 cols; the
  full URL lives on the card).
- NEW [u] copy-API-URL key on every Run & Operate tab (kept separate
  from [Y] so row-copy semantics are untouched); honest notify no-op
  when nothing serves.

F3b (readiness lag): api_booting only cleared when the HEAVY docker+
health batch re-ran, so " booting" outlived actual readiness by a poll
cycle (audit: >=25s stale across three surfaces). While booting, the
fast GPU tick now piggybacks a 1.5s /v1/models probe; on 200 it pulls
the next heavy poll forward (burst regime) — every booting surface
flips within a tick or two of the API answering. Bounded: fires only in
the booting state.

Verified: c3 pytest suite 516 passed; live serve via switch.sh prints
the ready-line with the real LAN IP + the neutral served name
(http://192.168.86.33:8020/v1 · qwen3.6-27b · no auth).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 09:09:12 +00:00
John Shojaei
c22a9d2d84 tools: add HTTP model-switch service (thin wrapper over switch.sh)
Adds a stdlib HTTP control plane that wraps scripts/switch.sh so a harness
can POST /switch and block until the new model is serving. Introduces no new
orchestration logic — switch.sh stays the single source of truth (registry
lookup, down/up, readiness).

- tools/model-switch/server.py: GET /healthz|/status|/models, POST /switch
  ({slug}|{model}); registry-validated; /health readiness (works with or
  without VLLM_API_KEY); single-flight lock; refuses to start unauthenticated
  on a non-loopback bind.
- scripts/systemd/club3090-model-switch.service: host daemon unit.
- scripts/tests/test-model-switch.sh: hermetic HTTP/auth/validation contract.
- docs/EXAMPLES.md, .env.example: usage + config.

Mirrors the existing stdlib HTTP style (services/studio/*); zero new deps.
Experimental/opt-in per the repo's staging convention.
2026-07-03 17:44:17 -07:00
noonghunna
e4d33e09be Merge pull request #546 from noonghunna/feat/agents-a1-promotion
Promote Agents-A1 to the catalog: vllm/agents-a1-dual (⚠️ production w/ caveats)
2026-07-03 08:53:11 +05:00
noonghunna
6926dfde48 Promote Agents-A1 to the catalog: vllm/agents-a1-dual (production w/ caveats)
InternScience's 35B agentic MoE (Qwen3-Next MoE arch; the card declares its
OWN base -> own model entry, not a qwen fine-tune slug), served from the
official FP8-dynamic compressed-tensors checkpoint on stock vLLM v0.24.0,
dual 3090 TP=2, full 262K. The agentic thinking-ON specialist.

First model onboarded END-TO-END through the Bring & Validate lane (T2
producer-zero): pull.sh route-C sibling swap (post deriver fix) ->
generate-compose -> documented swap -> full gate -> ④ vs-bar -> ⑤ scaffold.

Full gate PASS (rebench tag agents-a1-fp8-dual, 2026-07-03):
- bench 153.9/154.0 decode TPS (n=5, CV 0.1%, TTFT ~130ms), ~22.0 GB/card
- verify-stress 8/8 — staggered NIAH exact-recall to 240K (91%), VRAM Δ0
- soak PASS (0 growth, 0/100 silent-empty, 99.8% retention)
- 8-pack --full OFF 105/150 · ON 110/150 (post benchlocal #79+#81 harness):
  toolcall 15/15 OFF · IF 15/15 ON · cli-40 thinking-ON 23/40 (the stack's
  highest; base 17/40) · hermes 12/20 OFF, REGRESSES to 9/20 thinking-ON
  (verified model behavior — disclosed as caveat 1)

vs qwen3.6-35b-a3b: general capability TIES (ON 110=110), decode −13% —
NOT a general upgrade; reach for it on tool/CLI-agent work thinking-ON.

Catalog wiring: compose dual/fp8-dynamic/fp8.yml (non-default KV named per
convention; OWN port 8072 — sibling-shared ports masquerade the slug in
estate detection); model profile (geometry verified == 35B-A3B from
config.json; mtp_num_hidden_layers=0 — safetensors header shows 0 mtp
tensors, config's 1 is an inherited default); registry entry status=caveats;
kv-calc agents-a1 spec + alias (fit-all prices, calibration 100%);
weights.py aliases (hf auto-fetch); LiteLLM route :8072; BENCHMARKS section;
count bumps (registry 57, disk 58, models 11).

Full catalog suite 60/60 (1 = known worktree-fixture submit-bench).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-03 03:50:14 +00:00
noonghunna
15a4e75180 deriver: resolve bit-width from compressed-tensors config_groups
pull.sh (and c3's ① Bring pane) rejected EVERY llm-compressor checkpoint
(FP8-dynamic, INT8 W8A8, INT4 pack-quantized) with `quant-dtype-unknown
(bits undeterminable)`: _quant_bpw only read top-level bits keys + method-
name heuristics, and "compressed-tensors" as a method matches none — but
the bits live nested at config_groups.<g>.weights.num_bits. Parse them
(widest group wins — mixed-precision groups exist; the widest dominates
the VRAM footprint the fit-check prices).

Found by the T2 producer-zero dogfood (Agents-A1-FP8-dynamic): step ①
hard-stopped with arch=null + swap_path=null; post-fix the lane resolves
the arch and emits the honest route-C verdict (sibling qwen3.6-35b-a3b,
BRING_YOUR_OWN swap pointer) — validated live, and the swapped compose
boots + serves on stock v0.24.0 (Marlin weight-only FP8 MoE on sm_86).

+3 deriver test cases: FP8-dynamic (the A1 shape), INT4 pack-quantized,
mixed groups -> widest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-02 23:24:28 +00:00
noonghunna
05f1076265 fix(#535): fail fast + actionable when gpu_memory_utilization doesn't fit free VRAM
Switching gpu-mode ai-studio → gemma on a desktop rig (GNOME running on the GPUs,
~1.3 GiB/card) leaves < the 0.95 util budget free, so gemma failed vLLM's boot
free-memory check (`free < util × total`); with restart:unless-stopped it restart-
looped and switch.sh waited the full 600s READY_TIMEOUT for an endpoint that never
came. The existing gpu_preflight only enforces a coarse 80%-free floor (the reporter
had ~90% free), so it sailed through. qwen (0.92) fit → that path worked, which is
why *only* gemma broke.

- preflight.sh: new preflight_compose_gpu_fit — parses the effective
  GPU_MEMORY_UTILIZATION (env override or compose default) + TP card-count, compares
  per-card free VRAM to `util × total × 0.98` (~ vLLM's own mem_get_info total), and
  HARD-fails (unless --force) with an actionable message (free VRAM / lower util). A
  ~10s settle-retry covers teardown lag (docker `down` returns before CUDA frees).
- switch.sh: call it in up_variant (vllm-only, after down_running so the retry also
  covers the just-torn-down container). launch.sh delegates to switch.sh → covered.
- gpu-mode.sh: wait_gpu_vram_settle after scene teardown in the gemma handler, so the
  incoming TP=2 model boots into freed VRAM instead of ai-studio's residue.
- test-preflight-gpu-fit.sh: mocks nvidia-smi + the #535 numbers — fit / short+message
  / --force / env-override / single-card.

Full shell gate 60/60 (1 = known worktree-fixture). Validated end-to-end against the
real gemma compose: idle rig fits; simulated 22060 MiB free → instant "GPU 1 has 21.5
GiB free, needs ~22.3" + lower-util hint (matches vLLM's 22.38 threshold).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-02 12:31:35 +00:00
noonghunna
7e9f5cb930 serve: neutral primary model name (qwen3.6-27b / gemma-4-31b), keep -autoround alias
The shared served-model-name `qwen3.6-27b-autoround` (from #490) mislabels the
non-autoround 27b scenes (fp8 dual-max, lmcache): /v1/models advertises
"autoround" while `root` points at qwen3.6-27b-fp8. Same class on gemma-4-31b —
after the v0.24.0 consolidation the default is cyankiwi qat-AWQ-INT4 (bf16 KV),
yet the LiteLLM route still targeted `gemma-4-31b-autoround` (a latent #482 drift).

Fix WITHOUT breaking anything, via vLLM multi-served-name:
- Every 27b scene now serves `qwen3.6-27b <its-quant-name>`; every gemma-31b
  scene serves `gemma-4-31b <its-quant-name>`. The neutral name is PRIMARY
  (honest /v1/models id); the quant-specific name is retained as a live ALIAS.
- LiteLLM: add `qwen3.6-27b` / `gemma-4-31b` canonical public routes; keep the
  `-autoround` routes as back-compat aliases (same upstream). Repairs the gemma drift.
- Migrate our own MODEL= defaults + docs (bench/verify/quality/launch/setup, c3,
  tui-core, EXAMPLES, ...) to the neutral name. Weights slugs (`-autoround-int4`)
  untouched; CHANGELOG + results/ history left as-is.

Retiring the `-autoround` alias entirely is a deliberate later step once nothing
still asks for it.

Live-verified on-rig (single/minimal, stock v0.24.0): /v1/models lists BOTH names
(root=...-autoround-int4); chat to `qwen3.6-27b` AND `qwen3.6-27b-autoround` both
return 200; `qwen3.6-27b-fp8` correctly 404s. Full shell gate 59/59 (1 = known
worktree-fixture); c3 pytest 41 passed; served-name arg-order + YAML validated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-02 11:28:31 +00:00
noonghunna
ae4d1fcad6 patches: de-register vllm-marlin-pad (merged upstream, native in v0.24.0)
Our vllm#40361 sub-tile-n Marlin pad was closed-superseded by mgoin's
vllm#45295 (consolidated marlin_padded_nk across all dense Marlin paths),
native in vLLM v0.24.0. `vllm-stable` now pins v0.24.0 and the AutoRound
INT4 TP=2 path (vllm/dual) boots clean without the overlay (validated in
#533 Phase 0b). No live compose mounts the patch (archive-only), so this
is a tracking-only de-registration — the cleanup deferred from #533.

- patches.yml: qwen-vllm-marlin-pad -> deprecated (upstream.status
  open->merged, load_bearing_when [], delivery none, drift_guard null),
  mirroring the gemma-vllm-pr41800 merged-and-dropped precedent. Kept as
  history (foundational false; entry not deleted).
- arch_patches.yml: correct the stale kernel_constraints note (#40361 ->
  #45295 native in v0.24.0). required_patches / marlin_alignment_required
  unchanged: the alignment is a real arch property (now satisfied stock),
  and deprecated patches stay listed per the pr41800 precedent.
- UPSTREAM.md: mark the #40361 / #40354 / v0.24.0-bump marlin rows DONE
  (native in v0.24.0, patch de-registered, no live mount) and correct the
  stale "composes still mount it" line (all mounts are under _archive/).

Full shell gate green (59/59; the 1 = known worktree-fixture-absent
test-submit-bench). test-patch-attribution (reads both registry files) passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-02 10:03:44 +00:00
noonghunna
a6373533e1 Merge pull request #539 from noonghunna/feat/diffusiongemma-v0.24.0
DiffusionGemma → stock vLLM v0.24.0 (drop the :gemma branch digest)
2026-07-02 13:37:39 +05:00