The prior 4-card validation (@Whamp #446) was on the OLD int8-PTH KV. #595 flips
to fp8/e4m3; the fp8 config is validated on the 2-card dual-max proxy (all gates
green) but not yet re-confirmed at TP=4. Request a fresh 4-card report before
upgrading ⚠️ caveats -> ✅.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Comment-only polish on the KV flip (no runtime change):
- fp8 -> e4m3 runs at scale=1.0 (checkpoint is weight-only; calculate_kv_scales
is disabled on Qwen3-Next hybrid), not "loads the checkpoint's scales"
- must be `fp8` not `fp8_e5m2` (e5m2 hard-rejected with fp8 checkpoints)
- backend FlashInfer (int8-PTH is TRITON_ATTN-only) -> flat decode at depth
- quality 109 ties int8-PTH 107; soak-continuous PASS (0 growth, 509 MB margin)
- fix comparison-table column spacing
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
guybrush01 ran the full 8-pack (thinking-off) on dual-max on his 2x 5090:
102/150 — within +-5-7 noise of our 2x3090 fp8 (107). So FP8 is near-lossless
on native Blackwell too, which closes the 'route Blackwell -> FP8 weights'
recommendation gate (fp8 quality confirmed on BOTH Ampere-Marlin and native
Blackwell fp8).
- baselines.yml: guybrush's 2x5090-pcie dual-max submission gains
quality_8pk: 102/150 + quality_env (harness fingerprint).
- One dip noted: dataextract 9/15 (vs 13 elsewhere) — mostly verifier_fail,
the known DE brittleness cluster, not an obvious fp8 regression. thinking-on
run pending.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
@ryanmpelletier rebuilt the benchlocal sandboxes and ran the full 8-pack on
his 4x 3090 (v0.24.0, thinking-off): TOTAL 108/150 — within +-1 of the 109/150
we carry from the 2-card fast tier (vllm/dual). So multi-fast's quality is now
MEASURED TP-invariant on real 4-card hardware, not assumed.
- baselines.yml: ryan's 4x3090-pcie submission gains quality_8pk: 108/150 +
quality_env (harness provenance — the first quality ingest into a submission
row, the friction-#8 / slice-3e hook: a quality number carries its harness
fingerprint so it's reproducible).
- compose header Quality: 'open follow-up' -> the measured 108/150 confirmation.
Closes the last open item on the multi-fast promotion (bench + soak + quality
all confirmed on 4-card). Does NOT touch multi-max's caveat (that needs an
fp8 4-card report, not INT4).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The 4-card FP8 max tier is byte-identical to the now-production
vllm/qwen-27b-dual-max apart from TP=4 + gpu-count. We can't self-validate
(2-card dev rig), so it rides @Whamp's cross-rig full chain — #446, 4x 3090:
verify-full + verify-stress 7/7 + soak-continuous PASS (85/102).
Status experimental -> ⚠️ Production w/ caveats (registry 'caveats' + header
Caveats line + drift guard green). The caveat, stated honestly: that
validation was on an OLDER engine (pre-v0.24.0 pin) + a non-standard rig
(aikitoria P2P kernel, mixed x4/x16/x8/x16 lanes), single report — no clean
v0.24.0 4-card datapoint yet. A fresh one (like @ryanmpelletier's for
multi-fast) upgrades it to bare ✅.
No baseline row inducted: @Whamp's number is under-specified (engine version
unstated, non-standard rig), so inducting it as THE bar would mislead — the
status_note carries the provenance + caveats instead. No DEFAULTS change.
Removed the now-contradictory '🧪 Experimental until a ≥4x host validates it'
prose from the header.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The FP8 max-accuracy tier had bench (83.1/108.2) + quality (107/150) on the
2x3090 reference rig but no soak-continuous run — the one missing gate item.
Ran the full operational gate fresh on the v0.24.0 pin:
- verify-full 9/9
- verify-stress: all 6 rungs, fillable to 240,636 tok clean (91%)
- soak-continuous PASS: 0 err / 0 growth / 100% retention, p50 decode 85
That clears the production bar.
- registry + compose header: status experimental -> production (drift guard
green). status_note also de-staled: dropped the '~56 TPS' probe / 'slowest of
the three' framing (corrected to decode 83/108, the slow axis is PREFILL/TTFT
from MarlinFP8 W8A16 on Ampere, not decode) + the full-gate results.
- baselines.yml: row comment notes the soak PASS completing the gate.
- No DEFAULTS[(qwen,vllm,dual)] change -> vllm/dual (fast) stays the dual
default; dual-max just joins the actionable list as the max-fidelity tier.
Bonus for the 5090 crowd: dual-max is now a non-experimental config, so the
launch command drops --force (bash scripts/switch.sh vllm/qwen-27b-dual-max).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
#588 fixed the fp8-weights DeepGEMM crash but the _deepgemm_env gate (and my
guard) matched weights_variant == 'fp8' EXACTLY — missing agents-a1, whose
weights are 'fp8-dynamic' (compressed-tensors FP8). Those are genuine FP8
weights that route FP8 GEMM, so a 5090 running vllm/agents-a1-dual would still
hit the 'recipe not found' crash the whole framework is meant to prevent.
- _deepgemm_env gate: weights_variant == 'fp8' -> startswith('fp8'), catching
both 'fp8' and 'fp8-dynamic' (compressed-tensors). INT4/AWQ/W8A8/bf16 never
route FP8 GEMM so they stay correctly excluded.
- agents-a1 compose: add the - VLLM_USE_DEEP_GEMM pass-through.
- test-deepgemm-fp8-parity: broadened to startswith('fp8') so it now covers
agents-a1 (5 fp8-family composes) + any future fp8-* variant.
- test-launch-compat: new case — fp8-dynamic on 5090 injects VLLM_USE_DEEP_GEMM=0.
Now ALL fp8-family vLLM slugs are 5090-safe via the launcher (dual-max,
multi-max, dual-lmcache, diffusiongemma-dual, agents-a1). Convention: fp8
weight variants must be named fp8-* for the gate/guard to catch them.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
#580 added the consumer-Blackwell DeepGEMM auto-disable pass-through to
dual-max, but the other fp8-weights vLLM composes — hand-maintained
parallel copies, no extends/include — had silently drifted without it:
multi-max, dual-lmcache, diffusiongemma-dual. On a 5090 / PRO 6000 / GB10
(sm_120/121) the launcher's _deepgemm_env injects VLLM_USE_DEEP_GEMM=0, but
with no receiving line docker-compose drops it → the fp8-GEMM 'recipe not
found' boot crash #580 fixed for dual-max (disc #571).
- Add the pass-through (+ the shared comment) to all
three. All 4 fp8-weights vLLM composes now carry it.
- NEW test-deepgemm-fp8-parity: asserts every fp8-weights vLLM compose has
the pass-through — REDs on the exact drift class that caused this (verified
it fails when the line is removed, passes when restored). Scope matches the
_deepgemm_env gate (weights_variant==fp8 + vllm engine); INT4/AWQ/bf16 never
invoke DeepGEMM so they're correctly excluded.
Removes ONE 5090 blocker (fp8 crash) — NOT a blanket 'runs on 5090': beellama
composes stay broken on sm_120 (Anbeeld#85, CUDA 12.4), llama.cpp/INT4 on
sm_120 remain unvalidated (no 5090 on the reference rig).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
@ryanmpelletier's flawless 4x 3090 report (#584) is the cross-rig
validation the compose header required ('until a real >=4x 3090 host
validates it'): v0.24.0, no power caps, all x16, verify-full 9/9,
verify-stress needles clean to 240K, soak-continuous PASS, bench n=5.
Third concordant 4-card validation (after @alanspires 6xVFIO + @Whamp),
first on v0.24.0.
- registry + compose header: status experimental -> production (both, drift
guard green). No DEFAULTS[(qwen,vllm,multi4)] entry -> no resolver cascade;
promotion just moves it onto the actionable list. Quality is TP-invariant,
carried from the vllm/dual proxy (109/150); 4-card 8-pack is the open
confirmation (report.sh runs none).
- baselines.yml: vllm/qwen-27b-multi-fast gains its 4x3090-pcie submission
(tier: submitted; decode 74.76/90.83, prefill 1288->1175, KV pool
1.77M/6.77x, NIAH 240K). Submission-only by construction - the 2-card
reference rig can never produce a local 4-card row.
- BENCHMARKS.md: ryan's row (all-x16 74.8/90.8 beats @Whamp's mixed-lane
59/74 on the same fp8-KV config -> PCIe lane width matters for TP=4).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The dual-max FP8 2x3090 decode bench (83.1/108.2, v0.24.0, 2026-06-30) lived
only in BENCHMARKS.md:122 + disc #515 (c17490565) — never inducted into
baselines.yml, so c3 / the registry-emit join showed only guybrush01's
2x5090 submission and no local bar. Surfaced while diagnosing #585.
- baselines.yml: vllm/qwen-27b-dual-max gains its primary tier: local row
(83.1/108.2 · TTFT 158 · prefill 1364->875 · 8-pack 107/150 · NIAH 240K ·
v0.24.0 = current pin -> FRESH); guybrush's 2x5090 submission preserved.
- compose header + DUAL_CARD.md: the stale '~56 TPS' probe (and the now-false
'slowest of the three' framing) -> real decode 83/108; the genuine
tradeoffs (smallest KV pool 295K/1.13x, slowest prefill/TTFT 158ms from
FP8's compute-heavy Marlin W8A16 dequant) kept. BENCHMARKS.md already
corrected; this closes the two spots that still read ~56.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Three maintainer frictions from live round 2:
- label = topology/engine/model-quant ONLY — the serving-stem tail
duplicated the path axes and read as a second slug (some registry rows
carry a subpath in the file field: 'dual/piehsoft-q6k/mtp'). The stem
(basename only) is appended SOLELY to disambiguate genuine collisions
(two slugs sharing all three axes, e.g. fp8-mtp vs turbo)
- every input carries a TITLE (bare dropdowns read as anonymous fields):
field = Vertical(title + widget); titles toggle WITH their widgets so
the staged reveal stays intact
- NEW selected-slug detail card beside the HF-repo verdict: status · ctx
· port · drafter/vision · the shipped bar (TPS/8pk + provenance +
staleness dagger) · caveat note — follows the selection, hides on the
custom-slug sentinel
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Live-dogfood round 1: the slug default came from the rig-topology rule
(2 cards → dual), recommending a DUAL slug for a 5 GiB gguf — reads as
'the UI wants this on two cards'. New funnel_recommended: the smallest
fitting topology (options are size-floored + topology-sorted, so its
first group = the cheapest config that holds the artifact), preferring
the registry's curated default for the (family, topology) within it,
else the first functional-status option. The pick is starred (⭐) AND
pre-selected; the full filtered list stays reachable below it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Live-dogfood round 1 (Qwythos-9B-Claude-Mythos-5-1M-GGUF): the repo ships
base + MTP builds per quant (…-Q4_K_M.gguf / …-MTP-Q4_K_M.gguf); token-
keyed grouping merged them into ONE "2-part" variant with a summed, wrong
size (10.7 GiB shown for a 5.2 GiB pick) and no way to select just one.
Group by STEM instead (basename minus the -NNNNN-of-NNNNN part suffix) —
true multi-part shards share a stem so 'parts' still counts them; distinct
artifacts don't. Label = stem minus the repo-wide common prefix (standard
repos keep their plain quant token; multi-artifact repos keep the
distinguishing part: Q4_K_M vs MTP-Q4_K_M). Guard case added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The Bring page reveals nothing template-side until the artifact is known:
- [Inspect] runs the deriver's artifact inventory (HF metadata only, never
downloads) — a GGUF-only repo is a first-class bring, no longer
unsupported-format; lineage (base_model) rides the verdict (friction #11)
- GGUF repos: ALL discovered quants presented (sizes, multi-part counts;
mmproj split out) — the pick comes BEFORE any slug appears; the pick
Select starts BLANK so only a genuine user pick reveals stage 3
- slug Select: EVERY catalog option passing the ABSOLUTE artifact→engine
compat filter (GGUF never sees vLLM; safetensors never sees the llama.cpp
family), labeled topology-first (topology/engine/model-quant · serving),
sorted so models group under topology/engine; custom-slug escape kept
- topology floor (maintainer rule): hide topologies whose total VRAM can't
hold the weights — rig-relative (cards × min-card-VRAM × 0.90), never
fixed thresholds; ONE-directional (larger topologies never hidden —
small-quant-on-more-GPUs is legitimate); unhostable card-counts hidden;
unknown size/rig → no floor (never guess-hide)
- fit-check success surfaces the weights state: on disk → explicit ② Serve
handoff; absent → [D] downloads via the REAL pull.sh (SHA-verified,
streamed into the pane; disk write, no GPU claim), re-probed on completion
- presence probe mirrors downloader.py pull_dir/sanitize_slug with a
parity drift-guard test; HF_HOME pinned identically for probe + download
c3 suite 779/779 (11 new: staged reveal, gguf pick flow, compat filter,
label/sort, size floor, presence probe, weights line + handoff).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
artifact_inventory(api) enumerates a repo's servable artifacts WITHOUT
gating on format — a GGUF-only repo is a first-class bring (design §2b-1/2;
select_weight_files stays the vLLM/safetensors gate). GGUF variants are
enumerated at any depth, grouped by quant token with multi-part files
summed, mmproj projectors split out (never a variant); safetensors sets
reuse the adapter-excluding filter; cardData.base_model rides along
(friction #11 — lineage for ⑤'s taxonomy + credits). inspect_repo() wraps
the API fetch with structured errors; CLI: deriver.py --inventory <repo>
--json (the c3 Bring pane's INSPECT subprocess). Offline guard test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
catalog-baseline.sh extracts runner_version from the benchlocal results
JSON and stamps quality_env: { harness: "benchlocal-cli X.Y.Z" } next to
quality_8pk — the half-deployed-sandbox lesson (§2.1.4): a quality number
without its harness fingerprint is unreproducible. sandbox_digest rides
along once benchlocal emits it (upstream candidate). Arms on DIFFERENT
harness versions → loud warn + field omitted (the arm delta isn't
comparable). Guard validates the shape; fixtures updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
A NEW model has no same-model bar BY DEFINITION — the lane's primary case.
measure_vs_bar now accepts class_hint (①'s swap_path sibling, wired from
the session's fit-check) and falls back to the labeled CLASS bar:
bar_is_class + class_model ride the struct, a caveat says the deltas are
class-relative, engine-matching still applies. Without a hint the no-bar
caveat points at ① fit-check instead of dead-ending.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Bundle mode inducts a volunteer's rebench bundle into the slug's
submissions: map with provenance FROM THE BUNDLE — rig/power from rig.txt,
engine pin from container-config.json Config.Image — never from this rig
(nvidia-smi / resolve_variant_pin would stamp our fingerprint onto foreign
numbers). --source and --submitted-by are required, no $USER default.
- splice safety both directions: a primary re-induction preserves an
existing submissions map; bundle mode never touches the primary row
- one row per rig_class (newest replaces; history stays in git)
- multi-tag bundles refuse without --from-tag selection
- test-catalog-baseline: synthetic-bundle fixture covering refusals,
bundle-derived provenance, add/replace, submission-only entries,
splice preservation
Dogfood: first external row — guybrush01's #571 fp8w bundle lands as
vllm/qwen-27b-dual-max submissions[2x5090-pcie] (134.52/165.11 decode,
NIAH-clean 240,635 tok, v0.24.0, tier: submitted, TPS-only pending his
quality run).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Add the slice-3 trust-boundary schema to baselines.yml: every primary row
gains tier: local; an optional submissions: map (keyed by rig_class, e.g.
2x5090-pcie) carries cross-rig rows with required source + tier
submitted|reproduced. A slug may be submission-only (hardware we don't have).
- test-baselines: shared field validator across both row shapes; rig_class
key format + rig-field parity; tier enums; submission-only entries legal
- registry-emit _baseline_for: submissions ride the join with per-submission
staleness (pin comparison is rig-independent)
- c3: a submission-only baseline is NOT the bar (TPS column stays em-dash);
detail panel renders rig-labeled, tier-badged cross-rig lines, never merged
into the local bar
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Two fixes from guybrush01's fp8-weights 5090 run (disc #571), which needed both
to boot the fp8w arm on consumer Blackwell.
1) VLLM_USE_DEEP_GEMM: vLLM's DeepGEMM fp8-GEMM path is built for Hopper (sm_90)
+ datacenter Blackwell (sm_100/103). Consumer Blackwell (5090 / PRO 6000 /
GB10, sm_120/121) has no recipe and hard-fails "recipe not found" at boot.
Add _deepgemm_env: for fp8-weights slugs, inject VLLM_USE_DEEP_GEMM=0 on the
consumer SMs (sm_120/121 confirmed-broken; sm_89 Ada added proactively — it
routes fp8 via Marlin/CUTLASS so disabling is a harmless no-op that pre-empts
the same wall for 4090 owners). Hopper/datacenter untouched. dual/fp8/mtp.yml
gains a pass-through env; both launchers whitelist the export.
2) arch-ab.sh fp8w arm: add --force. vllm/qwen-27b-dual-max is status=experimental
so switch.sh gates it without --force.
test-launch-compat locks 5090/Ada-down / Hopper-keep / non-fp8-skip / user-pin.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The concurrency envelope spends the KV pool, but nothing sized the pool per
card: the composes default to --gpu-memory-utilization 0.92 and the launcher
never adjusted it. For discrete cards that's fine-to-conservative (their
mem_util_safe is 0.95-0.96, above the default). For DGX Spark it's a real
safety hole: its 128 GB is unified LPDDR5X shared with the Grace CPU/OS, so
mem_util_safe is 0.85 — booting at 0.92 would take ~118 GB and starve the OS.
The seeded Spark envelope row already assumes 0.85; the launch would wrongly
use 0.92.
Add _mem_util_env (same seam as _envelope_env): inject GPU_MEMORY_UTILIZATION
DOWNWARD only — when a detected card's mem_util_safe is below the compose's
registry mem_util. Heterogeneous rigs clamp to the lowest ceiling (one GMU
applies across all ranks; a unified-memory card forces the rig down). Today
this fires for exactly one card: Spark -> 0.85.
It deliberately NEVER raises above the tested default. A 3090/5090 could give
0.95-0.96, but that changes the validated Cliff-2b margin and risks boot-OOM,
so the upward move stays a validated opt-in on the soak protocol, not an
automatic bump (and the big-card concurrency ceiling is bandwidth-bound anyway,
so the extra pool mostly buys context headroom).
Both launchers whitelist the new export; test-launch-compat locks Spark-down /
discrete-no-raise / het-min / user-pin-wins.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Enhances concurrency-probe.sh into the tool that upgrades a `computed` envelope
row to `validated` (see /opt/ai/docs/phase2-soak-validation-protocol.md). All
additive — the plain fit-check stays the default behaviour.
The design's core is two row classes, two bars, and this probe serves both:
• pool-ceiling rows (5090): VALIDATE=1 fills each stream to the served
--max-model-len (or TARGET_CTX), runs 6 rounds, gates fit + >=98% TPS
retention. The value IS the kv-calc ceiling, so this is a fit+stability test.
• bandwidth-cap rows (PRO 6000 / Spark): SWEEP="4 8 12" SLUG=... TPS_FLOOR=15
reboots per N (vLLM can't hot-change max-num-seqs), probes decode-dominated,
and prints the throughput KNEE — the largest clean N whose per-stream decode
TPS clears the floor. A fit test is useless here (N=8 trivially fits 96 GB).
Key addition: streamed per-stream DECODE tok/s (TTFT-separated) — the only
honest throughput number at deep context, where prefill dominates wall time.
Retention drops the round-1 cudagraph warmup once >=4 rounds so it isn't
flattered. Machine-readable RESULT line drives the sweep's knee-finding.
SWEEP needs SLUG (refuses with exit 2 otherwise); SWEEP_DRY=1 plans the reboots
without booting. New test-concurrency-probe.sh guards syntax + refusal + dry-plan
(offline / CI-safe). Live-validated on the dev-rig Qwen: streaming TPS, floor
gate (FAIL on high floor), and VALIDATE fill all behaved.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Follow-up to #576. Adds the two >24 GB rows deferred there and fixes the GPU
detector wrinkle that made them impossible to reach.
Detector (launch_compat.py): the `sm >= 12 -> rtx-5090` catch-all collapsed
every Blackwell to a 32 GB 5090 — a 96 GB PRO 6000 and a 128 GB GB10 both
mis-detected. Split it by SM then VRAM: sm_121 -> dgx-spark, sm_120 + >=64 GB
-> rtx-6000-pro-blackwell, else rtx-5090. Also fixed the PRO 6000 name alias
(the shipped "6000 pro blackwell" never matched the real "RTX PRO 6000
Blackwell" word order; now "pro 6000", which does NOT catch the sm_89 RTX 6000
Ada). Locked with 5 detector regression cases.
New hardware profile: dgx-spark.yml (GB10, sm_12.1, 128 GB unified LPDDR5X,
conservative mem_util for CPU-shared memory, no nvfp4 KV — same FMHA gap as
consumer Blackwell).
Rows (all `computed`, live-verified injecting):
vllm/dual @ rtx-6000-pro-blackwell -> 8 (2x96 GB, NVLink)
vllm/minimal @ rtx-6000-pro-blackwell -> 16 (1x96 GB)
vllm/minimal @ dgx-spark -> 8 (1x128 GB)
These are OPERATIONAL CAPS, not raw pool ceilings: kv-calc gives >56 (PRO 6000
minimal) and >80 (Spark minimal), but decode concurrency is BANDWIDTH-bound
there, not pool-bound. PRO 6000 (~1.8 TB/s GDDR7 ≈ a 5090) lifts modestly above
the 5090; Spark (~273 GB/s, ≈1/7) does not — its huge pool buys context
headroom, not streams. Each row's basis cites both numbers. Spark-dual is
unseeded (2-unit clustering is ConnectX/RDMA, network-bound for TP). A soak on
real hardware refines each and upgrades it to validated.
test-profiles-compat: hardware count 9 -> 10 (dgx-spark).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Previously a mixed-card rig no-op'd the concurrency injection ("no single
card-class row applies"). But it does have a right answer: vLLM already sizes
the KV pool to min(free blocks) across TP ranks (the cache is symmetric-
sharded), so the smallest card dictates the pool. Clamping the envelope lookup
to the smallest-VRAM card therefore MIRRORS the engine — it's the real ceiling,
not a conservative guess — and stays safe for TP=1 too (the ceiling fits
whichever single card vLLM lands on, all >= the smallest).
Behaviour:
5090 + 3090 -> smallest 3090 has no row -> compose default (unchanged)
5090 + H100 -> smallest 5090 is seeded -> inject its ceiling (was: no-op)
2x 5090 -> homogeneous -> unchanged
The common 5090+3090 case still lands on the compose default (now for the
principled reason: the 24 GB card caps the pool), so nothing regresses; the
gain is heterogeneous rigs whose SMALLEST card is itself a seeded >24 GB class.
Fixtures lock both the new inject-on-smallest and the still-no-op paths.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Populate envelopes.yml with the first memory-envelope rows, computed (not
guessed) from the kv-calc pool ceiling at model-max context:
vllm/dual @ rtx-5090 (2x32 GB): 2 -> 4 concurrent full-262K sessions
vllm/minimal @ rtx-5090 (1x32 GB): 1 -> 9 concurrent full-32K sessions
Both are the no-preemption ceiling: N *full-context* sequences whose KV
blocks fit the pool. Concurrency is a capacity question kv-calc answers
deterministically (arithmetic, arch-independent), and max_num_seqs is a cap
with graceful preemption (never OOM), so an over-optimistic ceiling costs an
occasional preempt, not a crash. kv-calc is sm_120-calibrated (disc #571
paulp83's real-5090 verify-stress PASS), so a 32 GB projection is trusted
arithmetic. A concurrency-soak upgrades a `computed` row to `validated`.
Guard: test-envelopes now accepts a `computed` (kv-calc basis) block as
provenance alongside `validated` (soak) — a computed row must name its basis
(invocation + PASS/cap boundary). Added a fixture proving injection is
provenance-agnostic (the launcher reads max_num_seqs; only the guard cares
about provenance).
Multi-GPU: the seam already fires on homogeneous >2-card rigs (collapses N
identical cards to one class; verified vllm/dual injects on a 4x5090 spec) and
kv-calc computes TP=4 pools; the experimental TP=4 slugs are documented as
computable-but-deferred rather than seeded. Heterogeneous rigs no-op by design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Found testing vibethinker-3b (a reasoning model): with a small max_tokens
it spends the budget mid-<think> and returns HTTP 200 with 256 tokens but
EMPTY content (never emits the final answer). The old content-based check
wrongly flagged those streams as silent-empty failures. For a KV-pool
stress test the stream DID run (generated tokens, held KV) -> ok =
completion_tokens > 0; a true silent-empty is HTTP 200 with ZERO tokens
(what soak-test means). After the fix vibethinker sustains N=8 clean on a
3090 (8x the compose default, 0 post-warm growth).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
The #246 Phase 2 first pass — spend a bigger card's KV pool on
concurrency, weights-invariant. Ships the machinery + measurement; the
32 GB envelope rows land from volunteer probe runs (baselines-style
plumbing-first).
Why concurrency-only (agreed 2026-07-05, design doc): most headline
configs are already at 262K model-max on 24 GB (GGUF single, dual vLLM),
so the extra 32 GB pool serves MORE STREAMS, not more context. The
context lever (bigger max_model_len) only helps single-card vLLM and is
Cliff-2b-capped anyway — deferred.
- Composes: `--max-num-seqs` parametrized to `${MAX_NUM_SEQS:-N}` on the
two pilot slugs (dual=2, minimal=1). Weights-invariant, default kept.
- envelopes.yml (empty): measured per-(slug, card-class) `max_num_seqs`,
born-from-measurement (a `validated` block is required — no guesses).
- launch_compat.py `_envelope_env`: injects MAX_NUM_SEQS from a row when
it EXCEEDS the compose default; 24 GB / no-row / heterogeneous /
user-env-set -> no-op. Both launchers whitelist MAX_NUM_SEQS.
- test-envelopes: schema + injection contract (inject / no-row /
heterogeneous / user-env-wins / no-gain).
- concurrency-probe.sh: the measurement — N concurrent long streams x R
rounds, separating expected pool-fill from a real leak (post-warm
growth). Live-validated N=2 on the dual-3090 (0 MB post-warm growth,
all rounds clean — the shipped default is sound).
Reuses the Phase 1 injection seam; independent of the KV verdict.
Full scripts gate 67/67.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Adds a subsection to DTYPE_MATRIX's two-axes section clarifying that the
FMHA-kernel arch gating is a vLLM-family (vLLM + SGLang) phenomenon.
The llama.cpp family (mainline / ik-llama / beellama) is always-dequant:
KV quant is storage-only on every arch (dequant inside the FA kernel, no
FP8/FP4 tensor cores even on Hopper), so no consumer-vs-datacenter split
— q4_0 KV behaves the same on a 3090/4090/5090/Spark, which is why our
single-card GGUF configs hit 262K anywhere. GGUF weight quant is
dequant-to-FP16 too, so the native-FP8/NVFP4-weights win is vLLM-only.
Division of labor: native low-precision COMPUTE wins are vLLM-only (and
mostly datacenter for KV); the CAPACITY win (KV compression for long
ctx) is delivered arch-agnostically by the GGUF family — the right tool
for a consumer card that wants big context.
Verified: llama.cpp #22411 / #24109, ik_llama #1142.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Consolidates the validated findings from the #246 A/B + #571 volunteer
data + vLLM/TRT-LLM source review. New DTYPE_MATRIX section "Having the
Tensor Cores ≠ using them" makes the two load-bearing distinctions
explicit:
- Axis 1 — weights vs KV are gated SEPARATELY. FP8/FP4 *weights* work
wherever the TCs exist (sm_89+ FP8, all Blackwell FP4). KV is
different because it rides the attention (FMHA) kernel.
- Axis 2 — KV compute vs KV storage. Native FP8/FP4 *attention* compute
needs FA3 (Hopper sm_90 only) or trtllm-gen FMHA (datacenter Blackwell
sm_100/103 only). On EVERY consumer card — Ada 4090 (sm_89), consumer
Blackwell 5090/PRO-6000 (sm_120), DGX Spark GB10 (sm_121) — FP8 KV is
storage-only and nvfp4 KV doesn't work.
Consequences documented (both empirically confirmed): e4m3 ≡ e5m2 in
speed on consumer cards (86.73 vs 86.66 on a 5090, disc #571 — a
precision choice, not a perf lever); nvfp4 KV crashes on consumer
Blackwell (#43562). DGX Spark = same sm_12x family as the 5090.
Corrected now-wrong claims: DTYPE_MATRIX "Ada/Blackwell get a real win
on FP8 KV" (false — Hopper/DC-only) + the per-arch Ada/Blackwell-consumer
rows; QUANTIZATION + HARDWARE "native FP8 compute on sm_89+" → storage-
only + precision framing. Compute-win KV on consumer cards = INT8-PTH.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm