Files
club-3090/scripts/lib/profiles/baselines.yml
T
noonghunnaandClaude Fable 5 5ea64bc539 Re-gate beellama/gemma-dflash on the official pin; fix rebench-full n=3
The single-card Gemma-4 default was the catalog's only default with no
baseline row (wave-2 PRIORITY gap). Full gate on Anbeeld's OFFICIAL
v0.3.2-preview digest (the current beellama-local pin), rebench tag
gemma-dflash-regate:

- decode 44.91 narr / 79.76 code (n=5, CV 2.1%/5.6%), TTFT 115 ms
- NIAH ladder clean to 117,513 tok (91% of 128K); ceiling margin 929 MB
  (< 1024 MB bar -> caveat stays, but up from ~177 MB pre-v0.3.2)
- 8-pack current harness: 108/150 off / 113/150 on (+5 in-band wash,
  think-off default re-confirmed)
- soak PASS (p50 59.8, 100.6% retention, 0 growth, 0/100 silent-empty)

Row seeded via catalog-baseline.sh (first non-vLLM engine through the
producer flow); BENCHMARKS row + compose header (Quality + margin
numbers) updated to match.

Tooling fallout the gate exposed:
- rebench-full.sh hardcoded RUNS:-3/WARMUPS:-1, silently under-running
  the bench protocol on every orchestrated gate. The A1 "n=3 env leak"
  attribution was wrong -- it was this line (baselines comment
  corrected). Fixed to protocol 5/3; n=3 had flattered gemma-dflash
  code TPS by +8%.
- catalog-baseline.sh counted prefill-probe run-lines into its bench-n
  gate (read n=8 for an n=5 log); now reads the summary headers.

Full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 19:28:03 +00:00

294 lines
14 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# baselines.yml — the shipped catalog baseline ("the bar").
# ===========================================================================
# One row per registry slug: the ACCEPTED display projection of a validated
# gate run (full evidence stays in results/rebench/<source_tag>/ or the
# BENCHMARKS.md row it was reviewed from). PR-reviewed; written by
# scripts/catalog-baseline.sh (slice 2) or by hand at promotion time.
# Consumers read the registry-emit join, NEVER this file directly, and NEVER
# BENCHMARKS.md (which stays the public human/cross-rig ledger — publication,
# not machine source).
#
# Conventions:
# narr_tps / code_tps DECODE-class TPS means (bench.sh canonical prompts,
# warm, n>=3; the per-token rate, not wall).
# quality_8pk thinking-OFF arm "P/150"; _think_on = thinking-ON.
# 8-pack noise band is ±5–7 — separates classes, not
# fine rankings (badge semantics, design §2.1.4).
# ctx_validated highest ctx EXERCISED by the gate, with the recall
# verdict: {tokens: N, niah: "clean@XK" | "allocation-
# only"} — allocation alone is NOT validation (§2.1.2).
# engine_pin the pin the numbers were measured ON. If it differs
# from the slug's current engine pin, the row is
# PROVABLY STALE → c3 badges "re-bench owed" (§2.2).
# rig / power_cap_w hardware fingerprint class + per-card caps at bench.
# source_tag results/rebench/<tag>/ when a tag dir exists;
# omitted for rows reviewed from a BENCHMARKS.md entry
# (the row comment names the source).
#
# SEED WAVE 1 (2026-07-04): rows with airtight traceability only — a rebench
# tag on disk or an unambiguous decode-class BENCHMARKS.md row. Slugs still
# owed a row (evidence exists but needs review/archaeology) are listed at the
# bottom; DON'T guess numbers into them.
# SEED WAVE 2 (2026-07-04): the reviewed dispositions of the wave-1 gap list —
# 5 slugs seeded (each from a named BENCHMARKS.md row), the rest converted to
# explicit no-row gaps with their re-gate status (footer).
# ===========================================================================
schema_version: 1
baselines:
# ── Qwen3.6-27B ───────────────────────────────────────────────────────────
vllm/dual:
# BENCHMARKS.md 2026-06-30 row — the v0.22.0→v0.24.0 pin-bump validation
# gate (decode 70.7/93.5; MTP n=3 accept 3.51, KV pool 622K/2.37×,
# verify-full 8/8, verify-stress NIAH→240K, soak-continuous PASS).
# Measured ON the current vllm-stable pin → FRESH. Quality gate was
# `--quick` only (toolcall 11/15 · IF 15/15) — no 8-pack yet; add
# quality_8pk on the next full gate.
narr_tps: 70.7
code_tps: 93.5
ctx_validated: { tokens: 245760, niah: "clean@240K" }
date: 2026-06-30
engine_pin: "vllm/vllm-openai:v0.24.0"
rig: "2x3090-pcie"
power_cap_w: [370, 420]
submitted_by: "noonghunna"
vllm/qwen-27b-dual-fast:
# alias of vllm/dual (same compose_path dual/autoround-int4/fp8-mtp.yml,
# same port — the #340 tier naming). Row mirrors vllm/dual verbatim;
# update both together.
narr_tps: 70.7
code_tps: 93.5
ctx_validated: { tokens: 245760, niah: "clean@240K" }
date: 2026-06-30
engine_pin: "vllm/vllm-openai:v0.24.0"
rig: "2x3090-pcie"
power_cap_w: [370, 420]
submitted_by: "noonghunna"
llamacpp/default:
# BENCHMARKS.md 2026-05-23 row (decode, n=3, CV 0.6%/1.4%).
# Measured on the pre-2026-05-26 ROLLING tag (unknowable build) — the
# compose has pinned server-cuda-b9246 since → BORN-STALE, re-bench owed.
narr_tps: 50.27
code_tps: 58.92
date: 2026-05-23
engine_pin: "ghcr.io/ggml-org/llama.cpp:server-cuda"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
llamacpp/mtp:
# BENCHMARKS.md 2026-06-18 row (decode, n=5, CV <2%, thinking-off).
narr_tps: 47.9
code_tps: 55.3
date: 2026-06-18
engine_pin: "ghcr.io/ggml-org/llama.cpp:server-cuda-b9246"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
llamacpp/mtp-vision:
# BENCHMARKS.md 2026-05-20 row (decode, n=5, CV 1.6%/1.7%).
# Rolling-tag era (pre-b9246 pin) → BORN-STALE, re-bench owed.
narr_tps: 56.52
code_tps: 66.17
date: 2026-05-20
engine_pin: "ghcr.io/ggml-org/llama.cpp:server-cuda"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
ik-llama/iq4ks-mtp:
# BENCHMARKS.md 2026-05-23 row (decode 60.39/72.40, n=3; set+readback 370 W).
narr_tps: 60.39
code_tps: 72.4
date: 2026-05-23
# ik composes pin by digest; measurement-era pin presumed == current cu13
# digest (verify + exactify on the next gate via the induction tool).
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
ik-llama/iq4ks-two-stage:
# BENCHMARKS.md 2026-05-24 row (decode, n=3, CV 1.9%/5.3%).
narr_tps: 59.4
code_tps: 97.8
date: 2026-05-24
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
beellama/dflash:
# BENCHMARKS.md 2026-05-30 row (decode 50.4/101.3, n=5, CV 5.6%/7.4%).
# Measured on the pre-#296 noonghunna multiarch image → PROVABLY STALE vs
# the current v0.3.2-preview engine pin (never re-benched on it) — the
# staleness badge is CORRECT here, keep it until a re-bench.
narr_tps: 50.4
code_tps: 101.3
date: 2026-05-30
engine_pin: "ghcr.io/noonghunna/beellama-cpp:multiarch-b9459-07ac3ce"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
ik-llama/iq4ks-mtp-vision:
# BENCHMARKS.md 2026-05-25 vision re-tune row (PR #227/#437): decode
# CARRIED from ik-llama/iq4ks-mtp (≈60/72 — decode is ctx-independent;
# never re-benched with the mmproj loaded). verify-stress 8/8 @160K+1M-px
# was measured directly (recall to 147K, vision functional). Carried
# numbers qualify because the delta config doesn't touch the decode path;
# replace with measured values on the next gate.
narr_tps: 60.39
code_tps: 72.4
date: 2026-05-25
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
ik-llama/byteshape-iq4xs-mtp:
# BENCHMARKS.md row (decode 115.60/137.07, n=5) + 8-pack 110/150 off —
# community intake #293/#299 reproduced on our rig 2026-06-02.
narr_tps: 115.6
code_tps: 137.07
quality_8pk: "110/150"
date: 2026-06-02
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
# ── Qwen3.6-35B-A3B ──────────────────────────────────────────────────────
ik-llama/apex-fit-q8q5:
# BENCHMARKS.md 2026-05-28 row (decode 105.63/156.80, n=5, CV 3.0%/1.5%;
# @laurimyllari's --fit + asymmetric q8/q5 KV from #241). verify-full 8/8,
# verify-stress 8/8 incl. 180K NIAH, soak-continuous PASS. Quality in the
# row is deterministic-6 76/90 + sandbox packs — not an 8-pack /150, so no
# quality_8pk here (add on the next full gate).
narr_tps: 105.63
code_tps: 156.8
ctx_validated: { tokens: 184320, niah: "clean@180K" }
date: 2026-05-28
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
rig: "1x3090-pcie"
power_cap_w: [370]
submitted_by: "noonghunna"
vllm/qwen-35b-a3b-dual:
# BENCHMARKS.md promotion row (#259, decode 182.3/182.3; NIAH-clean 240K).
# Measured on v0.22.0; the slug now pins v0.24.0 (fast/max-tier era) →
# PROVABLY STALE — the born-stale demo row for the guard (§2.2).
# quality_8pk_think_on from the fresher #480-era measurement (incumbent
# baseline used in the Agents-A1 comparison).
narr_tps: 182.3
code_tps: 182.3
quality_8pk_think_on: "110/150"
ctx_validated: { tokens: 245760, niah: "clean@240K" }
date: 2026-05-30
engine_pin: "vllm/vllm-openai:v0.22.0"
rig: "2x3090-pcie"
power_cap_w: [370, 420]
submitted_by: "noonghunna"
# ── Agents-A1 ────────────────────────────────────────────────────────────
vllm/agents-a1-dual:
# rebench tag agents-a1-fp8-dual (2026-07-03 gate): bench n=3 CV 0.1%
# (n=3 came from rebench-full.sh's old RUNS:-3 default — NOT an env leak
# as first attributed; fixed to protocol n=5 2026-07-04. n=5 was
# misstated in early publications; artifact-
# verified by catalog-baseline.sh 2026-07-04. think-on 110 = the #81
# rescored total, MATERIALIZED into the tag JSONs 2026-07-04 via
# `benchlocal-cli rescore --in-place` — artifact and publication agree;
# see docs/QUALITY_TEST.md "Rescoring saved results");
# verify-stress 8/8 staggered NIAH to 240,635 tok exact-recall, VRAM Δ0;
# soak PASS (99.8% retention). Quality = post benchlocal #79+#81 harness:
# OFF 105/150 · ON 110/150 (cli-40 ON 23/40 fresh-image).
# Cross-rig confirmed 2026-07-04 (#552 @sumo-dandan, decode −1%).
# Decode pair order corrected in wave-2 from the tag artifact
# (narrative 153.97 / code 153.78 — the wave-1 seed had them swapped).
narr_tps: 154.0
code_tps: 153.8
ttft_ms: 130
quality_8pk: "105/150"
quality_8pk_think_on: "110/150"
ctx_validated: { tokens: 240635, niah: "clean@240K" }
date: 2026-07-03
engine_pin: "vllm/vllm-openai:v0.24.0"
rig: "2x3090-pcie"
power_cap_w: [370, 420]
source_tag: "agents-a1-fp8-dual"
submitted_by: "noonghunna"
# ── Gemma-4-31B ──────────────────────────────────────────────────────────
vllm/gemma-31b-dual:
# rebench tag gemma-31b-dual-bf16 (v0.24.0 consolidation gate #538/#539):
# decode 59.07/59.06, TTFT ~72 ms, soak PASS p50 58.69.
# TODO-review: quality_8pk + ctx_validated (gate artifacts have the NIAH
# ladder — extract the verdict; 224K claim pending #40391 for 262K).
narr_tps: 59.07
code_tps: 59.06
ttft_ms: 72
date: 2026-07-02
engine_pin: "vllm/vllm-openai:v0.24.0"
rig: "2x3090-pcie"
power_cap_w: [370, 420]
source_tag: "gemma-31b-dual-bf16"
submitted_by: "noonghunna"
# ── Gemma-4-26B-A4B ──────────────────────────────────────────────────────
vllm/gemma-26ba4b-single:
# BENCHMARKS.md 2026-06-06 row (decode 169.2/219.9, CV 1.9%/2.0%; rebench
# tag gemma-26ba4b-int8r): single-3090 INT8-PTH KV via the vendored
# #40391 overlay on stock v0.22.0 (`vllm-gemma-stable`) + external MTP
# drafter n=4. verify-full ✓, NIAH-clean to 161K (91% of 176K), soak PASS
# 0-growth. gemma-stable still pins v0.22.0 → FRESH (the v0.24.0
# consolidation moved the gemma-TEXT slugs; this one kept the overlay pin).
narr_tps: 169.2
code_tps: 219.9
quality_8pk: "98/150"
quality_8pk_think_on: "109/150"
ctx_validated: { tokens: 164864, niah: "clean@161K" }
date: 2026-06-06
engine_pin: "vllm/vllm-openai:v0.22.0"
rig: "1x3090-pcie"
power_cap_w: [370]
source_tag: "gemma-26ba4b-int8r"
submitted_by: "noonghunna"
beellama/gemma-dflash:
# inducted by catalog-baseline.sh from rebench tag gemma-dflash-regate (2026-07-04);
# evidence: verify-full pass · bench n=5 · quality both arms · NIAH ladder
narr_tps: 44.91
code_tps: 79.76
ttft_ms: 115
prefill_tps: { 10k: 949, 90k: 603 }
quality_8pk: "108/150"
quality_8pk_think_on: "113/150"
ctx_validated: { tokens: 117513, niah: "clean@118K" }
date: 2026-07-04
engine_pin: "ghcr.io/anbeeld/beellama.cpp@sha256:858e7cfbfb0d5d3d605723a82ace1474698de368b81bc310f4df4dadc6c1b7d2"
rig: "1x3090-pcie"
power_cap_w: [370]
source_tag: "gemma-dflash-regate"
submitted_by: "noonghunna"
# ───────────────────────────────────────────────────────────────────────────
# KNOWN GAPS — wave-2 dispositions (2026-07-04): each slug below has NO row
# on purpose; the disposition says what unlocks one. Do NOT guess numbers in.
# vllm/minimal — only ~32/~33 approximations exist (never a protocol bench);
# needs a proper gate run. No row until then.
# vllm/gemma-12b-dual-bf16-mtp · vllm/gemma-12b-single-int8-mtp —
# conflicting-era rows (2026-06-04) with ambiguous decode-vs-wall class,
# AND both slugs re-pinned in the v0.24.0 gemma consolidation (#538/#539)
# → any seed would be born-stale AND class-uncertain. Re-gate owed.
# beellama/gemma-dflash — no direct BENCHMARKS row under the slug (numbers
# live in promotion-era notes). PRIORITY re-gate: this is the single-card
# gemma DEFAULT — the catalog's recommended path with no bar is the worst
# gap on this list.
# llamacpp/deckard40B-dual-mtp — 41.6 tok/s MTP exists only in registry
# notes; no BENCHMARKS row. Needs a gate run.
# ───────────────────────────────────────────────────────────────────────────