The single-card Gemma-4 default was the catalog's only default with no baseline row (wave-2 PRIORITY gap). Full gate on Anbeeld's OFFICIAL v0.3.2-preview digest (the current beellama-local pin), rebench tag gemma-dflash-regate: - decode 44.91 narr / 79.76 code (n=5, CV 2.1%/5.6%), TTFT 115 ms - NIAH ladder clean to 117,513 tok (91% of 128K); ceiling margin 929 MB (< 1024 MB bar -> caveat stays, but up from ~177 MB pre-v0.3.2) - 8-pack current harness: 108/150 off / 113/150 on (+5 in-band wash, think-off default re-confirmed) - soak PASS (p50 59.8, 100.6% retention, 0 growth, 0/100 silent-empty) Row seeded via catalog-baseline.sh (first non-vLLM engine through the producer flow); BENCHMARKS row + compose header (Quality + margin numbers) updated to match. Tooling fallout the gate exposed: - rebench-full.sh hardcoded RUNS:-3/WARMUPS:-1, silently under-running the bench protocol on every orchestrated gate. The A1 "n=3 env leak" attribution was wrong -- it was this line (baselines comment corrected). Fixed to protocol 5/3; n=3 had flattered gemma-dflash code TPS by +8%. - catalog-baseline.sh counted prefill-probe run-lines into its bench-n gate (read n=8 for an n=5 log); now reads the summary headers. Full scripts gate 64/64. Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
294 lines
14 KiB
YAML
294 lines
14 KiB
YAML
# baselines.yml — the shipped catalog baseline ("the bar").
|
||
# ===========================================================================
|
||
# One row per registry slug: the ACCEPTED display projection of a validated
|
||
# gate run (full evidence stays in results/rebench/<source_tag>/ or the
|
||
# BENCHMARKS.md row it was reviewed from). PR-reviewed; written by
|
||
# scripts/catalog-baseline.sh (slice 2) or by hand at promotion time.
|
||
# Consumers read the registry-emit join, NEVER this file directly, and NEVER
|
||
# BENCHMARKS.md (which stays the public human/cross-rig ledger — publication,
|
||
# not machine source).
|
||
#
|
||
# Conventions:
|
||
# narr_tps / code_tps DECODE-class TPS means (bench.sh canonical prompts,
|
||
# warm, n>=3; the per-token rate, not wall).
|
||
# quality_8pk thinking-OFF arm "P/150"; _think_on = thinking-ON.
|
||
# 8-pack noise band is ±5–7 — separates classes, not
|
||
# fine rankings (badge semantics, design §2.1.4).
|
||
# ctx_validated highest ctx EXERCISED by the gate, with the recall
|
||
# verdict: {tokens: N, niah: "clean@XK" | "allocation-
|
||
# only"} — allocation alone is NOT validation (§2.1.2).
|
||
# engine_pin the pin the numbers were measured ON. If it differs
|
||
# from the slug's current engine pin, the row is
|
||
# PROVABLY STALE → c3 badges "re-bench owed" (§2.2).
|
||
# rig / power_cap_w hardware fingerprint class + per-card caps at bench.
|
||
# source_tag results/rebench/<tag>/ when a tag dir exists;
|
||
# omitted for rows reviewed from a BENCHMARKS.md entry
|
||
# (the row comment names the source).
|
||
#
|
||
# SEED WAVE 1 (2026-07-04): rows with airtight traceability only — a rebench
|
||
# tag on disk or an unambiguous decode-class BENCHMARKS.md row. Slugs still
|
||
# owed a row (evidence exists but needs review/archaeology) are listed at the
|
||
# bottom; DON'T guess numbers into them.
|
||
# SEED WAVE 2 (2026-07-04): the reviewed dispositions of the wave-1 gap list —
|
||
# 5 slugs seeded (each from a named BENCHMARKS.md row), the rest converted to
|
||
# explicit no-row gaps with their re-gate status (footer).
|
||
# ===========================================================================
|
||
schema_version: 1
|
||
baselines:
|
||
|
||
# ── Qwen3.6-27B ───────────────────────────────────────────────────────────
|
||
vllm/dual:
|
||
# BENCHMARKS.md 2026-06-30 row — the v0.22.0→v0.24.0 pin-bump validation
|
||
# gate (decode 70.7/93.5; MTP n=3 accept 3.51, KV pool 622K/2.37×,
|
||
# verify-full 8/8, verify-stress NIAH→240K, soak-continuous PASS).
|
||
# Measured ON the current vllm-stable pin → FRESH. Quality gate was
|
||
# `--quick` only (toolcall 11/15 · IF 15/15) — no 8-pack yet; add
|
||
# quality_8pk on the next full gate.
|
||
narr_tps: 70.7
|
||
code_tps: 93.5
|
||
ctx_validated: { tokens: 245760, niah: "clean@240K" }
|
||
date: 2026-06-30
|
||
engine_pin: "vllm/vllm-openai:v0.24.0"
|
||
rig: "2x3090-pcie"
|
||
power_cap_w: [370, 420]
|
||
submitted_by: "noonghunna"
|
||
|
||
vllm/qwen-27b-dual-fast:
|
||
# alias of vllm/dual (same compose_path dual/autoround-int4/fp8-mtp.yml,
|
||
# same port — the #340 tier naming). Row mirrors vllm/dual verbatim;
|
||
# update both together.
|
||
narr_tps: 70.7
|
||
code_tps: 93.5
|
||
ctx_validated: { tokens: 245760, niah: "clean@240K" }
|
||
date: 2026-06-30
|
||
engine_pin: "vllm/vllm-openai:v0.24.0"
|
||
rig: "2x3090-pcie"
|
||
power_cap_w: [370, 420]
|
||
submitted_by: "noonghunna"
|
||
|
||
llamacpp/default:
|
||
# BENCHMARKS.md 2026-05-23 row (decode, n=3, CV 0.6%/1.4%).
|
||
# Measured on the pre-2026-05-26 ROLLING tag (unknowable build) — the
|
||
# compose has pinned server-cuda-b9246 since → BORN-STALE, re-bench owed.
|
||
narr_tps: 50.27
|
||
code_tps: 58.92
|
||
date: 2026-05-23
|
||
engine_pin: "ghcr.io/ggml-org/llama.cpp:server-cuda"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
llamacpp/mtp:
|
||
# BENCHMARKS.md 2026-06-18 row (decode, n=5, CV <2%, thinking-off).
|
||
narr_tps: 47.9
|
||
code_tps: 55.3
|
||
date: 2026-06-18
|
||
engine_pin: "ghcr.io/ggml-org/llama.cpp:server-cuda-b9246"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
llamacpp/mtp-vision:
|
||
# BENCHMARKS.md 2026-05-20 row (decode, n=5, CV 1.6%/1.7%).
|
||
# Rolling-tag era (pre-b9246 pin) → BORN-STALE, re-bench owed.
|
||
narr_tps: 56.52
|
||
code_tps: 66.17
|
||
date: 2026-05-20
|
||
engine_pin: "ghcr.io/ggml-org/llama.cpp:server-cuda"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
ik-llama/iq4ks-mtp:
|
||
# BENCHMARKS.md 2026-05-23 row (decode 60.39/72.40, n=3; set+readback 370 W).
|
||
narr_tps: 60.39
|
||
code_tps: 72.4
|
||
date: 2026-05-23
|
||
# ik composes pin by digest; measurement-era pin presumed == current cu13
|
||
# digest (verify + exactify on the next gate via the induction tool).
|
||
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
ik-llama/iq4ks-two-stage:
|
||
# BENCHMARKS.md 2026-05-24 row (decode, n=3, CV 1.9%/5.3%).
|
||
narr_tps: 59.4
|
||
code_tps: 97.8
|
||
date: 2026-05-24
|
||
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
beellama/dflash:
|
||
# BENCHMARKS.md 2026-05-30 row (decode 50.4/101.3, n=5, CV 5.6%/7.4%).
|
||
# Measured on the pre-#296 noonghunna multiarch image → PROVABLY STALE vs
|
||
# the current v0.3.2-preview engine pin (never re-benched on it) — the
|
||
# staleness badge is CORRECT here, keep it until a re-bench.
|
||
narr_tps: 50.4
|
||
code_tps: 101.3
|
||
date: 2026-05-30
|
||
engine_pin: "ghcr.io/noonghunna/beellama-cpp:multiarch-b9459-07ac3ce"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
ik-llama/iq4ks-mtp-vision:
|
||
# BENCHMARKS.md 2026-05-25 vision re-tune row (PR #227/#437): decode
|
||
# CARRIED from ik-llama/iq4ks-mtp (≈60/72 — decode is ctx-independent;
|
||
# never re-benched with the mmproj loaded). verify-stress 8/8 @160K+1M-px
|
||
# was measured directly (recall to 147K, vision functional). Carried
|
||
# numbers qualify because the delta config doesn't touch the decode path;
|
||
# replace with measured values on the next gate.
|
||
narr_tps: 60.39
|
||
code_tps: 72.4
|
||
date: 2026-05-25
|
||
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
ik-llama/byteshape-iq4xs-mtp:
|
||
# BENCHMARKS.md row (decode 115.60/137.07, n=5) + 8-pack 110/150 off —
|
||
# community intake #293/#299 reproduced on our rig 2026-06-02.
|
||
narr_tps: 115.6
|
||
code_tps: 137.07
|
||
quality_8pk: "110/150"
|
||
date: 2026-06-02
|
||
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
# ── Qwen3.6-35B-A3B ──────────────────────────────────────────────────────
|
||
ik-llama/apex-fit-q8q5:
|
||
# BENCHMARKS.md 2026-05-28 row (decode 105.63/156.80, n=5, CV 3.0%/1.5%;
|
||
# @laurimyllari's --fit + asymmetric q8/q5 KV from #241). verify-full 8/8,
|
||
# verify-stress 8/8 incl. 180K NIAH, soak-continuous PASS. Quality in the
|
||
# row is deterministic-6 76/90 + sandbox packs — not an 8-pack /150, so no
|
||
# quality_8pk here (add on the next full gate).
|
||
narr_tps: 105.63
|
||
code_tps: 156.8
|
||
ctx_validated: { tokens: 184320, niah: "clean@180K" }
|
||
date: 2026-05-28
|
||
engine_pin: "ghcr.io/ikawrakow/ik-llama-cpp@sha256:5f914f1ccade922417af58c94bd1cbb558052c8852d86678ead3fe693eec0143"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
submitted_by: "noonghunna"
|
||
|
||
vllm/qwen-35b-a3b-dual:
|
||
# BENCHMARKS.md promotion row (#259, decode 182.3/182.3; NIAH-clean 240K).
|
||
# Measured on v0.22.0; the slug now pins v0.24.0 (fast/max-tier era) →
|
||
# PROVABLY STALE — the born-stale demo row for the guard (§2.2).
|
||
# quality_8pk_think_on from the fresher #480-era measurement (incumbent
|
||
# baseline used in the Agents-A1 comparison).
|
||
narr_tps: 182.3
|
||
code_tps: 182.3
|
||
quality_8pk_think_on: "110/150"
|
||
ctx_validated: { tokens: 245760, niah: "clean@240K" }
|
||
date: 2026-05-30
|
||
engine_pin: "vllm/vllm-openai:v0.22.0"
|
||
rig: "2x3090-pcie"
|
||
power_cap_w: [370, 420]
|
||
submitted_by: "noonghunna"
|
||
|
||
# ── Agents-A1 ────────────────────────────────────────────────────────────
|
||
vllm/agents-a1-dual:
|
||
# rebench tag agents-a1-fp8-dual (2026-07-03 gate): bench n=3 CV 0.1%
|
||
# (n=3 came from rebench-full.sh's old RUNS:-3 default — NOT an env leak
|
||
# as first attributed; fixed to protocol n=5 2026-07-04. n=5 was
|
||
# misstated in early publications; artifact-
|
||
# verified by catalog-baseline.sh 2026-07-04. think-on 110 = the #81
|
||
# rescored total, MATERIALIZED into the tag JSONs 2026-07-04 via
|
||
# `benchlocal-cli rescore --in-place` — artifact and publication agree;
|
||
# see docs/QUALITY_TEST.md "Rescoring saved results");
|
||
# verify-stress 8/8 staggered NIAH to 240,635 tok exact-recall, VRAM Δ0;
|
||
# soak PASS (99.8% retention). Quality = post benchlocal #79+#81 harness:
|
||
# OFF 105/150 · ON 110/150 (cli-40 ON 23/40 fresh-image).
|
||
# Cross-rig confirmed 2026-07-04 (#552 @sumo-dandan, decode −1%).
|
||
# Decode pair order corrected in wave-2 from the tag artifact
|
||
# (narrative 153.97 / code 153.78 — the wave-1 seed had them swapped).
|
||
narr_tps: 154.0
|
||
code_tps: 153.8
|
||
ttft_ms: 130
|
||
quality_8pk: "105/150"
|
||
quality_8pk_think_on: "110/150"
|
||
ctx_validated: { tokens: 240635, niah: "clean@240K" }
|
||
date: 2026-07-03
|
||
engine_pin: "vllm/vllm-openai:v0.24.0"
|
||
rig: "2x3090-pcie"
|
||
power_cap_w: [370, 420]
|
||
source_tag: "agents-a1-fp8-dual"
|
||
submitted_by: "noonghunna"
|
||
|
||
# ── Gemma-4-31B ──────────────────────────────────────────────────────────
|
||
vllm/gemma-31b-dual:
|
||
# rebench tag gemma-31b-dual-bf16 (v0.24.0 consolidation gate #538/#539):
|
||
# decode 59.07/59.06, TTFT ~72 ms, soak PASS p50 58.69.
|
||
# TODO-review: quality_8pk + ctx_validated (gate artifacts have the NIAH
|
||
# ladder — extract the verdict; 224K claim pending #40391 for 262K).
|
||
narr_tps: 59.07
|
||
code_tps: 59.06
|
||
ttft_ms: 72
|
||
date: 2026-07-02
|
||
engine_pin: "vllm/vllm-openai:v0.24.0"
|
||
rig: "2x3090-pcie"
|
||
power_cap_w: [370, 420]
|
||
source_tag: "gemma-31b-dual-bf16"
|
||
submitted_by: "noonghunna"
|
||
|
||
# ── Gemma-4-26B-A4B ──────────────────────────────────────────────────────
|
||
vllm/gemma-26ba4b-single:
|
||
# BENCHMARKS.md 2026-06-06 row (decode 169.2/219.9, CV 1.9%/2.0%; rebench
|
||
# tag gemma-26ba4b-int8r): single-3090 INT8-PTH KV via the vendored
|
||
# #40391 overlay on stock v0.22.0 (`vllm-gemma-stable`) + external MTP
|
||
# drafter n=4. verify-full ✓, NIAH-clean to 161K (91% of 176K), soak PASS
|
||
# 0-growth. gemma-stable still pins v0.22.0 → FRESH (the v0.24.0
|
||
# consolidation moved the gemma-TEXT slugs; this one kept the overlay pin).
|
||
narr_tps: 169.2
|
||
code_tps: 219.9
|
||
quality_8pk: "98/150"
|
||
quality_8pk_think_on: "109/150"
|
||
ctx_validated: { tokens: 164864, niah: "clean@161K" }
|
||
date: 2026-06-06
|
||
engine_pin: "vllm/vllm-openai:v0.22.0"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
source_tag: "gemma-26ba4b-int8r"
|
||
submitted_by: "noonghunna"
|
||
|
||
beellama/gemma-dflash:
|
||
# inducted by catalog-baseline.sh from rebench tag gemma-dflash-regate (2026-07-04);
|
||
# evidence: verify-full pass · bench n=5 · quality both arms · NIAH ladder
|
||
narr_tps: 44.91
|
||
code_tps: 79.76
|
||
ttft_ms: 115
|
||
prefill_tps: { 10k: 949, 90k: 603 }
|
||
quality_8pk: "108/150"
|
||
quality_8pk_think_on: "113/150"
|
||
ctx_validated: { tokens: 117513, niah: "clean@118K" }
|
||
date: 2026-07-04
|
||
engine_pin: "ghcr.io/anbeeld/beellama.cpp@sha256:858e7cfbfb0d5d3d605723a82ace1474698de368b81bc310f4df4dadc6c1b7d2"
|
||
rig: "1x3090-pcie"
|
||
power_cap_w: [370]
|
||
source_tag: "gemma-dflash-regate"
|
||
submitted_by: "noonghunna"
|
||
|
||
# ───────────────────────────────────────────────────────────────────────────
|
||
# KNOWN GAPS — wave-2 dispositions (2026-07-04): each slug below has NO row
|
||
# on purpose; the disposition says what unlocks one. Do NOT guess numbers in.
|
||
# vllm/minimal — only ~32/~33 approximations exist (never a protocol bench);
|
||
# needs a proper gate run. No row until then.
|
||
# vllm/gemma-12b-dual-bf16-mtp · vllm/gemma-12b-single-int8-mtp —
|
||
# conflicting-era rows (2026-06-04) with ambiguous decode-vs-wall class,
|
||
# AND both slugs re-pinned in the v0.24.0 gemma consolidation (#538/#539)
|
||
# → any seed would be born-stale AND class-uncertain. Re-gate owed.
|
||
# beellama/gemma-dflash — no direct BENCHMARKS row under the slug (numbers
|
||
# live in promotion-era notes). PRIORITY re-gate: this is the single-card
|
||
# gemma DEFAULT — the catalog's recommended path with no bar is the worst
|
||
# gap on this list.
|
||
# llamacpp/deckard40B-dual-mtp — 41.6 tok/s MTP exists only in registry
|
||
# notes; no BENCHMARKS row. Needs a gate run.
|
||
# ───────────────────────────────────────────────────────────────────────────
|