Commit Graph
5 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 1055c9162f fix(rebench-report): surface ceiling VRAM margin in verify-stress section (#184)
verify-stress.sh already hard-fails a ceiling free-VRAM margin below the
sustained-agent guard (default 1024 MB), but a margin that *passes* the guard
yet sits close to it can still OOM on a second run or server restart — the
marginal-but-passing case @laurimyllari hit on his 4090 at 160K (8/8 PASS, then
a second rebench OOMed). The report dropped the margin number entirely, so a
PASS read as comfortable headroom.

parse_verify_stress now extracts the deepest-rung free VRAM + guard from the
existing verify-stress output (both the thin-margin and ok lines), classifies
it (thin < guard / marginal < 2× guard / comfortable), and the verify-stress
report section prints e.g. "Ceiling VRAM margin: 1500 MB free (guard 1024 MB)
— ⚠ marginal: ... a second run or server restart may OOM. Lower CTX_SIZE ...".

rebench-report.py only — no change to verify-stress.sh's pass/fail logic.
Validated parse on comfortable / marginal / thin / no-margin logs.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-24 20:05:52 +00:00
noonghunna 2a148d702b feat(bench): surface prompt processing throughput
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m23s
Release / release (push) Failing after 45s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-15 09:03:10 +00:00
noonghunna be7f9aa143 feat(rebench-report): close 9 gaps — TL;DR + rig + timings + reproducer + delta + discuss variant
scripts/rebench-report.py:
  - Date extraction via regex (handles tag-with-or-without YYYY-MM-DD suffix),
    falls back to dir mtime.
  - Quality latency parser now reads nested latency.{p50,p95,mean} (was looking
    for non-existent flat fields).
  - New 'Paste-ready compose Quality: schema line' subsection grepped from
    quality-full.log.
  - Rig info section in Meta: hostname, GPU models, per-card power cap,
    parsed from rig.txt.
  - Auto-computed TL;DR section (3-6 bullets covering TPS, KV pool,
    verify-stress, quality, soak, aider).
  - Phase timings table from timings.json with grand total.
  - Reproducer-commands section with paste-ready bash.
  - --compare-to <tag-dir> flag for numerical deltas across runs; writes
    _internal.json sidecar so future runs can diff against this one.
  - REPORT-discuss.md trimmed variant for GitHub Discussion comments.
  - Aider per_language already fixed in prior commit; quant-string fix
    in discuss header (was rendering 'INT16' from --dtype float16 suffix).

scripts/rebench-full.sh:
  - Writes rig.txt (hostname + nvidia-smi -L + power cap) at preamble.
  - Writes timings.json incrementally as each phase completes, via
    record_timing helper called from run_step.
  - quality-full.log is the source for the Quality: one-liner so the
    rebench-report grep just works.

Validated against the synthesized Qwen INT8 PTH n=4 tag dir — all 9
sections render correctly.
2026-05-11 00:45:19 +00:00
noonghunna 7c4b310cca fix(rebench-report): parse aider upstream_per_exercise as dict (not list)
aider's benchmark.py emits per-exercise results as a dict keyed by
'<lang>/<exercise>' path, nested under verifier_trace.trace. Earlier
parser assumed a flat list of {language, passed} dicts and so produced
no per-language tally.

Also handle the dual nesting (verifier_trace.trace vs flat) and use the
path prefix as the language key. Fall back to dict-aware tally when
top-level pass_rate/passed_count/total_count are None.

Validated against today's Qwen INT8 PTH n=4 aider rerun (quality-2026-05-11
T00-12-30.json) — produces:
  Total: 19/30
  cpp: 3/5  go: 4/5  java: 4/5  javascript: 4/5  python: 3/5  rust: 1/5
2026-05-11 00:35:13 +00:00
noonghunna 18355f41f6 feat(rebench): add REPORT.md synthesizer + container/boot/GPU captures
scripts/rebench-report.py — parses raw artifacts (bench.log, verify-stress.log,
quality-full.json, soak.log, aider-polyglot.json, container-config.json,
vllm-boot.log) and renders a single REPORT.md at the top of the tag dir.

Sections: meta · config · performance · concurrency+VRAM · verify-stress
matrix · quality 8-pack table · failure examples · soak KPIs · aider per-language.

scripts/rebench-full.sh additions:
  - container-config.json (docker inspect snapshot)
  - vllm-boot.log (KV pool size + max concurrency + MTP detect, trimmed)
  - gpu-state-{start,end}.log (nvidia-smi snapshots)
  - final REPORT.md synthesis phase

Standalone re-render: python3 scripts/rebench-report.py results/rebench/<tag>/

Dry-run validated against today's running Qwen INT8 PTH compose — meta/config/
concurrency sections render cleanly with patches-mounted manifest + Genesis
detection working.
2026-05-11 00:28:30 +00:00