58 KiB
58 KiB
Changelog
Auto-generated from commit subjects by git-cliff on tag
push. Click any commit SHA below to see the full message body (why / how /
validation data) — those live in git log, not here. Don't hand-edit; the file
is regenerated on every tag.
Versioning: SemVer in 0.x — treat any minor bump as potentially breaking
until 1.0. Past CalVer tags (v2026.05.09, v2026.05.10) are preserved for
history; SemVer takes over from v0.3.0 onward.
| CalVer tag | SemVer equivalent | Date |
|---|---|---|
v2026.05.09 |
(≈ v0.1.0) | 2026-05-09 — first tagged release |
v2026.05.10 |
(≈ v0.2.0) | 2026-05-10 — stack reorg + Gemma 4 INT8 PTH unblock |
v0.5.0 — 2026-05-12
✨ Features
- feat(qwen): ship froggeric chat-template fixes as default-on (84498d4)
- feat(vllm): add PR #35936 required-tool fallback overlay (28b16b5)
- feat(qwen-tq3): add CLUB3090_TQ_K1_SKIP_MTP layer-filter for PR #40914 K+1 dispatch (6b2a7d5)
🎯 New models + serving paths
- compose(tq3-mtp-genesis): pin to Genesis v7.72.2 known-good vLLM nightly (570fa71)
📊 Benchmarks + cross-rig data
- bench(matrix): @ygafarov first heterogeneous Ampere + Blackwell eGPU dual (1770931)
📝 Documentation
- docs(dtype-matrix): more polish — RDNA naming, FP8 maturity caveats, AMD detection (62b3b45)
- docs(dtype-matrix): polish nuances + add Intel and AMD vendor sections (3d4548c)
- docs(dtype-matrix): per-arch hardware accelerator matrix for compose optimization (9c6d3cf)
- docs(faq): add 'INT8 PTH doesn't scale at concurrency — is that a bug?' (df53287)
- docs(tq3-mtp): writeup + charts for the Genesis-backed TQ3+MTP path (c2b1c93)
- docs(qwen-tq3): close round-4 — #40914 not shippable, route to nomtp + Genesis (9fba037)
- docs(qwen-tq3): re-tombstone tq3-mtp.yml after round-3 MTP-skip validation (063d3e9)
🧹 Maintenance
- refactor(qwen): rename int8-tq3 → tq3-* family + add no-MTP + Genesis variants (6182922)
[Pin: git checkout v0.5.0] · Full diff
v0.4.0 — 2026-05-11
✨ Features
- feat(rebench-report): close 9 gaps — TL;DR + rig + timings + reproducer + delta + discuss variant (be7f9aa)
- feat(rebench): add REPORT.md synthesizer + container/boot/GPU captures (18355f4)
- feat(rebench): halve default soak to 10 sessions × 5 turns (~15-20 min) (3406894)
- feat(rebench): one-shot canonical 5-step bench orchestrator (94a2522)
🐛 Bug fixes
- fix(switch): GPU memory pre-flight + widen RUNNING_PATTERN (4866913)
- fix(rebench-report): parse aider upstream_per_exercise as dict (not list) (7c4b310)
📊 Benchmarks + cross-rig data
- bench(head-to-head): matched-config rebench + Qwen INT8 PTH KV compose (755e519)
📝 Documentation
- docs(gemma-4-31b): document TQ3 Ampere FA2 head_dim wall + vendor #40108 overlay (f8c7066)
- docs(benchmarks): Qwen 3.6 27B vs Gemma 4 31B head-to-head on dual 3090 (edda3b3)
🧹 Maintenance
- chore(composes): bump Qwen pins → 1acd67a7, drop obsolete patch_tolist_cudagraph (16a1374)
- chore(cliff): skip auto-regen bot commits in changelog parser (a258e49)
[Pin: git checkout v0.4.0] · Full diff
v0.3.3 — 2026-05-10
🧹 Maintenance
- chore(changelog): subject-only rendering (drop commit body verbosity) (eeb946b)
[Pin: git checkout v0.3.3] · Full diff
v0.3.2 — 2026-05-10
✨ Features
- feat(quality-test): auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for localhost URLs (83bf73d)
🧹 Maintenance
- chore: trigger v0.3.2 release workflow (GitHub deduped previous tag push) (255c743)
- chore(changelog): automate CHANGELOG + release notes from commits via cliff (Option A) (64b0474)
[Pin: git checkout v0.3.2] · Full diff
v0.3.1 — 2026-05-10
🐛 Bug fixes
- fix(soak-helper): capture
delta.reasoningalongsidedelta.reasoning_content(88eb67a)
📝 Documentation
- docs(changelog): v0.3.1 entry for soak-helper delta.reasoning capture (9db8b26)
[Pin: git checkout v0.3.1] · Full diff
v0.3.0 — 2026-05-10
✨ Features
- feat(power-cap-sweep): --include-commit flag stamps club-3090 git SHA in report header (closes #112) (7d91ac7)
- feat(qwen3.6-27b): thinking OFF by default across all 21 composes (29d17ed)
- feat(setup): interactive MODEL_DIR prompt for fresh TTY users (3909c2d)
🐛 Bug fixes
- fix(qwen3.6-27b): use --default-chat-template-kwargs (not --chat-template-kwargs) (534d29f)
- fix: 4 stale refs missed in 2026-05-10 reorg push (caught by RobH589 #116) (cf7f195)
📝 Documentation
- docs(benchmarks): aider-polyglot-30 — Qwen 27B 20/30 (66.7%) > Gemma 4 31B 17/30 (56.7%) (e08988e)
- docs(recipes): use $MODEL_DIR placeholder + sensible cross-rig default (cc3a717)
- docs: use $MODEL_DIR placeholder, not the dev rig's /mnt/models/huggingface/ (fbf3431)
🧹 Other
- release: SemVer adoption + v0.3.0 changelog entry (7080f1f)
[Pin: git checkout v0.3.0] · Full diff
v2026.05.10 — 2026-05-10
✨ Features
- feat(gemma-4-31b): INT8 PTH KV unblocks 262K + AWQ + DFlash compose family (403b16f)
🎯 New models + serving paths
- compose: parametrize VLLM_ENFORCE_EAGER, KV_CACHE_DTYPE, P40/P82/PN54 across all variants (#110) (#110 by @easel)
- composes: refresh Quality lines with --full sandboxed (8-pack) results (9dea0eb)
- composes: add --full Quality lines on Qwen3.6-27B + Gemma 4 31B duals (26ff0e5)
🐛 Bug fixes
- fix: BIND_HOST opt-in + localhost script fixes (#109) (#109 by @easel)
📝 Documentation
- docs: WSL2 budget formula + Cliff 3 (DeltaNet SSM-state non-cacheable) (6e12700)
🛠️ Scripts + tooling
- quality-test.sh: --sandboxed-only passthrough (7020d96)
- quality-test.sh: --help, --pack passthrough, align with benchlocal-cli v0.5 (1be02d2)
- ci: replace Release Drafter with git-cliff for commit-based release notes (7002e6b)
🧹 Other
- reorg: services/ consolidation + gpu-mode under git + ComfyUI + pin tracker + path updates (00366a5)
- encourage-soak: template dropdown + script ergonomics + report reminder + Notes convention (c298b60)
- BENCHMARKS: add @ygafarov Strix-Halo + oculink-eGPU x4-PCIe single-3090 row (#113) (a589058)
[Pin: git checkout v2026.05.10] · Full diff
v2026.05.09 — 2026-05-09
⚠️ Cliffs, gotchas, regressions
- Merge v7.69-cliff2-test: ship Cliff 2 closure recipes (Balanced MTP + Max-context) (15b84df)
- v7.69 + #35975 + Codex P103 gate fix — Cliff 2 closure recipes (f6613c8)
- docs + charts: v7.66 + Cliff 1 mech B closed across all 4 TQ3 composes (ae4846f)
- PN30 dst-shaped temp fix: close DS conv state regression class on long-text (9af1a52)
- PN25 v3: close Cliff 1 mech B (club-3090#16) on long-text via setup-time Genesis backport (a62ad78)
- walk back: Cliff 1 mech B reproduces on real IDE-agent prompts (club-3090#16) (b62b6b1)
- Ship verified Cliff 1 closure on long-text 205K + long-vision 192K (287de1c)
- Cliff 1 P104 + P101 anchor fix outcomes (built on cliff1-fa-clamp branch) (e6570a7)
- Cliff 1 dual-mechanism: P101+P103 cross-rig test reveals FFN buffer cliff (573a377)
- Cliff 1 root cause revised: FA2 softmax_lse sized by max_seqlen (2d6b69d)
✨ Features
- feat(preflight): compose-dependency + HF_TOKEN + KV-format checks (#37, #47, #219) (b6c8708)
- feat(tools): kv-calc.py — predict per-card VRAM budget for Qwen3.6-27B (#226) (4e89c6a)
- feat(bounded-thinking): Phase 3 grammar A/B complete; DeepSeek scratchpad is the new recommended grammar (b956c85)
- feat(report.sh): --stress + --soak flags, --full now the canonical "everything" pass (8a29b95)
- feat(qwen3.6-27b/vllm): add dual4 + dual4-dflash composes (TP=4, 4×3090, #44) (#44 by @Whamp)
- feat(grammar-eval): land harness for Holiday tagline grammar A/B (7be8ecc)
- feat(soak-test): continuous-mode v2 fixtures + reproduces Cliff 2 at 25K accumulated context (8d5bfd8)
- feat(scripts): add soak-test.sh — runtime VRAM accretion validation (closes gap from #41) (563a39e)
- feat: detect repo drift in preflight + add scripts/update.sh (43fe2a4)
- feat(launch/switch): register vllm/dual-nvlink as a known variant (75de7c9)
- feat(preflight): warn when Genesis tree out of sync with setup.sh's declared pin (d552ed9)
- feat(scripts/report.sh): capture per-GPU PCIe lane width + Gen + bus ID (535be29)
- feat(scripts/report.sh): capture container-internal Python/CUDA versions (e491e07)
- feat(scripts): add report.sh — paste-ready triage report (31982f0)
- push long-text/bounded-thinking back to 185K + 0.975; long-vision stays 140K + 0.95 (df91d64)
- feat(vllm): structured-CoT bounded-thinking compose (cross-rig port) (3d151b9)
- Push verified ceilings: long-text 218K, long-vision 198K (f3e5b52)
- Verify 256K single-prompt prefill on dual.yml (Sandermage cross-rig) (5270d94)
🎯 New models + serving paths
- composes: formalize Status enum + Caveats field (100% coverage) (e1137d6)
- composes: rename dual4 → multi4 to align topology prefix with MULTI_CARD.md framing (d33e6f8)
- composes: complete profile-schema header rollout (8 more composes) (fca643d)
- compose: extend VLLM_ENFORCE_EAGER hook to dual / dual-nvlink / dual4 + HARDWARE.md docs (5ec40c6)
- compose: VLLM_ENFORCE_EAGER env hook + WSL2 .env docs (#99 by @easel)
- Add dual-nvlink-dflash-noviz compose variant (NVLink + DFlash N=5, 200K ctx, no vision) (63ab224)
- Add docker-compose.dual-nvlink-dflash.yml (#92) (#92 by @danbedford)
- composes: PYTORCH_CUDA_ALLOC_CONF env-override knob + WSL2 boot-crash docs (#84) (#84 by @easel)
- Add Gemma 4 + DFlash compose (vLLM PR #41703 Codex-rebased overlay) (#81) (#81 by @noonghunna)
- composes: env-override knobs MAX_MODEL_LEN + GPU_MEMORY_UTILIZATION (#79) (#79 by @noonghunna)
- add Gemma 4 31B + Google MTP drafter (first Ampere data) (#68) (#68 by @noonghunna)
- Add dual NVLINK Docker Compose setup for Qwen3.6-27B (1350450)
- Add llama.cpp compose + perf chart + Q3_K_XL bench data (39692c9)
- Add long-vision + long-text composes (formalize R3' / R3''' bench rows) (b641719)
🐛 Bug fixes
- fix: verify-full.sh broken pipe + llama-cpp DISABLE_THINKING env hook (8f103f3)
- fix(preflight): catch missing llama.cpp GGUF before container boot (#63) (#63 by @noonghunna)
- fix: remove thinking prompt from Carnice chat template + JSON tool format (3729144)
- fix: missing pipe in DUAL_CARD table row (a28ba38)
- fix(soak): flag silent-empty turns (HTTP 200 + 0 tokens) as warnings (f32d8a6)
- fix: 3 issues from community feedback (2f8ed19)
- fix(soak-test, switch): calibration + boot-progress UX from first cross-rig runs (8e9cf70)
- fix(dual-nvlink): rename to avoid collision + vendored Marlin path (147f2e3)
- fix(default compose): swap P65 (cudagraph workaround) → P67 (proper Triton kernel fix) (620d918)
- fix(long-text-no-mtp): drop P65 + P85 — missed in
a26e30b(22e6549) - fix(composes): drop GENESIS_ENABLE_P65 + P85 — out of sync with v7.69 dispatcher v2 (a26e30b)
- fix(setup.sh): auto-clone vllm-src Marlin patched fork (was manual step) (2e934ad)
- fix(docs): replace dead luce-spec/llama-cpp-dflash links with Luce-Org/lucebox-hub (e9c658c)
- fix(scripts): register vllm/long-text-no-mtp in switch.sh + launch.sh (1f09a05)
- fix(dflash): close the docs+setup gap that hit @lolren on club-3090#18 (eb54cf4)
- fix(verify): drop tail buffer on Genesis check 2 anchor (refines
95b0905) (f2c1433) - fix(verify+docs): close two items from troymroberts cross-rig validation (#25) (95b0905)
- fix(launch): pass per-variant URL + CONTAINER to verify-full.sh (#20) (77ca576)
- fix(docs): bump curl smoke-test max_tokens 30 → 200 (#14) (2f8bade)
- fix(vllm): fail fast when Genesis patches volume is empty (#13) (0df8f74)
- fix: address open issues #1, #4, #7 (ebacba1)
📊 Benchmarks + cross-rig data
- results: re-bench dual.yml + dual-dflash + dual-dflash-noviz on v0.20 (0bdcb69)
- results: dual-turbo re-bench with corrected env vars (PN22 / PN26 naming fix) (077228e)
📝 Documentation
- docs: add Discord invite to README + FAQ + issue template (c18257f)
- docs: refresh 4090 cross-rig knee with @laurimyllari's richer 38-cap sweep (20ca297)
- AGENTS.md: codify why patches/cache stay engine-level (not under a topology) (9fbce96)
- AGENTS.md: capture compose naming + profile schema + experimental-compose conventions (62e636c)
- docs+composes: align Gemma 4 compose names to Qwen's -.yml convention (fe86b48)
- docs: surface Gemma 4 31B + add at-a-glance profile schemas to canonical composes (4d7356a)
- docs: add Community projects section pointing at VykosX/club-3090-server (cd48764)
- docs: laptop EC-managed power + TQ3 vs fp8 KV naming-trap; verify-stress: auto-bump curl timeout under VLLM_ENFORCE_EAGER (fe23eff)
- docs + compose: ship Phase 2 INT8 PTH validation results — 262K Gemma 4 unblocked (1e1886a)
- docs(hardware): add Qwen3.6-35B-A3B (MoE) 3090 power-cap charts + comparison (ec27d75)
- docs(img): reposition freq-cap chart annotations to clear right margin (2f7eb44)
- docs(hardware): add 5090 clock-lock chart + Blackwell freq-cap section (119a5fa)
- docs(hardware): regen 3090 power-cap charts with SM clock + plateau evidence (9f77be7)
- docs(hardware): reconcile 230W vs 290W vs 330W sweet-spot story (a7a1d59)
- docs(hardware): @apnar prefill-heavy 5090 sweep — proves per-workload power ceiling (d5ef8c8)
- docs(hardware): correct 3090 cooling class — air, not water (1f94478)
- docs(hardware): embed 3090 + Qwen3.6 + llama.cpp power-cap chart (42afdbb)
- docs(hardware): embed 4090 + Qwen3.6 + llama.cpp power-cap chart (e70258c)
- docs(hardware): embed 5090 + Gemma 4 power-cap efficiency chart (8b1d51a)
- docs(engines): more honest vLLM GGUF status (2d9aa14)
- docs(engines): fix 12 corrupted table separators from ik_llama.cpp column add (8f5b924)
- docs(engines): add ik_llama.cpp as 5th column to comparison matrix (0959206)
- docs(hardware): add 5090 + Gemma 4 + MTP cross-rig anchor rows (apnar disc #86) (bef5701)
- docs: add INFERENCE_ENGINES.md feature matrix (vLLM/llama.cpp/SGLang/ktransformers) (dfceccb)
- docs: codify canonical power-cap-sweep command for cross-rig anchors (886b619)
- docs: BENCHMARKS rows + CHANGELOG entry for danbedford NVLink+DFlash variants (b893d60)
- docs(benchmarks): @apnar 5090 Gemma 4 MTP + DFlash rows (disc #67) (98b0601)
- docs(benchmarks): three cross-rig rows from 2026-05-07 reports (76aacdc)
- docs(benchmarks): add @aaronlockhartdev patched-P2P driver row (#91, disc #70) (4eea837)
- docs(upstream): note we filed cross-rig validation on vLLM PR #40391 (e46f1e8)
- docs(gemma-4): int8_per_token_head on Ampere — Codex investigation verdict (1c2c156)
- docs: surface host-build contributor flow + power-cap-sweep in README + CONTRIBUTING (9aa6cb2)
- docs(benchmarks): add @lamentofhighborne 1× 3090 llama.cpp MTP row (#85) (68dbfaf)
- docs(hardware): add @apnar's 5090 power-cap anchor + compute-saturation note (60d4df6)
- docs(upstream): correct Gemma 4 per-token-head KV row — upstream PR exists (eb9f955)
- docs(gemma-4): document fp8 + int8 KV exploration on Ampere — both blocked (bb07eb5)
- docs(gemma-4): empirical ctx ceilings + PR #41745 merge status (1038e5f)
- docs(power): add cooling caveat — 388W stock requires liquid cooling (b15c5e1)
- docs(power): revise default cap 230W → 330W per @syangsao cross-rig data (2fe017f)
- docs(benchmarks): correct V100 row VRAM 14.6→15.6 GB/card per @efschu (d7bffec)
- docs(benchmarks): add @efschu 2× Tesla V100 16GB row (first sm_70 Volta data) (9212c60)
- docs(benchmarks): @danbedford 2× 3090 cross-rig matrix (6 benches, controlled PCIe vs NVLink) (6e57215)
- docs(benchmarks): add @laurimyllari 4090 single-card vllm/long-text row (461c4d4)
- docs(benchmarks): add @lolren 2× 3090 + Ryzen 5950X cross-rig rows (3 variants) (34a2348)
- docs(benchmarks): add @apriori dual-dflash row (EPYC 7302P + Arch + 2× 3090) (344e595)
- docs(upstream): track llama.cpp MTP PR #22673 + non-adoption rationale (#64) (#64 by @noonghunna)
- docs(contributing): clarify issues-vs-discussions routing (#61) (#61 by @noonghunna)
- docs: add Carnice BF16MTP to DUAL_CARD, vllm README, and CHANGELOG (fbd3531)
- docs(runtimes): tighten Proxmox section — native venv works (#49) (a51202c)
- docs(hardware): note SM86 structural ~70% TG drop at 131K (cross-rig) (eb5cd70)
- docs: capture environmental footnotes — WSL2 TDR + Proxmox uvloop (#49, #50) (224ca71)
- docs(cliffs): add rig-class caveat — "known good" is rig-specific (#49) (53d5c6b)
- docs(multi-card): topology-aware pair selection on awkward GPU counts (#49) (8e60539)
- docs(benchmarks): walk back PFlash "shippable" framing — TTFT + NIAH ≠ full validation (#230, #231) (ccac1ff)
- docs(benchmarks): PFlash long-context bench — 131K source ceiling on 1× 3090 (#230) (ebca0c8)
- docs(benchmarks): K8V4 result + P2P-CNS finding on lucebox-hub dual-GPU (#229) (e78eaa1)
- docs(benchmarks): add lucebox-hub DFlash dual-GPU bench — no-op on 24 GB cards (#229) (cb089e1)
- docs(benchmarks): add @JusefPol's 2× 3090 + NVLink dual-nvlink row (#29, #31) (017d0d2)
- docs(lucebox): record PRs #78 + #80 — dual-GPU PFlash + DFlash split shipped (May 2026) (dec0f22)
- docs(sglang): refresh per-engine + comparison pages — DFlash + MTP native upstream as of May 2026 (ecc2d74)
- docs(structured-cot): soften Phase 3 framing per Codex v2-prompt validation (011d4cc)
- docs(cliffs/hardware): ground Cliff 2 + TQ3 explanations in published literature (9b370f5)
- docs: cross-reference TQ3→fp8 KV swap from CLIFFS, DUAL_CARD, dual-turbo.yml + CHANGELOG record (#47) (129a4f4)
- docs(hardware): 20 GB Ampere TP=2 needs fp8_e5m2 KV, not TQ3 (#47) (124f08c)
- docs(benchmarks): add @snoby's 2× 4090 dual-dflash-noviz row (#46) (fc4c061)
- docs: align bug-report + FAQ + MULTI_CARD with report.sh --full / --soak (b859630)
- docs(benchmarks): add Rig column for cross-rig contributions (d8e7f73)
- docs: add BENCHMARKS.md + extend grammar harness for full-bench mode (9043678)
- docs+gates: PR template, soak-continuous gate, Phase 2 grammar A/B (85a6ea8)
- docs: UPSTREAM tracker + SINGLE_CARD polish — close the cliff-2b research thread (451b9f3)
- docs: surface Cliff 2b multi-turn envelope + WHY TP=2 / llama.cpp escape (04764c5)
- docs(UPSTREAM): sync 3 upstream changes + add next-week revisit queue (4327fd3)
- docs: surface scripts/update.sh + repo-drift detection (bca5a06)
- docs(vllm-marlin-pad/README): add sanity-check procedure before image-bump syncs (1bb85fa)
- docs: add MULTI_CARD.md for 3+ GPU users (derived, untested locally) (75a64a6)
- docs(FAQ): add WSL2 RAM-constraint failure mode to troubleshooting (3bf7da7)
- docs: surface triage ladder at issue-filing time + add at-a-glance table (f55b0a7)
- docs(FAQ): add 5-step triage ladder before symptom-matching (9560efd)
- docs: add PFlash integration feasibility memo (Codex audit, 2026-05-02) (90a83a3)
- docs: route bug + bench templates through scripts/report.sh (b9a1305)
- docs(dual-card): substrate refs from v7.65/v7.66 → v7.69 (95b2c3b)
- docs: full sync to v7.69 + Cliff 2 60K closure recipes (f8c9c36)
- docs(UPSTREAM): track Pflash (Luce-Org prefill accelerator) — flagged by @troymroberts (#25) (e0e1752)
- docs(CLIFFS): note v7.68 cross-rig test outcome — 3 regressions, master stays on v7.66 (ae1b92f)
- docs: Genesis #14/#15 fixes shipped on Sandermage dev (P38B/P15B/PN25 pending v7.65) (60d7b02)
- docs(upstream): refresh tracker for v0.20 blockers, P38/FA varlen filings, v7.64 closures (f633fdb)
- docs+composes: refresh long-text/long-vision/bounded-thinking headers + max_tokens guidance (cc4f083)
- docs + bounded-thinking: roll new context defaults across user-facing surfaces (d803278)
- docs(compose): document Cliff 1 mech B real-workload gap + escape hatches (#16) (6bff99a)
- charts: add tweet-asset variant (single-card vLLM only, 2 bars) (f754669)
- charts: combined width 18 + 2-line group labels + dual VRAM title says vLLM (24c8a62)
- charts: fix layout overlap with Luce DFlash 7th bar (1ce7dc4)
- docs+charts: add Luce DFlash bench + watch entry; cautions in single-card chart (cf71feb)
- docs: demote 48K/tools-text/minimal to fallback; lead with long-* + llama.cpp (48f93e5)
- docs: fix stale chart ref in HARDWARE.md + delete obsolete vram-budget.svg (cc02699)
- docs: catch remaining stale 192K/205K refs in long-text.yml header (f00f279)
- docs: final cleanup pass on stale 192K/205K refs (a1fc225)
- docs+scripts+charts: propagate new ceilings (long-vision 198K, long-text 218K) (427d2f8)
- docs: record verified ceilings and bisection in CLIFFS + CHANGELOG (26e5f65)
- docs: note revised Cliff 1 diagnosis posted on Sandermage issue #11 (8d8968b)
- docs: link PN12 PR #13 + record independent validation pass (5e38365)
- docs: revise Cliff 1 analysis (PN12 anchor drift was the real bug) (13d325b)
- Document Cliff 1 205K closure (9f6182e)
- changelog: link P101 PR #12 in 2026-04-30 entry (90a03ce)
- docs: link P101 PR #12 in UPSTREAM and CLIFFS (d0d79b1)
- CLIFFS.md: post-2026-04-30 architectural-wall conclusion (8580dc6)
- CLIFFS.md: refine clamp formula + implementation shape (ChatGPT review) (da6393b)
- Add docs/CLIFFS.md — comprehensive prefill-cliff synopsis (b0eed46)
- LLAMA_CPP.md: add structural explanation of why prefill cliffs don't fire (17aff4c)
- Add docs/UPSTREAM.md + AGENTS.md (consolidate upstream tracking) (53d811d)
- Add docs/COMPARISONS.md — self-host vs cloud and other local options (297a982)
- Add docs/FAQ.md — common questions answered for tweet click-throughs (1b9374b)
- Add docs/EXAMPLES.md — client snippets + IDE / Open WebUI connection (91b817f)
- README: lead with two-routes framing (matches launch tweet) (710def5)
🔧 Pin bumps + upstream
- bump Genesis pin 753344b → fc89395 (v7.66 dev tip) (7a7efbe)
- v0.20 migration + Genesis v7.65 dev tip + cold-start cache + env-var alignment (5aa97a2)
- Genesis v7.62.x + PN8 on FP8 paths (closes Cliff 1 on tools-text) (51a4001)
🛠️ Scripts + tooling
- ci: add Release Drafter for CalVer release notes (c49db50)
- power-cap-sweep: also sum delta.reasoning (third field-path) (1528b59)
- power-cap-sweep: sum delta.reasoning_content alongside delta.content (71e5954)
- power-cap-sweep: clamp prefill calibration to model context window (32f924c)
- power-cap-sweep: plateau auto-detection + multi-mode chain docs (fd11ae6)
- power-cap-sweep: add SM/mem clock + throttle% + pstate sampling (ab2796d)
- power-cap-sweep: time-bounded prefill-heavy + decode-concurrent (Codex round 2) (1ede998)
- power-cap-sweep: time-bounded streaming bench (Codex Option A redesign) (7877c04)
- power-cap-sweep: 4 cross-card portability fixes (652103f)
- power-cap-sweep: env-overridable bench shape for decode-single mode (c638c30)
- setup.sh: auto-create .env for WSL2 boot-crash workaround (#60) (4861ee7)
- power-cap-sweep: --concurrency-stretch N flag for probing headroom past plateau pick (3991ecc)
- power-cap-sweep: plateau-detection auto-calibration (saturate headroomy GPUs) (29e7de5)
- report.sh: engine-aware Active container probes (vllm + llamacpp) (6fa66d2)
- report.sh: capture recently-exited containers' boot logs (#60) (cd980f6)
- setup.sh: add gemma-4-31b model support (#89) (dd3bccc)
- power-cap-sweep: --concurrency auto for workload-calibrated sweeps (Codex) (f811457)
- power-cap-sweep: --bench-runs N for variance mitigation (Codex) (f99fad3)
- power-cap-sweep: document decode-concurrent n=1 variance caveat (18c74de)
- power-cap-sweep: load-mode flag + concurrent/prefill modes (Codex iteration) (f387622)
- verify-stress: engine-aware diagnostic hints (closes #87) (4f01abb)
- power-cap-sweep: make CONTAINER optional for host engine builds (#85, #87) (2bb3cf7)
- scripts(verify-full, soak-test): decouple from docker/vLLM assumptions (#85, #87) (a8606e3)
- power-cap-sweep: fix stale summary footer + add compute-saturation note (8c26c4b)
- power-cap-sweep: reduce per-cap bench to ~30s for faster sweeps (a413321)
- power-cap-sweep: 10W default increment + under-load median power sampling (6d70b72)
- power-cap-sweep: auto-derive cap range from card's min/max power limits (e5c7a34)
- Add scripts/power-cap-sweep.sh — automated cross-rig power-cap A/B (#83) (#83 by @noonghunna)
- scripts: auto-detect running container + port in verify / bench (closes #52 promise) (29718ca)
- verify-stress: add 3 probes to cover the bug shapes we missed (5e745c5)
- Add scripts/health.sh — operational health check for running server (e7780c5)
- Split verify-full.sh → verify-full.sh (fast functional) + verify-stress.sh (boundary) (5060e22)
🧹 Maintenance
- restructure: promote topology to a directory level (single/dual/multi4) (acd7ffb)
- Drop vllm-gemma4-mtp overlay tree (merged upstream as #41745, validated) (aa99173)
- chore(gitignore): allow results/lucebox-*/ — evidence for BENCHMARKS lucebox row (030f780)
- chore(tools): commit residency-instrument as research tool with framing README (#41, #217) (ed05d1c)
- chore(results): commit grammar bench evidence + gitignore investigation artifacts (#217) (d82e898)
- refactor: vendor vllm#40361 Marlin patched files in-repo (drops /opt/ai/vllm-src/ host dep) (d8b341f)
- chore: untrack docs/diagnostics/, gitignore the path (3f18053)
- Remove no-genesis-mtp.yml (research artifact, not user-facing) (f4a28b1)
- Remove fast-chat.yml; extend P68/P69 disable to default (37a4895)
- Restructure docs around hardware axis: SINGLE_CARD.md + DUAL_CARD.md (26ac811)
- Audit + reconcile dual-card compose headers, patches README, setup output (0f33561)
🧹 Other
- benchmarks: add JDWarner #107 TB3 dual-eGPU + mixed-arch row (fa9df49)
- Rename gemma-mtp-fp8.yml → gemma-mtp-int8.yml to match Ampere reality (160e8fc)
- Two regressions caught + reframe Phase 2 around INT8 PTH (Ampere reality) (119f296)
- gemma-mtp-fp8: vendor rebased PR #40391 + stacked tool-parser fixes (#42006 + #41991) (f93d312)
- gemma-mtp: drop PR #41745 overlay + bump to post-merge nightly (595be8f)
- llama.cpp: --reasoning-format none default (opencode unblock, #97) (af00ab7)
- Set dual-nvlink-dflash-noviz --max-model-len default to 188000 (89c6862)
- patches: qwen3coder tool-parser deferred-commit sidecar (#72) (2e00b6d)
- TQ3 composes: propagate PN34 to remaining 4 (follow-up to #82 audit) (ab69f65)
- vllm/default: also enable P98 (belt+suspenders with PN34, follow-up to #82) (2c7efe6)
- vllm/default: add GENESIS_ENABLE_PN34_WORKSPACE_LOCK_RELAX=1 (#82) (3167497)
- add dual-nvlink-turbo variant (rebased on v7.72.2 master, sibling-table edits dropped) (#65) (#65 by @noonghunna)
- release(v7.72.2-uplift): Genesis pin + vLLM pin + sidecar consolidation (#59) (#59 by @noonghunna)
- carnice-bf16mtp: restore original template + qwen3_xml parser (d57579c)
- carnice-bf16mtp: JSON tool format + empty think block, no reasoning parser (a350df7)
- carnice-bf16mtp: add HF model URL to header (5da50ec)
- carnice-bf16mtp: formal narrative + code bench results (7fef94f)
- carnice-bf16mtp: 2 streams at 262K confirmed + formal bench numbers (66d42c7)
- carnice-bf16mtp: 65K context was config choice, not VRAM ceiling — bumped to 262K (1cf0cb2)
- Carnice-V2-27B + BF16 MTP overlay — new compose variant (bc28542)
- extend PN25 v3 + PN30 dst-shaped temp fix to all 4 TQ3 composes (b875624)
- Genesis pin d89a089 → 753344b + cross-rig validation of Sander's PN30/PN31 (2b5ab4d)
- cliffs: v0.20 unblock recipe + 50K-stress-PASSES finding (9506561)
- cliffs: document P38 silently no-op'd on TurboQuant KV path (91355b8)
- long-text/long-vision/bounded-thinking: middle-ground recovery 130K → 175K / 120K → 140K (383b5cc)
- long-text/long-vision: enable P37 + back off context for activation headroom (1a931b4)
- genesis: bump pin v7.62 → v7.64 + add compile-safe FFN sidecar (#16) (53d0663)
- Add local FA max seqlen clamp sidecar (9f06a0f)
- Fix local PN12 activation pool anchor (41eabac)
- CLIFFS: document PN12-is-partial finding (full stack still hits wall) (537875a)
- Add genesis #11 row to UPSTREAM.md (bb406f9)
- Add Max ctx column to TL;DR + perf-summary tables on both pages (e94c2e7)
- DUAL_CARD: promote perf chart to top, parallel to SINGLE_CARD (19fb8e7)
- Disable P68/P69 on long-vision, long-text, dual-turbo too (f0cbcc6)
- Disable Genesis P68/P69 in shipped composes (silent-stop bugfix) (aab8ff4)
- Split charts per GPU-count page; chart sources land in tools/charts/ (3742244)
- Move performance chart into docs/img/ alongside vram-budget-dual (2e3ae0c)
- UX polish: pre-flight checks + cards-first wizard + PNG embeds (abc06c3)
- FAQ: add VS Code Copilot LLM Gateway entry (f275bf5)
- Add CONTRIBUTING.md — what kind of PRs land cleanly (0c261eb)
- CHANGELOG: capture post-launch polish day in cross + per-model logs (1cc6ee6)
- Add launch.sh wizard + switch.sh stateless variant switcher (4b77ed5)
- Add per-card VRAM allocation diagram + reference from model README (88523b3)
- Cite Kaitchup Qwen3.6-27B GGUF eval as quant-quality lens (b7ef91f)
- Pin Genesis to exact tested commit + add .env.example + issue templates (ec704e4)
- Dual-card re-bench on club-3090 substrate + fix dual-turbo mount path (c701474)
- dual-turbo: switch kv-cache-dtype k8v4 → 3bit_nc to align with test findings (3e1f5f6)
- Pin Genesis version + fix MODEL_DIR defaults + clean stale headers (7f00e52)
- Fix .gitignore + add the entire models/ tree (initial commit was incomplete) (2511a98)
- Initial commit — club-3090: model-agnostic LLM serving recipes for RTX 3090 (3fa3333)
[Pin: git checkout v2026.05.09]