Docs: syangsao water-cooled byteshape cross-rig row (#445)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
This commit is contained in:
noonghunna
2026-07-04 23:06:33 +00:00
parent 055986a5d0
commit afe56e35a5

View File

@@ -374,6 +374,7 @@ First-pass numbers for the two MoE models onboarded in v0.7.3. **Preview** track
| **`qwen3.6-35b-a3b/ik-llama/single/mudler-apex-compact/mtp.yml` (TP=1, q4_0 KV, MTP n=5, --fit)** | @noonghunna (1× 3090, 370 W cap) | q4_0 / q4_0 | 200K | **96.47 / 144.01** (decode 98.54 / 150.98, n=5, CV 5.4% / 4.5%) | ~3000 @10K ~1283 @180K | (ik-llama MTP active, AL not in bench log) | (not reported) | 20.46 GB | 2026-05-28 | **First-ever real run of the existing EVAL-only compose** (APEX weights weren't on disk before today). Pre-`fit-mtp.yml` baseline, kept as the q4-KV / max-headroom variant. `verify-full` 8/8, `verify-stress` 8/8 incl. 180K NIAH at 91% of n_ctx. **Confirms ik built-in MTP on the 35B-A3B MoE does NOT pay the vLLM `51%` penalty** (compare row above) MoE×MTP is a vLLM-scheduler-specific problem, not architectural. TTFT 205 ms. Compose: `models/qwen3.6-35b-a3b/ik-llama/compose/single/mudler-apex-compact/mtp.yml`. |
| **`qwen3.6-35b-a3b/ik-llama/single/mudler-apex-compact/fit-mtp.yml` (TP=1, q8/q5 KV, MTP n=4)** ⭐⭐ | @noonghunna (1× 3090, 370 W cap) | **q8_0 (K) / q5_0 (V) — asymmetric K-high/V-low** | **196K** | **103.25 / 149.12** (decode 105.63 / 156.80, n=5, CV 3.0% / 1.5%) | ~3000 @10K ~1283 @180K | (ik built-in MTP active, AL not in bench log) | (not reported) | **21.09 GB** | 2026-05-28 | **Captures @laurimyllari's `--fit` + asymmetric KV config from [discussion #241](https://github.com/noonghunna/club-3090/discussions/241).** A/B vs `mtp.yml` row above (same engine, same weights, same rig + 370 W): **+7% narr / +4% code wall TPS** at **tighter CV** (3.0/1.5% vs 5.4/4.5%) for +0.6 GB VRAM the Anbeeld K-high/V-low pattern materialised. `verify-full` 8/8, `verify-stress` **8/8 incl. 180K NIAH** (91% of 196K, 26.4 GB margin, 1283 tok/s prefill at ceiling). Adds `--no-mmap` + `--cache-ram 4096` per his fit-mtp.yml. **Deterministic quality 76/90 = 84%** on PR #38-fixed verifiers (toolcall 14/15, instructfollow **15/15**, structoutput 14/15, dataextract 11/15, reasonmath 12/15, bugfind 10/15). **Sandbox-pack quality measured post-fix** (benchlocal-cli [#42](https://github.com/noonghunna/benchlocal-cli/pull/42) hermes thinking-on deterministic sampler + [#43](https://github.com/noonghunna/benchlocal-cli/pull/43) per-pack timeout defaults + [#44](https://github.com/noonghunna/benchlocal-cli/pull/44) aider git-checkout guard + club-3090 [#245](https://github.com/noonghunna/club-3090/pull/245) wrapper-omit): **hermesagent-20 11/20 (55%), cli-40 12/40 (30%), aider-polyglot-30 12/30 (40%)**. cli-40 retains 7 "timeout" failures but those are sandbox-internal agent-give-up (max latency 19 s, well under the 300 s budget) distinct from wall-clock, tracked in a follow-up benchlocal-cli issue. `soak-continuous` PASS (0 errors, 0/25 silent_empty, 0 VRAM growth, 100% retention, p50 decode 223 TPS). vs @laurimyllari on 4090 (205/254 TPS, 105-107/150 8-pack pre-PR #38): 3090 trails on TPS as expected (newer silicon + power), deterministic quality is on par. MoE × MTP no-penalty confirmed in ik-llama (cf. `preview-mtp.yml`). Compose: `models/qwen3.6-35b-a3b/ik-llama/compose/single/mudler-apex-compact/fit-mtp.yml`. Registry tag: `ik-llama/apex-fit-q8q5`. |
| **`qwen3.6-35b-a3b/ik-llama/single/byteshape-iq4xs/mtp.yml` (TP=1, q4_0 KV, MTP n=2, --fit)** | @noonghunna (1× 3090, 370 W) | q4_0 / q4_0 | **262K** (--fit auto) | **112.90 / 128.61** (decode 115.60 / 137.07, n=5, CV 1.9% / 1.8%) | ~1859 @95K ~1050 @240K | (ik built-in MTP n=2 active) | **110/150 (73%, off)** | **22.5 GB** | 2026-06-02 | **Community intake from PR #293 (@Rhonstin)** byteshape IQ4_XS 4.19bpw MoE GGUF (embedded MTP head), single-card 35B-A3B. First-party validated on 1× 3090: `verify-full` all-pass, `verify-stress` **8/8 incl. NIAH→240K** (91% of 262K, no Cliff; KV self 1584 MiB @ 262K MoE KV is cheap), `soak-continuous` **PASS** (0 err, 0 VRAM growth, 0/25 silent-empty, 100% retention, p50 223 TPS). **8-pack 110/150 (think-off)**: toolcall 15/15 · instructfollow 14/15 · structoutput 14/15 · dataextract 13/15 · reasonmath 13/15 · bugfind 13/15 · hermesagent-20 11/20 · cli-40 17/40 reproduces @Rhonstin's 111/150 within noise. q4/q4 KV buys **full 262K** (vs the apex sibling's q8/q5 @196K). Intake fixes vs #293: image cu13-server, port 8058. Compose: `models/qwen3.6-35b-a3b/ik-llama/compose/single/byteshape-iq4xs/mtp.yml`. Registry tag: `ik-llama/byteshape-iq4xs-mtp`. |
| `ik-llama/byteshape-iq4xs-mtp` (1× 3090 **water-cooled**) | @syangsao (1× 3090 water-cooled, **330 W cap** (default 390), driver 595.71.05, persistence on; [#445](https://github.com/noonghunna/club-3090/issues/445)) | q4_0 / q4_0 | **262K** | **133.05 / 154.81** (decode **135.51 / 160.90**, CV 1.4% / 1.1%, TTFT 135 ms) | 1941 @94K ~1130 @214K (NIAH-rung diagnostic) | ~22.5 GB | 2026-06-19 | **First water-cooled datapoint on the config — +17% decode over the air-cooled reference row above (115.6/137.1 @ 370 W) at 40 W LESS.** Sustained-clock effect: decode on this config is thermal/clock-bound enough that water cooling beats a higher power cap. verify-stress ladder **6/6 clean, fillable to 240,635 tok** (91% of 262K). Useful anchor for the power-sweep guidance in [docs/HARDWARE.md](docs/HARDWARE.md#power). |
| `Qwen3.6-35B-A3B-Cerebellum-v3` (1× 3090, mainline llama.cpp) **author data point, NOT a catalog compose** | @deucebucket (1× 3090, Bazzite bare metal, host llama.cpp b9603, 420 W limit / ~300316 W draw, desktop compositor on same GPU) | q8_0 / q8_0 | 131K (196K boots @ ~15.9 GB, fill-probed ~119K NIAH ladder not run) | **147.9 / 146.2** (decode 150.5 / 149.9; `report.sh --full` re-run 141.8 / 141.2) | | n/a (no drafter) | n/a | **~15.1 GB** (15,115 MiB @ 131K) | 2026-06-12 | **Author-reported; not club-validated or cataloged** declined as a catalog compose ([PR #393](https://github.com/noonghunna/club-3090/pull/393)): headline value is the lean ~15 GB footprint (fits a **16 GB card**, where the ik-llama siblings at 2122.5 GB don't), but we have no 16 GB card to validate/support that tier, and on 24 GB it's context/quality-dominated by `apex-fit-q8q5` / `byteshape-iq4xs`. Quant: **ablation-informed per-tensor mixed precision** (each tensor's Q2_K sensitivity measured, precision allocated under a size budget) 11.96 GB GGUF on a stock ggml-org image. Det-5-pack **63/75** (toolcall 13 · instructfollow 15 · structoutput 14 · dataextract 9 · reasonmath 12; benchlocal `--medium`, sandboxed packs not run no docker). TPS at `-np 4` provenance (single-card convention is `-np 1`). Drove the README `llama.cpp ❌→✅` correction (`559a3fe`). Weights + per-question evidence: [deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF](https://huggingface.co/deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF). |
| **`qwen3.6-35b-a3b/llama-cpp/dual/morikomorizz-q6kp/mtp.yml` (TP=2, q8_0 KV, MTP n=3)** 🧪 | @noonghunna (2× 3090 PCIe) | q8_0 / q8_0 | **262K** | **113.4 / ~150** (decode, n=3 shipped; wall 111.6 / ~144, CV<1% narr) | ~1820 @240K | (MTP draft-mtp n=3; acc narr 47% / code 83%) | n/a | **19.9 / 18.4 GB/card** | 2026-06-14 | **Uncensored "HauhauCS-Aggressive" 35B-A3B fine-tune** morikomorizz Q6_K_P GGUF with embedded nextn MTP head, on **mainline llama.cpp b9570**, dual-card layer-split `-ts 0.55,0.45` (rebalances the MTP draft card; even `1,1` skews ~3 GB at 262K). **MTP loads clean** (`speculative decoding context initialized`) the prior HauhauCS-MTP `ret=-3` was an ik-llama/older-build issue, not mainline. **n-sweep @262K (canonical bench.sh, same fresh container): n=3 113.4/~150 vs n=1 125.2/136.4 → n=3 = 9% prose / +10% code** (code-leaning default by maintainer request; n=1 is prose-best). An earlier rebench read 147/162 @ n=1 that **did NOT reproduce** on fresh boot (125 reproducible GPU-state TPS variance; the 147 is not trusted). `verify-stress` **8/8 incl. NIAH→240K** (91% of 262K), `soak` fresh 20×5 **PASS** (0 MiB growth, 0/100 silent-empty, p50 162.4 TPS, 99.6% retention). **8-pack think-OFF 103/150 (69%) · think-ON 105/150 (70%)** (wash; reasoning-on costs heavy latency hermes hit 300s for +2): toolcall 14/14 · instructfollow 12/14 · structoutput 14/14 · dataextract 11/8 · reasonmath 11/12 · bugfind 13/13 · hermesagent 11/11 (of 20) · cli 17/19 (of 40). Mid-band for the class (byteshape 110); uncensoring buys compliance, not capability. **REASONING=on default** (vs stack thinking-off). Community GGUF, **digest-UNPINNED** + uncensored 🧪. Compose: `models/qwen3.6-35b-a3b/llama-cpp/compose/dual/morikomorizz-q6kp/mtp.yml`. Registry tag: `llamacpp/hauhaucs-35ba3b-dual`. |
| **`gemma-4-26b-a4b/dual/awq.yml` (TP=2)** | @noonghunna (2× 3090 PCIe, 230 W cap) | bf16 | 32K | **138.88 / 138.67** (decode 139.92 / 139.98) | | n/a (no drafter) | n/a | **23.45 GB/card** | 2026-05-15 | **First v0.7.3 Gemma MoE production-track row.** Engine `vllm-nightly-clean` (nightly `bf610c2f`) + vendored [vLLM PR #40886](https://github.com/vllm-project/vllm/pull/40886) overlay (compressed-tensors AWQ MoE key remapping applied at boot via anchor-based Python patcher in `models/gemma-4-26b-a4b/vllm/patches/vllm-pr40886-awq-moe-keys/install.sh`). Weights: `cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit` (17 GB). CV **0.2% / 0.0%** extraordinarily stable, characteristic of MoE memory-bandwidth-bound decode. TTFT 53 ms. GPU 0 at 98% util / 360 W, GPU 1 at 59% / 299 W. **Identical narr/code wall TPS (139 / 139)** same MoE signature as Qwen 3.6 35B-A3B preview above (per-token weight reads dominate, prompt content distribution doesn't matter). ~76% of Qwen 35B-A3B preview's narr TPS consistent with 4 B vs 3 B active params per forward. Compose: `models/gemma-4-26b-a4b/vllm/compose/dual/awq/bf16.yml`. |