BENCHMARKS: add @guybrush01 dual-5090 cross-rig row (Qwen3.6-27B, first Blackwell)
First Blackwell / dual-5090 datapoint: qwen3.6-27b dual (AutoRound-INT4 + fp8 KV + MTP n=3) on vLLM v0.22.0 TP=2, sm_120 — 153.41/196.91 wall TPS (~2x the 3090 dual), verify-stress 8/8, soak PASS. Issue #474. Co-Authored-By: Claude Opus 4.8 <[email protected]> Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
6cdefab1f3
commit
2c686e96b6
@@ -124,6 +124,7 @@ Stock Qwen3.6-27B, single-card llama.cpp (`--reasoning on`), **FREE thinking-on
|
||||
| `dual-dflash.yml` | @apriori (2× 3090 + EPYC 7302P, Arch Linux, 230 W cap, NODE topology, no NVLink) | fp8 | 185K | **78.44 / 122.71** | — | ~24.0 GB | 2026-05-05 | **First EPYC + Arch cross-rig data on `dual-dflash`** — matches @noonghunna baseline within run-to-run CV (78/127 reference, narr drift +0.4 / code −3.4%). **PASSES continuous soak** (0 MiB VRAM growth, 0 errors, 0/25 silent-empty, 100% TPS retention) — first independent confirmation `dual-dflash` is Cliff 2b clean cross-rig. 3 turns >30s TTFT warning (informational). [Discussion #18](https://github.com/noonghunna/club-3090/discussions/18#discussioncomment-16819551). |
|
||||
| `dual-dflash-noviz.yml` | @noonghunna (2× 3090 PCIe) | fp8 | 200K | 78 / **127** | — | ~23.8 GB | 2026-04-29 | DFlash + no vision tower. +15K ctx vs `dual-dflash`. |
|
||||
| `dual-dflash-noviz.yml` | @snoby (2× **4090** PCIe — 5-GPU rig, GPUs 2,3, no NVLink, [#46](https://github.com/noonghunna/club-3090/issues/46)) | fp8 | **180K** | 92.55 / **148.99** | — | ~21.8 GB | 2026-05-04 | First non-3090 cross-rig data. **Required `max-model-len` drop from 200K→180K** vs 3090 baseline (boot OOM at 200K) — 4090 ctx-ceiling gotcha pending investigation. +17% TPS lift vs same compose on 3090 (78→92.55 narr / 127→148.99 code). |
|
||||
| `dual.yml` | @guybrush01 (2× **5090** PCIe 5.0 x8, Ryzen 9 9950X3D2, CachyOS, driver 610.43.02, 575 W cap, no NVLink) | fp8 | 262K | **153.41 / 196.91** (decode 154.62 / 200.13) | — | ~31 GB/card (934 MB free at 262K ceiling) | 2026-06-25 | **First Blackwell / dual-5090 cross-rig data.** Qwen3.6-27B dual (AutoRound-INT4 + fp8 KV + MTP n=3) on vLLM v0.22.0, TP=2, sm_120 — boots + serves the full 262K with **no changes**. **~2× the 3090 dual** (vs @noonghunna v7.72.2 ~89/115) at tiny TTFT (~51/53 ms), CV 1.5%/1.2%. MTP accept 2.71–3.03. verify-stress **8/8** (NIAH→240K, Cliff 2 clear); soak-continuous **PASS** (0 err, 0/25 silent-empty, 0 MiB growth, p50 decode 252). Blackwell symm-mem all-reduce N/A → PYNCCL fallback (custom all-reduce off anyway, as on our PCIe rigs). PCIe 5.0 x8 ≈ Gen4 x16. [Issue #474](https://github.com/noonghunna/club-3090/issues/474). |
|
||||
| `dual-nvlink.yml` | @JusefPol (2× 3090 PCIe x8 + **NVLink 4× bonded**, i7-11700K, 365 W/card) | fp8 | 262K | **108.81 / 138.55** | — | ~23.7 GB | 2026-05-04 | First NVLink cross-rig data. **+58% narr / +56% code TPS vs `dual.yml` PCIe-only baseline (69 / 89)** — NVLink reduces the per-token NCCL allreduce latency floor; compounds at multi-stream. verify-stress 8/8 PASS incl. 91K needle. **PASSES v2 continuous soak** (5 sessions × 5 turns, 0 MiB growth, 100% TPS retention). MTP n=3, 65–98% per-position accept. PR [#31](https://github.com/noonghunna/club-3090/pull/31). |
|
||||
| `dual-nvlink-turbo.yml` ⭐ | @danbedford (2× 3090 NVLink, 230W cap) | TQ3 | 262K | **102.34 / 133.98** | — | ~22.3 GB | 2026-05-05 | **v7.72.2-rebench** (image `nightly-01d4d1ad3`). 4-stream TurboQuant KV + NVLink. **+11% narr / +12% code vs same-rig PCIe `dual-turbo` (#73 below)** — controlled A/B on identical hardware, only `NCCL_P2P_LEVEL` differs. Custom all-reduce ENABLED (disabled on PCIe). CV 3.1% narr / 1.8% code. PR [#56](https://github.com/noonghunna/club-3090/pull/56) + [Issue #69](https://github.com/noonghunna/club-3090/issues/69). |
|
||||
| `dual.yml` | @danbedford (2× 3090 NVLink-cable-attached, run as PCIe via `NCCL_P2P_DISABLE=1`, 230W cap) | fp8 | 262K | **89.24 / 114.57** | — | ~23.7 GB | 2026-05-06 | **First controlled PCIe-vs-NVLink A/B on same rig** — pair with `dual-nvlink.yml` row immediately above. **+15% narr / +15% code lift from NVLink** (#74 102/132 vs this 89/115). CV 3.8%/2.5%. **Note: this corrects the "+58% narr / +56% code" claim from JusefPol's row** — that comparison conflated NVLink lift with v7.72.2 lift (his baseline was 2026-04-29 dual.yml at 69/89 on the older image). On a strictly v7.72.2-controlled comparison NVLink adds ~15%, not ~58%. [Issue #77](https://github.com/noonghunna/club-3090/issues/77). |
|
||||
|
||||
Reference in New Issue
Block a user