7 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 0f0a9c84c5 fix(scripts): report.sh shows human-readable version via git describe
Release / release (push) Failing after 51s
Per feedback from alanspires #127 ("the report doesn't show club-3090 version
in use?"): the line WAS there but printed as a bare SHA which is opaque to
anyone not knowing the SHA→tag mapping.

Changed:
  Before: - **club-3090:** `0d48dac` (branch: `master`)
  After:  - **club-3090:** `v0.6.2-4-ge299e70-dirty` (branch: `master`, SHA `e299e70`)

Uses `git describe --tags --always --dirty` for the version. Falls back to
raw SHA if no tags are reachable. SHA is still printed as a parenthetical
hint for log-cross-referencing.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-14 11:18:09 +00:00
noonghunnaandJohn Karabudak 3fe5aec314 Merge PR #128: unify dual-card composes with NVLink auto-detection
Closes #128.

@JohnTheNerd's contribution — discussed in #98. Net -639 lines across 22 files.

What lands:
- scripts/detect_nvlink.sh — runtime NVLink detection at compose boot via
  nvidia-smi topo -m. NVLINK_MODE=auto|force_on|force_off override.
- 4 Qwen NVLink composes (nvlink/turbo/dflash/dflash-noviz) collapsed
  from 871 lines into 32-line stubs that 'extends:' the unified compose
  with NVLINK_MODE=force_on. Back-compat preserved (same container_name,
  same ports — switch.sh routes unchanged).
- 4 Qwen non-NVLink composes (docker-compose, turbo, dflash, dflash-noviz)
  gain the unified entrypoint that sources detect_nvlink.sh and selects
  --disable-custom-all-reduce based on detection.
- Same pattern extended to all 7 Gemma dual composes.
- launch.sh + switch.sh: doc-only updates noting the new env var + that
  nvlink-* variants are stubs.
- docs/HARDWARE.md NVLink section rewritten.

Validated by author on their rig (full report.sh --full attached to PR).
Local smoke on canonical 2× 3090 PCIe (no NVLink bridge): unified dual.yml
boots clean, detect_nvlink.sh correctly reports 'PCIe topology (PHB) →
using PCIe mode', verify-full 8/8 PASS including MTP AL 2.67.

v0.6.1/v0.6.2 metadata preserved across all touched composes: Engine-profile
headers, ${TP:-N}/${PP:-1} env overrides, Requires-min-* fit metadata.

Co-Authored-By: John Karabudak <JohnTheNerd>
2026-05-14 11:17:17 +00:00
noonghunnaandClaude Opus 4.7 e41323ffa9 docs: cross-rig data — eddie 3090/3090Ti power-cap + alanspires 6×3090 VFIO
- HARDWARE.md: +2 rows in power-cap table — @eddietheengineer's 3090 air
  (vLLM dual MTP, knee 210W / 0.144 TPS/W) and 3090 Ti air (knee 200W /
  0.154 TPS/W, FIRST 3090 Ti data point on the matrix).
- BENCHMARKS.md: +1 multi4 row — @alanspires #127 on 6× 3090
  VFIO-passthrough + AMD EPYC 7313 host. TP=4 (Qwen num_kv_heads=4 doesn't
  divide 6 → GPUs 4-5 free). 74.93 / 92.85 TPS. First VFIO + virtualized
  data point; scripts (verify-full / verify-stress / soak) all PASS
  without modification.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-14 10:47:44 +00:00
John Karabudak e00626a50e feat: unify dual-card composes with NVLink auto-detection
As discussed in #98, collapse 8 Qwen dual-card compose files into 4.
NVLink presence is now detected at startup instead of requiring
a separate compose per interconnect. The 4 nvlink-*.yml files become
deprecated stubs that extend the unified compose with NVLINK_MODE=force_on.

Add entrypoint conditionals and env-var substitution so the same compose
selects the correct NCCL settings and --disable-custom-all-reduce flag based
on detected topology. Extend the same pattern to all 7 Gemma dual composes.
2026-05-14 02:36:15 -02:30
github-actions[bot] c181aec3cc chore(changelog): regenerate for v0.6.2 [skip ci] 2026-05-14 00:33:05 +00:00
noonghunna 98f0406d0f fix(launch): project TP greater than four
Release / release (push) Failing after 48s
2026-05-14 00:23:32 +00:00
github-actions[bot] 2a92a199ff chore(changelog): regenerate for v0.6.1 [skip ci] 2026-05-14 00:02:46 +00:00
26 changed files with 355 additions and 911 deletions
+7
View File
@@ -49,6 +49,13 @@
# NVIDIA_VISIBLE_DEVICES explicitly. Use comma-separated physical indices.
# NVIDIA_VISIBLE_DEVICES=0,1
# NVLINK_MODE=auto|force_on|force_off — NVLink auto-detection for dual-card composes
# auto (default): detects NVLink via nvidia-smi topo -m
# force_on: assume NVLink bridge present, set NVLink env vars
# force_off: force PCIe-only path even if NVLink detected
# This only affects dual-card (TP=2) composes. Single-card and multi4 are unaffected.
# NVLINK_MODE=auto
# -----------------------------------------------------------------------------
# vLLM tuning knobs
+3
View File
@@ -81,6 +81,8 @@ Primary serving model. Hybrid Qwen3-Next architecture (DeltaNet GDN + standard a
### Dual-card (2× RTX 3090, TP=2)
> NVLink auto-detection: dual-card composes now detect NVLink presence automatically. The `dual-nvlink*.yml` files are deprecated stubs that extend the unified compose with `NVLINK_MODE=force_on`. All NVLink bench rows below were measured with NVLink enabled (either via auto-detection or the deprecated stub). PCIe rows used `NCCL_P2P_DISABLE=1`.
| Compose | Rig | KV | Max ctx | Narr / Code TPS | Peak VRAM | Date | Notes |
|---|---|---|---:|---:|---:|---|---|
| `dual.yml` ⭐ | @noonghunna (2× 3090 PCIe, no NVLink) | fp8 | 262K (237K single-prompt verified) | 69 / 89 | ~23.6 GB | 2026-04-29 | tested 2-card baseline. fp8 KV, 2 streams, full feature set. **PASSES v2 continuous soak** (Cliff 2b clean). |
@@ -113,6 +115,7 @@ Primary serving model. Hybrid Qwen3-Next architecture (DeltaNet GDN + standard a
| Compose | Rig | KV | Max ctx | Narr / Code TPS | Peak VRAM | Date | Notes |
|---|---|---|---:|---:|---:|---|---|
| `multi4.yml` | @whamp (4× 3090 PCIe x4/x16/x8/x16, 300 W cap, no NVLink) | fp8 | 262K | 63 / 76 | ~23.5 GB | 2026-05-03 | TP=4 capacity king. **6.77× concurrency at 262K**. PASSES v2 continuous soak (20 sessions, 0 MiB growth, 90.8% TPS retention). PR [#44](https://github.com/noonghunna/club-3090/pull/44). |
| `multi4.yml` | [@alanspires #127](https://github.com/noonghunna/club-3090/issues/127) (**6× 3090 VFIO-passthrough**, AMD EPYC 7313 host, all cards 250W cap, no NVLink) | fp8 | 262K | **74.93 / 92.85** | ~21.7 GB | 2026-05-14 | TP=4 on a 6-card rig — GPUs 4-5 free (Qwen `num_kv_heads=4` doesn't divide 6, so TP=6 invalid). **First VFIO-passthrough + virtualized data point** — scripts (verify-full / verify-stress / soak) all PASS on the virt envelope without modification. KV pool 1,774,963 tokens at 6.77× concurrency. Soak p50 122.84 / p95 161.14 across 5 multi-turn sessions, 0 errors. |
| `multi4-dflash.yml` | @whamp (4× 3090 PCIe x4/x16/x8/x16, 300 W cap) | fp8 | 262K | 64 / **104** | ~22.0 GB | 2026-05-03 | TP=4 + DFlash. 2.27× concurrency at 262K. PASSES v2 continuous soak (5 sessions, 0 MiB growth, 100% TPS retention). **Bench-vs-soak inversion**: bench shows DFlash wins by 37% on short-prompt code, soak shows DFlash *loses* by 47% on multi-turn agent — DFlash AL likely collapses on mixed prompts. PR [#44](https://github.com/noonghunna/club-3090/pull/44). |
### Verify-stress + soak-continuous matrix
+31
View File
@@ -16,6 +16,37 @@ history; SemVer takes over from `v0.3.0` onward.
---
## v0.6.2 — 2026-05-14
### 🐛 Bug fixes
- fix(launch): project TP greater than four ([98f0406](https://github.com/noonghunna/club-3090/commit/98f0406d0f265767b4a6712a1eba293b1e2dc889))
[Pin: `git checkout v0.6.2`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.6.1...v0.6.2)
## v0.6.1 — 2026-05-14
### ✨ Features
- feat(launch): add hardware-aware launcher ([5882bbe](https://github.com/noonghunna/club-3090/commit/5882bbef6f5ed7e8ceea450fa3a6b167a1bf4926))
- feat(tools): extend kv-calc.py to multi-model (Qwen 3.6 + Gemma 4 31B) ([0d48dac](https://github.com/noonghunna/club-3090/commit/0d48dac818ba889f74b74df1c454bb961a123c37))
### 📝 Documentation
- docs: update launch.sh references for v0.6.1 wizard flow ([e299e70](https://github.com/noonghunna/club-3090/commit/e299e70451c8d146214a6560e582d0e174dd0ebc))
### 🧹 Other
- Merge codex/v0.6.1-launch into master ([056dcb6](https://github.com/noonghunna/club-3090/commit/056dcb643914fee6169b02b89cb420b038c29b0f))
[Pin: `git checkout v0.6.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.6.0...v0.6.1)
## v0.6.0 — 2026-05-13
+14 -3
View File
@@ -1,6 +1,8 @@
# Dual 3090 — what changes when you add the second card
You have **2× RTX 3090s, PCIe-only (no NVLink)**. This page is the front door for picking a config and knowing what dual-card unlocks vs single. Model-specific deep dives (quants, Genesis, engine internals) live in the model directory — links at the bottom.
You have **2× RTX 3090s**. This page is the front door for picking a config and knowing what dual-card unlocks vs single. Model-specific deep dives (quants, Genesis, engine internals) live in the model directory — links at the bottom.
**NVLink auto-detection** (since 2026-05-14): the dual-card composes now auto-detect whether an NVLink bridge is present. If you have one, you get the NVLink-optimized path automatically. If not, PCIe mode is used. Override with `NVLINK_MODE=force_on|force_off` in your `.env`. See the "NVLink auto-detection" section below.
> **Have 3+ GPUs?** See [`MULTI_CARD.md`](MULTI_CARD.md) — derivation of TP=4 / TP=8 configs from `dual.yml`, valid TP values for Qwen3.6-27B (1, 2, 4, 5, 8, 10), and what scales vs what doesn't.
@@ -134,9 +136,18 @@ git clone https://github.com/vllm-project/vllm.git /opt/ai/engines/vllm/primary
cd /opt/ai/engines/vllm/primary && git checkout main
```
### PCIe allreduce overhead (no NVLink)
### NVLink auto-detection
`--disable-custom-all-reduce` is set in all dual composes. Without it, vLLM tries to use a custom CUDA path that assumes NVLink topology and crashes. The trade is some allreduce latency on every layer, hence the per-stream TPS being lower than you'd see on an A100/A5000 dual setup with NVLink. Don't bother with NVLink bridges; this stack is intentionally PCIe-tested.
The dual-card composes automatically detect whether an NVLink bridge is installed and configure themselves accordingly. No separate compose files needed — `dual.yml`, `dual-turbo.yml`, `dual-dflash.yml`, and `dual-dflash-noviz.yml` all adapt to your hardware.
**How it works:** Each dual compose mounts `scripts/detect_nvlink.sh` and sources it in the entrypoint at container boot. The script checks `nvidia-smi topo -m` for NVLink links between GPUs, sets the correct NCCL env vars, and the entrypoint conditionally passes `--disable-custom-all-reduce` to vLLM.
**Override:** Set `NVLINK_MODE` in your `.env` (passed through to the container):
- `auto` (default) — detect via `nvidia-smi topo -m`
- `force_on` — assume NVLink bridge present, enable NVLink mode
- `force_off` — force PCIe-only path even if NVLink detected
Without NVLink, `--disable-custom-all-reduce` is passed to vLLM and `NCCL_P2P_DISABLE=1` is set. With NVLink, custom all-reduce is enabled and NCCL uses the NVLink path. The per-stream TPS difference is ~10-15% on dual 3090 (see cross-rig data in [BENCHMARKS.md](../BENCHMARKS.md)).
### `dual.yml` is Genesis-less by design
+7 -5
View File
@@ -67,13 +67,13 @@ Same `MAX_MODEL_LEN` / `GPU_MEMORY_UTILIZATION` env overrides apply for any setu
## NVLink
**Not required.** We've explicitly designed for PCIe-only consumer setups.
**Not required.** Dual-card composes auto-detect NVLink and configure themselves accordingly.
- 3090s have an NVLink connector but a **bridge has to be physically installed**. Most consumer setups don't have one. (Cost: ~$70-150 for a working 3-slot bridge if you wanted to add one.)
- Our composes set `NCCL_P2P_DISABLE=1` and avoid NVLink-dependent allreduce paths.
- **If you have NVLink installed and working**, single-stream TPS on dual-card will be ~1.6-1.8× single-card (vs ~1.05× without). Concurrent throughput scales similarly. Not a huge deal unless you really care about per-stream speed.
The user explicitly chose to operate without NVLink. Don't suggest adding one.
- **Auto-detection**: each dual compose sources `scripts/detect_nvlink.sh` in its entrypoint at boot. The script checks `nvidia-smi topo -m` and sets the correct NCCL env vars + vLLM flags.
- **Override**: set `NVLINK_MODE=force_on|force_off` in your `.env` to bypass auto-detection.
- Without NVLink (PCIe), `--disable-custom-all-reduce` is passed to vLLM and `NCCL_P2P_DISABLE=1` is set. With NVLink, custom all-reduce is enabled and NCCL uses the NVLink path.
- **If you have NVLink installed and working**, single-stream TPS on dual-card will be ~1.6-1.8× single-card (vs ~1.05× without). Measured NVLink lift is ~10-15% over PCIe on the same rig. See [BENCHMARKS.md](../BENCHMARKS.md) for cross-rig data.
---
@@ -162,6 +162,8 @@ How `decode-single` is timed (the new default since 2026-05-07):
| 3090 | water | llama.cpp default | Qwen3.6 27B Q3_K_XL | **330W** ⭐ | 36.35 | 36.26 | 0.110 | [@syangsao #58](https://github.com/noonghunna/club-3090/issues/58#issuecomment-4388766174) |
| 3090 | water | llama.cpp default | Qwen3.6 27B Q3_K_XL | 388W (stock) | 38.23 | 37.97 | 0.098 | [@syangsao #58](https://github.com/noonghunna/club-3090/issues/58#issuecomment-4388766174) |
| 3090 | air | llama.cpp default | Qwen3.6 27B Q3_K_XL | **290W** ⭐ | 32.26 | 32.17 | **0.111** | @noonghunna (this rig, 21-cap 10W sweep, time-bounded bench, SM 1380 MHz at sweet spot) |
| 3090 | air | vLLM dual + MTP | Qwen3.6 27B AutoRound | **210W** ⭐ | 30.23 | 30.57 | **0.144** | [@eddietheengineer #86](https://github.com/noonghunna/club-3090/discussions/86#discussioncomment-16918020) (vLLM dual MTP, 10-cap sweep 180-250W, knee identical to llama.cpp 290W in absolute draw ratio) |
| **3090 Ti** | air | vLLM dual + MTP | Qwen3.6 27B AutoRound | **200W** ⭐ | 30.72 | 31.24 | **0.154** | [@eddietheengineer #86](https://github.com/noonghunna/club-3090/discussions/86#discussioncomment-16918020) — **first 3090 Ti data point on this matrix.** Hits knee at lower cap than 3090 despite higher 480W stock TDP. Sub-knee plateau visible at 100-120W (SM stalls at 225-240 MHz, throttle 100%). |
| 3090 | air | llama.cpp default | Qwen3.6 27B Q3_K_XL | 370W (stock) | 34.66 | 34.67 | 0.104 | same — SM locks at 1560 MHz across 340-370W (boost-clock plateau) |
| 3090 | air | llama.cpp default | Qwen3.6 27B Q3_K_XL | 390W (max) | 36.26 | 36.06 | 0.093 | same — SM 1680 MHz at 388W draw |
| 3090 | air | llama.cpp `decode-concurrent` N=4 | Qwen3.6 27B Q3_K_XL | **290W** ⭐ | 31.74 | 29.98 | **0.110** | @noonghunna (this rig, 21-cap, 4-stream aggregate, 8m wall) |
+11 -3
View File
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (cyankiwi AWQ-4bit weights — different quant from AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: MTP n=4 default (n=8 single-stream code variant via MTP_N=8)
# KV: bfloat16 (AWQ doesn't use FP8 KV — bypasses PR #40391 entirely)
# Vision: yes
@@ -76,10 +76,13 @@ services:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
- ../../cache/torch_compile_awq:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton_awq:/root/.triton/cache
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -107,7 +110,13 @@ services:
- |
set -e
pip install --quiet --upgrade transformers==5.8.0
exec vllm serve "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
command:
- --host
@@ -124,7 +133,6 @@ services:
- "${PP:-1}"
- --dtype
- "${DTYPE:-bfloat16}"
- --disable-custom-all-reduce
- --trust-remote-code
- --enable-chunked-prefill
# ---- max-model-len ----
+16 -2
View File
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (Intel AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: MTP n=3 (Google's official `gemma-4-31B-it-assistant`, BF16 0.5B)
# KV: auto (1 byte/token, via vendored PR #40391 overlay)
# Vision: yes
@@ -121,11 +121,14 @@ services:
# ---- Gemma 4 tool-parser stacked fixes (PR #42006 + PR #41991) ----
# 1 file: MTP streaming multi-tool + parser bounds. Drop when both PRs merge.
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -142,6 +145,18 @@ services:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
command:
- --host
- 0.0.0.0
@@ -155,7 +170,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# ---- KV format ----
# Default `auto` runs on every consumer GPU including
# sm_86 Ampere (uses standard PyTorch torch.int8 ops, not the Triton
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (Intel AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: z-lab Gemma 4 DFlash n=7 (block-diffusion, BF16 KV in independent pool)
# KV: int8_per_token_head (target) + bfloat16 (drafter — independent pools)
# Vision: yes
@@ -142,11 +142,14 @@ services:
- ../../patches/vllm-gemma4-dflash-int8/v1/worker/gpu/attn_utils.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu/attn_utils.py:ro
- ../../patches/vllm-gemma4-dflash-int8/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
- ../../patches/vllm-gemma4-dflash-int8/v1/worker/kv_cache_shape_utils.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/kv_cache_shape_utils.py:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -172,7 +175,13 @@ services:
- |
set -e
pip install --quiet --upgrade transformers==5.8.0
exec vllm serve "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
command:
- --host
@@ -189,7 +198,6 @@ services:
- "${PP:-1}"
- --dtype
- bfloat16
- --disable-custom-all-reduce
# ---- KV format (per-token-head INT8 — Ampere-compatible) ----
- --kv-cache-dtype
- "${KV_DTYPE:-int8_per_token_head}"
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (Intel AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: z-lab Gemma 4 DFlash n=7 (block-diffusion drafter, vLLM PR #41703 vendored overlay)
# KV: bfloat16 (forced for drafter compatibility)
# Vision: yes
@@ -107,11 +107,14 @@ services:
- ../../patches/vllm-gemma4-dflash/v1/spec_decode/eagle.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/spec_decode/eagle.py:ro
- ../../patches/vllm-gemma4-dflash/v1/spec_decode/utils.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/spec_decode/utils.py:ro
- ../../patches/vllm-gemma4-dflash/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -137,7 +140,13 @@ services:
- |
set -e
pip install --quiet --upgrade transformers==5.8.0
exec vllm serve "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
command:
- --host
@@ -157,7 +166,6 @@ services:
# workaround until that lands).
- --dtype
- bfloat16
- --disable-custom-all-reduce
# max-model-len ceiling (TP=2, BF16 KV, mem-util 0.95):
# 32K ctx @ 0.92 → 38,339 token KV pool (shipped default — matches bench above)
# 48K ctx @ 0.95 → ~46,000 token KV pool (safe, ~2 GB headroom)
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (Intel AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink) — single-card boot-OOMs on 24 GB
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE) — single-card boot-OOMs on 24 GB
# Drafter: MTP n=3 (Google's official `gemma-4-31B-it-assistant`, BF16 0.5B)
# KV: bfloat16 (2 bytes/token)
# Vision: yes
@@ -67,10 +67,13 @@ services:
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -90,6 +93,18 @@ services:
# transformers 5.8.0+ ships in the post-Gemma4-merge nightly (verified
# 2026-05-08: nightly-1acd67a795... has transformers 5.8.0). Entrypoint
# upgrade dropped.
entrypoint:
- /bin/bash
- -c
- |
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
command:
- --host
- 0.0.0.0
@@ -103,7 +118,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# max-model-len ceiling (TP=2, BF16 KV):
# 32K ctx @ 0.92 → KV pool ~38K tokens (shipped default — matches bench above)
# 48K ctx @ 0.95 → KV pool ~50K tokens (safe — MTP drafter is 0.5B vs DFlash's 2.9B,
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (Intel AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: MTP n=4 (Google's official `gemma-4-31B-it-assistant`, BF16 0.5B)
# KV: turboquant_3bit_nc
# Status: ⛔ HARDWARE-BLOCKED on Ampere — see 2026-05-11 walkthrough below.
@@ -217,11 +217,14 @@ services:
# ---- Gemma 4 tool-parser stacked fixes (PR #42006 + PR #41991) ----
# 1 file: MTP streaming multi-tool + parser bounds. Drop when both PRs merge.
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -248,7 +251,13 @@ services:
echo "[club3090-tq3] Gate 3: dropping turboquant-vllm plugin..." >&2
pip uninstall -y turboquant-vllm 2>&1 | tail -1 >&2 || true
echo "[club3090-tq3] Launching vllm serve..." >&2
exec vllm serve "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
shm_size: "16gb"
ipc: host
@@ -272,7 +281,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# ---- 2026-05-11 workaround flags ----
# Force TURBOQUANT backend — required because vLLM has no per-layer
# backend routing (Gate 2). Without force, vLLM picks TRITON_ATTN
+16 -2
View File
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Gemma 4 31B (Intel AutoRound INT4)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: MTP n=3 (Google's official `gemma-4-31B-it-assistant`, BF16 0.5B)
# KV: int8_per_token_head (1 byte/token, via vendored PR #40391 overlay)
# Vision: yes
@@ -121,11 +121,14 @@ services:
# ---- Gemma 4 tool-parser stacked fixes (PR #42006 + PR #41991) ----
# 1 file: MTP streaming multi-tool + parser bounds. Drop when both PRs merge.
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -142,6 +145,18 @@ services:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve "$@"
else
exec vllm serve --disable-custom-all-reduce "$@"
fi
- --
command:
- --host
- 0.0.0.0
@@ -155,7 +170,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# ---- KV format ----
# Default `int8_per_token_head` runs on every consumer GPU including
# sm_86 Ampere (uses standard PyTorch torch.int8 ops, not the Triton
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: z-lab DFlash N=5 (block-diffusion drafter)
# KV: fp8_e5m2 (forced bf16 for drafter via --dtype bfloat16)
# Vision: no (vision tower dropped — frees ~0.5 GB/card for ctx)
@@ -81,10 +81,13 @@ services:
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -111,7 +114,13 @@ services:
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
else
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} --disable-custom-all-reduce "$@"
fi
- --
command:
- --model
@@ -126,7 +135,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-200000}"
- --gpu-memory-utilization
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: z-lab DFlash N=5 (block-diffusion drafter)
# KV: fp8_e5m2 (forced bf16 for drafter via --dtype bfloat16)
# Vision: yes
@@ -105,10 +105,13 @@ services:
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -135,7 +138,13 @@ services:
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
else
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} --disable-custom-all-reduce "$@"
fi
- --
command:
- --model
@@ -150,7 +159,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-185000}"
- --gpu-memory-utilization
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: MTP n=3 (built-in)
# KV: fp8_e5m2 (1 byte/token)
# Vision: yes
@@ -93,11 +93,13 @@ services:
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# PCIe-only stack — disable NCCL features that assume NVLink.
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -126,7 +128,13 @@ services:
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
else
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} --disable-custom-all-reduce "$@"
fi
- --
command:
- --model
@@ -141,7 +149,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
- --gpu-memory-utilization
@@ -1,197 +1,12 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 + NVLink (TP=2, NCCL P2P over NVLink)
# Drafter: z-lab DFlash N=5 (block-diffusion drafter)
# KV: FP16 (forced for DFlash drafter compatibility)
# Vision: no (frees ~0.5 GB/card for ctx)
# Max ctx: 200K (+15K vs dual-nvlink-dflash.yml; NVLink lifts max to 188K stable)
# Genesis: none
# Status: ✅ Production
# Best for: NVLink + peak code TPS, no-vision (max ctx headroom)
# ---------------------------------------------------------------------------
# Dual RTX 3090 with NVLink + DFlash text-only — TP=2 + DFlash N=5 + 200K ctx
# + NO vision + NVLink P2P for faster allreduce.
#
# Mirrors dual/dflash-noviz.yml but enables NCCL P2P over NVLink
# and re-enables vLLM's custom all-reduce kernel (which dual-dflash-noviz.yml
# disables for PCIe-only stacks). Combines DFlash N=5 draft model with NVLink
# bridge for maximum single-stream code throughput on 2x 3090 without vision.
#
# Status: COMMUNITY-CONTRIBUTED, EXPERIMENTAL.
# If you run this, please drop numbers in discussion #19 — paste-ready report
# via `bash scripts/report.sh`.
#
# What this gives you (vs the PCIe-only `dual-dflash-noviz.yml`):
# - NCCL P2P over NVLink (NCCL_P2P_LEVEL=NVL) — much faster allreduce on TP=2
# - Custom all-reduce ENABLED (--disable-custom-all-reduce removed) — NVLink
# makes vLLM's custom kernel a win where PCIe makes it a loss
# - PYTORCH_CUDA_ALLOC_CONF without `expandable_segments:True` — JusefPol
# reports it crashes on startup with NVLink wired up
#
# vs `dual/nvlink-dflash.yml` (the with-vision variant): drops
# MoonViT to free ~0.78 GiB per card → bumps max_model_len 185K → 200K
# (close to the absolute DFlash ceiling on dual-3090 with FP16 KV).
#
# Best for: long single-prompt text workloads (RAG / summarization / code
# review across very long codebases) where you don't need image input.
#
# KV cache: FP16 (default — DFlash needs head_size=256 + non-causal attention,
# and no Ampere backend supports that triple with fp8/turbo KV).
#
# ─── Prerequisite: download the DFlash draft model ──────────────────────
# Same as dual-dflash-noviz.yml — needs `z-lab/Qwen3.6-27B-DFlash` at
# `<MODEL_DIR>/qwen3.6-27b-dflash/`. Get it via:
#
# WITH_DFLASH_DRAFT=1 bash scripts/setup.sh qwen3.6-27b
#
# OR manually `hf download z-lab/Qwen3.6-27B-DFlash --local-dir <MODEL_DIR>/qwen3.6-27b-dflash`.
# If missing, vLLM falls back silently to baseline bf16 decode (~25 TPS
# instead of 125 TPS — reported by @lolren in club-3090#18). See
# `dual-dflash.yml` header for the under-training caveat.
#
# Dependencies:
# - 2x RTX 3090 (Ampere SM 8.6) WITH NVLink bridge installed and `nvidia-smi
# topo -m` showing `NV*` between GPU0 and GPU1
# - vLLM PR #40361 (Marlin pad-sub-tile-n) — patched files vendored in-repo
# at ../../patches/vllm-marlin-pad/. PR is open upstream; drop the mount when
# it lands. See ../../patches/vllm-marlin-pad/README.md.
#
# All dual-card variants in this dir:
#
# File Ctx Streams Narr/Code TPS KV Vision NVLink
# dual/docker-compose.yml (DEFAULT) 262K 2 69 / 89 fp8 ✓ not used
# multi4/docker-compose.yml 262K 4 63 / 76 fp8 ✓ not used (4x PCIe)
# multi4/dflash.yml 262K 2 64 / 104 FP16 ✓ not used (4x PCIe)
# dual/nvlink.yml 262K 2 (community) fp8 ✓ required
# dual/nvlink-turbo.yml 262K 4 101 / 133 TQ3 ✓ required
# dual/turbo.yml 262K 4 54 / 73 TQ3 ✓ not used
# dual/dflash.yml 185K 1 82 / 125 FP16 ✓ not used
# dual/dflash-noviz.yml 200K 1 78 / 127 FP16 ✗ not used
# dual/nvlink-dflash.yml 185K 1 (community) FP16 ✓ required
# dual/nvlink-dflash-noviz.yml 200K 1 (community) FP16 ✗ required
#
# To run:
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink-dflash-noviz.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# DEPRECATED (2026-05-14): use dual/dflash-noviz.yml — NVLink is auto-detected
# in the compose entrypoint. This stub forces NVLink mode for backward compatibility.
# Will be removed in a future release.
services:
vllm-qwen36-27b-dual-nvlink-dflash-noviz:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
extends:
file: dflash-noviz.yml
service: vllm-qwen36-27b-dual-dflash
container_name: vllm-qwen36-27b-dual-nvlink-dflash-noviz
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${PORT:-8019}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
# torch.compile + Triton kernel caches — first boot warms (~60-90 sec);
# subsequent boots reuse cached graphs. Pattern from Sander's PROD launch.
# Closes club-3090 #22.
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# Marlin pad-sub-tile-n (vLLM PR #40361) — vendored in this repo at
# ../../patches/vllm-marlin-pad/. Drops out when vllm#40361 lands upstream.
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
# which write to chat_completion/serving.py at vllm-import time. See
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
# dual-dflash-noviz.yml does for PCIe-only stacks) and pin the P2P level
# to NVLink so NCCL doesn't fall back to PCIe paths if topology query is fuzzy.
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_LEVEL=NVL
- VLLM_NO_USAGE_STATS=1
- VLLM_USE_FLASHINFER_SAMPLER=1
- OMP_NUM_THREADS=1
# JusefPol report (PR #31): expandable_segments=True crashes on startup
# with NVLink wired in. Keep max_split_size_mb cap, drop the rest.
- PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --model
- /root/.cache/huggingface/qwen3.6-27b-autoround-int4
- --served-model-name
- qwen3.6-27b-autoround
- --quantization
- auto_round
- --dtype
- bfloat16
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual-dflash-noviz.yml disables it
# because PCIe P2P bandwidth makes the NCCL fallback faster there.
- --max-model-len
- "${MAX_MODEL_LEN:-188000}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.95}"
- --max-num-seqs
- "1"
- --max-num-batched-tokens
- "8192"
# No --kv-cache-dtype: DFlash needs head_size=256 + non-causal attention,
# and no Ampere backend supports that triple with fp8/turbo KV. FP16 default
# is the only working choice (matches Qwen3.5-27B + DFlash row 4 = 89.7 TPS).
# --language-model-only frees ~0.78 GiB of MoonViT weights → 15K more KV ctx.
- --language-model-only
- --trust-remote-code
# froggeric chat-template override (see volumes block + docs/UPSTREAM.md)
- --chat-template
- /etc/qwen-froggeric-chat-template.jinja
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
- --speculative-config
- '{"method":"dflash","model":"/root/.cache/huggingface/qwen3.6-27b-dflash","num_speculative_tokens":5}'
- --host
- 0.0.0.0
- --port
- "8000"
- NVLINK_MODE=force_on
- PORT=8019
@@ -1,192 +1,12 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 + NVLink (TP=2, NCCL P2P over NVLink)
# Drafter: z-lab DFlash N=5 (block-diffusion drafter)
# KV: FP16 (forced for DFlash drafter compatibility)
# Vision: yes
# Max ctx: 185K
# Genesis: none — DFlash drafter handles spec-decode independently
# Status: ✅ Production
# Best for: NVLink + peak code TPS — +17% over PCIe-only dual-dflash.yml
# ---------------------------------------------------------------------------
# Dual RTX 3090 with NVLink + DFlash — TP=2 + DFlash N=5 spec-decode + 185K ctx
# + vision + NVLink P2P for faster allreduce.
#
# Mirrors dual/dflash.yml but enables NCCL P2P over NVLink and
# re-enables vLLM's custom all-reduce kernel (which dual-dflash.yml disables
# for PCIe-only stacks). Combines DFlash N=5 draft model with NVLink bridge
# for maximum single-stream code throughput on 2x 3090.
#
# Status: COMMUNITY-CONTRIBUTED, EXPERIMENTAL.
# If you run this, please drop numbers in discussion #19 — paste-ready report
# via `bash scripts/report.sh`.
#
# What this gives you (vs the PCIe-only `dual-dflash.yml`):
# - NCCL P2P over NVLink (NCCL_P2P_LEVEL=NVL) — much faster allreduce on TP=2
# - Custom all-reduce ENABLED (--disable-custom-all-reduce removed) — NVLink
# makes vLLM's custom kernel a win where PCIe makes it a loss
# - PYTORCH_CUDA_ALLOC_CONF without `expandable_segments:True` — JusefPol
# reports it crashes on startup with NVLink wired up
#
# What's intentionally NOT enabled:
# - TurboQuant KV — DFlash needs head_size=256 + non-causal attention,
# and no Ampere backend supports that triple with fp8/turbo KV.
#
# KV cache: FP16 (default — DFlash needs head_size=256 + non-causal attention,
# and no Ampere backend supports that triple with fp8/turbo KV).
#
# ─── Prerequisite: download the DFlash draft model ──────────────────────
# Same as dual-dflash.yml — needs `z-lab/Qwen3.6-27B-DFlash` at
# `<MODEL_DIR>/qwen3.6-27b-dflash/`. Get it via:
#
# WITH_DFLASH_DRAFT=1 bash scripts/setup.sh qwen3.6-27b
#
# OR manually `hf download z-lab/Qwen3.6-27B-DFlash --local-dir <MODEL_DIR>/qwen3.6-27b-dflash`.
# If missing, vLLM falls back silently to baseline bf16 decode (~25 TPS
# instead of 125 TPS — reported by @lolren in club-3090#18). See
# `dual-dflash.yml` header for the under-training caveat.
#
# Dependencies:
# - 2x RTX 3090 (Ampere SM 8.6) WITH NVLink bridge installed and `nvidia-smi
# topo -m` showing `NV*` between GPU0 and GPU1
# - vLLM PR #40361 (Marlin pad-sub-tile-n) — patched files vendored in-repo
# at ../../patches/vllm-marlin-pad/. PR is open upstream; drop the mount when
# it lands. See ../../patches/vllm-marlin-pad/README.md.
#
# All dual-card variants in this dir:
#
# File Ctx Streams Narr/Code TPS KV Vision NVLink
# dual/docker-compose.yml (DEFAULT) 262K 2 69 / 89 fp8 ✓ not used
# multi4/docker-compose.yml 262K 4 63 / 76 fp8 ✓ not used (4x PCIe)
# multi4/dflash.yml 262K 2 64 / 104 FP16 ✓ not used (4x PCIe)
# dual/nvlink.yml 262K 2 (community) fp8 ✓ required
# dual/nvlink-turbo.yml 262K 4 101 / 133 TQ3 ✓ required
# dual/turbo.yml 262K 4 54 / 73 TQ3 ✓ not used
# dual/dflash.yml 185K 1 82 / 125 FP16 ✓ not used
# dual/dflash-noviz.yml 200K 1 78 / 127 FP16 ✗ not used
# dual/nvlink-dflash.yml 185K 1 (community) FP16 ✓ required
#
# To run:
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink-dflash.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# DEPRECATED (2026-05-14): use dual/dflash.yml — NVLink is auto-detected
# in the compose entrypoint. This stub forces NVLink mode for backward compatibility.
# Will be removed in a future release.
services:
vllm-qwen36-27b-dual-nvlink-dflash:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
extends:
file: dflash.yml
service: vllm-qwen36-27b-dual-dflash
container_name: vllm-qwen36-27b-dual-nvlink-dflash
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${PORT:-8018}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
# torch.compile + Triton kernel caches — first boot warms (~60-90 sec);
# subsequent boots reuse cached graphs. Pattern from Sander's PROD launch.
# Closes club-3090 #22.
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# Marlin pad-sub-tile-n (vLLM PR #40361) — vendored in this repo at
# ../../patches/vllm-marlin-pad/. Drops out when vllm#40361 lands upstream.
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
# which write to chat_completion/serving.py at vllm-import time. See
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
# dual-dflash.yml does for PCIe-only stacks) and pin the P2P level to
# NVLink so NCCL doesn't fall back to PCIe paths if topology query is fuzzy.
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_LEVEL=NVL
- VLLM_NO_USAGE_STATS=1
- VLLM_USE_FLASHINFER_SAMPLER=1
- OMP_NUM_THREADS=1
# JusefPol report (PR #31): expandable_segments=True crashes on startup
# with NVLink wired in. Keep max_split_size_mb cap, drop the rest.
- PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
# hardware where graph capture causes OOM or instability (e.g. WSL2).
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --model
- /root/.cache/huggingface/qwen3.6-27b-autoround-int4
- --served-model-name
- qwen3.6-27b-autoround
- --quantization
- auto_round
- --dtype
- bfloat16
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual-dflash.yml disables it because
# PCIe P2P bandwidth makes the NCCL fallback faster there.
- --max-model-len
- "${MAX_MODEL_LEN:-185000}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.95}"
- --max-num-seqs
- "1"
- --max-num-batched-tokens
- "8192"
# No --kv-cache-dtype: DFlash needs head_size=256 + non-causal attention,
# and no Ampere backend supports that triple with fp8/turbo KV. FP16 default
# is the only working choice (matches Qwen3.5-27B + DFlash row 4 = 89.7 TPS).
# --language-model-only removed to enable MoonViT vision tower (2026-04-25 test).
- --trust-remote-code
# froggeric chat-template override (see volumes block + docs/UPSTREAM.md)
- --chat-template
- /etc/qwen-froggeric-chat-template.jinja
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
- --speculative-config
- '{"method":"dflash","model":"/root/.cache/huggingface/qwen3.6-27b-dflash","num_speculative_tokens":5}'
- --host
- 0.0.0.0
- --port
- "8000"
- NVLINK_MODE=force_on
- PORT=8018
@@ -1,300 +1,12 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 + NVLink (TP=2, NCCL P2P over NVLink)
# Drafter: MTP n=3 (built-in)
# KV: turboquant_3bit_nc (TQ3, 0.375 bytes/token)
# Vision: yes
# Max ctx: 262K
# Genesis: v7.72.2 (full PROD env stack)
# Status: ✅ Production
# Best for: NVLink + multi-tenant — +11% narr / +12% code over PCIe-only dual-turbo.yml
# ---------------------------------------------------------------------------
# Dual-card Turbo with NVLink — TP=2 + TurboQuant KV (turboquant_3bit_nc) + MTP n=3
# + Genesis v7.72.2 + NVLink P2P.
#
# Mirrors dual/turbo.yml but enables NCCL P2P over NVLink for
# faster allreduce, re-enables vLLM's custom all-reduce kernel, and drops
# expandable_segments (which JusefPol reports crashes with NVLink wired).
#
# Status: COMMUNITY-CONTRIBUTED, EXPERIMENTAL.
# If you run this, please drop numbers in discussion #19 — paste-ready report
# via `bash scripts/report.sh`.
#
# The 4-stream concurrent-serving NVLink variant. NVLink lifts the per-token
# allreduce-latency ceiling that caps PCIe-TP=2 decode; same TQ3 KV path so
# the 4-stream concurrency budget (lower KV-pool pressure) carries through.
#
# Measured (2× RTX 3090 w/ NVLink, Genesis v7.69, TQ3 KV — danbedford bench, 2026-05-04):
# 101.49 narr (CV 2.0%) / 133.20 code (CV 2.3%) wall TPS, VRAM 20.4 GB/card.
# NVLink gain over own PCIe-only dual-turbo.yml baseline (~90 / ~120, A/B
# tested on same rig): +12.6% narr, +10.7% code. Smaller delta than fp8's
# NVLink gain (compare dual-nvlink.yml vs dual.yml, +58% / +56%) because
# TQ3 already reduces per-token allreduce volume on PCIe.
#
# Genesis v7.69 P65 (cudagraph downgrade for spec-decode) makes MTP × TurboQuant
# work on dual-card TP=2 with vision + tools + 262K. Other relevant patches:
# - P4 hybrid turboquant support (replaces standalone PR #39931 patches)
# - P5 KV page-size unification for hybrid models
# - P64 streaming MTP tool-call edge case
# - P66 cudagraph_capture_sizes divisibility filter
# v7.72.2 native: PN34 (workspace_lock), P78 (tolist guard), PN35 (inputs_embeds).
#
# What this gives you (vs the PCIe-only `dual-turbo.yml`):
# - NCCL P2P over NVLink (NCCL_P2P_LEVEL=NVL) — faster allreduce on TP=2
# - Custom all-reduce ENABLED (--disable-custom-all-reduce removed) — NVLink
# makes vLLM's custom kernel a win where PCIe makes it a loss
# - PYTORCH_CUDA_ALLOC_CONF without `expandable_segments:True` — JusefPol
# reports it crashes on startup with NVLink wired up
#
# Dependencies:
# - 2× RTX 3090 (Ampere SM 8.6) WITH NVLink bridge installed and `nvidia-smi
# topo -m` showing `NV*` between GPU0 and GPU1
# - vLLM PR #40361 (Marlin pad-sub-tile-n) — patched files vendored in-repo
# at ../../patches/vllm-marlin-pad/. PR is open upstream; drop the mount when
# it lands. See ../../patches/vllm-marlin-pad/README.md.
#
# KV cache: turboquant_3bit_nc — 3-bit symmetric K and V (~3 bits average per
# token). Aligned with the single-card v714 default config we test extensively.
#
# To run:
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink-turbo.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# DEPRECATED (2026-05-14): use dual/turbo.yml — NVLink is auto-detected
# in the compose entrypoint. This stub forces NVLink mode for backward compatibility.
# Will be removed in a future release.
services:
vllm-qwen36-27b-dual-nvlink-turbo:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
extends:
file: turbo.yml
service: vllm-qwen36-27b-dual-turbo
container_name: vllm-qwen36-27b-dual-nvlink-turbo
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${PORT:-8017}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
# torch.compile + Triton kernel caches — first boot warms (~60-90 sec);
# subsequent boots reuse cached graphs. Pattern from Sander's PROD launch.
# Closes club-3090 #22.
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# Marlin pad-sub-tile-n (vLLM PR #40361) — vendored in this repo at
# ../../patches/vllm-marlin-pad/. Drops out when vllm#40361 lands upstream.
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
# Genesis modular package (auto-checked-out at pin in scripts/setup.sh).
# As of v7.72.2: PN35 supersedes patch_inputs_embeds_optional.py,
# P78 supersedes patch_tolist_cudagraph.py, PN34 supersedes
# patch_workspace_lock_disable.py — those local sidecars dropped
# from this compose 2026-05-05.
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
# which write to chat_completion/serving.py at vllm-import time. See
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
# dual-turbo.yml does for PCIe-only stacks) and pin the P2P level to
# NVLink so NCCL doesn't fall back to PCIe paths if topology query is fuzzy.
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_LEVEL=NVL
- VLLM_NO_USAGE_STATS=1
- VLLM_USE_FLASHINFER_SAMPLER=1
- OMP_NUM_THREADS=1
# JusefPol report (PR #31): expandable_segments=True crashes on startup
# with NVLink wired in. Keep max_split_size_mb cap, drop the rest.
- PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512
- VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
- VLLM_MARLIN_USE_ATOMIC_ADD=1
- TRITON_CACHE_DIR=/root/.triton/cache
- VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0
- VLLM_FLOAT32_MATMUL_PRECISION=high
# VLLM_SSM_CONV_STATE_LAYOUT=DS — RE-ENABLED with our local PN30
# dst-shaped temp fix (patch_pn30_dst_shaped_temp_fix.py, applied at
# setup time). Without our fix, Sander's PN30 a9977d8 corrupts DS row
# strides on spec-decode AL>1. With our fix, PN30 builds destination-
# shaped temp via collect_mamba_copy_meta. +6% TPS retained.
- VLLM_SSM_CONV_STATE_LAYOUT=DS
- VLLM_USE_FUSED_MOE_GROUPED_TOPK=1
- CUDA_DEVICE_MAX_CONNECTIONS=8
# FULL Genesis v7.69 PROD env-var set per Sandermage's
# bare_metal_27b_int4_TQ_k8v4.sh. Validated 2026-05-01 PM dual-3090:
# 116.59 code wall_TPS / 92.12 narrative wall_TPS (vs Sander's A5000
# 89.23 reference — +30.7% over).
- GENESIS_ENABLE_P4=1
- GENESIS_ENABLE_P58_ASYNC_PLACEHOLDER_FIX=1
- GENESIS_ENABLE_P60_GDN_NGRAM_FIX=1
- GENESIS_ENABLE_P60B_TRITON_KERNEL=1
- GENESIS_ENABLE_P61_QWEN3_MULTI_TOOL=1
- GENESIS_ENABLE_P61B_STREAMING_OVERLAP=1
- GENESIS_ENABLE_P62_STRUCT_OUT_SPEC_TIMING=1
- GENESIS_ENABLE_P64_QWEN3CODER_MTP_STREAMING=1
# P65 dropped 2026-05-03 (mutually exclusive with P67/P67b in v7.69)
- GENESIS_ENABLE_P66_CUDAGRAPH_SIZE_FILTER=1
- GENESIS_ENABLE_P67_TQ_MULTI_QUERY_KERNEL=1
- GENESIS_ENABLE_P68_AUTO_FORCE_TOOL=1
- GENESIS_ENABLE_P69_LONG_CTX_TOOL_REMINDER=1
- GENESIS_P68_P69_LONG_CTX_THRESHOLD_CHARS=50000
- GENESIS_ENABLE_P72_PROFILE_RUN_CAP=1
- GENESIS_PROFILE_RUN_CAP_M=4128
- GENESIS_ENABLE_P74_CHUNK_CLAMP=1
- GENESIS_ENABLE_P83=1
# P85 dropped 2026-05-03 (requires P84 which we don't enable in v7.69)
# P87 (marlin pad-sub-tile-n text-patch) disabled because the same fix is
# already vendored at ../../patches/vllm-marlin-pad/marlin.py and RO-mounted
# over the target file (lines 53-54). Letting Genesis re-do the patch fails
# with [Errno 30] read-only filesystem and `set -e` propagates exit-1 from
# `apply_all` before `vllm serve` runs (club-3090 #49).
- GENESIS_ENABLE_P87=0
- GENESIS_ENABLE_P91=1
- GENESIS_ENABLE_P94=1
- GENESIS_ENABLE_P98=1
# PN34: active env-opt-in workspace-lock relaxation. P98 above auto-skips
# on v0.20 (UNIFORM_SINGLE_TOKEN_DECODE drift-marker false-positive — see
# docs/UPSTREAM.md), so PN34 is what's actually firing today. Belt+suspenders
# pattern matches default + long-text composes. Propagated from #82 audit.
- GENESIS_ENABLE_PN34_WORKSPACE_LOCK_RELAX=1
- GENESIS_ENABLE_P99=1
- GENESIS_ENABLE_P100=1
- GENESIS_ENABLE_P101=1
- GENESIS_ENABLE_P103=1
- GENESIS_ENABLE_PN8_MTP_DRAFT_ONLINE_QUANT=1
- GENESIS_ENABLE_PN9_INDEPENDENT_DRAFTER_ATTN=1
- GENESIS_ENABLE_PN11_GDN_AB_CONTIGUOUS=1
- GENESIS_ENABLE_PN12_FFN_INTERMEDIATE_POOL=1
- GENESIS_ENABLE_PN13_CUDA_GRAPH_LAMBDA_ARITY=1
- GENESIS_ENABLE_PN14_TQ_DECODE_OOB_CLAMP=1
- GENESIS_ENABLE_PN17_FA2_LSE_CLAMP=1
- GENESIS_ENABLE_PN19_SCOPED_MAX_SPLIT=1
- GENESIS_ENABLE_PN22_LOCAL_ARGMAX_TP=1
- GENESIS_ENABLE_PN26_SPARSE_V=1
- GENESIS_ENABLE_PN59_STREAMING_GDN=1
- GENESIS_PN26_SPARSE_V_BLOCK_KV=8
- GENESIS_PN26_SPARSE_V_NUM_WARPS=4
- GENESIS_PN26_SPARSE_V_THRESHOLD=0.01
- GENESIS_ENABLE_P38B_COMPILE_SAFE=1
- GENESIS_ENABLE_P15B_FA_VARLEN_CLAMP=1
- GENESIS_ENABLE_PN25_SILU_INDUCTOR_SAFE=1
# PN30 — RE-ENABLED with our local dst-shaped temp fix
# (patch_pn30_dst_shaped_temp_fix.py, applied during setup.sh).
- GENESIS_ENABLE_PN30_DS_LAYOUT_SPEC_DECODE=1
- GENESIS_PREALLOC_TOKEN_BUDGET=4128
- GENESIS_BUFFER_MODE=shared
# P40 — TQ k8v4 GQA grouping kernel (+15-30% on compute-regime GPUs, L2>=24MB).
# Default off (RTX 3090: 6MB L2 = no gain). Enable on RTX 5090/A100/H100.
- GENESIS_ENABLE_P40=${GENESIS_ENABLE_P40:-0}
# PN54 — GDN contiguous-call deduplication (Cliff 2b OOM mitigation).
- GENESIS_ENABLE_PN54=${GENESIS_ENABLE_PN54:-0}
# Explicit OFFs to match Sandermage's PROD env-var set verbatim:
# P78 (P78_TOLIST_CAPTURE_GUARD) — superseded by our patch_tolist_cudagraph.py
# P81 (FP8 block-scaled M<=8) — FP8-specific, no-op on our TQ3 path
# P82 — biased on small-batch single-stream Lorbus INT4 + MTP K=3 (Sander PROD)
- GENESIS_ENABLE_P78_TOLIST_CAPTURE_GUARD=0
- GENESIS_ENABLE_P81_FP8_BLOCK_SCALED_M_LE_8=0
- GENESIS_ENABLE_P82=${GENESIS_ENABLE_P82:-0}
- GENESIS_P82_THRESHOLD_SINGLE=0.3
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
set -e
pip install xxhash pandas scipy -q
# Pre-flight: Genesis patches must be populated. Empty volume mount
# = silent no-op apply_all = boot fails later with cryptic upstream
# error (e.g. "TurboQuant KV not supported for hybrid models", #13).
if [ ! -f /usr/local/lib/python3.12/dist-packages/vllm/_genesis/patches/apply_all.py ]; then
echo "ERROR: Genesis patches missing — host volume models/qwen3.6-27b/vllm/patches/genesis/ is empty." >&2
echo " Run from repo root before 'docker compose up':" >&2
echo " bash scripts/setup.sh qwen3.6-27b" >&2
exit 1
fi
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
# to chat_completion/serving.py without hitting RO-mount errors.
bash /etc/club3090/install-pr35936.sh
python3 -m vllm._genesis.patches.apply_all
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
# Drops out when vllm-project/vllm lands the upstream fix.
python3 /patches/qwen3coder_tool_parser_deferred_commit.py
# were previously invoked here; superseded by Genesis natives in
# v7.72.2 (P78 + PN34). Mounts and invocations dropped 2026-05-05.
# VLLM_ENFORCE_EAGER=1 in compose/.env disables CUDA graphs — use on
# hardware where Cliff 2 GDN activation spikes occur at runtime.
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --model
- /root/.cache/huggingface/qwen3.6-27b-autoround-int4
- --served-model-name
- qwen3.6-27b-autoround
- --quantization
- auto_round
- --dtype
- float16
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual-turbo.yml disables it because
# PCIe P2P bandwidth makes the NCCL fallback faster there.
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.85}"
- --max-num-seqs
- "4"
- --max-num-batched-tokens
- "4128"
# TQ3 is the right pick on 24 GB / 3090 (smaller KV pool → more concurrency).
# On 20 GB Ampere (modded 3080 / cap'd 3090) override to fp8_e5m2 — TQ3's
# activation peak during DeltaNet GDN forward exceeds the per-card budget
# after TP=2 split and Cliff 2 fires at 90K. fp8_e5m2 trades KV-pool
# capacity for activation headroom on the smaller-VRAM sub-class. See
# docs/HARDWARE.md "Note for sub-24 GB cards" + #47 for cross-rig data.
- --kv-cache-dtype
- "${KV_CACHE_DTYPE:-turboquant_3bit_nc}"
- --trust-remote-code
# froggeric chat-template override (see volumes block + docs/UPSTREAM.md)
- --chat-template
- /etc/qwen-froggeric-chat-template.jinja
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
- --speculative-config
- '{"method":"mtp","num_speculative_tokens":3}'
- --host
- 0.0.0.0
- --port
- "8000"
- NVLINK_MODE=force_on
- PORT=8017
+8 -178
View File
@@ -1,182 +1,12 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 + NVLink bridge (TP=2, NCCL P2P over NVLink)
# Drafter: MTP n=3 (built-in)
# KV: fp8_e5m2
# Vision: yes
# Max ctx: 262K
# Genesis: v7.72.2
# Status: ✅ Production
# Best for: Users with the NVLink bridge — +15% narr / +15% code over PCIe-only dual.yml
# ---------------------------------------------------------------------------
# Dual RTX 3090 with NVLink — opt-in variant for users with the bridge installed.
# Mirrors dual/docker-compose.yml but enables NCCL P2P over NVLink for faster
# allreduce, and re-enables vLLM's custom all-reduce kernel (which `dual.yml`
# disables for PCIe-only stacks).
#
# Status: COMMUNITY-CONTRIBUTED, EXPERIMENTAL.
# First landed via PR #31 (JusefPol). Initial reports show stable boot +
# higher TPS than `dual.yml` on NVLink-bridged 2× 3090 rigs, but the maintainer
# doesn't have NVLink hardware to validate, so this hasn't been canonical-benched.
# If you run this, please drop numbers in discussion #19 — paste-ready report
# via `bash scripts/report.sh`.
#
# What this gives you (vs the PCIe-only `dual.yml`):
# - NCCL P2P over NVLink (NCCL_P2P_LEVEL=NVL) — much faster allreduce on TP=2
# - Custom all-reduce ENABLED (--disable-custom-all-reduce removed) — NVLink
# makes vLLM's custom kernel a win where PCIe makes it a loss
# - PYTORCH_CUDA_ALLOC_CONF without `expandable_segments:True` — JusefPol
# reports it crashes on startup with NVLink wired up
#
# What's intentionally NOT enabled:
# - TurboQuant KV — same rationale as dual.yml: fp8_e5m2 is plenty for 262K.
# For 4-stream concurrency at 262K, switch to dual/turbo.yml.
#
# Dependencies:
# - 2× RTX 3090 (Ampere SM 8.6) WITH NVLink bridge installed and `nvidia-smi
# topo -m` showing `NV*` between GPU0 and GPU1
# - vLLM PR #40361 (Marlin pad-sub-tile-n) — patched files vendored in-repo
# at ../../patches/vllm-marlin-pad/. PR is open upstream; drop the mount when
# it lands. See ../../patches/vllm-marlin-pad/README.md.
#
# All dual-card variants in this dir:
#
# File Ctx Streams Narr/Code TPS KV Vision NVLink
# dual/docker-compose.yml (DEFAULT) 262K 2 69 / 89 fp8 ✅ not used
# multi4/docker-compose.yml 262K 4 63 / 76 fp8 ✅ not used (4× PCIe)
# multi4/dflash.yml 262K 2 64 / 104 FP16 ✅ not used (4× PCIe)
# dual/nvlink.yml 262K 2 (community) fp8 ✅ required
# dual/turbo.yml 262K 4 54 / 73 TQ3 ✅ not used
# dual/dflash.yml 185K 1 82 / 125 FP16 ✅ not used
# dual/dflash-noviz.yml 200K 1 78 / 127 FP16 ❌ not used
# dual/nvlink-dflash.yml 185K 1 (community) FP16 ✅ required
# dual/nvlink-dflash-noviz.yml 188K 1 (community) FP16 ❌ required
#
# Run:
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# DEPRECATED (2026-05-14): use dual/docker-compose.yml — NVLink is auto-detected
# in the compose entrypoint. This stub forces NVLink mode for backward compatibility.
# Will be removed in a future release.
services:
vllm-qwen36-27b-dual-nvlink:
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
# TurboQuant), so it doesn't accumulate the same anchor-drift exposure
# the single-card project does. Pin if you hit a regression; otherwise
# ride the wave.
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
extends:
file: docker-compose.yml
service: vllm-qwen36-27b-dual
container_name: vllm-qwen36-27b-dual-nvlink
restart: "no"
ports:
- "${BIND_HOST:-0.0.0.0}:${PORT:-8014}:8000"
volumes:
- ${MODEL_DIR:-../../../../../models-cache}:/root/.cache/huggingface
# torch.compile + Triton kernel caches — first boot warms (~60-90 sec);
# subsequent boots reuse cached graphs. Pattern from Sander's PROD launch.
# Closes club-3090 #22.
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton:/root/.triton/cache
# Marlin pad-sub-tile-n patch (vLLM PR #40361) — required for TP=2
# on AutoRound W4A16 models where out-dim shards fall below 64.
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
# which write to chat_completion/serving.py at vllm-import time. See
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
# think before tool call, no-user-query crash, developer role, etc.).
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
# dual.yml does for PCIe-only stacks) and pin the P2P level to NVLink
# so NCCL doesn't fall back to PCIe paths if topology query is fuzzy.
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_LEVEL=NVL
- VLLM_NO_USAGE_STATS=1
- VLLM_USE_FLASHINFER_SAMPLER=1
- OMP_NUM_THREADS=1
# JusefPol report (PR #31): expandable_segments=True crashes on startup
# with NVLink wired in. Keep max_split_size_mb cap, drop the rest.
- PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
entrypoint:
- bash
- -c
- |
# VLLM_ENFORCE_EAGER=1 in compose/.env disables CUDA graphs — use on
# hardware where Cliff 2 GDN activation spikes occur at runtime
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
bash /etc/club3090/install-pr35936.sh
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --model
- /root/.cache/huggingface/qwen3.6-27b-autoround-int4
- --served-model-name
- qwen3.6-27b-autoround
- --quantization
- auto_round
- --dtype
- float16
- --tensor-parallel-size
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual.yml disables it because PCIe
# P2P bandwidth makes the NCCL fallback faster there.
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.92}"
- --max-num-seqs
- "2"
- --max-num-batched-tokens
- "8192"
- --kv-cache-dtype
- "${KV_CACHE_DTYPE:-fp8_e5m2}"
- --trust-remote-code
# froggeric chat-template override (see volumes block + docs/UPSTREAM.md)
- --chat-template
- /etc/qwen-froggeric-chat-template.jinja
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
- --speculative-config
- '{"method":"mtp","num_speculative_tokens":3}'
- --host
- 0.0.0.0
- --port
- "8000"
- NVLINK_MODE=force_on
- PORT=8014
+11 -3
View File
@@ -1,7 +1,7 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Lorbus AutoRound INT4 + BF16 mtp.fc preserved)
# Topology: Dual 3090 PCIe (TP=2, no NVLink)
# Topology: Dual 3090 (TP=2, NVLink auto-detected via NVLINK_MODE)
# Drafter: MTP n=3 (built-in)
# KV: turboquant_3bit_nc (TQ3, 0.375 bytes/token)
# Vision: yes
@@ -90,10 +90,13 @@ services:
# See docs/UPSTREAM.md "Community templates / model assets" + the row at
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
# NVLink auto-detection — runs inside container at boot.
- ../../../../../scripts/detect_nvlink.sh:/etc/club3090/detect_nvlink.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NVLINK_MODE=${NVLINK_MODE:-auto}
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
@@ -227,7 +230,13 @@ services:
# v7.72.2 (P78 + PN34). Mounts and invocations dropped 2026-05-05.
# VLLM_ENFORCE_EAGER=1 in compose/.env disables CUDA graphs — use on
# hardware where Cliff 2 GDN activation spikes occur at runtime.
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
# NVLink auto-detection (sets NCCL env vars, _NVLINK_ENABLED).
source /etc/club3090/detect_nvlink.sh
if [ "${_NVLINK_ENABLED:-0}" = "1" ]; then
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
else
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} --disable-custom-all-reduce "$@"
fi
- --
command:
- --model
@@ -242,7 +251,6 @@ services:
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
- --gpu-memory-utilization
+53
View File
@@ -0,0 +1,53 @@
#!/bin/bash
# NVLink auto-detection + override. Sources NVLINK_MODE from env (default: auto).
# Exports: _NVLINK_ENABLED (0 or 1), sets NCCL/PYTORCH env vars accordingly.
# Designed for dual-card (2x GPU) setups. Skips detection on >2 GPUs.
NVLINK_MODE="${NVLINK_MODE:-auto}"
case "$NVLINK_MODE" in
force_on)
_NVLINK_ENABLED=1
echo "[nvlink] NVLINK_MODE=force_on — enabling NVLink mode"
;;
force_off)
_NVLINK_ENABLED=0
echo "[nvlink] NVLINK_MODE=force_off — forcing PCIe mode"
;;
auto)
GPU_COUNT=$(nvidia-smi -L 2>/dev/null | grep -c 'GPU' || echo 0)
if [ "$GPU_COUNT" -gt 2 ]; then
_NVLINK_ENABLED=0
echo "[nvlink] $GPU_COUNT GPUs detected — skipping NVLink detection (dual-card only)"
elif [ "$GPU_COUNT" -eq 2 ]; then
LINK=$(nvidia-smi topo -m 2>/dev/null | awk '/^GPU0/{print $3}')
if [[ "$LINK" =~ ^NV[0-9]+$ ]]; then
_NVLINK_ENABLED=1
echo "[nvlink] detected NVLink ($LINK) between GPU0-GPU1 — enabling NVLink mode"
else
_NVLINK_ENABLED=0
echo "[nvlink] PCIe topology ($LINK) — using PCIe mode"
fi
else
_NVLINK_ENABLED=0
echo "[nvlink] $GPU_COUNT GPU(s) — skipping NVLink detection"
fi
;;
*)
echo "[nvlink] ERROR: invalid NVLINK_MODE=$NVLINK_MODE (must be auto|force_on|force_off)" >&2
exit 1
;;
esac
# Apply environment overrides based on detection result
if [ "$_NVLINK_ENABLED" -eq 1 ]; then
export NCCL_P2P_LEVEL=NVL
unset NCCL_P2P_DISABLE 2>/dev/null || true
export PYTORCH_CUDA_ALLOC_CONF="${PYTORCH_CUDA_ALLOC_CONF:-max_split_size_mb:512}"
echo "[nvlink] NVLink ENABLED — NCCL_P2P_LEVEL=NVL, custom all-reduce ON, expandable_segments OFF"
else
export NCCL_P2P_DISABLE=1
unset NCCL_P2P_LEVEL 2>/dev/null || true
export PYTORCH_CUDA_ALLOC_CONF="${PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True,max_split_size_mb:512}"
echo "[nvlink] NVLink DISABLED — NCCL_P2P_DISABLE=1, custom all-reduce OFF, expandable_segments ON"
fi
+6 -7
View File
@@ -23,6 +23,12 @@
# bash scripts/launch.sh --variant vllm/default
# bash scripts/launch.sh --variant llamacpp/default
# bash scripts/launch.sh --variant vllm/dual
#
# Env vars:
# NVLINK_MODE=auto|force_on|force_off — NVLink auto-detection for dual-card composes
# auto (default): detects NVLink via nvidia-smi topo -m
# force_on: assume NVLink bridge present, set NVLink env vars
# force_off: force PCIe-only path even if NVLink detected
set -euo pipefail
@@ -706,12 +712,6 @@ kv_projection() {
echo "[launch] KV projection only available for vLLM variants today." >&2
return 0
fi
if (( TP_VALUE > 4 )); then
echo "[launch] KV projection skipped: tools/kv-calc.py currently models TP up to 4." >&2
echo "[launch] Proceeding with launch-side head-divisibility validation only." >&2
return 0
fi
local kv_model="${mapping%%:*}" kv_compose="${mapping#*:}" kv_json status
if kv_json="$("${ROOT_DIR}/tools/kv-calc.py" --model "$kv_model" --compose "$kv_compose" --vram "$MIN_VRAM_GB" --tp "$TP_VALUE" --json 2>&1)"; then
status=0
@@ -783,7 +783,6 @@ if [[ -z "$VARIANT" ]]; then
VARIANT="${CANDIDATE_VARIANTS[0]}"
fi
echo "[launch] model: $(model_label "$MODEL_NAME")" >&2
echo "[launch] selected variant: ${VARIANT}" >&2
if (( HET_VRAM_MIXED == 1 && TP_VALUE > 1 )); then
echo "[launch] Note: heterogeneous TP is bottlenecked by the smallest selected card (${MIN_VRAM_GB} GB)." >&2
fi
+5 -1
View File
@@ -341,9 +341,13 @@ section "Container runtime"
section "Stack version"
{
if [[ -d .git ]]; then
# Prefer `git describe` for a human-readable version (e.g. v0.6.2-3-ge299e70,
# "3 commits past v0.6.2 at SHA e299e70"). Falls back to raw SHA if no tags
# are reachable (shallow clone, fresh repo).
version=$(git describe --tags --always --dirty 2>/dev/null)
commit=$(git rev-parse --short HEAD 2>/dev/null)
branch=$(git branch --show-current 2>/dev/null)
echo "- **club-3090:** \`${commit:-unknown}\` (branch: \`${branch:-detached}\`)"
echo "- **club-3090:** \`${version:-${commit:-unknown}}\` (branch: \`${branch:-detached}\`, SHA \`${commit:-unknown}\`)"
if ! git diff --quiet 2>/dev/null || ! git diff --cached --quiet 2>/dev/null; then
echo "- **Working tree:** ⚠ has uncommitted changes (run \`git status\` to inspect)"
fi
+4 -4
View File
@@ -31,10 +31,10 @@
# vllm/dual-turbo 262K + TQ3 + 4 streams + vision (multi-tenant)
# vllm/dual-dflash 185K + FP16 + DFlash N=5 + vision (peak code TPS)
# vllm/dual-dflash-noviz 200K + FP16 + DFlash N=5 + no vision (peak code, max ctx)
# vllm/dual-nvlink 262K + fp8 + 2 streams + vision (REQUIRES NVLink bridge — community/experimental)
# vllm/dual-nvlink-turbo 262K + TQ3 + 4 streams + vision (REQUIRES NVLink bridge — community/experimental)
# vllm/dual-nvlink-dflash 185K + FP16 + DFlash N=5 + vision (REQUIRES NVLink bridge — community/experimental)
# vllm/dual-nvlink-dflash-noviz 188K + FP16 + DFlash N=5 + no vision (REQUIRES NVLink bridge — community/experimental)
# vllm/dual-nvlink 262K + fp8 + 2 streams + vision (NVLink stub — auto-detected via dual/)
# vllm/dual-nvlink-turbo 262K + TQ3 + 4 streams + vision (NVLink stub — auto-detected via dual/)
# vllm/dual-nvlink-dflash 185K + FP16 + DFlash N=5 + vision (NVLink stub — auto-detected via dual/)
# vllm/dual-nvlink-dflash-noviz 188K + FP16 + DFlash N=5 + no vision (NVLink stub — auto-detected via dual/)
# vllm/gemma-mtp Gemma-4-31B + Google MTP drafter (32K, bf16 KV, vision — community/experimental, pre-merge)
#
# Single-card llama.cpp:
+27
View File
@@ -131,6 +131,7 @@ assert_contains "$out" "[model] SKIP_MODEL=1"
# skip prompts, select the expected variant, and export GPU / TP / PP envs.
mkdir -p "${TMP_DIR}/models/qwen3.6-27b-autoround-int4" \
"${TMP_DIR}/models/gemma-4-31b-autoround-int4"
FAKE_8X3090='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6,4:RTX_3090:24576:8.6,5:RTX_3090:24576:8.6,6:RTX_3090:24576:8.6,7:RTX_3090:24576:8.6'
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
@@ -143,6 +144,12 @@ out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1 --no-projection 2>&1)"
assert_contains "$out" "[launch] Tensor parallel TP=2"
assert_contains "$out" "SWITCHED vllm/dual CUDA=0,1 NVD=0,1 TP=2 PP=1"
selected_count="$(grep -c "\[launch\] selected variant:" <<< "$out" || true)"
if [[ "$selected_count" != "1" ]]; then
echo "ASSERTION FAILED: expected one selected-variant line, got ${selected_count}" >&2
echo "$out" >&2
exit 1
fi
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
@@ -162,6 +169,26 @@ if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6
fi
assert_contains "$out" "Valid TP values: 1 2 4"
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS="${FAKE_8X3090}" \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model gemma-4-31b --gpus 0,1,2,3,4,5,6,7 --tp 8 2>&1)"
assert_contains "$out" "[launch] Tensor parallel TP=8"
assert_contains "$out" "[launch] Suggested: vllm/gemma-mtp"
assert_contains "$out" "VRAM budget — per card"
assert_contains "$out" "Note: TP > 4 predictions are extrapolated"
assert_not_contains "$out" "KV projection skipped"
assert_contains "$out" "SWITCHED vllm/gemma-mtp CUDA=0,1,2,3,4,5,6,7 NVD=0,1,2,3,4,5,6,7 TP=8 PP=1"
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS="${FAKE_8X3090}" \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1,2,3,4,5,6,7 --tp 8 --no-projection 2>&1)"; then
echo "ASSERTION FAILED: invalid Qwen TP=8 unexpectedly succeeded" >&2
echo "$out" >&2
exit 1
fi
assert_contains "$out" "num_kv_heads does not divide TP=8"
assert_contains "$out" "Valid TP values: 1 2 4"
# TTY-backed no-arg setup supports the cosmetic but real "Both" choice by
# dispatching through the positional path for both model families.
if ! command -v script >/dev/null 2>&1; then
+30 -5
View File
@@ -1,5 +1,10 @@
#!/usr/bin/env python3
"""kv-calc.py — predict per-card VRAM budget for vLLM composes.
#!/bin/sh
''':'
exec python3 "$0" "$@"
':'''
from __future__ import annotations
__doc__ = """kv-calc.py — predict per-card VRAM budget for vLLM composes.
Predicts (per card, after TP split):
- Model weights
@@ -37,8 +42,6 @@ Usage:
bash tools/kv-calc.py --calibration # both models, grouped per-model
"""
from __future__ import annotations
import argparse
import json
import sys
@@ -60,6 +63,7 @@ QWEN36_27B = {
"num_attn_layers": 16, # full_attention layers
"num_attn_heads": 24,
"num_kv_heads": 4, # GQA
"valid_tp": [1, 2, 4],
"head_dim_attn": 256, # attention head dim
"linear_num_v_heads": 48, # GDN value heads
"linear_num_k_heads": 16, # GDN key heads (GQA-style at the GDN level too)
@@ -86,6 +90,7 @@ GEMMA4_31B = {
"num_sliding_attn_layers": 50, # sliding_attention (fixed window, head_dim=256)
"num_attn_heads": 32,
"num_kv_heads": 16, # GQA 2:1
"valid_tp": [1, 2, 4, 8, 16],
"head_dim_sliding": 256, # sliding_attention head dim
"global_head_dim": 512, # full_attention head dim (asymmetric)
"sliding_window": 1024,
@@ -357,6 +362,16 @@ def cudagraph_overhead_gb(mem_util, tp):
return base + tp_bump
def _validate_tp_for_spec(spec, tp):
valid_tp = spec.get("valid_tp")
if valid_tp and tp not in valid_tp:
raise ValueError(
f"TP={tp} invalid for {spec['model_id']} "
f"(num_kv_heads={spec['num_kv_heads']} cannot be divided across TP cleanly). "
f"Valid TP values: {valid_tp}"
)
def predict(
spec=QWEN36_27B,
kv_format="fp8_e5m2",
@@ -380,6 +395,8 @@ def predict(
drafter_gb: total drafter weight (MTP / DFlash) — split by TP.
dflash_draft_gb: legacy alias — folded into drafter_gb if set.
"""
_validate_tp_for_spec(spec, tp)
weights_gb = _weights_per_card_gb(spec, tp, weights_variant)
growing_b, sliding_b = kv_pool_per_card_bytes(
@@ -441,6 +458,8 @@ def predict(
notes.append("⚠ fp8_e4m3 on Ampere (sm_86): Triton `fp8e4nv` kernel unsupported; use int8_per_token_head instead (PR #40391 via #42102)")
if spec["model_family"] == "gemma4-swa-dense" and tp == 1 and vram_gb < 32:
notes.append("⚠ Gemma 4 31B TP=1 needs ≥32 GB VRAM; 24 GB Ampere boot-OOMs (model weights + drafter + min KV)")
if tp > 4:
notes.append("TP > 4 predictions are extrapolated; report deltas via scripts/report.sh --bench")
return Prediction(
model=spec["model_id"],
@@ -636,7 +655,7 @@ def main():
help="KV cache format. Default: from --compose, or fp8_e5m2.")
p.add_argument("--max-ctx", type=int, help="max_model_len. Default: from --compose, or 180000.")
p.add_argument("--max-num-seqs", type=int, help="max_num_seqs. Default: from --compose, or 1.")
p.add_argument("--tp", type=int, choices=[1, 2, 4], help="tensor_parallel_size. Default: from --compose, or 1.")
p.add_argument("--tp", type=int, choices=[1, 2, 4, 8, 16], help="tensor_parallel_size. Default: from --compose, or 1.")
p.add_argument("--mem-util", type=float, help="gpu_memory_utilization. Default: from --compose, or 0.95.")
p.add_argument("--vram", type=float, default=24, help="VRAM per card in GB. Default 24.")
p.add_argument("--mtp", action="store_true", default=None, help="MTP enabled (Qwen: n=3 built-in; Gemma: external drafter).")
@@ -690,6 +709,12 @@ def main():
weights_variant = args.weights_variant or "default"
header = f"Predicted budget — {model_key} custom config on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx:,}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
try:
_validate_tp_for_spec(spec, tp)
except ValueError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
return 2
if args.solve_max_ctx:
best = solve_max_ctx(
spec, kv_format=kv_format, max_num_seqs=max_num_seqs,