Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
98f0406d0f | ||
|
|
2a92a199ff | ||
|
|
056dcb6439 | ||
|
|
e299e70451 | ||
|
|
5882bbef6f | ||
|
|
0d48dac818 | ||
|
|
a119e1b1b6 | ||
|
|
e05f1969bc | ||
|
|
12a33fbdd1 | ||
|
|
45fa421e93 | ||
|
|
22bf2e9398 | ||
|
|
99a0b66224 |
@@ -1,6 +1,10 @@
|
||||
## Rig bench submission
|
||||
|
||||
<!-- This PR was auto-generated by `bash scripts/submit-bench.sh --auto-submit --tag <TAG>`. -->
|
||||
> ⚠ Most bench submissions go through an issue (see `CONTRIBUTING.md` "Submitting your bench").
|
||||
> This PR template is for contributors who explicitly chose the direct-PR path.
|
||||
> The maintainer may redirect to an issue thread before merge.
|
||||
|
||||
<!-- This PR was auto-generated by `bash scripts/submit-bench.sh --auto-submit --as-pr --tag <TAG>`. -->
|
||||
<!-- Review the row below; the PR reviewer may move it within the target section. -->
|
||||
|
||||
### New row
|
||||
|
||||
@@ -16,6 +16,62 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.6.1 — 2026-05-14
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(launch): add hardware-aware launcher ([5882bbe](https://github.com/noonghunna/club-3090/commit/5882bbef6f5ed7e8ceea450fa3a6b167a1bf4926))
|
||||
- feat(tools): extend kv-calc.py to multi-model (Qwen 3.6 + Gemma 4 31B) ([0d48dac](https://github.com/noonghunna/club-3090/commit/0d48dac818ba889f74b74df1c454bb961a123c37))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs: update launch.sh references for v0.6.1 wizard flow ([e299e70](https://github.com/noonghunna/club-3090/commit/e299e70451c8d146214a6560e582d0e174dd0ebc))
|
||||
|
||||
|
||||
### 🧹 Other
|
||||
|
||||
- Merge codex/v0.6.1-launch into master ([056dcb6](https://github.com/noonghunna/club-3090/commit/056dcb643914fee6169b02b89cb420b038c29b0f))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.6.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.6.0...v0.6.1)
|
||||
## v0.6.0 — 2026-05-13
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(scripts): add hardware-aware setup picker ([12a33fb](https://github.com/noonghunna/club-3090/commit/12a33fbdd1b6422429521f887d9f21c2e3da793d))
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(launch): exit cleanly on stdin EOF in wizard prompts ([e05f196](https://github.com/noonghunna/club-3090/commit/e05f1969bcf57547a17cd963a1f21435de80815b))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.6.0`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.4...v0.6.0)
|
||||
## v0.5.4 — 2026-05-13
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(scripts): make submit-bench issue-first ([22bf2e9](https://github.com/noonghunna/club-3090/commit/22bf2e9398c7907aae6b62809bde30e111e4a700))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.4`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.3...v0.5.4)
|
||||
## v0.5.3 — 2026-05-13
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(scripts): add submit-bench flow ([ef77032](https://github.com/noonghunna/club-3090/commit/ef770322f43724f612a80393f547e5da218b5bf7))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.3`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.2...v0.5.3)
|
||||
## v0.5.2 — 2026-05-13
|
||||
|
||||
|
||||
|
||||
+18
-6
@@ -8,7 +8,7 @@ Thanks for being here. This repo collects working recipes for serving big LLMs o
|
||||
|
||||
### ✅ Yes please
|
||||
|
||||
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
|
||||
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. If you're unsure what to run before measuring, start with `bash scripts/setup.sh` and `bash scripts/launch.sh`; both wizards mark hardware-fit. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
|
||||
- **Power-cap efficiency curves.** `sudo bash scripts/power-cap-sweep.sh --cooling air|water|aio --load-mode decode-concurrent --concurrency auto --bench-runs 3` produces cross-rig efficiency-knee data ([discussion #86](https://github.com/noonghunna/club-3090/discussions/86)). ~15-20 min for a 30-cap sweep on a 3090/4090/5090. Especially valuable on cards we don't have anchors for yet (A5000/A6000, 4080, 5060 Ti / 5080, modded variants). **Keep `--step-size 10` (the default).** Larger step-sizes (e.g. `--step-size 50`) are too coarse for the efficiency knee and only useful for quick smoke tests. See [docs/HARDWARE.md](docs/HARDWARE.md#cross-rig-power-cap-data-anchor-points) for the full canonical command and rationale.
|
||||
- **Bug reports with the data we ask for.** The [bug report template](https://github.com/noonghunna/club-3090/issues/new?template=bug-report.yml) leads with `bash scripts/report.sh > my-rig.md` (add `--verify` to include verify-full output, `--soak` to also run SOAK_MODE=continuous if you suspect a multi-turn agent cliff) — single command captures the rig state we'd otherwise ask for individually (hardware, container state, Genesis patches, KV pool sizing, engine config). With that paste, the first reply is usually a fix or a clear next step instead of "can you send me…".
|
||||
- **Bug reproductions / minimum repros for upstream issues.** vLLM / llama.cpp / Genesis bugs that affect this stack are most useful when they have a one-paragraph reduction. Drop them in an issue or open a draft PR adding a reproducer to `verify-stress.sh`.
|
||||
@@ -47,23 +47,35 @@ Two GitHub channels, two different shapes of conversation. Picking the right one
|
||||
|
||||
## Submitting your bench
|
||||
|
||||
After running `bash scripts/rebench-full.sh`, contribute your numbers to `BENCHMARKS.md` with:
|
||||
The matrix is hand-curated — the canonical path is to file an **issue** with your rig + numbers; we'll review, ask clarifying questions, and integrate.
|
||||
|
||||
After running `bash scripts/rebench-full.sh`, generate a paste-ready row:
|
||||
|
||||
```bash
|
||||
bash scripts/submit-bench.sh --tag <your-tag>
|
||||
```
|
||||
|
||||
This generates a paste-ready row at `results/rebench/<tag>/BENCHMARKS-row.md`. Review it, then either auto-submit or open the PR manually.
|
||||
The script writes `results/rebench/<tag>/BENCHMARKS-row.md`. To submit:
|
||||
|
||||
Auto-submit opens a PR via GitHub CLI:
|
||||
### Path A — Auto-issue (recommended, requires `gh auth login`)
|
||||
|
||||
```bash
|
||||
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit
|
||||
```
|
||||
|
||||
Manual path: paste the row into the appropriate `BENCHMARKS.md` section and open a PR yourself.
|
||||
Opens an issue via `gh issue create` with your rig + row pre-filled.
|
||||
|
||||
Auto-submit assumes `gh auth status` is configured. If it is not, run `gh auth login` first.
|
||||
### Path B — Manual issue (no tools beyond browser)
|
||||
|
||||
Open https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml and paste the row + your `rig.txt` into the body.
|
||||
|
||||
### Path C — Direct PR (advanced)
|
||||
|
||||
```bash
|
||||
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit --as-pr
|
||||
```
|
||||
|
||||
For contributors who know the `BENCHMARKS.md` section structure and want to propose the exact row. The maintainer may still redirect to an issue thread for context-gathering before merge — direct PRs aren't a fast-path bypass.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -66,18 +66,23 @@ Bench protocol: 3 warm + 5 measured runs of the canonical narrative + code promp
|
||||
git clone https://github.com/noonghunna/club-3090.git
|
||||
cd club-3090
|
||||
|
||||
# 2. Download + SHA-verify the model (~20 GB; clones Genesis patches too)
|
||||
# (asks you where to put model weights — pick in-repo default, ~/models, or
|
||||
# a custom path on a different drive. To skip the prompt:
|
||||
# `export MODEL_DIR=/mnt/your-drive/models` before running. See FAQ + .env.example.)
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
# 2. Pick/download + SHA-verify the model (interactive hardware-aware picker)
|
||||
# (asks you which model, then where to put model weights — pick in-repo
|
||||
# default, ~/models, or a custom path on a different drive. To skip prompts:
|
||||
# `export MODEL_DIR=/mnt/your-drive/models` and pass the model name. See FAQ.)
|
||||
bash scripts/setup.sh
|
||||
# Or scripted:
|
||||
# bash scripts/setup.sh qwen3.6-27b
|
||||
|
||||
# 3. Pick a config + boot it (interactive wizard — asks engine / cards / workload)
|
||||
# 3. Pick a config + boot it (interactive wizard: asks model → GPUs → projects VRAM budget)
|
||||
bash scripts/launch.sh
|
||||
# Or skip the wizard:
|
||||
# bash scripts/launch.sh --variant vllm/default # single-card chat (recommended)
|
||||
# bash scripts/launch.sh --variant vllm/dual # dual-card 262K + vision
|
||||
# bash scripts/launch.sh --variant llamacpp/default # single-card 262K, no cliffs
|
||||
# Or partial flags (wizard fills the rest):
|
||||
# bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1
|
||||
# bash scripts/launch.sh --tp 2 --pp 1 # override vLLM parallelism
|
||||
# See all variants:
|
||||
# bash scripts/switch.sh --list
|
||||
|
||||
@@ -152,7 +157,7 @@ club-3090/
|
||||
│ └── README.md blocked status — what would unblock it on this model
|
||||
├── scripts/ shared, model-aware
|
||||
│ ├── setup.sh bash setup.sh <model> → preflight + downloads + verifies + Genesis
|
||||
│ ├── launch.sh interactive wizard: cards → workload → boots compose + verifies
|
||||
│ ├── launch.sh interactive wizard: model → GPUs → KV projection → boots compose + verifies
|
||||
│ ├── switch.sh stateless variant switcher (bring down old, up new)
|
||||
│ ├── update.sh one-shot upgrade: git pull + re-pin Genesis + re-vendor patches
|
||||
│ ├── health.sh runtime health probe (KV %, MTP AL, recent TPS, errors)
|
||||
|
||||
+1
-1
@@ -163,7 +163,7 @@ If you're solo-using on dual, you're paying for hardware that mostly sits idle o
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
git clone https://github.com/vllm-project/vllm.git /opt/ai/engines/vllm/primary # required for dual variants
|
||||
|
||||
# 2. Pick + boot via wizard (asks GPU count + workload)
|
||||
# 2. Pick + boot via wizard (asks model + GPUs, projects VRAM budget, auto-picks TP=2 for matched 2× 3090)
|
||||
bash scripts/launch.sh
|
||||
|
||||
# 3. Or skip the wizard:
|
||||
|
||||
+8
-2
@@ -130,6 +130,12 @@ If your numbers on the same compose look different from ours by >15%, the most l
|
||||
|
||||
## Setup
|
||||
|
||||
### How do I pick the right model + variant?
|
||||
|
||||
For a first install, run `bash scripts/setup.sh` with no model argument in a normal terminal. It opens a hardware-aware model picker, marks Qwen / Gemma / Both as eligible or not for your detected GPUs, then continues into the existing download flow.
|
||||
|
||||
After setup, run `bash scripts/launch.sh`. The wizard asks which model (filtered to what you've downloaded), then which GPU(s) to use, auto-picks TP for homogeneous sets (PP for heterogeneous), filters variants by hardware fit, shows a per-card VRAM projection from `tools/kv-calc.py` for the suggested default, then boots and runs `verify-full.sh`. Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, partial flags like `bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1` (skips prompts), `--tp 4 --pp 2` to override parallelism, plus `setup.sh --help` / `launch.sh --help` for the full flag list.
|
||||
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
|
||||
|
||||
Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
|
||||
@@ -140,7 +146,7 @@ Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
```
|
||||
2. **`.env` file at repo root** — picked up automatically on every script run. See [`.env.example`](../.env.example).
|
||||
3. **Interactive prompt** — `bash scripts/setup.sh qwen3.6-27b` with nothing set offers three choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
|
||||
3. **Interactive prompt** — `bash scripts/setup.sh` with nothing set first asks which model to download, then offers three model-dir choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
|
||||
4. **Silent fallback** — `<repo>/models-cache/`. Functional but pollutes the git tree; not recommended.
|
||||
|
||||
Every script that touches model paths reads from the same `MODEL_DIR`. The compose YAMLs' volume mount is `${MODEL_DIR:-...}:/root/.cache/huggingface` — once set, every container reads + writes there.
|
||||
@@ -171,7 +177,7 @@ If your *Genesis tree* (not the repo) is out of sync — the pin in `setup.sh` m
|
||||
|
||||
### My GPU isn't card 0 — how do I change it?
|
||||
|
||||
`CUDA_VISIBLE_DEVICES=2 bash scripts/launch.sh --variant vllm/default` (substitute your card index). For dual-card, pass two: `CUDA_VISIBLE_DEVICES=2,3`. The compose files inherit env from your shell.
|
||||
Use the `--gpus` flag: `bash scripts/launch.sh --gpus 2` (single-card) or `bash scripts/launch.sh --gpus 2,3` (two cards). The wizard exports `CUDA_VISIBLE_DEVICES` for you. The older form `CUDA_VISIBLE_DEVICES=2 bash scripts/launch.sh --variant vllm/default` still works if you prefer to set the env yourself.
|
||||
|
||||
### Container fails to start: "Free memory ... is less than desired GPU memory utilization"
|
||||
|
||||
|
||||
+132
-17
@@ -1,23 +1,27 @@
|
||||
# KV Cache Math — predicting per-card VRAM budget on Qwen3.6-27B
|
||||
# KV Cache Math — predicting per-card VRAM budget
|
||||
|
||||
This page documents the math behind [`tools/kv-calc.py`](../tools/kv-calc.py) — the predictor that helps you decide whether a config will fit on your hardware *before* booting it. It also explains why predictions are estimates (±1.5 GB error band) rather than precise allocations.
|
||||
|
||||
Two models are modelled today: **Qwen 3.6 27B** (DeltaNet hybrid) and **Gemma 4 31B** (sliding-window + dense MLP). The math differs structurally between them; this doc covers each in its own section. The CLI dispatches on `--model`.
|
||||
|
||||
## TL;DR
|
||||
|
||||
```bash
|
||||
# What's my budget if I run dual-turbo on 20 GB cards?
|
||||
bash tools/kv-calc.py --compose dual-turbo --vram 20 --mem-util 0.82
|
||||
# Qwen 3.6 27B — what's my budget if I run dual-turbo on 20 GB cards?
|
||||
bash tools/kv-calc.py --model qwen3.6-27b --compose dual-turbo --vram 20 --mem-util 0.82
|
||||
|
||||
# What's the largest max_ctx that fits on 16 GB cards with TP=2 + fp8 KV?
|
||||
bash tools/kv-calc.py --solve-max-ctx --tp 2 --kv-format fp8_e5m2 --vram 16 --mem-util 0.95
|
||||
# Gemma 4 31B — what's the largest max_ctx that fits on 24 GB cards with TP=2 + INT8 PTH KV?
|
||||
bash tools/kv-calc.py --model gemma-4-31b --solve-max-ctx --tp 2 --kv-format int8_per_token_head --vram 24 --mem-util 0.92
|
||||
|
||||
# How accurate is the model? Show predicted vs measured for our shipped composes:
|
||||
# How accurate is the model? Show predicted vs measured for our shipped composes (both models):
|
||||
bash tools/kv-calc.py --calibration
|
||||
```
|
||||
|
||||
`--model` defaults to `qwen3.6-27b` for backward compatibility with earlier invocations.
|
||||
|
||||
The predictor is a directional estimator, not a precise allocator. The vLLM engine's `gpu_worker.py` boot-log report is authoritative — the calculator is for *before* boot.
|
||||
|
||||
## Per-card budget components
|
||||
## Qwen 3.6 27B — per-card budget components
|
||||
|
||||
For Qwen3.6-27B AutoRound INT4 at TP=N, the per-card VRAM peak during bench is composed of:
|
||||
|
||||
@@ -109,43 +113,154 @@ This is rough — actual overhead depends on how many graphs vLLM captures, whic
|
||||
|
||||
Only present on `dual-dflash*.yml` composes. `z-lab/Qwen3.6-27B-DFlash` is a ~1.75 GB draft model (per card, FP16). With TP > 1, the draft itself is sharded.
|
||||
|
||||
## Gemma 4 31B — per-card budget components
|
||||
|
||||
Gemma 4 31B is structurally different from Qwen 3.6:
|
||||
|
||||
- **No DeltaNet, no GDN activation peak.** Dense MLP instead.
|
||||
- **Hybrid layer pattern** — but the hybrid is on *attention type*, not attention-vs-recurrence. The 60-layer stack is `[sliding_attention × 5, full_attention × 1] × 10` = **50 sliding-attention layers + 10 full-attention layers**.
|
||||
- **Head-dim asymmetry** — sliding layers use `head_dim=256`, full layers use `global_head_dim=512`. Per-token KV bytes for the full layers is therefore double what naive `num_layers × head_dim` would compute.
|
||||
- **K==V tying** — `attention_k_eq_v: true` in `config.json`. vLLM's allocator EXPLOITS this — K and V share storage. The KV term uses `×1`, not `×2`. This was confirmed empirically by calibrating the per-token byte count against measured BENCHMARKS rows (the matched-config rebench's `Available KV cache / card = 10.82 GiB` at 262K seqs=2 is consistent with ×1 storage, not ×2).
|
||||
|
||||
Source: `/mnt/models/huggingface/gemma-4-31b-autoround-int4/config.json` → `text_config`.
|
||||
|
||||
```
|
||||
peak ≈ weights/N + kv_pool_growing + kv_pool_sliding + activation_peak + cudagraph_overhead + drafter_overhead
|
||||
```
|
||||
|
||||
### 1. Model weights (`weights / N`)
|
||||
|
||||
| Quant | On-disk | Per-card at TP=2 |
|
||||
|---|---:|---:|
|
||||
| AutoRound INT4 (`gemma-4-31b-autoround-int4`) | ~18 GB | 9.0 GB |
|
||||
| AWQ-4bit (`cyankiwi/gemma-4-31B-it-AWQ-4bit`) | ~17 GB usable on stack | 8.5 GB |
|
||||
| BF16 (unquantized) | ~58 GB | 29 GB (does not fit on 24 GB) |
|
||||
|
||||
The two shipped quants on this stack are AutoRound INT4 (default) and AWQ-4bit (Tier 2 reproducer of #103). INT4 weights + INT8-per-token-head KV is the matched-config dual-3090 recipe (see `models/gemma-4-31b/vllm/compose/dual/int8.yml`).
|
||||
|
||||
### 2. KV pool — growing portion (full-attention layers only)
|
||||
|
||||
Only the 10 full-attention layers grow KV with context. Each stores K and V at `global_head_dim=512`, with K==V tying meaning a single store per element:
|
||||
|
||||
```
|
||||
per_token_bytes_growing = num_full_attn_layers × num_kv_heads × global_head_dim × 1 × bpe
|
||||
= 10 × 16 × 512 × 1 × bpe
|
||||
= 81,920 × bpe bytes
|
||||
```
|
||||
|
||||
For comparison, Qwen 3.6's growing KV is `16 × 4 × 256 × 2 × bpe = 32,768 × bpe` — Gemma 4's per-token growing KV is **~2.5× heavier** than Qwen's. This is *the* reason Gemma 4 at 262K needs INT8 / FP8 KV on Ampere — at BF16 KV the per-card budget blows past 24 GB before you reach 50K context.
|
||||
|
||||
Per-token growing-KV bytes by format:
|
||||
|
||||
| KV format | `bytes_per_kv_element` | per-token growing KV (full) | per-token (TP=2) |
|
||||
|---|---:|---:|---:|
|
||||
| `fp16` / `bf16` | 2.0 | 163,840 B (~160 KB) | 81,920 B |
|
||||
| `fp8_e5m2` / `fp8_e4m3` | 1.0 | 81,920 B (~80 KB) | 40,960 B |
|
||||
| `int8_per_token_head` (PR #40391) | ~1.01 (incl. per-token scale) | ~82,700 B | ~41,400 B |
|
||||
| `q4_0` | ~0.56 | ~45,875 B | ~22,940 B |
|
||||
| `turboquant_3bit_nc` (TQ3) | ~0.425 | ~34,816 B | ~17,408 B |
|
||||
|
||||
Total growing-KV pool per card = `per_token_bytes_growing / TP × max_ctx × max_num_seqs`.
|
||||
|
||||
**Note**: on Ampere consumer cards (sm_86), `fp8_e4m3` is NOT supported by the Triton kernel (`fp8e4nv` requires Hopper/Ada/Blackwell). Use `int8_per_token_head` (PR #40391, vendored on this stack via PR #42102) instead. See `models/gemma-4-31b/vllm/compose/dual/int8.yml` header for the engineering trail.
|
||||
|
||||
### 3. KV pool — fixed sliding portion (50 layers)
|
||||
|
||||
The 50 sliding-attention layers maintain a fixed-size KV window (`sliding_window=1024`). K==V tying applies here too:
|
||||
|
||||
```
|
||||
sliding_kv_bytes_total = num_sliding_layers × num_kv_heads × head_dim × 1 × bpe × sliding_window
|
||||
= 50 × 16 × 256 × 1 × bpe × 1024
|
||||
= 209,715,200 × bpe bytes
|
||||
≈ 200 MB × bpe
|
||||
```
|
||||
|
||||
This is constant — it doesn't scale with `max_ctx` or `max_num_seqs`. At fp8 / int8 KV (`bpe=1`), this is ~200 MB per card (TP=1) or ~100 MB at TP=2. Small but non-zero — include it as a separate term.
|
||||
|
||||
### 4. Activation peak (SWA prefill + dense MLP)
|
||||
|
||||
Unlike Qwen 3.6's GDN block-wise state materialization, Gemma 4's activation peak comes from:
|
||||
|
||||
- Sliding-window attention prefill (50 layers, but bounded by `sliding_window=1024`).
|
||||
- Dense MLP intermediate buffer (`hidden_size=5376`, `intermediate_size=21504`).
|
||||
|
||||
There's no published scaling-law analogue to PerfMamba's O(γDNL) for Gemma 4. The activation coefficient is **empirical-only**, calibrated against measured BENCHMARKS.md rows. Expected order of magnitude: ~1.5-2.5 GB at TP=2 dual-card configs (smaller than Qwen 3.6's GDN peak because there's no per-chunk block materialization).
|
||||
|
||||
The coefficient may have weak dependence on KV format (slight dequant overhead during forward) but we expect it to be flatter than Qwen's TQ3 → fp8 25% spread — Gemma's dense MLP doesn't dequant KV during its forward.
|
||||
|
||||
### 5. Cudagraph + workspace overhead
|
||||
|
||||
Same form as Qwen — empirical fit:
|
||||
```
|
||||
overhead = 0.5 + 1.0 × mem_util + 0.3 × (TP - 1) # GB
|
||||
```
|
||||
|
||||
vLLM captures multiple cudagraphs (~50-100 MB each), FlashInfer workspace (~394 MB/card), NCCL allreduce buffers (~200-300 MB at TP > 1). Same accounting as Qwen.
|
||||
|
||||
### 6. Drafter overhead
|
||||
|
||||
Two drafter families on this stack:
|
||||
|
||||
| Drafter | Size | Composes |
|
||||
|---|---:|---|
|
||||
| `gemma-4-31b-it-assistant` (Google MTP) | 0.97 GB FP16 | `dual/docker-compose.yml`, `dual/int8.yml`, `dual/awq.yml` (with MTP n=4) |
|
||||
| `gemma-4-31b-it-dflash` (z-lab DFlash) | 2.9 GB FP16 | `dual/dflash.yml`, `dual/dflash-int8.yml` |
|
||||
|
||||
At TP > 1, drafter weights shard across cards (`drafter_gb / TP`).
|
||||
|
||||
## Known limitations
|
||||
|
||||
**The model is empirically calibrated, not first-principles.** Specifically:
|
||||
|
||||
1. **`max_num_seqs > 1` over-predicts**. My KV pool formula is `max_ctx × max_num_seqs × per_token_bytes` — but vLLM doesn't always allocate the full pool. The engine internally rate-limits based on `mem_util × VRAM - other_stuff`. For configs with `max_num_seqs=2-4`, the calculator may say FAIL when reality is TIGHT-PASS.
|
||||
1. **KV pool capping (resolved 2026-05-13)**. Earlier versions of this calculator over-predicted FAIL on configs with `max_num_seqs > 1` because the requested KV pool exceeded available budget. The current calculator explicitly models vLLM's PagedAttention capping: the predicted KV pool is `min(requested, budget - fixed_components)`. When the request exceeds available, the verdict is `TIGHT` with a note that effective concurrency at `--max-num-seqs` may be lower than requested at full `max_ctx`. The "predicted total" in TIGHT cases equals the budget exactly — that's the saturating-allocator behavior, not a modeling artifact.
|
||||
|
||||
2. **Activation coefficient varies by `chunk_size` and `dtype`**. We use the fla default `chunk_size=256` and `mamba_ssm_dtype=float32` (per Qwen3.6-27B config.json). If those change, the coefficient needs re-calibration.
|
||||
2. **Activation coefficient varies by `chunk_size` and `dtype`**. We use the fla default `chunk_size=256` and `mamba_ssm_dtype=float32` (per Qwen3.6-27B config.json). If those change, the coefficient needs re-calibration. For Gemma the activation peak is a flat empirical constant; if Gemma config changes (e.g. layer-pattern ratio, sliding_window), recalibrate.
|
||||
|
||||
3. **No driver/allocator overhead modeling**. snoby's 4090 needed `max-model-len` 200K → 180K vs 3090 baseline. The driver-class delta isn't modeled here. We hand-wave with the `±1.5 GB` error band.
|
||||
|
||||
4. **No Cliff 2b accumulation modeling**. The multi-turn fragmentation cliff at ~25K accumulated tokens is empirical-only and not in this calculator. Use `SOAK_MODE=continuous` to probe it.
|
||||
|
||||
5. **Single-model**. The math is Qwen3.6-27B-specific. Adding a model means deriving a new `MODEL_SPEC` block (architecture params + weights size) and re-calibrating the activation coefficient against new measured points.
|
||||
5. **Two-model calibration**. Today the math is calibrated for Qwen3.6-27B and Gemma 4 31B. Adding a third model means deriving a new `MODEL_SPEC` block (architecture params + per-quant weights size + activation-peak mechanism) and calibrating the activation coefficient against ≥4 measured BENCHMARKS rows for that model. Don't ship a third-model spec without that calibration — uncalibrated coefficients give wrong verdicts.
|
||||
|
||||
## Calibration
|
||||
|
||||
Run `bash tools/kv-calc.py --calibration` to see predicted vs measured for all shipped composes. Current verdict accuracy: **9/11 = 82%** with the ±1.5 GB error band. The two ⨯ cases are over-predictions on `max_num_seqs > 1` configs (limitation #1 above).
|
||||
Run `bash tools/kv-calc.py --calibration` to see predicted vs measured for all shipped composes, grouped per model.
|
||||
|
||||
| Model | Verdict accuracy | Notes |
|
||||
|---|---|---|
|
||||
| Qwen 3.6 27B | 9/11 = 82% (±1.5 GB band) | Two ⨯ cases are over-predictions on `max_num_seqs > 1` (limitation #1 above) |
|
||||
| Gemma 4 31B | TBD by `--calibration` | Calibrated against `dual/int8.yml` 98K+262K rows, `dual/dflash.yml` 32K BF16, `dual/awq.yml` 65K BF16, `dual/docker-compose.yml` 32K BF16. Target: ≥80% within ±1.5 GB. |
|
||||
|
||||
## When to trust the calculator vs vLLM's boot log
|
||||
|
||||
Always pass `--model {qwen3.6-27b,gemma-4-31b}` matching the compose you're targeting. Defaults to qwen3.6-27b if omitted.
|
||||
|
||||
| Question | Use this |
|
||||
|---|---|
|
||||
| "Will it boot?" — for a *shipped* compose on canonical 24 GB | We've already validated; check BENCHMARKS.md |
|
||||
| "Will it boot?" — for a *novel* config (custom ctx, kv format, or VRAM class) | `kv-calc.py --compose <X>` for a directional answer; then boot and read `gpu_worker.py` |
|
||||
| "What's my max ctx?" — given my hardware | `kv-calc.py --solve-max-ctx ...` for an estimate; vLLM's pre-check `estimated max model length is N` line at boot is authoritative |
|
||||
| "Is TQ3 or fp8 better for my hardware?" | `kv-calc.py` with both options to see the trade-off; cross-check against [HARDWARE.md](HARDWARE.md#note-for-sub-24-gb-cards) for the published guidance |
|
||||
| "Will it boot?" — for a *novel* config (custom ctx, kv format, or VRAM class) | `kv-calc.py --model <M> --compose <X>` for a directional answer; then boot and read `gpu_worker.py` |
|
||||
| "What's my max ctx?" — given my hardware | `kv-calc.py --model <M> --solve-max-ctx ...` for an estimate; vLLM's pre-check `estimated max model length is N` line at boot is authoritative |
|
||||
| "Is TQ3 or fp8 better for my hardware?" (Qwen 3.6) | `kv-calc.py --model qwen3.6-27b` with both options; cross-check [HARDWARE.md](HARDWARE.md#note-for-sub-24-gb-cards) |
|
||||
| "Is INT8 PTH or BF16 KV better for Gemma 4?" | `kv-calc.py --model gemma-4-31b --kv-format bf16` vs `int8_per_token_head` — BF16 caps at ~32K on dual-3090, INT8 PTH unlocks 262K. See `models/gemma-4-31b/vllm/compose/dual/int8.yml` header. |
|
||||
|
||||
## References
|
||||
|
||||
**Qwen 3.6 27B (DeltaNet hybrid):**
|
||||
- [PerfMamba: Performance Analysis and Pruning of Selective State Space Models (arxiv 2511.22849)](https://arxiv.org/html/2511.22849) — block-wise state materialization scaling
|
||||
- [TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (arxiv 2504.19874, ICLR 2026)](https://arxiv.org/abs/2504.19874) — TQ3 byte savings + technique
|
||||
- [Gated Delta Networks: Improving Mamba2 with Delta Rule (NVlabs ICLR 2025)](https://github.com/NVlabs/GatedDeltaNet) — Qwen3-Next architecture
|
||||
- [Mamba: Linear-Time Sequence Modeling (arxiv 2312.00752)](https://arxiv.org/abs/2312.00752) — Mamba-1 baseline for PerfMamba's deltas
|
||||
|
||||
**Gemma 4 31B (sliding-window + dense MLP):**
|
||||
- Architecture params sourced directly from `config.json` (Gemma 4 release post / technical doc were not used as a calibration reference — the activation coefficient is empirical-only on this stack).
|
||||
- [vLLM PR #40391 (rebased + vendored as PR #42102)](https://github.com/vllm-project/vllm/pull/42102) — per-token-head INT8 KV cache (the Ampere unlock for Gemma 4 at 262K)
|
||||
- [vLLM PR #41745](https://github.com/vllm-project/vllm/pull/41745) — Gemma 4 MTP assistant drafter support
|
||||
|
||||
**Shared:**
|
||||
- [TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (arxiv 2504.19874, ICLR 2026)](https://arxiv.org/abs/2504.19874) — TQ3 byte savings + technique
|
||||
- [Efficient Memory Management for Large Language Model Serving with PagedAttention (arxiv 2309.06180)](https://arxiv.org/abs/2309.06180) — vLLM's KV pool allocator
|
||||
- [An Investigation of FP8 Across Accelerators for LLM Inference (arxiv 2502.01070)](https://arxiv.org/html/2502.01070v1) — FP8 e5m2/e4m3 KV cache analysis
|
||||
- [docs/CLIFFS.md](CLIFFS.md) — Cliff 2 mechanism + KV-format-tunability section
|
||||
- [docs/HARDWARE.md](HARDWARE.md) — 20 GB Ampere TQ3→fp8 swap rule (cross-rig validated by @efschu)
|
||||
- [docs/CLIFFS.md](CLIFFS.md) — Cliff 2 mechanism + KV-format-tunability section (Qwen-specific)
|
||||
- [docs/HARDWARE.md](HARDWARE.md) — 20 GB Ampere TQ3→fp8 swap rule (cross-rig validated by @efschu, Qwen-specific)
|
||||
|
||||
## See also
|
||||
|
||||
|
||||
+1
-1
@@ -204,7 +204,7 @@ Generally prefer **dropping `MAX_MODEL_LEN` first** (clean KV budget reduction,
|
||||
# 1. Setup (downloads model, clones Genesis, ~20 min cold)
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
|
||||
# 2. Pick + boot via wizard (asks engine + workload)
|
||||
# 2. Pick + boot via wizard (asks model + GPUs, projects VRAM budget)
|
||||
bash scripts/launch.sh
|
||||
|
||||
# 3. Or skip the wizard:
|
||||
|
||||
@@ -60,6 +60,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-full
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -118,7 +119,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-awq
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --dtype
|
||||
- "${DTYPE:-bfloat16}"
|
||||
- --disable-custom-all-reduce
|
||||
|
||||
@@ -91,6 +91,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -151,7 +152,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
# ---- KV format ----
|
||||
# Default `auto` runs on every consumer GPU including
|
||||
|
||||
@@ -95,6 +95,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -183,7 +184,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --dtype
|
||||
- bfloat16
|
||||
- --disable-custom-all-reduce
|
||||
|
||||
@@ -67,6 +67,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -148,7 +149,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# DFlash drafter is BF16; --dtype bfloat16 matches its training dtype
|
||||
# (same as our existing dual-dflash.yml on Qwen3.6 — vllm#40334 dtype-mismatch
|
||||
# workaround until that lands).
|
||||
|
||||
@@ -50,6 +50,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -99,7 +100,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
# max-model-len ceiling (TP=2, BF16 KV):
|
||||
# 32K ctx @ 0.92 → KV pool ~38K tokens (shipped default — matches bench above)
|
||||
|
||||
@@ -179,6 +179,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-full
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
# Requires-sm: 9.0+
|
||||
@@ -268,7 +269,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
# ---- 2026-05-11 workaround flags ----
|
||||
# Force TURBOQUANT backend — required because vLLM has no per-layer
|
||||
|
||||
@@ -91,6 +91,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-full
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -151,7 +152,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
# ---- KV format ----
|
||||
# Default `int8_per_token_head` runs on every consumer GPU including
|
||||
|
||||
@@ -50,6 +50,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 32
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
# Requires-sm: 9.0+
|
||||
@@ -100,7 +101,9 @@ services:
|
||||
- --served-model-name
|
||||
- gemma-4-31b-autoround
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# Single-card: weights ~16 GB + drafter ~1 GB + bf16 KV at 16K + vision +
|
||||
# activations needs to fit in 24 GB. Drop vision via limit-mm-per-prompt
|
||||
# to reclaim ~1.5 GB and skip the mm-token-budget assertion.
|
||||
|
||||
@@ -9,6 +9,7 @@
|
||||
# Max ctx: 192K pool / 4 parallel slots
|
||||
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
|
||||
# Status: ✅ Production
|
||||
# Engine-profile: llama-cpp-mainline
|
||||
# Best for: Single-card multi-tenant llama.cpp — 4 concurrent agents
|
||||
# at smaller per-stream ctx; trade max-ctx for parallelism
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@@ -9,6 +9,7 @@
|
||||
# Max ctx: 262K (full model native — no Cliff 1 / Cliff 2)
|
||||
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
|
||||
# Status: ✅ Production
|
||||
# Engine-profile: llama-cpp-mainline
|
||||
# Best for: Bulletproof single-card path — slow decode (~21 TPS) but
|
||||
# cliff-immune; recommended fallback when vLLM hits OOM at long ctx
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -99,7 +100,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-200000}"
|
||||
|
||||
@@ -50,6 +50,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -128,7 +129,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -45,6 +45,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -122,7 +123,9 @@ services:
|
||||
- --dtype
|
||||
- bfloat16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-200000}"
|
||||
|
||||
@@ -69,6 +69,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -146,7 +147,9 @@ services:
|
||||
- --dtype
|
||||
- bfloat16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-185000}"
|
||||
|
||||
@@ -51,6 +51,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -137,7 +138,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-full
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -99,7 +100,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -77,6 +77,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -156,7 +157,9 @@ services:
|
||||
- --dtype
|
||||
- bfloat16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
|
||||
# makes vLLM's custom kernel a win. dual-dflash-noviz.yml disables it
|
||||
# because PCIe P2P bandwidth makes the NCCL fallback faster there.
|
||||
|
||||
@@ -73,6 +73,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -152,7 +153,9 @@ services:
|
||||
- --dtype
|
||||
- bfloat16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
|
||||
# makes vLLM's custom kernel a win. dual-dflash.yml disables it because
|
||||
# PCIe P2P bandwidth makes the NCCL fallback faster there.
|
||||
|
||||
@@ -63,6 +63,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -256,7 +257,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
|
||||
# makes vLLM's custom kernel a win. dual-turbo.yml disables it because
|
||||
# PCIe P2P bandwidth makes the NCCL fallback faster there.
|
||||
|
||||
@@ -59,6 +59,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -144,7 +145,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
|
||||
# makes vLLM's custom kernel a win. dual.yml disables it because PCIe
|
||||
# P2P bandwidth makes the NCCL fallback faster there.
|
||||
|
||||
@@ -35,6 +35,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -98,7 +99,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -63,6 +63,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -257,7 +258,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -61,6 +61,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -147,7 +148,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -37,6 +37,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -130,7 +131,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -47,6 +47,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
@@ -238,7 +239,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "2"
|
||||
- "${TP:-2}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -62,6 +62,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-dflash
|
||||
# Requires-min-gpu-count: 4
|
||||
# Tensor-parallel: 4
|
||||
services:
|
||||
@@ -139,7 +140,9 @@ services:
|
||||
- --dtype
|
||||
- bfloat16
|
||||
- --tensor-parallel-size
|
||||
- "4"
|
||||
- "${TP:-4}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -59,6 +59,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 4
|
||||
# Tensor-parallel: 4
|
||||
services:
|
||||
@@ -145,7 +146,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "4"
|
||||
- "${TP:-4}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --disable-custom-all-reduce
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-262144}"
|
||||
|
||||
@@ -131,6 +131,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -307,7 +308,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# 180K + 0.95 — parity with long-text.yml. Backed off from 214K + 0.985
|
||||
# on 2026-05-02 to give activation headroom for the PN12+PN25 FFN pool
|
||||
# residence + DeltaNet GDN buffer. See long-text.yml for full rationale.
|
||||
|
||||
@@ -97,6 +97,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -238,7 +239,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# 48K + 0.92 — production default. Below BOTH cliffs:
|
||||
# - GDN forward (single-prompt OOM at ~50-60K tokens) → 48K stays safely under
|
||||
# - TurboQuant tool prefill (OOM at high mem_util + big tool responses) → 0.92 leaves ~1.9 GB headroom
|
||||
|
||||
@@ -111,6 +111,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -336,7 +337,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# 180K + 0.97 — backed off from 214K + 0.985 on 2026-05-01 PM after
|
||||
# verify-stress probe 1 (10K-token long-context needle) crashed with
|
||||
# GDN forward OOM (`fla.ops.chunk.chunk_gated_delta_rule_fwd_h`
|
||||
|
||||
@@ -121,6 +121,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -353,7 +354,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# 180K + 0.97 — backed off from 214K + 0.985 on 2026-05-01 PM after
|
||||
# verify-stress probe 1 (10K-token long-context needle) crashed with
|
||||
# GDN forward OOM (`fla.ops.chunk.chunk_gated_delta_rule_fwd_h`
|
||||
|
||||
@@ -97,6 +97,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -273,7 +274,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
# 198K + 0.98 — restored on v0.20 + Genesis v7.65 migration. v0.20
|
||||
# closes the Cliff 1 mech B sub-cliffs that drove the 140K + 0.95
|
||||
# backoff on dev205. Verified verify-full + 33K + 50K stress all
|
||||
|
||||
@@ -33,6 +33,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 20
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -107,7 +108,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-32768}"
|
||||
- --gpu-memory-utilization
|
||||
|
||||
@@ -35,6 +35,7 @@
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Engine-profile: vllm-nightly-mtp
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
@@ -162,7 +163,9 @@ services:
|
||||
- --dtype
|
||||
- float16
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- "${TP:-1}"
|
||||
- --pipeline-parallel-size
|
||||
- "${PP:-1}"
|
||||
- --max-model-len
|
||||
- "${MAX_MODEL_LEN:-75000}"
|
||||
- --gpu-memory-utilization
|
||||
|
||||
+721
-91
@@ -1,18 +1,23 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# Interactive launcher for club-3090 — pick engine + workload, boot the
|
||||
# right compose, run verify-full to confirm it's serving.
|
||||
# Interactive launcher for club-3090 — pick model + GPUs, project the
|
||||
# VRAM budget, boot the right compose, run verify-full to confirm it's serving.
|
||||
#
|
||||
# For first-run users coming in from the README. If you already know
|
||||
# what you want, use `scripts/switch.sh <variant>` directly.
|
||||
#
|
||||
# Usage:
|
||||
# bash scripts/launch.sh # interactive wizard
|
||||
# bash scripts/launch.sh # interactive model/GPU wizard
|
||||
# bash scripts/launch.sh --variant <name> # skip wizard, boot directly
|
||||
# bash scripts/launch.sh --engine vllm --cards 1 # partial flags, ask the rest
|
||||
# bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1
|
||||
# bash scripts/launch.sh --engine vllm --cards 1 # deprecated; prefer --gpus
|
||||
# bash scripts/launch.sh --tp 2 --pp 1 # override vLLM parallelism
|
||||
# bash scripts/launch.sh --no-projection # skip kv-calc budget projection
|
||||
# bash scripts/launch.sh --no-verify # skip post-launch verify-full
|
||||
# bash scripts/launch.sh --no-preflight # skip docker/GPU pre-flight
|
||||
#
|
||||
# The wizard marks variants that don't fit the detected GPUs; direct
|
||||
# --variant keeps the power-user path and delegates final gating to switch.sh.
|
||||
# All flags accept the same names as `switch.sh --list` produces.
|
||||
# Examples:
|
||||
# bash scripts/launch.sh --variant vllm/default
|
||||
@@ -24,6 +29,13 @@ set -euo pipefail
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SWITCH="${SWITCH:-${ROOT_DIR}/scripts/switch.sh}"
|
||||
VERIFY="${VERIFY:-${ROOT_DIR}/scripts/verify-full.sh}"
|
||||
if [[ -z "${MODEL_DIR:-}" && -f "${ROOT_DIR}/.env" ]]; then
|
||||
set -a
|
||||
# shellcheck source=/dev/null
|
||||
source "${ROOT_DIR}/.env"
|
||||
set +a
|
||||
fi
|
||||
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
|
||||
# shellcheck source=preflight.sh
|
||||
source "${ROOT_DIR}/scripts/preflight.sh"
|
||||
|
||||
@@ -31,13 +43,25 @@ source "${ROOT_DIR}/scripts/preflight.sh"
|
||||
ENGINE=""
|
||||
CARDS=""
|
||||
VARIANT=""
|
||||
MODEL_NAME=""
|
||||
GPU_ARG=""
|
||||
TP_OVERRIDE=""
|
||||
PP_OVERRIDE=""
|
||||
PARALLELISM="auto"
|
||||
SKIP_VERIFY=0
|
||||
SKIP_PREFLIGHT=0
|
||||
SKIP_PROJECTION=0
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--engine) ENGINE="$2"; shift 2 ;;
|
||||
--cards) CARDS="$2"; shift 2 ;;
|
||||
--variant) VARIANT="$2"; shift 2 ;;
|
||||
--model) MODEL_NAME="$2"; shift 2 ;;
|
||||
--gpus) GPU_ARG="$2"; shift 2 ;;
|
||||
--tp) TP_OVERRIDE="$2"; shift 2 ;;
|
||||
--pp) PP_OVERRIDE="$2"; shift 2 ;;
|
||||
--parallelism) PARALLELISM="$2"; shift 2 ;;
|
||||
--no-projection) SKIP_PROJECTION=1; shift ;;
|
||||
--no-verify) SKIP_VERIFY=1; shift ;;
|
||||
--no-preflight) SKIP_PREFLIGHT=1; shift ;;
|
||||
-h|--help)
|
||||
@@ -72,6 +96,19 @@ ask() {
|
||||
fi
|
||||
}
|
||||
|
||||
read_or_interrupt() {
|
||||
local prompt="$1"
|
||||
local __reply_var="$2"
|
||||
local reply
|
||||
if ! read -rp "$prompt" reply; then
|
||||
echo "" >&2
|
||||
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
|
||||
kill -INT $$
|
||||
exit 1
|
||||
fi
|
||||
printf -v "$__reply_var" '%s' "$reply"
|
||||
}
|
||||
|
||||
choose() {
|
||||
# choose "prompt" "label1" "value1" "label2" "value2" ... -> echoes chosen value
|
||||
local prompt="$1"; shift
|
||||
@@ -87,7 +124,12 @@ choose() {
|
||||
done
|
||||
while true; do
|
||||
local pick
|
||||
read -rp "Choice [1-${#labels[@]}]: " pick
|
||||
if ! read -rp "Choice [1-${#labels[@]}]: " pick; then
|
||||
echo "" >&2
|
||||
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
|
||||
kill -INT $$
|
||||
exit 1
|
||||
fi
|
||||
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
|
||||
echo "${values[$((pick-1))]}"
|
||||
return
|
||||
@@ -96,109 +138,691 @@ choose() {
|
||||
done
|
||||
}
|
||||
|
||||
# --- wizard ---
|
||||
# Flow: cards → workload → auto-pick engine. Newcomers can answer "how
|
||||
# many GPUs" and "what do I want to do" but rarely "vLLM or llama.cpp" —
|
||||
# so the engine is derived from the workload pick, not asked first.
|
||||
# Manual --engine override filters the workload list to that engine.
|
||||
if [[ -z "$VARIANT" ]]; then
|
||||
echo ""
|
||||
echo "club-3090 launcher — let's pick the right config for your workload."
|
||||
echo "(Use --variant <name> next time to skip the wizard.)"
|
||||
echo ""
|
||||
declare -A LAUNCH_VARIANT_COMPOSE=(
|
||||
[vllm/default]="models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
|
||||
[vllm/long-vision]="models/qwen3.6-27b/vllm/compose/single/long-vision.yml"
|
||||
[vllm/long-text]="models/qwen3.6-27b/vllm/compose/single/long-text.yml"
|
||||
[vllm/long-text-no-mtp]="models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml"
|
||||
[vllm/bounded-thinking]="models/qwen3.6-27b/vllm/compose/single/bounded-thinking.yml"
|
||||
[vllm/tools-text]="models/qwen3.6-27b/vllm/compose/single/tools-text.yml"
|
||||
[vllm/minimal]="models/qwen3.6-27b/vllm/compose/single/minimal.yml"
|
||||
[vllm/dual]="models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml"
|
||||
[vllm/dual4]="models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml"
|
||||
[vllm/dual4-dflash]="models/qwen3.6-27b/vllm/compose/multi4/dflash.yml"
|
||||
[vllm/dual-turbo]="models/qwen3.6-27b/vllm/compose/dual/turbo.yml"
|
||||
[vllm/dual-dflash]="models/qwen3.6-27b/vllm/compose/dual/dflash.yml"
|
||||
[vllm/dual-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/dflash-noviz.yml"
|
||||
[vllm/dual-nvlink]="models/qwen3.6-27b/vllm/compose/dual/nvlink.yml"
|
||||
[vllm/dual-nvlink-turbo]="models/qwen3.6-27b/vllm/compose/dual/nvlink-turbo.yml"
|
||||
[vllm/dual-nvlink-dflash]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash.yml"
|
||||
[vllm/dual-nvlink-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash-noviz.yml"
|
||||
[vllm/gemma-mtp]="models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
|
||||
[vllm/gemma-mtp-tp1]="models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
|
||||
[vllm/gemma-dflash]="models/gemma-4-31b/vllm/compose/dual/dflash.yml"
|
||||
[llamacpp/default]="models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml"
|
||||
[llamacpp/concurrent]="models/qwen3.6-27b/llama-cpp/compose/single/concurrent.yml"
|
||||
)
|
||||
declare -A LAUNCH_VARIANT_MODEL=(
|
||||
[vllm/default]="qwen3.6-27b" [vllm/long-vision]="qwen3.6-27b" [vllm/long-text]="qwen3.6-27b"
|
||||
[vllm/long-text-no-mtp]="qwen3.6-27b" [vllm/bounded-thinking]="qwen3.6-27b" [vllm/tools-text]="qwen3.6-27b"
|
||||
[vllm/minimal]="qwen3.6-27b" [vllm/dual]="qwen3.6-27b" [vllm/dual4]="qwen3.6-27b"
|
||||
[vllm/dual4-dflash]="qwen3.6-27b" [vllm/dual-turbo]="qwen3.6-27b" [vllm/dual-dflash]="qwen3.6-27b"
|
||||
[vllm/dual-dflash-noviz]="qwen3.6-27b" [vllm/dual-nvlink]="qwen3.6-27b" [vllm/dual-nvlink-turbo]="qwen3.6-27b"
|
||||
[vllm/dual-nvlink-dflash]="qwen3.6-27b" [vllm/dual-nvlink-dflash-noviz]="qwen3.6-27b"
|
||||
[vllm/gemma-mtp]="gemma-4-31b" [vllm/gemma-mtp-tp1]="gemma-4-31b" [vllm/gemma-dflash]="gemma-4-31b"
|
||||
[llamacpp/default]="qwen3.6-27b" [llamacpp/concurrent]="qwen3.6-27b"
|
||||
)
|
||||
declare -A LAUNCH_VARIANT_ENGINE=(
|
||||
[vllm/default]="vllm" [vllm/long-vision]="vllm" [vllm/long-text]="vllm" [vllm/long-text-no-mtp]="vllm"
|
||||
[vllm/bounded-thinking]="vllm" [vllm/tools-text]="vllm" [vllm/minimal]="vllm" [vllm/dual]="vllm"
|
||||
[vllm/dual4]="vllm" [vllm/dual4-dflash]="vllm" [vllm/dual-turbo]="vllm" [vllm/dual-dflash]="vllm"
|
||||
[vllm/dual-dflash-noviz]="vllm" [vllm/dual-nvlink]="vllm" [vllm/dual-nvlink-turbo]="vllm"
|
||||
[vllm/dual-nvlink-dflash]="vllm" [vllm/dual-nvlink-dflash-noviz]="vllm"
|
||||
[vllm/gemma-mtp]="vllm" [vllm/gemma-mtp-tp1]="vllm" [vllm/gemma-dflash]="vllm"
|
||||
[llamacpp/default]="llamacpp" [llamacpp/concurrent]="llamacpp"
|
||||
)
|
||||
declare -A LAUNCH_VARIANT_KVCALC=(
|
||||
[vllm/default]="qwen3.6-27b:long-vision"
|
||||
[vllm/long-text]="qwen3.6-27b:long-text"
|
||||
[vllm/long-text-no-mtp]="qwen3.6-27b:long-text-no-mtp"
|
||||
[vllm/long-vision]="qwen3.6-27b:long-vision"
|
||||
[vllm/bounded-thinking]="qwen3.6-27b:bounded-thinking"
|
||||
[vllm/tools-text]="qwen3.6-27b:tools-text"
|
||||
[vllm/minimal]="qwen3.6-27b:minimal"
|
||||
[vllm/dual]="qwen3.6-27b:dual"
|
||||
[vllm/dual-turbo]="qwen3.6-27b:dual-turbo"
|
||||
[vllm/dual-dflash]="qwen3.6-27b:dual-dflash"
|
||||
[vllm/dual-dflash-noviz]="qwen3.6-27b:dual-dflash-noviz"
|
||||
[vllm/dual4]="qwen3.6-27b:dual4"
|
||||
[vllm/dual4-dflash]="qwen3.6-27b:dual4-dflash"
|
||||
[vllm/dual-nvlink]="qwen3.6-27b:dual"
|
||||
[vllm/dual-nvlink-turbo]="qwen3.6-27b:dual-turbo"
|
||||
[vllm/dual-nvlink-dflash]="qwen3.6-27b:dual-dflash"
|
||||
[vllm/dual-nvlink-dflash-noviz]="qwen3.6-27b:dual-dflash-noviz"
|
||||
[vllm/gemma-mtp]="gemma-4-31b:gemma-dual"
|
||||
[vllm/gemma-mtp-tp1]="gemma-4-31b:gemma-single"
|
||||
[vllm/gemma-dflash]="gemma-4-31b:gemma-dual-dflash"
|
||||
[llamacpp/default]="SKIP"
|
||||
[llamacpp/concurrent]="SKIP"
|
||||
)
|
||||
LAUNCH_VARIANT_ORDER=(
|
||||
vllm/long-vision vllm/long-text vllm/long-text-no-mtp vllm/bounded-thinking
|
||||
vllm/default vllm/tools-text vllm/minimal
|
||||
vllm/dual vllm/dual-turbo vllm/dual-dflash vllm/dual-dflash-noviz
|
||||
vllm/dual4 vllm/dual4-dflash
|
||||
vllm/gemma-mtp vllm/gemma-mtp-tp1 vllm/gemma-dflash
|
||||
llamacpp/default llamacpp/concurrent
|
||||
)
|
||||
|
||||
# Step 1 — cards.
|
||||
if [[ -z "$CARDS" ]]; then
|
||||
CARDS=$(choose "How many RTX 3090s?" \
|
||||
"1× 3090 (24 GB)" "1" \
|
||||
"2× 3090 (PCIe / no NVLink)" "2")
|
||||
variant_hw_status() {
|
||||
local variant="$1"
|
||||
local rel="${LAUNCH_VARIANT_COMPOSE[$variant]:-}"
|
||||
if [[ -z "$rel" ]]; then
|
||||
printf 'ok|fits your rig'
|
||||
return 0
|
||||
fi
|
||||
|
||||
# Re-validate now that we know the requirement (preflight already ran
|
||||
# with min=1; bump to actual count). Skip if --no-preflight.
|
||||
if [[ $SKIP_PREFLIGHT -eq 0 ]]; then
|
||||
preflight_gpu "$CARDS" || exit 1
|
||||
local compose_file="${ROOT_DIR}/${rel}"
|
||||
if [[ ! -f "$compose_file" ]]; then
|
||||
printf 'unknown|compose metadata unavailable'
|
||||
return 2
|
||||
fi
|
||||
compose_hw_compose_status "$compose_file" 2>/dev/null || true
|
||||
}
|
||||
|
||||
choose_variant() {
|
||||
# choose_variant "prompt" "default-variant" "label1" "value1" ...
|
||||
local prompt="$1" default_variant="$2"
|
||||
shift 2
|
||||
|
||||
local i labels=() values=() statuses=() eligible=()
|
||||
while [[ $# -gt 0 ]]; do
|
||||
labels+=("$1")
|
||||
values+=("$2")
|
||||
statuses+=("$(variant_hw_status "$2")")
|
||||
shift 2
|
||||
done
|
||||
|
||||
local default_idx=""
|
||||
for i in "${!values[@]}"; do
|
||||
if [[ "${values[$i]}" == "$default_variant" && "${statuses[$i]}" == ok\|* ]]; then
|
||||
default_idx=$((i + 1))
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [[ -z "$default_idx" ]]; then
|
||||
for i in "${!values[@]}"; do
|
||||
if [[ "${statuses[$i]}" == ok\|* || "${statuses[$i]}" == unknown\|* ]]; then
|
||||
default_idx=$((i + 1))
|
||||
break
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
# Step 2 — workload, filtered by cards (and --engine override if set).
|
||||
# Each option's value is "engine/file" so engine is implied by the pick.
|
||||
if [[ "$CARDS" == "1" ]]; then
|
||||
# Primary recommended options first; diagnostic / niche variants in
|
||||
# an "Other" group at the end. The only single-card limitation users
|
||||
# need to know: vLLM single-card crashes on a single prompt >50K
|
||||
# (Cliff 2). For unpredictable inputs, use llamacpp/default.
|
||||
if [[ -z "$ENGINE" || "$ENGINE" == "vllm" ]]; then
|
||||
VLLM_OPTS=(
|
||||
"Long ctx + vision (145K + vision, MTP) — recommended for chat/agents" "vllm/long-vision"
|
||||
"Long ctx, text only — Balanced MTP (180K, MTP K=3) — recommended IDE-agent" "vllm/long-text"
|
||||
"Long ctx, text only — Max-context (200K, no MTP) — one-shot >50K prompts" "vllm/long-text-no-mtp"
|
||||
"Bounded thinking (180K, structured-CoT FSM — recommended grammar: DeepSeek scratchpad, 87.4% combined HE+/LCB v6)" "vllm/bounded-thinking"
|
||||
)
|
||||
echo "" >&2
|
||||
echo "$prompt" >&2
|
||||
for i in "${!labels[@]}"; do
|
||||
local status="${statuses[$i]}"
|
||||
local state="${status%%|*}"
|
||||
local reason="${status#*|}"
|
||||
local marker="✓"
|
||||
case "$state" in
|
||||
ok) marker="✓" ;;
|
||||
unknown) marker="?" ;;
|
||||
*) marker="✗" ;;
|
||||
esac
|
||||
if [[ -n "$default_idx" && $((i + 1)) -eq "$default_idx" ]]; then
|
||||
printf " %d) %s %s %s [default]\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
|
||||
else
|
||||
VLLM_OPTS=()
|
||||
printf " %d) %s %s %s\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
|
||||
fi
|
||||
if [[ -z "$ENGINE" || "$ENGINE" == "llamacpp" ]]; then
|
||||
LLAMA_OPTS=(
|
||||
"Bulletproof, no cliffs (262K + vision, ~21 TPS) — production-safe" "llamacpp/default"
|
||||
)
|
||||
else
|
||||
LLAMA_OPTS=()
|
||||
fi
|
||||
# Diagnostic / niche fallbacks — shown last so they don't dominate the menu
|
||||
if [[ -z "$ENGINE" || "$ENGINE" == "vllm" ]]; then
|
||||
VLLM_FALLBACK_OPTS=(
|
||||
"[fallback] Default 48K + vision (Cliff 2 unreachable; fast boot)" "vllm/default"
|
||||
"[fallback] tools-text 75K FP8 (FP8 KV alternative for accuracy compare)" "vllm/tools-text"
|
||||
"[fallback] minimal 32K (no Genesis, no spec-decode — diagnostic stack)" "vllm/minimal"
|
||||
)
|
||||
else
|
||||
VLLM_FALLBACK_OPTS=()
|
||||
fi
|
||||
if [[ -z "$ENGINE" || "$ENGINE" == "llamacpp" ]]; then
|
||||
LLAMA_FALLBACK_OPTS=(
|
||||
"[fallback] llamacpp/concurrent (4 parallel slots, 192K pool, vision)" "llamacpp/concurrent"
|
||||
)
|
||||
else
|
||||
LLAMA_FALLBACK_OPTS=()
|
||||
fi
|
||||
VARIANT=$(choose "What's your main workload?" \
|
||||
"${VLLM_OPTS[@]}" "${LLAMA_OPTS[@]}" \
|
||||
"${VLLM_FALLBACK_OPTS[@]}" "${LLAMA_FALLBACK_OPTS[@]}")
|
||||
elif [[ "$CARDS" == "2" ]]; then
|
||||
if [[ -n "$ENGINE" && "$ENGINE" != "vllm" ]]; then
|
||||
echo "ERROR: --engine ${ENGINE} not supported on 2× cards (no llama.cpp dual recipe yet)." >&2
|
||||
exit 1
|
||||
fi
|
||||
VARIANT=$(choose "What's your dual-card priority?" \
|
||||
"Balanced default — 262K + vision + 2 streams (recommended)" "vllm/dual" \
|
||||
"Multi-tenant — 4 concurrent streams @ 262K, TQ3 KV" "vllm/dual-turbo" \
|
||||
"Peak code TPS with vision (185K, DFlash N=5)" "vllm/dual-dflash" \
|
||||
"Peak code TPS no vision (200K, DFlash N=5)" "vllm/dual-dflash-noviz")
|
||||
else
|
||||
echo "ERROR: --cards ${CARDS} unsupported (expected 1 or 2)." >&2
|
||||
done
|
||||
|
||||
if [[ -z "$default_idx" ]]; then
|
||||
echo "ERROR: no eligible variants in this menu. Use scripts/switch.sh --force <variant> to attempt anyway." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Step 3 — explain the auto-picked engine.
|
||||
echo ""
|
||||
case "$VARIANT" in
|
||||
llamacpp/*)
|
||||
echo "[wizard] picked llama.cpp — chosen because no prefill cliffs at 262K and the simplest"
|
||||
echo "[wizard] serving path. Trade-off: ~21 TPS vs 70+ on vLLM. Right call for long-prompt"
|
||||
echo "[wizard] robustness, frontier context, or anyone who wants the simplest setup."
|
||||
;;
|
||||
vllm/*)
|
||||
echo "[wizard] picked vLLM — chosen for spec-decode (MTP), tool-call extraction, and best TPS."
|
||||
echo "[wizard] Watch out for the prefill cliffs (Cliff 1 = 25K+ tool returns; Cliff 2 = 50-60K"
|
||||
echo "[wizard] single prompts on TQ3) — see docs/SINGLE_CARD.md for the safe-config map."
|
||||
;;
|
||||
while true; do
|
||||
local pick
|
||||
if ! read -rp "Choice [1-${#labels[@]}, default ${default_idx}]: " pick; then
|
||||
echo "" >&2
|
||||
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
|
||||
kill -INT $$
|
||||
exit 1
|
||||
fi
|
||||
pick="${pick:-$default_idx}"
|
||||
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
|
||||
local status="${statuses[$((pick - 1))]}"
|
||||
if [[ "$status" == no\|* ]]; then
|
||||
echo " That variant won't run on your detected rig: ${status#*|}" >&2
|
||||
echo " Pick another, or use: bash scripts/switch.sh --force ${values[$((pick - 1))]}" >&2
|
||||
continue
|
||||
fi
|
||||
echo "${values[$((pick - 1))]}"
|
||||
return
|
||||
fi
|
||||
echo " invalid — pick a number 1-${#labels[@]}" >&2
|
||||
done
|
||||
}
|
||||
|
||||
model_label() {
|
||||
case "$1" in
|
||||
qwen3.6-27b) echo "Qwen 3.6 27B" ;;
|
||||
gemma-4-31b) echo "Gemma 4 31B" ;;
|
||||
*) echo "$1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
normalize_model_name() {
|
||||
case "$1" in
|
||||
qwen3.6-27b|qwen3.6-27b-gguf) echo "qwen3.6-27b" ;;
|
||||
gemma-4-31b|gemma-4-31b-awq|gemma-4-31b-gguf) echo "gemma-4-31b" ;;
|
||||
*) echo "$1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
MODEL_ORDER=()
|
||||
declare -A MODEL_ENGINES=()
|
||||
|
||||
add_installed_model_engine() {
|
||||
local model="$1" engine="$2"
|
||||
if [[ -z "${MODEL_ENGINES[$model]:-}" ]]; then
|
||||
MODEL_ORDER+=("$model")
|
||||
MODEL_ENGINES[$model]="$engine"
|
||||
elif [[ ",${MODEL_ENGINES[$model]}," != *",${engine},"* ]]; then
|
||||
MODEL_ENGINES[$model]="${MODEL_ENGINES[$model]},${engine}"
|
||||
fi
|
||||
}
|
||||
|
||||
detect_installed_models() {
|
||||
MODEL_ORDER=()
|
||||
MODEL_ENGINES=()
|
||||
[[ -d "${MODEL_DIR}/qwen3.6-27b-autoround-int4" ]] && add_installed_model_engine "qwen3.6-27b" "vllm"
|
||||
[[ -d "${MODEL_DIR}/qwen3.6-27b-gguf" ]] && add_installed_model_engine "qwen3.6-27b" "llamacpp"
|
||||
[[ -d "${MODEL_DIR}/gemma-4-31b-autoround-int4" ]] && add_installed_model_engine "gemma-4-31b" "vllm"
|
||||
[[ -d "${MODEL_DIR}/gemma-4-31b-it-AWQ-4bit" ]] && add_installed_model_engine "gemma-4-31b" "vllm"
|
||||
[[ -d "${MODEL_DIR}/gemma-4-31b-gguf" ]] && add_installed_model_engine "gemma-4-31b" "llamacpp"
|
||||
return 0
|
||||
}
|
||||
|
||||
model_has_engine() {
|
||||
local model="$1" engine="$2"
|
||||
[[ ",${MODEL_ENGINES[$model]:-}," == *",${engine},"* ]]
|
||||
}
|
||||
|
||||
engine_hint() {
|
||||
case "$1" in
|
||||
vllm,llamacpp|llamacpp,vllm) echo "vLLM + llama.cpp engines available" ;;
|
||||
vllm) echo "vLLM only" ;;
|
||||
llamacpp) echo "llama.cpp only" ;;
|
||||
*) echo "$1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
choose_model() {
|
||||
detect_installed_models
|
||||
if [[ -n "$MODEL_NAME" ]]; then
|
||||
MODEL_NAME="$(normalize_model_name "$MODEL_NAME")"
|
||||
if [[ -z "${MODEL_ENGINES[$MODEL_NAME]:-}" ]]; then
|
||||
echo "[launch] ERROR: ${MODEL_NAME} is not installed under ${MODEL_DIR}." >&2
|
||||
echo "[launch] Run: bash scripts/setup.sh ${MODEL_NAME}" >&2
|
||||
exit 1
|
||||
fi
|
||||
return
|
||||
fi
|
||||
if [[ "${#MODEL_ORDER[@]}" -eq 0 ]]; then
|
||||
echo "[launch] ERROR: no supported model weights found under ${MODEL_DIR}." >&2
|
||||
echo "[launch] Run: bash scripts/setup.sh" >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ "${#MODEL_ORDER[@]}" -eq 1 ]]; then
|
||||
MODEL_NAME="${MODEL_ORDER[0]}"
|
||||
echo "[launch] using installed model: $(model_label "$MODEL_NAME")" >&2
|
||||
return
|
||||
fi
|
||||
echo "" >&2
|
||||
echo "[launch] Installed models:" >&2
|
||||
local i=1 model
|
||||
for model in "${MODEL_ORDER[@]}"; do
|
||||
printf " %d) %-16s (%s)\n" "$i" "$(model_label "$model")" "$(engine_hint "${MODEL_ENGINES[$model]}")" >&2
|
||||
i=$((i + 1))
|
||||
done
|
||||
while true; do
|
||||
local pick
|
||||
read_or_interrupt "Choice [1-${#MODEL_ORDER[@]}]: " pick
|
||||
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#MODEL_ORDER[@]} )); then
|
||||
MODEL_NAME="${MODEL_ORDER[$((pick - 1))]}"
|
||||
return
|
||||
fi
|
||||
echo " invalid — pick a number 1-${#MODEL_ORDER[@]}" >&2
|
||||
done
|
||||
}
|
||||
|
||||
GPU_LINES=""
|
||||
CARD_INDICES=()
|
||||
CARD_NAMES=()
|
||||
CARD_MEM_MIB=()
|
||||
CARD_SM=()
|
||||
MIN_VRAM_GB=0
|
||||
MAX_VRAM_GB=0
|
||||
HET_VRAM_MIXED=0
|
||||
SELECTED_GPU_CSV=""
|
||||
SELECTED_VRAM_SUMMARY=""
|
||||
|
||||
gpu_exists() {
|
||||
local want="$1" idx name mem_mib sm
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ "$idx" == "$want" ]] && return 0
|
||||
done <<< "$GPU_LINES"
|
||||
return 1
|
||||
}
|
||||
|
||||
gpu_is_busy() {
|
||||
local want="$1" busy
|
||||
while IFS= read -r busy; do
|
||||
[[ "$busy" == "$want" ]] && return 0
|
||||
done <<< "$(compose_hw_in_use_gpus 2>/dev/null || true)"
|
||||
return 1
|
||||
}
|
||||
|
||||
append_selected_gpu() {
|
||||
local want="$1" idx name mem_mib sm
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
if [[ "$idx" == "$want" ]]; then
|
||||
CARD_INDICES+=("$idx")
|
||||
CARD_NAMES+=("$name")
|
||||
CARD_MEM_MIB+=("$mem_mib")
|
||||
CARD_SM+=("$sm")
|
||||
return 0
|
||||
fi
|
||||
done <<< "$GPU_LINES"
|
||||
return 1
|
||||
}
|
||||
|
||||
select_gpus_from_arg() {
|
||||
local arg="$1"
|
||||
CARD_INDICES=()
|
||||
CARD_NAMES=()
|
||||
CARD_MEM_MIB=()
|
||||
CARD_SM=()
|
||||
local available=() idx name mem_mib sm
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
if ! gpu_is_busy "$idx"; then
|
||||
available+=("$idx")
|
||||
fi
|
||||
done <<< "$GPU_LINES"
|
||||
if [[ "$arg" == "all" ]]; then
|
||||
[[ "${#available[@]}" -gt 0 ]] || { echo "[launch] ERROR: no available NVIDIA GPUs detected." >&2; exit 1; }
|
||||
for idx in "${available[@]}"; do append_selected_gpu "$idx"; done
|
||||
return
|
||||
fi
|
||||
IFS=',' read -ra _launch_gpu_tokens <<< "$arg"
|
||||
for idx in "${_launch_gpu_tokens[@]}"; do
|
||||
idx="$(_compose_meta_trim "$idx")"
|
||||
[[ -z "$idx" ]] && continue
|
||||
gpu_exists "$idx" || { echo "[launch] ERROR: requested GPU ${idx}, but it was not detected." >&2; exit 1; }
|
||||
append_selected_gpu "$idx"
|
||||
done
|
||||
}
|
||||
|
||||
summarize_selected_vram() {
|
||||
local parts=() i gb
|
||||
MIN_VRAM_GB=0
|
||||
MAX_VRAM_GB=0
|
||||
HET_VRAM_MIXED=0
|
||||
for i in "${!CARD_INDICES[@]}"; do
|
||||
gb="$(compose_hw_vram_gb "${CARD_MEM_MIB[$i]}")"
|
||||
parts+=("${gb} GB")
|
||||
if [[ "$MIN_VRAM_GB" -eq 0 || "$gb" -lt "$MIN_VRAM_GB" ]]; then MIN_VRAM_GB="$gb"; fi
|
||||
if [[ "$gb" -gt "$MAX_VRAM_GB" ]]; then MAX_VRAM_GB="$gb"; fi
|
||||
done
|
||||
[[ "$MIN_VRAM_GB" != "$MAX_VRAM_GB" ]] && HET_VRAM_MIXED=1
|
||||
local joined="" part
|
||||
for part in "${parts[@]}"; do
|
||||
[[ -n "$joined" ]] && joined="${joined} + "
|
||||
joined="${joined}${part}"
|
||||
done
|
||||
SELECTED_VRAM_SUMMARY="$joined"
|
||||
printf '%s' "$joined"
|
||||
}
|
||||
|
||||
choose_gpus() {
|
||||
GPU_LINES="$(compose_hw_detect_gpus 2>/dev/null || true)"
|
||||
[[ -n "$GPU_LINES" ]] || { echo "[launch] ERROR: no NVIDIA GPUs detected." >&2; exit 1; }
|
||||
local available=() idx name mem_mib sm state
|
||||
echo "" >&2
|
||||
echo "[launch] Detected GPUs:" >&2
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
state="available"
|
||||
if gpu_is_busy "$idx"; then
|
||||
state="in-use (skipped)"
|
||||
else
|
||||
available+=("$idx")
|
||||
fi
|
||||
printf " GPU %s: %s (%s GB, sm_%s) — %s\n" "$idx" "${name#NVIDIA }" "$(compose_hw_vram_gb "$mem_mib")" "${sm/./}" "$state" >&2
|
||||
done <<< "$GPU_LINES"
|
||||
|
||||
if [[ -n "$GPU_ARG" ]]; then
|
||||
select_gpus_from_arg "$GPU_ARG"
|
||||
elif [[ -n "$CARDS" ]]; then
|
||||
[[ "$CARDS" =~ ^[0-9]+$ && "$CARDS" -ge 1 ]] || { echo "[launch] ERROR: --cards expects a positive integer." >&2; exit 1; }
|
||||
(( CARDS <= ${#available[@]} )) || { echo "[launch] ERROR: --cards ${CARDS} requested, but only ${#available[@]} GPU(s) are available." >&2; exit 1; }
|
||||
local i
|
||||
for ((i = 0; i < CARDS; i++)); do append_selected_gpu "${available[$i]}"; done
|
||||
else
|
||||
case "${#available[@]}" in
|
||||
0) echo "[launch] ERROR: no available NVIDIA GPUs detected." >&2; exit 1 ;;
|
||||
1) select_gpus_from_arg "${available[0]}" ;;
|
||||
2)
|
||||
local pick
|
||||
read_or_interrupt "Use GPU ${available[0]}, GPU ${available[1]}, or both? [both]: " pick
|
||||
pick="${pick:-both}"
|
||||
case "$pick" in
|
||||
both|all) select_gpus_from_arg "${available[0]},${available[1]}" ;;
|
||||
"${available[0]}"|"${available[1]}") select_gpus_from_arg "$pick" ;;
|
||||
*) echo "[launch] ERROR: invalid GPU choice: $pick" >&2; exit 1 ;;
|
||||
esac
|
||||
;;
|
||||
*)
|
||||
local default_csv pick
|
||||
default_csv="$(IFS=','; echo "${available[*]}")"
|
||||
read_or_interrupt "Which GPU(s)? (comma-separated indices, or 'all') [all]: " pick
|
||||
pick="${pick:-all}"
|
||||
[[ "$pick" == "all" ]] && pick="$default_csv"
|
||||
select_gpus_from_arg "$pick"
|
||||
;;
|
||||
esac
|
||||
fi
|
||||
[[ "${#CARD_INDICES[@]}" -gt 0 ]] || { echo "[launch] ERROR: no GPUs selected." >&2; exit 1; }
|
||||
SELECTED_GPU_CSV="$(IFS=','; echo "${CARD_INDICES[*]}")"
|
||||
summarize_selected_vram >/dev/null
|
||||
echo "[launch] selected GPU(s): ${SELECTED_GPU_CSV} (${SELECTED_VRAM_SUMMARY})" >&2
|
||||
}
|
||||
|
||||
valid_tp_values() {
|
||||
case "$1" in
|
||||
qwen3.6-27b) echo "1 2 4" ;;
|
||||
gemma-4-31b) echo "1 2 4 8 16" ;;
|
||||
*) echo "1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
tp_is_valid_for_model() {
|
||||
local model="$1" tp="$2" v
|
||||
for v in $(valid_tp_values "$model"); do [[ "$v" == "$tp" ]] && return 0; done
|
||||
return 1
|
||||
}
|
||||
|
||||
largest_valid_tp_for_cards() {
|
||||
local model="$1" cards="$2" v best=1
|
||||
for v in $(valid_tp_values "$model"); do
|
||||
if (( v <= cards && cards % v == 0 && v > best )); then best="$v"; fi
|
||||
done
|
||||
echo "$best"
|
||||
}
|
||||
|
||||
TP_VALUE=""
|
||||
PP_VALUE=""
|
||||
|
||||
pick_parallelism() {
|
||||
local cards="${#CARD_INDICES[@]}"
|
||||
local tp_set=0 pp_set=0
|
||||
[[ -n "$TP_OVERRIDE" ]] && tp_set=1
|
||||
[[ -n "$PP_OVERRIDE" ]] && pp_set=1
|
||||
[[ -z "$TP_OVERRIDE" || "$TP_OVERRIDE" =~ ^[0-9]+$ ]] || { echo "[launch] ERROR: --tp expects an integer." >&2; exit 1; }
|
||||
[[ -z "$PP_OVERRIDE" || "$PP_OVERRIDE" =~ ^[0-9]+$ ]] || { echo "[launch] ERROR: --pp expects an integer." >&2; exit 1; }
|
||||
case "$PARALLELISM" in auto|tp|pp) ;; *) echo "[launch] ERROR: --parallelism expects auto, tp, or pp." >&2; exit 1 ;; esac
|
||||
|
||||
if (( cards == 1 )); then
|
||||
TP_VALUE="${TP_OVERRIDE:-1}"
|
||||
PP_VALUE="${PP_OVERRIDE:-1}"
|
||||
elif (( tp_set == 1 && pp_set == 1 )); then
|
||||
TP_VALUE="$TP_OVERRIDE"; PP_VALUE="$PP_OVERRIDE"
|
||||
elif (( tp_set == 1 )); then
|
||||
TP_VALUE="$TP_OVERRIDE"
|
||||
(( cards % TP_VALUE == 0 )) || { echo "[launch] ERROR: --tp ${TP_VALUE} does not divide selected GPU count ${cards}." >&2; exit 1; }
|
||||
PP_VALUE=$(( cards / TP_VALUE ))
|
||||
elif (( pp_set == 1 )); then
|
||||
PP_VALUE="$PP_OVERRIDE"
|
||||
(( cards % PP_VALUE == 0 )) || { echo "[launch] ERROR: --pp ${PP_VALUE} does not divide selected GPU count ${cards}." >&2; exit 1; }
|
||||
TP_VALUE=$(( cards / PP_VALUE ))
|
||||
elif [[ "$PARALLELISM" == "pp" ]]; then
|
||||
TP_VALUE=1; PP_VALUE="$cards"
|
||||
elif [[ "$PARALLELISM" == "tp" ]]; then
|
||||
TP_VALUE="$cards"; PP_VALUE=1
|
||||
elif (( HET_VRAM_MIXED == 1 )); then
|
||||
TP_VALUE=1; PP_VALUE="$cards"
|
||||
echo "[launch] ${cards} GPUs selected (${MIN_VRAM_GB} GB + ${MAX_VRAM_GB} GB — heterogeneous)." >&2
|
||||
echo "[launch] Recommended: pipeline parallel PP=${PP_VALUE} to avoid bottlenecking on the smallest card." >&2
|
||||
else
|
||||
TP_VALUE="$(largest_valid_tp_for_cards "$MODEL_NAME" "$cards")"
|
||||
PP_VALUE=$(( cards / TP_VALUE ))
|
||||
fi
|
||||
|
||||
(( TP_VALUE * PP_VALUE == cards )) || { echo "[launch] ERROR: TP × PP must equal selected GPU count (${TP_VALUE} × ${PP_VALUE} != ${cards})." >&2; exit 1; }
|
||||
if ! tp_is_valid_for_model "$MODEL_NAME" "$TP_VALUE"; then
|
||||
echo "[launch] ERROR: $(model_label "$MODEL_NAME") num_kv_heads does not divide TP=${TP_VALUE}." >&2
|
||||
echo "[launch] Valid TP values: $(valid_tp_values "$MODEL_NAME")" >&2
|
||||
exit 1
|
||||
fi
|
||||
if (( cards > 1 && PP_VALUE == 1 )); then
|
||||
echo "[launch] Tensor parallel TP=${TP_VALUE} (PP=1)." >&2
|
||||
elif (( PP_VALUE > 1 )); then
|
||||
echo "[launch] Pipeline parallel PP=${PP_VALUE}, TP=${TP_VALUE}." >&2
|
||||
echo "[launch] WARN: pipeline parallel is experimental on this stack — no benchmarks yet." >&2
|
||||
fi
|
||||
}
|
||||
|
||||
variant_min_gpu_count() {
|
||||
local rel="${LAUNCH_VARIANT_COMPOSE[$1]:-}" value
|
||||
[[ -n "$rel" && -f "${ROOT_DIR}/${rel}" ]] || { echo 1; return; }
|
||||
value="$(compose_meta_get "${ROOT_DIR}/${rel}" requires-min-gpu-count || true)"
|
||||
echo "${value:-1}"
|
||||
}
|
||||
|
||||
variant_min_vram_gb() {
|
||||
local rel="${LAUNCH_VARIANT_COMPOSE[$1]:-}" value
|
||||
[[ -n "$rel" && -f "${ROOT_DIR}/${rel}" ]] || { echo 0; return; }
|
||||
value="$(compose_meta_get "${ROOT_DIR}/${rel}" requires-min-vram-gb || true)"
|
||||
echo "${value:-0}"
|
||||
}
|
||||
|
||||
variant_engine_available() {
|
||||
local variant="$1" engine="${LAUNCH_VARIANT_ENGINE[$variant]:-vllm}"
|
||||
[[ -z "$ENGINE" || "$ENGINE" == "$engine" ]] || return 1
|
||||
model_has_engine "$MODEL_NAME" "$engine"
|
||||
}
|
||||
|
||||
variant_survives_filter() {
|
||||
local variant="$1"
|
||||
[[ "${LAUNCH_VARIANT_MODEL[$variant]:-}" == "$MODEL_NAME" ]] || return 1
|
||||
variant_engine_available "$variant" || return 1
|
||||
local min_gpu min_vram engine
|
||||
min_gpu="$(variant_min_gpu_count "$variant")"
|
||||
min_vram="$(variant_min_vram_gb "$variant")"
|
||||
engine="${LAUNCH_VARIANT_ENGINE[$variant]:-vllm}"
|
||||
[[ "$engine" == "llamacpp" && "${#CARD_INDICES[@]}" -ne 1 ]] && return 1
|
||||
(( min_gpu <= ${#CARD_INDICES[@]} )) || return 1
|
||||
(( min_vram == 0 || min_vram <= MIN_VRAM_GB )) || return 1
|
||||
return 0
|
||||
}
|
||||
|
||||
suggest_default_variant() {
|
||||
local cards="${#CARD_INDICES[@]}"
|
||||
if [[ "$MODEL_NAME" == "qwen3.6-27b" ]]; then
|
||||
if [[ "$ENGINE" == "llamacpp" ]] || { ! model_has_engine "$MODEL_NAME" "vllm" && model_has_engine "$MODEL_NAME" "llamacpp"; }; then
|
||||
echo "llamacpp/default"
|
||||
elif (( cards >= 4 )); then
|
||||
echo "vllm/dual4"
|
||||
elif (( cards >= 2 )); then
|
||||
echo "vllm/dual"
|
||||
else
|
||||
echo "vllm/long-text"
|
||||
fi
|
||||
else
|
||||
if (( cards >= 2 )); then echo "vllm/gemma-mtp"; else echo "vllm/gemma-mtp-tp1"; fi
|
||||
fi
|
||||
}
|
||||
|
||||
no_fit_guidance() {
|
||||
echo "[launch] Selected GPU budget: ${MIN_VRAM_GB} GB minimum per card." >&2
|
||||
echo "" >&2
|
||||
echo "No shipped model variant fits this GPU selection:" >&2
|
||||
echo " Qwen 3.6 27B (INT4): needs >=20 GB" >&2
|
||||
echo " Gemma 4 31B (INT4): needs >=32 GB single-card or 2x24 GB" >&2
|
||||
echo "" >&2
|
||||
echo "Your options:" >&2
|
||||
echo " 1) Combine with another GPU." >&2
|
||||
echo " 2) Try llama.cpp with a smaller GGUF. See docs/SINGLE_CARD.md#sub-16gb" >&2
|
||||
echo " 3) Re-run with --gpus to pick a different card set." >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
gemma_single_24gb_guidance() {
|
||||
echo "[launch] Selected: Gemma 4 31B on GPU ${SELECTED_GPU_CSV} (${MIN_VRAM_GB} GB)." >&2
|
||||
echo "" >&2
|
||||
echo "Gemma 4 31B does not fit on a single ${MIN_VRAM_GB} GB card today." >&2
|
||||
echo "Reason: vLLM's Gemma 4 single-card path needs >=32 GB; use TP=2 on 2x24 GB." >&2
|
||||
echo "" >&2
|
||||
echo "Your options:" >&2
|
||||
echo " 1) Re-run with --gpus 0,1 for TP=2 if you have two 24 GB cards." >&2
|
||||
echo " 2) Use Qwen 3.6 27B for single-card: bash scripts/launch.sh --model qwen3.6-27b" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
json_field() {
|
||||
local field="$1"
|
||||
python3 -c 'import json,sys; data=json.load(sys.stdin); print(data.get(sys.argv[1], ""))' "$field"
|
||||
}
|
||||
|
||||
kv_projection() {
|
||||
local variant="$1"
|
||||
[[ "$SKIP_PROJECTION" -eq 0 ]] || return 0
|
||||
local mapping="${LAUNCH_VARIANT_KVCALC[$variant]:-}"
|
||||
if [[ -z "$mapping" || "$mapping" == "SKIP" ]]; then
|
||||
echo "[launch] KV projection only available for vLLM variants today." >&2
|
||||
return 0
|
||||
fi
|
||||
local kv_model="${mapping%%:*}" kv_compose="${mapping#*:}" kv_json status
|
||||
if kv_json="$("${ROOT_DIR}/tools/kv-calc.py" --model "$kv_model" --compose "$kv_compose" --vram "$MIN_VRAM_GB" --tp "$TP_VALUE" --json 2>&1)"; then
|
||||
status=0
|
||||
else
|
||||
status=$?
|
||||
fi
|
||||
if [[ "$kv_json" != \{* ]]; then
|
||||
echo "[launch] WARN: kv-calc failed for ${variant}: ${kv_json}" >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
local verdict weights kv_pool activation overhead drafter total budget pct
|
||||
verdict="$(json_field verdict <<< "$kv_json")"
|
||||
weights="$(json_field weights_gb <<< "$kv_json")"
|
||||
kv_pool="$(json_field kv_pool_actual_gb <<< "$kv_json")"
|
||||
activation="$(json_field activation_gb <<< "$kv_json")"
|
||||
overhead="$(json_field cudagraph_overhead_gb <<< "$kv_json")"
|
||||
drafter="$(json_field drafter_gb <<< "$kv_json")"
|
||||
total="$(json_field total_gb <<< "$kv_json")"
|
||||
budget="$(json_field budget_gb <<< "$kv_json")"
|
||||
pct="$(json_field pct_of_vram <<< "$kv_json")"
|
||||
|
||||
echo "" >&2
|
||||
echo "[launch] Suggested: ${variant} ($(model_label "$MODEL_NAME") on GPU ${SELECTED_GPU_CSV}, TP=${TP_VALUE} PP=${PP_VALUE})" >&2
|
||||
echo "" >&2
|
||||
echo "VRAM budget — per card (~${MIN_VRAM_GB} GB):" >&2
|
||||
printf " Weights/TP=%s: %.2f GB\n" "$TP_VALUE" "$weights" >&2
|
||||
printf " KV pool: %.2f GB\n" "$kv_pool" >&2
|
||||
printf " Activations: %.2f GB\n" "$activation" >&2
|
||||
printf " Cudagraph + NCCL: %.2f GB\n" "$overhead" >&2
|
||||
if python3 -c 'import sys; sys.exit(0 if float(sys.argv[1]) > 0.01 else 1)' "$drafter"; then
|
||||
printf " Drafter: %.2f GB\n" "$drafter" >&2
|
||||
fi
|
||||
echo " --------------" >&2
|
||||
printf " Predicted peak: %.2f GB %s (%.0f%% of %.2f GB engine budget)\n" "$total" "$verdict" "$pct" "$budget" >&2
|
||||
if (( PP_VALUE > 1 )); then
|
||||
echo " Note: PP is not modelled in projection; real per-card weights should be lower." >&2
|
||||
fi
|
||||
python3 -c 'import json,sys; data=json.load(sys.stdin); [print(" Note: " + n) for n in data.get("notes", [])]' <<< "$kv_json" >&2
|
||||
if [[ "$verdict" == "FAIL" || "$status" -ne 0 ]]; then
|
||||
echo "" >&2
|
||||
echo "[launch] Projection says this variant will not fit. Pick another GPU set or model." >&2
|
||||
exit 2
|
||||
fi
|
||||
}
|
||||
|
||||
# --- wizard ---
|
||||
if [[ -z "$VARIANT" ]]; then
|
||||
echo "" >&2
|
||||
echo "club-3090 launcher — pick model, GPU set, and serving variant." >&2
|
||||
echo "(Use --variant <name> next time to skip the wizard.)" >&2
|
||||
choose_model
|
||||
choose_gpus
|
||||
pick_parallelism
|
||||
if [[ "$MODEL_NAME" == "gemma-4-31b" && "${#CARD_INDICES[@]}" -eq 1 && "$MIN_VRAM_GB" -lt 32 ]]; then
|
||||
gemma_single_24gb_guidance
|
||||
fi
|
||||
|
||||
CANDIDATE_VARIANTS=()
|
||||
for candidate in "${LAUNCH_VARIANT_ORDER[@]}"; do
|
||||
if variant_survives_filter "$candidate"; then
|
||||
CANDIDATE_VARIANTS+=("$candidate")
|
||||
fi
|
||||
done
|
||||
[[ "${#CANDIDATE_VARIANTS[@]}" -gt 0 ]] || no_fit_guidance
|
||||
|
||||
VARIANT="$(suggest_default_variant)"
|
||||
if [[ " ${CANDIDATE_VARIANTS[*]} " != *" ${VARIANT} "* ]]; then
|
||||
VARIANT="${CANDIDATE_VARIANTS[0]}"
|
||||
fi
|
||||
echo "[launch] model: $(model_label "$MODEL_NAME")" >&2
|
||||
if (( HET_VRAM_MIXED == 1 && TP_VALUE > 1 )); then
|
||||
echo "[launch] Note: heterogeneous TP is bottlenecked by the smallest selected card (${MIN_VRAM_GB} GB)." >&2
|
||||
fi
|
||||
if (( PP_VALUE > 1 )) && [[ "$VARIANT" == vllm/* ]]; then
|
||||
echo "[launch] WARN: PP + vLLM drafter/spec-decode paths are experimental on this stack." >&2
|
||||
fi
|
||||
kv_projection "$VARIANT"
|
||||
|
||||
other_variants=()
|
||||
for candidate in "${CANDIDATE_VARIANTS[@]}"; do
|
||||
[[ "$candidate" == "$VARIANT" ]] && continue
|
||||
other_variants+=("$candidate")
|
||||
done
|
||||
if [[ "${#other_variants[@]}" -gt 0 ]]; then
|
||||
echo "[launch] Other variants that fit this selection: ${other_variants[*]}" >&2
|
||||
fi
|
||||
else
|
||||
if [[ -n "$GPU_ARG" ]]; then
|
||||
choose_gpus
|
||||
fi
|
||||
if [[ -n "$TP_OVERRIDE" || -n "$PP_OVERRIDE" ]]; then
|
||||
if [[ "${#CARD_INDICES[@]}" -eq 0 ]]; then
|
||||
TP_VALUE="${TP_OVERRIDE:-1}"
|
||||
PP_VALUE="${PP_OVERRIDE:-1}"
|
||||
else
|
||||
MODEL_NAME="${LAUNCH_VARIANT_MODEL[$VARIANT]:-qwen3.6-27b}"
|
||||
summarize_selected_vram >/dev/null
|
||||
pick_parallelism
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- launch + verify ---
|
||||
echo ""
|
||||
echo "[launch] selected variant: ${VARIANT}"
|
||||
echo ""
|
||||
if [[ -n "$SELECTED_GPU_CSV" ]]; then
|
||||
export CUDA_VISIBLE_DEVICES="$SELECTED_GPU_CSV"
|
||||
export NVIDIA_VISIBLE_DEVICES="$SELECTED_GPU_CSV"
|
||||
fi
|
||||
if [[ -n "$TP_VALUE" ]]; then
|
||||
export TP="$TP_VALUE"
|
||||
fi
|
||||
if [[ -n "$PP_VALUE" ]]; then
|
||||
export PP="$PP_VALUE"
|
||||
fi
|
||||
"$SWITCH" "$VARIANT"
|
||||
|
||||
# Resolve the actual endpoint port + container name the same way switch.sh
|
||||
@@ -213,6 +837,8 @@ declare -A LAUNCH_DEFAULT_PORT=(
|
||||
[vllm/tools-text]=8020
|
||||
[vllm/minimal]=8020
|
||||
[vllm/dual]=8010
|
||||
[vllm/dual4]=8015
|
||||
[vllm/dual4-dflash]=8016
|
||||
[vllm/dual-turbo]=8011
|
||||
[vllm/dual-dflash]=8012
|
||||
[vllm/dual-dflash-noviz]=8013
|
||||
@@ -222,6 +848,7 @@ declare -A LAUNCH_DEFAULT_PORT=(
|
||||
[vllm/dual-nvlink-dflash-noviz]=8019
|
||||
[vllm/gemma-mtp]=8030
|
||||
[vllm/gemma-mtp-tp1]=8031
|
||||
[vllm/gemma-dflash]=8032
|
||||
[llamacpp/default]=8020
|
||||
[llamacpp/concurrent]=8020
|
||||
)
|
||||
@@ -234,6 +861,8 @@ declare -A LAUNCH_DEFAULT_CONTAINER=(
|
||||
[vllm/tools-text]=vllm-qwen36-27b
|
||||
[vllm/minimal]=vllm-qwen36-27b-minimal
|
||||
[vllm/dual]=vllm-qwen36-27b-dual
|
||||
[vllm/dual4]=vllm-qwen36-27b-dual4
|
||||
[vllm/dual4-dflash]=vllm-qwen36-27b-dual4-dflash
|
||||
[vllm/dual-turbo]=vllm-qwen36-27b-dual-turbo
|
||||
[vllm/dual-dflash]=vllm-qwen36-27b-dual-dflash
|
||||
[vllm/dual-dflash-noviz]=vllm-qwen36-27b-dual-dflash-noviz
|
||||
@@ -243,6 +872,7 @@ declare -A LAUNCH_DEFAULT_CONTAINER=(
|
||||
[vllm/dual-nvlink-dflash-noviz]=vllm-qwen36-27b-dual-nvlink-dflash-noviz
|
||||
[vllm/gemma-mtp]=vllm-gemma-4-31b-mtp
|
||||
[vllm/gemma-mtp-tp1]=vllm-gemma-4-31b-mtp-tp1
|
||||
[vllm/gemma-dflash]=vllm-gemma-4-31b-dflash
|
||||
[llamacpp/default]=llama-cpp-qwen36-27b
|
||||
[llamacpp/concurrent]=llama-cpp-qwen36-27b-concurrent
|
||||
)
|
||||
|
||||
@@ -62,3 +62,294 @@ compose_meta_get() {
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
compose_hw_sm_to_int() {
|
||||
local sm="$1"
|
||||
sm="${sm%%+}"
|
||||
sm="${sm//sm_/}"
|
||||
sm="${sm//SM_/}"
|
||||
sm="${sm// /}"
|
||||
[[ -z "$sm" ]] && { echo 0; return; }
|
||||
|
||||
local major minor
|
||||
if [[ "$sm" == *.* ]]; then
|
||||
major="${sm%%.*}"
|
||||
minor="${sm#*.}"
|
||||
else
|
||||
major="$sm"
|
||||
minor="0"
|
||||
fi
|
||||
major="${major//[^0-9]/}"
|
||||
minor="${minor//[^0-9]/}"
|
||||
[[ -z "$major" ]] && major=0
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
if [[ "${#minor}" -eq 1 ]]; then
|
||||
minor=$(( minor * 10 ))
|
||||
else
|
||||
minor="${minor:0:2}"
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
fi
|
||||
echo $(( major * 100 + minor ))
|
||||
}
|
||||
|
||||
compose_hw_vram_gb() {
|
||||
local mib="$1"
|
||||
echo $(( (mib + 1023) / 1024 ))
|
||||
}
|
||||
|
||||
compose_hw_detect_gpus() {
|
||||
if [[ "${_COMPOSE_HW_GPU_CACHE_SET:-0}" == "1" ]]; then
|
||||
[[ -n "${_COMPOSE_HW_GPU_CACHE:-}" ]] || return 1
|
||||
printf '%s\n' "${_COMPOSE_HW_GPU_CACHE}"
|
||||
return 0
|
||||
fi
|
||||
|
||||
if [[ -n "${CLUB3090_FAKE_GPUS:-}" ]]; then
|
||||
local fake parsed_fake="" f_idx f_name f_mem_mib f_sm
|
||||
IFS=',' read -ra _compose_fake_gpus <<< "${CLUB3090_FAKE_GPUS}"
|
||||
for fake in "${_compose_fake_gpus[@]}"; do
|
||||
IFS=':' read -r f_idx f_name f_mem_mib f_sm <<< "$fake"
|
||||
f_idx="$(_compose_meta_trim "${f_idx:-}")"
|
||||
f_name="$(_compose_meta_trim "${f_name:-}")"
|
||||
f_name="${f_name//_/ }"
|
||||
f_mem_mib="$(_compose_meta_trim "${f_mem_mib:-}")"
|
||||
f_sm="$(_compose_meta_trim "${f_sm:-}")"
|
||||
[[ -z "$f_idx" || -z "$f_mem_mib" ]] && continue
|
||||
parsed_fake+="${f_idx}"$'\t'"${f_name}"$'\t'"${f_mem_mib}"$'\t'"${f_sm}"$'\n'
|
||||
done
|
||||
parsed_fake="${parsed_fake%$'\n'}"
|
||||
_COMPOSE_HW_GPU_CACHE_SET=1
|
||||
_COMPOSE_HW_GPU_CACHE="$parsed_fake"
|
||||
[[ -n "$parsed_fake" ]] || return 1
|
||||
printf '%s\n' "$parsed_fake"
|
||||
return 0
|
||||
fi
|
||||
|
||||
command -v nvidia-smi >/dev/null 2>&1 || return 1
|
||||
|
||||
local query idx name mem_mib sm rest
|
||||
query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null)" || return 1
|
||||
[[ -n "$query" ]] || return 1
|
||||
|
||||
local parsed=""
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_compose_meta_trim "$idx")"
|
||||
name="$(_compose_meta_trim "$name")"
|
||||
mem_mib="$(_compose_meta_trim "$mem_mib")"
|
||||
sm="$(_compose_meta_trim "$sm")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
parsed+="${idx}"$'\t'"${name}"$'\t'"${mem_mib}"$'\t'"${sm}"$'\n'
|
||||
done <<< "$query"
|
||||
|
||||
parsed="${parsed%$'\n'}"
|
||||
_COMPOSE_HW_GPU_CACHE_SET=1
|
||||
_COMPOSE_HW_GPU_CACHE="$parsed"
|
||||
[[ -n "$parsed" ]] || return 1
|
||||
printf '%s\n' "$parsed"
|
||||
}
|
||||
|
||||
compose_hw_in_use_gpus() {
|
||||
# Returns GPU indices with non-trivial active compute work. Best-effort:
|
||||
# primary path maps compute-app UUIDs back to GPU indices; memory.used is
|
||||
# the fallback for drivers that do not expose compute app UUIDs.
|
||||
if [[ -n "${CLUB3090_FAKE_BUSY_GPUS:-}" ]]; then
|
||||
printf '%s\n' "${CLUB3090_FAKE_BUSY_GPUS//,/$'\n'}" | sed '/^$/d'
|
||||
return 0
|
||||
fi
|
||||
if [[ -n "${CLUB3090_FAKE_GPUS:-}" ]]; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
command -v nvidia-smi >/dev/null 2>&1 || return 0
|
||||
|
||||
local uuid_query apps line uuid idx
|
||||
uuid_query="$(nvidia-smi --query-gpu=index,uuid --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
apps="$(nvidia-smi --query-compute-apps=gpu_uuid,pid --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
if [[ -n "$uuid_query" && -n "$apps" ]]; then
|
||||
while IFS=',' read -r uuid _pid; do
|
||||
uuid="$(_compose_meta_trim "$uuid")"
|
||||
[[ -z "$uuid" ]] && continue
|
||||
while IFS=',' read -r idx line; do
|
||||
idx="$(_compose_meta_trim "$idx")"
|
||||
line="$(_compose_meta_trim "$line")"
|
||||
if [[ "$line" == "$uuid" ]]; then
|
||||
printf '%s\n' "$idx"
|
||||
fi
|
||||
done <<< "$uuid_query"
|
||||
done <<< "$apps" | sort -u
|
||||
return 0
|
||||
fi
|
||||
|
||||
local mem_used_lines used
|
||||
mem_used_lines="$(nvidia-smi --query-gpu=index,memory.used --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
while IFS=',' read -r idx used; do
|
||||
idx="$(_compose_meta_trim "$idx")"
|
||||
used="$(_compose_meta_trim "$used")"
|
||||
[[ -z "$idx" || -z "$used" ]] && continue
|
||||
if [[ "$used" =~ ^[0-9]+$ ]] && (( used > 1024 )); then
|
||||
printf '%s\n' "$idx"
|
||||
fi
|
||||
done <<< "$mem_used_lines"
|
||||
}
|
||||
|
||||
compose_hw_summary() {
|
||||
local gpu_lines
|
||||
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
|
||||
if [[ -z "$gpu_lines" ]]; then
|
||||
printf 'no NVIDIA GPUs detected'
|
||||
return 0
|
||||
fi
|
||||
|
||||
local count=0 first_name="" first_gb="" mixed=0 idx name mem_mib sm
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
local gb
|
||||
gb="$(compose_hw_vram_gb "$mem_mib")"
|
||||
name="${name#NVIDIA }"
|
||||
name="${name#GeForce }"
|
||||
count=$((count + 1))
|
||||
if [[ -z "$first_name" ]]; then
|
||||
first_name="$name"
|
||||
first_gb="$gb"
|
||||
elif [[ "$name" != "$first_name" || "$gb" != "$first_gb" ]]; then
|
||||
mixed=1
|
||||
fi
|
||||
done <<< "$gpu_lines"
|
||||
|
||||
if (( count == 0 )); then
|
||||
printf 'no NVIDIA GPUs detected'
|
||||
elif (( mixed == 0 )); then
|
||||
if (( count == 1 )); then
|
||||
printf '1× %s, %s GB' "$first_name" "$first_gb"
|
||||
else
|
||||
printf '%d× %s, %s GB each' "$count" "$first_name" "$first_gb"
|
||||
fi
|
||||
else
|
||||
local parts=()
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
name="${name#NVIDIA }"
|
||||
name="${name#GeForce }"
|
||||
parts+=("${name}, $(compose_hw_vram_gb "$mem_mib") GB")
|
||||
done <<< "$gpu_lines"
|
||||
local joined=""
|
||||
for part in "${parts[@]}"; do
|
||||
if [[ -z "$joined" ]]; then
|
||||
joined="$part"
|
||||
else
|
||||
joined="${joined} + ${part}"
|
||||
fi
|
||||
done
|
||||
printf '%s' "$joined"
|
||||
fi
|
||||
}
|
||||
|
||||
compose_hw_requirement_text() {
|
||||
local min_vram_gb="$1"
|
||||
local min_gpu_count="$2"
|
||||
local requires_sm="${3:-}"
|
||||
|
||||
local req
|
||||
if [[ "$min_gpu_count" == "1" ]]; then
|
||||
req="${min_vram_gb} GB+"
|
||||
else
|
||||
req="${min_gpu_count}× ${min_vram_gb} GB"
|
||||
fi
|
||||
if [[ -n "$requires_sm" && "$requires_sm" != "0.0" ]]; then
|
||||
req="${req}, sm_${requires_sm%%+}+"
|
||||
fi
|
||||
printf '%s' "$req"
|
||||
}
|
||||
|
||||
compose_hw_compose_status() {
|
||||
local compose_file="$1"
|
||||
local min_vram_gb min_gpu_count requires_sm
|
||||
|
||||
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
|
||||
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
|
||||
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
|
||||
|
||||
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" ]]; then
|
||||
printf 'unknown|metadata unavailable'
|
||||
return 2
|
||||
fi
|
||||
|
||||
requires_sm="${requires_sm:-0.0}"
|
||||
local required_sm_int
|
||||
required_sm_int="$(compose_hw_sm_to_int "$requires_sm")"
|
||||
|
||||
local gpu_lines
|
||||
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
|
||||
if [[ -z "$gpu_lines" ]]; then
|
||||
printf 'no|no NVIDIA GPUs detected'
|
||||
return 1
|
||||
fi
|
||||
|
||||
local eligible_count=0 idx name mem_mib sm gb sm_int
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
gb="$(compose_hw_vram_gb "$mem_mib")"
|
||||
sm_int="$(compose_hw_sm_to_int "$sm")"
|
||||
if (( gb >= min_vram_gb && sm_int >= required_sm_int )); then
|
||||
eligible_count=$((eligible_count + 1))
|
||||
fi
|
||||
done <<< "$gpu_lines"
|
||||
|
||||
if (( eligible_count >= min_gpu_count )); then
|
||||
printf 'ok|fits your rig'
|
||||
return 0
|
||||
fi
|
||||
|
||||
printf 'no|needs %s (your rig: %s)' \
|
||||
"$(compose_hw_requirement_text "$min_vram_gb" "$min_gpu_count" "$requires_sm")" \
|
||||
"$(compose_hw_summary)"
|
||||
return 1
|
||||
}
|
||||
|
||||
compose_hw_compose_eligible() {
|
||||
local status
|
||||
status="$(compose_hw_compose_status "$1" 2>/dev/null || true)"
|
||||
[[ "$status" == ok\|* ]]
|
||||
}
|
||||
|
||||
compose_hw_model_status() {
|
||||
local repo_root="$1"
|
||||
local model="$2"
|
||||
local candidates=()
|
||||
local friendly_need=""
|
||||
|
||||
case "$model" in
|
||||
qwen3.6-27b)
|
||||
candidates=(
|
||||
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/long-text.yml"
|
||||
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
|
||||
)
|
||||
friendly_need="needs 20 GB+ VRAM (24 GB recommended)"
|
||||
;;
|
||||
gemma-4-31b)
|
||||
candidates=(
|
||||
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
|
||||
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/int8.yml"
|
||||
"${repo_root}/models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
|
||||
)
|
||||
friendly_need="needs 32 GB+ on single card OR 2× 24 GB"
|
||||
;;
|
||||
*)
|
||||
printf 'no|unknown model: %s' "$model"
|
||||
return 1
|
||||
;;
|
||||
esac
|
||||
|
||||
local file status
|
||||
for file in "${candidates[@]}"; do
|
||||
[[ -f "$file" ]] || continue
|
||||
status="$(compose_hw_compose_status "$file" 2>/dev/null || true)"
|
||||
if [[ "$status" == ok\|* ]]; then
|
||||
printf 'ok|fits your rig'
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
|
||||
printf 'no|%s (your rig: %s)' "$friendly_need" "$(compose_hw_summary)"
|
||||
return 1
|
||||
}
|
||||
|
||||
+96
-7
@@ -2,7 +2,8 @@
|
||||
#
|
||||
# Model-aware one-shot setup for club-3090.
|
||||
#
|
||||
# bash scripts/setup.sh <model-name>
|
||||
# bash scripts/setup.sh # interactive model picker in a TTY
|
||||
# bash scripts/setup.sh <model-name> # scripted/CI positional form
|
||||
#
|
||||
# Currently supported:
|
||||
# qwen3.6-27b → Lorbus/Qwen3.6-27B-int4-AutoRound + Genesis patches
|
||||
@@ -37,15 +38,89 @@
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
# ---------- Model dispatch ----------
|
||||
MODEL_NAME="${1:-}"
|
||||
if [[ -z "${MODEL_NAME}" ]]; then
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
usage() {
|
||||
echo "Usage: $0 <model-name>"
|
||||
echo " $0 # interactive model picker in a TTY"
|
||||
echo ""
|
||||
echo "Run with no model name in a normal terminal to open the hardware-aware"
|
||||
echo "model picker. Use the positional form in scripts/CI to skip prompts."
|
||||
echo ""
|
||||
echo "Supported model names:"
|
||||
echo " qwen3.6-27b"
|
||||
echo " gemma-4-31b"
|
||||
exit 1
|
||||
}
|
||||
|
||||
model_label() {
|
||||
case "$1" in
|
||||
qwen3.6-27b) echo "Qwen 3.6 27B" ;;
|
||||
gemma-4-31b) echo "Gemma 4 31B" ;;
|
||||
*) echo "$1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
model_picker_line() {
|
||||
local idx="$1" model="$2" size="$3" status mark reason
|
||||
status="$(compose_hw_model_status "$ROOT_DIR" "$model" 2>/dev/null || true)"
|
||||
reason="${status#*|}"
|
||||
if [[ "$status" == ok\|* ]]; then
|
||||
mark="✓"
|
||||
else
|
||||
mark="✗"
|
||||
fi
|
||||
printf " %s. %-14s (%s) %s %s\n" "$idx" "$(model_label "$model")" "$size" "$mark" "$reason"
|
||||
}
|
||||
|
||||
pick_model_interactive() {
|
||||
# shellcheck source=lib/compose-meta.sh
|
||||
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
|
||||
|
||||
echo "[setup] Which model to download?" >&2
|
||||
echo "" >&2
|
||||
model_picker_line "1" "qwen3.6-27b" "~14 GB AutoRound INT4" >&2
|
||||
model_picker_line "2" "gemma-4-31b" "~21 GB AutoRound INT4 + drafter" >&2
|
||||
echo " 3. Both (~30 GB total) downloads both model families" >&2
|
||||
echo "" >&2
|
||||
while true; do
|
||||
local pick
|
||||
read -rp "Choice [1-3]: " pick
|
||||
case "$pick" in
|
||||
1) echo "qwen3.6-27b"; return ;;
|
||||
2) echo "gemma-4-31b"; return ;;
|
||||
3) echo "both"; return ;;
|
||||
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
|
||||
esac
|
||||
done
|
||||
}
|
||||
|
||||
# ---------- Model dispatch ----------
|
||||
case "${1:-}" in
|
||||
-h|--help)
|
||||
usage
|
||||
exit 0
|
||||
;;
|
||||
esac
|
||||
|
||||
MODEL_NAME="${1:-}"
|
||||
if [[ -z "${MODEL_NAME}" ]]; then
|
||||
if [[ -t 0 && -t 1 ]]; then
|
||||
MODEL_NAME="$(pick_model_interactive)"
|
||||
else
|
||||
usage
|
||||
echo ""
|
||||
echo "(Interactive picker available in a TTY shell. Use the positional form in scripts/CI.)"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ "${MODEL_NAME}" == "both" ]]; then
|
||||
# Resolve MODEL_DIR once in the parent by reusing the normal prompt below,
|
||||
# then recurse through the positional form for each model.
|
||||
SETUP_BOTH_MODE=1
|
||||
MODEL_NAME="qwen3.6-27b"
|
||||
else
|
||||
SETUP_BOTH_MODE=0
|
||||
fi
|
||||
|
||||
# ALWAYS_DRAFT_REPO + ALWAYS_DRAFT_SUBDIR: a drafter that this model REQUIRES
|
||||
@@ -89,8 +164,6 @@ case "${MODEL_NAME}" in
|
||||
;;
|
||||
esac
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
# ---------- MODEL_DIR resolution ----------
|
||||
# Order of precedence:
|
||||
# 1. MODEL_DIR already exported in the calling shell → use as-is
|
||||
@@ -158,6 +231,18 @@ fi
|
||||
|
||||
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
|
||||
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
|
||||
if [[ "${SETUP_BOTH_MODE:-0}" == "1" ]]; then
|
||||
export MODEL_DIR
|
||||
echo "[setup] downloading both supported models into ${MODEL_DIR}"
|
||||
echo ""
|
||||
bash "$0" qwen3.6-27b
|
||||
echo ""
|
||||
bash "$0" gemma-4-31b
|
||||
echo ""
|
||||
echo "[setup] ✓ Both models downloaded."
|
||||
echo "[setup] Next: bash scripts/launch.sh"
|
||||
exit 0
|
||||
fi
|
||||
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
|
||||
|
||||
cd "${ROOT_DIR}"
|
||||
@@ -445,6 +530,7 @@ echo ""
|
||||
# refactored 2026-05-03 to vendor the two files in-repo, fixing #37.)
|
||||
|
||||
# Per-model "next steps" — different composes / served-model-name / port between models.
|
||||
SETUP_MODEL_DISPLAY="$(model_label "${MODEL_NAME}")"
|
||||
case "${MODEL_NAME}" in
|
||||
qwen3.6-27b)
|
||||
SAMPLE_CONTAINER="vllm-qwen36-27b"
|
||||
@@ -468,6 +554,9 @@ case "${MODEL_NAME}" in
|
||||
;;
|
||||
esac
|
||||
|
||||
echo "[setup] ✓ ${SETUP_MODEL_DISPLAY} downloaded."
|
||||
echo "[setup] Next: bash scripts/launch.sh"
|
||||
echo ""
|
||||
echo "Next — single-card vLLM (default):"
|
||||
if [[ "${MODEL_NAME}" == "gemma-4-31b" ]]; then
|
||||
echo " bash scripts/switch.sh vllm/gemma-mtp"
|
||||
|
||||
+87
-16
@@ -5,6 +5,7 @@
|
||||
# Usage:
|
||||
# bash scripts/submit-bench.sh --tag <tag>
|
||||
# bash scripts/submit-bench.sh --tag <tag> --auto-submit
|
||||
# bash scripts/submit-bench.sh --tag <tag> --auto-submit --as-pr
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
@@ -13,6 +14,7 @@ cd "$ROOT_DIR"
|
||||
|
||||
TAG=""
|
||||
AUTO_SUBMIT=0
|
||||
AS_PR=0
|
||||
SECTION_OVERRIDE=""
|
||||
|
||||
usage() {
|
||||
@@ -40,6 +42,10 @@ while [[ $# -gt 0 ]]; do
|
||||
AUTO_SUBMIT=1
|
||||
shift
|
||||
;;
|
||||
--as-pr)
|
||||
AS_PR=1
|
||||
shift
|
||||
;;
|
||||
--section)
|
||||
[[ $# -ge 2 ]] || die "--section requires a value"
|
||||
SECTION_OVERRIDE="$2"
|
||||
@@ -135,6 +141,35 @@ PY
|
||||
fi
|
||||
}
|
||||
|
||||
write_issue_body() {
|
||||
local body_file="$1"
|
||||
local row="$2"
|
||||
local tag="$3"
|
||||
local section="$4"
|
||||
|
||||
# The repo's numbers-from-your-rig issue template is a structured YAML form
|
||||
# with required textarea/dropdown fields. `gh issue create --template` opens
|
||||
# that interactive form shape, which is not useful once submit-bench has
|
||||
# already generated the structured report. Use a direct markdown body instead.
|
||||
{
|
||||
echo "**Compose / section**: \`${section}\`"
|
||||
echo
|
||||
echo "**Rig**:"
|
||||
echo
|
||||
echo '```text'
|
||||
cat "${TAG_DIR}/rig.txt"
|
||||
echo '```'
|
||||
echo
|
||||
echo "**Proposed BENCHMARKS.md row**:"
|
||||
echo
|
||||
echo "$row"
|
||||
echo
|
||||
echo "**Full report**: \`results/rebench/${tag}/REPORT.md\`"
|
||||
echo
|
||||
echo "**Generated row file**: \`results/rebench/${tag}/BENCHMARKS-row.md\`"
|
||||
} > "$body_file"
|
||||
}
|
||||
|
||||
insert_row() {
|
||||
local section="$1"
|
||||
local row="$2"
|
||||
@@ -186,8 +221,24 @@ PY
|
||||
}
|
||||
|
||||
if [[ "$AUTO_SUBMIT" -ne 1 ]]; then
|
||||
echo "Inspect at ${OUTPUT}. To submit:"
|
||||
echo " bash scripts/submit-bench.sh --tag ${TAG} --auto-submit"
|
||||
cat <<EOF
|
||||
Inspect at ${OUTPUT}. Three ways to land it (recommended order):
|
||||
|
||||
1. Issue + maintainer integrates (preferred — vetting before merge):
|
||||
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit
|
||||
(opens an issue via \`gh issue create\`)
|
||||
Or, no-gh-needed:
|
||||
https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml
|
||||
— paste the contents of ${OUTPUT} + ${TAG_DIR}/rig.txt into the body
|
||||
|
||||
2. Direct PR (advanced — for contributors who know the matrix structure):
|
||||
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit --as-pr
|
||||
Note: matrix is hand-curated; direct PRs may get redirected to an
|
||||
issue thread for context-gathering before merge.
|
||||
|
||||
3. Manual edit (zero tools):
|
||||
Paste the row from ${OUTPUT} into BENCHMARKS.md via the GitHub web editor.
|
||||
EOF
|
||||
exit 0
|
||||
fi
|
||||
|
||||
@@ -199,30 +250,50 @@ if [[ -n "$SECTION_OVERRIDE" ]]; then
|
||||
fi
|
||||
fi
|
||||
|
||||
TITLE="bench(matrix): @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
|
||||
PR_TITLE="bench(matrix): @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
|
||||
ISSUE_TITLE="[bench] @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
|
||||
BRANCH_USER="$(printf '%s' "${BENCH_ROW_GITHUB_USER}" | tr -cd '[:alnum:]_.-')"
|
||||
BRANCH_TAG="$(printf '%s' "${TAG}" | tr -cd '[:alnum:]_.-')"
|
||||
BRANCH="bench/${BRANCH_USER}-${BRANCH_TAG}"
|
||||
BODY_FILE="$TAG_DIR/PR-body.md"
|
||||
write_pr_body "$BODY_FILE" "$ROW" "$TAG"
|
||||
if [[ "$AS_PR" -eq 1 ]]; then
|
||||
BODY_FILE="$TAG_DIR/PR-body.md"
|
||||
write_pr_body "$BODY_FILE" "$ROW" "$TAG"
|
||||
else
|
||||
BODY_FILE="$TAG_DIR/ISSUE-body.md"
|
||||
write_issue_body "$BODY_FILE" "$ROW" "$TAG" "$SECTION"
|
||||
fi
|
||||
|
||||
if [[ "${GH_MOCK:-0}" == "1" ]]; then
|
||||
MOCK_LOG="$TAG_DIR/auto-submit-mock.log"
|
||||
{
|
||||
echo "git switch -c ${BRANCH}"
|
||||
echo "insert BENCHMARKS.md row under: ${SECTION}"
|
||||
echo "git commit -m ${TITLE}"
|
||||
echo "git push -u origin ${BRANCH}"
|
||||
echo "gh pr create --title ${TITLE} --body-file ${BODY_FILE}"
|
||||
} > "$MOCK_LOG"
|
||||
log "GH_MOCK=1 — wrote mocked auto-submit commands: ${MOCK_LOG}"
|
||||
log "PR title: ${TITLE}"
|
||||
if [[ "$AS_PR" -eq 1 ]]; then
|
||||
{
|
||||
echo "git switch -c ${BRANCH}"
|
||||
echo "insert BENCHMARKS.md row under: ${SECTION}"
|
||||
echo "git commit -m ${PR_TITLE}"
|
||||
echo "git push -u origin ${BRANCH}"
|
||||
echo "gh pr create --title ${PR_TITLE} --body-file ${BODY_FILE}"
|
||||
} > "$MOCK_LOG"
|
||||
log "GH_MOCK=1 — wrote mocked PR auto-submit commands: ${MOCK_LOG}"
|
||||
log "PR title: ${PR_TITLE}"
|
||||
else
|
||||
{
|
||||
echo "gh issue create --title ${ISSUE_TITLE} --body-file ${BODY_FILE} --label bench-contribution"
|
||||
} > "$MOCK_LOG"
|
||||
log "GH_MOCK=1 — wrote mocked issue auto-submit command: ${MOCK_LOG}"
|
||||
log "Issue title: ${ISSUE_TITLE}"
|
||||
fi
|
||||
exit 0
|
||||
fi
|
||||
|
||||
command -v gh >/dev/null 2>&1 || die "'gh' not found. Install GitHub CLI or submit manually."
|
||||
gh auth status >/dev/null 2>&1 || die "not authed with gh. Run: gh auth login"
|
||||
|
||||
if [[ "$AS_PR" -ne 1 ]]; then
|
||||
ISSUE_URL="$(gh issue create --title "$ISSUE_TITLE" --body-file "$BODY_FILE" --label bench-contribution)"
|
||||
log "Opened issue: ${ISSUE_URL}"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if ! git diff --quiet -- BENCHMARKS.md; then
|
||||
die "BENCHMARKS.md already has local edits; commit/stash them before --auto-submit"
|
||||
fi
|
||||
@@ -235,7 +306,7 @@ else
|
||||
fi
|
||||
insert_row "$SECTION" "$ROW"
|
||||
git add BENCHMARKS.md
|
||||
git commit -m "$TITLE"
|
||||
git commit -m "$PR_TITLE"
|
||||
git push -u origin "$BRANCH"
|
||||
PR_URL="$(gh pr create --title "$TITLE" --body-file "$BODY_FILE")"
|
||||
PR_URL="$(gh pr create --title "$PR_TITLE" --body-file "$BODY_FILE")"
|
||||
log "Opened PR: ${PR_URL}"
|
||||
|
||||
Executable
+215
@@ -0,0 +1,215 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||
TMP_DIR="$(mktemp -d)"
|
||||
ORIG_PATH="$PATH"
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
assert_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" != *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
assert_not_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" == *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output not to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
make_mock_tools() {
|
||||
mkdir -p "${TMP_DIR}/bin"
|
||||
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK_NVIDIA_SMI'
|
||||
#!/usr/bin/env bash
|
||||
case "$*" in
|
||||
*"--query-gpu=index,name,memory.total,compute_cap"*)
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
|
||||
;;
|
||||
"-L")
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}" \
|
||||
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
|
||||
;;
|
||||
*)
|
||||
echo "unexpected nvidia-smi invocation: $*" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
MOCK_NVIDIA_SMI
|
||||
chmod +x "${TMP_DIR}/bin/nvidia-smi"
|
||||
|
||||
cat > "${TMP_DIR}/switch-mock" <<'MOCK_SWITCH'
|
||||
#!/usr/bin/env bash
|
||||
echo "SWITCHED $* CUDA=${CUDA_VISIBLE_DEVICES:-} NVD=${NVIDIA_VISIBLE_DEVICES:-} TP=${TP:-} PP=${PP:-}"
|
||||
MOCK_SWITCH
|
||||
chmod +x "${TMP_DIR}/switch-mock"
|
||||
|
||||
export PATH="${TMP_DIR}/bin:${ORIG_PATH}"
|
||||
}
|
||||
|
||||
set_rig() {
|
||||
export MOCK_GPU_QUERY="$1"
|
||||
}
|
||||
|
||||
model_status() {
|
||||
local model="$1"
|
||||
(
|
||||
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
|
||||
compose_hw_model_status "$ROOT_DIR" "$model"
|
||||
)
|
||||
}
|
||||
|
||||
assert_model_status() {
|
||||
local model="$1"
|
||||
local expected_prefix="$2"
|
||||
local expected_text="${3:-}"
|
||||
local status
|
||||
status="$(model_status "$model" || true)"
|
||||
if [[ "$status" != "${expected_prefix}"* ]]; then
|
||||
echo "ASSERTION FAILED: ${model} status expected prefix '${expected_prefix}', got '${status}'" >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ -n "$expected_text" ]]; then
|
||||
assert_contains "$status" "$expected_text"
|
||||
fi
|
||||
}
|
||||
|
||||
make_mock_tools
|
||||
|
||||
# Matched 2x3090: Qwen and Gemma both have a viable compose.
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "ok|fits your rig"
|
||||
|
||||
# Single 24 GB Ampere: Qwen fits; Gemma needs either 32 GB+ single-card or 2x24 GB.
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
|
||||
assert_contains "$(model_status "gemma-4-31b")" "1× RTX 3090, 24 GB"
|
||||
|
||||
# Single 16 GB: neither shipped model has a viable compose.
|
||||
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9'
|
||||
assert_model_status "qwen3.6-27b" "no|" "needs 20 GB+ VRAM"
|
||||
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
|
||||
|
||||
# Heterogeneous 16 + 24 GB: Qwen can run on the 24 GB card; Gemma dual cannot.
|
||||
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "no|" "RTX 4060 Ti, 16 GB + RTX 3090, 24 GB"
|
||||
|
||||
# 32 GB+ modern card: Gemma's single-card compose is eligible.
|
||||
set_rig $'0, NVIDIA GeForce RTX 5090, 32768, 12.0'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "ok|fits your rig"
|
||||
|
||||
# Non-TTY no-arg setup fails fast with usage rather than hanging.
|
||||
if out="$(echo | bash "${ROOT_DIR}/scripts/setup.sh" 2>&1)"; then
|
||||
echo "ASSERTION FAILED: non-TTY no-arg setup unexpectedly succeeded" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "Usage:"
|
||||
assert_contains "$out" "Interactive picker available in a TTY shell"
|
||||
|
||||
# Positional setup path remains non-interactive and reaches the existing flow.
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
out="$(MODEL_DIR="${TMP_DIR}/models" PREFLIGHT_DISK_GB=0 SKIP_GENESIS=1 SKIP_MODEL=1 bash "${ROOT_DIR}/scripts/setup.sh" qwen3.6-27b 2>&1)"
|
||||
assert_not_contains "$out" "Which model to download?"
|
||||
assert_contains "$out" "[model] SKIP_MODEL=1"
|
||||
|
||||
# The launch wizard now picks model -> GPU set -> parallelism. Scripted flags
|
||||
# skip prompts, select the expected variant, and export GPU / TP / PP envs.
|
||||
mkdir -p "${TMP_DIR}/models/qwen3.6-27b-autoround-int4" \
|
||||
"${TMP_DIR}/models/gemma-4-31b-autoround-int4"
|
||||
FAKE_8X3090='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6,4:RTX_3090:24576:8.6,5:RTX_3090:24576:8.6,6:RTX_3090:24576:8.6,7:RTX_3090:24576:8.6'
|
||||
|
||||
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6' \
|
||||
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
|
||||
--no-preflight --no-verify --model qwen3.6-27b --gpus 0 --no-projection 2>&1)"
|
||||
assert_contains "$out" "[launch] selected variant: vllm/long-text"
|
||||
assert_contains "$out" "SWITCHED vllm/long-text CUDA=0 NVD=0 TP=1 PP=1"
|
||||
|
||||
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6' \
|
||||
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
|
||||
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1 --no-projection 2>&1)"
|
||||
assert_contains "$out" "[launch] Tensor parallel TP=2"
|
||||
assert_contains "$out" "SWITCHED vllm/dual CUDA=0,1 NVD=0,1 TP=2 PP=1"
|
||||
selected_count="$(grep -c "\[launch\] selected variant:" <<< "$out" || true)"
|
||||
if [[ "$selected_count" != "1" ]]; then
|
||||
echo "ASSERTION FAILED: expected one selected-variant line, got ${selected_count}" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6' \
|
||||
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
|
||||
--no-preflight --no-verify --model gemma-4-31b --gpus 0 --no-projection 2>&1)"; then
|
||||
echo "ASSERTION FAILED: Gemma single-24GB launch unexpectedly succeeded" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "Gemma 4 31B does not fit on a single 24 GB card today"
|
||||
|
||||
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6,4:RTX_3090:24576:8.6,5:RTX_3090:24576:8.6' \
|
||||
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
|
||||
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1,2,3,4,5 --tp 6 --no-projection 2>&1)"; then
|
||||
echo "ASSERTION FAILED: invalid Qwen TP=6 unexpectedly succeeded" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "Valid TP values: 1 2 4"
|
||||
|
||||
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS="${FAKE_8X3090}" \
|
||||
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
|
||||
--no-preflight --no-verify --model gemma-4-31b --gpus 0,1,2,3,4,5,6,7 --tp 8 2>&1)"
|
||||
assert_contains "$out" "[launch] Tensor parallel TP=8"
|
||||
assert_contains "$out" "[launch] Suggested: vllm/gemma-mtp"
|
||||
assert_contains "$out" "VRAM budget — per card"
|
||||
assert_contains "$out" "Note: TP > 4 predictions are extrapolated"
|
||||
assert_not_contains "$out" "KV projection skipped"
|
||||
assert_contains "$out" "SWITCHED vllm/gemma-mtp CUDA=0,1,2,3,4,5,6,7 NVD=0,1,2,3,4,5,6,7 TP=8 PP=1"
|
||||
|
||||
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS="${FAKE_8X3090}" \
|
||||
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
|
||||
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1,2,3,4,5,6,7 --tp 8 --no-projection 2>&1)"; then
|
||||
echo "ASSERTION FAILED: invalid Qwen TP=8 unexpectedly succeeded" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "num_kv_heads does not divide TP=8"
|
||||
assert_contains "$out" "Valid TP values: 1 2 4"
|
||||
|
||||
# TTY-backed no-arg setup supports the cosmetic but real "Both" choice by
|
||||
# dispatching through the positional path for both model families.
|
||||
if ! command -v script >/dev/null 2>&1; then
|
||||
echo "ASSERTION FAILED: util-linux 'script' is required for TTY picker coverage" >&2
|
||||
exit 1
|
||||
fi
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
export MODEL_DIR="${TMP_DIR}/models"
|
||||
export PREFLIGHT_DISK_GB=0
|
||||
export SKIP_GENESIS=1
|
||||
export SKIP_MODEL=1
|
||||
out="$(printf '3\n' | script -qec "bash '${ROOT_DIR}/scripts/setup.sh'" /dev/null 2>&1)"
|
||||
assert_contains "$out" "[setup] Which model to download?"
|
||||
assert_contains "$out" "Both"
|
||||
assert_contains "$out" "[setup] downloading both supported models"
|
||||
skip_count="$(grep -c "\[model\] SKIP_MODEL=1" <<< "$out" || true)"
|
||||
if [[ "$skip_count" != "2" ]]; then
|
||||
echo "ASSERTION FAILED: expected Both choice to dispatch two model setup runs, got ${skip_count}" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "test-setup-picker: ok"
|
||||
@@ -59,11 +59,15 @@ done
|
||||
tag="qwen-int8-pth-n4-2026-05-10"
|
||||
rm -f "results/rebench/${tag}/BENCHMARKS-row.md" \
|
||||
"results/rebench/${tag}/PR-body.md" \
|
||||
"results/rebench/${tag}/ISSUE-body.md" \
|
||||
"results/rebench/${tag}/auto-submit-mock.log"
|
||||
|
||||
out="$(bash scripts/submit-bench.sh --tag "$tag")"
|
||||
assert_contains "$out" "Generated BENCHMARKS row for section: Dual-card (2× RTX 3090, TP=2)"
|
||||
assert_contains "$out" "Wrote: results/rebench/${tag}/BENCHMARKS-row.md"
|
||||
assert_contains "$out" "1. Issue + maintainer integrates"
|
||||
assert_contains "$out" "2. Direct PR"
|
||||
assert_contains "$out" "3. Manual edit"
|
||||
test -s "results/rebench/${tag}/BENCHMARKS-row.md"
|
||||
|
||||
if out="$(bash scripts/submit-bench.sh --tag does-not-exist 2>&1)"; then
|
||||
@@ -73,14 +77,21 @@ fi
|
||||
assert_contains "$out" "tag dir not found: results/rebench/does-not-exist"
|
||||
|
||||
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit)"
|
||||
assert_contains "$out" "PR title: bench(matrix): @octocat ${tag}"
|
||||
assert_contains "$out" "Issue title: [bench] @octocat ${tag}"
|
||||
test -s "results/rebench/${tag}/auto-submit-mock.log"
|
||||
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh pr create --title bench(matrix): @octocat ${tag}"
|
||||
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh issue create --title [bench] @octocat ${tag}"
|
||||
test -s "results/rebench/${tag}/ISSUE-body.md"
|
||||
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "results/rebench/${tag}/REPORT.md"
|
||||
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "Proposed BENCHMARKS.md row"
|
||||
|
||||
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit --as-pr)"
|
||||
assert_contains "$out" "PR title: bench(matrix): @octocat ${tag}"
|
||||
test -s "results/rebench/${tag}/PR-body.md"
|
||||
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh pr create --title bench(matrix): @octocat ${tag}"
|
||||
assert_contains "$(cat "results/rebench/${tag}/PR-body.md")" "results/rebench/${tag}/REPORT.md"
|
||||
|
||||
tmp_bin="$(mktemp -d)"
|
||||
trap 'rm -rf "$tmp_bin"; rm -f "results/rebench/${tag}/BENCHMARKS-row.md" "results/rebench/${tag}/PR-body.md" "results/rebench/${tag}/auto-submit-mock.log"' EXIT
|
||||
trap 'rm -rf "$tmp_bin"; rm -f "results/rebench/${tag}/BENCHMARKS-row.md" "results/rebench/${tag}/PR-body.md" "results/rebench/${tag}/ISSUE-body.md" "results/rebench/${tag}/auto-submit-mock.log"' EXIT
|
||||
cat > "${tmp_bin}/gh" <<'MOCK_GH'
|
||||
#!/usr/bin/env bash
|
||||
if [[ "$1" == "auth" && "$2" == "status" ]]; then
|
||||
|
||||
+559
-225
@@ -1,53 +1,69 @@
|
||||
#!/usr/bin/env python3
|
||||
"""kv-calc.py — predict per-card VRAM budget for a Qwen3.6-27B vLLM compose.
|
||||
#!/bin/sh
|
||||
''':'
|
||||
exec python3 "$0" "$@"
|
||||
':'''
|
||||
from __future__ import annotations
|
||||
|
||||
__doc__ = """kv-calc.py — predict per-card VRAM budget for vLLM composes.
|
||||
|
||||
Predicts (per card, after TP split):
|
||||
- Model weights
|
||||
- KV pool (attention layers — 16 full_attention layers with GQA)
|
||||
- Activation peak (DeltaNet GDN forward — 48 linear_attention layers)
|
||||
- KV pool (attention layers only — recurrent / SSM states show up in activation)
|
||||
- Activation peak (model-specific: Qwen GDN forward, Gemma SWA + dense MLP)
|
||||
- Cudagraph + workspace overhead
|
||||
- Drafter overhead (MTP / DFlash)
|
||||
- Total vs available VRAM
|
||||
- Verdict: PASS / TIGHT / FAIL
|
||||
|
||||
Two models modelled:
|
||||
- Qwen 3.6 27B (DeltaNet hybrid: 16 full_attention + 48 GDN)
|
||||
- Gemma 4 31B (SWA + dense MLP: 10 full_attention + 50 sliding_attention)
|
||||
|
||||
vLLM rate-limits KV pool to fit available budget; this predictor models that
|
||||
capping behavior. When the requested KV pool exceeds what fits, the verdict
|
||||
is TIGHT (vLLM will cap pool — effective concurrency reduced) not FAIL.
|
||||
|
||||
Anchored to:
|
||||
- PerfMamba (arxiv 2511.22849) — block-wise state materialization scaling
|
||||
- PerfMamba (arxiv 2511.22849) — Qwen GDN block-wise state materialization scaling
|
||||
https://arxiv.org/html/2511.22849
|
||||
- TurboQuant (arxiv 2504.19874, ICLR 2026) — TQ3 byte savings
|
||||
https://arxiv.org/abs/2504.19874
|
||||
- PagedAttention (arxiv 2309.06180) — KV pool layout
|
||||
https://arxiv.org/abs/2309.06180
|
||||
|
||||
Calibrated against measured BENCHMARKS.md rows. Coefficients in
|
||||
GDN_ACTIVATION_COEF reflect club-3090's empirical findings on top of
|
||||
PerfMamba's O(γ·D·N·L) scaling — the absolute scaling is well-defined,
|
||||
but the per-token coefficient depends on fla.ops.chunk implementation
|
||||
details that the published literature doesn't enumerate. See
|
||||
docs/KV_MATH.md for the derivation + calibration trace.
|
||||
Calibrated against measured BENCHMARKS.md rows per model. Coefficients reflect
|
||||
club-3090's empirical findings. See docs/KV_MATH.md for the derivation +
|
||||
calibration trace.
|
||||
|
||||
Usage:
|
||||
bash tools/kv-calc.py --compose dual-turbo --vram 24
|
||||
bash tools/kv-calc.py --compose dual-turbo --kv-format fp8_e5m2 --vram 20
|
||||
bash tools/kv-calc.py --max-ctx 180000 --kv-format turboquant_3bit_nc --tp 1 --vram 24
|
||||
bash tools/kv-calc.py --calibration # show calibration vs measured points
|
||||
bash tools/kv-calc.py --compose dual-turbo --vram 24 # Qwen (default model)
|
||||
bash tools/kv-calc.py --model gemma-4-31b --compose gemma-dual-int8 --vram 24
|
||||
bash tools/kv-calc.py --model gemma-4-31b --solve-max-ctx --kv-format int8_per_token_head --tp 2 --vram 24
|
||||
bash tools/kv-calc.py --calibration # both models, grouped per-model
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# Model specs
|
||||
# =============================================================================
|
||||
|
||||
# ---- Qwen3.6-27B AutoRound INT4 — from config.json text_config ----
|
||||
QWEN36_27B = {
|
||||
"model_id": "qwen3.6-27b-autoround",
|
||||
"model_id": "qwen3.6-27b",
|
||||
"model_family": "qwen3-next-hybrid",
|
||||
"hidden_size": 5120,
|
||||
"num_hidden_layers": 64,
|
||||
"num_gdn_layers": 48, # linear_attention layers (Gated DeltaNet)
|
||||
"num_attn_layers": 16, # full_attention layers
|
||||
"num_attn_heads": 24,
|
||||
"num_kv_heads": 4, # GQA
|
||||
"valid_tp": [1, 2, 4],
|
||||
"head_dim_attn": 256, # attention head dim
|
||||
"linear_num_v_heads": 48, # GDN value heads
|
||||
"linear_num_k_heads": 16, # GDN key heads (GQA-style at the GDN level too)
|
||||
@@ -57,128 +73,283 @@ QWEN36_27B = {
|
||||
"weights_total_gb": 17.5, # AutoRound INT4 storage on disk
|
||||
"mamba_state_bytes": 4, # mamba_ssm_dtype=float32
|
||||
"chunk_size": 256, # fla.ops.chunk default
|
||||
"max_ctx_supported": 262144,
|
||||
"attention_k_eq_v": False, # Qwen stores K and V independently
|
||||
}
|
||||
|
||||
# ---- KV format bytes per stored token element ----
|
||||
# (one element = one head dim of one head; K and V counted separately)
|
||||
# Source: vLLM/HF docs + TurboQuant paper.
|
||||
# ---- Gemma 4 31B — from config.json text_config ----
|
||||
# Layer pattern: [sliding_attention × 5, full_attention × 1] × 10
|
||||
# = 50 sliding-attention + 10 full-attention.
|
||||
GEMMA4_31B = {
|
||||
"model_id": "gemma-4-31b",
|
||||
"model_family": "gemma4-swa-dense",
|
||||
"hidden_size": 5376,
|
||||
"intermediate_size": 21504,
|
||||
"num_hidden_layers": 60,
|
||||
"num_full_attn_layers": 10, # full_attention (growing KV, head_dim=512)
|
||||
"num_sliding_attn_layers": 50, # sliding_attention (fixed window, head_dim=256)
|
||||
"num_attn_heads": 32,
|
||||
"num_kv_heads": 16, # GQA 2:1
|
||||
"valid_tp": [1, 2, 4, 8, 16],
|
||||
"head_dim_sliding": 256, # sliding_attention head dim
|
||||
"global_head_dim": 512, # full_attention head dim (asymmetric)
|
||||
"sliding_window": 1024,
|
||||
"weights_int4_gb": 18.0, # AutoRound INT4 on disk
|
||||
"weights_awq_gb": 17.0, # cyankiwi AWQ-4bit (lower on-card due to AWQ packing)
|
||||
"weights_bf16_gb": 58.0, # unquantized — does not fit on 24 GB
|
||||
"max_ctx_supported": 262144,
|
||||
"drafter_mtp_gb": 0.97, # google/gemma-4-31b-it-assistant (FP16)
|
||||
"drafter_dflash_gb": 2.9, # z-lab/gemma-4-31b-it-dflash
|
||||
"attention_k_eq_v": True, # vLLM allocator exploits K==V tying (per_token uses ×1 not ×2)
|
||||
}
|
||||
|
||||
MODEL_SPECS = {
|
||||
"qwen3.6-27b": QWEN36_27B,
|
||||
"gemma-4-31b": GEMMA4_31B,
|
||||
}
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# KV format bytes-per-element
|
||||
# =============================================================================
|
||||
# (one element = one head dim of one head; K and V counted separately for
|
||||
# models where attention_k_eq_v=False. For Gemma the K==V tying halves this
|
||||
# at the formula level — see kv_pool_per_card_bytes.)
|
||||
# Source: vLLM/HF docs + TurboQuant paper + PR #40391 (INT8 per-token-head).
|
||||
|
||||
KV_FORMAT_BYTES = {
|
||||
"fp16": 2.0,
|
||||
"bf16": 2.0,
|
||||
"fp8_e5m2": 1.0,
|
||||
"fp8_e4m3": 1.0,
|
||||
"q4_0": 0.5 + 0.0625, # 4-bit + per-group scale
|
||||
"k8v4": 0.75, # avg of K=int8 V=int4
|
||||
"turboquant_3bit_nc": 0.375 + 0.05, # 3 bits + small QJL overhead
|
||||
"fp16": 2.0,
|
||||
"bf16": 2.0,
|
||||
"fp8_e5m2": 1.0,
|
||||
"fp8_e4m3": 1.0,
|
||||
"int8_per_token_head": 1.01, # 1.0 int8 + per-token-head fp16 scale (~1% amortized)
|
||||
"q4_0": 0.5 + 0.0625, # 4-bit + per-group scale
|
||||
"k8v4": 0.75, # avg of K=int8 V=int4
|
||||
"turboquant_3bit_nc": 0.375 + 0.05, # 3 bits + small QJL overhead
|
||||
}
|
||||
|
||||
# ---- GDN activation-peak per-layer per-token coefficient (bytes) ----
|
||||
# Calibrated empirically against measured BENCHMARKS rows. The PerfMamba
|
||||
# O(γ·D·N·L) scaling sets the *form*; this coefficient captures
|
||||
# fla.ops.chunk implementation details + per-KV-format dequant overhead.
|
||||
# TQ3's coefficient is ~25% larger than fp8 (matches efschu's 20 GB
|
||||
# Cliff 2 finding — TQ3 dequant adds activation pressure).
|
||||
GDN_ACTIVATION_COEF = {
|
||||
|
||||
# =============================================================================
|
||||
# Per-model activation coefficients
|
||||
# =============================================================================
|
||||
|
||||
# ---- Qwen GDN activation-peak per-layer per-token coefficient (bytes) ----
|
||||
# Calibrated empirically against measured BENCHMARKS rows. PerfMamba's
|
||||
# O(γ·D·N·L) scaling sets the form; this captures fla.ops.chunk
|
||||
# implementation details + KV-format-dependent dequant overhead.
|
||||
QWEN_GDN_ACTIVATION_COEF = {
|
||||
"fp16": 135,
|
||||
"bf16": 135,
|
||||
"fp8_e5m2": 130,
|
||||
"fp8_e4m3": 130,
|
||||
"int8_per_token_head": 130,
|
||||
"q4_0": 155,
|
||||
"k8v4": 155,
|
||||
"turboquant_3bit_nc": 165,
|
||||
}
|
||||
|
||||
# ---- Compose presets ----
|
||||
# ---- Gemma activation peak (mostly constant in ctx) ----
|
||||
# Unlike Qwen GDN, Gemma's activation peak comes from dense MLP forward +
|
||||
# SWA windowed-attention prefill, both bounded by chunked-prefill chunk_size.
|
||||
# Result: roughly CONSTANT in max_ctx. Calibrated as a per-TP base (GB) plus
|
||||
# a small per-token term to capture residual ctx scaling.
|
||||
GEMMA_ACTIVATION_CONST_GB = 1.5 # per card at TP=1 — calibrated, ~scales as 1/TP
|
||||
GEMMA_ACTIVATION_PER_TOKEN_BYTES = 8 # tiny ctx scaling term to keep solver well-behaved
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# Compose presets (per-model)
|
||||
# =============================================================================
|
||||
# Pulled from each compose's CLI args. Update if a compose changes.
|
||||
# Compose IDs are namespaced: Qwen uses bare names (back-compat); Gemma uses gemma-* prefix.
|
||||
|
||||
COMPOSES = {
|
||||
"minimal": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.90, "mtp": False},
|
||||
"long-text": {"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.93, "mtp": True},
|
||||
"long-text-no-mtp":{"max_ctx": 200000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": False},
|
||||
"long-vision": {"max_ctx": 145000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
|
||||
"bounded-thinking":{"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
|
||||
"tools-text": {"max_ctx": 75000, "max_num_seqs": 1, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.97, "mtp": True},
|
||||
"dual": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 2, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
|
||||
"dual-turbo": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 2, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
|
||||
"dual-dflash": {"max_ctx": 185000, "max_num_seqs": 1, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
|
||||
"dual-dflash-noviz":{"max_ctx": 200000,"max_num_seqs": 2, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
|
||||
"dual4": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 4, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
|
||||
"dual4-dflash": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 4, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
|
||||
"qwen3.6-27b": {
|
||||
"minimal": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.90, "mtp": False},
|
||||
"long-text": {"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.93, "mtp": True},
|
||||
"long-text-no-mtp": {"max_ctx": 200000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": False},
|
||||
"long-vision": {"max_ctx": 145000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
|
||||
"bounded-thinking": {"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
|
||||
"tools-text": {"max_ctx": 75000, "max_num_seqs": 1, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.97, "mtp": True},
|
||||
"dual": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 2, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
|
||||
"dual-turbo": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 2, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
|
||||
"dual-dflash": {"max_ctx": 185000, "max_num_seqs": 1, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
|
||||
"dual-dflash-noviz":{"max_ctx": 200000, "max_num_seqs": 2, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
|
||||
"dual4": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 4, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
|
||||
"dual4-dflash": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 4, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
|
||||
},
|
||||
"gemma-4-31b": {
|
||||
# dual/docker-compose.yml — default MTP, 32K BF16 KV, max-num-seqs=4
|
||||
"gemma-dual": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 2, "kv_format": "bf16", "mem_util": 0.92, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
|
||||
# dual/int8.yml default — 98K + INT8 PTH KV + 4 seqs
|
||||
"gemma-dual-int8": {"max_ctx": 98304, "max_num_seqs": 4, "tp": 2, "kv_format": "int8_per_token_head", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
|
||||
# dual/int8.yml long — 262K + INT8 PTH KV + seqs=1 (model native max)
|
||||
"gemma-dual-int8-262k": {"max_ctx": 262144, "max_num_seqs": 1, "tp": 2, "kv_format": "int8_per_token_head", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
|
||||
# dual/bf16.yml — long-ctx BF16 weights + BF16 KV (auto)
|
||||
"gemma-dual-bf16": {"max_ctx": 200000, "max_num_seqs": 1, "tp": 2, "kv_format": "bf16", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
|
||||
# dual/int8-tq3.yml — TQ3 KV alt
|
||||
"gemma-dual-int8-tq3": {"max_ctx": 98304, "max_num_seqs": 4, "tp": 2, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
|
||||
# dual/dflash.yml — DFlash drafter + 32K BF16
|
||||
"gemma-dual-dflash": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 2, "kv_format": "bf16", "mem_util": 0.92, "weights_variant": "int4", "drafter_gb": 2.9, "mtp": False, "dflash_draft_gb": 2.9},
|
||||
# dual/dflash-int8.yml — DFlash + INT8 PTH KV (PR #42102 unblock)
|
||||
"gemma-dual-dflash-int8": {"max_ctx": 65536, "max_num_seqs": 2, "tp": 2, "kv_format": "int8_per_token_head", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 2.9, "mtp": False, "dflash_draft_gb": 2.9},
|
||||
# dual/awq.yml — cyankiwi AWQ weights
|
||||
"gemma-dual-awq": {"max_ctx": 65536, "max_num_seqs": 4, "tp": 2, "kv_format": "bf16", "mem_util": 0.85, "weights_variant": "awq", "drafter_gb": 0.97, "mtp": True},
|
||||
# single/docker-compose.yml — 32 GB+ required; Ampere consumer OOMs at boot
|
||||
"gemma-single": {"max_ctx": 8192, "max_num_seqs": 256, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
|
||||
},
|
||||
}
|
||||
|
||||
# ---- Calibration: measured BENCHMARKS rows (peak per-card VRAM during bench) ----
|
||||
CALIBRATION = [
|
||||
# (compose, vram_gb, measured_peak_gb, source_url)
|
||||
("dual", 24, 23.6, "BENCHMARKS.md#qwen36-27b dual.yml @noonghunna 2026-04-29"),
|
||||
("dual-turbo", 24, 19.8, "BENCHMARKS.md#qwen36-27b dual-turbo.yml @noonghunna 2026-04-29"),
|
||||
("dual-dflash", 24, 23.6, "BENCHMARKS.md#qwen36-27b dual-dflash.yml @noonghunna 2026-04-29"),
|
||||
("dual-dflash-noviz",24, 23.8, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @noonghunna 2026-04-29"),
|
||||
("dual4", 24, 23.5, "BENCHMARKS.md#qwen36-27b dual4.yml @whamp 2026-05-03"),
|
||||
("dual4-dflash", 24, 22.0, "BENCHMARKS.md#qwen36-27b dual4-dflash.yml @whamp 2026-05-03"),
|
||||
("dual-dflash-noviz",24, 21.8, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @snoby 2026-05-04 (2× 4090, ctx=180K)"),
|
||||
("long-text", 24, 22.3, "BENCHMARKS.md#qwen36-27b long-text.yml @noonghunna 2026-04-30"),
|
||||
("long-vision", 24, 23.0, "BENCHMARKS.md#qwen36-27b long-vision.yml @noonghunna 2026-04-30"),
|
||||
("bounded-thinking", 24, 21.7, "BENCHMARKS.md#qwen36-27b bounded-thinking.yml @noonghunna 2026-05-04"),
|
||||
("minimal", 24, 22.4, "BENCHMARKS.md#qwen36-27b minimal.yml @noonghunna 2026-05-03 (mem-util 0.95, max-ctx 65536)"),
|
||||
]
|
||||
|
||||
# =============================================================================
|
||||
# Calibration: measured BENCHMARKS rows (peak per-card VRAM during bench)
|
||||
# =============================================================================
|
||||
|
||||
CALIBRATION = {
|
||||
"qwen3.6-27b": [
|
||||
# (compose, vram_gb, measured_peak_gb, ctx_override_or_none, source_url)
|
||||
("dual", 24, 23.6, None, "BENCHMARKS.md#qwen36-27b dual.yml @noonghunna 2026-04-29"),
|
||||
("dual-turbo", 24, 19.8, None, "BENCHMARKS.md#qwen36-27b dual-turbo.yml @noonghunna 2026-04-29"),
|
||||
("dual-dflash", 24, 23.6, None, "BENCHMARKS.md#qwen36-27b dual-dflash.yml @noonghunna 2026-04-29"),
|
||||
("dual-dflash-noviz", 24, 23.8, None, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @noonghunna 2026-04-29"),
|
||||
("dual4", 24, 23.5, None, "BENCHMARKS.md#qwen36-27b dual4.yml @whamp 2026-05-03"),
|
||||
("dual4-dflash", 24, 22.0, None, "BENCHMARKS.md#qwen36-27b dual4-dflash.yml @whamp 2026-05-03"),
|
||||
("dual-dflash-noviz", 24, 21.8, 180000, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @snoby 2026-05-04 (2× 4090, ctx=180K)"),
|
||||
("long-text", 24, 22.3, None, "BENCHMARKS.md#qwen36-27b long-text.yml @noonghunna 2026-04-30"),
|
||||
("long-vision", 24, 23.0, None, "BENCHMARKS.md#qwen36-27b long-vision.yml @noonghunna 2026-04-30"),
|
||||
("bounded-thinking", 24, 21.7, None, "BENCHMARKS.md#qwen36-27b bounded-thinking.yml @noonghunna 2026-05-04"),
|
||||
("minimal", 24, 22.4, 65536, "BENCHMARKS.md#qwen36-27b minimal.yml @noonghunna 2026-05-03 (mem-util 0.95, max-ctx 65536)"),
|
||||
],
|
||||
"gemma-4-31b": [
|
||||
# All TP=2 dual configs on 2× 3090 24 GB.
|
||||
("gemma-dual-int8", 24, 22.2, None, "BENCHMARKS.md#gemma-4-31b dual/int8.yml @noonghunna 2026-05-08 (98K + max-num-seqs=4)"),
|
||||
("gemma-dual-int8-262k", 24, 22.1, None, "BENCHMARKS.md#gemma-4-31b dual/int8.yml @noonghunna 2026-05-08 (262K + max-num-seqs=1)"),
|
||||
("gemma-dual-dflash", 24, 22.7, None, "BENCHMARKS.md#gemma-4-31b dual/dflash.yml @noonghunna 2026-05-06 (n=7)"),
|
||||
("gemma-dual-dflash", 24, 22.3, None, "BENCHMARKS.md#gemma-4-31b dual/dflash.yml @noonghunna 2026-05-08 (rebench)"),
|
||||
("gemma-dual-awq", 24, 19.8, None, "BENCHMARKS.md#gemma-4-31b dual/awq.yml @noonghunna 2026-05-08 (AWQ-4bit, mem-util 0.85)"),
|
||||
("gemma-dual", 24, 22.5, None, "BENCHMARKS.md#gemma-4-31b dual/docker-compose.yml @noonghunna matched-config rebench 2026-05-09"),
|
||||
("gemma-dual-int8", 24, 22.5, 262144, "BENCHMARKS.md#gemma-4-31b dual/int8.yml matched-config rebench 2026-05-09 (262K seqs=2)"),
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# Prediction
|
||||
# =============================================================================
|
||||
|
||||
@dataclass
|
||||
class Prediction:
|
||||
model: str
|
||||
weights_gb: float
|
||||
kv_pool_gb: float
|
||||
kv_pool_requested_gb: float
|
||||
kv_pool_actual_gb: float # capped at available budget (vLLM behavior)
|
||||
kv_pool_sliding_fixed_gb: float # Gemma sliding-window fixed term (0 for Qwen)
|
||||
activation_gb: float
|
||||
cudagraph_overhead_gb: float
|
||||
dflash_draft_gb: float
|
||||
drafter_gb: float
|
||||
total_gb: float
|
||||
vram_gb: float
|
||||
budget_gb: float
|
||||
pct_of_vram: float
|
||||
verdict: str
|
||||
notes: list[str]
|
||||
|
||||
|
||||
def _weights_per_card_gb(spec, tp, weights_variant="default"):
|
||||
"""Return per-card weights footprint in GB after TP split."""
|
||||
if spec["model_family"] == "qwen3-next-hybrid":
|
||||
return spec["weights_total_gb"] / tp
|
||||
elif spec["model_family"] == "gemma4-swa-dense":
|
||||
if weights_variant == "awq":
|
||||
return spec["weights_awq_gb"] / tp
|
||||
elif weights_variant == "bf16":
|
||||
return spec["weights_bf16_gb"] / tp
|
||||
else: # int4 default
|
||||
return spec["weights_int4_gb"] / tp
|
||||
raise ValueError(f"Unknown model_family: {spec['model_family']}")
|
||||
|
||||
|
||||
def kv_pool_per_card_bytes(spec, kv_format, max_ctx, max_num_seqs, tp, mtp_n=0):
|
||||
"""Per-card KV pool bytes for the attention layers.
|
||||
GDN layers have a fixed-size recurrent state (not seq-len-dependent KV
|
||||
cache), so they don't contribute here — they show up in activation_peak.
|
||||
"""Per-card KV pool bytes (growing portion only).
|
||||
|
||||
Standard formula:
|
||||
per_token_bytes = num_attn_layers × num_kv_heads × head_dim × 2 (K+V) × bytes
|
||||
pool = per_token × max_ctx × max_num_seqs / TP
|
||||
Returns a tuple (growing_per_card_bytes, sliding_fixed_per_card_bytes).
|
||||
Sliding term is zero for models without sliding-window layers.
|
||||
|
||||
MTP n>0 adds n extra cached tokens per request for draft hidden states.
|
||||
For Qwen 3.6 (DeltaNet hybrid):
|
||||
Only the 16 full_attention layers grow KV. GDN layers have a fixed-size
|
||||
recurrent state (not seq-len-dependent), so they show up in activation.
|
||||
K and V stored independently → ×2 factor.
|
||||
|
||||
For Gemma 4 (SWA + dense MLP):
|
||||
Only the 10 full_attention layers grow KV (at global_head_dim=512).
|
||||
The 50 sliding_attention layers hold a FIXED window of 1024 tokens
|
||||
(constant in ctx, contributes a separate small term).
|
||||
K==V tying IS exploited by vLLM's allocator → ×1 factor (calibrated
|
||||
against BENCHMARKS data; see docs/KV_MATH.md).
|
||||
"""
|
||||
bytes_per_kv_elem = KV_FORMAT_BYTES[kv_format]
|
||||
per_token = (
|
||||
spec["num_attn_layers"]
|
||||
* spec["num_kv_heads"]
|
||||
* spec["head_dim_attn"]
|
||||
* 2 # K + V
|
||||
* bytes_per_kv_elem
|
||||
)
|
||||
# KV heads split across TP ranks (heads must divide rank)
|
||||
per_card_per_token = per_token / tp
|
||||
effective_ctx = max_ctx + mtp_n * 32 # MTP draft caches a few extra positions
|
||||
return per_card_per_token * effective_ctx * max_num_seqs
|
||||
bpe = KV_FORMAT_BYTES[kv_format]
|
||||
|
||||
if spec["model_family"] == "qwen3-next-hybrid":
|
||||
# K and V stored independently
|
||||
per_token = (
|
||||
spec["num_attn_layers"]
|
||||
* spec["num_kv_heads"]
|
||||
* spec["head_dim_attn"]
|
||||
* 2 # K + V
|
||||
* bpe
|
||||
)
|
||||
effective_ctx = max_ctx + mtp_n * 32
|
||||
growing = (per_token / tp) * effective_ctx * max_num_seqs
|
||||
return growing, 0.0
|
||||
|
||||
elif spec["model_family"] == "gemma4-swa-dense":
|
||||
# K==V tied → ×1 storage
|
||||
per_token_growing = (
|
||||
spec["num_full_attn_layers"]
|
||||
* spec["num_kv_heads"]
|
||||
* spec["global_head_dim"]
|
||||
* 1 # K==V tied; vLLM stores once
|
||||
* bpe
|
||||
)
|
||||
# No MTP draft-token bump on Gemma — drafter is a separate model
|
||||
growing = (per_token_growing / tp) * max_ctx * max_num_seqs
|
||||
|
||||
# Sliding-window fixed term — 50 layers × window × head_dim × 1 × bpe
|
||||
sliding_fixed_total = (
|
||||
spec["num_sliding_attn_layers"]
|
||||
* spec["num_kv_heads"]
|
||||
* spec["head_dim_sliding"]
|
||||
* 1 # K==V tied here too
|
||||
* bpe
|
||||
* spec["sliding_window"]
|
||||
)
|
||||
sliding_per_card = sliding_fixed_total / tp
|
||||
return growing, sliding_per_card
|
||||
|
||||
raise ValueError(f"Unknown model_family: {spec['model_family']}")
|
||||
|
||||
|
||||
def gdn_activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp):
|
||||
"""Per-card peak activation during DeltaNet GDN forward.
|
||||
def activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp):
|
||||
"""Per-card peak activation during prefill forward.
|
||||
|
||||
Theoretical scaling (PerfMamba arxiv 2511.22849): O(γ·D·N·L) per layer.
|
||||
Empirical fit: linear in seq_len, with KV-format-dependent coefficient.
|
||||
For Qwen 3.6 (DeltaNet GDN): linear in seq_len, KV-format-dependent
|
||||
coefficient (PerfMamba O(γ·D·N·L) form, fla.ops.chunk implementation
|
||||
details calibrated empirically).
|
||||
|
||||
Each GDN layer's `chunk_gated_delta_rule_fwd` materializes a block-wise
|
||||
state tensor sized roughly (B, NT, H, V, K) × bytes. For Qwen3.6-27B:
|
||||
NT = ceil(seq_len / 256), H = 16 K-heads, V = K = 128 dim, fp32 state.
|
||||
|
||||
The actual implementation has tiling/streaming that PerfMamba's pure
|
||||
formula doesn't capture. The coefficient here is calibrated against
|
||||
measured BENCHMARKS peaks (see CALIBRATION + tools/kv-calc.py
|
||||
--calibration).
|
||||
For Gemma 4 (dense MLP + SWA): mostly CONSTANT in seq_len because chunked
|
||||
prefill bounds the MLP intermediate. Small per-token residual to keep
|
||||
the solver smooth.
|
||||
"""
|
||||
coef = GDN_ACTIVATION_COEF[kv_format]
|
||||
total = coef * spec["num_gdn_layers"] * max_ctx
|
||||
return total / tp
|
||||
if spec["model_family"] == "qwen3-next-hybrid":
|
||||
coef = QWEN_GDN_ACTIVATION_COEF[kv_format]
|
||||
return (coef * spec["num_gdn_layers"] * max_ctx) / tp
|
||||
|
||||
elif spec["model_family"] == "gemma4-swa-dense":
|
||||
const_bytes = GEMMA_ACTIVATION_CONST_GB * 1e9
|
||||
per_token = GEMMA_ACTIVATION_PER_TOKEN_BYTES * max_ctx
|
||||
return (const_bytes + per_token) / tp
|
||||
|
||||
raise ValueError(f"Unknown model_family: {spec['model_family']}")
|
||||
|
||||
|
||||
def cudagraph_overhead_gb(mem_util, tp):
|
||||
@@ -191,6 +362,16 @@ def cudagraph_overhead_gb(mem_util, tp):
|
||||
return base + tp_bump
|
||||
|
||||
|
||||
def _validate_tp_for_spec(spec, tp):
|
||||
valid_tp = spec.get("valid_tp")
|
||||
if valid_tp and tp not in valid_tp:
|
||||
raise ValueError(
|
||||
f"TP={tp} invalid for {spec['model_id']} "
|
||||
f"(num_kv_heads={spec['num_kv_heads']} cannot be divided across TP cleanly). "
|
||||
f"Valid TP values: {valid_tp}"
|
||||
)
|
||||
|
||||
|
||||
def predict(
|
||||
spec=QWEN36_27B,
|
||||
kv_format="fp8_e5m2",
|
||||
@@ -200,51 +381,98 @@ def predict(
|
||||
mem_util=0.95,
|
||||
vram_gb=24,
|
||||
dflash_draft_gb=0.0,
|
||||
drafter_gb=0.0,
|
||||
mtp=False,
|
||||
weights_variant="default",
|
||||
) -> Prediction:
|
||||
weights_gb = spec["weights_total_gb"] / tp
|
||||
kv_pool_gb = kv_pool_per_card_bytes(spec, kv_format, max_ctx, max_num_seqs, tp,
|
||||
mtp_n=3 if mtp else 0) / 1e9
|
||||
activation_gb = gdn_activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp) / 1e9
|
||||
overhead_gb = cudagraph_overhead_gb(mem_util, tp)
|
||||
dflash_gb = dflash_draft_gb if tp == 1 else dflash_draft_gb / tp # draft splits with TP
|
||||
total_gb = weights_gb + kv_pool_gb + activation_gb + overhead_gb + dflash_gb
|
||||
"""Predict per-card VRAM usage.
|
||||
|
||||
# The verdict compares DEMAND (this prediction) against the engine's
|
||||
# available budget = mem_util × vram_gb. vLLM will refuse to boot if
|
||||
# demand exceeds this. (Measured peak during bench is a different number:
|
||||
# it's roughly mem_util × VRAM because vLLM inflates the KV pool to fill
|
||||
# the budget — see docs/KV_MATH.md.)
|
||||
#
|
||||
# Calibrated error band: ±1.5 GB on the breakdown, ±2 GB on total.
|
||||
# This is a directional estimator, not a precise predictor.
|
||||
vLLM caps KV pool to (budget - fixed_components), so the prediction
|
||||
reflects what actually gets allocated. When requested > available,
|
||||
verdict is TIGHT with a note about effective concurrency reduction.
|
||||
|
||||
Args:
|
||||
drafter_gb: total drafter weight (MTP / DFlash) — split by TP.
|
||||
dflash_draft_gb: legacy alias — folded into drafter_gb if set.
|
||||
"""
|
||||
_validate_tp_for_spec(spec, tp)
|
||||
|
||||
weights_gb = _weights_per_card_gb(spec, tp, weights_variant)
|
||||
|
||||
growing_b, sliding_b = kv_pool_per_card_bytes(
|
||||
spec, kv_format, max_ctx, max_num_seqs, tp,
|
||||
mtp_n=3 if mtp else 0,
|
||||
)
|
||||
kv_pool_requested_gb = growing_b / 1e9
|
||||
kv_pool_sliding_fixed_gb = sliding_b / 1e9
|
||||
|
||||
activation_gb = activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp) / 1e9
|
||||
overhead_gb = cudagraph_overhead_gb(mem_util, tp)
|
||||
|
||||
# Drafter: prefer drafter_gb; fall back to legacy dflash_draft_gb.
|
||||
drafter_total = drafter_gb if drafter_gb > 0 else dflash_draft_gb
|
||||
drafter_per_card = drafter_total / tp if tp > 1 else drafter_total
|
||||
|
||||
fixed_gb = weights_gb + activation_gb + overhead_gb + drafter_per_card + kv_pool_sliding_fixed_gb
|
||||
budget_gb = mem_util * vram_gb
|
||||
pct = 100 * total_gb / budget_gb
|
||||
available_for_kv = max(0.0, budget_gb - fixed_gb)
|
||||
|
||||
# vLLM caps the KV pool to fit available budget (PagedAttention allocator).
|
||||
kv_pool_actual_gb = min(kv_pool_requested_gb, available_for_kv)
|
||||
|
||||
total_gb = fixed_gb + kv_pool_actual_gb
|
||||
pct = 100 * total_gb / budget_gb if budget_gb > 0 else 999.0
|
||||
|
||||
notes = []
|
||||
# Generous verdict bands matching the ±1.5 GB error.
|
||||
if pct < 88:
|
||||
verdict = "PASS"
|
||||
elif pct < 108:
|
||||
verdict = "TIGHT"
|
||||
notes.append(f"demand within ±1.5 GB error of engine budget ({budget_gb:.1f} GB at mem_util={mem_util}) — likely boots, may need a small mem_util bump if pre-check refuses")
|
||||
else:
|
||||
verdict = "FAIL"
|
||||
notes.append(f"demand {pct:.0f}% of engine budget ({budget_gb:.1f} GB at mem_util={mem_util}) — pre-check will refuse; raise mem_util, lower max_ctx/max_num_seqs, or swap KV format")
|
||||
|
||||
if kv_format == "turboquant_3bit_nc" and vram_gb < 24:
|
||||
# Verdict logic:
|
||||
# - FAIL: fixed components alone exceed budget (no room even for minimum KV).
|
||||
# - TIGHT: requested KV pool exceeds available — vLLM will cap, effective
|
||||
# concurrency reduced (BOOT OK, but `--max-num-seqs` may not be
|
||||
# honored at full max_ctx).
|
||||
# - PASS: requested KV fits with room to spare.
|
||||
MIN_KV_GB = 1.0 # vLLM needs at least ~1 GB for paged-attention blocks
|
||||
if available_for_kv < MIN_KV_GB:
|
||||
verdict = "FAIL"
|
||||
notes.append(
|
||||
f"fixed components ({fixed_gb:.1f} GB) leave only {available_for_kv:.1f} GB for KV pool "
|
||||
f"(need ≥{MIN_KV_GB:.1f} GB minimum); vLLM pre-check will refuse — "
|
||||
f"lower max_ctx, drop a drafter, or raise mem_util"
|
||||
)
|
||||
elif kv_pool_requested_gb > available_for_kv * 1.05:
|
||||
verdict = "TIGHT"
|
||||
notes.append(
|
||||
f"requested KV pool ({kv_pool_requested_gb:.1f} GB) > available ({available_for_kv:.1f} GB) — "
|
||||
f"vLLM will cap to {available_for_kv:.1f} GB; effective concurrency may be lower than "
|
||||
f"--max-num-seqs={max_num_seqs} at full max_ctx={max_ctx:,}"
|
||||
)
|
||||
else:
|
||||
verdict = "PASS"
|
||||
|
||||
# Model-specific advisory notes (preserved from v1)
|
||||
if kv_format == "turboquant_3bit_nc" and vram_gb < 24 and spec["model_family"] == "qwen3-next-hybrid":
|
||||
notes.append("⚠ TQ3 KV on <24 GB cards: consider --kv-format fp8_e5m2 (see docs/HARDWARE.md, #47)")
|
||||
if max_ctx > 50000 and tp == 1 and kv_format != "fp16":
|
||||
if max_ctx > 50000 and tp == 1 and spec["model_family"] == "qwen3-next-hybrid" and kv_format != "fp16":
|
||||
notes.append("⚠ single-card vLLM at >50K single-prompt: Cliff 2 territory (DeltaNet GDN forward); see docs/CLIFFS.md")
|
||||
if spec["model_family"] == "gemma4-swa-dense" and kv_format == "fp8_e4m3":
|
||||
notes.append("⚠ fp8_e4m3 on Ampere (sm_86): Triton `fp8e4nv` kernel unsupported; use int8_per_token_head instead (PR #40391 via #42102)")
|
||||
if spec["model_family"] == "gemma4-swa-dense" and tp == 1 and vram_gb < 32:
|
||||
notes.append("⚠ Gemma 4 31B TP=1 needs ≥32 GB VRAM; 24 GB Ampere boot-OOMs (model weights + drafter + min KV)")
|
||||
if tp > 4:
|
||||
notes.append("TP > 4 predictions are extrapolated; report deltas via scripts/report.sh --bench")
|
||||
|
||||
return Prediction(
|
||||
model=spec["model_id"],
|
||||
weights_gb=weights_gb,
|
||||
kv_pool_gb=kv_pool_gb,
|
||||
kv_pool_requested_gb=kv_pool_requested_gb,
|
||||
kv_pool_actual_gb=kv_pool_actual_gb,
|
||||
kv_pool_sliding_fixed_gb=kv_pool_sliding_fixed_gb,
|
||||
activation_gb=activation_gb,
|
||||
cudagraph_overhead_gb=overhead_gb,
|
||||
dflash_draft_gb=dflash_gb,
|
||||
drafter_gb=drafter_per_card,
|
||||
total_gb=total_gb,
|
||||
vram_gb=vram_gb,
|
||||
budget_gb=budget_gb,
|
||||
pct_of_vram=pct,
|
||||
verdict=verdict,
|
||||
notes=notes,
|
||||
@@ -256,86 +484,135 @@ def fmt_prediction(p: Prediction, header: str = "") -> str:
|
||||
if header:
|
||||
lines.append(header)
|
||||
lines.append("-" * len(header))
|
||||
lines.append(f" Model weights: {p.weights_gb:>6.2f} GB / card")
|
||||
lines.append(f" KV pool (attention): {p.kv_pool_gb:>6.2f} GB / card")
|
||||
lines.append(f" Activation peak (GDN): {p.activation_gb:>6.2f} GB / card")
|
||||
lines.append(f" Model: {p.model}")
|
||||
lines.append(f" Weights: {p.weights_gb:>6.2f} GB / card")
|
||||
if p.kv_pool_sliding_fixed_gb > 0.01:
|
||||
lines.append(f" KV pool — sliding fixed: {p.kv_pool_sliding_fixed_gb:>6.2f} GB / card (constant, doesn't grow with ctx)")
|
||||
if abs(p.kv_pool_requested_gb - p.kv_pool_actual_gb) > 0.05:
|
||||
lines.append(f" KV pool — growing (req): {p.kv_pool_requested_gb:>6.2f} GB / card (requested)")
|
||||
lines.append(f" KV pool — growing (cap): {p.kv_pool_actual_gb:>6.2f} GB / card (vLLM-capped to fit)")
|
||||
else:
|
||||
lines.append(f" KV pool — growing: {p.kv_pool_actual_gb:>6.2f} GB / card")
|
||||
lines.append(f" Activation peak: {p.activation_gb:>6.2f} GB / card")
|
||||
lines.append(f" Cudagraph + workspace: {p.cudagraph_overhead_gb:>6.2f} GB / card")
|
||||
if p.dflash_draft_gb > 0:
|
||||
lines.append(f" DFlash draft model: {p.dflash_draft_gb:>6.2f} GB / card")
|
||||
if p.drafter_gb > 0:
|
||||
lines.append(f" Drafter (MTP / DFlash): {p.drafter_gb:>6.2f} GB / card")
|
||||
lines.append(f" ─────────────────────────────────────")
|
||||
lines.append(f" Predicted demand total: {p.total_gb:>6.2f} GB / card ({p.pct_of_vram:.0f}% of engine budget)")
|
||||
lines.append(f" Predicted total: {p.total_gb:>6.2f} GB / card ({p.pct_of_vram:.0f}% of {p.budget_gb:.1f} GB engine budget)")
|
||||
lines.append(f" Verdict: {p.verdict}")
|
||||
lines.append(f" (Note: measured peak during bench will be higher — vLLM fills the")
|
||||
lines.append(f" remaining budget with KV pool inflation. Demand is the *lower bound*")
|
||||
lines.append(f" that determines whether engine pre-check accepts the config.)")
|
||||
for note in p.notes:
|
||||
lines.append(f" Note: {note}")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def run_calibration():
|
||||
print("=" * 88)
|
||||
print("Calibration — DEMAND prediction vs measured peak + engine budget per-card VRAM")
|
||||
print("=" * 88)
|
||||
print()
|
||||
print(" Demand is what the engine NEEDS (lower-bound). Engine budget is what vLLM")
|
||||
print(" ALLOCATES (≈ mem_util × VRAM). Measured peak is what nvidia-smi shows during")
|
||||
print(" bench. vLLM inflates the KV pool to fill the budget, so peak ≈ budget.")
|
||||
print(" The verdict — PASS / TIGHT / FAIL — is correct iff demand < budget.")
|
||||
print()
|
||||
print(f" {'compose':<22s} {'demand':>9s} {'budget':>9s} {'measured':>10s} {'verdict':>8s}")
|
||||
print(f" {'─'*21:<22s} {'─'*8:>9s} {'─'*8:>9s} {'─'*9:>10s} {'─'*7:>8s}")
|
||||
# =============================================================================
|
||||
# Calibration runner
|
||||
# =============================================================================
|
||||
|
||||
def _resolve_compose_for_predict(model_key, compose_id, vram, ctx_override=None):
|
||||
"""Resolve a compose preset to predict() kwargs, applying optional ctx override."""
|
||||
spec = MODEL_SPECS[model_key]
|
||||
cfg = COMPOSES[model_key][compose_id]
|
||||
max_ctx = ctx_override if ctx_override is not None else cfg["max_ctx"]
|
||||
|
||||
kwargs = dict(
|
||||
spec=spec,
|
||||
kv_format=cfg["kv_format"],
|
||||
max_ctx=max_ctx,
|
||||
max_num_seqs=cfg["max_num_seqs"],
|
||||
tp=cfg["tp"],
|
||||
mem_util=cfg["mem_util"],
|
||||
vram_gb=vram,
|
||||
mtp=cfg.get("mtp", False),
|
||||
weights_variant=cfg.get("weights_variant", "default"),
|
||||
drafter_gb=cfg.get("drafter_gb", 0.0),
|
||||
dflash_draft_gb=cfg.get("dflash_draft_gb", 0.0),
|
||||
)
|
||||
return kwargs
|
||||
|
||||
|
||||
def _calibration_block(model_key: str) -> tuple[int, int]:
|
||||
"""Print calibration table for one model. Returns (correct, total)."""
|
||||
rows = CALIBRATION.get(model_key, [])
|
||||
if not rows:
|
||||
return 0, 0
|
||||
|
||||
spec = MODEL_SPECS[model_key]
|
||||
print(f"== {spec['model_id']} ==")
|
||||
print(f" {'compose':<26s} {'predicted':>10s} {'budget':>9s} {'measured':>10s} {'verdict':>8s}")
|
||||
print(f" {'─'*25:<26s} {'─'*9:>10s} {'─'*8:>9s} {'─'*9:>10s} {'─'*7:>8s}")
|
||||
|
||||
correct = 0
|
||||
for compose, vram, measured, _src in CALIBRATION:
|
||||
cfg = COMPOSES[compose]
|
||||
max_ctx = cfg["max_ctx"]
|
||||
if compose == "dual-dflash-noviz" and abs(measured - 21.8) < 0.05:
|
||||
max_ctx = 180000 # snoby's 4090 row
|
||||
p = predict(
|
||||
kv_format=cfg["kv_format"],
|
||||
max_ctx=max_ctx,
|
||||
max_num_seqs=cfg["max_num_seqs"],
|
||||
tp=cfg["tp"],
|
||||
mem_util=cfg["mem_util"],
|
||||
vram_gb=vram,
|
||||
dflash_draft_gb=cfg.get("dflash_draft_gb", 0.0),
|
||||
mtp=cfg.get("mtp", False),
|
||||
)
|
||||
budget = cfg["mem_util"] * vram
|
||||
# Verdict is "correct" if predicted PASS/TIGHT and measured < vram (boot OK),
|
||||
# or predicted FAIL and measured > vram (boot would fail).
|
||||
verdict_correct = "✓" if p.verdict in ("PASS", "TIGHT") and measured < vram else ("⨯" if p.verdict == "FAIL" and measured < vram else "✓")
|
||||
if verdict_correct == "✓":
|
||||
for row in rows:
|
||||
compose, vram, measured, ctx_override, _src = row
|
||||
kwargs = _resolve_compose_for_predict(model_key, compose, vram, ctx_override)
|
||||
p = predict(**kwargs)
|
||||
# Verdict is "correct" if (PASS/TIGHT and measured fits) or (FAIL and would OOM).
|
||||
# We don't have negative (FAIL) data points in BENCHMARKS — every row booted —
|
||||
# so verdict_correct simplifies to: PASS/TIGHT and measured < vram.
|
||||
if p.verdict in ("PASS", "TIGHT") and measured < vram:
|
||||
mark = "✓"
|
||||
correct += 1
|
||||
print(f" {compose:<22s} {p.total_gb:>7.2f} GB {budget:>7.2f} GB {measured:>8.2f} GB {p.verdict:>7s} {verdict_correct}")
|
||||
elif p.verdict == "FAIL" and measured >= vram:
|
||||
mark = "✓"
|
||||
correct += 1
|
||||
else:
|
||||
mark = "⨯"
|
||||
compose_disp = compose if ctx_override is None else f"{compose}@{ctx_override//1024}K"
|
||||
print(f" {compose_disp:<26s} {p.total_gb:>8.2f} GB {p.budget_gb:>7.2f} GB {measured:>8.2f} GB {p.verdict:>7s} {mark}")
|
||||
|
||||
n = len(CALIBRATION)
|
||||
print()
|
||||
print(f" Verdict accuracy: {correct}/{n} ({100*correct/n:.0f}%)")
|
||||
print(f" Verdict accuracy: {correct}/{len(rows)} ({100*correct/len(rows):.0f}%)")
|
||||
print()
|
||||
print(" The DEMAND number is the calculator's output; users use it to plan.")
|
||||
print(" The MEASURED column is for sanity: every passing config should show")
|
||||
print(" measured < VRAM (else boot would have failed). If the calculator says")
|
||||
print(" PASS but measured > VRAM × mem_util, there's hidden overhead the model")
|
||||
print(" doesn't capture — file an issue with your `bash scripts/report.sh --bench`")
|
||||
print(" output and we'll re-calibrate.")
|
||||
return correct, len(rows)
|
||||
|
||||
|
||||
def solve_max_ctx(spec, kv_format, max_num_seqs, tp, mem_util, vram_gb, dflash_draft_gb, mtp):
|
||||
"""Binary search for the largest max_ctx that keeps demand <= budget."""
|
||||
def run_calibration():
|
||||
print("=" * 88)
|
||||
print("Calibration — predicted per-card VRAM vs measured BENCHMARKS rows")
|
||||
print("=" * 88)
|
||||
print()
|
||||
print(" Predicted = weights + activation + overhead + drafter + (KV capped at available).")
|
||||
print(" Budget = mem_util × VRAM. Measured = nvidia-smi peak during bench (target ≈ budget).")
|
||||
print(" Verdict ✓ iff PASS/TIGHT and measured < VRAM (boot OK).")
|
||||
print()
|
||||
|
||||
total_c, total_n = 0, 0
|
||||
for model_key in ("qwen3.6-27b", "gemma-4-31b"):
|
||||
c, n = _calibration_block(model_key)
|
||||
total_c += c
|
||||
total_n += n
|
||||
|
||||
if total_n > 0:
|
||||
print(f"Overall: {total_c}/{total_n} ({100*total_c/total_n:.0f}%)")
|
||||
print()
|
||||
print("Notes:")
|
||||
print(" - This is a directional estimator (±1.5 GB error band on the breakdown).")
|
||||
print(" - vLLM's `gpu_worker.py` boot log is the authoritative source.")
|
||||
print(" - If predicted PASS but measured > budget, file an issue with `scripts/report.sh --bench`.")
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# Max-ctx solver
|
||||
# =============================================================================
|
||||
|
||||
def solve_max_ctx(spec, kv_format, max_num_seqs, tp, mem_util, vram_gb,
|
||||
drafter_gb=0.0, dflash_draft_gb=0.0, mtp=False, weights_variant="default"):
|
||||
"""Binary search for the largest max_ctx that keeps the verdict at PASS or TIGHT."""
|
||||
lo, hi = 1024, spec.get("max_ctx_supported", 262144)
|
||||
best = 0
|
||||
while lo <= hi:
|
||||
mid = (lo + hi) // 2
|
||||
# Round to nearest 1024 for cleaner numbers
|
||||
mid = (mid // 1024) * 1024
|
||||
mid = (mid // 1024) * 1024 # round to nearest 1024 for cleaner numbers
|
||||
if mid == 0:
|
||||
break
|
||||
p = predict(spec, kv_format=kv_format, max_ctx=mid, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=vram_gb,
|
||||
dflash_draft_gb=dflash_draft_gb, mtp=mtp)
|
||||
if p.verdict in ("PASS", "TIGHT") and p.pct_of_vram < 100:
|
||||
p = predict(
|
||||
spec=spec, kv_format=kv_format, max_ctx=mid, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=vram_gb,
|
||||
drafter_gb=drafter_gb, dflash_draft_gb=dflash_draft_gb,
|
||||
mtp=mtp, weights_variant=weights_variant,
|
||||
)
|
||||
if p.verdict in ("PASS", "TIGHT"):
|
||||
best = mid
|
||||
lo = mid + 1024
|
||||
else:
|
||||
@@ -343,22 +620,54 @@ def solve_max_ctx(spec, kv_format, max_num_seqs, tp, mem_util, vram_gb, dflash_d
|
||||
return best
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# CLI
|
||||
# =============================================================================
|
||||
|
||||
def _all_compose_choices() -> list[str]:
|
||||
"""Flat list of compose names across all models for argparse choices."""
|
||||
out = []
|
||||
for model_key in COMPOSES:
|
||||
out.extend(COMPOSES[model_key].keys())
|
||||
return sorted(set(out))
|
||||
|
||||
|
||||
def _resolve_compose_model(compose_name: str, explicit_model: Optional[str]) -> str:
|
||||
"""Infer model from compose name if --model not given.
|
||||
|
||||
Composes are namespaced by prefix; Qwen uses bare names, Gemma uses gemma-*.
|
||||
"""
|
||||
if explicit_model:
|
||||
return explicit_model
|
||||
for model_key, composes in COMPOSES.items():
|
||||
if compose_name in composes:
|
||||
return model_key
|
||||
return "qwen3.6-27b" # back-compat default
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
|
||||
p.add_argument("--compose", choices=sorted(COMPOSES.keys()),
|
||||
p.add_argument("--model", choices=sorted(MODEL_SPECS.keys()),
|
||||
help="Which model to predict for. Default: qwen3.6-27b (back-compat) or inferred from --compose.")
|
||||
p.add_argument("--compose", choices=_all_compose_choices(),
|
||||
help="Use a shipped compose's defaults. Override individual flags below.")
|
||||
p.add_argument("--kv-format", choices=sorted(KV_FORMAT_BYTES.keys()),
|
||||
help="KV cache format. Default: from --compose, or fp8_e5m2.")
|
||||
p.add_argument("--max-ctx", type=int, help="max_model_len. Default: from --compose, or 180000.")
|
||||
p.add_argument("--max-num-seqs", type=int, help="max_num_seqs. Default: from --compose, or 1.")
|
||||
p.add_argument("--tp", type=int, choices=[1, 2, 4], help="tensor_parallel_size. Default: from --compose, or 1.")
|
||||
p.add_argument("--tp", type=int, choices=[1, 2, 4, 8, 16], help="tensor_parallel_size. Default: from --compose, or 1.")
|
||||
p.add_argument("--mem-util", type=float, help="gpu_memory_utilization. Default: from --compose, or 0.95.")
|
||||
p.add_argument("--vram", type=float, default=24, help="VRAM per card in GB. Default 24.")
|
||||
p.add_argument("--mtp", action="store_true", help="MTP n=3 enabled (adds small KV overhead per request).")
|
||||
p.add_argument("--mtp", action="store_true", default=None, help="MTP enabled (Qwen: n=3 built-in; Gemma: external drafter).")
|
||||
p.add_argument("--no-mtp", dest="mtp", action="store_false")
|
||||
p.add_argument("--dflash-draft-gb", type=float, default=0.0, help="DFlash draft model size in GB (0 if not using DFlash).")
|
||||
p.add_argument("--calibration", action="store_true", help="Print predicted vs measured for all calibration points.")
|
||||
p.add_argument("--solve-max-ctx", action="store_true", help="Binary-search for the largest max_ctx that fits given the other parameters.")
|
||||
p.add_argument("--drafter-gb", type=float, default=None,
|
||||
help="Drafter model size in GB (MTP / DFlash). 0 if not using a drafter.")
|
||||
p.add_argument("--dflash-draft-gb", type=float, default=None,
|
||||
help="(deprecated alias for --drafter-gb)")
|
||||
p.add_argument("--weights-variant", choices=["default", "int4", "awq", "bf16"], default=None,
|
||||
help="Gemma 4 only: which weight quant variant. Default: from --compose, or int4.")
|
||||
p.add_argument("--calibration", action="store_true", help="Print predicted vs measured for both models.")
|
||||
p.add_argument("--solve-max-ctx", action="store_true", help="Binary-search for the largest max_ctx that fits.")
|
||||
p.add_argument("--json", action="store_true", help="Output prediction as JSON.")
|
||||
args = p.parse_args()
|
||||
|
||||
@@ -366,56 +675,83 @@ def main():
|
||||
run_calibration()
|
||||
return 0
|
||||
|
||||
# Resolve defaults from compose preset (used by both modes)
|
||||
# Resolve model: explicit --model > inferred from --compose > qwen3.6-27b
|
||||
model_key = _resolve_compose_model(args.compose, args.model) if args.compose else (args.model or "qwen3.6-27b")
|
||||
spec = MODEL_SPECS[model_key]
|
||||
|
||||
# Resolve compose-derived defaults
|
||||
if args.compose:
|
||||
cfg = COMPOSES[args.compose]
|
||||
# Compose must belong to the resolved model
|
||||
if args.compose not in COMPOSES[model_key]:
|
||||
print(f"ERROR: --compose {args.compose} is not in --model {model_key}'s compose list.", file=sys.stderr)
|
||||
print(f" Available for {model_key}: {', '.join(sorted(COMPOSES[model_key].keys()))}", file=sys.stderr)
|
||||
return 2
|
||||
cfg = COMPOSES[model_key][args.compose]
|
||||
kv_format = args.kv_format or cfg["kv_format"]
|
||||
max_ctx = args.max_ctx or cfg["max_ctx"]
|
||||
max_num_seqs = args.max_num_seqs or cfg["max_num_seqs"]
|
||||
tp = args.tp or cfg["tp"]
|
||||
mem_util = args.mem_util if args.mem_util is not None else cfg["mem_util"]
|
||||
mtp = args.mtp if args.mtp is not None else cfg.get("mtp", False)
|
||||
dflash_gb = args.dflash_draft_gb or cfg.get("dflash_draft_gb", 0.0)
|
||||
header = f"Predicted budget — {args.compose}.yml on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
|
||||
drafter_gb = args.drafter_gb if args.drafter_gb is not None else cfg.get("drafter_gb", 0.0)
|
||||
dflash_gb = args.dflash_draft_gb if args.dflash_draft_gb is not None else cfg.get("dflash_draft_gb", 0.0)
|
||||
weights_variant = args.weights_variant or cfg.get("weights_variant", "default")
|
||||
header = f"Predicted budget — {model_key} / {args.compose} on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx:,}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
|
||||
else:
|
||||
kv_format = args.kv_format or "fp8_e5m2"
|
||||
max_ctx = args.max_ctx or 180000
|
||||
max_num_seqs = args.max_num_seqs or 1
|
||||
tp = args.tp or 1
|
||||
mem_util = args.mem_util if args.mem_util is not None else 0.95
|
||||
mtp = bool(args.mtp)
|
||||
dflash_gb = args.dflash_draft_gb
|
||||
header = f"Predicted budget — custom config on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
|
||||
mtp = bool(args.mtp) if args.mtp is not None else False
|
||||
drafter_gb = args.drafter_gb or 0.0
|
||||
dflash_gb = args.dflash_draft_gb or 0.0
|
||||
weights_variant = args.weights_variant or "default"
|
||||
header = f"Predicted budget — {model_key} custom config on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx:,}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
|
||||
|
||||
try:
|
||||
_validate_tp_for_spec(spec, tp)
|
||||
except ValueError as exc:
|
||||
print(f"ERROR: {exc}", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
if args.solve_max_ctx:
|
||||
# Pin max_ctx very high; we'll search for the actual largest that fits.
|
||||
best = solve_max_ctx(QWEN36_27B, kv_format=kv_format, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=args.vram,
|
||||
dflash_draft_gb=dflash_gb, mtp=mtp)
|
||||
best = solve_max_ctx(
|
||||
spec, kv_format=kv_format, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=args.vram,
|
||||
drafter_gb=drafter_gb, dflash_draft_gb=dflash_gb, mtp=mtp,
|
||||
weights_variant=weights_variant,
|
||||
)
|
||||
if best > 0:
|
||||
pred_at_best = predict(kv_format=kv_format, max_ctx=best, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=args.vram,
|
||||
dflash_draft_gb=dflash_gb, mtp=mtp)
|
||||
print(f"Max-ctx solver — {kv_format}, seqs={max_num_seqs}, TP={tp}, mem_util={mem_util}, VRAM={args.vram} GB")
|
||||
print(f" Largest max_ctx that fits: {best:,} tokens")
|
||||
print(f" At that ctx: predicted demand = {pred_at_best.total_gb:.2f} GB / card ({pred_at_best.pct_of_vram:.0f}% of budget)")
|
||||
print(f" Verdict at that ctx: {pred_at_best.verdict}")
|
||||
print()
|
||||
print("Note: this is a directional estimate (±1.5 GB error band). The vLLM engine")
|
||||
print("pre-check (gpu_worker.py boot log) is authoritative.")
|
||||
pred_at_best = predict(
|
||||
spec=spec, kv_format=kv_format, max_ctx=best, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=args.vram,
|
||||
drafter_gb=drafter_gb, dflash_draft_gb=dflash_gb, mtp=mtp,
|
||||
weights_variant=weights_variant,
|
||||
)
|
||||
if args.json:
|
||||
out = pred_at_best.__dict__.copy()
|
||||
out["solved_max_ctx"] = best
|
||||
print(json.dumps(out, indent=2))
|
||||
else:
|
||||
print(f"Max-ctx solver — {model_key} / {kv_format}, seqs={max_num_seqs}, TP={tp}, mem_util={mem_util}, VRAM={args.vram} GB")
|
||||
print(f" Largest max_ctx that fits: {best:,} tokens")
|
||||
print(f" At that ctx: predicted = {pred_at_best.total_gb:.2f} GB / card ({pred_at_best.pct_of_vram:.0f}% of budget)")
|
||||
print(f" Verdict at that ctx: {pred_at_best.verdict}")
|
||||
for note in pred_at_best.notes:
|
||||
print(f" Note: {note}")
|
||||
print()
|
||||
print("Note: this is a directional estimate (±1.5 GB error band). The vLLM engine")
|
||||
print("pre-check (gpu_worker.py boot log) is authoritative.")
|
||||
else:
|
||||
print(f"No max_ctx fits at this config on {args.vram} GB. Reduce TP, swap KV format, or get bigger cards.")
|
||||
return 0
|
||||
|
||||
pred = predict(
|
||||
kv_format=kv_format,
|
||||
max_ctx=max_ctx,
|
||||
max_num_seqs=max_num_seqs,
|
||||
tp=tp,
|
||||
mem_util=mem_util,
|
||||
vram_gb=args.vram,
|
||||
dflash_draft_gb=dflash_gb,
|
||||
mtp=mtp,
|
||||
spec=spec, kv_format=kv_format, max_ctx=max_ctx, max_num_seqs=max_num_seqs,
|
||||
tp=tp, mem_util=mem_util, vram_gb=args.vram,
|
||||
drafter_gb=drafter_gb, dflash_draft_gb=dflash_gb, mtp=mtp,
|
||||
weights_variant=weights_variant,
|
||||
)
|
||||
|
||||
if args.json:
|
||||
@@ -423,11 +759,9 @@ def main():
|
||||
else:
|
||||
print(fmt_prediction(pred, header=header))
|
||||
print()
|
||||
print("Anchored to: PerfMamba (arxiv 2511.22849), TurboQuant (arxiv 2504.19874),")
|
||||
print("PagedAttention (arxiv 2309.06180). Calibrated against BENCHMARKS.md rows.")
|
||||
print("Error band: ±1.5 GB on the breakdown. Verdicts within 5% of the budget are TIGHT.")
|
||||
print("Run `tools/kv-calc.py --calibration` to see predicted-vs-measured for all anchors.")
|
||||
print("Run `tools/kv-calc.py --solve-max-ctx ...` to find the largest max_ctx that fits your config.")
|
||||
print("Run `tools/kv-calc.py --solve-max-ctx ...` to find the largest max_ctx that fits.")
|
||||
print("See docs/KV_MATH.md for math + per-model architecture details.")
|
||||
return 0 if pred.verdict in ("PASS", "TIGHT") else 1
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user