12 Commits
Author SHA1 Message Date
noonghunna 98f0406d0f fix(launch): project TP greater than four
Release / release (push) Failing after 48s
2026-05-14 00:23:32 +00:00
github-actions[bot] 2a92a199ff chore(changelog): regenerate for v0.6.1 [skip ci] 2026-05-14 00:02:46 +00:00
noonghunnaandCodex (via codex/v0.6.1-launch branch) 056dcb6439 Merge codex/v0.6.1-launch into master
Release / release (push) Failing after 51s
v0.6.1 — hardware-aware launch.sh with KV projection + TP/PP env overrides.

Brings in:
- feat(launch): hardware-aware launcher (model picker, GPU picker, TP/PP
  auto-pick, KV projection via kv-calc.py, sub-fit guidance, --variant
  back-compat)
- All vLLM composes accept ${TP:-N} / ${PP:-1} env overrides + carry
  Engine-profile metadata header
- llama.cpp composes tagged with Engine-profile metadata
- compose-meta.sh extended with fake-GPU + busy-GPU support for tests
- docs: launch.sh references updated across README, FAQ, SINGLE_CARD,
  DUAL_CARD to describe the new wizard

Co-Authored-By: Codex (via codex/v0.6.1-launch branch)
2026-05-14 00:02:08 +00:00
noonghunnaandClaude Opus 4.7 e299e70451 docs: update launch.sh references for v0.6.1 wizard flow
- README.md: example block now shows model + GPUs flow with new --model /
  --gpus / --tp / --pp flag examples; scripts/ layout description updated to
  "model → GPUs → KV projection".
- docs/FAQ.md: rewrite the "how do I pick model + variant?" answer to
  describe the new wizard (model picker → GPU picker → TP/PP auto-pick →
  kv-calc projection → boot). Replace CUDA_VISIBLE_DEVICES card-selection
  guidance with --gpus flag, preserving the env-var path for compatibility.
- docs/SINGLE_CARD.md, docs/DUAL_CARD.md: fix inline wizard descriptions
  ("asks engine + workload" / "asks GPU count + workload" → "asks model +
  GPUs, projects VRAM budget").

All --variant <name> back-compat invocations across the docs continue to
work; no flag-renaming churn surfaced to users.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-14 00:00:39 +00:00
noonghunna 5882bbef6f feat(launch): add hardware-aware launcher 2026-05-13 23:43:32 +00:00
noonghunnaandClaude Opus 4.7 0d48dac818 feat(tools): extend kv-calc.py to multi-model (Qwen 3.6 + Gemma 4 31B)
- Refactor MODEL_SPECS / COMPOSES / CALIBRATION into per-model dicts
- Add --model flag (defaults qwen3.6-27b for back-compat; infers from --compose)
- Add Gemma 4 31B architecture (10 full + 50 sliding layers, K==V tying,
  global_head_dim=512 asymmetry)
- KV pool capping models vLLM's PagedAttention allocator: predicted KV is
  min(requested, available_budget); verdict becomes TIGHT when capped
  (resolves the original limitation #1 — max_num_seqs>1 over-prediction)
- 8 Gemma composes registered; calibration accuracy 18/18 across both models
  (was 9/11 Qwen-only)
- Update docs/KV_MATH.md: per-model sections, K==V tying empirical finding,
  multi-model calibration table, limitation #1 marked resolved

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-13 23:15:13 +00:00
github-actions[bot] a119e1b1b6 chore(changelog): regenerate for v0.6.0 [skip ci] 2026-05-13 20:45:23 +00:00
noonghunnaandClaude Opus 4.7 e05f1969bc fix(launch): exit cleanly on stdin EOF in wizard prompts
Release / release (push) Failing after 52s
The two `read -rp` loops in `choose()` and the variant-selection step
spun infinitely when stdin closed mid-prompt (piped input shorter than
the wizard asks, or test/CI invocations). Now both loops check `read`'s
exit code; on EOF, print a clear message and SIGINT the parent so the
whole process exits cleanly (exit 130).

Doesn't affect interactive users — they type real input. Affects only
non-TTY piped invocations of the wizard, which should use --variant
<name> to skip the wizard entirely.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-13 20:44:35 +00:00
noonghunna 12a33fbdd1 feat(scripts): add hardware-aware setup picker 2026-05-13 20:44:35 +00:00
github-actions[bot] 45fa421e93 chore(changelog): regenerate for v0.5.4 [skip ci] 2026-05-13 18:41:00 +00:00
noonghunna 22bf2e9398 fix(scripts): make submit-bench issue-first
Release / release (push) Failing after 53s
2026-05-13 18:40:43 +00:00
github-actions[bot] 99a0b66224 chore(changelog): regenerate for v0.5.3 [skip ci] 2026-05-13 18:24:40 +00:00
49 changed files with 2346 additions and 409 deletions
+5 -1
View File
@@ -1,6 +1,10 @@
## Rig bench submission
<!-- This PR was auto-generated by `bash scripts/submit-bench.sh --auto-submit --tag <TAG>`. -->
> ⚠ Most bench submissions go through an issue (see `CONTRIBUTING.md` "Submitting your bench").
> This PR template is for contributors who explicitly chose the direct-PR path.
> The maintainer may redirect to an issue thread before merge.
<!-- This PR was auto-generated by `bash scripts/submit-bench.sh --auto-submit --as-pr --tag <TAG>`. -->
<!-- Review the row below; the PR reviewer may move it within the target section. -->
### New row
+56
View File
@@ -16,6 +16,62 @@ history; SemVer takes over from `v0.3.0` onward.
---
## v0.6.1 — 2026-05-14
### ✨ Features
- feat(launch): add hardware-aware launcher ([5882bbe](https://github.com/noonghunna/club-3090/commit/5882bbef6f5ed7e8ceea450fa3a6b167a1bf4926))
- feat(tools): extend kv-calc.py to multi-model (Qwen 3.6 + Gemma 4 31B) ([0d48dac](https://github.com/noonghunna/club-3090/commit/0d48dac818ba889f74b74df1c454bb961a123c37))
### 📝 Documentation
- docs: update launch.sh references for v0.6.1 wizard flow ([e299e70](https://github.com/noonghunna/club-3090/commit/e299e70451c8d146214a6560e582d0e174dd0ebc))
### 🧹 Other
- Merge codex/v0.6.1-launch into master ([056dcb6](https://github.com/noonghunna/club-3090/commit/056dcb643914fee6169b02b89cb420b038c29b0f))
[Pin: `git checkout v0.6.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.6.0...v0.6.1)
## v0.6.0 — 2026-05-13
### ✨ Features
- feat(scripts): add hardware-aware setup picker ([12a33fb](https://github.com/noonghunna/club-3090/commit/12a33fbdd1b6422429521f887d9f21c2e3da793d))
### 🐛 Bug fixes
- fix(launch): exit cleanly on stdin EOF in wizard prompts ([e05f196](https://github.com/noonghunna/club-3090/commit/e05f1969bcf57547a17cd963a1f21435de80815b))
[Pin: `git checkout v0.6.0`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.4...v0.6.0)
## v0.5.4 — 2026-05-13
### 🐛 Bug fixes
- fix(scripts): make submit-bench issue-first ([22bf2e9](https://github.com/noonghunna/club-3090/commit/22bf2e9398c7907aae6b62809bde30e111e4a700))
[Pin: `git checkout v0.5.4`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.3...v0.5.4)
## v0.5.3 — 2026-05-13
### ✨ Features
- feat(scripts): add submit-bench flow ([ef77032](https://github.com/noonghunna/club-3090/commit/ef770322f43724f612a80393f547e5da218b5bf7))
[Pin: `git checkout v0.5.3`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.2...v0.5.3)
## v0.5.2 — 2026-05-13
+18 -6
View File
@@ -8,7 +8,7 @@ Thanks for being here. This repo collects working recipes for serving big LLMs o
### ✅ Yes please
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. If you're unsure what to run before measuring, start with `bash scripts/setup.sh` and `bash scripts/launch.sh`; both wizards mark hardware-fit. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
- **Power-cap efficiency curves.** `sudo bash scripts/power-cap-sweep.sh --cooling air|water|aio --load-mode decode-concurrent --concurrency auto --bench-runs 3` produces cross-rig efficiency-knee data ([discussion #86](https://github.com/noonghunna/club-3090/discussions/86)). ~15-20 min for a 30-cap sweep on a 3090/4090/5090. Especially valuable on cards we don't have anchors for yet (A5000/A6000, 4080, 5060 Ti / 5080, modded variants). **Keep `--step-size 10` (the default).** Larger step-sizes (e.g. `--step-size 50`) are too coarse for the efficiency knee and only useful for quick smoke tests. See [docs/HARDWARE.md](docs/HARDWARE.md#cross-rig-power-cap-data-anchor-points) for the full canonical command and rationale.
- **Bug reports with the data we ask for.** The [bug report template](https://github.com/noonghunna/club-3090/issues/new?template=bug-report.yml) leads with `bash scripts/report.sh > my-rig.md` (add `--verify` to include verify-full output, `--soak` to also run SOAK_MODE=continuous if you suspect a multi-turn agent cliff) — single command captures the rig state we'd otherwise ask for individually (hardware, container state, Genesis patches, KV pool sizing, engine config). With that paste, the first reply is usually a fix or a clear next step instead of "can you send me…".
- **Bug reproductions / minimum repros for upstream issues.** vLLM / llama.cpp / Genesis bugs that affect this stack are most useful when they have a one-paragraph reduction. Drop them in an issue or open a draft PR adding a reproducer to `verify-stress.sh`.
@@ -47,23 +47,35 @@ Two GitHub channels, two different shapes of conversation. Picking the right one
## Submitting your bench
After running `bash scripts/rebench-full.sh`, contribute your numbers to `BENCHMARKS.md` with:
The matrix is hand-curated — the canonical path is to file an **issue** with your rig + numbers; we'll review, ask clarifying questions, and integrate.
After running `bash scripts/rebench-full.sh`, generate a paste-ready row:
```bash
bash scripts/submit-bench.sh --tag <your-tag>
```
This generates a paste-ready row at `results/rebench/<tag>/BENCHMARKS-row.md`. Review it, then either auto-submit or open the PR manually.
The script writes `results/rebench/<tag>/BENCHMARKS-row.md`. To submit:
Auto-submit opens a PR via GitHub CLI:
### Path A — Auto-issue (recommended, requires `gh auth login`)
```bash
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit
```
Manual path: paste the row into the appropriate `BENCHMARKS.md` section and open a PR yourself.
Opens an issue via `gh issue create` with your rig + row pre-filled.
Auto-submit assumes `gh auth status` is configured. If it is not, run `gh auth login` first.
### Path B — Manual issue (no tools beyond browser)
Open https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml and paste the row + your `rig.txt` into the body.
### Path C — Direct PR (advanced)
```bash
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit --as-pr
```
For contributors who know the `BENCHMARKS.md` section structure and want to propose the exact row. The maintainer may still redirect to an issue thread for context-gathering before merge — direct PRs aren't a fast-path bypass.
---
+12 -7
View File
@@ -66,18 +66,23 @@ Bench protocol: 3 warm + 5 measured runs of the canonical narrative + code promp
git clone https://github.com/noonghunna/club-3090.git
cd club-3090
# 2. Download + SHA-verify the model (~20 GB; clones Genesis patches too)
# (asks you where to put model weights — pick in-repo default, ~/models, or
# a custom path on a different drive. To skip the prompt:
# `export MODEL_DIR=/mnt/your-drive/models` before running. See FAQ + .env.example.)
bash scripts/setup.sh qwen3.6-27b
# 2. Pick/download + SHA-verify the model (interactive hardware-aware picker)
# (asks you which model, then where to put model weights — pick in-repo
# default, ~/models, or a custom path on a different drive. To skip prompts:
# `export MODEL_DIR=/mnt/your-drive/models` and pass the model name. See FAQ.)
bash scripts/setup.sh
# Or scripted:
# bash scripts/setup.sh qwen3.6-27b
# 3. Pick a config + boot it (interactive wizard — asks engine / cards / workload)
# 3. Pick a config + boot it (interactive wizard: asks model → GPUs → projects VRAM budget)
bash scripts/launch.sh
# Or skip the wizard:
# bash scripts/launch.sh --variant vllm/default # single-card chat (recommended)
# bash scripts/launch.sh --variant vllm/dual # dual-card 262K + vision
# bash scripts/launch.sh --variant llamacpp/default # single-card 262K, no cliffs
# Or partial flags (wizard fills the rest):
# bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1
# bash scripts/launch.sh --tp 2 --pp 1 # override vLLM parallelism
# See all variants:
# bash scripts/switch.sh --list
@@ -152,7 +157,7 @@ club-3090/
│ └── README.md blocked status — what would unblock it on this model
├── scripts/ shared, model-aware
│ ├── setup.sh bash setup.sh <model> → preflight + downloads + verifies + Genesis
│ ├── launch.sh interactive wizard: cards → workload → boots compose + verifies
│ ├── launch.sh interactive wizard: model → GPUs → KV projection → boots compose + verifies
│ ├── switch.sh stateless variant switcher (bring down old, up new)
│ ├── update.sh one-shot upgrade: git pull + re-pin Genesis + re-vendor patches
│ ├── health.sh runtime health probe (KV %, MTP AL, recent TPS, errors)
+1 -1
View File
@@ -163,7 +163,7 @@ If you're solo-using on dual, you're paying for hardware that mostly sits idle o
bash scripts/setup.sh qwen3.6-27b
git clone https://github.com/vllm-project/vllm.git /opt/ai/engines/vllm/primary # required for dual variants
# 2. Pick + boot via wizard (asks GPU count + workload)
# 2. Pick + boot via wizard (asks model + GPUs, projects VRAM budget, auto-picks TP=2 for matched 2× 3090)
bash scripts/launch.sh
# 3. Or skip the wizard:
+8 -2
View File
@@ -130,6 +130,12 @@ If your numbers on the same compose look different from ours by >15%, the most l
## Setup
### How do I pick the right model + variant?
For a first install, run `bash scripts/setup.sh` with no model argument in a normal terminal. It opens a hardware-aware model picker, marks Qwen / Gemma / Both as eligible or not for your detected GPUs, then continues into the existing download flow.
After setup, run `bash scripts/launch.sh`. The wizard asks which model (filtered to what you've downloaded), then which GPU(s) to use, auto-picks TP for homogeneous sets (PP for heterogeneous), filters variants by hardware fit, shows a per-card VRAM projection from `tools/kv-calc.py` for the suggested default, then boots and runs `verify-full.sh`. Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, partial flags like `bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1` (skips prompts), `--tp 4 --pp 2` to override parallelism, plus `setup.sh --help` / `launch.sh --help` for the full flag list.
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
@@ -140,7 +146,7 @@ Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
bash scripts/setup.sh qwen3.6-27b
```
2. **`.env` file at repo root** — picked up automatically on every script run. See [`.env.example`](../.env.example).
3. **Interactive prompt** — `bash scripts/setup.sh qwen3.6-27b` with nothing set offers three choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
3. **Interactive prompt** — `bash scripts/setup.sh` with nothing set first asks which model to download, then offers three model-dir choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
4. **Silent fallback** — `<repo>/models-cache/`. Functional but pollutes the git tree; not recommended.
Every script that touches model paths reads from the same `MODEL_DIR`. The compose YAMLs' volume mount is `${MODEL_DIR:-...}:/root/.cache/huggingface` — once set, every container reads + writes there.
@@ -171,7 +177,7 @@ If your *Genesis tree* (not the repo) is out of sync — the pin in `setup.sh` m
### My GPU isn't card 0 — how do I change it?
`CUDA_VISIBLE_DEVICES=2 bash scripts/launch.sh --variant vllm/default` (substitute your card index). For dual-card, pass two: `CUDA_VISIBLE_DEVICES=2,3`. The compose files inherit env from your shell.
Use the `--gpus` flag: `bash scripts/launch.sh --gpus 2` (single-card) or `bash scripts/launch.sh --gpus 2,3` (two cards). The wizard exports `CUDA_VISIBLE_DEVICES` for you. The older form `CUDA_VISIBLE_DEVICES=2 bash scripts/launch.sh --variant vllm/default` still works if you prefer to set the env yourself.
### Container fails to start: "Free memory ... is less than desired GPU memory utilization"
+132 -17
View File
@@ -1,23 +1,27 @@
# KV Cache Math — predicting per-card VRAM budget on Qwen3.6-27B
# KV Cache Math — predicting per-card VRAM budget
This page documents the math behind [`tools/kv-calc.py`](../tools/kv-calc.py) — the predictor that helps you decide whether a config will fit on your hardware *before* booting it. It also explains why predictions are estimates (±1.5 GB error band) rather than precise allocations.
Two models are modelled today: **Qwen 3.6 27B** (DeltaNet hybrid) and **Gemma 4 31B** (sliding-window + dense MLP). The math differs structurally between them; this doc covers each in its own section. The CLI dispatches on `--model`.
## TL;DR
```bash
# What's my budget if I run dual-turbo on 20 GB cards?
bash tools/kv-calc.py --compose dual-turbo --vram 20 --mem-util 0.82
# Qwen 3.6 27B — what's my budget if I run dual-turbo on 20 GB cards?
bash tools/kv-calc.py --model qwen3.6-27b --compose dual-turbo --vram 20 --mem-util 0.82
# What's the largest max_ctx that fits on 16 GB cards with TP=2 + fp8 KV?
bash tools/kv-calc.py --solve-max-ctx --tp 2 --kv-format fp8_e5m2 --vram 16 --mem-util 0.95
# Gemma 4 31B — what's the largest max_ctx that fits on 24 GB cards with TP=2 + INT8 PTH KV?
bash tools/kv-calc.py --model gemma-4-31b --solve-max-ctx --tp 2 --kv-format int8_per_token_head --vram 24 --mem-util 0.92
# How accurate is the model? Show predicted vs measured for our shipped composes:
# How accurate is the model? Show predicted vs measured for our shipped composes (both models):
bash tools/kv-calc.py --calibration
```
`--model` defaults to `qwen3.6-27b` for backward compatibility with earlier invocations.
The predictor is a directional estimator, not a precise allocator. The vLLM engine's `gpu_worker.py` boot-log report is authoritative — the calculator is for *before* boot.
## Per-card budget components
## Qwen 3.6 27B — per-card budget components
For Qwen3.6-27B AutoRound INT4 at TP=N, the per-card VRAM peak during bench is composed of:
@@ -109,43 +113,154 @@ This is rough — actual overhead depends on how many graphs vLLM captures, whic
Only present on `dual-dflash*.yml` composes. `z-lab/Qwen3.6-27B-DFlash` is a ~1.75 GB draft model (per card, FP16). With TP > 1, the draft itself is sharded.
## Gemma 4 31B — per-card budget components
Gemma 4 31B is structurally different from Qwen 3.6:
- **No DeltaNet, no GDN activation peak.** Dense MLP instead.
- **Hybrid layer pattern** — but the hybrid is on *attention type*, not attention-vs-recurrence. The 60-layer stack is `[sliding_attention × 5, full_attention × 1] × 10` = **50 sliding-attention layers + 10 full-attention layers**.
- **Head-dim asymmetry** — sliding layers use `head_dim=256`, full layers use `global_head_dim=512`. Per-token KV bytes for the full layers is therefore double what naive `num_layers × head_dim` would compute.
- **K==V tying** — `attention_k_eq_v: true` in `config.json`. vLLM's allocator EXPLOITS this — K and V share storage. The KV term uses `×1`, not `×2`. This was confirmed empirically by calibrating the per-token byte count against measured BENCHMARKS rows (the matched-config rebench's `Available KV cache / card = 10.82 GiB` at 262K seqs=2 is consistent with ×1 storage, not ×2).
Source: `/mnt/models/huggingface/gemma-4-31b-autoround-int4/config.json` → `text_config`.
```
peak ≈ weights/N + kv_pool_growing + kv_pool_sliding + activation_peak + cudagraph_overhead + drafter_overhead
```
### 1. Model weights (`weights / N`)
| Quant | On-disk | Per-card at TP=2 |
|---|---:|---:|
| AutoRound INT4 (`gemma-4-31b-autoround-int4`) | ~18 GB | 9.0 GB |
| AWQ-4bit (`cyankiwi/gemma-4-31B-it-AWQ-4bit`) | ~17 GB usable on stack | 8.5 GB |
| BF16 (unquantized) | ~58 GB | 29 GB (does not fit on 24 GB) |
The two shipped quants on this stack are AutoRound INT4 (default) and AWQ-4bit (Tier 2 reproducer of #103). INT4 weights + INT8-per-token-head KV is the matched-config dual-3090 recipe (see `models/gemma-4-31b/vllm/compose/dual/int8.yml`).
### 2. KV pool — growing portion (full-attention layers only)
Only the 10 full-attention layers grow KV with context. Each stores K and V at `global_head_dim=512`, with K==V tying meaning a single store per element:
```
per_token_bytes_growing = num_full_attn_layers × num_kv_heads × global_head_dim × 1 × bpe
= 10 × 16 × 512 × 1 × bpe
= 81,920 × bpe bytes
```
For comparison, Qwen 3.6's growing KV is `16 × 4 × 256 × 2 × bpe = 32,768 × bpe` — Gemma 4's per-token growing KV is **~2.5× heavier** than Qwen's. This is *the* reason Gemma 4 at 262K needs INT8 / FP8 KV on Ampere — at BF16 KV the per-card budget blows past 24 GB before you reach 50K context.
Per-token growing-KV bytes by format:
| KV format | `bytes_per_kv_element` | per-token growing KV (full) | per-token (TP=2) |
|---|---:|---:|---:|
| `fp16` / `bf16` | 2.0 | 163,840 B (~160 KB) | 81,920 B |
| `fp8_e5m2` / `fp8_e4m3` | 1.0 | 81,920 B (~80 KB) | 40,960 B |
| `int8_per_token_head` (PR #40391) | ~1.01 (incl. per-token scale) | ~82,700 B | ~41,400 B |
| `q4_0` | ~0.56 | ~45,875 B | ~22,940 B |
| `turboquant_3bit_nc` (TQ3) | ~0.425 | ~34,816 B | ~17,408 B |
Total growing-KV pool per card = `per_token_bytes_growing / TP × max_ctx × max_num_seqs`.
**Note**: on Ampere consumer cards (sm_86), `fp8_e4m3` is NOT supported by the Triton kernel (`fp8e4nv` requires Hopper/Ada/Blackwell). Use `int8_per_token_head` (PR #40391, vendored on this stack via PR #42102) instead. See `models/gemma-4-31b/vllm/compose/dual/int8.yml` header for the engineering trail.
### 3. KV pool — fixed sliding portion (50 layers)
The 50 sliding-attention layers maintain a fixed-size KV window (`sliding_window=1024`). K==V tying applies here too:
```
sliding_kv_bytes_total = num_sliding_layers × num_kv_heads × head_dim × 1 × bpe × sliding_window
= 50 × 16 × 256 × 1 × bpe × 1024
= 209,715,200 × bpe bytes
≈ 200 MB × bpe
```
This is constant — it doesn't scale with `max_ctx` or `max_num_seqs`. At fp8 / int8 KV (`bpe=1`), this is ~200 MB per card (TP=1) or ~100 MB at TP=2. Small but non-zero — include it as a separate term.
### 4. Activation peak (SWA prefill + dense MLP)
Unlike Qwen 3.6's GDN block-wise state materialization, Gemma 4's activation peak comes from:
- Sliding-window attention prefill (50 layers, but bounded by `sliding_window=1024`).
- Dense MLP intermediate buffer (`hidden_size=5376`, `intermediate_size=21504`).
There's no published scaling-law analogue to PerfMamba's O(γDNL) for Gemma 4. The activation coefficient is **empirical-only**, calibrated against measured BENCHMARKS.md rows. Expected order of magnitude: ~1.5-2.5 GB at TP=2 dual-card configs (smaller than Qwen 3.6's GDN peak because there's no per-chunk block materialization).
The coefficient may have weak dependence on KV format (slight dequant overhead during forward) but we expect it to be flatter than Qwen's TQ3 → fp8 25% spread — Gemma's dense MLP doesn't dequant KV during its forward.
### 5. Cudagraph + workspace overhead
Same form as Qwen — empirical fit:
```
overhead = 0.5 + 1.0 × mem_util + 0.3 × (TP - 1) # GB
```
vLLM captures multiple cudagraphs (~50-100 MB each), FlashInfer workspace (~394 MB/card), NCCL allreduce buffers (~200-300 MB at TP > 1). Same accounting as Qwen.
### 6. Drafter overhead
Two drafter families on this stack:
| Drafter | Size | Composes |
|---|---:|---|
| `gemma-4-31b-it-assistant` (Google MTP) | 0.97 GB FP16 | `dual/docker-compose.yml`, `dual/int8.yml`, `dual/awq.yml` (with MTP n=4) |
| `gemma-4-31b-it-dflash` (z-lab DFlash) | 2.9 GB FP16 | `dual/dflash.yml`, `dual/dflash-int8.yml` |
At TP > 1, drafter weights shard across cards (`drafter_gb / TP`).
## Known limitations
**The model is empirically calibrated, not first-principles.** Specifically:
1. **`max_num_seqs > 1` over-predicts**. My KV pool formula is `max_ctx × max_num_seqs × per_token_bytes` — but vLLM doesn't always allocate the full pool. The engine internally rate-limits based on `mem_util × VRAM - other_stuff`. For configs with `max_num_seqs=2-4`, the calculator may say FAIL when reality is TIGHT-PASS.
1. **KV pool capping (resolved 2026-05-13)**. Earlier versions of this calculator over-predicted FAIL on configs with `max_num_seqs > 1` because the requested KV pool exceeded available budget. The current calculator explicitly models vLLM's PagedAttention capping: the predicted KV pool is `min(requested, budget - fixed_components)`. When the request exceeds available, the verdict is `TIGHT` with a note that effective concurrency at `--max-num-seqs` may be lower than requested at full `max_ctx`. The "predicted total" in TIGHT cases equals the budget exactly — that's the saturating-allocator behavior, not a modeling artifact.
2. **Activation coefficient varies by `chunk_size` and `dtype`**. We use the fla default `chunk_size=256` and `mamba_ssm_dtype=float32` (per Qwen3.6-27B config.json). If those change, the coefficient needs re-calibration.
2. **Activation coefficient varies by `chunk_size` and `dtype`**. We use the fla default `chunk_size=256` and `mamba_ssm_dtype=float32` (per Qwen3.6-27B config.json). If those change, the coefficient needs re-calibration. For Gemma the activation peak is a flat empirical constant; if Gemma config changes (e.g. layer-pattern ratio, sliding_window), recalibrate.
3. **No driver/allocator overhead modeling**. snoby's 4090 needed `max-model-len` 200K → 180K vs 3090 baseline. The driver-class delta isn't modeled here. We hand-wave with the `±1.5 GB` error band.
4. **No Cliff 2b accumulation modeling**. The multi-turn fragmentation cliff at ~25K accumulated tokens is empirical-only and not in this calculator. Use `SOAK_MODE=continuous` to probe it.
5. **Single-model**. The math is Qwen3.6-27B-specific. Adding a model means deriving a new `MODEL_SPEC` block (architecture params + weights size) and re-calibrating the activation coefficient against new measured points.
5. **Two-model calibration**. Today the math is calibrated for Qwen3.6-27B and Gemma 4 31B. Adding a third model means deriving a new `MODEL_SPEC` block (architecture params + per-quant weights size + activation-peak mechanism) and calibrating the activation coefficient against ≥4 measured BENCHMARKS rows for that model. Don't ship a third-model spec without that calibration — uncalibrated coefficients give wrong verdicts.
## Calibration
Run `bash tools/kv-calc.py --calibration` to see predicted vs measured for all shipped composes. Current verdict accuracy: **9/11 = 82%** with the ±1.5 GB error band. The two ⨯ cases are over-predictions on `max_num_seqs > 1` configs (limitation #1 above).
Run `bash tools/kv-calc.py --calibration` to see predicted vs measured for all shipped composes, grouped per model.
| Model | Verdict accuracy | Notes |
|---|---|---|
| Qwen 3.6 27B | 9/11 = 82% (±1.5 GB band) | Two ⨯ cases are over-predictions on `max_num_seqs > 1` (limitation #1 above) |
| Gemma 4 31B | TBD by `--calibration` | Calibrated against `dual/int8.yml` 98K+262K rows, `dual/dflash.yml` 32K BF16, `dual/awq.yml` 65K BF16, `dual/docker-compose.yml` 32K BF16. Target: ≥80% within ±1.5 GB. |
## When to trust the calculator vs vLLM's boot log
Always pass `--model {qwen3.6-27b,gemma-4-31b}` matching the compose you're targeting. Defaults to qwen3.6-27b if omitted.
| Question | Use this |
|---|---|
| "Will it boot?" — for a *shipped* compose on canonical 24 GB | We've already validated; check BENCHMARKS.md |
| "Will it boot?" — for a *novel* config (custom ctx, kv format, or VRAM class) | `kv-calc.py --compose <X>` for a directional answer; then boot and read `gpu_worker.py` |
| "What's my max ctx?" — given my hardware | `kv-calc.py --solve-max-ctx ...` for an estimate; vLLM's pre-check `estimated max model length is N` line at boot is authoritative |
| "Is TQ3 or fp8 better for my hardware?" | `kv-calc.py` with both options to see the trade-off; cross-check against [HARDWARE.md](HARDWARE.md#note-for-sub-24-gb-cards) for the published guidance |
| "Will it boot?" — for a *novel* config (custom ctx, kv format, or VRAM class) | `kv-calc.py --model <M> --compose <X>` for a directional answer; then boot and read `gpu_worker.py` |
| "What's my max ctx?" — given my hardware | `kv-calc.py --model <M> --solve-max-ctx ...` for an estimate; vLLM's pre-check `estimated max model length is N` line at boot is authoritative |
| "Is TQ3 or fp8 better for my hardware?" (Qwen 3.6) | `kv-calc.py --model qwen3.6-27b` with both options; cross-check [HARDWARE.md](HARDWARE.md#note-for-sub-24-gb-cards) |
| "Is INT8 PTH or BF16 KV better for Gemma 4?" | `kv-calc.py --model gemma-4-31b --kv-format bf16` vs `int8_per_token_head` — BF16 caps at ~32K on dual-3090, INT8 PTH unlocks 262K. See `models/gemma-4-31b/vllm/compose/dual/int8.yml` header. |
## References
**Qwen 3.6 27B (DeltaNet hybrid):**
- [PerfMamba: Performance Analysis and Pruning of Selective State Space Models (arxiv 2511.22849)](https://arxiv.org/html/2511.22849) — block-wise state materialization scaling
- [TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (arxiv 2504.19874, ICLR 2026)](https://arxiv.org/abs/2504.19874) — TQ3 byte savings + technique
- [Gated Delta Networks: Improving Mamba2 with Delta Rule (NVlabs ICLR 2025)](https://github.com/NVlabs/GatedDeltaNet) — Qwen3-Next architecture
- [Mamba: Linear-Time Sequence Modeling (arxiv 2312.00752)](https://arxiv.org/abs/2312.00752) — Mamba-1 baseline for PerfMamba's deltas
**Gemma 4 31B (sliding-window + dense MLP):**
- Architecture params sourced directly from `config.json` (Gemma 4 release post / technical doc were not used as a calibration reference — the activation coefficient is empirical-only on this stack).
- [vLLM PR #40391 (rebased + vendored as PR #42102)](https://github.com/vllm-project/vllm/pull/42102) — per-token-head INT8 KV cache (the Ampere unlock for Gemma 4 at 262K)
- [vLLM PR #41745](https://github.com/vllm-project/vllm/pull/41745) — Gemma 4 MTP assistant drafter support
**Shared:**
- [TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (arxiv 2504.19874, ICLR 2026)](https://arxiv.org/abs/2504.19874) — TQ3 byte savings + technique
- [Efficient Memory Management for Large Language Model Serving with PagedAttention (arxiv 2309.06180)](https://arxiv.org/abs/2309.06180) — vLLM's KV pool allocator
- [An Investigation of FP8 Across Accelerators for LLM Inference (arxiv 2502.01070)](https://arxiv.org/html/2502.01070v1) — FP8 e5m2/e4m3 KV cache analysis
- [docs/CLIFFS.md](CLIFFS.md) — Cliff 2 mechanism + KV-format-tunability section
- [docs/HARDWARE.md](HARDWARE.md) — 20 GB Ampere TQ3→fp8 swap rule (cross-rig validated by @efschu)
- [docs/CLIFFS.md](CLIFFS.md) — Cliff 2 mechanism + KV-format-tunability section (Qwen-specific)
- [docs/HARDWARE.md](HARDWARE.md) — 20 GB Ampere TQ3→fp8 swap rule (cross-rig validated by @efschu, Qwen-specific)
## See also
+1 -1
View File
@@ -204,7 +204,7 @@ Generally prefer **dropping `MAX_MODEL_LEN` first** (clean KV budget reduction,
# 1. Setup (downloads model, clones Genesis, ~20 min cold)
bash scripts/setup.sh qwen3.6-27b
# 2. Pick + boot via wizard (asks engine + workload)
# 2. Pick + boot via wizard (asks model + GPUs, projects VRAM budget)
bash scripts/launch.sh
# 3. Or skip the wizard:
+4 -1
View File
@@ -60,6 +60,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-full
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -118,7 +119,9 @@ services:
- --served-model-name
- gemma-4-31b-awq
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --dtype
- "${DTYPE:-bfloat16}"
- --disable-custom-all-reduce
@@ -91,6 +91,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -151,7 +152,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# ---- KV format ----
# Default `auto` runs on every consumer GPU including
@@ -95,6 +95,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -183,7 +184,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --dtype
- bfloat16
- --disable-custom-all-reduce
@@ -67,6 +67,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -148,7 +149,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# DFlash drafter is BF16; --dtype bfloat16 matches its training dtype
# (same as our existing dual-dflash.yml on Qwen3.6 — vllm#40334 dtype-mismatch
# workaround until that lands).
@@ -50,6 +50,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -99,7 +100,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# max-model-len ceiling (TP=2, BF16 KV):
# 32K ctx @ 0.92 → KV pool ~38K tokens (shipped default — matches bench above)
@@ -179,6 +179,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-full
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 9.0+
@@ -268,7 +269,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# ---- 2026-05-11 workaround flags ----
# Force TURBOQUANT backend — required because vLLM has no per-layer
@@ -91,6 +91,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-full
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -151,7 +152,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
# ---- KV format ----
# Default `int8_per_token_head` runs on every consumer GPU including
@@ -50,6 +50,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 32
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
# Requires-sm: 9.0+
@@ -100,7 +101,9 @@ services:
- --served-model-name
- gemma-4-31b-autoround
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
# Single-card: weights ~16 GB + drafter ~1 GB + bf16 KV at 16K + vision +
# activations needs to fit in 24 GB. Drop vision via limit-mm-per-prompt
# to reclaim ~1.5 GB and skip the mm-token-budget assertion.
@@ -9,6 +9,7 @@
# Max ctx: 192K pool / 4 parallel slots
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
# Status: ✅ Production
# Engine-profile: llama-cpp-mainline
# Best for: Single-card multi-tenant llama.cpp — 4 concurrent agents
# at smaller per-stream ctx; trade max-ctx for parallelism
# ---------------------------------------------------------------------------
@@ -9,6 +9,7 @@
# Max ctx: 262K (full model native — no Cliff 1 / Cliff 2)
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
# Status: ✅ Production
# Engine-profile: llama-cpp-mainline
# Best for: Bulletproof single-card path — slow decode (~21 TPS) but
# cliff-immune; recommended fallback when vLLM hits OOM at long ctx
# ---------------------------------------------------------------------------
@@ -22,6 +22,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -99,7 +100,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-200000}"
@@ -50,6 +50,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -128,7 +129,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -45,6 +45,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -122,7 +123,9 @@ services:
- --dtype
- bfloat16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-200000}"
@@ -69,6 +69,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -146,7 +147,9 @@ services:
- --dtype
- bfloat16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-185000}"
@@ -51,6 +51,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -137,7 +138,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -22,6 +22,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-full
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -99,7 +100,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -77,6 +77,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -156,7 +157,9 @@ services:
- --dtype
- bfloat16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual-dflash-noviz.yml disables it
# because PCIe P2P bandwidth makes the NCCL fallback faster there.
@@ -73,6 +73,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -152,7 +153,9 @@ services:
- --dtype
- bfloat16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual-dflash.yml disables it because
# PCIe P2P bandwidth makes the NCCL fallback faster there.
@@ -63,6 +63,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -256,7 +257,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual-turbo.yml disables it because
# PCIe P2P bandwidth makes the NCCL fallback faster there.
@@ -59,6 +59,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -144,7 +145,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
# Custom all-reduce ENABLED (no --disable-custom-all-reduce) — NVLink
# makes vLLM's custom kernel a win. dual.yml disables it because PCIe
# P2P bandwidth makes the NCCL fallback faster there.
@@ -35,6 +35,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -98,7 +99,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -63,6 +63,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -257,7 +258,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -61,6 +61,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -147,7 +148,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -37,6 +37,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -130,7 +131,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -47,6 +47,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
@@ -238,7 +239,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "2"
- "${TP:-2}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -62,6 +62,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-dflash
# Requires-min-gpu-count: 4
# Tensor-parallel: 4
services:
@@ -139,7 +140,9 @@ services:
- --dtype
- bfloat16
- --tensor-parallel-size
- "4"
- "${TP:-4}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -59,6 +59,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 4
# Tensor-parallel: 4
services:
@@ -145,7 +146,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "4"
- "${TP:-4}"
- --pipeline-parallel-size
- "${PP:-1}"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-262144}"
@@ -131,6 +131,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -307,7 +308,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
# 180K + 0.95 — parity with long-text.yml. Backed off from 214K + 0.985
# on 2026-05-02 to give activation headroom for the PN12+PN25 FFN pool
# residence + DeltaNet GDN buffer. See long-text.yml for full rationale.
@@ -97,6 +97,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -238,7 +239,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
# 48K + 0.92 — production default. Below BOTH cliffs:
# - GDN forward (single-prompt OOM at ~50-60K tokens) → 48K stays safely under
# - TurboQuant tool prefill (OOM at high mem_util + big tool responses) → 0.92 leaves ~1.9 GB headroom
@@ -111,6 +111,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -336,7 +337,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
# 180K + 0.97 — backed off from 214K + 0.985 on 2026-05-01 PM after
# verify-stress probe 1 (10K-token long-context needle) crashed with
# GDN forward OOM (`fla.ops.chunk.chunk_gated_delta_rule_fwd_h`
@@ -121,6 +121,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -353,7 +354,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
# 180K + 0.97 — backed off from 214K + 0.985 on 2026-05-01 PM after
# verify-stress probe 1 (10K-token long-context needle) crashed with
# GDN forward OOM (`fla.ops.chunk.chunk_gated_delta_rule_fwd_h`
@@ -97,6 +97,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -273,7 +274,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
# 198K + 0.98 — restored on v0.20 + Genesis v7.65 migration. v0.20
# closes the Cliff 1 mech B sub-cliffs that drove the 140K + 0.95
# backoff on dev205. Verified verify-full + 33K + 50K stress all
@@ -33,6 +33,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 20
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -107,7 +108,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-32768}"
- --gpu-memory-utilization
@@ -35,6 +35,7 @@
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Engine-profile: vllm-nightly-mtp
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
@@ -162,7 +163,9 @@ services:
- --dtype
- float16
- --tensor-parallel-size
- "1"
- "${TP:-1}"
- --pipeline-parallel-size
- "${PP:-1}"
- --max-model-len
- "${MAX_MODEL_LEN:-75000}"
- --gpu-memory-utilization
+721 -91
View File
@@ -1,18 +1,23 @@
#!/usr/bin/env bash
#
# Interactive launcher for club-3090 — pick engine + workload, boot the
# right compose, run verify-full to confirm it's serving.
# Interactive launcher for club-3090 — pick model + GPUs, project the
# VRAM budget, boot the right compose, run verify-full to confirm it's serving.
#
# For first-run users coming in from the README. If you already know
# what you want, use `scripts/switch.sh <variant>` directly.
#
# Usage:
# bash scripts/launch.sh # interactive wizard
# bash scripts/launch.sh # interactive model/GPU wizard
# bash scripts/launch.sh --variant <name> # skip wizard, boot directly
# bash scripts/launch.sh --engine vllm --cards 1 # partial flags, ask the rest
# bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1
# bash scripts/launch.sh --engine vllm --cards 1 # deprecated; prefer --gpus
# bash scripts/launch.sh --tp 2 --pp 1 # override vLLM parallelism
# bash scripts/launch.sh --no-projection # skip kv-calc budget projection
# bash scripts/launch.sh --no-verify # skip post-launch verify-full
# bash scripts/launch.sh --no-preflight # skip docker/GPU pre-flight
#
# The wizard marks variants that don't fit the detected GPUs; direct
# --variant keeps the power-user path and delegates final gating to switch.sh.
# All flags accept the same names as `switch.sh --list` produces.
# Examples:
# bash scripts/launch.sh --variant vllm/default
@@ -24,6 +29,13 @@ set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
SWITCH="${SWITCH:-${ROOT_DIR}/scripts/switch.sh}"
VERIFY="${VERIFY:-${ROOT_DIR}/scripts/verify-full.sh}"
if [[ -z "${MODEL_DIR:-}" && -f "${ROOT_DIR}/.env" ]]; then
set -a
# shellcheck source=/dev/null
source "${ROOT_DIR}/.env"
set +a
fi
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
# shellcheck source=preflight.sh
source "${ROOT_DIR}/scripts/preflight.sh"
@@ -31,13 +43,25 @@ source "${ROOT_DIR}/scripts/preflight.sh"
ENGINE=""
CARDS=""
VARIANT=""
MODEL_NAME=""
GPU_ARG=""
TP_OVERRIDE=""
PP_OVERRIDE=""
PARALLELISM="auto"
SKIP_VERIFY=0
SKIP_PREFLIGHT=0
SKIP_PROJECTION=0
while [[ $# -gt 0 ]]; do
case "$1" in
--engine) ENGINE="$2"; shift 2 ;;
--cards) CARDS="$2"; shift 2 ;;
--variant) VARIANT="$2"; shift 2 ;;
--model) MODEL_NAME="$2"; shift 2 ;;
--gpus) GPU_ARG="$2"; shift 2 ;;
--tp) TP_OVERRIDE="$2"; shift 2 ;;
--pp) PP_OVERRIDE="$2"; shift 2 ;;
--parallelism) PARALLELISM="$2"; shift 2 ;;
--no-projection) SKIP_PROJECTION=1; shift ;;
--no-verify) SKIP_VERIFY=1; shift ;;
--no-preflight) SKIP_PREFLIGHT=1; shift ;;
-h|--help)
@@ -72,6 +96,19 @@ ask() {
fi
}
read_or_interrupt() {
local prompt="$1"
local __reply_var="$2"
local reply
if ! read -rp "$prompt" reply; then
echo "" >&2
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
kill -INT $$
exit 1
fi
printf -v "$__reply_var" '%s' "$reply"
}
choose() {
# choose "prompt" "label1" "value1" "label2" "value2" ... -> echoes chosen value
local prompt="$1"; shift
@@ -87,7 +124,12 @@ choose() {
done
while true; do
local pick
read -rp "Choice [1-${#labels[@]}]: " pick
if ! read -rp "Choice [1-${#labels[@]}]: " pick; then
echo "" >&2
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
kill -INT $$
exit 1
fi
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
echo "${values[$((pick-1))]}"
return
@@ -96,109 +138,691 @@ choose() {
done
}
# --- wizard ---
# Flow: cards → workload → auto-pick engine. Newcomers can answer "how
# many GPUs" and "what do I want to do" but rarely "vLLM or llama.cpp" —
# so the engine is derived from the workload pick, not asked first.
# Manual --engine override filters the workload list to that engine.
if [[ -z "$VARIANT" ]]; then
echo ""
echo "club-3090 launcher — let's pick the right config for your workload."
echo "(Use --variant <name> next time to skip the wizard.)"
echo ""
declare -A LAUNCH_VARIANT_COMPOSE=(
[vllm/default]="models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
[vllm/long-vision]="models/qwen3.6-27b/vllm/compose/single/long-vision.yml"
[vllm/long-text]="models/qwen3.6-27b/vllm/compose/single/long-text.yml"
[vllm/long-text-no-mtp]="models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml"
[vllm/bounded-thinking]="models/qwen3.6-27b/vllm/compose/single/bounded-thinking.yml"
[vllm/tools-text]="models/qwen3.6-27b/vllm/compose/single/tools-text.yml"
[vllm/minimal]="models/qwen3.6-27b/vllm/compose/single/minimal.yml"
[vllm/dual]="models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml"
[vllm/dual4]="models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml"
[vllm/dual4-dflash]="models/qwen3.6-27b/vllm/compose/multi4/dflash.yml"
[vllm/dual-turbo]="models/qwen3.6-27b/vllm/compose/dual/turbo.yml"
[vllm/dual-dflash]="models/qwen3.6-27b/vllm/compose/dual/dflash.yml"
[vllm/dual-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/dflash-noviz.yml"
[vllm/dual-nvlink]="models/qwen3.6-27b/vllm/compose/dual/nvlink.yml"
[vllm/dual-nvlink-turbo]="models/qwen3.6-27b/vllm/compose/dual/nvlink-turbo.yml"
[vllm/dual-nvlink-dflash]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash.yml"
[vllm/dual-nvlink-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash-noviz.yml"
[vllm/gemma-mtp]="models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
[vllm/gemma-mtp-tp1]="models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
[vllm/gemma-dflash]="models/gemma-4-31b/vllm/compose/dual/dflash.yml"
[llamacpp/default]="models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml"
[llamacpp/concurrent]="models/qwen3.6-27b/llama-cpp/compose/single/concurrent.yml"
)
declare -A LAUNCH_VARIANT_MODEL=(
[vllm/default]="qwen3.6-27b" [vllm/long-vision]="qwen3.6-27b" [vllm/long-text]="qwen3.6-27b"
[vllm/long-text-no-mtp]="qwen3.6-27b" [vllm/bounded-thinking]="qwen3.6-27b" [vllm/tools-text]="qwen3.6-27b"
[vllm/minimal]="qwen3.6-27b" [vllm/dual]="qwen3.6-27b" [vllm/dual4]="qwen3.6-27b"
[vllm/dual4-dflash]="qwen3.6-27b" [vllm/dual-turbo]="qwen3.6-27b" [vllm/dual-dflash]="qwen3.6-27b"
[vllm/dual-dflash-noviz]="qwen3.6-27b" [vllm/dual-nvlink]="qwen3.6-27b" [vllm/dual-nvlink-turbo]="qwen3.6-27b"
[vllm/dual-nvlink-dflash]="qwen3.6-27b" [vllm/dual-nvlink-dflash-noviz]="qwen3.6-27b"
[vllm/gemma-mtp]="gemma-4-31b" [vllm/gemma-mtp-tp1]="gemma-4-31b" [vllm/gemma-dflash]="gemma-4-31b"
[llamacpp/default]="qwen3.6-27b" [llamacpp/concurrent]="qwen3.6-27b"
)
declare -A LAUNCH_VARIANT_ENGINE=(
[vllm/default]="vllm" [vllm/long-vision]="vllm" [vllm/long-text]="vllm" [vllm/long-text-no-mtp]="vllm"
[vllm/bounded-thinking]="vllm" [vllm/tools-text]="vllm" [vllm/minimal]="vllm" [vllm/dual]="vllm"
[vllm/dual4]="vllm" [vllm/dual4-dflash]="vllm" [vllm/dual-turbo]="vllm" [vllm/dual-dflash]="vllm"
[vllm/dual-dflash-noviz]="vllm" [vllm/dual-nvlink]="vllm" [vllm/dual-nvlink-turbo]="vllm"
[vllm/dual-nvlink-dflash]="vllm" [vllm/dual-nvlink-dflash-noviz]="vllm"
[vllm/gemma-mtp]="vllm" [vllm/gemma-mtp-tp1]="vllm" [vllm/gemma-dflash]="vllm"
[llamacpp/default]="llamacpp" [llamacpp/concurrent]="llamacpp"
)
declare -A LAUNCH_VARIANT_KVCALC=(
[vllm/default]="qwen3.6-27b:long-vision"
[vllm/long-text]="qwen3.6-27b:long-text"
[vllm/long-text-no-mtp]="qwen3.6-27b:long-text-no-mtp"
[vllm/long-vision]="qwen3.6-27b:long-vision"
[vllm/bounded-thinking]="qwen3.6-27b:bounded-thinking"
[vllm/tools-text]="qwen3.6-27b:tools-text"
[vllm/minimal]="qwen3.6-27b:minimal"
[vllm/dual]="qwen3.6-27b:dual"
[vllm/dual-turbo]="qwen3.6-27b:dual-turbo"
[vllm/dual-dflash]="qwen3.6-27b:dual-dflash"
[vllm/dual-dflash-noviz]="qwen3.6-27b:dual-dflash-noviz"
[vllm/dual4]="qwen3.6-27b:dual4"
[vllm/dual4-dflash]="qwen3.6-27b:dual4-dflash"
[vllm/dual-nvlink]="qwen3.6-27b:dual"
[vllm/dual-nvlink-turbo]="qwen3.6-27b:dual-turbo"
[vllm/dual-nvlink-dflash]="qwen3.6-27b:dual-dflash"
[vllm/dual-nvlink-dflash-noviz]="qwen3.6-27b:dual-dflash-noviz"
[vllm/gemma-mtp]="gemma-4-31b:gemma-dual"
[vllm/gemma-mtp-tp1]="gemma-4-31b:gemma-single"
[vllm/gemma-dflash]="gemma-4-31b:gemma-dual-dflash"
[llamacpp/default]="SKIP"
[llamacpp/concurrent]="SKIP"
)
LAUNCH_VARIANT_ORDER=(
vllm/long-vision vllm/long-text vllm/long-text-no-mtp vllm/bounded-thinking
vllm/default vllm/tools-text vllm/minimal
vllm/dual vllm/dual-turbo vllm/dual-dflash vllm/dual-dflash-noviz
vllm/dual4 vllm/dual4-dflash
vllm/gemma-mtp vllm/gemma-mtp-tp1 vllm/gemma-dflash
llamacpp/default llamacpp/concurrent
)
# Step 1 — cards.
if [[ -z "$CARDS" ]]; then
CARDS=$(choose "How many RTX 3090s?" \
"1× 3090 (24 GB)" "1" \
"2× 3090 (PCIe / no NVLink)" "2")
variant_hw_status() {
local variant="$1"
local rel="${LAUNCH_VARIANT_COMPOSE[$variant]:-}"
if [[ -z "$rel" ]]; then
printf 'ok|fits your rig'
return 0
fi
# Re-validate now that we know the requirement (preflight already ran
# with min=1; bump to actual count). Skip if --no-preflight.
if [[ $SKIP_PREFLIGHT -eq 0 ]]; then
preflight_gpu "$CARDS" || exit 1
local compose_file="${ROOT_DIR}/${rel}"
if [[ ! -f "$compose_file" ]]; then
printf 'unknown|compose metadata unavailable'
return 2
fi
compose_hw_compose_status "$compose_file" 2>/dev/null || true
}
choose_variant() {
# choose_variant "prompt" "default-variant" "label1" "value1" ...
local prompt="$1" default_variant="$2"
shift 2
local i labels=() values=() statuses=() eligible=()
while [[ $# -gt 0 ]]; do
labels+=("$1")
values+=("$2")
statuses+=("$(variant_hw_status "$2")")
shift 2
done
local default_idx=""
for i in "${!values[@]}"; do
if [[ "${values[$i]}" == "$default_variant" && "${statuses[$i]}" == ok\|* ]]; then
default_idx=$((i + 1))
break
fi
done
if [[ -z "$default_idx" ]]; then
for i in "${!values[@]}"; do
if [[ "${statuses[$i]}" == ok\|* || "${statuses[$i]}" == unknown\|* ]]; then
default_idx=$((i + 1))
break
fi
done
fi
# Step 2 — workload, filtered by cards (and --engine override if set).
# Each option's value is "engine/file" so engine is implied by the pick.
if [[ "$CARDS" == "1" ]]; then
# Primary recommended options first; diagnostic / niche variants in
# an "Other" group at the end. The only single-card limitation users
# need to know: vLLM single-card crashes on a single prompt >50K
# (Cliff 2). For unpredictable inputs, use llamacpp/default.
if [[ -z "$ENGINE" || "$ENGINE" == "vllm" ]]; then
VLLM_OPTS=(
"Long ctx + vision (145K + vision, MTP) — recommended for chat/agents" "vllm/long-vision"
"Long ctx, text only — Balanced MTP (180K, MTP K=3) — recommended IDE-agent" "vllm/long-text"
"Long ctx, text only — Max-context (200K, no MTP) — one-shot >50K prompts" "vllm/long-text-no-mtp"
"Bounded thinking (180K, structured-CoT FSM — recommended grammar: DeepSeek scratchpad, 87.4% combined HE+/LCB v6)" "vllm/bounded-thinking"
)
echo "" >&2
echo "$prompt" >&2
for i in "${!labels[@]}"; do
local status="${statuses[$i]}"
local state="${status%%|*}"
local reason="${status#*|}"
local marker="✓"
case "$state" in
ok) marker="✓" ;;
unknown) marker="?" ;;
*) marker="✗" ;;
esac
if [[ -n "$default_idx" && $((i + 1)) -eq "$default_idx" ]]; then
printf " %d) %s %s %s [default]\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
else
VLLM_OPTS=()
printf " %d) %s %s %s\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
fi
if [[ -z "$ENGINE" || "$ENGINE" == "llamacpp" ]]; then
LLAMA_OPTS=(
"Bulletproof, no cliffs (262K + vision, ~21 TPS) — production-safe" "llamacpp/default"
)
else
LLAMA_OPTS=()
fi
# Diagnostic / niche fallbacks — shown last so they don't dominate the menu
if [[ -z "$ENGINE" || "$ENGINE" == "vllm" ]]; then
VLLM_FALLBACK_OPTS=(
"[fallback] Default 48K + vision (Cliff 2 unreachable; fast boot)" "vllm/default"
"[fallback] tools-text 75K FP8 (FP8 KV alternative for accuracy compare)" "vllm/tools-text"
"[fallback] minimal 32K (no Genesis, no spec-decode — diagnostic stack)" "vllm/minimal"
)
else
VLLM_FALLBACK_OPTS=()
fi
if [[ -z "$ENGINE" || "$ENGINE" == "llamacpp" ]]; then
LLAMA_FALLBACK_OPTS=(
"[fallback] llamacpp/concurrent (4 parallel slots, 192K pool, vision)" "llamacpp/concurrent"
)
else
LLAMA_FALLBACK_OPTS=()
fi
VARIANT=$(choose "What's your main workload?" \
"${VLLM_OPTS[@]}" "${LLAMA_OPTS[@]}" \
"${VLLM_FALLBACK_OPTS[@]}" "${LLAMA_FALLBACK_OPTS[@]}")
elif [[ "$CARDS" == "2" ]]; then
if [[ -n "$ENGINE" && "$ENGINE" != "vllm" ]]; then
echo "ERROR: --engine ${ENGINE} not supported on 2× cards (no llama.cpp dual recipe yet)." >&2
exit 1
fi
VARIANT=$(choose "What's your dual-card priority?" \
"Balanced default — 262K + vision + 2 streams (recommended)" "vllm/dual" \
"Multi-tenant — 4 concurrent streams @ 262K, TQ3 KV" "vllm/dual-turbo" \
"Peak code TPS with vision (185K, DFlash N=5)" "vllm/dual-dflash" \
"Peak code TPS no vision (200K, DFlash N=5)" "vllm/dual-dflash-noviz")
else
echo "ERROR: --cards ${CARDS} unsupported (expected 1 or 2)." >&2
done
if [[ -z "$default_idx" ]]; then
echo "ERROR: no eligible variants in this menu. Use scripts/switch.sh --force <variant> to attempt anyway." >&2
exit 1
fi
# Step 3 — explain the auto-picked engine.
echo ""
case "$VARIANT" in
llamacpp/*)
echo "[wizard] picked llama.cpp — chosen because no prefill cliffs at 262K and the simplest"
echo "[wizard] serving path. Trade-off: ~21 TPS vs 70+ on vLLM. Right call for long-prompt"
echo "[wizard] robustness, frontier context, or anyone who wants the simplest setup."
;;
vllm/*)
echo "[wizard] picked vLLM — chosen for spec-decode (MTP), tool-call extraction, and best TPS."
echo "[wizard] Watch out for the prefill cliffs (Cliff 1 = 25K+ tool returns; Cliff 2 = 50-60K"
echo "[wizard] single prompts on TQ3) — see docs/SINGLE_CARD.md for the safe-config map."
;;
while true; do
local pick
if ! read -rp "Choice [1-${#labels[@]}, default ${default_idx}]: " pick; then
echo "" >&2
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
kill -INT $$
exit 1
fi
pick="${pick:-$default_idx}"
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
local status="${statuses[$((pick - 1))]}"
if [[ "$status" == no\|* ]]; then
echo " That variant won't run on your detected rig: ${status#*|}" >&2
echo " Pick another, or use: bash scripts/switch.sh --force ${values[$((pick - 1))]}" >&2
continue
fi
echo "${values[$((pick - 1))]}"
return
fi
echo " invalid — pick a number 1-${#labels[@]}" >&2
done
}
model_label() {
case "$1" in
qwen3.6-27b) echo "Qwen 3.6 27B" ;;
gemma-4-31b) echo "Gemma 4 31B" ;;
*) echo "$1" ;;
esac
}
normalize_model_name() {
case "$1" in
qwen3.6-27b|qwen3.6-27b-gguf) echo "qwen3.6-27b" ;;
gemma-4-31b|gemma-4-31b-awq|gemma-4-31b-gguf) echo "gemma-4-31b" ;;
*) echo "$1" ;;
esac
}
MODEL_ORDER=()
declare -A MODEL_ENGINES=()
add_installed_model_engine() {
local model="$1" engine="$2"
if [[ -z "${MODEL_ENGINES[$model]:-}" ]]; then
MODEL_ORDER+=("$model")
MODEL_ENGINES[$model]="$engine"
elif [[ ",${MODEL_ENGINES[$model]}," != *",${engine},"* ]]; then
MODEL_ENGINES[$model]="${MODEL_ENGINES[$model]},${engine}"
fi
}
detect_installed_models() {
MODEL_ORDER=()
MODEL_ENGINES=()
[[ -d "${MODEL_DIR}/qwen3.6-27b-autoround-int4" ]] && add_installed_model_engine "qwen3.6-27b" "vllm"
[[ -d "${MODEL_DIR}/qwen3.6-27b-gguf" ]] && add_installed_model_engine "qwen3.6-27b" "llamacpp"
[[ -d "${MODEL_DIR}/gemma-4-31b-autoround-int4" ]] && add_installed_model_engine "gemma-4-31b" "vllm"
[[ -d "${MODEL_DIR}/gemma-4-31b-it-AWQ-4bit" ]] && add_installed_model_engine "gemma-4-31b" "vllm"
[[ -d "${MODEL_DIR}/gemma-4-31b-gguf" ]] && add_installed_model_engine "gemma-4-31b" "llamacpp"
return 0
}
model_has_engine() {
local model="$1" engine="$2"
[[ ",${MODEL_ENGINES[$model]:-}," == *",${engine},"* ]]
}
engine_hint() {
case "$1" in
vllm,llamacpp|llamacpp,vllm) echo "vLLM + llama.cpp engines available" ;;
vllm) echo "vLLM only" ;;
llamacpp) echo "llama.cpp only" ;;
*) echo "$1" ;;
esac
}
choose_model() {
detect_installed_models
if [[ -n "$MODEL_NAME" ]]; then
MODEL_NAME="$(normalize_model_name "$MODEL_NAME")"
if [[ -z "${MODEL_ENGINES[$MODEL_NAME]:-}" ]]; then
echo "[launch] ERROR: ${MODEL_NAME} is not installed under ${MODEL_DIR}." >&2
echo "[launch] Run: bash scripts/setup.sh ${MODEL_NAME}" >&2
exit 1
fi
return
fi
if [[ "${#MODEL_ORDER[@]}" -eq 0 ]]; then
echo "[launch] ERROR: no supported model weights found under ${MODEL_DIR}." >&2
echo "[launch] Run: bash scripts/setup.sh" >&2
exit 1
fi
if [[ "${#MODEL_ORDER[@]}" -eq 1 ]]; then
MODEL_NAME="${MODEL_ORDER[0]}"
echo "[launch] using installed model: $(model_label "$MODEL_NAME")" >&2
return
fi
echo "" >&2
echo "[launch] Installed models:" >&2
local i=1 model
for model in "${MODEL_ORDER[@]}"; do
printf " %d) %-16s (%s)\n" "$i" "$(model_label "$model")" "$(engine_hint "${MODEL_ENGINES[$model]}")" >&2
i=$((i + 1))
done
while true; do
local pick
read_or_interrupt "Choice [1-${#MODEL_ORDER[@]}]: " pick
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#MODEL_ORDER[@]} )); then
MODEL_NAME="${MODEL_ORDER[$((pick - 1))]}"
return
fi
echo " invalid — pick a number 1-${#MODEL_ORDER[@]}" >&2
done
}
GPU_LINES=""
CARD_INDICES=()
CARD_NAMES=()
CARD_MEM_MIB=()
CARD_SM=()
MIN_VRAM_GB=0
MAX_VRAM_GB=0
HET_VRAM_MIXED=0
SELECTED_GPU_CSV=""
SELECTED_VRAM_SUMMARY=""
gpu_exists() {
local want="$1" idx name mem_mib sm
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ "$idx" == "$want" ]] && return 0
done <<< "$GPU_LINES"
return 1
}
gpu_is_busy() {
local want="$1" busy
while IFS= read -r busy; do
[[ "$busy" == "$want" ]] && return 0
done <<< "$(compose_hw_in_use_gpus 2>/dev/null || true)"
return 1
}
append_selected_gpu() {
local want="$1" idx name mem_mib sm
while IFS=$'\t' read -r idx name mem_mib sm; do
if [[ "$idx" == "$want" ]]; then
CARD_INDICES+=("$idx")
CARD_NAMES+=("$name")
CARD_MEM_MIB+=("$mem_mib")
CARD_SM+=("$sm")
return 0
fi
done <<< "$GPU_LINES"
return 1
}
select_gpus_from_arg() {
local arg="$1"
CARD_INDICES=()
CARD_NAMES=()
CARD_MEM_MIB=()
CARD_SM=()
local available=() idx name mem_mib sm
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
if ! gpu_is_busy "$idx"; then
available+=("$idx")
fi
done <<< "$GPU_LINES"
if [[ "$arg" == "all" ]]; then
[[ "${#available[@]}" -gt 0 ]] || { echo "[launch] ERROR: no available NVIDIA GPUs detected." >&2; exit 1; }
for idx in "${available[@]}"; do append_selected_gpu "$idx"; done
return
fi
IFS=',' read -ra _launch_gpu_tokens <<< "$arg"
for idx in "${_launch_gpu_tokens[@]}"; do
idx="$(_compose_meta_trim "$idx")"
[[ -z "$idx" ]] && continue
gpu_exists "$idx" || { echo "[launch] ERROR: requested GPU ${idx}, but it was not detected." >&2; exit 1; }
append_selected_gpu "$idx"
done
}
summarize_selected_vram() {
local parts=() i gb
MIN_VRAM_GB=0
MAX_VRAM_GB=0
HET_VRAM_MIXED=0
for i in "${!CARD_INDICES[@]}"; do
gb="$(compose_hw_vram_gb "${CARD_MEM_MIB[$i]}")"
parts+=("${gb} GB")
if [[ "$MIN_VRAM_GB" -eq 0 || "$gb" -lt "$MIN_VRAM_GB" ]]; then MIN_VRAM_GB="$gb"; fi
if [[ "$gb" -gt "$MAX_VRAM_GB" ]]; then MAX_VRAM_GB="$gb"; fi
done
[[ "$MIN_VRAM_GB" != "$MAX_VRAM_GB" ]] && HET_VRAM_MIXED=1
local joined="" part
for part in "${parts[@]}"; do
[[ -n "$joined" ]] && joined="${joined} + "
joined="${joined}${part}"
done
SELECTED_VRAM_SUMMARY="$joined"
printf '%s' "$joined"
}
choose_gpus() {
GPU_LINES="$(compose_hw_detect_gpus 2>/dev/null || true)"
[[ -n "$GPU_LINES" ]] || { echo "[launch] ERROR: no NVIDIA GPUs detected." >&2; exit 1; }
local available=() idx name mem_mib sm state
echo "" >&2
echo "[launch] Detected GPUs:" >&2
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
state="available"
if gpu_is_busy "$idx"; then
state="in-use (skipped)"
else
available+=("$idx")
fi
printf " GPU %s: %s (%s GB, sm_%s) — %s\n" "$idx" "${name#NVIDIA }" "$(compose_hw_vram_gb "$mem_mib")" "${sm/./}" "$state" >&2
done <<< "$GPU_LINES"
if [[ -n "$GPU_ARG" ]]; then
select_gpus_from_arg "$GPU_ARG"
elif [[ -n "$CARDS" ]]; then
[[ "$CARDS" =~ ^[0-9]+$ && "$CARDS" -ge 1 ]] || { echo "[launch] ERROR: --cards expects a positive integer." >&2; exit 1; }
(( CARDS <= ${#available[@]} )) || { echo "[launch] ERROR: --cards ${CARDS} requested, but only ${#available[@]} GPU(s) are available." >&2; exit 1; }
local i
for ((i = 0; i < CARDS; i++)); do append_selected_gpu "${available[$i]}"; done
else
case "${#available[@]}" in
0) echo "[launch] ERROR: no available NVIDIA GPUs detected." >&2; exit 1 ;;
1) select_gpus_from_arg "${available[0]}" ;;
2)
local pick
read_or_interrupt "Use GPU ${available[0]}, GPU ${available[1]}, or both? [both]: " pick
pick="${pick:-both}"
case "$pick" in
both|all) select_gpus_from_arg "${available[0]},${available[1]}" ;;
"${available[0]}"|"${available[1]}") select_gpus_from_arg "$pick" ;;
*) echo "[launch] ERROR: invalid GPU choice: $pick" >&2; exit 1 ;;
esac
;;
*)
local default_csv pick
default_csv="$(IFS=','; echo "${available[*]}")"
read_or_interrupt "Which GPU(s)? (comma-separated indices, or 'all') [all]: " pick
pick="${pick:-all}"
[[ "$pick" == "all" ]] && pick="$default_csv"
select_gpus_from_arg "$pick"
;;
esac
fi
[[ "${#CARD_INDICES[@]}" -gt 0 ]] || { echo "[launch] ERROR: no GPUs selected." >&2; exit 1; }
SELECTED_GPU_CSV="$(IFS=','; echo "${CARD_INDICES[*]}")"
summarize_selected_vram >/dev/null
echo "[launch] selected GPU(s): ${SELECTED_GPU_CSV} (${SELECTED_VRAM_SUMMARY})" >&2
}
valid_tp_values() {
case "$1" in
qwen3.6-27b) echo "1 2 4" ;;
gemma-4-31b) echo "1 2 4 8 16" ;;
*) echo "1" ;;
esac
}
tp_is_valid_for_model() {
local model="$1" tp="$2" v
for v in $(valid_tp_values "$model"); do [[ "$v" == "$tp" ]] && return 0; done
return 1
}
largest_valid_tp_for_cards() {
local model="$1" cards="$2" v best=1
for v in $(valid_tp_values "$model"); do
if (( v <= cards && cards % v == 0 && v > best )); then best="$v"; fi
done
echo "$best"
}
TP_VALUE=""
PP_VALUE=""
pick_parallelism() {
local cards="${#CARD_INDICES[@]}"
local tp_set=0 pp_set=0
[[ -n "$TP_OVERRIDE" ]] && tp_set=1
[[ -n "$PP_OVERRIDE" ]] && pp_set=1
[[ -z "$TP_OVERRIDE" || "$TP_OVERRIDE" =~ ^[0-9]+$ ]] || { echo "[launch] ERROR: --tp expects an integer." >&2; exit 1; }
[[ -z "$PP_OVERRIDE" || "$PP_OVERRIDE" =~ ^[0-9]+$ ]] || { echo "[launch] ERROR: --pp expects an integer." >&2; exit 1; }
case "$PARALLELISM" in auto|tp|pp) ;; *) echo "[launch] ERROR: --parallelism expects auto, tp, or pp." >&2; exit 1 ;; esac
if (( cards == 1 )); then
TP_VALUE="${TP_OVERRIDE:-1}"
PP_VALUE="${PP_OVERRIDE:-1}"
elif (( tp_set == 1 && pp_set == 1 )); then
TP_VALUE="$TP_OVERRIDE"; PP_VALUE="$PP_OVERRIDE"
elif (( tp_set == 1 )); then
TP_VALUE="$TP_OVERRIDE"
(( cards % TP_VALUE == 0 )) || { echo "[launch] ERROR: --tp ${TP_VALUE} does not divide selected GPU count ${cards}." >&2; exit 1; }
PP_VALUE=$(( cards / TP_VALUE ))
elif (( pp_set == 1 )); then
PP_VALUE="$PP_OVERRIDE"
(( cards % PP_VALUE == 0 )) || { echo "[launch] ERROR: --pp ${PP_VALUE} does not divide selected GPU count ${cards}." >&2; exit 1; }
TP_VALUE=$(( cards / PP_VALUE ))
elif [[ "$PARALLELISM" == "pp" ]]; then
TP_VALUE=1; PP_VALUE="$cards"
elif [[ "$PARALLELISM" == "tp" ]]; then
TP_VALUE="$cards"; PP_VALUE=1
elif (( HET_VRAM_MIXED == 1 )); then
TP_VALUE=1; PP_VALUE="$cards"
echo "[launch] ${cards} GPUs selected (${MIN_VRAM_GB} GB + ${MAX_VRAM_GB} GB — heterogeneous)." >&2
echo "[launch] Recommended: pipeline parallel PP=${PP_VALUE} to avoid bottlenecking on the smallest card." >&2
else
TP_VALUE="$(largest_valid_tp_for_cards "$MODEL_NAME" "$cards")"
PP_VALUE=$(( cards / TP_VALUE ))
fi
(( TP_VALUE * PP_VALUE == cards )) || { echo "[launch] ERROR: TP × PP must equal selected GPU count (${TP_VALUE} × ${PP_VALUE} != ${cards})." >&2; exit 1; }
if ! tp_is_valid_for_model "$MODEL_NAME" "$TP_VALUE"; then
echo "[launch] ERROR: $(model_label "$MODEL_NAME") num_kv_heads does not divide TP=${TP_VALUE}." >&2
echo "[launch] Valid TP values: $(valid_tp_values "$MODEL_NAME")" >&2
exit 1
fi
if (( cards > 1 && PP_VALUE == 1 )); then
echo "[launch] Tensor parallel TP=${TP_VALUE} (PP=1)." >&2
elif (( PP_VALUE > 1 )); then
echo "[launch] Pipeline parallel PP=${PP_VALUE}, TP=${TP_VALUE}." >&2
echo "[launch] WARN: pipeline parallel is experimental on this stack — no benchmarks yet." >&2
fi
}
variant_min_gpu_count() {
local rel="${LAUNCH_VARIANT_COMPOSE[$1]:-}" value
[[ -n "$rel" && -f "${ROOT_DIR}/${rel}" ]] || { echo 1; return; }
value="$(compose_meta_get "${ROOT_DIR}/${rel}" requires-min-gpu-count || true)"
echo "${value:-1}"
}
variant_min_vram_gb() {
local rel="${LAUNCH_VARIANT_COMPOSE[$1]:-}" value
[[ -n "$rel" && -f "${ROOT_DIR}/${rel}" ]] || { echo 0; return; }
value="$(compose_meta_get "${ROOT_DIR}/${rel}" requires-min-vram-gb || true)"
echo "${value:-0}"
}
variant_engine_available() {
local variant="$1" engine="${LAUNCH_VARIANT_ENGINE[$variant]:-vllm}"
[[ -z "$ENGINE" || "$ENGINE" == "$engine" ]] || return 1
model_has_engine "$MODEL_NAME" "$engine"
}
variant_survives_filter() {
local variant="$1"
[[ "${LAUNCH_VARIANT_MODEL[$variant]:-}" == "$MODEL_NAME" ]] || return 1
variant_engine_available "$variant" || return 1
local min_gpu min_vram engine
min_gpu="$(variant_min_gpu_count "$variant")"
min_vram="$(variant_min_vram_gb "$variant")"
engine="${LAUNCH_VARIANT_ENGINE[$variant]:-vllm}"
[[ "$engine" == "llamacpp" && "${#CARD_INDICES[@]}" -ne 1 ]] && return 1
(( min_gpu <= ${#CARD_INDICES[@]} )) || return 1
(( min_vram == 0 || min_vram <= MIN_VRAM_GB )) || return 1
return 0
}
suggest_default_variant() {
local cards="${#CARD_INDICES[@]}"
if [[ "$MODEL_NAME" == "qwen3.6-27b" ]]; then
if [[ "$ENGINE" == "llamacpp" ]] || { ! model_has_engine "$MODEL_NAME" "vllm" && model_has_engine "$MODEL_NAME" "llamacpp"; }; then
echo "llamacpp/default"
elif (( cards >= 4 )); then
echo "vllm/dual4"
elif (( cards >= 2 )); then
echo "vllm/dual"
else
echo "vllm/long-text"
fi
else
if (( cards >= 2 )); then echo "vllm/gemma-mtp"; else echo "vllm/gemma-mtp-tp1"; fi
fi
}
no_fit_guidance() {
echo "[launch] Selected GPU budget: ${MIN_VRAM_GB} GB minimum per card." >&2
echo "" >&2
echo "No shipped model variant fits this GPU selection:" >&2
echo " Qwen 3.6 27B (INT4): needs >=20 GB" >&2
echo " Gemma 4 31B (INT4): needs >=32 GB single-card or 2x24 GB" >&2
echo "" >&2
echo "Your options:" >&2
echo " 1) Combine with another GPU." >&2
echo " 2) Try llama.cpp with a smaller GGUF. See docs/SINGLE_CARD.md#sub-16gb" >&2
echo " 3) Re-run with --gpus to pick a different card set." >&2
exit 2
}
gemma_single_24gb_guidance() {
echo "[launch] Selected: Gemma 4 31B on GPU ${SELECTED_GPU_CSV} (${MIN_VRAM_GB} GB)." >&2
echo "" >&2
echo "Gemma 4 31B does not fit on a single ${MIN_VRAM_GB} GB card today." >&2
echo "Reason: vLLM's Gemma 4 single-card path needs >=32 GB; use TP=2 on 2x24 GB." >&2
echo "" >&2
echo "Your options:" >&2
echo " 1) Re-run with --gpus 0,1 for TP=2 if you have two 24 GB cards." >&2
echo " 2) Use Qwen 3.6 27B for single-card: bash scripts/launch.sh --model qwen3.6-27b" >&2
exit 2
}
json_field() {
local field="$1"
python3 -c 'import json,sys; data=json.load(sys.stdin); print(data.get(sys.argv[1], ""))' "$field"
}
kv_projection() {
local variant="$1"
[[ "$SKIP_PROJECTION" -eq 0 ]] || return 0
local mapping="${LAUNCH_VARIANT_KVCALC[$variant]:-}"
if [[ -z "$mapping" || "$mapping" == "SKIP" ]]; then
echo "[launch] KV projection only available for vLLM variants today." >&2
return 0
fi
local kv_model="${mapping%%:*}" kv_compose="${mapping#*:}" kv_json status
if kv_json="$("${ROOT_DIR}/tools/kv-calc.py" --model "$kv_model" --compose "$kv_compose" --vram "$MIN_VRAM_GB" --tp "$TP_VALUE" --json 2>&1)"; then
status=0
else
status=$?
fi
if [[ "$kv_json" != \{* ]]; then
echo "[launch] WARN: kv-calc failed for ${variant}: ${kv_json}" >&2
return 0
fi
local verdict weights kv_pool activation overhead drafter total budget pct
verdict="$(json_field verdict <<< "$kv_json")"
weights="$(json_field weights_gb <<< "$kv_json")"
kv_pool="$(json_field kv_pool_actual_gb <<< "$kv_json")"
activation="$(json_field activation_gb <<< "$kv_json")"
overhead="$(json_field cudagraph_overhead_gb <<< "$kv_json")"
drafter="$(json_field drafter_gb <<< "$kv_json")"
total="$(json_field total_gb <<< "$kv_json")"
budget="$(json_field budget_gb <<< "$kv_json")"
pct="$(json_field pct_of_vram <<< "$kv_json")"
echo "" >&2
echo "[launch] Suggested: ${variant} ($(model_label "$MODEL_NAME") on GPU ${SELECTED_GPU_CSV}, TP=${TP_VALUE} PP=${PP_VALUE})" >&2
echo "" >&2
echo "VRAM budget — per card (~${MIN_VRAM_GB} GB):" >&2
printf " Weights/TP=%s: %.2f GB\n" "$TP_VALUE" "$weights" >&2
printf " KV pool: %.2f GB\n" "$kv_pool" >&2
printf " Activations: %.2f GB\n" "$activation" >&2
printf " Cudagraph + NCCL: %.2f GB\n" "$overhead" >&2
if python3 -c 'import sys; sys.exit(0 if float(sys.argv[1]) > 0.01 else 1)' "$drafter"; then
printf " Drafter: %.2f GB\n" "$drafter" >&2
fi
echo " --------------" >&2
printf " Predicted peak: %.2f GB %s (%.0f%% of %.2f GB engine budget)\n" "$total" "$verdict" "$pct" "$budget" >&2
if (( PP_VALUE > 1 )); then
echo " Note: PP is not modelled in projection; real per-card weights should be lower." >&2
fi
python3 -c 'import json,sys; data=json.load(sys.stdin); [print(" Note: " + n) for n in data.get("notes", [])]' <<< "$kv_json" >&2
if [[ "$verdict" == "FAIL" || "$status" -ne 0 ]]; then
echo "" >&2
echo "[launch] Projection says this variant will not fit. Pick another GPU set or model." >&2
exit 2
fi
}
# --- wizard ---
if [[ -z "$VARIANT" ]]; then
echo "" >&2
echo "club-3090 launcher — pick model, GPU set, and serving variant." >&2
echo "(Use --variant <name> next time to skip the wizard.)" >&2
choose_model
choose_gpus
pick_parallelism
if [[ "$MODEL_NAME" == "gemma-4-31b" && "${#CARD_INDICES[@]}" -eq 1 && "$MIN_VRAM_GB" -lt 32 ]]; then
gemma_single_24gb_guidance
fi
CANDIDATE_VARIANTS=()
for candidate in "${LAUNCH_VARIANT_ORDER[@]}"; do
if variant_survives_filter "$candidate"; then
CANDIDATE_VARIANTS+=("$candidate")
fi
done
[[ "${#CANDIDATE_VARIANTS[@]}" -gt 0 ]] || no_fit_guidance
VARIANT="$(suggest_default_variant)"
if [[ " ${CANDIDATE_VARIANTS[*]} " != *" ${VARIANT} "* ]]; then
VARIANT="${CANDIDATE_VARIANTS[0]}"
fi
echo "[launch] model: $(model_label "$MODEL_NAME")" >&2
if (( HET_VRAM_MIXED == 1 && TP_VALUE > 1 )); then
echo "[launch] Note: heterogeneous TP is bottlenecked by the smallest selected card (${MIN_VRAM_GB} GB)." >&2
fi
if (( PP_VALUE > 1 )) && [[ "$VARIANT" == vllm/* ]]; then
echo "[launch] WARN: PP + vLLM drafter/spec-decode paths are experimental on this stack." >&2
fi
kv_projection "$VARIANT"
other_variants=()
for candidate in "${CANDIDATE_VARIANTS[@]}"; do
[[ "$candidate" == "$VARIANT" ]] && continue
other_variants+=("$candidate")
done
if [[ "${#other_variants[@]}" -gt 0 ]]; then
echo "[launch] Other variants that fit this selection: ${other_variants[*]}" >&2
fi
else
if [[ -n "$GPU_ARG" ]]; then
choose_gpus
fi
if [[ -n "$TP_OVERRIDE" || -n "$PP_OVERRIDE" ]]; then
if [[ "${#CARD_INDICES[@]}" -eq 0 ]]; then
TP_VALUE="${TP_OVERRIDE:-1}"
PP_VALUE="${PP_OVERRIDE:-1}"
else
MODEL_NAME="${LAUNCH_VARIANT_MODEL[$VARIANT]:-qwen3.6-27b}"
summarize_selected_vram >/dev/null
pick_parallelism
fi
fi
fi
# --- launch + verify ---
echo ""
echo "[launch] selected variant: ${VARIANT}"
echo ""
if [[ -n "$SELECTED_GPU_CSV" ]]; then
export CUDA_VISIBLE_DEVICES="$SELECTED_GPU_CSV"
export NVIDIA_VISIBLE_DEVICES="$SELECTED_GPU_CSV"
fi
if [[ -n "$TP_VALUE" ]]; then
export TP="$TP_VALUE"
fi
if [[ -n "$PP_VALUE" ]]; then
export PP="$PP_VALUE"
fi
"$SWITCH" "$VARIANT"
# Resolve the actual endpoint port + container name the same way switch.sh
@@ -213,6 +837,8 @@ declare -A LAUNCH_DEFAULT_PORT=(
[vllm/tools-text]=8020
[vllm/minimal]=8020
[vllm/dual]=8010
[vllm/dual4]=8015
[vllm/dual4-dflash]=8016
[vllm/dual-turbo]=8011
[vllm/dual-dflash]=8012
[vllm/dual-dflash-noviz]=8013
@@ -222,6 +848,7 @@ declare -A LAUNCH_DEFAULT_PORT=(
[vllm/dual-nvlink-dflash-noviz]=8019
[vllm/gemma-mtp]=8030
[vllm/gemma-mtp-tp1]=8031
[vllm/gemma-dflash]=8032
[llamacpp/default]=8020
[llamacpp/concurrent]=8020
)
@@ -234,6 +861,8 @@ declare -A LAUNCH_DEFAULT_CONTAINER=(
[vllm/tools-text]=vllm-qwen36-27b
[vllm/minimal]=vllm-qwen36-27b-minimal
[vllm/dual]=vllm-qwen36-27b-dual
[vllm/dual4]=vllm-qwen36-27b-dual4
[vllm/dual4-dflash]=vllm-qwen36-27b-dual4-dflash
[vllm/dual-turbo]=vllm-qwen36-27b-dual-turbo
[vllm/dual-dflash]=vllm-qwen36-27b-dual-dflash
[vllm/dual-dflash-noviz]=vllm-qwen36-27b-dual-dflash-noviz
@@ -243,6 +872,7 @@ declare -A LAUNCH_DEFAULT_CONTAINER=(
[vllm/dual-nvlink-dflash-noviz]=vllm-qwen36-27b-dual-nvlink-dflash-noviz
[vllm/gemma-mtp]=vllm-gemma-4-31b-mtp
[vllm/gemma-mtp-tp1]=vllm-gemma-4-31b-mtp-tp1
[vllm/gemma-dflash]=vllm-gemma-4-31b-dflash
[llamacpp/default]=llama-cpp-qwen36-27b
[llamacpp/concurrent]=llama-cpp-qwen36-27b-concurrent
)
+291
View File
@@ -62,3 +62,294 @@ compose_meta_get() {
return 1
}
compose_hw_sm_to_int() {
local sm="$1"
sm="${sm%%+}"
sm="${sm//sm_/}"
sm="${sm//SM_/}"
sm="${sm// /}"
[[ -z "$sm" ]] && { echo 0; return; }
local major minor
if [[ "$sm" == *.* ]]; then
major="${sm%%.*}"
minor="${sm#*.}"
else
major="$sm"
minor="0"
fi
major="${major//[^0-9]/}"
minor="${minor//[^0-9]/}"
[[ -z "$major" ]] && major=0
[[ -z "$minor" ]] && minor=0
if [[ "${#minor}" -eq 1 ]]; then
minor=$(( minor * 10 ))
else
minor="${minor:0:2}"
[[ -z "$minor" ]] && minor=0
fi
echo $(( major * 100 + minor ))
}
compose_hw_vram_gb() {
local mib="$1"
echo $(( (mib + 1023) / 1024 ))
}
compose_hw_detect_gpus() {
if [[ "${_COMPOSE_HW_GPU_CACHE_SET:-0}" == "1" ]]; then
[[ -n "${_COMPOSE_HW_GPU_CACHE:-}" ]] || return 1
printf '%s\n' "${_COMPOSE_HW_GPU_CACHE}"
return 0
fi
if [[ -n "${CLUB3090_FAKE_GPUS:-}" ]]; then
local fake parsed_fake="" f_idx f_name f_mem_mib f_sm
IFS=',' read -ra _compose_fake_gpus <<< "${CLUB3090_FAKE_GPUS}"
for fake in "${_compose_fake_gpus[@]}"; do
IFS=':' read -r f_idx f_name f_mem_mib f_sm <<< "$fake"
f_idx="$(_compose_meta_trim "${f_idx:-}")"
f_name="$(_compose_meta_trim "${f_name:-}")"
f_name="${f_name//_/ }"
f_mem_mib="$(_compose_meta_trim "${f_mem_mib:-}")"
f_sm="$(_compose_meta_trim "${f_sm:-}")"
[[ -z "$f_idx" || -z "$f_mem_mib" ]] && continue
parsed_fake+="${f_idx}"$'\t'"${f_name}"$'\t'"${f_mem_mib}"$'\t'"${f_sm}"$'\n'
done
parsed_fake="${parsed_fake%$'\n'}"
_COMPOSE_HW_GPU_CACHE_SET=1
_COMPOSE_HW_GPU_CACHE="$parsed_fake"
[[ -n "$parsed_fake" ]] || return 1
printf '%s\n' "$parsed_fake"
return 0
fi
command -v nvidia-smi >/dev/null 2>&1 || return 1
local query idx name mem_mib sm rest
query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null)" || return 1
[[ -n "$query" ]] || return 1
local parsed=""
while IFS=',' read -r idx name mem_mib sm rest; do
idx="$(_compose_meta_trim "$idx")"
name="$(_compose_meta_trim "$name")"
mem_mib="$(_compose_meta_trim "$mem_mib")"
sm="$(_compose_meta_trim "$sm")"
[[ -z "$idx" || -z "$mem_mib" ]] && continue
parsed+="${idx}"$'\t'"${name}"$'\t'"${mem_mib}"$'\t'"${sm}"$'\n'
done <<< "$query"
parsed="${parsed%$'\n'}"
_COMPOSE_HW_GPU_CACHE_SET=1
_COMPOSE_HW_GPU_CACHE="$parsed"
[[ -n "$parsed" ]] || return 1
printf '%s\n' "$parsed"
}
compose_hw_in_use_gpus() {
# Returns GPU indices with non-trivial active compute work. Best-effort:
# primary path maps compute-app UUIDs back to GPU indices; memory.used is
# the fallback for drivers that do not expose compute app UUIDs.
if [[ -n "${CLUB3090_FAKE_BUSY_GPUS:-}" ]]; then
printf '%s\n' "${CLUB3090_FAKE_BUSY_GPUS//,/$'\n'}" | sed '/^$/d'
return 0
fi
if [[ -n "${CLUB3090_FAKE_GPUS:-}" ]]; then
return 0
fi
command -v nvidia-smi >/dev/null 2>&1 || return 0
local uuid_query apps line uuid idx
uuid_query="$(nvidia-smi --query-gpu=index,uuid --format=csv,noheader,nounits 2>/dev/null || true)"
apps="$(nvidia-smi --query-compute-apps=gpu_uuid,pid --format=csv,noheader,nounits 2>/dev/null || true)"
if [[ -n "$uuid_query" && -n "$apps" ]]; then
while IFS=',' read -r uuid _pid; do
uuid="$(_compose_meta_trim "$uuid")"
[[ -z "$uuid" ]] && continue
while IFS=',' read -r idx line; do
idx="$(_compose_meta_trim "$idx")"
line="$(_compose_meta_trim "$line")"
if [[ "$line" == "$uuid" ]]; then
printf '%s\n' "$idx"
fi
done <<< "$uuid_query"
done <<< "$apps" | sort -u
return 0
fi
local mem_used_lines used
mem_used_lines="$(nvidia-smi --query-gpu=index,memory.used --format=csv,noheader,nounits 2>/dev/null || true)"
while IFS=',' read -r idx used; do
idx="$(_compose_meta_trim "$idx")"
used="$(_compose_meta_trim "$used")"
[[ -z "$idx" || -z "$used" ]] && continue
if [[ "$used" =~ ^[0-9]+$ ]] && (( used > 1024 )); then
printf '%s\n' "$idx"
fi
done <<< "$mem_used_lines"
}
compose_hw_summary() {
local gpu_lines
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
if [[ -z "$gpu_lines" ]]; then
printf 'no NVIDIA GPUs detected'
return 0
fi
local count=0 first_name="" first_gb="" mixed=0 idx name mem_mib sm
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
local gb
gb="$(compose_hw_vram_gb "$mem_mib")"
name="${name#NVIDIA }"
name="${name#GeForce }"
count=$((count + 1))
if [[ -z "$first_name" ]]; then
first_name="$name"
first_gb="$gb"
elif [[ "$name" != "$first_name" || "$gb" != "$first_gb" ]]; then
mixed=1
fi
done <<< "$gpu_lines"
if (( count == 0 )); then
printf 'no NVIDIA GPUs detected'
elif (( mixed == 0 )); then
if (( count == 1 )); then
printf '1× %s, %s GB' "$first_name" "$first_gb"
else
printf '%d× %s, %s GB each' "$count" "$first_name" "$first_gb"
fi
else
local parts=()
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
name="${name#NVIDIA }"
name="${name#GeForce }"
parts+=("${name}, $(compose_hw_vram_gb "$mem_mib") GB")
done <<< "$gpu_lines"
local joined=""
for part in "${parts[@]}"; do
if [[ -z "$joined" ]]; then
joined="$part"
else
joined="${joined} + ${part}"
fi
done
printf '%s' "$joined"
fi
}
compose_hw_requirement_text() {
local min_vram_gb="$1"
local min_gpu_count="$2"
local requires_sm="${3:-}"
local req
if [[ "$min_gpu_count" == "1" ]]; then
req="${min_vram_gb} GB+"
else
req="${min_gpu_count}× ${min_vram_gb} GB"
fi
if [[ -n "$requires_sm" && "$requires_sm" != "0.0" ]]; then
req="${req}, sm_${requires_sm%%+}+"
fi
printf '%s' "$req"
}
compose_hw_compose_status() {
local compose_file="$1"
local min_vram_gb min_gpu_count requires_sm
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" ]]; then
printf 'unknown|metadata unavailable'
return 2
fi
requires_sm="${requires_sm:-0.0}"
local required_sm_int
required_sm_int="$(compose_hw_sm_to_int "$requires_sm")"
local gpu_lines
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
if [[ -z "$gpu_lines" ]]; then
printf 'no|no NVIDIA GPUs detected'
return 1
fi
local eligible_count=0 idx name mem_mib sm gb sm_int
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
gb="$(compose_hw_vram_gb "$mem_mib")"
sm_int="$(compose_hw_sm_to_int "$sm")"
if (( gb >= min_vram_gb && sm_int >= required_sm_int )); then
eligible_count=$((eligible_count + 1))
fi
done <<< "$gpu_lines"
if (( eligible_count >= min_gpu_count )); then
printf 'ok|fits your rig'
return 0
fi
printf 'no|needs %s (your rig: %s)' \
"$(compose_hw_requirement_text "$min_vram_gb" "$min_gpu_count" "$requires_sm")" \
"$(compose_hw_summary)"
return 1
}
compose_hw_compose_eligible() {
local status
status="$(compose_hw_compose_status "$1" 2>/dev/null || true)"
[[ "$status" == ok\|* ]]
}
compose_hw_model_status() {
local repo_root="$1"
local model="$2"
local candidates=()
local friendly_need=""
case "$model" in
qwen3.6-27b)
candidates=(
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/long-text.yml"
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
)
friendly_need="needs 20 GB+ VRAM (24 GB recommended)"
;;
gemma-4-31b)
candidates=(
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/int8.yml"
"${repo_root}/models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
)
friendly_need="needs 32 GB+ on single card OR 2× 24 GB"
;;
*)
printf 'no|unknown model: %s' "$model"
return 1
;;
esac
local file status
for file in "${candidates[@]}"; do
[[ -f "$file" ]] || continue
status="$(compose_hw_compose_status "$file" 2>/dev/null || true)"
if [[ "$status" == ok\|* ]]; then
printf 'ok|fits your rig'
return 0
fi
done
printf 'no|%s (your rig: %s)' "$friendly_need" "$(compose_hw_summary)"
return 1
}
+96 -7
View File
@@ -2,7 +2,8 @@
#
# Model-aware one-shot setup for club-3090.
#
# bash scripts/setup.sh <model-name>
# bash scripts/setup.sh # interactive model picker in a TTY
# bash scripts/setup.sh <model-name> # scripted/CI positional form
#
# Currently supported:
# qwen3.6-27b → Lorbus/Qwen3.6-27B-int4-AutoRound + Genesis patches
@@ -37,15 +38,89 @@
set -euo pipefail
# ---------- Model dispatch ----------
MODEL_NAME="${1:-}"
if [[ -z "${MODEL_NAME}" ]]; then
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
usage() {
echo "Usage: $0 <model-name>"
echo " $0 # interactive model picker in a TTY"
echo ""
echo "Run with no model name in a normal terminal to open the hardware-aware"
echo "model picker. Use the positional form in scripts/CI to skip prompts."
echo ""
echo "Supported model names:"
echo " qwen3.6-27b"
echo " gemma-4-31b"
exit 1
}
model_label() {
case "$1" in
qwen3.6-27b) echo "Qwen 3.6 27B" ;;
gemma-4-31b) echo "Gemma 4 31B" ;;
*) echo "$1" ;;
esac
}
model_picker_line() {
local idx="$1" model="$2" size="$3" status mark reason
status="$(compose_hw_model_status "$ROOT_DIR" "$model" 2>/dev/null || true)"
reason="${status#*|}"
if [[ "$status" == ok\|* ]]; then
mark="✓"
else
mark="✗"
fi
printf " %s. %-14s (%s) %s %s\n" "$idx" "$(model_label "$model")" "$size" "$mark" "$reason"
}
pick_model_interactive() {
# shellcheck source=lib/compose-meta.sh
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
echo "[setup] Which model to download?" >&2
echo "" >&2
model_picker_line "1" "qwen3.6-27b" "~14 GB AutoRound INT4" >&2
model_picker_line "2" "gemma-4-31b" "~21 GB AutoRound INT4 + drafter" >&2
echo " 3. Both (~30 GB total) downloads both model families" >&2
echo "" >&2
while true; do
local pick
read -rp "Choice [1-3]: " pick
case "$pick" in
1) echo "qwen3.6-27b"; return ;;
2) echo "gemma-4-31b"; return ;;
3) echo "both"; return ;;
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
esac
done
}
# ---------- Model dispatch ----------
case "${1:-}" in
-h|--help)
usage
exit 0
;;
esac
MODEL_NAME="${1:-}"
if [[ -z "${MODEL_NAME}" ]]; then
if [[ -t 0 && -t 1 ]]; then
MODEL_NAME="$(pick_model_interactive)"
else
usage
echo ""
echo "(Interactive picker available in a TTY shell. Use the positional form in scripts/CI.)"
exit 1
fi
fi
if [[ "${MODEL_NAME}" == "both" ]]; then
# Resolve MODEL_DIR once in the parent by reusing the normal prompt below,
# then recurse through the positional form for each model.
SETUP_BOTH_MODE=1
MODEL_NAME="qwen3.6-27b"
else
SETUP_BOTH_MODE=0
fi
# ALWAYS_DRAFT_REPO + ALWAYS_DRAFT_SUBDIR: a drafter that this model REQUIRES
@@ -89,8 +164,6 @@ case "${MODEL_NAME}" in
;;
esac
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
# ---------- MODEL_DIR resolution ----------
# Order of precedence:
# 1. MODEL_DIR already exported in the calling shell → use as-is
@@ -158,6 +231,18 @@ fi
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
if [[ "${SETUP_BOTH_MODE:-0}" == "1" ]]; then
export MODEL_DIR
echo "[setup] downloading both supported models into ${MODEL_DIR}"
echo ""
bash "$0" qwen3.6-27b
echo ""
bash "$0" gemma-4-31b
echo ""
echo "[setup] ✓ Both models downloaded."
echo "[setup] Next: bash scripts/launch.sh"
exit 0
fi
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
cd "${ROOT_DIR}"
@@ -445,6 +530,7 @@ echo ""
# refactored 2026-05-03 to vendor the two files in-repo, fixing #37.)
# Per-model "next steps" — different composes / served-model-name / port between models.
SETUP_MODEL_DISPLAY="$(model_label "${MODEL_NAME}")"
case "${MODEL_NAME}" in
qwen3.6-27b)
SAMPLE_CONTAINER="vllm-qwen36-27b"
@@ -468,6 +554,9 @@ case "${MODEL_NAME}" in
;;
esac
echo "[setup] ✓ ${SETUP_MODEL_DISPLAY} downloaded."
echo "[setup] Next: bash scripts/launch.sh"
echo ""
echo "Next — single-card vLLM (default):"
if [[ "${MODEL_NAME}" == "gemma-4-31b" ]]; then
echo " bash scripts/switch.sh vllm/gemma-mtp"
+87 -16
View File
@@ -5,6 +5,7 @@
# Usage:
# bash scripts/submit-bench.sh --tag <tag>
# bash scripts/submit-bench.sh --tag <tag> --auto-submit
# bash scripts/submit-bench.sh --tag <tag> --auto-submit --as-pr
set -euo pipefail
@@ -13,6 +14,7 @@ cd "$ROOT_DIR"
TAG=""
AUTO_SUBMIT=0
AS_PR=0
SECTION_OVERRIDE=""
usage() {
@@ -40,6 +42,10 @@ while [[ $# -gt 0 ]]; do
AUTO_SUBMIT=1
shift
;;
--as-pr)
AS_PR=1
shift
;;
--section)
[[ $# -ge 2 ]] || die "--section requires a value"
SECTION_OVERRIDE="$2"
@@ -135,6 +141,35 @@ PY
fi
}
write_issue_body() {
local body_file="$1"
local row="$2"
local tag="$3"
local section="$4"
# The repo's numbers-from-your-rig issue template is a structured YAML form
# with required textarea/dropdown fields. `gh issue create --template` opens
# that interactive form shape, which is not useful once submit-bench has
# already generated the structured report. Use a direct markdown body instead.
{
echo "**Compose / section**: \`${section}\`"
echo
echo "**Rig**:"
echo
echo '```text'
cat "${TAG_DIR}/rig.txt"
echo '```'
echo
echo "**Proposed BENCHMARKS.md row**:"
echo
echo "$row"
echo
echo "**Full report**: \`results/rebench/${tag}/REPORT.md\`"
echo
echo "**Generated row file**: \`results/rebench/${tag}/BENCHMARKS-row.md\`"
} > "$body_file"
}
insert_row() {
local section="$1"
local row="$2"
@@ -186,8 +221,24 @@ PY
}
if [[ "$AUTO_SUBMIT" -ne 1 ]]; then
echo "Inspect at ${OUTPUT}. To submit:"
echo " bash scripts/submit-bench.sh --tag ${TAG} --auto-submit"
cat <<EOF
Inspect at ${OUTPUT}. Three ways to land it (recommended order):
1. Issue + maintainer integrates (preferred — vetting before merge):
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit
(opens an issue via \`gh issue create\`)
Or, no-gh-needed:
https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml
— paste the contents of ${OUTPUT} + ${TAG_DIR}/rig.txt into the body
2. Direct PR (advanced — for contributors who know the matrix structure):
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit --as-pr
Note: matrix is hand-curated; direct PRs may get redirected to an
issue thread for context-gathering before merge.
3. Manual edit (zero tools):
Paste the row from ${OUTPUT} into BENCHMARKS.md via the GitHub web editor.
EOF
exit 0
fi
@@ -199,30 +250,50 @@ if [[ -n "$SECTION_OVERRIDE" ]]; then
fi
fi
TITLE="bench(matrix): @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
PR_TITLE="bench(matrix): @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
ISSUE_TITLE="[bench] @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
BRANCH_USER="$(printf '%s' "${BENCH_ROW_GITHUB_USER}" | tr -cd '[:alnum:]_.-')"
BRANCH_TAG="$(printf '%s' "${TAG}" | tr -cd '[:alnum:]_.-')"
BRANCH="bench/${BRANCH_USER}-${BRANCH_TAG}"
BODY_FILE="$TAG_DIR/PR-body.md"
write_pr_body "$BODY_FILE" "$ROW" "$TAG"
if [[ "$AS_PR" -eq 1 ]]; then
BODY_FILE="$TAG_DIR/PR-body.md"
write_pr_body "$BODY_FILE" "$ROW" "$TAG"
else
BODY_FILE="$TAG_DIR/ISSUE-body.md"
write_issue_body "$BODY_FILE" "$ROW" "$TAG" "$SECTION"
fi
if [[ "${GH_MOCK:-0}" == "1" ]]; then
MOCK_LOG="$TAG_DIR/auto-submit-mock.log"
{
echo "git switch -c ${BRANCH}"
echo "insert BENCHMARKS.md row under: ${SECTION}"
echo "git commit -m ${TITLE}"
echo "git push -u origin ${BRANCH}"
echo "gh pr create --title ${TITLE} --body-file ${BODY_FILE}"
} > "$MOCK_LOG"
log "GH_MOCK=1 — wrote mocked auto-submit commands: ${MOCK_LOG}"
log "PR title: ${TITLE}"
if [[ "$AS_PR" -eq 1 ]]; then
{
echo "git switch -c ${BRANCH}"
echo "insert BENCHMARKS.md row under: ${SECTION}"
echo "git commit -m ${PR_TITLE}"
echo "git push -u origin ${BRANCH}"
echo "gh pr create --title ${PR_TITLE} --body-file ${BODY_FILE}"
} > "$MOCK_LOG"
log "GH_MOCK=1 — wrote mocked PR auto-submit commands: ${MOCK_LOG}"
log "PR title: ${PR_TITLE}"
else
{
echo "gh issue create --title ${ISSUE_TITLE} --body-file ${BODY_FILE} --label bench-contribution"
} > "$MOCK_LOG"
log "GH_MOCK=1 — wrote mocked issue auto-submit command: ${MOCK_LOG}"
log "Issue title: ${ISSUE_TITLE}"
fi
exit 0
fi
command -v gh >/dev/null 2>&1 || die "'gh' not found. Install GitHub CLI or submit manually."
gh auth status >/dev/null 2>&1 || die "not authed with gh. Run: gh auth login"
if [[ "$AS_PR" -ne 1 ]]; then
ISSUE_URL="$(gh issue create --title "$ISSUE_TITLE" --body-file "$BODY_FILE" --label bench-contribution)"
log "Opened issue: ${ISSUE_URL}"
exit 0
fi
if ! git diff --quiet -- BENCHMARKS.md; then
die "BENCHMARKS.md already has local edits; commit/stash them before --auto-submit"
fi
@@ -235,7 +306,7 @@ else
fi
insert_row "$SECTION" "$ROW"
git add BENCHMARKS.md
git commit -m "$TITLE"
git commit -m "$PR_TITLE"
git push -u origin "$BRANCH"
PR_URL="$(gh pr create --title "$TITLE" --body-file "$BODY_FILE")"
PR_URL="$(gh pr create --title "$PR_TITLE" --body-file "$BODY_FILE")"
log "Opened PR: ${PR_URL}"
+215
View File
@@ -0,0 +1,215 @@
#!/usr/bin/env bash
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
TMP_DIR="$(mktemp -d)"
ORIG_PATH="$PATH"
trap 'rm -rf "$TMP_DIR"' EXIT
assert_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" != *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
assert_not_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" == *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output not to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
make_mock_tools() {
mkdir -p "${TMP_DIR}/bin"
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK_NVIDIA_SMI'
#!/usr/bin/env bash
case "$*" in
*"--query-gpu=index,name,memory.total,compute_cap"*)
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
;;
"-L")
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}" \
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
;;
*)
echo "unexpected nvidia-smi invocation: $*" >&2
exit 2
;;
esac
MOCK_NVIDIA_SMI
chmod +x "${TMP_DIR}/bin/nvidia-smi"
cat > "${TMP_DIR}/switch-mock" <<'MOCK_SWITCH'
#!/usr/bin/env bash
echo "SWITCHED $* CUDA=${CUDA_VISIBLE_DEVICES:-} NVD=${NVIDIA_VISIBLE_DEVICES:-} TP=${TP:-} PP=${PP:-}"
MOCK_SWITCH
chmod +x "${TMP_DIR}/switch-mock"
export PATH="${TMP_DIR}/bin:${ORIG_PATH}"
}
set_rig() {
export MOCK_GPU_QUERY="$1"
}
model_status() {
local model="$1"
(
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
compose_hw_model_status "$ROOT_DIR" "$model"
)
}
assert_model_status() {
local model="$1"
local expected_prefix="$2"
local expected_text="${3:-}"
local status
status="$(model_status "$model" || true)"
if [[ "$status" != "${expected_prefix}"* ]]; then
echo "ASSERTION FAILED: ${model} status expected prefix '${expected_prefix}', got '${status}'" >&2
exit 1
fi
if [[ -n "$expected_text" ]]; then
assert_contains "$status" "$expected_text"
fi
}
make_mock_tools
# Matched 2x3090: Qwen and Gemma both have a viable compose.
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "ok|fits your rig"
# Single 24 GB Ampere: Qwen fits; Gemma needs either 32 GB+ single-card or 2x24 GB.
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
assert_contains "$(model_status "gemma-4-31b")" "1× RTX 3090, 24 GB"
# Single 16 GB: neither shipped model has a viable compose.
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9'
assert_model_status "qwen3.6-27b" "no|" "needs 20 GB+ VRAM"
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
# Heterogeneous 16 + 24 GB: Qwen can run on the 24 GB card; Gemma dual cannot.
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "no|" "RTX 4060 Ti, 16 GB + RTX 3090, 24 GB"
# 32 GB+ modern card: Gemma's single-card compose is eligible.
set_rig $'0, NVIDIA GeForce RTX 5090, 32768, 12.0'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "ok|fits your rig"
# Non-TTY no-arg setup fails fast with usage rather than hanging.
if out="$(echo | bash "${ROOT_DIR}/scripts/setup.sh" 2>&1)"; then
echo "ASSERTION FAILED: non-TTY no-arg setup unexpectedly succeeded" >&2
echo "$out" >&2
exit 1
fi
assert_contains "$out" "Usage:"
assert_contains "$out" "Interactive picker available in a TTY shell"
# Positional setup path remains non-interactive and reaches the existing flow.
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
out="$(MODEL_DIR="${TMP_DIR}/models" PREFLIGHT_DISK_GB=0 SKIP_GENESIS=1 SKIP_MODEL=1 bash "${ROOT_DIR}/scripts/setup.sh" qwen3.6-27b 2>&1)"
assert_not_contains "$out" "Which model to download?"
assert_contains "$out" "[model] SKIP_MODEL=1"
# The launch wizard now picks model -> GPU set -> parallelism. Scripted flags
# skip prompts, select the expected variant, and export GPU / TP / PP envs.
mkdir -p "${TMP_DIR}/models/qwen3.6-27b-autoround-int4" \
"${TMP_DIR}/models/gemma-4-31b-autoround-int4"
FAKE_8X3090='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6,4:RTX_3090:24576:8.6,5:RTX_3090:24576:8.6,6:RTX_3090:24576:8.6,7:RTX_3090:24576:8.6'
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0 --no-projection 2>&1)"
assert_contains "$out" "[launch] selected variant: vllm/long-text"
assert_contains "$out" "SWITCHED vllm/long-text CUDA=0 NVD=0 TP=1 PP=1"
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1 --no-projection 2>&1)"
assert_contains "$out" "[launch] Tensor parallel TP=2"
assert_contains "$out" "SWITCHED vllm/dual CUDA=0,1 NVD=0,1 TP=2 PP=1"
selected_count="$(grep -c "\[launch\] selected variant:" <<< "$out" || true)"
if [[ "$selected_count" != "1" ]]; then
echo "ASSERTION FAILED: expected one selected-variant line, got ${selected_count}" >&2
echo "$out" >&2
exit 1
fi
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model gemma-4-31b --gpus 0 --no-projection 2>&1)"; then
echo "ASSERTION FAILED: Gemma single-24GB launch unexpectedly succeeded" >&2
echo "$out" >&2
exit 1
fi
assert_contains "$out" "Gemma 4 31B does not fit on a single 24 GB card today"
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6,2:RTX_3090:24576:8.6,3:RTX_3090:24576:8.6,4:RTX_3090:24576:8.6,5:RTX_3090:24576:8.6' \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1,2,3,4,5 --tp 6 --no-projection 2>&1)"; then
echo "ASSERTION FAILED: invalid Qwen TP=6 unexpectedly succeeded" >&2
echo "$out" >&2
exit 1
fi
assert_contains "$out" "Valid TP values: 1 2 4"
out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS="${FAKE_8X3090}" \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model gemma-4-31b --gpus 0,1,2,3,4,5,6,7 --tp 8 2>&1)"
assert_contains "$out" "[launch] Tensor parallel TP=8"
assert_contains "$out" "[launch] Suggested: vllm/gemma-mtp"
assert_contains "$out" "VRAM budget — per card"
assert_contains "$out" "Note: TP > 4 predictions are extrapolated"
assert_not_contains "$out" "KV projection skipped"
assert_contains "$out" "SWITCHED vllm/gemma-mtp CUDA=0,1,2,3,4,5,6,7 NVD=0,1,2,3,4,5,6,7 TP=8 PP=1"
if out="$(MODEL_DIR="${TMP_DIR}/models" CLUB3090_FAKE_GPUS="${FAKE_8X3090}" \
SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" \
--no-preflight --no-verify --model qwen3.6-27b --gpus 0,1,2,3,4,5,6,7 --tp 8 --no-projection 2>&1)"; then
echo "ASSERTION FAILED: invalid Qwen TP=8 unexpectedly succeeded" >&2
echo "$out" >&2
exit 1
fi
assert_contains "$out" "num_kv_heads does not divide TP=8"
assert_contains "$out" "Valid TP values: 1 2 4"
# TTY-backed no-arg setup supports the cosmetic but real "Both" choice by
# dispatching through the positional path for both model families.
if ! command -v script >/dev/null 2>&1; then
echo "ASSERTION FAILED: util-linux 'script' is required for TTY picker coverage" >&2
exit 1
fi
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
export MODEL_DIR="${TMP_DIR}/models"
export PREFLIGHT_DISK_GB=0
export SKIP_GENESIS=1
export SKIP_MODEL=1
out="$(printf '3\n' | script -qec "bash '${ROOT_DIR}/scripts/setup.sh'" /dev/null 2>&1)"
assert_contains "$out" "[setup] Which model to download?"
assert_contains "$out" "Both"
assert_contains "$out" "[setup] downloading both supported models"
skip_count="$(grep -c "\[model\] SKIP_MODEL=1" <<< "$out" || true)"
if [[ "$skip_count" != "2" ]]; then
echo "ASSERTION FAILED: expected Both choice to dispatch two model setup runs, got ${skip_count}" >&2
echo "--- output ---" >&2
echo "$out" >&2
exit 1
fi
echo "test-setup-picker: ok"
+14 -3
View File
@@ -59,11 +59,15 @@ done
tag="qwen-int8-pth-n4-2026-05-10"
rm -f "results/rebench/${tag}/BENCHMARKS-row.md" \
"results/rebench/${tag}/PR-body.md" \
"results/rebench/${tag}/ISSUE-body.md" \
"results/rebench/${tag}/auto-submit-mock.log"
out="$(bash scripts/submit-bench.sh --tag "$tag")"
assert_contains "$out" "Generated BENCHMARKS row for section: Dual-card (2× RTX 3090, TP=2)"
assert_contains "$out" "Wrote: results/rebench/${tag}/BENCHMARKS-row.md"
assert_contains "$out" "1. Issue + maintainer integrates"
assert_contains "$out" "2. Direct PR"
assert_contains "$out" "3. Manual edit"
test -s "results/rebench/${tag}/BENCHMARKS-row.md"
if out="$(bash scripts/submit-bench.sh --tag does-not-exist 2>&1)"; then
@@ -73,14 +77,21 @@ fi
assert_contains "$out" "tag dir not found: results/rebench/does-not-exist"
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit)"
assert_contains "$out" "PR title: bench(matrix): @octocat ${tag}"
assert_contains "$out" "Issue title: [bench] @octocat ${tag}"
test -s "results/rebench/${tag}/auto-submit-mock.log"
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh pr create --title bench(matrix): @octocat ${tag}"
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh issue create --title [bench] @octocat ${tag}"
test -s "results/rebench/${tag}/ISSUE-body.md"
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "results/rebench/${tag}/REPORT.md"
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "Proposed BENCHMARKS.md row"
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit --as-pr)"
assert_contains "$out" "PR title: bench(matrix): @octocat ${tag}"
test -s "results/rebench/${tag}/PR-body.md"
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh pr create --title bench(matrix): @octocat ${tag}"
assert_contains "$(cat "results/rebench/${tag}/PR-body.md")" "results/rebench/${tag}/REPORT.md"
tmp_bin="$(mktemp -d)"
trap 'rm -rf "$tmp_bin"; rm -f "results/rebench/${tag}/BENCHMARKS-row.md" "results/rebench/${tag}/PR-body.md" "results/rebench/${tag}/auto-submit-mock.log"' EXIT
trap 'rm -rf "$tmp_bin"; rm -f "results/rebench/${tag}/BENCHMARKS-row.md" "results/rebench/${tag}/PR-body.md" "results/rebench/${tag}/ISSUE-body.md" "results/rebench/${tag}/auto-submit-mock.log"' EXIT
cat > "${tmp_bin}/gh" <<'MOCK_GH'
#!/usr/bin/env bash
if [[ "$1" == "auth" && "$2" == "status" ]]; then
+559 -225
View File
@@ -1,53 +1,69 @@
#!/usr/bin/env python3
"""kv-calc.py — predict per-card VRAM budget for a Qwen3.6-27B vLLM compose.
#!/bin/sh
''':'
exec python3 "$0" "$@"
':'''
from __future__ import annotations
__doc__ = """kv-calc.py — predict per-card VRAM budget for vLLM composes.
Predicts (per card, after TP split):
- Model weights
- KV pool (attention layers — 16 full_attention layers with GQA)
- Activation peak (DeltaNet GDN forward — 48 linear_attention layers)
- KV pool (attention layers only — recurrent / SSM states show up in activation)
- Activation peak (model-specific: Qwen GDN forward, Gemma SWA + dense MLP)
- Cudagraph + workspace overhead
- Drafter overhead (MTP / DFlash)
- Total vs available VRAM
- Verdict: PASS / TIGHT / FAIL
Two models modelled:
- Qwen 3.6 27B (DeltaNet hybrid: 16 full_attention + 48 GDN)
- Gemma 4 31B (SWA + dense MLP: 10 full_attention + 50 sliding_attention)
vLLM rate-limits KV pool to fit available budget; this predictor models that
capping behavior. When the requested KV pool exceeds what fits, the verdict
is TIGHT (vLLM will cap pool — effective concurrency reduced) not FAIL.
Anchored to:
- PerfMamba (arxiv 2511.22849) — block-wise state materialization scaling
- PerfMamba (arxiv 2511.22849) — Qwen GDN block-wise state materialization scaling
https://arxiv.org/html/2511.22849
- TurboQuant (arxiv 2504.19874, ICLR 2026) — TQ3 byte savings
https://arxiv.org/abs/2504.19874
- PagedAttention (arxiv 2309.06180) — KV pool layout
https://arxiv.org/abs/2309.06180
Calibrated against measured BENCHMARKS.md rows. Coefficients in
GDN_ACTIVATION_COEF reflect club-3090's empirical findings on top of
PerfMamba's O(γ·D·N·L) scaling — the absolute scaling is well-defined,
but the per-token coefficient depends on fla.ops.chunk implementation
details that the published literature doesn't enumerate. See
docs/KV_MATH.md for the derivation + calibration trace.
Calibrated against measured BENCHMARKS.md rows per model. Coefficients reflect
club-3090's empirical findings. See docs/KV_MATH.md for the derivation +
calibration trace.
Usage:
bash tools/kv-calc.py --compose dual-turbo --vram 24
bash tools/kv-calc.py --compose dual-turbo --kv-format fp8_e5m2 --vram 20
bash tools/kv-calc.py --max-ctx 180000 --kv-format turboquant_3bit_nc --tp 1 --vram 24
bash tools/kv-calc.py --calibration # show calibration vs measured points
bash tools/kv-calc.py --compose dual-turbo --vram 24 # Qwen (default model)
bash tools/kv-calc.py --model gemma-4-31b --compose gemma-dual-int8 --vram 24
bash tools/kv-calc.py --model gemma-4-31b --solve-max-ctx --kv-format int8_per_token_head --tp 2 --vram 24
bash tools/kv-calc.py --calibration # both models, grouped per-model
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass
from typing import Optional
# =============================================================================
# Model specs
# =============================================================================
# ---- Qwen3.6-27B AutoRound INT4 — from config.json text_config ----
QWEN36_27B = {
"model_id": "qwen3.6-27b-autoround",
"model_id": "qwen3.6-27b",
"model_family": "qwen3-next-hybrid",
"hidden_size": 5120,
"num_hidden_layers": 64,
"num_gdn_layers": 48, # linear_attention layers (Gated DeltaNet)
"num_attn_layers": 16, # full_attention layers
"num_attn_heads": 24,
"num_kv_heads": 4, # GQA
"valid_tp": [1, 2, 4],
"head_dim_attn": 256, # attention head dim
"linear_num_v_heads": 48, # GDN value heads
"linear_num_k_heads": 16, # GDN key heads (GQA-style at the GDN level too)
@@ -57,128 +73,283 @@ QWEN36_27B = {
"weights_total_gb": 17.5, # AutoRound INT4 storage on disk
"mamba_state_bytes": 4, # mamba_ssm_dtype=float32
"chunk_size": 256, # fla.ops.chunk default
"max_ctx_supported": 262144,
"attention_k_eq_v": False, # Qwen stores K and V independently
}
# ---- KV format bytes per stored token element ----
# (one element = one head dim of one head; K and V counted separately)
# Source: vLLM/HF docs + TurboQuant paper.
# ---- Gemma 4 31B — from config.json text_config ----
# Layer pattern: [sliding_attention × 5, full_attention × 1] × 10
# = 50 sliding-attention + 10 full-attention.
GEMMA4_31B = {
"model_id": "gemma-4-31b",
"model_family": "gemma4-swa-dense",
"hidden_size": 5376,
"intermediate_size": 21504,
"num_hidden_layers": 60,
"num_full_attn_layers": 10, # full_attention (growing KV, head_dim=512)
"num_sliding_attn_layers": 50, # sliding_attention (fixed window, head_dim=256)
"num_attn_heads": 32,
"num_kv_heads": 16, # GQA 2:1
"valid_tp": [1, 2, 4, 8, 16],
"head_dim_sliding": 256, # sliding_attention head dim
"global_head_dim": 512, # full_attention head dim (asymmetric)
"sliding_window": 1024,
"weights_int4_gb": 18.0, # AutoRound INT4 on disk
"weights_awq_gb": 17.0, # cyankiwi AWQ-4bit (lower on-card due to AWQ packing)
"weights_bf16_gb": 58.0, # unquantized — does not fit on 24 GB
"max_ctx_supported": 262144,
"drafter_mtp_gb": 0.97, # google/gemma-4-31b-it-assistant (FP16)
"drafter_dflash_gb": 2.9, # z-lab/gemma-4-31b-it-dflash
"attention_k_eq_v": True, # vLLM allocator exploits K==V tying (per_token uses ×1 not ×2)
}
MODEL_SPECS = {
"qwen3.6-27b": QWEN36_27B,
"gemma-4-31b": GEMMA4_31B,
}
# =============================================================================
# KV format bytes-per-element
# =============================================================================
# (one element = one head dim of one head; K and V counted separately for
# models where attention_k_eq_v=False. For Gemma the K==V tying halves this
# at the formula level — see kv_pool_per_card_bytes.)
# Source: vLLM/HF docs + TurboQuant paper + PR #40391 (INT8 per-token-head).
KV_FORMAT_BYTES = {
"fp16": 2.0,
"bf16": 2.0,
"fp8_e5m2": 1.0,
"fp8_e4m3": 1.0,
"q4_0": 0.5 + 0.0625, # 4-bit + per-group scale
"k8v4": 0.75, # avg of K=int8 V=int4
"turboquant_3bit_nc": 0.375 + 0.05, # 3 bits + small QJL overhead
"fp16": 2.0,
"bf16": 2.0,
"fp8_e5m2": 1.0,
"fp8_e4m3": 1.0,
"int8_per_token_head": 1.01, # 1.0 int8 + per-token-head fp16 scale (~1% amortized)
"q4_0": 0.5 + 0.0625, # 4-bit + per-group scale
"k8v4": 0.75, # avg of K=int8 V=int4
"turboquant_3bit_nc": 0.375 + 0.05, # 3 bits + small QJL overhead
}
# ---- GDN activation-peak per-layer per-token coefficient (bytes) ----
# Calibrated empirically against measured BENCHMARKS rows. The PerfMamba
# O(γ·D·N·L) scaling sets the *form*; this coefficient captures
# fla.ops.chunk implementation details + per-KV-format dequant overhead.
# TQ3's coefficient is ~25% larger than fp8 (matches efschu's 20 GB
# Cliff 2 finding — TQ3 dequant adds activation pressure).
GDN_ACTIVATION_COEF = {
# =============================================================================
# Per-model activation coefficients
# =============================================================================
# ---- Qwen GDN activation-peak per-layer per-token coefficient (bytes) ----
# Calibrated empirically against measured BENCHMARKS rows. PerfMamba's
# O(γ·D·N·L) scaling sets the form; this captures fla.ops.chunk
# implementation details + KV-format-dependent dequant overhead.
QWEN_GDN_ACTIVATION_COEF = {
"fp16": 135,
"bf16": 135,
"fp8_e5m2": 130,
"fp8_e4m3": 130,
"int8_per_token_head": 130,
"q4_0": 155,
"k8v4": 155,
"turboquant_3bit_nc": 165,
}
# ---- Compose presets ----
# ---- Gemma activation peak (mostly constant in ctx) ----
# Unlike Qwen GDN, Gemma's activation peak comes from dense MLP forward +
# SWA windowed-attention prefill, both bounded by chunked-prefill chunk_size.
# Result: roughly CONSTANT in max_ctx. Calibrated as a per-TP base (GB) plus
# a small per-token term to capture residual ctx scaling.
GEMMA_ACTIVATION_CONST_GB = 1.5 # per card at TP=1 — calibrated, ~scales as 1/TP
GEMMA_ACTIVATION_PER_TOKEN_BYTES = 8 # tiny ctx scaling term to keep solver well-behaved
# =============================================================================
# Compose presets (per-model)
# =============================================================================
# Pulled from each compose's CLI args. Update if a compose changes.
# Compose IDs are namespaced: Qwen uses bare names (back-compat); Gemma uses gemma-* prefix.
COMPOSES = {
"minimal": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.90, "mtp": False},
"long-text": {"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.93, "mtp": True},
"long-text-no-mtp":{"max_ctx": 200000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": False},
"long-vision": {"max_ctx": 145000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
"bounded-thinking":{"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
"tools-text": {"max_ctx": 75000, "max_num_seqs": 1, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.97, "mtp": True},
"dual": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 2, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
"dual-turbo": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 2, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
"dual-dflash": {"max_ctx": 185000, "max_num_seqs": 1, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
"dual-dflash-noviz":{"max_ctx": 200000,"max_num_seqs": 2, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
"dual4": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 4, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
"dual4-dflash": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 4, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
"qwen3.6-27b": {
"minimal": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.90, "mtp": False},
"long-text": {"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.93, "mtp": True},
"long-text-no-mtp": {"max_ctx": 200000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": False},
"long-vision": {"max_ctx": 145000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
"bounded-thinking": {"max_ctx": 180000, "max_num_seqs": 1, "tp": 1, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
"tools-text": {"max_ctx": 75000, "max_num_seqs": 1, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.97, "mtp": True},
"dual": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 2, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
"dual-turbo": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 2, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "mtp": True},
"dual-dflash": {"max_ctx": 185000, "max_num_seqs": 1, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
"dual-dflash-noviz":{"max_ctx": 200000, "max_num_seqs": 2, "tp": 2, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
"dual4": {"max_ctx": 262144, "max_num_seqs": 4, "tp": 4, "kv_format": "fp8_e5m2", "mem_util": 0.95, "mtp": True},
"dual4-dflash": {"max_ctx": 262144, "max_num_seqs": 2, "tp": 4, "kv_format": "fp16", "mem_util": 0.95, "mtp": False, "dflash_draft_gb": 1.75},
},
"gemma-4-31b": {
# dual/docker-compose.yml — default MTP, 32K BF16 KV, max-num-seqs=4
"gemma-dual": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 2, "kv_format": "bf16", "mem_util": 0.92, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
# dual/int8.yml default — 98K + INT8 PTH KV + 4 seqs
"gemma-dual-int8": {"max_ctx": 98304, "max_num_seqs": 4, "tp": 2, "kv_format": "int8_per_token_head", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
# dual/int8.yml long — 262K + INT8 PTH KV + seqs=1 (model native max)
"gemma-dual-int8-262k": {"max_ctx": 262144, "max_num_seqs": 1, "tp": 2, "kv_format": "int8_per_token_head", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
# dual/bf16.yml — long-ctx BF16 weights + BF16 KV (auto)
"gemma-dual-bf16": {"max_ctx": 200000, "max_num_seqs": 1, "tp": 2, "kv_format": "bf16", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
# dual/int8-tq3.yml — TQ3 KV alt
"gemma-dual-int8-tq3": {"max_ctx": 98304, "max_num_seqs": 4, "tp": 2, "kv_format": "turboquant_3bit_nc", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
# dual/dflash.yml — DFlash drafter + 32K BF16
"gemma-dual-dflash": {"max_ctx": 32768, "max_num_seqs": 4, "tp": 2, "kv_format": "bf16", "mem_util": 0.92, "weights_variant": "int4", "drafter_gb": 2.9, "mtp": False, "dflash_draft_gb": 2.9},
# dual/dflash-int8.yml — DFlash + INT8 PTH KV (PR #42102 unblock)
"gemma-dual-dflash-int8": {"max_ctx": 65536, "max_num_seqs": 2, "tp": 2, "kv_format": "int8_per_token_head", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 2.9, "mtp": False, "dflash_draft_gb": 2.9},
# dual/awq.yml — cyankiwi AWQ weights
"gemma-dual-awq": {"max_ctx": 65536, "max_num_seqs": 4, "tp": 2, "kv_format": "bf16", "mem_util": 0.85, "weights_variant": "awq", "drafter_gb": 0.97, "mtp": True},
# single/docker-compose.yml — 32 GB+ required; Ampere consumer OOMs at boot
"gemma-single": {"max_ctx": 8192, "max_num_seqs": 256, "tp": 1, "kv_format": "fp8_e5m2", "mem_util": 0.95, "weights_variant": "int4", "drafter_gb": 0.97, "mtp": True},
},
}
# ---- Calibration: measured BENCHMARKS rows (peak per-card VRAM during bench) ----
CALIBRATION = [
# (compose, vram_gb, measured_peak_gb, source_url)
("dual", 24, 23.6, "BENCHMARKS.md#qwen36-27b dual.yml @noonghunna 2026-04-29"),
("dual-turbo", 24, 19.8, "BENCHMARKS.md#qwen36-27b dual-turbo.yml @noonghunna 2026-04-29"),
("dual-dflash", 24, 23.6, "BENCHMARKS.md#qwen36-27b dual-dflash.yml @noonghunna 2026-04-29"),
("dual-dflash-noviz",24, 23.8, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @noonghunna 2026-04-29"),
("dual4", 24, 23.5, "BENCHMARKS.md#qwen36-27b dual4.yml @whamp 2026-05-03"),
("dual4-dflash", 24, 22.0, "BENCHMARKS.md#qwen36-27b dual4-dflash.yml @whamp 2026-05-03"),
("dual-dflash-noviz",24, 21.8, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @snoby 2026-05-04 (2× 4090, ctx=180K)"),
("long-text", 24, 22.3, "BENCHMARKS.md#qwen36-27b long-text.yml @noonghunna 2026-04-30"),
("long-vision", 24, 23.0, "BENCHMARKS.md#qwen36-27b long-vision.yml @noonghunna 2026-04-30"),
("bounded-thinking", 24, 21.7, "BENCHMARKS.md#qwen36-27b bounded-thinking.yml @noonghunna 2026-05-04"),
("minimal", 24, 22.4, "BENCHMARKS.md#qwen36-27b minimal.yml @noonghunna 2026-05-03 (mem-util 0.95, max-ctx 65536)"),
]
# =============================================================================
# Calibration: measured BENCHMARKS rows (peak per-card VRAM during bench)
# =============================================================================
CALIBRATION = {
"qwen3.6-27b": [
# (compose, vram_gb, measured_peak_gb, ctx_override_or_none, source_url)
("dual", 24, 23.6, None, "BENCHMARKS.md#qwen36-27b dual.yml @noonghunna 2026-04-29"),
("dual-turbo", 24, 19.8, None, "BENCHMARKS.md#qwen36-27b dual-turbo.yml @noonghunna 2026-04-29"),
("dual-dflash", 24, 23.6, None, "BENCHMARKS.md#qwen36-27b dual-dflash.yml @noonghunna 2026-04-29"),
("dual-dflash-noviz", 24, 23.8, None, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @noonghunna 2026-04-29"),
("dual4", 24, 23.5, None, "BENCHMARKS.md#qwen36-27b dual4.yml @whamp 2026-05-03"),
("dual4-dflash", 24, 22.0, None, "BENCHMARKS.md#qwen36-27b dual4-dflash.yml @whamp 2026-05-03"),
("dual-dflash-noviz", 24, 21.8, 180000, "BENCHMARKS.md#qwen36-27b dual-dflash-noviz.yml @snoby 2026-05-04 (2× 4090, ctx=180K)"),
("long-text", 24, 22.3, None, "BENCHMARKS.md#qwen36-27b long-text.yml @noonghunna 2026-04-30"),
("long-vision", 24, 23.0, None, "BENCHMARKS.md#qwen36-27b long-vision.yml @noonghunna 2026-04-30"),
("bounded-thinking", 24, 21.7, None, "BENCHMARKS.md#qwen36-27b bounded-thinking.yml @noonghunna 2026-05-04"),
("minimal", 24, 22.4, 65536, "BENCHMARKS.md#qwen36-27b minimal.yml @noonghunna 2026-05-03 (mem-util 0.95, max-ctx 65536)"),
],
"gemma-4-31b": [
# All TP=2 dual configs on 2× 3090 24 GB.
("gemma-dual-int8", 24, 22.2, None, "BENCHMARKS.md#gemma-4-31b dual/int8.yml @noonghunna 2026-05-08 (98K + max-num-seqs=4)"),
("gemma-dual-int8-262k", 24, 22.1, None, "BENCHMARKS.md#gemma-4-31b dual/int8.yml @noonghunna 2026-05-08 (262K + max-num-seqs=1)"),
("gemma-dual-dflash", 24, 22.7, None, "BENCHMARKS.md#gemma-4-31b dual/dflash.yml @noonghunna 2026-05-06 (n=7)"),
("gemma-dual-dflash", 24, 22.3, None, "BENCHMARKS.md#gemma-4-31b dual/dflash.yml @noonghunna 2026-05-08 (rebench)"),
("gemma-dual-awq", 24, 19.8, None, "BENCHMARKS.md#gemma-4-31b dual/awq.yml @noonghunna 2026-05-08 (AWQ-4bit, mem-util 0.85)"),
("gemma-dual", 24, 22.5, None, "BENCHMARKS.md#gemma-4-31b dual/docker-compose.yml @noonghunna matched-config rebench 2026-05-09"),
("gemma-dual-int8", 24, 22.5, 262144, "BENCHMARKS.md#gemma-4-31b dual/int8.yml matched-config rebench 2026-05-09 (262K seqs=2)"),
],
}
# =============================================================================
# Prediction
# =============================================================================
@dataclass
class Prediction:
model: str
weights_gb: float
kv_pool_gb: float
kv_pool_requested_gb: float
kv_pool_actual_gb: float # capped at available budget (vLLM behavior)
kv_pool_sliding_fixed_gb: float # Gemma sliding-window fixed term (0 for Qwen)
activation_gb: float
cudagraph_overhead_gb: float
dflash_draft_gb: float
drafter_gb: float
total_gb: float
vram_gb: float
budget_gb: float
pct_of_vram: float
verdict: str
notes: list[str]
def _weights_per_card_gb(spec, tp, weights_variant="default"):
"""Return per-card weights footprint in GB after TP split."""
if spec["model_family"] == "qwen3-next-hybrid":
return spec["weights_total_gb"] / tp
elif spec["model_family"] == "gemma4-swa-dense":
if weights_variant == "awq":
return spec["weights_awq_gb"] / tp
elif weights_variant == "bf16":
return spec["weights_bf16_gb"] / tp
else: # int4 default
return spec["weights_int4_gb"] / tp
raise ValueError(f"Unknown model_family: {spec['model_family']}")
def kv_pool_per_card_bytes(spec, kv_format, max_ctx, max_num_seqs, tp, mtp_n=0):
"""Per-card KV pool bytes for the attention layers.
GDN layers have a fixed-size recurrent state (not seq-len-dependent KV
cache), so they don't contribute here — they show up in activation_peak.
"""Per-card KV pool bytes (growing portion only).
Standard formula:
per_token_bytes = num_attn_layers × num_kv_heads × head_dim × 2 (K+V) × bytes
pool = per_token × max_ctx × max_num_seqs / TP
Returns a tuple (growing_per_card_bytes, sliding_fixed_per_card_bytes).
Sliding term is zero for models without sliding-window layers.
MTP n>0 adds n extra cached tokens per request for draft hidden states.
For Qwen 3.6 (DeltaNet hybrid):
Only the 16 full_attention layers grow KV. GDN layers have a fixed-size
recurrent state (not seq-len-dependent), so they show up in activation.
K and V stored independently → ×2 factor.
For Gemma 4 (SWA + dense MLP):
Only the 10 full_attention layers grow KV (at global_head_dim=512).
The 50 sliding_attention layers hold a FIXED window of 1024 tokens
(constant in ctx, contributes a separate small term).
K==V tying IS exploited by vLLM's allocator → ×1 factor (calibrated
against BENCHMARKS data; see docs/KV_MATH.md).
"""
bytes_per_kv_elem = KV_FORMAT_BYTES[kv_format]
per_token = (
spec["num_attn_layers"]
* spec["num_kv_heads"]
* spec["head_dim_attn"]
* 2 # K + V
* bytes_per_kv_elem
)
# KV heads split across TP ranks (heads must divide rank)
per_card_per_token = per_token / tp
effective_ctx = max_ctx + mtp_n * 32 # MTP draft caches a few extra positions
return per_card_per_token * effective_ctx * max_num_seqs
bpe = KV_FORMAT_BYTES[kv_format]
if spec["model_family"] == "qwen3-next-hybrid":
# K and V stored independently
per_token = (
spec["num_attn_layers"]
* spec["num_kv_heads"]
* spec["head_dim_attn"]
* 2 # K + V
* bpe
)
effective_ctx = max_ctx + mtp_n * 32
growing = (per_token / tp) * effective_ctx * max_num_seqs
return growing, 0.0
elif spec["model_family"] == "gemma4-swa-dense":
# K==V tied → ×1 storage
per_token_growing = (
spec["num_full_attn_layers"]
* spec["num_kv_heads"]
* spec["global_head_dim"]
* 1 # K==V tied; vLLM stores once
* bpe
)
# No MTP draft-token bump on Gemma — drafter is a separate model
growing = (per_token_growing / tp) * max_ctx * max_num_seqs
# Sliding-window fixed term — 50 layers × window × head_dim × 1 × bpe
sliding_fixed_total = (
spec["num_sliding_attn_layers"]
* spec["num_kv_heads"]
* spec["head_dim_sliding"]
* 1 # K==V tied here too
* bpe
* spec["sliding_window"]
)
sliding_per_card = sliding_fixed_total / tp
return growing, sliding_per_card
raise ValueError(f"Unknown model_family: {spec['model_family']}")
def gdn_activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp):
"""Per-card peak activation during DeltaNet GDN forward.
def activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp):
"""Per-card peak activation during prefill forward.
Theoretical scaling (PerfMamba arxiv 2511.22849): O(γ·D·N·L) per layer.
Empirical fit: linear in seq_len, with KV-format-dependent coefficient.
For Qwen 3.6 (DeltaNet GDN): linear in seq_len, KV-format-dependent
coefficient (PerfMamba O(γ·D·N·L) form, fla.ops.chunk implementation
details calibrated empirically).
Each GDN layer's `chunk_gated_delta_rule_fwd` materializes a block-wise
state tensor sized roughly (B, NT, H, V, K) × bytes. For Qwen3.6-27B:
NT = ceil(seq_len / 256), H = 16 K-heads, V = K = 128 dim, fp32 state.
The actual implementation has tiling/streaming that PerfMamba's pure
formula doesn't capture. The coefficient here is calibrated against
measured BENCHMARKS peaks (see CALIBRATION + tools/kv-calc.py
--calibration).
For Gemma 4 (dense MLP + SWA): mostly CONSTANT in seq_len because chunked
prefill bounds the MLP intermediate. Small per-token residual to keep
the solver smooth.
"""
coef = GDN_ACTIVATION_COEF[kv_format]
total = coef * spec["num_gdn_layers"] * max_ctx
return total / tp
if spec["model_family"] == "qwen3-next-hybrid":
coef = QWEN_GDN_ACTIVATION_COEF[kv_format]
return (coef * spec["num_gdn_layers"] * max_ctx) / tp
elif spec["model_family"] == "gemma4-swa-dense":
const_bytes = GEMMA_ACTIVATION_CONST_GB * 1e9
per_token = GEMMA_ACTIVATION_PER_TOKEN_BYTES * max_ctx
return (const_bytes + per_token) / tp
raise ValueError(f"Unknown model_family: {spec['model_family']}")
def cudagraph_overhead_gb(mem_util, tp):
@@ -191,6 +362,16 @@ def cudagraph_overhead_gb(mem_util, tp):
return base + tp_bump
def _validate_tp_for_spec(spec, tp):
valid_tp = spec.get("valid_tp")
if valid_tp and tp not in valid_tp:
raise ValueError(
f"TP={tp} invalid for {spec['model_id']} "
f"(num_kv_heads={spec['num_kv_heads']} cannot be divided across TP cleanly). "
f"Valid TP values: {valid_tp}"
)
def predict(
spec=QWEN36_27B,
kv_format="fp8_e5m2",
@@ -200,51 +381,98 @@ def predict(
mem_util=0.95,
vram_gb=24,
dflash_draft_gb=0.0,
drafter_gb=0.0,
mtp=False,
weights_variant="default",
) -> Prediction:
weights_gb = spec["weights_total_gb"] / tp
kv_pool_gb = kv_pool_per_card_bytes(spec, kv_format, max_ctx, max_num_seqs, tp,
mtp_n=3 if mtp else 0) / 1e9
activation_gb = gdn_activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp) / 1e9
overhead_gb = cudagraph_overhead_gb(mem_util, tp)
dflash_gb = dflash_draft_gb if tp == 1 else dflash_draft_gb / tp # draft splits with TP
total_gb = weights_gb + kv_pool_gb + activation_gb + overhead_gb + dflash_gb
"""Predict per-card VRAM usage.
# The verdict compares DEMAND (this prediction) against the engine's
# available budget = mem_util × vram_gb. vLLM will refuse to boot if
# demand exceeds this. (Measured peak during bench is a different number:
# it's roughly mem_util × VRAM because vLLM inflates the KV pool to fill
# the budget — see docs/KV_MATH.md.)
#
# Calibrated error band: ±1.5 GB on the breakdown, ±2 GB on total.
# This is a directional estimator, not a precise predictor.
vLLM caps KV pool to (budget - fixed_components), so the prediction
reflects what actually gets allocated. When requested > available,
verdict is TIGHT with a note about effective concurrency reduction.
Args:
drafter_gb: total drafter weight (MTP / DFlash) — split by TP.
dflash_draft_gb: legacy alias — folded into drafter_gb if set.
"""
_validate_tp_for_spec(spec, tp)
weights_gb = _weights_per_card_gb(spec, tp, weights_variant)
growing_b, sliding_b = kv_pool_per_card_bytes(
spec, kv_format, max_ctx, max_num_seqs, tp,
mtp_n=3 if mtp else 0,
)
kv_pool_requested_gb = growing_b / 1e9
kv_pool_sliding_fixed_gb = sliding_b / 1e9
activation_gb = activation_peak_per_card_bytes(spec, kv_format, max_ctx, tp) / 1e9
overhead_gb = cudagraph_overhead_gb(mem_util, tp)
# Drafter: prefer drafter_gb; fall back to legacy dflash_draft_gb.
drafter_total = drafter_gb if drafter_gb > 0 else dflash_draft_gb
drafter_per_card = drafter_total / tp if tp > 1 else drafter_total
fixed_gb = weights_gb + activation_gb + overhead_gb + drafter_per_card + kv_pool_sliding_fixed_gb
budget_gb = mem_util * vram_gb
pct = 100 * total_gb / budget_gb
available_for_kv = max(0.0, budget_gb - fixed_gb)
# vLLM caps the KV pool to fit available budget (PagedAttention allocator).
kv_pool_actual_gb = min(kv_pool_requested_gb, available_for_kv)
total_gb = fixed_gb + kv_pool_actual_gb
pct = 100 * total_gb / budget_gb if budget_gb > 0 else 999.0
notes = []
# Generous verdict bands matching the ±1.5 GB error.
if pct < 88:
verdict = "PASS"
elif pct < 108:
verdict = "TIGHT"
notes.append(f"demand within ±1.5 GB error of engine budget ({budget_gb:.1f} GB at mem_util={mem_util}) — likely boots, may need a small mem_util bump if pre-check refuses")
else:
verdict = "FAIL"
notes.append(f"demand {pct:.0f}% of engine budget ({budget_gb:.1f} GB at mem_util={mem_util}) — pre-check will refuse; raise mem_util, lower max_ctx/max_num_seqs, or swap KV format")
if kv_format == "turboquant_3bit_nc" and vram_gb < 24:
# Verdict logic:
# - FAIL: fixed components alone exceed budget (no room even for minimum KV).
# - TIGHT: requested KV pool exceeds available — vLLM will cap, effective
# concurrency reduced (BOOT OK, but `--max-num-seqs` may not be
# honored at full max_ctx).
# - PASS: requested KV fits with room to spare.
MIN_KV_GB = 1.0 # vLLM needs at least ~1 GB for paged-attention blocks
if available_for_kv < MIN_KV_GB:
verdict = "FAIL"
notes.append(
f"fixed components ({fixed_gb:.1f} GB) leave only {available_for_kv:.1f} GB for KV pool "
f"(need ≥{MIN_KV_GB:.1f} GB minimum); vLLM pre-check will refuse — "
f"lower max_ctx, drop a drafter, or raise mem_util"
)
elif kv_pool_requested_gb > available_for_kv * 1.05:
verdict = "TIGHT"
notes.append(
f"requested KV pool ({kv_pool_requested_gb:.1f} GB) > available ({available_for_kv:.1f} GB) — "
f"vLLM will cap to {available_for_kv:.1f} GB; effective concurrency may be lower than "
f"--max-num-seqs={max_num_seqs} at full max_ctx={max_ctx:,}"
)
else:
verdict = "PASS"
# Model-specific advisory notes (preserved from v1)
if kv_format == "turboquant_3bit_nc" and vram_gb < 24 and spec["model_family"] == "qwen3-next-hybrid":
notes.append("⚠ TQ3 KV on <24 GB cards: consider --kv-format fp8_e5m2 (see docs/HARDWARE.md, #47)")
if max_ctx > 50000 and tp == 1 and kv_format != "fp16":
if max_ctx > 50000 and tp == 1 and spec["model_family"] == "qwen3-next-hybrid" and kv_format != "fp16":
notes.append("⚠ single-card vLLM at >50K single-prompt: Cliff 2 territory (DeltaNet GDN forward); see docs/CLIFFS.md")
if spec["model_family"] == "gemma4-swa-dense" and kv_format == "fp8_e4m3":
notes.append("⚠ fp8_e4m3 on Ampere (sm_86): Triton `fp8e4nv` kernel unsupported; use int8_per_token_head instead (PR #40391 via #42102)")
if spec["model_family"] == "gemma4-swa-dense" and tp == 1 and vram_gb < 32:
notes.append("⚠ Gemma 4 31B TP=1 needs ≥32 GB VRAM; 24 GB Ampere boot-OOMs (model weights + drafter + min KV)")
if tp > 4:
notes.append("TP > 4 predictions are extrapolated; report deltas via scripts/report.sh --bench")
return Prediction(
model=spec["model_id"],
weights_gb=weights_gb,
kv_pool_gb=kv_pool_gb,
kv_pool_requested_gb=kv_pool_requested_gb,
kv_pool_actual_gb=kv_pool_actual_gb,
kv_pool_sliding_fixed_gb=kv_pool_sliding_fixed_gb,
activation_gb=activation_gb,
cudagraph_overhead_gb=overhead_gb,
dflash_draft_gb=dflash_gb,
drafter_gb=drafter_per_card,
total_gb=total_gb,
vram_gb=vram_gb,
budget_gb=budget_gb,
pct_of_vram=pct,
verdict=verdict,
notes=notes,
@@ -256,86 +484,135 @@ def fmt_prediction(p: Prediction, header: str = "") -> str:
if header:
lines.append(header)
lines.append("-" * len(header))
lines.append(f" Model weights: {p.weights_gb:>6.2f} GB / card")
lines.append(f" KV pool (attention): {p.kv_pool_gb:>6.2f} GB / card")
lines.append(f" Activation peak (GDN): {p.activation_gb:>6.2f} GB / card")
lines.append(f" Model: {p.model}")
lines.append(f" Weights: {p.weights_gb:>6.2f} GB / card")
if p.kv_pool_sliding_fixed_gb > 0.01:
lines.append(f" KV pool — sliding fixed: {p.kv_pool_sliding_fixed_gb:>6.2f} GB / card (constant, doesn't grow with ctx)")
if abs(p.kv_pool_requested_gb - p.kv_pool_actual_gb) > 0.05:
lines.append(f" KV pool — growing (req): {p.kv_pool_requested_gb:>6.2f} GB / card (requested)")
lines.append(f" KV pool — growing (cap): {p.kv_pool_actual_gb:>6.2f} GB / card (vLLM-capped to fit)")
else:
lines.append(f" KV pool — growing: {p.kv_pool_actual_gb:>6.2f} GB / card")
lines.append(f" Activation peak: {p.activation_gb:>6.2f} GB / card")
lines.append(f" Cudagraph + workspace: {p.cudagraph_overhead_gb:>6.2f} GB / card")
if p.dflash_draft_gb > 0:
lines.append(f" DFlash draft model: {p.dflash_draft_gb:>6.2f} GB / card")
if p.drafter_gb > 0:
lines.append(f" Drafter (MTP / DFlash): {p.drafter_gb:>6.2f} GB / card")
lines.append(f" ─────────────────────────────────────")
lines.append(f" Predicted demand total: {p.total_gb:>6.2f} GB / card ({p.pct_of_vram:.0f}% of engine budget)")
lines.append(f" Predicted total: {p.total_gb:>6.2f} GB / card ({p.pct_of_vram:.0f}% of {p.budget_gb:.1f} GB engine budget)")
lines.append(f" Verdict: {p.verdict}")
lines.append(f" (Note: measured peak during bench will be higher — vLLM fills the")
lines.append(f" remaining budget with KV pool inflation. Demand is the *lower bound*")
lines.append(f" that determines whether engine pre-check accepts the config.)")
for note in p.notes:
lines.append(f" Note: {note}")
return "\n".join(lines)
def run_calibration():
print("=" * 88)
print("Calibration — DEMAND prediction vs measured peak + engine budget per-card VRAM")
print("=" * 88)
print()
print(" Demand is what the engine NEEDS (lower-bound). Engine budget is what vLLM")
print(" ALLOCATES (≈ mem_util × VRAM). Measured peak is what nvidia-smi shows during")
print(" bench. vLLM inflates the KV pool to fill the budget, so peak ≈ budget.")
print(" The verdict — PASS / TIGHT / FAIL — is correct iff demand < budget.")
print()
print(f" {'compose':<22s} {'demand':>9s} {'budget':>9s} {'measured':>10s} {'verdict':>8s}")
print(f" {'─'*21:<22s} {'─'*8:>9s} {'─'*8:>9s} {'─'*9:>10s} {'─'*7:>8s}")
# =============================================================================
# Calibration runner
# =============================================================================
def _resolve_compose_for_predict(model_key, compose_id, vram, ctx_override=None):
"""Resolve a compose preset to predict() kwargs, applying optional ctx override."""
spec = MODEL_SPECS[model_key]
cfg = COMPOSES[model_key][compose_id]
max_ctx = ctx_override if ctx_override is not None else cfg["max_ctx"]
kwargs = dict(
spec=spec,
kv_format=cfg["kv_format"],
max_ctx=max_ctx,
max_num_seqs=cfg["max_num_seqs"],
tp=cfg["tp"],
mem_util=cfg["mem_util"],
vram_gb=vram,
mtp=cfg.get("mtp", False),
weights_variant=cfg.get("weights_variant", "default"),
drafter_gb=cfg.get("drafter_gb", 0.0),
dflash_draft_gb=cfg.get("dflash_draft_gb", 0.0),
)
return kwargs
def _calibration_block(model_key: str) -> tuple[int, int]:
"""Print calibration table for one model. Returns (correct, total)."""
rows = CALIBRATION.get(model_key, [])
if not rows:
return 0, 0
spec = MODEL_SPECS[model_key]
print(f"== {spec['model_id']} ==")
print(f" {'compose':<26s} {'predicted':>10s} {'budget':>9s} {'measured':>10s} {'verdict':>8s}")
print(f" {'─'*25:<26s} {'─'*9:>10s} {'─'*8:>9s} {'─'*9:>10s} {'─'*7:>8s}")
correct = 0
for compose, vram, measured, _src in CALIBRATION:
cfg = COMPOSES[compose]
max_ctx = cfg["max_ctx"]
if compose == "dual-dflash-noviz" and abs(measured - 21.8) < 0.05:
max_ctx = 180000 # snoby's 4090 row
p = predict(
kv_format=cfg["kv_format"],
max_ctx=max_ctx,
max_num_seqs=cfg["max_num_seqs"],
tp=cfg["tp"],
mem_util=cfg["mem_util"],
vram_gb=vram,
dflash_draft_gb=cfg.get("dflash_draft_gb", 0.0),
mtp=cfg.get("mtp", False),
)
budget = cfg["mem_util"] * vram
# Verdict is "correct" if predicted PASS/TIGHT and measured < vram (boot OK),
# or predicted FAIL and measured > vram (boot would fail).
verdict_correct = "✓" if p.verdict in ("PASS", "TIGHT") and measured < vram else ("⨯" if p.verdict == "FAIL" and measured < vram else "✓")
if verdict_correct == "✓":
for row in rows:
compose, vram, measured, ctx_override, _src = row
kwargs = _resolve_compose_for_predict(model_key, compose, vram, ctx_override)
p = predict(**kwargs)
# Verdict is "correct" if (PASS/TIGHT and measured fits) or (FAIL and would OOM).
# We don't have negative (FAIL) data points in BENCHMARKS — every row booted —
# so verdict_correct simplifies to: PASS/TIGHT and measured < vram.
if p.verdict in ("PASS", "TIGHT") and measured < vram:
mark = "✓"
correct += 1
print(f" {compose:<22s} {p.total_gb:>7.2f} GB {budget:>7.2f} GB {measured:>8.2f} GB {p.verdict:>7s} {verdict_correct}")
elif p.verdict == "FAIL" and measured >= vram:
mark = "✓"
correct += 1
else:
mark = "⨯"
compose_disp = compose if ctx_override is None else f"{compose}@{ctx_override//1024}K"
print(f" {compose_disp:<26s} {p.total_gb:>8.2f} GB {p.budget_gb:>7.2f} GB {measured:>8.2f} GB {p.verdict:>7s} {mark}")
n = len(CALIBRATION)
print()
print(f" Verdict accuracy: {correct}/{n} ({100*correct/n:.0f}%)")
print(f" Verdict accuracy: {correct}/{len(rows)} ({100*correct/len(rows):.0f}%)")
print()
print(" The DEMAND number is the calculator's output; users use it to plan.")
print(" The MEASURED column is for sanity: every passing config should show")
print(" measured < VRAM (else boot would have failed). If the calculator says")
print(" PASS but measured > VRAM × mem_util, there's hidden overhead the model")
print(" doesn't capture — file an issue with your `bash scripts/report.sh --bench`")
print(" output and we'll re-calibrate.")
return correct, len(rows)
def solve_max_ctx(spec, kv_format, max_num_seqs, tp, mem_util, vram_gb, dflash_draft_gb, mtp):
"""Binary search for the largest max_ctx that keeps demand <= budget."""
def run_calibration():
print("=" * 88)
print("Calibration — predicted per-card VRAM vs measured BENCHMARKS rows")
print("=" * 88)
print()
print(" Predicted = weights + activation + overhead + drafter + (KV capped at available).")
print(" Budget = mem_util × VRAM. Measured = nvidia-smi peak during bench (target ≈ budget).")
print(" Verdict ✓ iff PASS/TIGHT and measured < VRAM (boot OK).")
print()
total_c, total_n = 0, 0
for model_key in ("qwen3.6-27b", "gemma-4-31b"):
c, n = _calibration_block(model_key)
total_c += c
total_n += n
if total_n > 0:
print(f"Overall: {total_c}/{total_n} ({100*total_c/total_n:.0f}%)")
print()
print("Notes:")
print(" - This is a directional estimator (±1.5 GB error band on the breakdown).")
print(" - vLLM's `gpu_worker.py` boot log is the authoritative source.")
print(" - If predicted PASS but measured > budget, file an issue with `scripts/report.sh --bench`.")
# =============================================================================
# Max-ctx solver
# =============================================================================
def solve_max_ctx(spec, kv_format, max_num_seqs, tp, mem_util, vram_gb,
drafter_gb=0.0, dflash_draft_gb=0.0, mtp=False, weights_variant="default"):
"""Binary search for the largest max_ctx that keeps the verdict at PASS or TIGHT."""
lo, hi = 1024, spec.get("max_ctx_supported", 262144)
best = 0
while lo <= hi:
mid = (lo + hi) // 2
# Round to nearest 1024 for cleaner numbers
mid = (mid // 1024) * 1024
mid = (mid // 1024) * 1024 # round to nearest 1024 for cleaner numbers
if mid == 0:
break
p = predict(spec, kv_format=kv_format, max_ctx=mid, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=vram_gb,
dflash_draft_gb=dflash_draft_gb, mtp=mtp)
if p.verdict in ("PASS", "TIGHT") and p.pct_of_vram < 100:
p = predict(
spec=spec, kv_format=kv_format, max_ctx=mid, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=vram_gb,
drafter_gb=drafter_gb, dflash_draft_gb=dflash_draft_gb,
mtp=mtp, weights_variant=weights_variant,
)
if p.verdict in ("PASS", "TIGHT"):
best = mid
lo = mid + 1024
else:
@@ -343,22 +620,54 @@ def solve_max_ctx(spec, kv_format, max_num_seqs, tp, mem_util, vram_gb, dflash_d
return best
# =============================================================================
# CLI
# =============================================================================
def _all_compose_choices() -> list[str]:
"""Flat list of compose names across all models for argparse choices."""
out = []
for model_key in COMPOSES:
out.extend(COMPOSES[model_key].keys())
return sorted(set(out))
def _resolve_compose_model(compose_name: str, explicit_model: Optional[str]) -> str:
"""Infer model from compose name if --model not given.
Composes are namespaced by prefix; Qwen uses bare names, Gemma uses gemma-*.
"""
if explicit_model:
return explicit_model
for model_key, composes in COMPOSES.items():
if compose_name in composes:
return model_key
return "qwen3.6-27b" # back-compat default
def main():
p = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
p.add_argument("--compose", choices=sorted(COMPOSES.keys()),
p.add_argument("--model", choices=sorted(MODEL_SPECS.keys()),
help="Which model to predict for. Default: qwen3.6-27b (back-compat) or inferred from --compose.")
p.add_argument("--compose", choices=_all_compose_choices(),
help="Use a shipped compose's defaults. Override individual flags below.")
p.add_argument("--kv-format", choices=sorted(KV_FORMAT_BYTES.keys()),
help="KV cache format. Default: from --compose, or fp8_e5m2.")
p.add_argument("--max-ctx", type=int, help="max_model_len. Default: from --compose, or 180000.")
p.add_argument("--max-num-seqs", type=int, help="max_num_seqs. Default: from --compose, or 1.")
p.add_argument("--tp", type=int, choices=[1, 2, 4], help="tensor_parallel_size. Default: from --compose, or 1.")
p.add_argument("--tp", type=int, choices=[1, 2, 4, 8, 16], help="tensor_parallel_size. Default: from --compose, or 1.")
p.add_argument("--mem-util", type=float, help="gpu_memory_utilization. Default: from --compose, or 0.95.")
p.add_argument("--vram", type=float, default=24, help="VRAM per card in GB. Default 24.")
p.add_argument("--mtp", action="store_true", help="MTP n=3 enabled (adds small KV overhead per request).")
p.add_argument("--mtp", action="store_true", default=None, help="MTP enabled (Qwen: n=3 built-in; Gemma: external drafter).")
p.add_argument("--no-mtp", dest="mtp", action="store_false")
p.add_argument("--dflash-draft-gb", type=float, default=0.0, help="DFlash draft model size in GB (0 if not using DFlash).")
p.add_argument("--calibration", action="store_true", help="Print predicted vs measured for all calibration points.")
p.add_argument("--solve-max-ctx", action="store_true", help="Binary-search for the largest max_ctx that fits given the other parameters.")
p.add_argument("--drafter-gb", type=float, default=None,
help="Drafter model size in GB (MTP / DFlash). 0 if not using a drafter.")
p.add_argument("--dflash-draft-gb", type=float, default=None,
help="(deprecated alias for --drafter-gb)")
p.add_argument("--weights-variant", choices=["default", "int4", "awq", "bf16"], default=None,
help="Gemma 4 only: which weight quant variant. Default: from --compose, or int4.")
p.add_argument("--calibration", action="store_true", help="Print predicted vs measured for both models.")
p.add_argument("--solve-max-ctx", action="store_true", help="Binary-search for the largest max_ctx that fits.")
p.add_argument("--json", action="store_true", help="Output prediction as JSON.")
args = p.parse_args()
@@ -366,56 +675,83 @@ def main():
run_calibration()
return 0
# Resolve defaults from compose preset (used by both modes)
# Resolve model: explicit --model > inferred from --compose > qwen3.6-27b
model_key = _resolve_compose_model(args.compose, args.model) if args.compose else (args.model or "qwen3.6-27b")
spec = MODEL_SPECS[model_key]
# Resolve compose-derived defaults
if args.compose:
cfg = COMPOSES[args.compose]
# Compose must belong to the resolved model
if args.compose not in COMPOSES[model_key]:
print(f"ERROR: --compose {args.compose} is not in --model {model_key}'s compose list.", file=sys.stderr)
print(f" Available for {model_key}: {', '.join(sorted(COMPOSES[model_key].keys()))}", file=sys.stderr)
return 2
cfg = COMPOSES[model_key][args.compose]
kv_format = args.kv_format or cfg["kv_format"]
max_ctx = args.max_ctx or cfg["max_ctx"]
max_num_seqs = args.max_num_seqs or cfg["max_num_seqs"]
tp = args.tp or cfg["tp"]
mem_util = args.mem_util if args.mem_util is not None else cfg["mem_util"]
mtp = args.mtp if args.mtp is not None else cfg.get("mtp", False)
dflash_gb = args.dflash_draft_gb or cfg.get("dflash_draft_gb", 0.0)
header = f"Predicted budget — {args.compose}.yml on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
drafter_gb = args.drafter_gb if args.drafter_gb is not None else cfg.get("drafter_gb", 0.0)
dflash_gb = args.dflash_draft_gb if args.dflash_draft_gb is not None else cfg.get("dflash_draft_gb", 0.0)
weights_variant = args.weights_variant or cfg.get("weights_variant", "default")
header = f"Predicted budget — {model_key} / {args.compose} on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx:,}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
else:
kv_format = args.kv_format or "fp8_e5m2"
max_ctx = args.max_ctx or 180000
max_num_seqs = args.max_num_seqs or 1
tp = args.tp or 1
mem_util = args.mem_util if args.mem_util is not None else 0.95
mtp = bool(args.mtp)
dflash_gb = args.dflash_draft_gb
header = f"Predicted budget — custom config on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
mtp = bool(args.mtp) if args.mtp is not None else False
drafter_gb = args.drafter_gb or 0.0
dflash_gb = args.dflash_draft_gb or 0.0
weights_variant = args.weights_variant or "default"
header = f"Predicted budget — {model_key} custom config on {args.vram} GB VRAM (kv={kv_format}, ctx={max_ctx:,}, seqs={max_num_seqs}, TP={tp}, mem={mem_util})"
try:
_validate_tp_for_spec(spec, tp)
except ValueError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
return 2
if args.solve_max_ctx:
# Pin max_ctx very high; we'll search for the actual largest that fits.
best = solve_max_ctx(QWEN36_27B, kv_format=kv_format, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=args.vram,
dflash_draft_gb=dflash_gb, mtp=mtp)
best = solve_max_ctx(
spec, kv_format=kv_format, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=args.vram,
drafter_gb=drafter_gb, dflash_draft_gb=dflash_gb, mtp=mtp,
weights_variant=weights_variant,
)
if best > 0:
pred_at_best = predict(kv_format=kv_format, max_ctx=best, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=args.vram,
dflash_draft_gb=dflash_gb, mtp=mtp)
print(f"Max-ctx solver — {kv_format}, seqs={max_num_seqs}, TP={tp}, mem_util={mem_util}, VRAM={args.vram} GB")
print(f" Largest max_ctx that fits: {best:,} tokens")
print(f" At that ctx: predicted demand = {pred_at_best.total_gb:.2f} GB / card ({pred_at_best.pct_of_vram:.0f}% of budget)")
print(f" Verdict at that ctx: {pred_at_best.verdict}")
print()
print("Note: this is a directional estimate (±1.5 GB error band). The vLLM engine")
print("pre-check (gpu_worker.py boot log) is authoritative.")
pred_at_best = predict(
spec=spec, kv_format=kv_format, max_ctx=best, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=args.vram,
drafter_gb=drafter_gb, dflash_draft_gb=dflash_gb, mtp=mtp,
weights_variant=weights_variant,
)
if args.json:
out = pred_at_best.__dict__.copy()
out["solved_max_ctx"] = best
print(json.dumps(out, indent=2))
else:
print(f"Max-ctx solver — {model_key} / {kv_format}, seqs={max_num_seqs}, TP={tp}, mem_util={mem_util}, VRAM={args.vram} GB")
print(f" Largest max_ctx that fits: {best:,} tokens")
print(f" At that ctx: predicted = {pred_at_best.total_gb:.2f} GB / card ({pred_at_best.pct_of_vram:.0f}% of budget)")
print(f" Verdict at that ctx: {pred_at_best.verdict}")
for note in pred_at_best.notes:
print(f" Note: {note}")
print()
print("Note: this is a directional estimate (±1.5 GB error band). The vLLM engine")
print("pre-check (gpu_worker.py boot log) is authoritative.")
else:
print(f"No max_ctx fits at this config on {args.vram} GB. Reduce TP, swap KV format, or get bigger cards.")
return 0
pred = predict(
kv_format=kv_format,
max_ctx=max_ctx,
max_num_seqs=max_num_seqs,
tp=tp,
mem_util=mem_util,
vram_gb=args.vram,
dflash_draft_gb=dflash_gb,
mtp=mtp,
spec=spec, kv_format=kv_format, max_ctx=max_ctx, max_num_seqs=max_num_seqs,
tp=tp, mem_util=mem_util, vram_gb=args.vram,
drafter_gb=drafter_gb, dflash_draft_gb=dflash_gb, mtp=mtp,
weights_variant=weights_variant,
)
if args.json:
@@ -423,11 +759,9 @@ def main():
else:
print(fmt_prediction(pred, header=header))
print()
print("Anchored to: PerfMamba (arxiv 2511.22849), TurboQuant (arxiv 2504.19874),")
print("PagedAttention (arxiv 2309.06180). Calibrated against BENCHMARKS.md rows.")
print("Error band: ±1.5 GB on the breakdown. Verdicts within 5% of the budget are TIGHT.")
print("Run `tools/kv-calc.py --calibration` to see predicted-vs-measured for all anchors.")
print("Run `tools/kv-calc.py --solve-max-ctx ...` to find the largest max_ctx that fits your config.")
print("Run `tools/kv-calc.py --solve-max-ctx ...` to find the largest max_ctx that fits.")
print("See docs/KV_MATH.md for math + per-model architecture details.")
return 0 if pred.verdict in ("PASS", "TIGHT") else 1