9 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 e05f1969bc fix(launch): exit cleanly on stdin EOF in wizard prompts
Release / release (push) Failing after 52s
The two `read -rp` loops in `choose()` and the variant-selection step
spun infinitely when stdin closed mid-prompt (piped input shorter than
the wizard asks, or test/CI invocations). Now both loops check `read`'s
exit code; on EOF, print a clear message and SIGINT the parent so the
whole process exits cleanly (exit 130).

Doesn't affect interactive users — they type real input. Affects only
non-TTY piped invocations of the wizard, which should use --variant
<name> to skip the wizard entirely.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-13 20:44:35 +00:00
noonghunna 12a33fbdd1 feat(scripts): add hardware-aware setup picker 2026-05-13 20:44:35 +00:00
github-actions[bot] 45fa421e93 chore(changelog): regenerate for v0.5.4 [skip ci] 2026-05-13 18:41:00 +00:00
noonghunna 22bf2e9398 fix(scripts): make submit-bench issue-first
Release / release (push) Failing after 53s
2026-05-13 18:40:43 +00:00
github-actions[bot] 99a0b66224 chore(changelog): regenerate for v0.5.3 [skip ci] 2026-05-13 18:24:40 +00:00
noonghunna ef770322f4 feat(scripts): add submit-bench flow
Release / release (push) Failing after 51s
2026-05-13 18:24:14 +00:00
github-actions[bot] 301083491b chore(changelog): regenerate for v0.5.2 [skip ci] 2026-05-13 18:04:00 +00:00
noonghunna 26985527f7 Add hardware-aware compose preflight
Release / release (push) Failing after 50s
2026-05-13 18:03:37 +00:00
github-actions[bot] 642dfba8ef chore(changelog): regenerate for v0.5.1 [skip ci] 2026-05-13 17:02:10 +00:00
50 changed files with 2376 additions and 42 deletions
+10 -5
View File
@@ -38,11 +38,16 @@
# GPU selection
# -----------------------------------------------------------------------------
# Which GPUs to expose to docker. Single-card composes use just `0`,
# dual-card composes use `0,1`. The compose files set this themselves;
# override here if your physical layout differs (e.g. you run dual-card
# on cards 2,3).
# CUDA_VISIBLE_DEVICES=0,1
# Which GPU(s) to expose to docker.
#
# scripts/switch.sh auto-selects the largest eligible card for TP=1 vLLM
# composes. Override with CLUB3090_GPU when you know which physical GPU should
# run a single-card compose (for example, a 24 GB card beside a 16 GB card).
# CLUB3090_GPU=1
#
# For direct `docker compose ... up` runs, or TP>=2 non-default layouts, set
# NVIDIA_VISIBLE_DEVICES explicitly. Use comma-separated physical indices.
# NVIDIA_VISIBLE_DEVICES=0,1
# -----------------------------------------------------------------------------
@@ -0,0 +1,20 @@
## Rig bench submission
> ⚠ Most bench submissions go through an issue (see `CONTRIBUTING.md` "Submitting your bench").
> This PR template is for contributors who explicitly chose the direct-PR path.
> The maintainer may redirect to an issue thread before merge.
<!-- This PR was auto-generated by `bash scripts/submit-bench.sh --auto-submit --as-pr --tag <TAG>`. -->
<!-- Review the row below; the PR reviewer may move it within the target section. -->
### New row
<!-- The generated BENCHMARKS.md row goes here -->
### Rig
<!-- Output of `bash scripts/report.sh` (redacted) -->
### Full results
See `results/rebench/<TAG>/REPORT.md` for the full per-phase breakdown.
+46
View File
@@ -16,6 +16,52 @@ history; SemVer takes over from `v0.3.0` onward.
---
## v0.5.4 — 2026-05-13
### 🐛 Bug fixes
- fix(scripts): make submit-bench issue-first ([22bf2e9](https://github.com/noonghunna/club-3090/commit/22bf2e9398c7907aae6b62809bde30e111e4a700))
[Pin: `git checkout v0.5.4`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.3...v0.5.4)
## v0.5.3 — 2026-05-13
### ✨ Features
- feat(scripts): add submit-bench flow ([ef77032](https://github.com/noonghunna/club-3090/commit/ef770322f43724f612a80393f547e5da218b5bf7))
[Pin: `git checkout v0.5.3`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.2...v0.5.3)
## v0.5.2 — 2026-05-13
### 🎯 New models + serving paths
- Add hardware-aware compose preflight ([2698552](https://github.com/noonghunna/club-3090/commit/26985527f75d8da2a32a8a2f985989d5dcf9e89a))
[Pin: `git checkout v0.5.2`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.1...v0.5.2)
## v0.5.1 — 2026-05-13
### 🐛 Bug fixes
- fix(qwen): PR #35936 overlay — sidecar pattern to resolve Genesis RO-mount conflict ([6617e1e](https://github.com/noonghunna/club-3090/commit/6617e1e090a6f52708aaf83821a92b612c4ad869))
### 📝 Documentation
- docs: clarify MODEL_DIR — second drive / HF cache / Windows-WSL ([1678ca0](https://github.com/noonghunna/club-3090/commit/1678ca0c8ba43ea09fad9073639a48337f2ff163))
- docs(upstream): correct stale vllm#40807 row + add #40798/#42215 row ([14ffe45](https://github.com/noonghunna/club-3090/commit/14ffe45667fb0aa292839be05ae3ad0d139d3a04))
[Pin: `git checkout v0.5.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.0...v0.5.1)
## v0.5.0 — 2026-05-12
+35 -1
View File
@@ -8,7 +8,7 @@ Thanks for being here. This repo collects working recipes for serving big LLMs o
### ✅ Yes please
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. If you're unsure what to run before measuring, start with `bash scripts/setup.sh` and `bash scripts/launch.sh`; both wizards mark hardware-fit. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
- **Power-cap efficiency curves.** `sudo bash scripts/power-cap-sweep.sh --cooling air|water|aio --load-mode decode-concurrent --concurrency auto --bench-runs 3` produces cross-rig efficiency-knee data ([discussion #86](https://github.com/noonghunna/club-3090/discussions/86)). ~15-20 min for a 30-cap sweep on a 3090/4090/5090. Especially valuable on cards we don't have anchors for yet (A5000/A6000, 4080, 5060 Ti / 5080, modded variants). **Keep `--step-size 10` (the default).** Larger step-sizes (e.g. `--step-size 50`) are too coarse for the efficiency knee and only useful for quick smoke tests. See [docs/HARDWARE.md](docs/HARDWARE.md#cross-rig-power-cap-data-anchor-points) for the full canonical command and rationale.
- **Bug reports with the data we ask for.** The [bug report template](https://github.com/noonghunna/club-3090/issues/new?template=bug-report.yml) leads with `bash scripts/report.sh > my-rig.md` (add `--verify` to include verify-full output, `--soak` to also run SOAK_MODE=continuous if you suspect a multi-turn agent cliff) — single command captures the rig state we'd otherwise ask for individually (hardware, container state, Genesis patches, KV pool sizing, engine config). With that paste, the first reply is usually a fix or a clear next step instead of "can you send me…".
- **Bug reproductions / minimum repros for upstream issues.** vLLM / llama.cpp / Genesis bugs that affect this stack are most useful when they have a one-paragraph reduction. Drop them in an issue or open a draft PR adding a reproducer to `verify-stress.sh`.
@@ -45,6 +45,40 @@ Two GitHub channels, two different shapes of conversation. Picking the right one
---
## Submitting your bench
The matrix is hand-curated — the canonical path is to file an **issue** with your rig + numbers; we'll review, ask clarifying questions, and integrate.
After running `bash scripts/rebench-full.sh`, generate a paste-ready row:
```bash
bash scripts/submit-bench.sh --tag <your-tag>
```
The script writes `results/rebench/<tag>/BENCHMARKS-row.md`. To submit:
### Path A — Auto-issue (recommended, requires `gh auth login`)
```bash
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit
```
Opens an issue via `gh issue create` with your rig + row pre-filled.
### Path B — Manual issue (no tools beyond browser)
Open https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml and paste the row + your `rig.txt` into the body.
### Path C — Direct PR (advanced)
```bash
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit --as-pr
```
For contributors who know the `BENCHMARKS.md` section structure and want to propose the exact row. The maintainer may still redirect to an issue thread for context-gathering before merge — direct PRs aren't a fast-path bypass.
---
## Process for non-trivial changes
1. **Open an issue first** for anything bigger than a typo fix or a one-line measurement contribution. We'll either align on shape or explain why we'd land it differently — saves you a wasted afternoon.
+8 -6
View File
@@ -66,13 +66,15 @@ Bench protocol: 3 warm + 5 measured runs of the canonical narrative + code promp
git clone https://github.com/noonghunna/club-3090.git
cd club-3090
# 2. Download + SHA-verify the model (~20 GB; clones Genesis patches too)
# (asks you where to put model weights — pick in-repo default, ~/models, or
# a custom path on a different drive. To skip the prompt:
# `export MODEL_DIR=/mnt/your-drive/models` before running. See FAQ + .env.example.)
bash scripts/setup.sh qwen3.6-27b
# 2. Pick/download + SHA-verify the model (interactive hardware-aware picker)
# (asks you which model, then where to put model weights — pick in-repo
# default, ~/models, or a custom path on a different drive. To skip prompts:
# `export MODEL_DIR=/mnt/your-drive/models` and pass the model name. See FAQ.)
bash scripts/setup.sh
# Or scripted:
# bash scripts/setup.sh qwen3.6-27b
# 3. Pick a config + boot it (interactive wizard — asks engine / cards / workload)
# 3. Pick a config + boot it (interactive hardware-aware wizard — asks cards / workload)
bash scripts/launch.sh
# Or skip the wizard:
# bash scripts/launch.sh --variant vllm/default # single-card chat (recommended)
+7 -1
View File
@@ -130,6 +130,12 @@ If your numbers on the same compose look different from ours by >15%, the most l
## Setup
### How do I pick the right model + variant?
For a first install, run `bash scripts/setup.sh` with no model argument in a normal terminal. It opens a hardware-aware model picker, marks Qwen / Gemma / Both as eligible or not for your detected GPUs, then continues into the existing download flow.
After setup, run `bash scripts/launch.sh`. Its existing cards + workload wizard now marks compose variants with hardware fit and picks the recommended default for the rig (`vllm/long-text` on one 24 GB card, `vllm/dual` on matched 2× 3090). Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, plus `setup.sh --help` / `launch.sh --help`.
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
@@ -140,7 +146,7 @@ Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
bash scripts/setup.sh qwen3.6-27b
```
2. **`.env` file at repo root** — picked up automatically on every script run. See [`.env.example`](../.env.example).
3. **Interactive prompt** — `bash scripts/setup.sh qwen3.6-27b` with nothing set offers three choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
3. **Interactive prompt** — `bash scripts/setup.sh` with nothing set first asks which model to download, then offers three model-dir choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
4. **Silent fallback** — `<repo>/models-cache/`. Functional but pollutes the git tree; not recommended.
Every script that touches model paths reads from the same `MODEL_DIR`. The compose YAMLs' volume mount is `${MODEL_DIR:-...}:/root/.cache/huggingface` — once set, every container reads + writes there.
+16
View File
@@ -33,6 +33,22 @@ The recipes are written against 3090 specifically but should work on:
**Won't work:** anything with <20 GB VRAM (3060, 3070, stock 3080, 3080 Ti). The 27B model in INT4 is ~18 GB — KV pool + activations push past 24 GB on smaller cards even with aggressive quantization. **Modded 20 GB 3080s do work** (see row above) — the mod gives them enough headroom for the 27B + TQ K8V4 KV path on TP=2, with `mem-util=0.82` to absorb cudagraph profiling overhead.
### Mismatched / heterogeneous GPUs
`scripts/switch.sh` reads hardware metadata from the vLLM compose headers before starting Docker. The preflight checks required GPU count, per-GPU VRAM, tensor parallel size, and any hard SM floor.
For TP=1 vLLM composes, `switch.sh` auto-selects the largest eligible GPU and exports it through `NVIDIA_VISIBLE_DEVICES`. On a mixed 16 GB + 24 GB rig, `bash scripts/switch.sh vllm/default` should pick the 24 GB card instead of trying to boot on GPU 0 blindly.
Overrides:
```bash
CLUB3090_GPU=1 bash scripts/switch.sh vllm/default
NVIDIA_VISIBLE_DEVICES=2,3 bash scripts/switch.sh vllm/dual
bash scripts/switch.sh --force vllm/gemma-mtp-tp1
```
Use `--force` only when you are intentionally testing an unsupported combo. Example: `vllm/gemma-mtp-tp1` is now preflight-blocked on a 24 GB 3090 because the compose is preserved for 32 GB / newer-SM single-card rigs.
### Note for sub-24 GB cards
On 20 GB cards (modded 3080) the cudagraph-profiling overhead is a meaningful slice of available VRAM. Drop `--gpu-memory-utilization` to **0.82** (vs shipped 0.95 for 24 GB). vLLM nightly's `gpu_worker.py` reports the equivalent effective KV size in the boot log; tune to keep activation headroom for the ~15K tool-prefill peak (verify-full check 8). Credit: [@troymroberts](https://github.com/troymroberts).
@@ -58,6 +58,10 @@
# - max-num-seqs default 4 (3dluvr's 1 is single-stream-only); override
# MAX_NUM_SEQS=1 for max-context single-stream like 3dluvr.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-gemma-4-31b-awq:
# Latest nightly that contains PR #41745 (Gemma 4 MTP). AWQ doesn't need
@@ -72,6 +76,7 @@ services:
- ../../cache/torch_compile_awq:/root/.cache/vllm/torch_compile_cache
- ../../cache/triton_awq:/root/.triton/cache
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -89,6 +89,10 @@
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-gemma-4-31b-mtp-bf16:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -118,6 +122,7 @@ services:
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -93,6 +93,10 @@
# 1. PR #41703 z-lab DFlash drafter — see ../../patches/vllm-gemma4-dflash/README
# 2. PR #40391 per-token-head — see ../../patches/vllm-gemma4-dflash-int8/README
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-gemma-4-31b-dflash-int8:
# SAME pin as dual-dflash.yml — DFlash overlay was rebased to e47c98ef.
@@ -139,6 +143,7 @@ services:
- ../../patches/vllm-gemma4-dflash-int8/v1/worker/kv_cache_shape_utils.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/kv_cache_shape_utils.py:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -65,6 +65,10 @@
# - default (auto = bfloat16) → sidesteps all the above.
# Smaller KV pool than fp8 would give but it's the only Ampere-shippable.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-gemma-4-31b-dflash:
# Nightly bumped to 2026-05-06 to match rebase target proximity (Codex
@@ -104,6 +108,7 @@ services:
- ../../patches/vllm-gemma4-dflash/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -48,6 +48,10 @@
# - default (auto = bfloat16) → sidesteps all fp8 kernel paths
# Smaller KV pool than fp8 but at 32K test ctx that's not the bottleneck.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-gemma-4-31b-mtp:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -63,6 +67,7 @@ services:
- ../../cache/triton:/root/.triton/cache
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -177,6 +177,11 @@
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
# Requires-sm: 9.0+
services:
vllm-gemma-4-31b-mtp-int8-tq3:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -213,6 +218,7 @@ services:
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -89,6 +89,10 @@
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-gemma-4-31b-mtp-int8:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -118,6 +122,7 @@ services:
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
# --------------------------------------------------------------------
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -48,6 +48,11 @@
# - default (auto = bfloat16) → sidesteps all fp8 kernel paths
# Smaller KV pool than fp8 but at 32K test ctx that's not the bottleneck.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 32
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
# Requires-sm: 9.0+
services:
vllm-gemma-4-31b-mtp-tp1:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -63,6 +68,7 @@ services:
- ../../cache/triton:/root/.triton/cache
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -20,6 +20,10 @@
# Run:
# docker compose -f dual/bf16.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-bf16:
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/bf16.yml — 2026-05-08
@@ -59,6 +63,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -48,6 +48,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/carnice-bf16mtp.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-carnice-bf16mtp:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -82,6 +86,7 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# PCIe-only stack — disable NCCL features that assume NVLink.
@@ -43,6 +43,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/dflash-noviz.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-dflash:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -77,6 +81,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -67,6 +67,10 @@
#
# docker compose -f dual/dflash.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-dflash:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -101,6 +105,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -49,6 +49,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/docker-compose.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual:
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
@@ -89,6 +93,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# PCIe-only stack — disable NCCL features that assume NVLink.
@@ -20,6 +20,10 @@
# Run:
# docker compose -f dual/int8.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-int8:
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8.yml — 2026-05-08
@@ -59,6 +63,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -75,6 +75,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink-dflash-noviz.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-nvlink-dflash-noviz:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -109,6 +113,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
@@ -71,6 +71,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink-dflash.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-nvlink-dflash:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -105,6 +109,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
@@ -61,6 +61,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink-turbo.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-nvlink-turbo:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -102,6 +106,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
@@ -57,6 +57,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/nvlink.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-nvlink:
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
@@ -95,6 +99,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# NVLink bridge present — let NCCL use P2P (don't disable it the way
@@ -33,6 +33,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/qwopus-bf16mtp.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwopus-bf16mtp:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -58,6 +62,7 @@ services:
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -61,6 +61,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/tq3-mtp-genesis.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-tq3-mtp-genesis:
# Pinned to a Genesis v7.72.2 known-good vLLM nightly. The canonical
@@ -101,6 +105,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -59,6 +59,10 @@
# Run (only when all 5 upstream fixes are available):
# docker compose -f dual/tq3-mtp.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-int8-tq3:
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8-tq3.yml — 2026-05-08
@@ -106,6 +110,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- VLLM_ENFORCE_EAGER=1
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
@@ -35,6 +35,10 @@
# To run:
# docker compose -f dual/tq3-nomtp.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-int8-tq3-nomtp:
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8-tq3.yml — 2026-05-08
@@ -90,6 +94,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -45,6 +45,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f dual/turbo.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 2
# Tensor-parallel: 2
services:
vllm-qwen36-27b-dual-turbo:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -86,6 +90,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -60,6 +60,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f multi4/dflash.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 4
# Tensor-parallel: 4
services:
vllm-qwen36-27b-multi4-dflash:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -94,6 +98,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
@@ -57,6 +57,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f multi4/docker-compose.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 4
# Tensor-parallel: 4
services:
vllm-qwen36-27b-multi4:
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
@@ -97,6 +101,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
# PCIe-only stack — disable NCCL features that assume NVLink.
@@ -129,6 +129,10 @@
# Then send requests with structured_outputs in extra_body — see
# docs/STRUCTURED_COT.md for client examples.
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b-bounded-thinking:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -163,6 +167,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -95,6 +95,10 @@
# bash ../scripts/setup.sh # ensures Genesis tree at v7.14+ layout
# cd compose && docker compose up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -134,6 +138,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -109,6 +109,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f single/long-text.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b-long-text-no-mtp:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -151,6 +155,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -119,6 +119,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f single/long-text.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b-long-text:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -161,6 +165,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -95,6 +95,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f single/long-vision.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b-long-vision:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -129,6 +133,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -31,6 +31,10 @@
# Run:
# cd compose && docker compose -f single/minimal.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 20
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b-minimal:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -61,6 +65,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
@@ -33,6 +33,10 @@
# cd <repo>/models/qwen3.6-27b/vllm/compose
# docker compose -f single/tools-text.yml up -d
# ===========================================================================
# Hardware metadata (parsed by scripts/preflight.sh):
# Requires-min-vram-gb: 24
# Requires-min-gpu-count: 1
# Tensor-parallel: 1
services:
vllm-qwen36-27b:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
@@ -66,6 +70,7 @@ services:
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
# - CUDA_VISIBLE_DEVICES=0
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_WORKER_MULTIPROC_METHOD=spawn
+145 -6
View File
@@ -7,12 +7,14 @@
# what you want, use `scripts/switch.sh <variant>` directly.
#
# Usage:
# bash scripts/launch.sh # interactive wizard
# bash scripts/launch.sh # interactive hardware-aware wizard
# bash scripts/launch.sh --variant <name> # skip wizard, boot directly
# bash scripts/launch.sh --engine vllm --cards 1 # partial flags, ask the rest
# bash scripts/launch.sh --no-verify # skip post-launch verify-full
# bash scripts/launch.sh --no-preflight # skip docker/GPU pre-flight
#
# The wizard marks variants that don't fit the detected GPUs; direct
# --variant keeps the power-user path and delegates final gating to switch.sh.
# All flags accept the same names as `switch.sh --list` produces.
# Examples:
# bash scripts/launch.sh --variant vllm/default
@@ -87,7 +89,12 @@ choose() {
done
while true; do
local pick
read -rp "Choice [1-${#labels[@]}]: " pick
if ! read -rp "Choice [1-${#labels[@]}]: " pick; then
echo "" >&2
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
kill -INT $$
exit 1
fi
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
echo "${values[$((pick-1))]}"
return
@@ -96,6 +103,121 @@ choose() {
done
}
declare -A LAUNCH_VARIANT_COMPOSE=(
[vllm/default]="models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
[vllm/long-vision]="models/qwen3.6-27b/vllm/compose/single/long-vision.yml"
[vllm/long-text]="models/qwen3.6-27b/vllm/compose/single/long-text.yml"
[vllm/long-text-no-mtp]="models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml"
[vllm/bounded-thinking]="models/qwen3.6-27b/vllm/compose/single/bounded-thinking.yml"
[vllm/tools-text]="models/qwen3.6-27b/vllm/compose/single/tools-text.yml"
[vllm/minimal]="models/qwen3.6-27b/vllm/compose/single/minimal.yml"
[vllm/dual]="models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml"
[vllm/dual4]="models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml"
[vllm/dual4-dflash]="models/qwen3.6-27b/vllm/compose/multi4/dflash.yml"
[vllm/dual-turbo]="models/qwen3.6-27b/vllm/compose/dual/turbo.yml"
[vllm/dual-dflash]="models/qwen3.6-27b/vllm/compose/dual/dflash.yml"
[vllm/dual-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/dflash-noviz.yml"
[vllm/dual-nvlink]="models/qwen3.6-27b/vllm/compose/dual/nvlink.yml"
[vllm/dual-nvlink-turbo]="models/qwen3.6-27b/vllm/compose/dual/nvlink-turbo.yml"
[vllm/dual-nvlink-dflash]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash.yml"
[vllm/dual-nvlink-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash-noviz.yml"
[vllm/gemma-mtp]="models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
[vllm/gemma-mtp-tp1]="models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
[vllm/gemma-dflash]="models/gemma-4-31b/vllm/compose/dual/dflash.yml"
)
variant_hw_status() {
local variant="$1"
local rel="${LAUNCH_VARIANT_COMPOSE[$variant]:-}"
if [[ -z "$rel" ]]; then
printf 'ok|fits your rig'
return 0
fi
local compose_file="${ROOT_DIR}/${rel}"
if [[ ! -f "$compose_file" ]]; then
printf 'unknown|compose metadata unavailable'
return 2
fi
compose_hw_compose_status "$compose_file" 2>/dev/null || true
}
choose_variant() {
# choose_variant "prompt" "default-variant" "label1" "value1" ...
local prompt="$1" default_variant="$2"
shift 2
local i labels=() values=() statuses=() eligible=()
while [[ $# -gt 0 ]]; do
labels+=("$1")
values+=("$2")
statuses+=("$(variant_hw_status "$2")")
shift 2
done
local default_idx=""
for i in "${!values[@]}"; do
if [[ "${values[$i]}" == "$default_variant" && "${statuses[$i]}" == ok\|* ]]; then
default_idx=$((i + 1))
break
fi
done
if [[ -z "$default_idx" ]]; then
for i in "${!values[@]}"; do
if [[ "${statuses[$i]}" == ok\|* || "${statuses[$i]}" == unknown\|* ]]; then
default_idx=$((i + 1))
break
fi
done
fi
echo "" >&2
echo "$prompt" >&2
for i in "${!labels[@]}"; do
local status="${statuses[$i]}"
local state="${status%%|*}"
local reason="${status#*|}"
local marker="✓"
case "$state" in
ok) marker="✓" ;;
unknown) marker="?" ;;
*) marker="✗" ;;
esac
if [[ -n "$default_idx" && $((i + 1)) -eq "$default_idx" ]]; then
printf " %d) %s %s %s [default]\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
else
printf " %d) %s %s %s\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
fi
done
if [[ -z "$default_idx" ]]; then
echo "ERROR: no eligible variants in this menu. Use scripts/switch.sh --force <variant> to attempt anyway." >&2
exit 1
fi
while true; do
local pick
if ! read -rp "Choice [1-${#labels[@]}, default ${default_idx}]: " pick; then
echo "" >&2
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
kill -INT $$
exit 1
fi
pick="${pick:-$default_idx}"
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
local status="${statuses[$((pick - 1))]}"
if [[ "$status" == no\|* ]]; then
echo " That variant won't run on your detected rig: ${status#*|}" >&2
echo " Pick another, or use: bash scripts/switch.sh --force ${values[$((pick - 1))]}" >&2
continue
fi
echo "${values[$((pick - 1))]}"
return
fi
echo " invalid — pick a number 1-${#labels[@]}" >&2
done
}
# --- wizard ---
# Flow: cards → workload → auto-pick engine. Newcomers can answer "how
# many GPUs" and "what do I want to do" but rarely "vLLM or llama.cpp" —
@@ -151,8 +273,15 @@ if [[ -z "$VARIANT" ]]; then
"[fallback] tools-text 75K FP8 (FP8 KV alternative for accuracy compare)" "vllm/tools-text"
"[fallback] minimal 32K (no Genesis, no spec-decode — diagnostic stack)" "vllm/minimal"
)
VLLM_DUAL_OPTS=(
"[2-card] vllm/dual — 262K + vision + 2 streams" "vllm/dual"
"[2-card] vllm/dual-turbo — 4 streams @ 262K, TQ3 KV" "vllm/dual-turbo"
"[2-card] vllm/dual-dflash — peak code TPS with vision" "vllm/dual-dflash"
"[2-card] vllm/gemma-mtp — Gemma 4 dual-card default" "vllm/gemma-mtp"
)
else
VLLM_FALLBACK_OPTS=()
VLLM_DUAL_OPTS=()
fi
if [[ -z "$ENGINE" || "$ENGINE" == "llamacpp" ]]; then
LLAMA_FALLBACK_OPTS=(
@@ -161,19 +290,29 @@ if [[ -z "$VARIANT" ]]; then
else
LLAMA_FALLBACK_OPTS=()
fi
VARIANT=$(choose "What's your main workload?" \
if [[ -n "$ENGINE" && "$ENGINE" != "vllm" && "$ENGINE" != "llamacpp" ]]; then
echo "ERROR: --engine ${ENGINE} unsupported (expected vllm or llamacpp)." >&2
exit 1
fi
if [[ "$ENGINE" == "llamacpp" ]]; then
DEFAULT_VARIANT="llamacpp/default"
else
DEFAULT_VARIANT="vllm/long-text"
fi
VARIANT=$(choose_variant "What's your main workload?" "$DEFAULT_VARIANT" \
"${VLLM_OPTS[@]}" "${LLAMA_OPTS[@]}" \
"${VLLM_FALLBACK_OPTS[@]}" "${LLAMA_FALLBACK_OPTS[@]}")
"${VLLM_FALLBACK_OPTS[@]}" "${VLLM_DUAL_OPTS[@]}" "${LLAMA_FALLBACK_OPTS[@]}")
elif [[ "$CARDS" == "2" ]]; then
if [[ -n "$ENGINE" && "$ENGINE" != "vllm" ]]; then
echo "ERROR: --engine ${ENGINE} not supported on 2× cards (no llama.cpp dual recipe yet)." >&2
exit 1
fi
VARIANT=$(choose "What's your dual-card priority?" \
VARIANT=$(choose_variant "What's your dual-card priority?" "vllm/dual" \
"Balanced default — 262K + vision + 2 streams (recommended)" "vllm/dual" \
"Multi-tenant — 4 concurrent streams @ 262K, TQ3 KV" "vllm/dual-turbo" \
"Peak code TPS with vision (185K, DFlash N=5)" "vllm/dual-dflash" \
"Peak code TPS no vision (200K, DFlash N=5)" "vllm/dual-dflash-noviz")
"Peak code TPS no vision (200K, DFlash N=5)" "vllm/dual-dflash-noviz" \
"Gemma 4 dual-card default — 32K + vision + MTP" "vllm/gemma-mtp")
else
echo "ERROR: --cards ${CARDS} unsupported (expected 1 or 2)." >&2
exit 1
+416
View File
@@ -0,0 +1,416 @@
#!/usr/bin/env bash
#
# Formatter for one-row BENCHMARKS.md submissions from results/rebench/<tag>/.
#
# Public functions:
# bench_row_format <rebench-tag-dir>
# bench_row_section <rebench-tag-dir>
# bench_row_fixtures
_BENCH_ROW_LIB_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
_BENCH_ROW_ROOT="$(cd -- "${_BENCH_ROW_LIB_DIR}/../.." && pwd)"
bench_row_fixtures() {
local tag
for tag in \
qwen-int8-pth-n4-2026-05-10 \
qwen-bf16-n4-2026-05-11 \
qwen-int8-tq3-n3-2026-05-11 \
qwen-tq3-mtp-genesis-2026-05-11 \
gemma-int8-pth-n4-2026-05-11 \
gemma-bf16-n4-2026-05-11; do
if [[ -d "${_BENCH_ROW_ROOT}/results/rebench/${tag}" ]]; then
printf '%s\n' "${_BENCH_ROW_ROOT}/results/rebench/${tag}"
fi
done
}
bench_row_section() {
_bench_row_python section "$1"
}
bench_row_format() {
_bench_row_python row "$1"
}
bench_row_rig_shortname() {
_bench_row_python rig-short "$1"
}
_bench_row_python() {
local mode="$1"
local tag_dir="$2"
BENCH_ROW_REPO_ROOT="${_BENCH_ROW_ROOT}" python3 - "$mode" "$tag_dir" <<'PY'
from __future__ import annotations
import json
import os
import re
import sys
from pathlib import Path
from typing import Any
MODE = sys.argv[1]
TAG_DIR = Path(sys.argv[2]).resolve()
ROOT = Path(os.environ.get("BENCH_ROW_REPO_ROOT", ".")).resolve()
def die(msg: str) -> None:
print(f"[bench-row] ERROR: {msg}", file=sys.stderr)
raise SystemExit(1)
def read_text(path: Path) -> str:
try:
return path.read_text(errors="replace")
except Exception:
return ""
def read_json(path: Path) -> Any:
try:
return json.loads(path.read_text(errors="replace"))
except Exception:
return None
def require_file(name: str) -> Path:
path = TAG_DIR / name
if not path.is_file():
die(f"missing required artifact: {path}")
return path
def first_container(blob: Any) -> dict[str, Any]:
if isinstance(blob, list) and blob:
return blob[0] if isinstance(blob[0], dict) else {}
return blob if isinstance(blob, dict) else {}
def flag(cmd: list[str], name: str) -> str:
try:
i = cmd.index(name)
return str(cmd[i + 1])
except Exception:
return "?"
def rel(path: str) -> str:
if not path:
return ""
p = Path(path)
try:
return str(p.resolve().relative_to(ROOT))
except Exception:
return str(p)
def infer_compose_path(container_name: str, served: str, tp: str) -> str:
name = container_name.lstrip("/")
is_gemma = "gemma" in name or "gemma" in served
model_root = "models/gemma-4-31b/vllm/compose" if is_gemma else "models/qwen3.6-27b/vllm/compose"
mapping = {
"dual-int8-tq3": "dual/int8-tq3.yml",
"dual-tq3-mtp-genesis": "dual/tq3-mtp-genesis.yml",
"dual-tq3-nomtp": "dual/tq3-nomtp.yml",
"dual-tq3-mtp": "dual/tq3-mtp.yml",
"dual-int8": "dual/int8.yml",
"dual-bf16": "dual/bf16.yml",
"dual-dflash-noviz": "dual/dflash-noviz.yml",
"dual-dflash": "dual/dflash.yml",
"dual-turbo": "dual/turbo.yml",
"dual": "dual/docker-compose.yml",
"minimal": "single/minimal.yml",
"tools-text": "single/tools-text.yml",
"long-text-no-mtp": "single/long-text-no-mtp.yml",
"long-text": "single/long-text.yml",
"long-vision": "single/long-vision.yml",
}
for needle, suffix in mapping.items():
if needle in name:
return f"{model_root}/{suffix}"
if tp == "4":
return "models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml"
if tp == "2":
return f"{model_root}/dual/docker-compose.yml"
return f"{model_root}/single/docker-compose.yml"
def compose_display(compose_path: str, served: str) -> str:
path = compose_path.replace("\\", "/")
parts = path.split("/")
base = parts[-1] if parts else path
parent = parts[-2] if len(parts) >= 2 else ""
is_gemma = "gemma" in served or "gemma-4-31b" in path
if base == "docker-compose.yml":
if parent == "dual":
return "dual.yml"
if parent == "multi4":
return "dual4.yml"
if parent == "single":
return "vllm/gemma-mtp-tp1" if is_gemma else "vllm/default"
if parent in {"dual", "multi4"} and not base.startswith(f"{parent}-"):
return f"{parent}-{base}"
return base
def parse_rig(rig_txt: str) -> dict[str, Any]:
out: dict[str, Any] = {"gpus": []}
for raw in rig_txt.splitlines():
line = raw.strip()
if not line:
continue
if line.startswith("hostname:"):
out["hostname"] = line.split(":", 1)[1].strip()
elif line.startswith("GPU "):
gpu = line.split(":", 1)[1].split("(UUID", 1)[0].strip()
out["gpus"].append(gpu)
elif line.startswith("power_cap_w:"):
out["power_cap_w"] = line.split(":", 1)[1].strip()
return out
def simplify_gpu(name: str) -> str:
name = re.sub(r"^NVIDIA\s+", "", name)
name = re.sub(r"^GeForce\s+", "", name)
name = re.sub(r"^RTX\s+", "", name)
return name.strip()
def rig_shape(rig: dict[str, Any]) -> str:
gpus = [simplify_gpu(g) for g in rig.get("gpus") or []]
if not gpus:
shape = "rig"
elif len(set(gpus)) == 1:
shape = f"{len(gpus)}× {gpus[0]}"
else:
shape = " + ".join(gpus)
power = str(rig.get("power_cap_w") or "").strip()
if power:
try:
power = f"{float(power):.0f} W/card"
except Exception:
power = f"{power} W/card"
return f"{shape}, {power}"
return shape
def rig_cell(rig: dict[str, Any]) -> str:
user = os.environ.get("BENCH_ROW_GITHUB_USER", "").strip().lstrip("@") or "your-handle"
return f"@{user} ({rig_shape(rig)})"
def short_date(tag: str, report: str) -> str:
m = re.search(r"(20\d{2}-\d{2}-\d{2})", tag)
if m:
return m.group(1)
m = re.search(r"\*\*Date:\*\*\s*(20\d{2}-\d{2}-\d{2})", report)
return m.group(1) if m else "—"
def kv_display(raw: str, served: str) -> str:
raw = (raw or "?").strip("`")
lowered = raw.lower()
if lowered in {"turboquant_3bit_nc", "tq3"}:
return "TQ3"
if lowered in {"fp8_e5m2", "fp8", "fp8_e4m3"}:
return "fp8"
if lowered in {"bfloat16", "bf16"}:
return "bf16"
if lowered == "auto":
return "bf16"
if lowered == "int8_per_token_head":
return "int8_per_token_head"
return raw or "?"
def fmt_ctx(value: str) -> str:
try:
n = int(str(value).replace(",", ""))
except Exception:
return str(value or "?")
if n >= 1000:
return f"{round(n / 1000):.0f}K"
return str(n)
def fmt_tps(value: Any) -> str:
try:
return f"{float(value):.2f}"
except Exception:
return "?"
def parse_mib(text: str) -> int | None:
m = re.search(r"([\d.]+)\s*MiB", str(text))
if not m:
return None
try:
return int(float(m.group(1)))
except Exception:
return None
def peak_vram(internal: dict[str, Any], gpu_count: int) -> str:
vals: list[int] = []
for g in (((internal.get("bench") or {}).get("gpu_state")) or []):
mib = parse_mib(g.get("mem_used", ""))
if mib is not None:
vals.append(mib)
if not vals:
return "TBD"
suffix = "/card" if gpu_count > 1 else ""
return f"{max(vals) / 1024:.1f} GB{suffix}"
def parse_jsonish(value: str) -> dict[str, Any]:
try:
parsed = json.loads(value)
return parsed if isinstance(parsed, dict) else {}
except Exception:
return {}
def verify_summary(report: str) -> str:
m = re.search(r"Verify-stress:\s+\*\*(\d+/\d+)\*\*", report)
if m:
return m.group(1)
m = re.search(r"\*\*Overall:\*\*\s+(PASS|FAIL|\?)", report)
return m.group(1) if m else "?"
def soak_note(soak: dict[str, Any]) -> str:
verdict = str(soak.get("verdict") or "").upper()
silent = str(soak.get("silent_empty") or "")
growth = str(soak.get("max_growth_mib") or "")
if not verdict:
return "Soak: —"
if verdict == "PASS" and (silent.startswith("0 ") or silent.startswith("0/") or silent == "0"):
return "Soak: ✓ PASS"
if verdict == "PASS":
return f"Soak: ⚠ borderline ({silent or growth})"
return f"Soak: ✗ {verdict}"
def section_name(compose_path: str, served: str, tp: str, container: str) -> str:
if "gemma" in served or "gemma" in container or "gemma-4-31b" in compose_path:
return "Gemma 4 31B (community-experimental)"
path = compose_path.replace("\\", "/")
if "llama-cpp" in path or "llama-cpp" in container:
return "Single-card (1× RTX 3090) — llama.cpp"
if "/multi4/" in path or tp == "4":
return "Quad-card (4× RTX 3090, TP=4)"
if "/dual/" in path or tp == "2":
return "Dual-card (2× RTX 3090, TP=2)"
return "Single-card (1× RTX 3090) — vLLM"
def load() -> dict[str, Any]:
if not TAG_DIR.is_dir():
die(f"tag dir not found: {TAG_DIR}")
internal = read_json(require_file("_internal.json"))
if not isinstance(internal, dict):
die(f"invalid JSON artifact: {TAG_DIR / '_internal.json'}")
report = read_text(require_file("REPORT.md"))
config_blob = first_container(read_json(require_file("container-config.json")))
cfg = config_blob.get("Config") or {}
labels = cfg.get("Labels") or {}
cmd = cfg.get("Cmd") or []
env = cfg.get("Env") or []
name = str(config_blob.get("Name") or "").lstrip("/")
served = flag(cmd, "--served-model-name")
tp = flag(cmd, "--tensor-parallel-size")
compose = rel(labels.get("com.docker.compose.project.config_files") or "")
if not compose:
compose = infer_compose_path(name, served, tp)
rig = parse_rig(read_text(require_file("rig.txt")))
tag = TAG_DIR.name
spec = parse_jsonish(flag(cmd, "--speculative-config"))
return {
"tag": tag,
"report": report,
"internal": internal,
"container": {
"name": name,
"served": served,
"tp": tp,
"compose": compose,
"kv": flag(cmd, "--kv-cache-dtype"),
"max_ctx": flag(cmd, "--max-model-len"),
"max_num_seqs": flag(cmd, "--max-num-seqs"),
"mem_util": flag(cmd, "--gpu-memory-utilization"),
"image": cfg.get("Image", ""),
"spec": spec,
"genesis": any(str(e).startswith("GENESIS_") for e in env),
},
"rig": rig,
"date": short_date(tag, report),
}
def format_row(data: dict[str, Any]) -> str:
c = data["container"]
internal = data["internal"]
bench = internal.get("bench") or {}
narrative = bench.get("narrative") or {}
code = bench.get("code") or {}
mtp = bench.get("mtp") or {}
quality = internal.get("quality") or {}
aider = internal.get("aider") or {}
soak = internal.get("soak") or {}
rig = data["rig"]
gpu_count = max(len(rig.get("gpus") or []), 1)
section = section_name(c["compose"], c["served"], c["tp"], c["name"])
verify = verify_summary(data["report"])
compose = compose_display(c["compose"], c["served"])
kv = kv_display(c["kv"], c["served"])
max_ctx = fmt_ctx(c["max_ctx"])
tps = f"**{fmt_tps(narrative.get('wall_tps_mean'))} / {fmt_tps(code.get('wall_tps_mean'))}**"
peak = peak_vram(internal, gpu_count)
spec_n = c["spec"].get("num_speculative_tokens")
notes = [soak_note(soak), f"verify-stress {verify}"]
if quality.get("total_passed") is not None:
notes.append(f"quality {quality.get('total_passed')}/{quality.get('total_total')}")
if aider.get("total_count"):
notes.append(f"aider {aider.get('passed_count')}/{aider.get('total_count')}")
if mtp.get("mean_accept_length") is not None:
n_text = f" n={spec_n}" if spec_n is not None else ""
notes.append(
f"MTP{n_text} AL {float(mtp['mean_accept_length']):.2f}, "
f"accept {float(mtp.get('avg_accept_rate', 0)):.1f}%"
)
if c.get("genesis"):
notes.append("Genesis on")
note_cell = "; ".join(notes) + f". Report: `results/rebench/{data['tag']}/REPORT.md`."
if section == "Gemma 4 31B (community-experimental)":
al = f"{float(mtp['mean_accept_length']):.2f}" if mtp.get("mean_accept_length") is not None else "—"
per_pos = str(mtp.get("per_position") or "—")
return (
f"| `{compose}` | {rig_cell(rig)} | {kv} | {max_ctx} | {tps} | "
f"{al} | {per_pos} | {peak} | {data['date']} | {note_cell} |"
)
return (
f"| `{compose}` | {rig_cell(rig)} | {kv} | {max_ctx} | {tps} | "
f"{peak} | {data['date']} | {note_cell} |"
)
data = load()
if MODE == "section":
c = data["container"]
print(section_name(c["compose"], c["served"], c["tp"], c["name"]))
elif MODE == "row":
print(format_row(data))
elif MODE == "rig-short":
print(data["tag"])
else:
die(f"unknown mode: {MODE}")
PY
}
+290
View File
@@ -0,0 +1,290 @@
#!/usr/bin/env bash
#
# Tiny parser for hardware metadata stored as compose header comments.
#
# Expected form:
# # Requires-min-vram-gb: 24
# # Requires-min-gpu-count: 2
# # Tensor-parallel: 2
# # Requires-sm: 9.0+
#
# This intentionally does not parse YAML. These fields are comments so that
# older docker compose versions and direct `docker compose -f ... up` flows keep
# working unchanged.
_compose_meta_trim() {
local value="$1"
value="${value#"${value%%[![:space:]]*}"}"
value="${value%"${value##*[![:space:]]}"}"
printf '%s' "$value"
}
_compose_meta_norm_key() {
local key="$1"
key="$(_compose_meta_trim "$key")"
key="${key//_/-}"
key="${key// /-}"
printf '%s' "$key" | tr '[:upper:]' '[:lower:]'
}
_compose_meta_wants_key() {
local requested="$(_compose_meta_norm_key "$1")"
local candidate="$(_compose_meta_norm_key "$2")"
case "$requested" in
min-vram-gb) requested="requires-min-vram-gb" ;;
min-gpu-count) requested="requires-min-gpu-count" ;;
tp) requested="tensor-parallel" ;;
sm) requested="requires-sm" ;;
esac
[[ "$candidate" == "$requested" ]]
}
compose_meta_get() {
local compose_file="$1"
local field="$2"
[[ -f "$compose_file" ]] || return 1
local line key value
while IFS= read -r line; do
[[ "$line" =~ ^[[:space:]]*# ]] || continue
line="${line#*\#}"
[[ "$line" == *:* ]] || continue
key="${line%%:*}"
value="${line#*:}"
if _compose_meta_wants_key "$field" "$key"; then
_compose_meta_trim "$value"
return 0
fi
done < "$compose_file"
return 1
}
compose_hw_sm_to_int() {
local sm="$1"
sm="${sm%%+}"
sm="${sm//sm_/}"
sm="${sm//SM_/}"
sm="${sm// /}"
[[ -z "$sm" ]] && { echo 0; return; }
local major minor
if [[ "$sm" == *.* ]]; then
major="${sm%%.*}"
minor="${sm#*.}"
else
major="$sm"
minor="0"
fi
major="${major//[^0-9]/}"
minor="${minor//[^0-9]/}"
[[ -z "$major" ]] && major=0
[[ -z "$minor" ]] && minor=0
if [[ "${#minor}" -eq 1 ]]; then
minor=$(( minor * 10 ))
else
minor="${minor:0:2}"
[[ -z "$minor" ]] && minor=0
fi
echo $(( major * 100 + minor ))
}
compose_hw_vram_gb() {
local mib="$1"
echo $(( (mib + 1023) / 1024 ))
}
compose_hw_detect_gpus() {
if [[ "${_COMPOSE_HW_GPU_CACHE_SET:-0}" == "1" ]]; then
[[ -n "${_COMPOSE_HW_GPU_CACHE:-}" ]] || return 1
printf '%s\n' "${_COMPOSE_HW_GPU_CACHE}"
return 0
fi
command -v nvidia-smi >/dev/null 2>&1 || return 1
local query idx name mem_mib sm rest
query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null)" || return 1
[[ -n "$query" ]] || return 1
local parsed=""
while IFS=',' read -r idx name mem_mib sm rest; do
idx="$(_compose_meta_trim "$idx")"
name="$(_compose_meta_trim "$name")"
mem_mib="$(_compose_meta_trim "$mem_mib")"
sm="$(_compose_meta_trim "$sm")"
[[ -z "$idx" || -z "$mem_mib" ]] && continue
parsed+="${idx}"$'\t'"${name}"$'\t'"${mem_mib}"$'\t'"${sm}"$'\n'
done <<< "$query"
parsed="${parsed%$'\n'}"
_COMPOSE_HW_GPU_CACHE_SET=1
_COMPOSE_HW_GPU_CACHE="$parsed"
[[ -n "$parsed" ]] || return 1
printf '%s\n' "$parsed"
}
compose_hw_summary() {
local gpu_lines
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
if [[ -z "$gpu_lines" ]]; then
printf 'no NVIDIA GPUs detected'
return 0
fi
local count=0 first_name="" first_gb="" mixed=0 idx name mem_mib sm
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
local gb
gb="$(compose_hw_vram_gb "$mem_mib")"
name="${name#NVIDIA }"
name="${name#GeForce }"
count=$((count + 1))
if [[ -z "$first_name" ]]; then
first_name="$name"
first_gb="$gb"
elif [[ "$name" != "$first_name" || "$gb" != "$first_gb" ]]; then
mixed=1
fi
done <<< "$gpu_lines"
if (( count == 0 )); then
printf 'no NVIDIA GPUs detected'
elif (( mixed == 0 )); then
if (( count == 1 )); then
printf '1× %s, %s GB' "$first_name" "$first_gb"
else
printf '%d× %s, %s GB each' "$count" "$first_name" "$first_gb"
fi
else
local parts=()
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
name="${name#NVIDIA }"
name="${name#GeForce }"
parts+=("${name}, $(compose_hw_vram_gb "$mem_mib") GB")
done <<< "$gpu_lines"
local joined=""
for part in "${parts[@]}"; do
if [[ -z "$joined" ]]; then
joined="$part"
else
joined="${joined} + ${part}"
fi
done
printf '%s' "$joined"
fi
}
compose_hw_requirement_text() {
local min_vram_gb="$1"
local min_gpu_count="$2"
local requires_sm="${3:-}"
local req
if [[ "$min_gpu_count" == "1" ]]; then
req="${min_vram_gb} GB+"
else
req="${min_gpu_count}× ${min_vram_gb} GB"
fi
if [[ -n "$requires_sm" && "$requires_sm" != "0.0" ]]; then
req="${req}, sm_${requires_sm%%+}+"
fi
printf '%s' "$req"
}
compose_hw_compose_status() {
local compose_file="$1"
local min_vram_gb min_gpu_count requires_sm
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" ]]; then
printf 'unknown|metadata unavailable'
return 2
fi
requires_sm="${requires_sm:-0.0}"
local required_sm_int
required_sm_int="$(compose_hw_sm_to_int "$requires_sm")"
local gpu_lines
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
if [[ -z "$gpu_lines" ]]; then
printf 'no|no NVIDIA GPUs detected'
return 1
fi
local eligible_count=0 idx name mem_mib sm gb sm_int
while IFS=$'\t' read -r idx name mem_mib sm; do
[[ -z "$idx" ]] && continue
gb="$(compose_hw_vram_gb "$mem_mib")"
sm_int="$(compose_hw_sm_to_int "$sm")"
if (( gb >= min_vram_gb && sm_int >= required_sm_int )); then
eligible_count=$((eligible_count + 1))
fi
done <<< "$gpu_lines"
if (( eligible_count >= min_gpu_count )); then
printf 'ok|fits your rig'
return 0
fi
printf 'no|needs %s (your rig: %s)' \
"$(compose_hw_requirement_text "$min_vram_gb" "$min_gpu_count" "$requires_sm")" \
"$(compose_hw_summary)"
return 1
}
compose_hw_compose_eligible() {
local status
status="$(compose_hw_compose_status "$1" 2>/dev/null || true)"
[[ "$status" == ok\|* ]]
}
compose_hw_model_status() {
local repo_root="$1"
local model="$2"
local candidates=()
local friendly_need=""
case "$model" in
qwen3.6-27b)
candidates=(
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/long-text.yml"
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
)
friendly_need="needs 20 GB+ VRAM (24 GB recommended)"
;;
gemma-4-31b)
candidates=(
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/int8.yml"
"${repo_root}/models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
)
friendly_need="needs 32 GB+ on single card OR 2× 24 GB"
;;
*)
printf 'no|unknown model: %s' "$model"
return 1
;;
esac
local file status
for file in "${candidates[@]}"; do
[[ -f "$file" ]] || continue
status="$(compose_hw_compose_status "$file" 2>/dev/null || true)"
if [[ "$status" == ok\|* ]]; then
printf 'ok|fits your rig'
return 0
fi
done
printf 'no|%s (your rig: %s)' "$friendly_need" "$(compose_hw_summary)"
return 1
}
+315 -3
View File
@@ -12,6 +12,7 @@
# preflight_running — warn if a club-3090 container is already up
# preflight_genesis_pin — warn if on-disk Genesis tree differs from setup.sh's pin
# preflight_repo_drift — warn if local HEAD is behind origin/master
# preflight_compose_hardware— check compose VRAM/GPU-count/SM metadata
#
# Style: each function prints one or more "[preflight] ..." lines.
# Hard failures get a one-line ERROR + a "Fix:" hint.
@@ -19,6 +20,12 @@
# Avoid double-sourcing.
[[ -n "${_PREFLIGHT_LOADED:-}" ]] && return 0
_PREFLIGHT_LOADED=1
_PREFLIGHT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
if [[ -f "${_PREFLIGHT_DIR}/lib/compose-meta.sh" ]]; then
# shellcheck source=lib/compose-meta.sh
source "${_PREFLIGHT_DIR}/lib/compose-meta.sh"
fi
preflight_docker() {
if ! command -v docker >/dev/null 2>&1; then
@@ -72,6 +79,301 @@ preflight_gpu() {
return 0
}
_preflight_trim() {
local value="$1"
value="${value#"${value%%[![:space:]]*}"}"
value="${value%"${value##*[![:space:]]}"}"
printf '%s' "$value"
}
_preflight_csv_token() {
local value="$1"
value="$(_preflight_trim "$value")"
printf '%s' "$value"
}
_preflight_selector() {
if [[ -n "${CLUB3090_GPU:-}" ]]; then
printf '%s' "${CLUB3090_GPU}"
elif [[ -n "${NVIDIA_VISIBLE_DEVICES:-}" && "${NVIDIA_VISIBLE_DEVICES}" != "all" && "${NVIDIA_VISIBLE_DEVICES}" != "void" ]]; then
printf '%s' "${NVIDIA_VISIBLE_DEVICES}"
elif [[ -n "${CUDA_VISIBLE_DEVICES:-}" && "${CUDA_VISIBLE_DEVICES}" != "all" && "${CUDA_VISIBLE_DEVICES}" != "void" ]]; then
printf '%s' "${CUDA_VISIBLE_DEVICES}"
fi
}
_preflight_selector_is_specific() {
local selector="${1:-}"
[[ -n "$selector" && "$selector" != "all" && "$selector" != "void" ]]
}
_preflight_selector_allows_index() {
local selector="$1"
local idx="$2"
local token
if ! _preflight_selector_is_specific "$selector"; then
return 0
fi
IFS=',' read -ra _preflight_selector_tokens <<< "$selector"
for token in "${_preflight_selector_tokens[@]}"; do
token="$(_preflight_trim "$token")"
[[ "$token" == "$idx" ]] && return 0
done
return 1
}
_preflight_selector_first_numeric() {
local selector="$1"
local token
IFS=',' read -ra _preflight_selector_tokens <<< "$selector"
for token in "${_preflight_selector_tokens[@]}"; do
token="$(_preflight_trim "$token")"
if [[ "$token" =~ ^[0-9]+$ ]]; then
printf '%s' "$token"
return 0
fi
done
return 1
}
_preflight_sm_to_int() {
local sm="$1"
sm="${sm%%+}"
sm="${sm//sm_/}"
sm="${sm//SM_/}"
sm="${sm// /}"
[[ -z "$sm" ]] && { echo 0; return; }
local major minor
if [[ "$sm" == *.* ]]; then
major="${sm%%.*}"
minor="${sm#*.}"
else
major="$sm"
minor="0"
fi
major="${major//[^0-9]/}"
minor="${minor//[^0-9]/}"
[[ -z "$major" ]] && major=0
[[ -z "$minor" ]] && minor=0
if [[ "${#minor}" -eq 1 ]]; then
minor=$(( minor * 10 ))
else
minor="${minor:0:2}"
[[ -z "$minor" ]] && minor=0
fi
echo $(( major * 100 + minor ))
}
_preflight_vram_gb() {
local mib="$1"
echo $(( (mib + 1023) / 1024 ))
}
_preflight_hardware_suggestions() {
local variant="${1:-}"
echo "[preflight]" >&2
echo "[preflight] Suggested next steps:" >&2
echo "[preflight] - Pick a compose that matches the detected GPU VRAM/topology." >&2
if [[ "$variant" == vllm/gemma-mtp-tp1 ]]; then
echo "[preflight] - On 2x 24 GB cards, use: bash scripts/switch.sh vllm/gemma-mtp" >&2
fi
echo "[preflight] - On a single 24 GB card, start with: bash scripts/switch.sh vllm/default" >&2
echo "[preflight] - For maximum compatibility, use: bash scripts/switch.sh llamacpp/default" >&2
echo "[preflight] - Explicit bypass: bash scripts/switch.sh --force ${variant:-<variant>}" >&2
}
# preflight_compose_hardware <compose_file> [variant] [force]
#
# Reads compose header metadata and checks the target host before docker compose
# starts. This is intentionally conservative:
# - Missing metadata warns and allows the boot.
# - TP=1 composes auto-select the largest eligible GPU unless the user set
# CLUB3090_GPU, CUDA_VISIBLE_DEVICES, or NVIDIA_VISIBLE_DEVICES.
# - TP>=2 composes hard-fail only on insufficient GPU count or hard SM gates;
# heterogeneous VRAM below the requested floor warns because advanced users
# may be validating sub-24 GB configs with tuned memory-utilization.
preflight_compose_hardware() {
local compose_file="$1"
local variant="${2:-}"
local force="${3:-${FORCE:-0}}"
if [[ "${PREFLIGHT_NO_HARDWARE:-0}" == "1" ]]; then
return 0
fi
if [[ "$force" == "1" || "${FORCE:-0}" == "1" ]]; then
echo "[preflight] hardware: skipped (--force/FORCE=1)"
return 0
fi
if [[ ! -f "$compose_file" ]]; then
echo "[preflight] ERROR: compose file not found: $compose_file" >&2
return 1
fi
if ! command -v nvidia-smi >/dev/null 2>&1; then
echo "[preflight] WARN: nvidia-smi not found; skipping compose hardware metadata check." >&2
return 0
fi
if ! declare -F compose_meta_get >/dev/null 2>&1; then
echo "[preflight] WARN: compose metadata parser unavailable; skipping hardware metadata check." >&2
return 0
fi
local min_vram_gb min_gpu_count tp requires_sm
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
tp="$(compose_meta_get "$compose_file" tensor-parallel || true)"
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" || -z "$tp" ]]; then
echo "[preflight] WARN: compose has no hardware metadata; allowing boot: $compose_file" >&2
return 0
fi
requires_sm="${requires_sm:-0.0}"
local required_sm_int
required_sm_int="$(_preflight_sm_to_int "$requires_sm")"
local gpu_query
gpu_query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null || true)"
if [[ -z "$gpu_query" ]]; then
echo "[preflight] WARN: could not query GPU VRAM/SM via nvidia-smi; skipping hardware metadata check." >&2
return 0
fi
local selector
selector="$(_preflight_selector || true)"
local total_count=0 selected_count=0 eligible_count=0 selected_below_vram=0 selected_below_sm=0
local best_idx="" best_name="" best_mib=0 best_sm=""
local first_idx="" first_name="" first_mib=0 first_sm=""
local idx name mem_mib sm rest vram_gb sm_int
while IFS=',' read -r idx name mem_mib sm rest; do
idx="$(_preflight_csv_token "$idx")"
name="$(_preflight_csv_token "$name")"
mem_mib="$(_preflight_csv_token "$mem_mib")"
sm="$(_preflight_csv_token "$sm")"
[[ -z "$idx" || -z "$mem_mib" ]] && continue
total_count=$(( total_count + 1 ))
_preflight_selector_allows_index "$selector" "$idx" || continue
selected_count=$(( selected_count + 1 ))
if [[ -z "$first_idx" ]]; then
first_idx="$idx"
first_name="$name"
first_mib="$mem_mib"
first_sm="$sm"
fi
vram_gb="$(_preflight_vram_gb "$mem_mib")"
sm_int="$(_preflight_sm_to_int "$sm")"
if (( vram_gb < min_vram_gb )); then
selected_below_vram=1
fi
if (( sm_int < required_sm_int )); then
selected_below_sm=1
fi
if (( vram_gb >= min_vram_gb && sm_int >= required_sm_int )); then
eligible_count=$(( eligible_count + 1 ))
if (( mem_mib > best_mib )); then
best_idx="$idx"
best_name="$name"
best_mib="$mem_mib"
best_sm="$sm"
fi
fi
done <<< "$gpu_query"
if (( total_count == 0 )); then
echo "[preflight] ERROR: no NVIDIA GPUs detected." >&2
_preflight_hardware_suggestions "$variant"
return 1
fi
if (( selected_count == 0 )); then
echo "[preflight] ERROR: GPU selector '${selector}' did not match any detected GPU index." >&2
_preflight_hardware_suggestions "$variant"
return 1
fi
local requires_sm_display="${requires_sm%%+}"
local sm_label=""
if (( required_sm_int > 0 )); then
sm_label=", sm_${requires_sm_display}+"
fi
if (( tp <= 1 )); then
if _preflight_selector_is_specific "$selector"; then
local first_vram_gb first_sm_int
first_vram_gb="$(_preflight_vram_gb "$first_mib")"
first_sm_int="$(_preflight_sm_to_int "$first_sm")"
if (( first_vram_gb < min_vram_gb || first_sm_int < required_sm_int )); then
echo "[preflight] ERROR: ${variant:-compose} requires one GPU with >=${min_vram_gb} GB VRAM${sm_label}." >&2
echo "[preflight] Explicit selector '${selector}' starts with GPU ${first_idx}: ${first_name}, ${first_vram_gb} GB, sm_${first_sm}." >&2
_preflight_hardware_suggestions "$variant"
return 1
fi
export CLUB3090_GPU="${CLUB3090_GPU:-$selector}"
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-$selector}"
export NVIDIA_VISIBLE_DEVICES="${NVIDIA_VISIBLE_DEVICES:-$selector}"
echo "[preflight] hardware: ${variant:-compose} TP=1 requires >=${min_vram_gb} GB${sm_label}; using explicit GPU ${first_idx} (${first_vram_gb} GB, sm_${first_sm})"
return 0
fi
if (( eligible_count == 0 )); then
echo "[preflight] ERROR: ${variant:-compose} requires one GPU with >=${min_vram_gb} GB VRAM${sm_label}; none found." >&2
echo "[preflight] Detected GPUs:" >&2
while IFS=',' read -r idx name mem_mib sm rest; do
idx="$(_preflight_csv_token "$idx")"
name="$(_preflight_csv_token "$name")"
mem_mib="$(_preflight_csv_token "$mem_mib")"
sm="$(_preflight_csv_token "$sm")"
[[ -z "$idx" || -z "$mem_mib" ]] && continue
echo "[preflight] GPU ${idx}: ${name}, $(_preflight_vram_gb "$mem_mib") GB, sm_${sm}" >&2
done <<< "$gpu_query"
_preflight_hardware_suggestions "$variant"
return 1
fi
export CLUB3090_GPU="$best_idx"
export CUDA_VISIBLE_DEVICES="$best_idx"
export NVIDIA_VISIBLE_DEVICES="$best_idx"
echo "[preflight] hardware: ${variant:-compose} TP=1 requires >=${min_vram_gb} GB${sm_label}; auto-selected GPU ${best_idx} ($(_preflight_vram_gb "$best_mib") GB, sm_${best_sm})"
return 0
fi
if (( selected_count < min_gpu_count )); then
echo "[preflight] ERROR: ${variant:-compose} requires ${min_gpu_count} visible GPU(s) for TP=${tp}; found ${selected_count}." >&2
_preflight_hardware_suggestions "$variant"
return 1
fi
if (( selected_below_sm == 1 )); then
echo "[preflight] ERROR: ${variant:-compose} requires sm_${requires_sm_display}+ on visible GPUs." >&2
while IFS=',' read -r idx name mem_mib sm rest; do
idx="$(_preflight_csv_token "$idx")"
name="$(_preflight_csv_token "$name")"
mem_mib="$(_preflight_csv_token "$mem_mib")"
sm="$(_preflight_csv_token "$sm")"
_preflight_selector_allows_index "$selector" "$idx" || continue
echo "[preflight] GPU ${idx}: ${name}, $(_preflight_vram_gb "$mem_mib") GB, sm_${sm}" >&2
done <<< "$gpu_query"
_preflight_hardware_suggestions "$variant"
return 1
fi
if (( selected_below_vram == 1 )); then
echo "[preflight] WARN: ${variant:-compose} requires >=${min_vram_gb} GB per visible GPU for TP=${tp}, but at least one selected GPU is smaller." >&2
echo "[preflight] Continuing because TP>=2 sub-24 GB rigs may use tuned gpu-memory-utilization/KV settings." >&2
fi
echo "[preflight] hardware: ${variant:-compose} TP=${tp} requires ${min_gpu_count} GPU(s), >=${min_vram_gb} GB each${sm_label}; ${selected_count} visible GPU(s) detected"
return 0
}
preflight_disk() {
local path="$1"
local need_gb="$2"
@@ -412,9 +714,19 @@ preflight_kv_format_hint() {
return 0
fi
# Detect smallest VRAM among visible cards (the TP-split ceiling).
local min_vram_mib
min_vram_mib="$(nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | sort -n | head -1)"
# Detect smallest VRAM among selected/visible cards (the TP-split ceiling).
local min_vram_mib="" mem_query selector idx mem_mib
selector="$(_preflight_selector || true)"
mem_query="$(nvidia-smi --query-gpu=index,memory.total --format=csv,noheader,nounits 2>/dev/null || true)"
while IFS=',' read -r idx mem_mib; do
idx="$(_preflight_csv_token "$idx")"
mem_mib="$(_preflight_csv_token "$mem_mib")"
[[ -z "$idx" || -z "$mem_mib" ]] && continue
_preflight_selector_allows_index "$selector" "$idx" || continue
if [[ -z "$min_vram_mib" || "$mem_mib" -lt "$min_vram_mib" ]]; then
min_vram_mib="$mem_mib"
fi
done <<< "$mem_query"
if [[ -z "$min_vram_mib" ]] || [[ "$min_vram_mib" -ge 24000 ]]; then
return 0 # 24 GB+ cards — TQ3 is the right pick, no hint needed
fi
+3
View File
@@ -273,3 +273,6 @@ echo " verify-stress: tail -5 $OUT_DIR/verify-stress.log"
echo " quality: grep '^Quality:' $OUT_DIR/quality-full.log"
echo " soak: grep -E 'verdict|silent_empty|p50_decode' $OUT_DIR/soak.log"
echo " aider: grep 'aider-polyglot-30' $OUT_DIR/aider-polyglot.log"
echo
echo "To submit your numbers (review then PR):"
echo " bash scripts/submit-bench.sh --tag $TAG"
+96 -7
View File
@@ -2,7 +2,8 @@
#
# Model-aware one-shot setup for club-3090.
#
# bash scripts/setup.sh <model-name>
# bash scripts/setup.sh # interactive model picker in a TTY
# bash scripts/setup.sh <model-name> # scripted/CI positional form
#
# Currently supported:
# qwen3.6-27b → Lorbus/Qwen3.6-27B-int4-AutoRound + Genesis patches
@@ -37,15 +38,89 @@
set -euo pipefail
# ---------- Model dispatch ----------
MODEL_NAME="${1:-}"
if [[ -z "${MODEL_NAME}" ]]; then
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
usage() {
echo "Usage: $0 <model-name>"
echo " $0 # interactive model picker in a TTY"
echo ""
echo "Run with no model name in a normal terminal to open the hardware-aware"
echo "model picker. Use the positional form in scripts/CI to skip prompts."
echo ""
echo "Supported model names:"
echo " qwen3.6-27b"
echo " gemma-4-31b"
exit 1
}
model_label() {
case "$1" in
qwen3.6-27b) echo "Qwen 3.6 27B" ;;
gemma-4-31b) echo "Gemma 4 31B" ;;
*) echo "$1" ;;
esac
}
model_picker_line() {
local idx="$1" model="$2" size="$3" status mark reason
status="$(compose_hw_model_status "$ROOT_DIR" "$model" 2>/dev/null || true)"
reason="${status#*|}"
if [[ "$status" == ok\|* ]]; then
mark="✓"
else
mark="✗"
fi
printf " %s. %-14s (%s) %s %s\n" "$idx" "$(model_label "$model")" "$size" "$mark" "$reason"
}
pick_model_interactive() {
# shellcheck source=lib/compose-meta.sh
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
echo "[setup] Which model to download?" >&2
echo "" >&2
model_picker_line "1" "qwen3.6-27b" "~14 GB AutoRound INT4" >&2
model_picker_line "2" "gemma-4-31b" "~21 GB AutoRound INT4 + drafter" >&2
echo " 3. Both (~30 GB total) downloads both model families" >&2
echo "" >&2
while true; do
local pick
read -rp "Choice [1-3]: " pick
case "$pick" in
1) echo "qwen3.6-27b"; return ;;
2) echo "gemma-4-31b"; return ;;
3) echo "both"; return ;;
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
esac
done
}
# ---------- Model dispatch ----------
case "${1:-}" in
-h|--help)
usage
exit 0
;;
esac
MODEL_NAME="${1:-}"
if [[ -z "${MODEL_NAME}" ]]; then
if [[ -t 0 && -t 1 ]]; then
MODEL_NAME="$(pick_model_interactive)"
else
usage
echo ""
echo "(Interactive picker available in a TTY shell. Use the positional form in scripts/CI.)"
exit 1
fi
fi
if [[ "${MODEL_NAME}" == "both" ]]; then
# Resolve MODEL_DIR once in the parent by reusing the normal prompt below,
# then recurse through the positional form for each model.
SETUP_BOTH_MODE=1
MODEL_NAME="qwen3.6-27b"
else
SETUP_BOTH_MODE=0
fi
# ALWAYS_DRAFT_REPO + ALWAYS_DRAFT_SUBDIR: a drafter that this model REQUIRES
@@ -89,8 +164,6 @@ case "${MODEL_NAME}" in
;;
esac
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
# ---------- MODEL_DIR resolution ----------
# Order of precedence:
# 1. MODEL_DIR already exported in the calling shell → use as-is
@@ -158,6 +231,18 @@ fi
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
if [[ "${SETUP_BOTH_MODE:-0}" == "1" ]]; then
export MODEL_DIR
echo "[setup] downloading both supported models into ${MODEL_DIR}"
echo ""
bash "$0" qwen3.6-27b
echo ""
bash "$0" gemma-4-31b
echo ""
echo "[setup] ✓ Both models downloaded."
echo "[setup] Next: bash scripts/launch.sh"
exit 0
fi
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
cd "${ROOT_DIR}"
@@ -445,6 +530,7 @@ echo ""
# refactored 2026-05-03 to vendor the two files in-repo, fixing #37.)
# Per-model "next steps" — different composes / served-model-name / port between models.
SETUP_MODEL_DISPLAY="$(model_label "${MODEL_NAME}")"
case "${MODEL_NAME}" in
qwen3.6-27b)
SAMPLE_CONTAINER="vllm-qwen36-27b"
@@ -468,6 +554,9 @@ case "${MODEL_NAME}" in
;;
esac
echo "[setup] ✓ ${SETUP_MODEL_DISPLAY} downloaded."
echo "[setup] Next: bash scripts/launch.sh"
echo ""
echo "Next — single-card vLLM (default):"
if [[ "${MODEL_NAME}" == "gemma-4-31b" ]]; then
echo " bash scripts/switch.sh vllm/gemma-mtp"
+312
View File
@@ -0,0 +1,312 @@
#!/usr/bin/env bash
#
# Generate or submit a BENCHMARKS.md row from results/rebench/<tag>/.
#
# Usage:
# bash scripts/submit-bench.sh --tag <tag>
# bash scripts/submit-bench.sh --tag <tag> --auto-submit
# bash scripts/submit-bench.sh --tag <tag> --auto-submit --as-pr
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$ROOT_DIR"
TAG=""
AUTO_SUBMIT=0
AS_PR=0
SECTION_OVERRIDE=""
usage() {
sed -n '2,18p' "$0" | sed 's/^# \{0,1\}//'
exit 0
}
die() {
echo "[submit-bench] ERROR: $*" >&2
exit 1
}
log() {
echo "[submit-bench] $*"
}
while [[ $# -gt 0 ]]; do
case "$1" in
--tag)
[[ $# -ge 2 ]] || die "--tag requires a value"
TAG="$2"
shift 2
;;
--auto-submit)
AUTO_SUBMIT=1
shift
;;
--as-pr)
AS_PR=1
shift
;;
--section)
[[ $# -ge 2 ]] || die "--section requires a value"
SECTION_OVERRIDE="$2"
shift 2
;;
-h|--help)
usage
;;
*)
die "unknown arg: $1"
;;
esac
done
[[ -n "$TAG" ]] || die "--tag <tag> is required"
TAG_DIR="results/rebench/${TAG}"
[[ -d "$TAG_DIR" ]] || die "tag dir not found: ${TAG_DIR}"
[[ -f "$TAG_DIR/REPORT.md" ]] || die "missing required artifact: ${TAG_DIR}/REPORT.md"
[[ -f "$TAG_DIR/_internal.json" ]] || die "missing required artifact: ${TAG_DIR}/_internal.json"
[[ -f "$TAG_DIR/container-config.json" ]] || die "missing required artifact: ${TAG_DIR}/container-config.json"
[[ -f "$TAG_DIR/rig.txt" ]] || die "missing required artifact: ${TAG_DIR}/rig.txt"
# shellcheck source=lib/bench-row-formatter.sh
source "$ROOT_DIR/scripts/lib/bench-row-formatter.sh"
github_user_for_row() {
if [[ -n "${BENCH_ROW_GITHUB_USER:-}" ]]; then
printf '%s' "${BENCH_ROW_GITHUB_USER#@}"
return 0
fi
if [[ "${GH_MOCK:-0}" == "1" ]]; then
printf '%s' "${GH_MOCK_USER:-mock-user}"
return 0
fi
if command -v gh >/dev/null 2>&1 && gh auth status >/dev/null 2>&1; then
gh api user --jq .login 2>/dev/null || true
fi
}
if [[ "$AUTO_SUBMIT" -eq 1 ]]; then
GH_USER="$(github_user_for_row)"
[[ -n "$GH_USER" ]] || die "not authed with gh. Run: gh auth login"
export BENCH_ROW_GITHUB_USER="$GH_USER"
fi
ROW="$(bench_row_format "$TAG_DIR")"
SECTION="${SECTION_OVERRIDE:-$(bench_row_section "$TAG_DIR")}"
OUTPUT="$TAG_DIR/BENCHMARKS-row.md"
printf '%s\n' "$ROW" > "$OUTPUT"
log "Generated BENCHMARKS row for section: ${SECTION}"
log "Wrote: ${OUTPUT}"
echo
valid_sections() {
rg -n '^(##|###) ' BENCHMARKS.md | sed 's/^/[submit-bench] /' >&2 || true
}
write_pr_body() {
local body_file="$1"
local row="$2"
local tag="$3"
local template=".github/PULL_REQUEST_TEMPLATE/bench-row.md"
if [[ -f "$template" ]]; then
python3 - "$template" "$body_file" "$tag" "$row" <<'PY'
from pathlib import Path
import sys
template, body_file, tag, row = sys.argv[1:5]
text = Path(template).read_text()
text = text.replace("<TAG>", tag)
text = text.replace("<!-- The generated BENCHMARKS.md row goes here -->", row)
text = text.replace(
"<!-- Output of `bash scripts/report.sh` (redacted) -->",
f"See `results/rebench/{tag}/rig.txt`.",
)
Path(body_file).write_text(text)
PY
else
{
echo "## Rig bench submission"
echo
echo "### New row"
echo
echo "$row"
echo
echo "### Full results"
echo
echo "See \`results/rebench/${tag}/REPORT.md\`."
} > "$body_file"
fi
}
write_issue_body() {
local body_file="$1"
local row="$2"
local tag="$3"
local section="$4"
# The repo's numbers-from-your-rig issue template is a structured YAML form
# with required textarea/dropdown fields. `gh issue create --template` opens
# that interactive form shape, which is not useful once submit-bench has
# already generated the structured report. Use a direct markdown body instead.
{
echo "**Compose / section**: \`${section}\`"
echo
echo "**Rig**:"
echo
echo '```text'
cat "${TAG_DIR}/rig.txt"
echo '```'
echo
echo "**Proposed BENCHMARKS.md row**:"
echo
echo "$row"
echo
echo "**Full report**: \`results/rebench/${tag}/REPORT.md\`"
echo
echo "**Generated row file**: \`results/rebench/${tag}/BENCHMARKS-row.md\`"
} > "$body_file"
}
insert_row() {
local section="$1"
local row="$2"
python3 - "$section" "$row" <<'PY'
from __future__ import annotations
import sys
from pathlib import Path
section = sys.argv[1]
row = sys.argv[2]
path = Path("BENCHMARKS.md")
lines = path.read_text().splitlines()
heading_idx = None
for i, line in enumerate(lines):
if line.strip() in {f"## {section}", f"### {section}"}:
heading_idx = i
break
if heading_idx is None:
print(f"[submit-bench] ERROR: section not found in BENCHMARKS.md: {section}", file=sys.stderr)
raise SystemExit(1)
table_start = None
for i in range(heading_idx + 1, len(lines)):
if lines[i].startswith("#"):
break
if lines[i].startswith("|"):
table_start = i
break
if table_start is None:
print(f"[submit-bench] ERROR: no markdown table found below section: {section}", file=sys.stderr)
raise SystemExit(1)
insert_at = table_start
for i in range(table_start, len(lines)):
line = lines[i]
if line.startswith("#"):
break
if line.startswith("|"):
insert_at = i + 1
continue
if insert_at > table_start:
break
lines.insert(insert_at, row)
path.write_text("\n".join(lines) + "\n")
PY
}
if [[ "$AUTO_SUBMIT" -ne 1 ]]; then
cat <<EOF
Inspect at ${OUTPUT}. Three ways to land it (recommended order):
1. Issue + maintainer integrates (preferred — vetting before merge):
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit
(opens an issue via \`gh issue create\`)
Or, no-gh-needed:
https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml
— paste the contents of ${OUTPUT} + ${TAG_DIR}/rig.txt into the body
2. Direct PR (advanced — for contributors who know the matrix structure):
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit --as-pr
Note: matrix is hand-curated; direct PRs may get redirected to an
issue thread for context-gathering before merge.
3. Manual edit (zero tools):
Paste the row from ${OUTPUT} into BENCHMARKS.md via the GitHub web editor.
EOF
exit 0
fi
if [[ -n "$SECTION_OVERRIDE" ]]; then
if ! rg -q "^(##|###) ${SECTION_OVERRIDE//\//\\/}$" BENCHMARKS.md; then
echo "[submit-bench] Known sections:" >&2
valid_sections
die "section override not found: ${SECTION_OVERRIDE}"
fi
fi
PR_TITLE="bench(matrix): @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
ISSUE_TITLE="[bench] @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
BRANCH_USER="$(printf '%s' "${BENCH_ROW_GITHUB_USER}" | tr -cd '[:alnum:]_.-')"
BRANCH_TAG="$(printf '%s' "${TAG}" | tr -cd '[:alnum:]_.-')"
BRANCH="bench/${BRANCH_USER}-${BRANCH_TAG}"
if [[ "$AS_PR" -eq 1 ]]; then
BODY_FILE="$TAG_DIR/PR-body.md"
write_pr_body "$BODY_FILE" "$ROW" "$TAG"
else
BODY_FILE="$TAG_DIR/ISSUE-body.md"
write_issue_body "$BODY_FILE" "$ROW" "$TAG" "$SECTION"
fi
if [[ "${GH_MOCK:-0}" == "1" ]]; then
MOCK_LOG="$TAG_DIR/auto-submit-mock.log"
if [[ "$AS_PR" -eq 1 ]]; then
{
echo "git switch -c ${BRANCH}"
echo "insert BENCHMARKS.md row under: ${SECTION}"
echo "git commit -m ${PR_TITLE}"
echo "git push -u origin ${BRANCH}"
echo "gh pr create --title ${PR_TITLE} --body-file ${BODY_FILE}"
} > "$MOCK_LOG"
log "GH_MOCK=1 — wrote mocked PR auto-submit commands: ${MOCK_LOG}"
log "PR title: ${PR_TITLE}"
else
{
echo "gh issue create --title ${ISSUE_TITLE} --body-file ${BODY_FILE} --label bench-contribution"
} > "$MOCK_LOG"
log "GH_MOCK=1 — wrote mocked issue auto-submit command: ${MOCK_LOG}"
log "Issue title: ${ISSUE_TITLE}"
fi
exit 0
fi
command -v gh >/dev/null 2>&1 || die "'gh' not found. Install GitHub CLI or submit manually."
gh auth status >/dev/null 2>&1 || die "not authed with gh. Run: gh auth login"
if [[ "$AS_PR" -ne 1 ]]; then
ISSUE_URL="$(gh issue create --title "$ISSUE_TITLE" --body-file "$BODY_FILE" --label bench-contribution)"
log "Opened issue: ${ISSUE_URL}"
exit 0
fi
if ! git diff --quiet -- BENCHMARKS.md; then
die "BENCHMARKS.md already has local edits; commit/stash them before --auto-submit"
fi
git fetch origin master >/dev/null 2>&1 || log "WARN: git fetch origin master failed; continuing from current branch"
if git show-ref --verify --quiet refs/remotes/origin/master; then
git switch -c "$BRANCH" origin/master
else
git switch -c "$BRANCH"
fi
insert_row "$SECTION" "$ROW"
git add BENCHMARKS.md
git commit -m "$PR_TITLE"
git push -u origin "$BRANCH"
PR_URL="$(gh pr create --title "$PR_TITLE" --body-file "$BODY_FILE")"
log "Opened PR: ${PR_URL}"
+35 -13
View File
@@ -9,6 +9,7 @@
# Usage:
# bash scripts/switch.sh <variant> # switch + tail until ready
# bash scripts/switch.sh <variant> --no-wait # switch and return immediately
# bash scripts/switch.sh --force <variant> # skip hardware/free-VRAM preflight
# bash scripts/switch.sh --list # show all variants
# bash scripts/switch.sh --down # just bring down whatever's up
#
@@ -42,6 +43,8 @@
#
# Env overrides (rarely needed):
# COMPOSE_BIN Default: "docker compose" (set to e.g. "podman compose" if needed)
# CLUB3090_GPU Single-card GPU index override, e.g. "1" on a hetero rig
# FORCE Set to 1 to skip hardware/free-VRAM preflight
# READY_URL Default: http://localhost:8020/v1/models
# READY_TIMEOUT Default: 600 (seconds — longer for cold cudagraph capture)
@@ -169,15 +172,23 @@ gpu_preflight() {
return
fi
# Free MiB per GPU. Tolerate small overhead (driver, X server) — abort
# if any GPU has <80% of its total memory free.
# if any selected GPU has <80% of its total memory free.
local mem_query
mem_query=$(nvidia-smi --query-gpu=index,memory.free,memory.total --format=csv,noheader,nounits 2>/dev/null) || return
local selector="${NVIDIA_VISIBLE_DEVICES:-${CUDA_VISIBLE_DEVICES:-}}"
local selector_specific=0
if [[ -n "$selector" && "$selector" != "all" && "$selector" != "void" ]]; then
selector_specific=1
fi
local bad=0
while IFS=',' read -r idx free total; do
free=$(echo "$free" | tr -d ' ')
total=$(echo "$total" | tr -d ' ')
idx=$(echo "$idx" | tr -d ' ')
[[ -z "$free" || -z "$total" ]] && continue
if [[ "$selector_specific" -eq 1 && ",${selector}," != *",${idx},"* ]]; then
continue
fi
# Require ≥80% free. Compose default gpu-memory-utilization is 0.92.
local need=$(( total * 80 / 100 ))
if [[ "$free" -lt "$need" ]]; then
@@ -241,8 +252,12 @@ up_variant() {
preflight_genesis_pin "${ROOT_DIR}" || true
preflight_repo_drift "${ROOT_DIR}" || true
preflight_compose_deps "${full_dir}/${file}" || exit 1
if [[ "$eng" == "vllm" ]]; then
preflight_compose_hardware "${full_dir}/${file}" "$v" "${FORCE:-0}" || exit 1
fi
preflight_kv_format_hint "${full_dir}/${file}" || true
fi
gpu_preflight
echo "[switch] bringing up: ${v} (${dir}/${file})"
(cd "${full_dir}" && ${COMPOSE_BIN} -f "${file}" up -d)
@@ -319,24 +334,31 @@ wait_ready() {
# --- arg parsing ---
WAIT=1
case "${1:-}" in
-h|--help|"") usage ;;
--list) list_variants ;;
--down) down_running; exit 0 ;;
esac
VARIANT="$1"
shift || true
for arg in "$@"; do
case "$arg" in
FORCE="${FORCE:-0}"
VARIANT=""
while [[ $# -gt 0 ]]; do
case "$1" in
-h|--help) usage ;;
--list) list_variants ;;
--down) down_running; exit 0 ;;
--no-wait) WAIT=0 ;;
*) echo "Unknown flag: $arg"; exit 1 ;;
--force) FORCE=1 ;;
--*) echo "Unknown flag: $1"; exit 1 ;;
*)
if [[ -n "$VARIANT" ]]; then
echo "ERROR: multiple variants supplied: '${VARIANT}' and '$1'" >&2
exit 1
fi
VARIANT="$1"
;;
esac
shift
done
[[ -n "$VARIANT" ]] || usage
resolve_ready_url "${VARIANT}"
down_running
gpu_preflight
up_variant "${VARIANT}"
[[ $WAIT -eq 1 ]] && wait_ready
echo "[switch] done. Try: curl -s ${READY_URL%/v1/models}/v1/models | jq ."
+187
View File
@@ -0,0 +1,187 @@
#!/usr/bin/env bash
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
TMP_DIR="$(mktemp -d)"
trap 'rm -rf "$TMP_DIR"' EXIT
make_mock_nvidia_smi() {
mkdir -p "${TMP_DIR}/bin"
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK'
#!/usr/bin/env bash
case "$*" in
*"--query-gpu=index,name,memory.total,compute_cap"*)
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
;;
*"--query-gpu=index,memory.total"*)
printf '%s\n' "${MOCK_GPU_MEM_QUERY:-${MOCK_GPU_QUERY:?}}" \
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $3); print $1 ", " $3}'
;;
*"--query-gpu=index,memory.free,memory.total"*)
printf '%s\n' "${MOCK_GPU_FREE_QUERY:?MOCK_GPU_FREE_QUERY not set}"
;;
"-L")
printf '%s\n' "${MOCK_GPU_QUERY:?}" \
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
;;
*)
echo "unexpected nvidia-smi invocation: $*" >&2
exit 2
;;
esac
MOCK
chmod +x "${TMP_DIR}/bin/nvidia-smi"
export PATH="${TMP_DIR}/bin:${PATH}"
}
make_compose() {
local path="$1"
local min_vram="$2"
local min_gpu="$3"
local tp="$4"
local sm="${5:-}"
{
echo "# Hardware metadata (test fixture):"
echo "# Requires-min-vram-gb: ${min_vram}"
echo "# Requires-min-gpu-count: ${min_gpu}"
echo "# Tensor-parallel: ${tp}"
if [[ -n "$sm" ]]; then
echo "# Requires-sm: ${sm}"
fi
echo "services: {}"
} > "$path"
}
assert_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" != *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
run_case() {
local compose="$1"
local variant="$2"
local force="${3:-0}"
(
unset CLUB3090_GPU CUDA_VISIBLE_DEVICES NVIDIA_VISIBLE_DEVICES FORCE
source "${ROOT_DIR}/scripts/preflight.sh"
if preflight_compose_hardware "$compose" "$variant" "$force"; then
echo "STATUS=ok"
echo "NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-}"
else
rc=$?
echo "STATUS=fail:${rc}"
exit "$rc"
fi
) 2>&1
}
expect_failure() {
local compose="$1"
local variant="$2"
local force="${3:-0}"
local output
if output="$(run_case "$compose" "$variant" "$force")"; then
echo "ASSERTION FAILED: expected preflight failure for ${variant}" >&2
echo "--- output ---" >&2
echo "$output" >&2
exit 1
fi
printf '%s' "$output"
}
make_mock_nvidia_smi
single_compose="${TMP_DIR}/single.yml"
dual_compose="${TMP_DIR}/dual.yml"
quad_compose="${TMP_DIR}/quad.yml"
gemma_single="${TMP_DIR}/gemma-single.yml"
missing_meta="${TMP_DIR}/missing.yml"
make_compose "$single_compose" 24 1 1
make_compose "$dual_compose" 24 2 2
make_compose "$quad_compose" 24 4 4
make_compose "$gemma_single" 32 1 1 "9.0+"
echo "services: {}" > "$missing_meta"
# 1. Matched 2x3090 + TP=1 compose: pass, deterministic GPU 0 selection.
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576\n1, 24000, 24576'
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
out="$(run_case "$single_compose" "vllm/default")"
assert_contains "$out" "auto-selected GPU 0"
assert_contains "$out" "NVIDIA_VISIBLE_DEVICES=0"
# 2. Matched 2x3090 + TP=2 compose: pass.
out="$(run_case "$dual_compose" "vllm/dual")"
assert_contains "$out" "TP=2 requires 2 GPU(s)"
assert_contains "$out" "STATUS=ok"
# 3. Matched 2x3090 + TP=4 compose: hard fail.
out="$(expect_failure "$quad_compose" "vllm/dual4")"
assert_contains "$out" "requires 4 visible GPU(s)"
# 4. 1x3090 + TP=2 compose: hard fail.
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576'
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
out="$(expect_failure "$dual_compose" "vllm/dual")"
assert_contains "$out" "requires 2 visible GPU(s)"
# 5. 1x3090 + TP=1, 24 GB floor: pass.
out="$(run_case "$single_compose" "vllm/default")"
assert_contains "$out" "auto-selected GPU 0"
assert_contains "$out" "STATUS=ok"
# 6. 16 GB + 24 GB + TP=1: auto-select the 24 GB card.
MOCK_GPU_QUERY=$'0, RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
MOCK_GPU_FREE_QUERY=$'0, 16000, 16384\n1, 24000, 24576'
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
out="$(run_case "$single_compose" "vllm/default")"
assert_contains "$out" "auto-selected GPU 1"
assert_contains "$out" "NVIDIA_VISIBLE_DEVICES=1"
# 7. 16 GB + 24 GB + TP=2: warn, then proceed for tuned sub-24 GB rigs.
out="$(run_case "$dual_compose" "vllm/dual")"
assert_contains "$out" "WARN:"
assert_contains "$out" "TP=2"
assert_contains "$out" "STATUS=ok"
# 8. 1x3090 + TP=1 compose with 32 GB + sm_9.0+ floor: hard fail.
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576'
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
out="$(expect_failure "$gemma_single" "vllm/gemma-mtp-tp1")"
assert_contains "$out" "requires one GPU with >=32 GB VRAM, sm_9.0+"
# 9. 1xH100 + TP=1 compose with sm_9.0+ floor: pass.
MOCK_GPU_QUERY=$'0, NVIDIA H100 80GB HBM3, 81920, 9.0'
MOCK_GPU_FREE_QUERY=$'0, 80000, 81920'
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
out="$(run_case "$gemma_single" "vllm/gemma-mtp-tp1")"
assert_contains "$out" "auto-selected GPU 0"
assert_contains "$out" "STATUS=ok"
# 10. --force skips the hardware gate.
out="$(run_case "$gemma_single" "vllm/gemma-mtp-tp1" 1)"
assert_contains "$out" "hardware: skipped"
assert_contains "$out" "STATUS=ok"
# 11. Missing metadata warns and allows.
out="$(run_case "$missing_meta" "vllm/local")"
assert_contains "$out" "no hardware metadata"
assert_contains "$out" "STATUS=ok"
echo "test-preflight-vram: ok"
+162
View File
@@ -0,0 +1,162 @@
#!/usr/bin/env bash
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
TMP_DIR="$(mktemp -d)"
ORIG_PATH="$PATH"
trap 'rm -rf "$TMP_DIR"' EXIT
assert_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" != *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
assert_not_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" == *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output not to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
make_mock_tools() {
mkdir -p "${TMP_DIR}/bin"
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK_NVIDIA_SMI'
#!/usr/bin/env bash
case "$*" in
*"--query-gpu=index,name,memory.total,compute_cap"*)
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
;;
"-L")
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}" \
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
;;
*)
echo "unexpected nvidia-smi invocation: $*" >&2
exit 2
;;
esac
MOCK_NVIDIA_SMI
chmod +x "${TMP_DIR}/bin/nvidia-smi"
cat > "${TMP_DIR}/switch-mock" <<'MOCK_SWITCH'
#!/usr/bin/env bash
echo "SWITCHED $*"
MOCK_SWITCH
chmod +x "${TMP_DIR}/switch-mock"
export PATH="${TMP_DIR}/bin:${ORIG_PATH}"
}
set_rig() {
export MOCK_GPU_QUERY="$1"
}
model_status() {
local model="$1"
(
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
compose_hw_model_status "$ROOT_DIR" "$model"
)
}
assert_model_status() {
local model="$1"
local expected_prefix="$2"
local expected_text="${3:-}"
local status
status="$(model_status "$model" || true)"
if [[ "$status" != "${expected_prefix}"* ]]; then
echo "ASSERTION FAILED: ${model} status expected prefix '${expected_prefix}', got '${status}'" >&2
exit 1
fi
if [[ -n "$expected_text" ]]; then
assert_contains "$status" "$expected_text"
fi
}
make_mock_tools
# Matched 2x3090: Qwen and Gemma both have a viable compose.
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "ok|fits your rig"
# Single 24 GB Ampere: Qwen fits; Gemma needs either 32 GB+ single-card or 2x24 GB.
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
assert_contains "$(model_status "gemma-4-31b")" "1× RTX 3090, 24 GB"
# Single 16 GB: neither shipped model has a viable compose.
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9'
assert_model_status "qwen3.6-27b" "no|" "needs 20 GB+ VRAM"
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
# Heterogeneous 16 + 24 GB: Qwen can run on the 24 GB card; Gemma dual cannot.
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "no|" "RTX 4060 Ti, 16 GB + RTX 3090, 24 GB"
# 32 GB+ modern card: Gemma's single-card compose is eligible.
set_rig $'0, NVIDIA GeForce RTX 5090, 32768, 12.0'
assert_model_status "qwen3.6-27b" "ok|fits your rig"
assert_model_status "gemma-4-31b" "ok|fits your rig"
# Non-TTY no-arg setup fails fast with usage rather than hanging.
if out="$(echo | bash "${ROOT_DIR}/scripts/setup.sh" 2>&1)"; then
echo "ASSERTION FAILED: non-TTY no-arg setup unexpectedly succeeded" >&2
echo "$out" >&2
exit 1
fi
assert_contains "$out" "Usage:"
assert_contains "$out" "Interactive picker available in a TTY shell"
# Positional setup path remains non-interactive and reaches the existing flow.
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
out="$(MODEL_DIR="${TMP_DIR}/models" PREFLIGHT_DISK_GB=0 SKIP_GENESIS=1 SKIP_MODEL=1 bash "${ROOT_DIR}/scripts/setup.sh" qwen3.6-27b 2>&1)"
assert_not_contains "$out" "Which model to download?"
assert_contains "$out" "[model] SKIP_MODEL=1"
# The launch wizard's variant-display step marks hardware viability and defaults
# to the single-card long-text recommendation on a single 24 GB card.
out="$(printf '\n' | SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" --no-preflight --no-verify --engine vllm --cards 1 2>&1)"
assert_contains "$out" "Long ctx, text only — Balanced MTP"
assert_contains "$out" "[default]"
assert_contains "$out" "vllm/dual"
assert_contains "$out" "✗ needs 2× 24 GB"
assert_contains "$out" "SWITCHED vllm/long-text"
# TTY-backed no-arg setup supports the cosmetic but real "Both" choice by
# dispatching through the positional path for both model families.
if ! command -v script >/dev/null 2>&1; then
echo "ASSERTION FAILED: util-linux 'script' is required for TTY picker coverage" >&2
exit 1
fi
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
export MODEL_DIR="${TMP_DIR}/models"
export PREFLIGHT_DISK_GB=0
export SKIP_GENESIS=1
export SKIP_MODEL=1
out="$(printf '3\n' | script -qec "bash '${ROOT_DIR}/scripts/setup.sh'" /dev/null 2>&1)"
assert_contains "$out" "[setup] Which model to download?"
assert_contains "$out" "Both"
assert_contains "$out" "[setup] downloading both supported models"
skip_count="$(grep -c "\[model\] SKIP_MODEL=1" <<< "$out" || true)"
if [[ "$skip_count" != "2" ]]; then
echo "ASSERTION FAILED: expected Both choice to dispatch two model setup runs, got ${skip_count}" >&2
echo "--- output ---" >&2
echo "$out" >&2
exit 1
fi
echo "test-setup-picker: ok"
+111
View File
@@ -0,0 +1,111 @@
#!/usr/bin/env bash
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
cd "$ROOT_DIR"
# shellcheck source=../lib/bench-row-formatter.sh
source "$ROOT_DIR/scripts/lib/bench-row-formatter.sh"
assert_contains() {
local haystack="$1"
local needle="$2"
if [[ "$haystack" != *"$needle"* ]]; then
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
echo "--- output ---" >&2
echo "$haystack" >&2
exit 1
fi
}
assert_columns() {
local row="$1"
local expected="$2"
local cols
cols="$(awk -F'|' '{print NF - 2}' <<< "$row")"
if [[ "$cols" != "$expected" ]]; then
echo "ASSERTION FAILED: expected $expected columns, got $cols" >&2
echo "$row" >&2
exit 1
fi
}
fixtures=()
while IFS= read -r fixture; do
fixtures+=("$fixture")
done < <(bench_row_fixtures)
if [[ "${#fixtures[@]}" -lt 6 ]]; then
echo "ASSERTION FAILED: expected at least 6 submit-bench fixtures, found ${#fixtures[@]}" >&2
exit 1
fi
for dir in "${fixtures[@]}"; do
section="$(bench_row_section "$dir")"
row="$(bench_row_format "$dir")"
[[ "$row" == \|*\| ]] || {
echo "ASSERTION FAILED: row is not a markdown table row for $dir" >&2
echo "$row" >&2
exit 1
}
assert_contains "$row" "Report: \`results/rebench/${dir##*/}/REPORT.md\`"
if [[ "$section" == "Gemma 4 31B (community-experimental)" ]]; then
assert_columns "$row" 10
else
assert_columns "$row" 8
fi
done
tag="qwen-int8-pth-n4-2026-05-10"
rm -f "results/rebench/${tag}/BENCHMARKS-row.md" \
"results/rebench/${tag}/PR-body.md" \
"results/rebench/${tag}/ISSUE-body.md" \
"results/rebench/${tag}/auto-submit-mock.log"
out="$(bash scripts/submit-bench.sh --tag "$tag")"
assert_contains "$out" "Generated BENCHMARKS row for section: Dual-card (2× RTX 3090, TP=2)"
assert_contains "$out" "Wrote: results/rebench/${tag}/BENCHMARKS-row.md"
assert_contains "$out" "1. Issue + maintainer integrates"
assert_contains "$out" "2. Direct PR"
assert_contains "$out" "3. Manual edit"
test -s "results/rebench/${tag}/BENCHMARKS-row.md"
if out="$(bash scripts/submit-bench.sh --tag does-not-exist 2>&1)"; then
echo "ASSERTION FAILED: missing tag unexpectedly succeeded" >&2
exit 1
fi
assert_contains "$out" "tag dir not found: results/rebench/does-not-exist"
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit)"
assert_contains "$out" "Issue title: [bench] @octocat ${tag}"
test -s "results/rebench/${tag}/auto-submit-mock.log"
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh issue create --title [bench] @octocat ${tag}"
test -s "results/rebench/${tag}/ISSUE-body.md"
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "results/rebench/${tag}/REPORT.md"
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "Proposed BENCHMARKS.md row"
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit --as-pr)"
assert_contains "$out" "PR title: bench(matrix): @octocat ${tag}"
test -s "results/rebench/${tag}/PR-body.md"
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh pr create --title bench(matrix): @octocat ${tag}"
assert_contains "$(cat "results/rebench/${tag}/PR-body.md")" "results/rebench/${tag}/REPORT.md"
tmp_bin="$(mktemp -d)"
trap 'rm -rf "$tmp_bin"; rm -f "results/rebench/${tag}/BENCHMARKS-row.md" "results/rebench/${tag}/PR-body.md" "results/rebench/${tag}/ISSUE-body.md" "results/rebench/${tag}/auto-submit-mock.log"' EXIT
cat > "${tmp_bin}/gh" <<'MOCK_GH'
#!/usr/bin/env bash
if [[ "$1" == "auth" && "$2" == "status" ]]; then
exit 1
fi
echo "unexpected gh call: $*" >&2
exit 2
MOCK_GH
chmod +x "${tmp_bin}/gh"
if out="$(PATH="${tmp_bin}:${PATH}" bash scripts/submit-bench.sh --tag "$tag" --auto-submit 2>&1)"; then
echo "ASSERTION FAILED: unauthenticated gh path unexpectedly succeeded" >&2
exit 1
fi
assert_contains "$out" "not authed with gh. Run: gh auth login"
echo "test-submit-bench: ok"