Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
22bf2e9398 | ||
|
|
99a0b66224 | ||
|
|
ef770322f4 | ||
|
|
301083491b | ||
|
|
26985527f7 | ||
|
|
642dfba8ef | ||
|
|
6617e1e090 | ||
|
|
1678ca0c8b | ||
|
|
14ffe45667 | ||
|
|
722f998ff3 |
+10
-5
@@ -38,11 +38,16 @@
|
||||
# GPU selection
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
# Which GPUs to expose to docker. Single-card composes use just `0`,
|
||||
# dual-card composes use `0,1`. The compose files set this themselves;
|
||||
# override here if your physical layout differs (e.g. you run dual-card
|
||||
# on cards 2,3).
|
||||
# CUDA_VISIBLE_DEVICES=0,1
|
||||
# Which GPU(s) to expose to docker.
|
||||
#
|
||||
# scripts/switch.sh auto-selects the largest eligible card for TP=1 vLLM
|
||||
# composes. Override with CLUB3090_GPU when you know which physical GPU should
|
||||
# run a single-card compose (for example, a 24 GB card beside a 16 GB card).
|
||||
# CLUB3090_GPU=1
|
||||
#
|
||||
# For direct `docker compose ... up` runs, or TP>=2 non-default layouts, set
|
||||
# NVIDIA_VISIBLE_DEVICES explicitly. Use comma-separated physical indices.
|
||||
# NVIDIA_VISIBLE_DEVICES=0,1
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
## Rig bench submission
|
||||
|
||||
> ⚠ Most bench submissions go through an issue (see `CONTRIBUTING.md` "Submitting your bench").
|
||||
> This PR template is for contributors who explicitly chose the direct-PR path.
|
||||
> The maintainer may redirect to an issue thread before merge.
|
||||
|
||||
<!-- This PR was auto-generated by `bash scripts/submit-bench.sh --auto-submit --as-pr --tag <TAG>`. -->
|
||||
<!-- Review the row below; the PR reviewer may move it within the target section. -->
|
||||
|
||||
### New row
|
||||
|
||||
<!-- The generated BENCHMARKS.md row goes here -->
|
||||
|
||||
### Rig
|
||||
|
||||
<!-- Output of `bash scripts/report.sh` (redacted) -->
|
||||
|
||||
### Full results
|
||||
|
||||
See `results/rebench/<TAG>/REPORT.md` for the full per-phase breakdown.
|
||||
@@ -16,6 +16,80 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.5.3 — 2026-05-13
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(scripts): add submit-bench flow ([ef77032](https://github.com/noonghunna/club-3090/commit/ef770322f43724f612a80393f547e5da218b5bf7))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.3`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.2...v0.5.3)
|
||||
## v0.5.2 — 2026-05-13
|
||||
|
||||
|
||||
### 🎯 New models + serving paths
|
||||
|
||||
- Add hardware-aware compose preflight ([2698552](https://github.com/noonghunna/club-3090/commit/26985527f75d8da2a32a8a2f985989d5dcf9e89a))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.2`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.1...v0.5.2)
|
||||
## v0.5.1 — 2026-05-13
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(qwen): PR #35936 overlay — sidecar pattern to resolve Genesis RO-mount conflict ([6617e1e](https://github.com/noonghunna/club-3090/commit/6617e1e090a6f52708aaf83821a92b612c4ad869))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs: clarify MODEL_DIR — second drive / HF cache / Windows-WSL ([1678ca0](https://github.com/noonghunna/club-3090/commit/1678ca0c8ba43ea09fad9073639a48337f2ff163))
|
||||
- docs(upstream): correct stale vllm#40807 row + add #40798/#42215 row ([14ffe45](https://github.com/noonghunna/club-3090/commit/14ffe45667fb0aa292839be05ae3ad0d139d3a04))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.0...v0.5.1)
|
||||
## v0.5.0 — 2026-05-12
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(qwen): ship froggeric chat-template fixes as default-on ([84498d4](https://github.com/noonghunna/club-3090/commit/84498d47aaf7a2fdb7c0203d53bb64414a64b6c1))
|
||||
- feat(vllm): add PR #35936 required-tool fallback overlay ([28b16b5](https://github.com/noonghunna/club-3090/commit/28b16b5dc9d602a8e1c4b8d5496aa82cbca7f95d))
|
||||
- feat(qwen-tq3): add CLUB3090_TQ_K1_SKIP_MTP layer-filter for PR #40914 K+1 dispatch ([6b2a7d5](https://github.com/noonghunna/club-3090/commit/6b2a7d553b6164bb4854e80fc0acdf8dcec18a87))
|
||||
|
||||
|
||||
### 🎯 New models + serving paths
|
||||
|
||||
- compose(tq3-mtp-genesis): pin to Genesis v7.72.2 known-good vLLM nightly ([570fa71](https://github.com/noonghunna/club-3090/commit/570fa71240a12ab642693c0570851611e933f8c4))
|
||||
|
||||
|
||||
### 📊 Benchmarks + cross-rig data
|
||||
|
||||
- bench(matrix): @ygafarov first heterogeneous Ampere + Blackwell eGPU dual ([1770931](https://github.com/noonghunna/club-3090/commit/1770931729a354bad319c8b58bdee143fe6ebce2))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs(dtype-matrix): more polish — RDNA naming, FP8 maturity caveats, AMD detection ([62b3b45](https://github.com/noonghunna/club-3090/commit/62b3b455a9ea5146cbf5576febd0c8b25a8c0fa1))
|
||||
- docs(dtype-matrix): polish nuances + add Intel and AMD vendor sections ([3d4548c](https://github.com/noonghunna/club-3090/commit/3d4548c50422da07c16fcba2a59d6f42f268355b))
|
||||
- docs(dtype-matrix): per-arch hardware accelerator matrix for compose optimization ([9c6d3cf](https://github.com/noonghunna/club-3090/commit/9c6d3cfba1de9d9079d6eafc9ff68c352cac7197))
|
||||
- docs(faq): add 'INT8 PTH doesn't scale at concurrency — is that a bug?' ([df53287](https://github.com/noonghunna/club-3090/commit/df53287b1c26ec83be8a30eec24baf2bddc993eb))
|
||||
- docs(tq3-mtp): writeup + charts for the Genesis-backed TQ3+MTP path ([c2b1c93](https://github.com/noonghunna/club-3090/commit/c2b1c93872f84fa9afa3bbe41360dc42be28c066))
|
||||
- docs(qwen-tq3): close round-4 — #40914 not shippable, route to nomtp + Genesis ([9fba037](https://github.com/noonghunna/club-3090/commit/9fba03788e30151f9bf8c85f279260696954d094))
|
||||
- docs(qwen-tq3): re-tombstone tq3-mtp.yml after round-3 MTP-skip validation ([063d3e9](https://github.com/noonghunna/club-3090/commit/063d3e943ce8da9cc69bd30c68b87426aca6202e))
|
||||
|
||||
|
||||
### 🧹 Maintenance
|
||||
|
||||
- refactor(qwen): rename int8-tq3 → tq3-* family + add no-MTP + Genesis variants ([6182922](https://github.com/noonghunna/club-3090/commit/6182922225dfeee1c28084d1ff917bfd25539520))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.0`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.4.0...v0.5.0)
|
||||
## v0.4.0 — 2026-05-11
|
||||
|
||||
|
||||
|
||||
@@ -45,6 +45,40 @@ Two GitHub channels, two different shapes of conversation. Picking the right one
|
||||
|
||||
---
|
||||
|
||||
## Submitting your bench
|
||||
|
||||
The matrix is hand-curated — the canonical path is to file an **issue** with your rig + numbers; we'll review, ask clarifying questions, and integrate.
|
||||
|
||||
After running `bash scripts/rebench-full.sh`, generate a paste-ready row:
|
||||
|
||||
```bash
|
||||
bash scripts/submit-bench.sh --tag <your-tag>
|
||||
```
|
||||
|
||||
The script writes `results/rebench/<tag>/BENCHMARKS-row.md`. To submit:
|
||||
|
||||
### Path A — Auto-issue (recommended, requires `gh auth login`)
|
||||
|
||||
```bash
|
||||
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit
|
||||
```
|
||||
|
||||
Opens an issue via `gh issue create` with your rig + row pre-filled.
|
||||
|
||||
### Path B — Manual issue (no tools beyond browser)
|
||||
|
||||
Open https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml and paste the row + your `rig.txt` into the body.
|
||||
|
||||
### Path C — Direct PR (advanced)
|
||||
|
||||
```bash
|
||||
bash scripts/submit-bench.sh --tag <your-tag> --auto-submit --as-pr
|
||||
```
|
||||
|
||||
For contributors who know the `BENCHMARKS.md` section structure and want to propose the exact row. The maintainer may still redirect to an issue thread for context-gathering before merge — direct PRs aren't a fast-path bypass.
|
||||
|
||||
---
|
||||
|
||||
## Process for non-trivial changes
|
||||
|
||||
1. **Open an issue first** for anything bigger than a typo fix or a one-line measurement contribution. We'll either align on shape or explain why we'd land it differently — saves you a wasted afternoon.
|
||||
|
||||
@@ -67,6 +67,9 @@ git clone https://github.com/noonghunna/club-3090.git
|
||||
cd club-3090
|
||||
|
||||
# 2. Download + SHA-verify the model (~20 GB; clones Genesis patches too)
|
||||
# (asks you where to put model weights — pick in-repo default, ~/models, or
|
||||
# a custom path on a different drive. To skip the prompt:
|
||||
# `export MODEL_DIR=/mnt/your-drive/models` before running. See FAQ + .env.example.)
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
|
||||
# 3. Pick a config + boot it (interactive wizard — asks engine / cards / workload)
|
||||
|
||||
+19
-2
@@ -130,9 +130,26 @@ If your numbers on the same compose look different from ours by >15%, the most l
|
||||
|
||||
## Setup
|
||||
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go?
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
|
||||
|
||||
`<repo>/models-cache/` by default. Override with `MODEL_DIR=/path/to/your/scratch bash scripts/setup.sh qwen3.6-27b`. See [`.env.example`](../.env.example) for all env vars.
|
||||
Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
|
||||
|
||||
1. **`MODEL_DIR` env var in your shell** — takes precedence over everything:
|
||||
```bash
|
||||
export MODEL_DIR=/mnt/your-second-drive/models
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
```
|
||||
2. **`.env` file at repo root** — picked up automatically on every script run. See [`.env.example`](../.env.example).
|
||||
3. **Interactive prompt** — `bash scripts/setup.sh qwen3.6-27b` with nothing set offers three choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
|
||||
4. **Silent fallback** — `<repo>/models-cache/`. Functional but pollutes the git tree; not recommended.
|
||||
|
||||
Every script that touches model paths reads from the same `MODEL_DIR`. The compose YAMLs' volume mount is `${MODEL_DIR:-...}:/root/.cache/huggingface` — once set, every container reads + writes there.
|
||||
|
||||
**HF env-var integration** — we don't directly respect `HF_HOME` / `HF_HUB_CACHE` because we mount a host directory INTO the container's `/root/.cache/huggingface`, not the host's HF cache. The internal layout inside `MODEL_DIR` matches HF's repo-cache convention (`<MODEL_DIR>/<repo-subdir>/`), so models downloaded by `setup.sh` are byte-compatible with anything that reads HF's local cache layout. Two clean workarounds if you already have an HF cache you want to reuse:
|
||||
- Set `MODEL_DIR=$HF_HOME/hub`
|
||||
- Or symlink between them
|
||||
|
||||
**On Windows / WSL2** — same mechanism. Docker Desktop handles path translation. Use Windows paths (`D:\models`) from PowerShell or WSL paths (`/mnt/d/models`) from WSL. If you flip between Linux and Windows on the same rig, point `MODEL_DIR` at a drive both OSes can see — the model files themselves are OS-agnostic.
|
||||
|
||||
### How do I keep my install up-to-date?
|
||||
|
||||
|
||||
@@ -33,6 +33,22 @@ The recipes are written against 3090 specifically but should work on:
|
||||
|
||||
**Won't work:** anything with <20 GB VRAM (3060, 3070, stock 3080, 3080 Ti). The 27B model in INT4 is ~18 GB — KV pool + activations push past 24 GB on smaller cards even with aggressive quantization. **Modded 20 GB 3080s do work** (see row above) — the mod gives them enough headroom for the 27B + TQ K8V4 KV path on TP=2, with `mem-util=0.82` to absorb cudagraph profiling overhead.
|
||||
|
||||
### Mismatched / heterogeneous GPUs
|
||||
|
||||
`scripts/switch.sh` reads hardware metadata from the vLLM compose headers before starting Docker. The preflight checks required GPU count, per-GPU VRAM, tensor parallel size, and any hard SM floor.
|
||||
|
||||
For TP=1 vLLM composes, `switch.sh` auto-selects the largest eligible GPU and exports it through `NVIDIA_VISIBLE_DEVICES`. On a mixed 16 GB + 24 GB rig, `bash scripts/switch.sh vllm/default` should pick the 24 GB card instead of trying to boot on GPU 0 blindly.
|
||||
|
||||
Overrides:
|
||||
|
||||
```bash
|
||||
CLUB3090_GPU=1 bash scripts/switch.sh vllm/default
|
||||
NVIDIA_VISIBLE_DEVICES=2,3 bash scripts/switch.sh vllm/dual
|
||||
bash scripts/switch.sh --force vllm/gemma-mtp-tp1
|
||||
```
|
||||
|
||||
Use `--force` only when you are intentionally testing an unsupported combo. Example: `vllm/gemma-mtp-tp1` is now preflight-blocked on a 24 GB 3090 because the compose is preserved for 32 GB / newer-SM single-card rigs.
|
||||
|
||||
### Note for sub-24 GB cards
|
||||
|
||||
On 20 GB cards (modded 3080) the cudagraph-profiling overhead is a meaningful slice of available VRAM. Drop `--gpu-memory-utilization` to **0.82** (vs shipped 0.95 for 24 GB). vLLM nightly's `gpu_worker.py` reports the equivalent effective KV size in the boot log; tune to keep activation headroom for the ~15K tool-prefill peak (verify-full check 8). Credit: [@troymroberts](https://github.com/troymroberts).
|
||||
|
||||
+2
-1
@@ -69,7 +69,8 @@ Run `bash scripts/maintenance/list-image-pins.sh` for a live snapshot.
|
||||
|---|---|---|---|
|
||||
| [#35936](https://github.com/vllm-project/vllm/pull/35936) — `tool_choice="required"` falls back to configured tool parser | 🟡 Open / **local overlay active** | Qwen3-Coder with `--tool-call-parser qwen3_coder` emits XML-style tool calls. On pinned nightly `1acd67a79`, non-streaming `tool_choice="required"` validates JSON only, bypasses the configured parser, and returns `tool_calls=[]`. MLS-Bench hits this when `thinking.enabled=false`. | Vendored overlay: [`models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md). Drop when #35936 or equivalent lands in our pinned image. |
|
||||
| [#40361](https://github.com/vllm-project/vllm/pull/40361) — Marlin pad-sub-tile-n | 🟡 Open, mergeable, **stale 13d** (last update 2026-04-20) | All 4 dual-card composes + `dual-nvlink.yml` + `dual-nvlink-turbo.yml` mount the patched files vendored in-repo at `models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/`. Drops out as a setup dependency when this merges + propagates. | Vendored mount: see [`models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md). Queued for rebase + ping next week (see "Active follow-ups" table above). |
|
||||
| [#40807](https://github.com/vllm-project/vllm/issues/40807) — `.tolist()` cudagraph crash on continuation-prefill | ⚫ Local workaround | Single-card TQ3 + spec-decode + chunked-prefill blocked without it. We ship a file-edit patch. | `patch_tolist_cudagraph.py` runs in `setup.sh`. Drop when upstream fixes the sync. |
|
||||
| [#40807](https://github.com/vllm-project/vllm/issues/40807) — `.tolist()` cudagraph crash on continuation-prefill | ✅ **Retired locally** (2026-05-05 Genesis v7.72.2 bump) — Genesis ships [P78 `TOLIST_CAPTURE_GUARD`](../models/qwen3.6-27b/vllm/compose/dual/tq3-mtp-genesis.yml) as the equivalent fix. Currently disabled (`=0`) on `tq3-mtp-genesis.yml` after rebench-full leg 6 (2026-05-11) passed clean with it off — apparent root cause is now covered by Genesis PN34 (workspace-lock relax) + post-#41434 attention rework. Non-Genesis composes on `1acd67a79` pin run without any guard for this bug; unvalidated at long-context TurboQuant chunked-prefill (worth testing per cferra's vllm#41403 validation pass — see vllm#40798 row below). | None active. Drop the Genesis env var permanently if a future v7.73.x rebench leaves it OFF without regression. |
|
||||
| [#40798](https://github.com/vllm-project/vllm/pull/40798) + [#42215](https://github.com/vllm-project/vllm/pull/42215) — share decode scratch workspace pre-CUDA-graph + decode-kernel warmup | 🟡 Open, validated cross-rig | Pair closes the `AssertionError: Workspace is locked but allocation requires NMB` crash that fires at `turboquant_attn.py:_continuation_prefill` for ≥48K-token chunked-prefill with TurboQuant KV. Independently validated on 2× 3090 sm_86 by cferra (vllm#41403 [comment](https://github.com/vllm-project/vllm/issues/41403#issuecomment-4435164709), 2026-05-12). | Genesis [PN34 `WORKSPACE_LOCK_RELAX`](../models/qwen3.6-27b/vllm/compose/dual/tq3-mtp-genesis.yml) addresses the same symptom via a different mechanism (relax-lock vs reserve-before-capture). On non-Genesis composes (`1acd67a79` pin) we currently have no guard — re-validate against this PR pair once they propagate to a nightly we pin to, then A/B PN34 vs upstream. |
|
||||
| [#40849](https://github.com/vllm-project/vllm/pull/40849) — MTP draft online-quant propagation | 🟡 Open / Genesis backport active | Closes Cliff 1 on FP8+MTP path (`tools-text.yml`). | Genesis PN8 backport: `GENESIS_ENABLE_PN8_MTP_DRAFT_ONLINE_QUANT=1`. |
|
||||
| [#40914](https://github.com/vllm-project/vllm/pull/40914) — Sandermage K+1 verify routing | 🟡 Open, ❌ negative on our Qwen3.6-27B stack | **Reframed 2026-05-11:** the synthetic `seq_lens` K+1 route is not the P67-equivalent we need here. Local rebase on post-#41434 nightly made MTP acceptance look perfect (AL=4.0 / ~100%) but produced `!`-flood needle corruption plus tool/multi-turn timeouts. Dropping it improved verify-stress from 3/7 to 5/7, but TQ3/TQ4/k8v4 + MTP still fail long-context needles. | Do not ship Genesis-free TQ+MTP on #40914 alone. Use `dual/tq3-nomtp.yml` without Genesis, or `dual/tq3-mtp-genesis.yml` with Genesis P67/P67b. |
|
||||
| [#40334](https://github.com/vllm-project/vllm/pull/40334) — DFlash `combine_hidden_states` dtype mismatch | 🟡 Open | All `dual-dflash*.yml` need `--dtype bfloat16` flag to work around. | Composes set `--dtype bfloat16`. Drop when this lands. |
|
||||
|
||||
@@ -58,6 +58,10 @@
|
||||
# - max-num-seqs default 4 (3dluvr's 1 is single-stream-only); override
|
||||
# MAX_NUM_SEQS=1 for max-context single-stream like 3dluvr.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-awq:
|
||||
# Latest nightly that contains PR #41745 (Gemma 4 MTP). AWQ doesn't need
|
||||
@@ -72,6 +76,7 @@ services:
|
||||
- ../../cache/torch_compile_awq:/root/.cache/vllm/torch_compile_cache
|
||||
- ../../cache/triton_awq:/root/.triton/cache
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -89,6 +89,10 @@
|
||||
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
|
||||
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-bf16:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -118,6 +122,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -93,6 +93,10 @@
|
||||
# 1. PR #41703 z-lab DFlash drafter — see ../../patches/vllm-gemma4-dflash/README
|
||||
# 2. PR #40391 per-token-head — see ../../patches/vllm-gemma4-dflash-int8/README
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-dflash-int8:
|
||||
# SAME pin as dual-dflash.yml — DFlash overlay was rebased to e47c98ef.
|
||||
@@ -139,6 +143,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-dflash-int8/v1/worker/kv_cache_shape_utils.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/kv_cache_shape_utils.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -65,6 +65,10 @@
|
||||
# - default (auto = bfloat16) → sidesteps all the above.
|
||||
# Smaller KV pool than fp8 would give but it's the only Ampere-shippable.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-dflash:
|
||||
# Nightly bumped to 2026-05-06 to match rebase target proximity (Codex
|
||||
@@ -104,6 +108,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-dflash/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -48,6 +48,10 @@
|
||||
# - default (auto = bfloat16) → sidesteps all fp8 kernel paths
|
||||
# Smaller KV pool than fp8 but at 32K test ctx that's not the bottleneck.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -63,6 +67,7 @@ services:
|
||||
- ../../cache/triton:/root/.triton/cache
|
||||
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -177,6 +177,11 @@
|
||||
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
|
||||
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
# Requires-sm: 9.0+
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-int8-tq3:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -213,6 +218,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -89,6 +89,10 @@
|
||||
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
|
||||
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-int8:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -118,6 +122,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -48,6 +48,11 @@
|
||||
# - default (auto = bfloat16) → sidesteps all fp8 kernel paths
|
||||
# Smaller KV pool than fp8 but at 32K test ctx that's not the bottleneck.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 32
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
# Requires-sm: 9.0+
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-tp1:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -63,6 +68,7 @@ services:
|
||||
- ../../cache/triton:/root/.triton/cache
|
||||
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -20,6 +20,10 @@
|
||||
# Run:
|
||||
# docker compose -f dual/bf16.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-bf16:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/bf16.yml — 2026-05-08
|
||||
@@ -44,8 +48,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -53,6 +63,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -74,6 +85,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -48,6 +48,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/carnice-bf16mtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-carnice-bf16mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -73,9 +77,16 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# PCIe-only stack — disable NCCL features that assume NVLink.
|
||||
@@ -103,6 +114,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -43,6 +43,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/dflash-noviz.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -62,8 +66,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -71,6 +81,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -97,6 +108,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -67,6 +67,10 @@
|
||||
#
|
||||
# docker compose -f dual/dflash.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -86,8 +90,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -95,6 +105,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -121,6 +132,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -49,6 +49,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/docker-compose.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual:
|
||||
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
|
||||
@@ -74,8 +78,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -83,6 +93,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# PCIe-only stack — disable NCCL features that assume NVLink.
|
||||
@@ -112,6 +123,8 @@ services:
|
||||
# hardware where Cliff 2 GDN activation spikes occur at runtime
|
||||
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
|
||||
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -20,6 +20,10 @@
|
||||
# Run:
|
||||
# docker compose -f dual/int8.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-int8:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8.yml — 2026-05-08
|
||||
@@ -44,8 +48,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -53,6 +63,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -74,6 +85,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -75,6 +75,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink-dflash-noviz.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink-dflash-noviz:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -94,8 +98,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -103,6 +113,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
@@ -131,6 +142,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -71,6 +71,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink-dflash.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -90,8 +94,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -99,6 +109,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
@@ -127,6 +138,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -61,6 +61,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink-turbo.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink-turbo:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -87,8 +91,14 @@ services:
|
||||
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
|
||||
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -96,6 +106,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
@@ -222,6 +233,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -57,6 +57,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink:
|
||||
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
|
||||
@@ -80,8 +84,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -89,6 +99,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
@@ -119,6 +130,8 @@ services:
|
||||
# hardware where Cliff 2 GDN activation spikes occur at runtime
|
||||
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
|
||||
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -33,6 +33,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/qwopus-bf16mtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwopus-bf16mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -49,9 +53,16 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -69,6 +80,14 @@ services:
|
||||
- driver: nvidia
|
||||
count: all
|
||||
capabilities: [gpu]
|
||||
entrypoint:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
- --model
|
||||
- /root/.cache/huggingface/qwopus3.6-27b-int4-recipe-d-bf16mtp
|
||||
|
||||
@@ -61,6 +61,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/tq3-mtp-genesis.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-tq3-mtp-genesis:
|
||||
# Pinned to a Genesis v7.72.2 known-good vLLM nightly. The canonical
|
||||
@@ -101,6 +105,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -194,7 +199,14 @@ services:
|
||||
# PN54 — GDN contiguous-call deduplication (Cliff 2b OOM mitigation).
|
||||
- GENESIS_ENABLE_PN54=${GENESIS_ENABLE_PN54:-0}
|
||||
# Explicit OFFs to match Sandermage's PROD env-var set verbatim:
|
||||
# P78 (P78_TOLIST_CAPTURE_GUARD) — superseded by our patch_tolist_cudagraph.py
|
||||
# P78 (P78_TOLIST_CAPTURE_GUARD) — vllm#40807 `.tolist()` cudagraph guard.
|
||||
# Was OFF here because we previously shipped our own patch_tolist_cudagraph.py
|
||||
# sidecar covering the same bug; that sidecar was retired 2026-05-05 (Genesis
|
||||
# v7.72.2 bump). We keep P78=0 because rebench-full leg 6 (2026-05-11) passed
|
||||
# clean with it off — 7/7 verify-stress, 0/100 silent-empty soak, 18/30 aider.
|
||||
# PN34 (workspace-lock relax) + post-#41434 attention rework appear to cover
|
||||
# the symptom on this Genesis-pinned nightly. Flip to 1 if a future bench
|
||||
# surfaces .tolist() cudagraph corruption.
|
||||
# P81 (FP8 block-scaled M<=8) — FP8-specific, no-op on our TQ3 path
|
||||
# P82 — biased on small-batch single-stream Lorbus INT4 + MTP K=3 (Sander PROD)
|
||||
- GENESIS_ENABLE_P78_TOLIST_CAPTURE_GUARD=0
|
||||
|
||||
@@ -59,6 +59,10 @@
|
||||
# Run (only when all 5 upstream fixes are available):
|
||||
# docker compose -f dual/tq3-mtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-int8-tq3:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8-tq3.yml — 2026-05-08
|
||||
@@ -91,8 +95,14 @@ services:
|
||||
# cudagraph path off the table entirely.
|
||||
- ../../patches/vllm-pr40798-rebased/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -100,6 +110,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- VLLM_ENFORCE_EAGER=1
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
@@ -122,6 +133,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -35,6 +35,10 @@
|
||||
# To run:
|
||||
# docker compose -f dual/tq3-nomtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-int8-tq3-nomtp:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8-tq3.yml — 2026-05-08
|
||||
@@ -75,8 +79,14 @@ services:
|
||||
# used; nightly already has the kernel + caller for the rest.
|
||||
- ../../patches/vllm-pr40798-rebased/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -84,6 +94,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -105,6 +116,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -45,6 +45,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/turbo.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-turbo:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -71,8 +75,14 @@ services:
|
||||
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
|
||||
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -80,6 +90,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -204,6 +215,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -60,6 +60,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f multi4/dflash.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 4
|
||||
# Tensor-parallel: 4
|
||||
services:
|
||||
vllm-qwen36-27b-multi4-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -79,8 +83,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -88,6 +98,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
@@ -114,6 +125,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -57,6 +57,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f multi4/docker-compose.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 4
|
||||
# Tensor-parallel: 4
|
||||
services:
|
||||
vllm-qwen36-27b-multi4:
|
||||
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
|
||||
@@ -82,8 +86,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -91,6 +101,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# PCIe-only stack — disable NCCL features that assume NVLink.
|
||||
@@ -120,6 +131,8 @@ services:
|
||||
# hardware where Cliff 2 GDN activation spikes occur at runtime
|
||||
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
|
||||
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -129,6 +129,10 @@
|
||||
# Then send requests with structured_outputs in extra_body — see
|
||||
# docs/STRUCTURED_COT.md for client examples.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-bounded-thinking:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -148,8 +152,14 @@ services:
|
||||
# workspace_lock_disable — relaxes vllm#39226 strict assertion until
|
||||
# Sandermage's P98 marker fix lands. v0.20-only requirement.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -157,6 +167,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
@@ -275,6 +286,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -95,6 +95,10 @@
|
||||
# bash ../scripts/setup.sh # ensures Genesis tree at v7.14+ layout
|
||||
# cd compose && docker compose up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -119,8 +123,14 @@ services:
|
||||
# workspace_lock_disable — relaxes vllm#39226 strict assertion until
|
||||
# Sandermage's P98 marker fix lands. v0.20-only requirement.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -128,6 +138,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
@@ -206,6 +217,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -109,6 +109,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/long-text.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-long-text-no-mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -136,8 +140,14 @@ services:
|
||||
# surface natively (PN12 native + PN25 + PN17 + P15B). See migration
|
||||
# commit history if you need to resurrect.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -145,6 +155,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
@@ -300,6 +311,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -119,6 +119,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/long-text.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-long-text:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -146,8 +150,14 @@ services:
|
||||
# surface natively (PN12 native + PN25 + PN17 + P15B). See migration
|
||||
# commit history if you need to resurrect.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -155,6 +165,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
@@ -317,6 +328,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -95,6 +95,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/long-vision.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-long-vision:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -114,8 +118,14 @@ services:
|
||||
# workspace_lock_disable — relaxes vllm#39226 strict assertion until
|
||||
# Sandermage's P98 marker fix lands. v0.20-only requirement.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -123,6 +133,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
@@ -241,6 +252,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -31,6 +31,10 @@
|
||||
# Run:
|
||||
# cd compose && docker compose -f single/minimal.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 20
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-minimal:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -46,8 +50,14 @@ services:
|
||||
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
|
||||
- ../../cache/triton:/root/.triton/cache
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -55,6 +65,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
@@ -82,6 +93,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -33,6 +33,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/tools-text.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -51,8 +55,14 @@ services:
|
||||
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
|
||||
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -60,6 +70,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
@@ -130,6 +141,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -49,10 +49,48 @@ This overlay therefore lands the PR's intent in **both** places in
|
||||
as before so XML-emitting paths keep working.
|
||||
|
||||
The chat completion file is vendored unchanged from the pinned nightly as
|
||||
a matching drop-in mount target. The PR's streaming-side hunks target a
|
||||
pre-parser-manager code path that has been refactored on `nightly-1acd67a79`;
|
||||
non-streaming clients (MLS-Bench, our curl repro, most agent harnesses)
|
||||
exercise only the engine-side fix.
|
||||
a future-ready slot. The PR's streaming-side hunks target a pre-parser-manager
|
||||
code path that has been refactored on `nightly-1acd67a79`; non-streaming
|
||||
clients (MLS-Bench, our curl repro, most agent harnesses) exercise only the
|
||||
engine-side fix.
|
||||
|
||||
## Installation: sidecar pattern (v0.5.1+)
|
||||
|
||||
The overlay is **not** mounted directly at vLLM's site-packages paths. Instead:
|
||||
|
||||
1. Both files are bind-mounted at side paths under `/etc/club3090/`:
|
||||
- `pr35936-chat-completion-serving.py`
|
||||
- `pr35936-engine-serving.py`
|
||||
2. `install.sh` is bind-mounted at `/etc/club3090/install-pr35936.sh`
|
||||
3. The compose entrypoint invokes `bash /etc/club3090/install-pr35936.sh`
|
||||
**before** any other patch step (`python3 -m vllm._genesis.patches.apply_all`
|
||||
in Genesis-loaded composes, or `exec vllm serve` in Genesis-less ones).
|
||||
4. `install.sh` copies our files into vLLM's site-packages with `cp`. The
|
||||
destination becomes a writable file in the container's RW layer.
|
||||
|
||||
### Why a sidecar instead of an RO bind-mount
|
||||
|
||||
v0.5.0 originally mounted both files directly at vLLM's site-packages paths
|
||||
with `:ro`. That broke 8 Genesis-loaded composes because Genesis P64
|
||||
(qwen3coder MTP streaming early-return), P68 (auto force tool_choice=required),
|
||||
and P69 (long-context tool-format reminder) all write hooks into
|
||||
`chat_completion/serving.py` at vllm-import time. The RO mount blocked those
|
||||
writes with `Errno 30: Read-only file system`, and Genesis explicitly warned
|
||||
"partial state risk; container should be torn down." Reported by @ygafarov in
|
||||
[#120](https://github.com/noonghunna/club-3090/issues/120#issuecomment-4443236686);
|
||||
sidecar pattern shipped in v0.5.1.
|
||||
|
||||
The sidecar resolves it cleanly:
|
||||
- Our patched files land in the container's RW layer (not bind-mounted RO)
|
||||
- Genesis can write its hooks on top freely
|
||||
- Host patches dir stays canonical (no Genesis hooks bleeding into our git tree)
|
||||
- Each container restart starts fresh; Genesis re-applies on a clean copy
|
||||
|
||||
If a future patch to our overlay also touches `chat_completion/serving.py`
|
||||
(e.g. when PR #35936's streaming hunks land for a nightly we pin to), the
|
||||
sidecar layout supports it without further changes — our patched content
|
||||
sits on disk in the install.sh source path, gets installed before Genesis
|
||||
runs, Genesis layers its hooks on top.
|
||||
|
||||
## Validation
|
||||
|
||||
@@ -68,6 +106,12 @@ End-to-end checked 2026-05-12:
|
||||
agent completes loop with non-zero steps (was "No action returned after
|
||||
3 attempts" pre-overlay).
|
||||
|
||||
Sidecar pattern validated 2026-05-13 on `single/long-text.yml` (Genesis-loaded):
|
||||
- `install.sh` runs before Genesis: "chat_completion/serving.py installed from /etc/club3090/..."
|
||||
- Genesis P64 succeeds: "P64 applied: 2 files modified, 0 idempotent" (was failing with `Read-only file system` pre-v0.5.1)
|
||||
- Genesis P68/P69 succeed: "Hook injected into create_chat_completion" (was failing pre-v0.5.1)
|
||||
- Zero `Read-only file system` errors anywhere in the boot log.
|
||||
|
||||
## Drop Trigger
|
||||
|
||||
Remove this overlay when vLLM PR #35936, or an equivalent fix, is merged
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install PR #35936 required-tool fallback overlay into vLLM's site-packages.
|
||||
#
|
||||
# WHY a sidecar instead of an RO bind-mount:
|
||||
# Genesis P64 (qwen3coder MTP streaming early-return), P68 (auto-force
|
||||
# tool_choice=required), and P69 (long-context tool-format reminder) all
|
||||
# write hooks INTO chat_completion/serving.py at vllm-import time. An RO
|
||||
# bind-mount blocks those writes with `Errno 30: Read-only file system`,
|
||||
# and Genesis explicitly warns "partial state risk; container should be
|
||||
# torn down."
|
||||
#
|
||||
# Sidecar copies our files into the container's RW layer BEFORE Genesis
|
||||
# runs, so Genesis can write its hooks freely on top. Same pattern we
|
||||
# used for `patch_tolist_cudagraph.py` before Genesis absorbed it.
|
||||
#
|
||||
# Today chat_completion/serving.py is byte-identical to the upstream
|
||||
# nightly (PR #35936's streaming hunks don't apply on our pin), so
|
||||
# Genesis writes its hooks onto a clean copy. Slot stays future-ready
|
||||
# for when streaming hunks land — our changes will sit underneath
|
||||
# Genesis writes.
|
||||
#
|
||||
# Idempotent: cp overwrites unconditionally (per-boot fresh copy is
|
||||
# correct behaviour — Genesis re-applies on each fresh container).
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
SRC_CC="${CLUB3090_PR35936_CHAT_COMPLETION_SRC:-/etc/club3090/pr35936-chat-completion-serving.py}"
|
||||
DST_CC="/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py"
|
||||
|
||||
SRC_ENGINE="${CLUB3090_PR35936_ENGINE_SRC:-/etc/club3090/pr35936-engine-serving.py}"
|
||||
DST_ENGINE="/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py"
|
||||
|
||||
if [ -r "$SRC_CC" ]; then
|
||||
cp "$SRC_CC" "$DST_CC"
|
||||
echo "[club3090/pr35936] chat_completion/serving.py installed from $SRC_CC" >&2
|
||||
else
|
||||
echo "[club3090/pr35936] WARN: $SRC_CC not found; chat_completion/serving.py left untouched" >&2
|
||||
fi
|
||||
|
||||
if [ -r "$SRC_ENGINE" ]; then
|
||||
cp "$SRC_ENGINE" "$DST_ENGINE"
|
||||
echo "[club3090/pr35936] engine/serving.py installed from $SRC_ENGINE" >&2
|
||||
else
|
||||
echo "[club3090/pr35936] WARN: $SRC_ENGINE not found; engine/serving.py left untouched (PR #35936 fix INACTIVE)" >&2
|
||||
fi
|
||||
@@ -0,0 +1,416 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# Formatter for one-row BENCHMARKS.md submissions from results/rebench/<tag>/.
|
||||
#
|
||||
# Public functions:
|
||||
# bench_row_format <rebench-tag-dir>
|
||||
# bench_row_section <rebench-tag-dir>
|
||||
# bench_row_fixtures
|
||||
|
||||
_BENCH_ROW_LIB_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
_BENCH_ROW_ROOT="$(cd -- "${_BENCH_ROW_LIB_DIR}/../.." && pwd)"
|
||||
|
||||
bench_row_fixtures() {
|
||||
local tag
|
||||
for tag in \
|
||||
qwen-int8-pth-n4-2026-05-10 \
|
||||
qwen-bf16-n4-2026-05-11 \
|
||||
qwen-int8-tq3-n3-2026-05-11 \
|
||||
qwen-tq3-mtp-genesis-2026-05-11 \
|
||||
gemma-int8-pth-n4-2026-05-11 \
|
||||
gemma-bf16-n4-2026-05-11; do
|
||||
if [[ -d "${_BENCH_ROW_ROOT}/results/rebench/${tag}" ]]; then
|
||||
printf '%s\n' "${_BENCH_ROW_ROOT}/results/rebench/${tag}"
|
||||
fi
|
||||
done
|
||||
}
|
||||
|
||||
bench_row_section() {
|
||||
_bench_row_python section "$1"
|
||||
}
|
||||
|
||||
bench_row_format() {
|
||||
_bench_row_python row "$1"
|
||||
}
|
||||
|
||||
bench_row_rig_shortname() {
|
||||
_bench_row_python rig-short "$1"
|
||||
}
|
||||
|
||||
_bench_row_python() {
|
||||
local mode="$1"
|
||||
local tag_dir="$2"
|
||||
BENCH_ROW_REPO_ROOT="${_BENCH_ROW_ROOT}" python3 - "$mode" "$tag_dir" <<'PY'
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
MODE = sys.argv[1]
|
||||
TAG_DIR = Path(sys.argv[2]).resolve()
|
||||
ROOT = Path(os.environ.get("BENCH_ROW_REPO_ROOT", ".")).resolve()
|
||||
|
||||
|
||||
def die(msg: str) -> None:
|
||||
print(f"[bench-row] ERROR: {msg}", file=sys.stderr)
|
||||
raise SystemExit(1)
|
||||
|
||||
|
||||
def read_text(path: Path) -> str:
|
||||
try:
|
||||
return path.read_text(errors="replace")
|
||||
except Exception:
|
||||
return ""
|
||||
|
||||
|
||||
def read_json(path: Path) -> Any:
|
||||
try:
|
||||
return json.loads(path.read_text(errors="replace"))
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def require_file(name: str) -> Path:
|
||||
path = TAG_DIR / name
|
||||
if not path.is_file():
|
||||
die(f"missing required artifact: {path}")
|
||||
return path
|
||||
|
||||
|
||||
def first_container(blob: Any) -> dict[str, Any]:
|
||||
if isinstance(blob, list) and blob:
|
||||
return blob[0] if isinstance(blob[0], dict) else {}
|
||||
return blob if isinstance(blob, dict) else {}
|
||||
|
||||
|
||||
def flag(cmd: list[str], name: str) -> str:
|
||||
try:
|
||||
i = cmd.index(name)
|
||||
return str(cmd[i + 1])
|
||||
except Exception:
|
||||
return "?"
|
||||
|
||||
|
||||
def rel(path: str) -> str:
|
||||
if not path:
|
||||
return ""
|
||||
p = Path(path)
|
||||
try:
|
||||
return str(p.resolve().relative_to(ROOT))
|
||||
except Exception:
|
||||
return str(p)
|
||||
|
||||
|
||||
def infer_compose_path(container_name: str, served: str, tp: str) -> str:
|
||||
name = container_name.lstrip("/")
|
||||
is_gemma = "gemma" in name or "gemma" in served
|
||||
model_root = "models/gemma-4-31b/vllm/compose" if is_gemma else "models/qwen3.6-27b/vllm/compose"
|
||||
|
||||
mapping = {
|
||||
"dual-int8-tq3": "dual/int8-tq3.yml",
|
||||
"dual-tq3-mtp-genesis": "dual/tq3-mtp-genesis.yml",
|
||||
"dual-tq3-nomtp": "dual/tq3-nomtp.yml",
|
||||
"dual-tq3-mtp": "dual/tq3-mtp.yml",
|
||||
"dual-int8": "dual/int8.yml",
|
||||
"dual-bf16": "dual/bf16.yml",
|
||||
"dual-dflash-noviz": "dual/dflash-noviz.yml",
|
||||
"dual-dflash": "dual/dflash.yml",
|
||||
"dual-turbo": "dual/turbo.yml",
|
||||
"dual": "dual/docker-compose.yml",
|
||||
"minimal": "single/minimal.yml",
|
||||
"tools-text": "single/tools-text.yml",
|
||||
"long-text-no-mtp": "single/long-text-no-mtp.yml",
|
||||
"long-text": "single/long-text.yml",
|
||||
"long-vision": "single/long-vision.yml",
|
||||
}
|
||||
for needle, suffix in mapping.items():
|
||||
if needle in name:
|
||||
return f"{model_root}/{suffix}"
|
||||
if tp == "4":
|
||||
return "models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml"
|
||||
if tp == "2":
|
||||
return f"{model_root}/dual/docker-compose.yml"
|
||||
return f"{model_root}/single/docker-compose.yml"
|
||||
|
||||
|
||||
def compose_display(compose_path: str, served: str) -> str:
|
||||
path = compose_path.replace("\\", "/")
|
||||
parts = path.split("/")
|
||||
base = parts[-1] if parts else path
|
||||
parent = parts[-2] if len(parts) >= 2 else ""
|
||||
is_gemma = "gemma" in served or "gemma-4-31b" in path
|
||||
|
||||
if base == "docker-compose.yml":
|
||||
if parent == "dual":
|
||||
return "dual.yml"
|
||||
if parent == "multi4":
|
||||
return "dual4.yml"
|
||||
if parent == "single":
|
||||
return "vllm/gemma-mtp-tp1" if is_gemma else "vllm/default"
|
||||
if parent in {"dual", "multi4"} and not base.startswith(f"{parent}-"):
|
||||
return f"{parent}-{base}"
|
||||
return base
|
||||
|
||||
|
||||
def parse_rig(rig_txt: str) -> dict[str, Any]:
|
||||
out: dict[str, Any] = {"gpus": []}
|
||||
for raw in rig_txt.splitlines():
|
||||
line = raw.strip()
|
||||
if not line:
|
||||
continue
|
||||
if line.startswith("hostname:"):
|
||||
out["hostname"] = line.split(":", 1)[1].strip()
|
||||
elif line.startswith("GPU "):
|
||||
gpu = line.split(":", 1)[1].split("(UUID", 1)[0].strip()
|
||||
out["gpus"].append(gpu)
|
||||
elif line.startswith("power_cap_w:"):
|
||||
out["power_cap_w"] = line.split(":", 1)[1].strip()
|
||||
return out
|
||||
|
||||
|
||||
def simplify_gpu(name: str) -> str:
|
||||
name = re.sub(r"^NVIDIA\s+", "", name)
|
||||
name = re.sub(r"^GeForce\s+", "", name)
|
||||
name = re.sub(r"^RTX\s+", "", name)
|
||||
return name.strip()
|
||||
|
||||
|
||||
def rig_shape(rig: dict[str, Any]) -> str:
|
||||
gpus = [simplify_gpu(g) for g in rig.get("gpus") or []]
|
||||
if not gpus:
|
||||
shape = "rig"
|
||||
elif len(set(gpus)) == 1:
|
||||
shape = f"{len(gpus)}× {gpus[0]}"
|
||||
else:
|
||||
shape = " + ".join(gpus)
|
||||
power = str(rig.get("power_cap_w") or "").strip()
|
||||
if power:
|
||||
try:
|
||||
power = f"{float(power):.0f} W/card"
|
||||
except Exception:
|
||||
power = f"{power} W/card"
|
||||
return f"{shape}, {power}"
|
||||
return shape
|
||||
|
||||
|
||||
def rig_cell(rig: dict[str, Any]) -> str:
|
||||
user = os.environ.get("BENCH_ROW_GITHUB_USER", "").strip().lstrip("@") or "your-handle"
|
||||
return f"@{user} ({rig_shape(rig)})"
|
||||
|
||||
|
||||
def short_date(tag: str, report: str) -> str:
|
||||
m = re.search(r"(20\d{2}-\d{2}-\d{2})", tag)
|
||||
if m:
|
||||
return m.group(1)
|
||||
m = re.search(r"\*\*Date:\*\*\s*(20\d{2}-\d{2}-\d{2})", report)
|
||||
return m.group(1) if m else "—"
|
||||
|
||||
|
||||
def kv_display(raw: str, served: str) -> str:
|
||||
raw = (raw or "?").strip("`")
|
||||
lowered = raw.lower()
|
||||
if lowered in {"turboquant_3bit_nc", "tq3"}:
|
||||
return "TQ3"
|
||||
if lowered in {"fp8_e5m2", "fp8", "fp8_e4m3"}:
|
||||
return "fp8"
|
||||
if lowered in {"bfloat16", "bf16"}:
|
||||
return "bf16"
|
||||
if lowered == "auto":
|
||||
return "bf16"
|
||||
if lowered == "int8_per_token_head":
|
||||
return "int8_per_token_head"
|
||||
return raw or "?"
|
||||
|
||||
|
||||
def fmt_ctx(value: str) -> str:
|
||||
try:
|
||||
n = int(str(value).replace(",", ""))
|
||||
except Exception:
|
||||
return str(value or "?")
|
||||
if n >= 1000:
|
||||
return f"{round(n / 1000):.0f}K"
|
||||
return str(n)
|
||||
|
||||
|
||||
def fmt_tps(value: Any) -> str:
|
||||
try:
|
||||
return f"{float(value):.2f}"
|
||||
except Exception:
|
||||
return "?"
|
||||
|
||||
|
||||
def parse_mib(text: str) -> int | None:
|
||||
m = re.search(r"([\d.]+)\s*MiB", str(text))
|
||||
if not m:
|
||||
return None
|
||||
try:
|
||||
return int(float(m.group(1)))
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def peak_vram(internal: dict[str, Any], gpu_count: int) -> str:
|
||||
vals: list[int] = []
|
||||
for g in (((internal.get("bench") or {}).get("gpu_state")) or []):
|
||||
mib = parse_mib(g.get("mem_used", ""))
|
||||
if mib is not None:
|
||||
vals.append(mib)
|
||||
if not vals:
|
||||
return "TBD"
|
||||
suffix = "/card" if gpu_count > 1 else ""
|
||||
return f"{max(vals) / 1024:.1f} GB{suffix}"
|
||||
|
||||
|
||||
def parse_jsonish(value: str) -> dict[str, Any]:
|
||||
try:
|
||||
parsed = json.loads(value)
|
||||
return parsed if isinstance(parsed, dict) else {}
|
||||
except Exception:
|
||||
return {}
|
||||
|
||||
|
||||
def verify_summary(report: str) -> str:
|
||||
m = re.search(r"Verify-stress:\s+\*\*(\d+/\d+)\*\*", report)
|
||||
if m:
|
||||
return m.group(1)
|
||||
m = re.search(r"\*\*Overall:\*\*\s+(PASS|FAIL|\?)", report)
|
||||
return m.group(1) if m else "?"
|
||||
|
||||
|
||||
def soak_note(soak: dict[str, Any]) -> str:
|
||||
verdict = str(soak.get("verdict") or "").upper()
|
||||
silent = str(soak.get("silent_empty") or "")
|
||||
growth = str(soak.get("max_growth_mib") or "")
|
||||
if not verdict:
|
||||
return "Soak: —"
|
||||
if verdict == "PASS" and (silent.startswith("0 ") or silent.startswith("0/") or silent == "0"):
|
||||
return "Soak: ✓ PASS"
|
||||
if verdict == "PASS":
|
||||
return f"Soak: ⚠ borderline ({silent or growth})"
|
||||
return f"Soak: ✗ {verdict}"
|
||||
|
||||
|
||||
def section_name(compose_path: str, served: str, tp: str, container: str) -> str:
|
||||
if "gemma" in served or "gemma" in container or "gemma-4-31b" in compose_path:
|
||||
return "Gemma 4 31B (community-experimental)"
|
||||
path = compose_path.replace("\\", "/")
|
||||
if "llama-cpp" in path or "llama-cpp" in container:
|
||||
return "Single-card (1× RTX 3090) — llama.cpp"
|
||||
if "/multi4/" in path or tp == "4":
|
||||
return "Quad-card (4× RTX 3090, TP=4)"
|
||||
if "/dual/" in path or tp == "2":
|
||||
return "Dual-card (2× RTX 3090, TP=2)"
|
||||
return "Single-card (1× RTX 3090) — vLLM"
|
||||
|
||||
|
||||
def load() -> dict[str, Any]:
|
||||
if not TAG_DIR.is_dir():
|
||||
die(f"tag dir not found: {TAG_DIR}")
|
||||
internal = read_json(require_file("_internal.json"))
|
||||
if not isinstance(internal, dict):
|
||||
die(f"invalid JSON artifact: {TAG_DIR / '_internal.json'}")
|
||||
report = read_text(require_file("REPORT.md"))
|
||||
config_blob = first_container(read_json(require_file("container-config.json")))
|
||||
cfg = config_blob.get("Config") or {}
|
||||
labels = cfg.get("Labels") or {}
|
||||
cmd = cfg.get("Cmd") or []
|
||||
env = cfg.get("Env") or []
|
||||
name = str(config_blob.get("Name") or "").lstrip("/")
|
||||
served = flag(cmd, "--served-model-name")
|
||||
tp = flag(cmd, "--tensor-parallel-size")
|
||||
compose = rel(labels.get("com.docker.compose.project.config_files") or "")
|
||||
if not compose:
|
||||
compose = infer_compose_path(name, served, tp)
|
||||
rig = parse_rig(read_text(require_file("rig.txt")))
|
||||
tag = TAG_DIR.name
|
||||
spec = parse_jsonish(flag(cmd, "--speculative-config"))
|
||||
|
||||
return {
|
||||
"tag": tag,
|
||||
"report": report,
|
||||
"internal": internal,
|
||||
"container": {
|
||||
"name": name,
|
||||
"served": served,
|
||||
"tp": tp,
|
||||
"compose": compose,
|
||||
"kv": flag(cmd, "--kv-cache-dtype"),
|
||||
"max_ctx": flag(cmd, "--max-model-len"),
|
||||
"max_num_seqs": flag(cmd, "--max-num-seqs"),
|
||||
"mem_util": flag(cmd, "--gpu-memory-utilization"),
|
||||
"image": cfg.get("Image", ""),
|
||||
"spec": spec,
|
||||
"genesis": any(str(e).startswith("GENESIS_") for e in env),
|
||||
},
|
||||
"rig": rig,
|
||||
"date": short_date(tag, report),
|
||||
}
|
||||
|
||||
|
||||
def format_row(data: dict[str, Any]) -> str:
|
||||
c = data["container"]
|
||||
internal = data["internal"]
|
||||
bench = internal.get("bench") or {}
|
||||
narrative = bench.get("narrative") or {}
|
||||
code = bench.get("code") or {}
|
||||
mtp = bench.get("mtp") or {}
|
||||
quality = internal.get("quality") or {}
|
||||
aider = internal.get("aider") or {}
|
||||
soak = internal.get("soak") or {}
|
||||
rig = data["rig"]
|
||||
gpu_count = max(len(rig.get("gpus") or []), 1)
|
||||
section = section_name(c["compose"], c["served"], c["tp"], c["name"])
|
||||
verify = verify_summary(data["report"])
|
||||
compose = compose_display(c["compose"], c["served"])
|
||||
kv = kv_display(c["kv"], c["served"])
|
||||
max_ctx = fmt_ctx(c["max_ctx"])
|
||||
tps = f"**{fmt_tps(narrative.get('wall_tps_mean'))} / {fmt_tps(code.get('wall_tps_mean'))}**"
|
||||
peak = peak_vram(internal, gpu_count)
|
||||
spec_n = c["spec"].get("num_speculative_tokens")
|
||||
notes = [soak_note(soak), f"verify-stress {verify}"]
|
||||
if quality.get("total_passed") is not None:
|
||||
notes.append(f"quality {quality.get('total_passed')}/{quality.get('total_total')}")
|
||||
if aider.get("total_count"):
|
||||
notes.append(f"aider {aider.get('passed_count')}/{aider.get('total_count')}")
|
||||
if mtp.get("mean_accept_length") is not None:
|
||||
n_text = f" n={spec_n}" if spec_n is not None else ""
|
||||
notes.append(
|
||||
f"MTP{n_text} AL {float(mtp['mean_accept_length']):.2f}, "
|
||||
f"accept {float(mtp.get('avg_accept_rate', 0)):.1f}%"
|
||||
)
|
||||
if c.get("genesis"):
|
||||
notes.append("Genesis on")
|
||||
|
||||
note_cell = "; ".join(notes) + f". Report: `results/rebench/{data['tag']}/REPORT.md`."
|
||||
|
||||
if section == "Gemma 4 31B (community-experimental)":
|
||||
al = f"{float(mtp['mean_accept_length']):.2f}" if mtp.get("mean_accept_length") is not None else "—"
|
||||
per_pos = str(mtp.get("per_position") or "—")
|
||||
return (
|
||||
f"| `{compose}` | {rig_cell(rig)} | {kv} | {max_ctx} | {tps} | "
|
||||
f"{al} | {per_pos} | {peak} | {data['date']} | {note_cell} |"
|
||||
)
|
||||
|
||||
return (
|
||||
f"| `{compose}` | {rig_cell(rig)} | {kv} | {max_ctx} | {tps} | "
|
||||
f"{peak} | {data['date']} | {note_cell} |"
|
||||
)
|
||||
|
||||
|
||||
data = load()
|
||||
if MODE == "section":
|
||||
c = data["container"]
|
||||
print(section_name(c["compose"], c["served"], c["tp"], c["name"]))
|
||||
elif MODE == "row":
|
||||
print(format_row(data))
|
||||
elif MODE == "rig-short":
|
||||
print(data["tag"])
|
||||
else:
|
||||
die(f"unknown mode: {MODE}")
|
||||
PY
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# Tiny parser for hardware metadata stored as compose header comments.
|
||||
#
|
||||
# Expected form:
|
||||
# # Requires-min-vram-gb: 24
|
||||
# # Requires-min-gpu-count: 2
|
||||
# # Tensor-parallel: 2
|
||||
# # Requires-sm: 9.0+
|
||||
#
|
||||
# This intentionally does not parse YAML. These fields are comments so that
|
||||
# older docker compose versions and direct `docker compose -f ... up` flows keep
|
||||
# working unchanged.
|
||||
|
||||
_compose_meta_trim() {
|
||||
local value="$1"
|
||||
value="${value#"${value%%[![:space:]]*}"}"
|
||||
value="${value%"${value##*[![:space:]]}"}"
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
_compose_meta_norm_key() {
|
||||
local key="$1"
|
||||
key="$(_compose_meta_trim "$key")"
|
||||
key="${key//_/-}"
|
||||
key="${key// /-}"
|
||||
printf '%s' "$key" | tr '[:upper:]' '[:lower:]'
|
||||
}
|
||||
|
||||
_compose_meta_wants_key() {
|
||||
local requested="$(_compose_meta_norm_key "$1")"
|
||||
local candidate="$(_compose_meta_norm_key "$2")"
|
||||
|
||||
case "$requested" in
|
||||
min-vram-gb) requested="requires-min-vram-gb" ;;
|
||||
min-gpu-count) requested="requires-min-gpu-count" ;;
|
||||
tp) requested="tensor-parallel" ;;
|
||||
sm) requested="requires-sm" ;;
|
||||
esac
|
||||
|
||||
[[ "$candidate" == "$requested" ]]
|
||||
}
|
||||
|
||||
compose_meta_get() {
|
||||
local compose_file="$1"
|
||||
local field="$2"
|
||||
|
||||
[[ -f "$compose_file" ]] || return 1
|
||||
|
||||
local line key value
|
||||
while IFS= read -r line; do
|
||||
[[ "$line" =~ ^[[:space:]]*# ]] || continue
|
||||
line="${line#*\#}"
|
||||
[[ "$line" == *:* ]] || continue
|
||||
key="${line%%:*}"
|
||||
value="${line#*:}"
|
||||
if _compose_meta_wants_key "$field" "$key"; then
|
||||
_compose_meta_trim "$value"
|
||||
return 0
|
||||
fi
|
||||
done < "$compose_file"
|
||||
|
||||
return 1
|
||||
}
|
||||
+315
-3
@@ -12,6 +12,7 @@
|
||||
# preflight_running — warn if a club-3090 container is already up
|
||||
# preflight_genesis_pin — warn if on-disk Genesis tree differs from setup.sh's pin
|
||||
# preflight_repo_drift — warn if local HEAD is behind origin/master
|
||||
# preflight_compose_hardware— check compose VRAM/GPU-count/SM metadata
|
||||
#
|
||||
# Style: each function prints one or more "[preflight] ..." lines.
|
||||
# Hard failures get a one-line ERROR + a "Fix:" hint.
|
||||
@@ -19,6 +20,12 @@
|
||||
# Avoid double-sourcing.
|
||||
[[ -n "${_PREFLIGHT_LOADED:-}" ]] && return 0
|
||||
_PREFLIGHT_LOADED=1
|
||||
_PREFLIGHT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
if [[ -f "${_PREFLIGHT_DIR}/lib/compose-meta.sh" ]]; then
|
||||
# shellcheck source=lib/compose-meta.sh
|
||||
source "${_PREFLIGHT_DIR}/lib/compose-meta.sh"
|
||||
fi
|
||||
|
||||
preflight_docker() {
|
||||
if ! command -v docker >/dev/null 2>&1; then
|
||||
@@ -72,6 +79,301 @@ preflight_gpu() {
|
||||
return 0
|
||||
}
|
||||
|
||||
_preflight_trim() {
|
||||
local value="$1"
|
||||
value="${value#"${value%%[![:space:]]*}"}"
|
||||
value="${value%"${value##*[![:space:]]}"}"
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
_preflight_csv_token() {
|
||||
local value="$1"
|
||||
value="$(_preflight_trim "$value")"
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
_preflight_selector() {
|
||||
if [[ -n "${CLUB3090_GPU:-}" ]]; then
|
||||
printf '%s' "${CLUB3090_GPU}"
|
||||
elif [[ -n "${NVIDIA_VISIBLE_DEVICES:-}" && "${NVIDIA_VISIBLE_DEVICES}" != "all" && "${NVIDIA_VISIBLE_DEVICES}" != "void" ]]; then
|
||||
printf '%s' "${NVIDIA_VISIBLE_DEVICES}"
|
||||
elif [[ -n "${CUDA_VISIBLE_DEVICES:-}" && "${CUDA_VISIBLE_DEVICES}" != "all" && "${CUDA_VISIBLE_DEVICES}" != "void" ]]; then
|
||||
printf '%s' "${CUDA_VISIBLE_DEVICES}"
|
||||
fi
|
||||
}
|
||||
|
||||
_preflight_selector_is_specific() {
|
||||
local selector="${1:-}"
|
||||
[[ -n "$selector" && "$selector" != "all" && "$selector" != "void" ]]
|
||||
}
|
||||
|
||||
_preflight_selector_allows_index() {
|
||||
local selector="$1"
|
||||
local idx="$2"
|
||||
local token
|
||||
|
||||
if ! _preflight_selector_is_specific "$selector"; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
IFS=',' read -ra _preflight_selector_tokens <<< "$selector"
|
||||
for token in "${_preflight_selector_tokens[@]}"; do
|
||||
token="$(_preflight_trim "$token")"
|
||||
[[ "$token" == "$idx" ]] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
_preflight_selector_first_numeric() {
|
||||
local selector="$1"
|
||||
local token
|
||||
|
||||
IFS=',' read -ra _preflight_selector_tokens <<< "$selector"
|
||||
for token in "${_preflight_selector_tokens[@]}"; do
|
||||
token="$(_preflight_trim "$token")"
|
||||
if [[ "$token" =~ ^[0-9]+$ ]]; then
|
||||
printf '%s' "$token"
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
_preflight_sm_to_int() {
|
||||
local sm="$1"
|
||||
sm="${sm%%+}"
|
||||
sm="${sm//sm_/}"
|
||||
sm="${sm//SM_/}"
|
||||
sm="${sm// /}"
|
||||
[[ -z "$sm" ]] && { echo 0; return; }
|
||||
|
||||
local major minor
|
||||
if [[ "$sm" == *.* ]]; then
|
||||
major="${sm%%.*}"
|
||||
minor="${sm#*.}"
|
||||
else
|
||||
major="$sm"
|
||||
minor="0"
|
||||
fi
|
||||
major="${major//[^0-9]/}"
|
||||
minor="${minor//[^0-9]/}"
|
||||
[[ -z "$major" ]] && major=0
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
if [[ "${#minor}" -eq 1 ]]; then
|
||||
minor=$(( minor * 10 ))
|
||||
else
|
||||
minor="${minor:0:2}"
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
fi
|
||||
echo $(( major * 100 + minor ))
|
||||
}
|
||||
|
||||
_preflight_vram_gb() {
|
||||
local mib="$1"
|
||||
echo $(( (mib + 1023) / 1024 ))
|
||||
}
|
||||
|
||||
_preflight_hardware_suggestions() {
|
||||
local variant="${1:-}"
|
||||
|
||||
echo "[preflight]" >&2
|
||||
echo "[preflight] Suggested next steps:" >&2
|
||||
echo "[preflight] - Pick a compose that matches the detected GPU VRAM/topology." >&2
|
||||
if [[ "$variant" == vllm/gemma-mtp-tp1 ]]; then
|
||||
echo "[preflight] - On 2x 24 GB cards, use: bash scripts/switch.sh vllm/gemma-mtp" >&2
|
||||
fi
|
||||
echo "[preflight] - On a single 24 GB card, start with: bash scripts/switch.sh vllm/default" >&2
|
||||
echo "[preflight] - For maximum compatibility, use: bash scripts/switch.sh llamacpp/default" >&2
|
||||
echo "[preflight] - Explicit bypass: bash scripts/switch.sh --force ${variant:-<variant>}" >&2
|
||||
}
|
||||
|
||||
# preflight_compose_hardware <compose_file> [variant] [force]
|
||||
#
|
||||
# Reads compose header metadata and checks the target host before docker compose
|
||||
# starts. This is intentionally conservative:
|
||||
# - Missing metadata warns and allows the boot.
|
||||
# - TP=1 composes auto-select the largest eligible GPU unless the user set
|
||||
# CLUB3090_GPU, CUDA_VISIBLE_DEVICES, or NVIDIA_VISIBLE_DEVICES.
|
||||
# - TP>=2 composes hard-fail only on insufficient GPU count or hard SM gates;
|
||||
# heterogeneous VRAM below the requested floor warns because advanced users
|
||||
# may be validating sub-24 GB configs with tuned memory-utilization.
|
||||
preflight_compose_hardware() {
|
||||
local compose_file="$1"
|
||||
local variant="${2:-}"
|
||||
local force="${3:-${FORCE:-0}}"
|
||||
|
||||
if [[ "${PREFLIGHT_NO_HARDWARE:-0}" == "1" ]]; then
|
||||
return 0
|
||||
fi
|
||||
if [[ "$force" == "1" || "${FORCE:-0}" == "1" ]]; then
|
||||
echo "[preflight] hardware: skipped (--force/FORCE=1)"
|
||||
return 0
|
||||
fi
|
||||
if [[ ! -f "$compose_file" ]]; then
|
||||
echo "[preflight] ERROR: compose file not found: $compose_file" >&2
|
||||
return 1
|
||||
fi
|
||||
if ! command -v nvidia-smi >/dev/null 2>&1; then
|
||||
echo "[preflight] WARN: nvidia-smi not found; skipping compose hardware metadata check." >&2
|
||||
return 0
|
||||
fi
|
||||
if ! declare -F compose_meta_get >/dev/null 2>&1; then
|
||||
echo "[preflight] WARN: compose metadata parser unavailable; skipping hardware metadata check." >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
local min_vram_gb min_gpu_count tp requires_sm
|
||||
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
|
||||
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
|
||||
tp="$(compose_meta_get "$compose_file" tensor-parallel || true)"
|
||||
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
|
||||
|
||||
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" || -z "$tp" ]]; then
|
||||
echo "[preflight] WARN: compose has no hardware metadata; allowing boot: $compose_file" >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
requires_sm="${requires_sm:-0.0}"
|
||||
local required_sm_int
|
||||
required_sm_int="$(_preflight_sm_to_int "$requires_sm")"
|
||||
|
||||
local gpu_query
|
||||
gpu_query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
if [[ -z "$gpu_query" ]]; then
|
||||
echo "[preflight] WARN: could not query GPU VRAM/SM via nvidia-smi; skipping hardware metadata check." >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
local selector
|
||||
selector="$(_preflight_selector || true)"
|
||||
|
||||
local total_count=0 selected_count=0 eligible_count=0 selected_below_vram=0 selected_below_sm=0
|
||||
local best_idx="" best_name="" best_mib=0 best_sm=""
|
||||
local first_idx="" first_name="" first_mib=0 first_sm=""
|
||||
local idx name mem_mib sm rest vram_gb sm_int
|
||||
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
name="$(_preflight_csv_token "$name")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
sm="$(_preflight_csv_token "$sm")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
total_count=$(( total_count + 1 ))
|
||||
_preflight_selector_allows_index "$selector" "$idx" || continue
|
||||
|
||||
selected_count=$(( selected_count + 1 ))
|
||||
if [[ -z "$first_idx" ]]; then
|
||||
first_idx="$idx"
|
||||
first_name="$name"
|
||||
first_mib="$mem_mib"
|
||||
first_sm="$sm"
|
||||
fi
|
||||
|
||||
vram_gb="$(_preflight_vram_gb "$mem_mib")"
|
||||
sm_int="$(_preflight_sm_to_int "$sm")"
|
||||
|
||||
if (( vram_gb < min_vram_gb )); then
|
||||
selected_below_vram=1
|
||||
fi
|
||||
if (( sm_int < required_sm_int )); then
|
||||
selected_below_sm=1
|
||||
fi
|
||||
|
||||
if (( vram_gb >= min_vram_gb && sm_int >= required_sm_int )); then
|
||||
eligible_count=$(( eligible_count + 1 ))
|
||||
if (( mem_mib > best_mib )); then
|
||||
best_idx="$idx"
|
||||
best_name="$name"
|
||||
best_mib="$mem_mib"
|
||||
best_sm="$sm"
|
||||
fi
|
||||
fi
|
||||
done <<< "$gpu_query"
|
||||
|
||||
if (( total_count == 0 )); then
|
||||
echo "[preflight] ERROR: no NVIDIA GPUs detected." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
if (( selected_count == 0 )); then
|
||||
echo "[preflight] ERROR: GPU selector '${selector}' did not match any detected GPU index." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
|
||||
local requires_sm_display="${requires_sm%%+}"
|
||||
local sm_label=""
|
||||
if (( required_sm_int > 0 )); then
|
||||
sm_label=", sm_${requires_sm_display}+"
|
||||
fi
|
||||
|
||||
if (( tp <= 1 )); then
|
||||
if _preflight_selector_is_specific "$selector"; then
|
||||
local first_vram_gb first_sm_int
|
||||
first_vram_gb="$(_preflight_vram_gb "$first_mib")"
|
||||
first_sm_int="$(_preflight_sm_to_int "$first_sm")"
|
||||
if (( first_vram_gb < min_vram_gb || first_sm_int < required_sm_int )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires one GPU with >=${min_vram_gb} GB VRAM${sm_label}." >&2
|
||||
echo "[preflight] Explicit selector '${selector}' starts with GPU ${first_idx}: ${first_name}, ${first_vram_gb} GB, sm_${first_sm}." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
export CLUB3090_GPU="${CLUB3090_GPU:-$selector}"
|
||||
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-$selector}"
|
||||
export NVIDIA_VISIBLE_DEVICES="${NVIDIA_VISIBLE_DEVICES:-$selector}"
|
||||
echo "[preflight] hardware: ${variant:-compose} TP=1 requires >=${min_vram_gb} GB${sm_label}; using explicit GPU ${first_idx} (${first_vram_gb} GB, sm_${first_sm})"
|
||||
return 0
|
||||
fi
|
||||
|
||||
if (( eligible_count == 0 )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires one GPU with >=${min_vram_gb} GB VRAM${sm_label}; none found." >&2
|
||||
echo "[preflight] Detected GPUs:" >&2
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
name="$(_preflight_csv_token "$name")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
sm="$(_preflight_csv_token "$sm")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
echo "[preflight] GPU ${idx}: ${name}, $(_preflight_vram_gb "$mem_mib") GB, sm_${sm}" >&2
|
||||
done <<< "$gpu_query"
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
|
||||
export CLUB3090_GPU="$best_idx"
|
||||
export CUDA_VISIBLE_DEVICES="$best_idx"
|
||||
export NVIDIA_VISIBLE_DEVICES="$best_idx"
|
||||
echo "[preflight] hardware: ${variant:-compose} TP=1 requires >=${min_vram_gb} GB${sm_label}; auto-selected GPU ${best_idx} ($(_preflight_vram_gb "$best_mib") GB, sm_${best_sm})"
|
||||
return 0
|
||||
fi
|
||||
|
||||
if (( selected_count < min_gpu_count )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires ${min_gpu_count} visible GPU(s) for TP=${tp}; found ${selected_count}." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
if (( selected_below_sm == 1 )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires sm_${requires_sm_display}+ on visible GPUs." >&2
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
name="$(_preflight_csv_token "$name")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
sm="$(_preflight_csv_token "$sm")"
|
||||
_preflight_selector_allows_index "$selector" "$idx" || continue
|
||||
echo "[preflight] GPU ${idx}: ${name}, $(_preflight_vram_gb "$mem_mib") GB, sm_${sm}" >&2
|
||||
done <<< "$gpu_query"
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
if (( selected_below_vram == 1 )); then
|
||||
echo "[preflight] WARN: ${variant:-compose} requires >=${min_vram_gb} GB per visible GPU for TP=${tp}, but at least one selected GPU is smaller." >&2
|
||||
echo "[preflight] Continuing because TP>=2 sub-24 GB rigs may use tuned gpu-memory-utilization/KV settings." >&2
|
||||
fi
|
||||
|
||||
echo "[preflight] hardware: ${variant:-compose} TP=${tp} requires ${min_gpu_count} GPU(s), >=${min_vram_gb} GB each${sm_label}; ${selected_count} visible GPU(s) detected"
|
||||
return 0
|
||||
}
|
||||
|
||||
preflight_disk() {
|
||||
local path="$1"
|
||||
local need_gb="$2"
|
||||
@@ -412,9 +714,19 @@ preflight_kv_format_hint() {
|
||||
return 0
|
||||
fi
|
||||
|
||||
# Detect smallest VRAM among visible cards (the TP-split ceiling).
|
||||
local min_vram_mib
|
||||
min_vram_mib="$(nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | sort -n | head -1)"
|
||||
# Detect smallest VRAM among selected/visible cards (the TP-split ceiling).
|
||||
local min_vram_mib="" mem_query selector idx mem_mib
|
||||
selector="$(_preflight_selector || true)"
|
||||
mem_query="$(nvidia-smi --query-gpu=index,memory.total --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
while IFS=',' read -r idx mem_mib; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
_preflight_selector_allows_index "$selector" "$idx" || continue
|
||||
if [[ -z "$min_vram_mib" || "$mem_mib" -lt "$min_vram_mib" ]]; then
|
||||
min_vram_mib="$mem_mib"
|
||||
fi
|
||||
done <<< "$mem_query"
|
||||
if [[ -z "$min_vram_mib" ]] || [[ "$min_vram_mib" -ge 24000 ]]; then
|
||||
return 0 # 24 GB+ cards — TQ3 is the right pick, no hint needed
|
||||
fi
|
||||
|
||||
@@ -273,3 +273,6 @@ echo " verify-stress: tail -5 $OUT_DIR/verify-stress.log"
|
||||
echo " quality: grep '^Quality:' $OUT_DIR/quality-full.log"
|
||||
echo " soak: grep -E 'verdict|silent_empty|p50_decode' $OUT_DIR/soak.log"
|
||||
echo " aider: grep 'aider-polyglot-30' $OUT_DIR/aider-polyglot.log"
|
||||
echo
|
||||
echo "To submit your numbers (review then PR):"
|
||||
echo " bash scripts/submit-bench.sh --tag $TAG"
|
||||
|
||||
Executable
+312
@@ -0,0 +1,312 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# Generate or submit a BENCHMARKS.md row from results/rebench/<tag>/.
|
||||
#
|
||||
# Usage:
|
||||
# bash scripts/submit-bench.sh --tag <tag>
|
||||
# bash scripts/submit-bench.sh --tag <tag> --auto-submit
|
||||
# bash scripts/submit-bench.sh --tag <tag> --auto-submit --as-pr
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
cd "$ROOT_DIR"
|
||||
|
||||
TAG=""
|
||||
AUTO_SUBMIT=0
|
||||
AS_PR=0
|
||||
SECTION_OVERRIDE=""
|
||||
|
||||
usage() {
|
||||
sed -n '2,18p' "$0" | sed 's/^# \{0,1\}//'
|
||||
exit 0
|
||||
}
|
||||
|
||||
die() {
|
||||
echo "[submit-bench] ERROR: $*" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
log() {
|
||||
echo "[submit-bench] $*"
|
||||
}
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--tag)
|
||||
[[ $# -ge 2 ]] || die "--tag requires a value"
|
||||
TAG="$2"
|
||||
shift 2
|
||||
;;
|
||||
--auto-submit)
|
||||
AUTO_SUBMIT=1
|
||||
shift
|
||||
;;
|
||||
--as-pr)
|
||||
AS_PR=1
|
||||
shift
|
||||
;;
|
||||
--section)
|
||||
[[ $# -ge 2 ]] || die "--section requires a value"
|
||||
SECTION_OVERRIDE="$2"
|
||||
shift 2
|
||||
;;
|
||||
-h|--help)
|
||||
usage
|
||||
;;
|
||||
*)
|
||||
die "unknown arg: $1"
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
[[ -n "$TAG" ]] || die "--tag <tag> is required"
|
||||
|
||||
TAG_DIR="results/rebench/${TAG}"
|
||||
[[ -d "$TAG_DIR" ]] || die "tag dir not found: ${TAG_DIR}"
|
||||
[[ -f "$TAG_DIR/REPORT.md" ]] || die "missing required artifact: ${TAG_DIR}/REPORT.md"
|
||||
[[ -f "$TAG_DIR/_internal.json" ]] || die "missing required artifact: ${TAG_DIR}/_internal.json"
|
||||
[[ -f "$TAG_DIR/container-config.json" ]] || die "missing required artifact: ${TAG_DIR}/container-config.json"
|
||||
[[ -f "$TAG_DIR/rig.txt" ]] || die "missing required artifact: ${TAG_DIR}/rig.txt"
|
||||
|
||||
# shellcheck source=lib/bench-row-formatter.sh
|
||||
source "$ROOT_DIR/scripts/lib/bench-row-formatter.sh"
|
||||
|
||||
github_user_for_row() {
|
||||
if [[ -n "${BENCH_ROW_GITHUB_USER:-}" ]]; then
|
||||
printf '%s' "${BENCH_ROW_GITHUB_USER#@}"
|
||||
return 0
|
||||
fi
|
||||
if [[ "${GH_MOCK:-0}" == "1" ]]; then
|
||||
printf '%s' "${GH_MOCK_USER:-mock-user}"
|
||||
return 0
|
||||
fi
|
||||
if command -v gh >/dev/null 2>&1 && gh auth status >/dev/null 2>&1; then
|
||||
gh api user --jq .login 2>/dev/null || true
|
||||
fi
|
||||
}
|
||||
|
||||
if [[ "$AUTO_SUBMIT" -eq 1 ]]; then
|
||||
GH_USER="$(github_user_for_row)"
|
||||
[[ -n "$GH_USER" ]] || die "not authed with gh. Run: gh auth login"
|
||||
export BENCH_ROW_GITHUB_USER="$GH_USER"
|
||||
fi
|
||||
|
||||
ROW="$(bench_row_format "$TAG_DIR")"
|
||||
SECTION="${SECTION_OVERRIDE:-$(bench_row_section "$TAG_DIR")}"
|
||||
OUTPUT="$TAG_DIR/BENCHMARKS-row.md"
|
||||
printf '%s\n' "$ROW" > "$OUTPUT"
|
||||
|
||||
log "Generated BENCHMARKS row for section: ${SECTION}"
|
||||
log "Wrote: ${OUTPUT}"
|
||||
echo
|
||||
|
||||
valid_sections() {
|
||||
rg -n '^(##|###) ' BENCHMARKS.md | sed 's/^/[submit-bench] /' >&2 || true
|
||||
}
|
||||
|
||||
write_pr_body() {
|
||||
local body_file="$1"
|
||||
local row="$2"
|
||||
local tag="$3"
|
||||
local template=".github/PULL_REQUEST_TEMPLATE/bench-row.md"
|
||||
|
||||
if [[ -f "$template" ]]; then
|
||||
python3 - "$template" "$body_file" "$tag" "$row" <<'PY'
|
||||
from pathlib import Path
|
||||
import sys
|
||||
|
||||
template, body_file, tag, row = sys.argv[1:5]
|
||||
text = Path(template).read_text()
|
||||
text = text.replace("<TAG>", tag)
|
||||
text = text.replace("<!-- The generated BENCHMARKS.md row goes here -->", row)
|
||||
text = text.replace(
|
||||
"<!-- Output of `bash scripts/report.sh` (redacted) -->",
|
||||
f"See `results/rebench/{tag}/rig.txt`.",
|
||||
)
|
||||
Path(body_file).write_text(text)
|
||||
PY
|
||||
else
|
||||
{
|
||||
echo "## Rig bench submission"
|
||||
echo
|
||||
echo "### New row"
|
||||
echo
|
||||
echo "$row"
|
||||
echo
|
||||
echo "### Full results"
|
||||
echo
|
||||
echo "See \`results/rebench/${tag}/REPORT.md\`."
|
||||
} > "$body_file"
|
||||
fi
|
||||
}
|
||||
|
||||
write_issue_body() {
|
||||
local body_file="$1"
|
||||
local row="$2"
|
||||
local tag="$3"
|
||||
local section="$4"
|
||||
|
||||
# The repo's numbers-from-your-rig issue template is a structured YAML form
|
||||
# with required textarea/dropdown fields. `gh issue create --template` opens
|
||||
# that interactive form shape, which is not useful once submit-bench has
|
||||
# already generated the structured report. Use a direct markdown body instead.
|
||||
{
|
||||
echo "**Compose / section**: \`${section}\`"
|
||||
echo
|
||||
echo "**Rig**:"
|
||||
echo
|
||||
echo '```text'
|
||||
cat "${TAG_DIR}/rig.txt"
|
||||
echo '```'
|
||||
echo
|
||||
echo "**Proposed BENCHMARKS.md row**:"
|
||||
echo
|
||||
echo "$row"
|
||||
echo
|
||||
echo "**Full report**: \`results/rebench/${tag}/REPORT.md\`"
|
||||
echo
|
||||
echo "**Generated row file**: \`results/rebench/${tag}/BENCHMARKS-row.md\`"
|
||||
} > "$body_file"
|
||||
}
|
||||
|
||||
insert_row() {
|
||||
local section="$1"
|
||||
local row="$2"
|
||||
python3 - "$section" "$row" <<'PY'
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
section = sys.argv[1]
|
||||
row = sys.argv[2]
|
||||
path = Path("BENCHMARKS.md")
|
||||
lines = path.read_text().splitlines()
|
||||
|
||||
heading_idx = None
|
||||
for i, line in enumerate(lines):
|
||||
if line.strip() in {f"## {section}", f"### {section}"}:
|
||||
heading_idx = i
|
||||
break
|
||||
if heading_idx is None:
|
||||
print(f"[submit-bench] ERROR: section not found in BENCHMARKS.md: {section}", file=sys.stderr)
|
||||
raise SystemExit(1)
|
||||
|
||||
table_start = None
|
||||
for i in range(heading_idx + 1, len(lines)):
|
||||
if lines[i].startswith("#"):
|
||||
break
|
||||
if lines[i].startswith("|"):
|
||||
table_start = i
|
||||
break
|
||||
if table_start is None:
|
||||
print(f"[submit-bench] ERROR: no markdown table found below section: {section}", file=sys.stderr)
|
||||
raise SystemExit(1)
|
||||
|
||||
insert_at = table_start
|
||||
for i in range(table_start, len(lines)):
|
||||
line = lines[i]
|
||||
if line.startswith("#"):
|
||||
break
|
||||
if line.startswith("|"):
|
||||
insert_at = i + 1
|
||||
continue
|
||||
if insert_at > table_start:
|
||||
break
|
||||
|
||||
lines.insert(insert_at, row)
|
||||
path.write_text("\n".join(lines) + "\n")
|
||||
PY
|
||||
}
|
||||
|
||||
if [[ "$AUTO_SUBMIT" -ne 1 ]]; then
|
||||
cat <<EOF
|
||||
Inspect at ${OUTPUT}. Three ways to land it (recommended order):
|
||||
|
||||
1. Issue + maintainer integrates (preferred — vetting before merge):
|
||||
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit
|
||||
(opens an issue via \`gh issue create\`)
|
||||
Or, no-gh-needed:
|
||||
https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml
|
||||
— paste the contents of ${OUTPUT} + ${TAG_DIR}/rig.txt into the body
|
||||
|
||||
2. Direct PR (advanced — for contributors who know the matrix structure):
|
||||
bash scripts/submit-bench.sh --tag ${TAG} --auto-submit --as-pr
|
||||
Note: matrix is hand-curated; direct PRs may get redirected to an
|
||||
issue thread for context-gathering before merge.
|
||||
|
||||
3. Manual edit (zero tools):
|
||||
Paste the row from ${OUTPUT} into BENCHMARKS.md via the GitHub web editor.
|
||||
EOF
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ -n "$SECTION_OVERRIDE" ]]; then
|
||||
if ! rg -q "^(##|###) ${SECTION_OVERRIDE//\//\\/}$" BENCHMARKS.md; then
|
||||
echo "[submit-bench] Known sections:" >&2
|
||||
valid_sections
|
||||
die "section override not found: ${SECTION_OVERRIDE}"
|
||||
fi
|
||||
fi
|
||||
|
||||
PR_TITLE="bench(matrix): @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
|
||||
ISSUE_TITLE="[bench] @${BENCH_ROW_GITHUB_USER} $(bench_row_rig_shortname "$TAG_DIR")"
|
||||
BRANCH_USER="$(printf '%s' "${BENCH_ROW_GITHUB_USER}" | tr -cd '[:alnum:]_.-')"
|
||||
BRANCH_TAG="$(printf '%s' "${TAG}" | tr -cd '[:alnum:]_.-')"
|
||||
BRANCH="bench/${BRANCH_USER}-${BRANCH_TAG}"
|
||||
if [[ "$AS_PR" -eq 1 ]]; then
|
||||
BODY_FILE="$TAG_DIR/PR-body.md"
|
||||
write_pr_body "$BODY_FILE" "$ROW" "$TAG"
|
||||
else
|
||||
BODY_FILE="$TAG_DIR/ISSUE-body.md"
|
||||
write_issue_body "$BODY_FILE" "$ROW" "$TAG" "$SECTION"
|
||||
fi
|
||||
|
||||
if [[ "${GH_MOCK:-0}" == "1" ]]; then
|
||||
MOCK_LOG="$TAG_DIR/auto-submit-mock.log"
|
||||
if [[ "$AS_PR" -eq 1 ]]; then
|
||||
{
|
||||
echo "git switch -c ${BRANCH}"
|
||||
echo "insert BENCHMARKS.md row under: ${SECTION}"
|
||||
echo "git commit -m ${PR_TITLE}"
|
||||
echo "git push -u origin ${BRANCH}"
|
||||
echo "gh pr create --title ${PR_TITLE} --body-file ${BODY_FILE}"
|
||||
} > "$MOCK_LOG"
|
||||
log "GH_MOCK=1 — wrote mocked PR auto-submit commands: ${MOCK_LOG}"
|
||||
log "PR title: ${PR_TITLE}"
|
||||
else
|
||||
{
|
||||
echo "gh issue create --title ${ISSUE_TITLE} --body-file ${BODY_FILE} --label bench-contribution"
|
||||
} > "$MOCK_LOG"
|
||||
log "GH_MOCK=1 — wrote mocked issue auto-submit command: ${MOCK_LOG}"
|
||||
log "Issue title: ${ISSUE_TITLE}"
|
||||
fi
|
||||
exit 0
|
||||
fi
|
||||
|
||||
command -v gh >/dev/null 2>&1 || die "'gh' not found. Install GitHub CLI or submit manually."
|
||||
gh auth status >/dev/null 2>&1 || die "not authed with gh. Run: gh auth login"
|
||||
|
||||
if [[ "$AS_PR" -ne 1 ]]; then
|
||||
ISSUE_URL="$(gh issue create --title "$ISSUE_TITLE" --body-file "$BODY_FILE" --label bench-contribution)"
|
||||
log "Opened issue: ${ISSUE_URL}"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if ! git diff --quiet -- BENCHMARKS.md; then
|
||||
die "BENCHMARKS.md already has local edits; commit/stash them before --auto-submit"
|
||||
fi
|
||||
|
||||
git fetch origin master >/dev/null 2>&1 || log "WARN: git fetch origin master failed; continuing from current branch"
|
||||
if git show-ref --verify --quiet refs/remotes/origin/master; then
|
||||
git switch -c "$BRANCH" origin/master
|
||||
else
|
||||
git switch -c "$BRANCH"
|
||||
fi
|
||||
insert_row "$SECTION" "$ROW"
|
||||
git add BENCHMARKS.md
|
||||
git commit -m "$PR_TITLE"
|
||||
git push -u origin "$BRANCH"
|
||||
PR_URL="$(gh pr create --title "$PR_TITLE" --body-file "$BODY_FILE")"
|
||||
log "Opened PR: ${PR_URL}"
|
||||
+35
-13
@@ -9,6 +9,7 @@
|
||||
# Usage:
|
||||
# bash scripts/switch.sh <variant> # switch + tail until ready
|
||||
# bash scripts/switch.sh <variant> --no-wait # switch and return immediately
|
||||
# bash scripts/switch.sh --force <variant> # skip hardware/free-VRAM preflight
|
||||
# bash scripts/switch.sh --list # show all variants
|
||||
# bash scripts/switch.sh --down # just bring down whatever's up
|
||||
#
|
||||
@@ -42,6 +43,8 @@
|
||||
#
|
||||
# Env overrides (rarely needed):
|
||||
# COMPOSE_BIN Default: "docker compose" (set to e.g. "podman compose" if needed)
|
||||
# CLUB3090_GPU Single-card GPU index override, e.g. "1" on a hetero rig
|
||||
# FORCE Set to 1 to skip hardware/free-VRAM preflight
|
||||
# READY_URL Default: http://localhost:8020/v1/models
|
||||
# READY_TIMEOUT Default: 600 (seconds — longer for cold cudagraph capture)
|
||||
|
||||
@@ -169,15 +172,23 @@ gpu_preflight() {
|
||||
return
|
||||
fi
|
||||
# Free MiB per GPU. Tolerate small overhead (driver, X server) — abort
|
||||
# if any GPU has <80% of its total memory free.
|
||||
# if any selected GPU has <80% of its total memory free.
|
||||
local mem_query
|
||||
mem_query=$(nvidia-smi --query-gpu=index,memory.free,memory.total --format=csv,noheader,nounits 2>/dev/null) || return
|
||||
local selector="${NVIDIA_VISIBLE_DEVICES:-${CUDA_VISIBLE_DEVICES:-}}"
|
||||
local selector_specific=0
|
||||
if [[ -n "$selector" && "$selector" != "all" && "$selector" != "void" ]]; then
|
||||
selector_specific=1
|
||||
fi
|
||||
local bad=0
|
||||
while IFS=',' read -r idx free total; do
|
||||
free=$(echo "$free" | tr -d ' ')
|
||||
total=$(echo "$total" | tr -d ' ')
|
||||
idx=$(echo "$idx" | tr -d ' ')
|
||||
[[ -z "$free" || -z "$total" ]] && continue
|
||||
if [[ "$selector_specific" -eq 1 && ",${selector}," != *",${idx},"* ]]; then
|
||||
continue
|
||||
fi
|
||||
# Require ≥80% free. Compose default gpu-memory-utilization is 0.92.
|
||||
local need=$(( total * 80 / 100 ))
|
||||
if [[ "$free" -lt "$need" ]]; then
|
||||
@@ -241,8 +252,12 @@ up_variant() {
|
||||
preflight_genesis_pin "${ROOT_DIR}" || true
|
||||
preflight_repo_drift "${ROOT_DIR}" || true
|
||||
preflight_compose_deps "${full_dir}/${file}" || exit 1
|
||||
if [[ "$eng" == "vllm" ]]; then
|
||||
preflight_compose_hardware "${full_dir}/${file}" "$v" "${FORCE:-0}" || exit 1
|
||||
fi
|
||||
preflight_kv_format_hint "${full_dir}/${file}" || true
|
||||
fi
|
||||
gpu_preflight
|
||||
|
||||
echo "[switch] bringing up: ${v} (${dir}/${file})"
|
||||
(cd "${full_dir}" && ${COMPOSE_BIN} -f "${file}" up -d)
|
||||
@@ -319,24 +334,31 @@ wait_ready() {
|
||||
|
||||
# --- arg parsing ---
|
||||
WAIT=1
|
||||
case "${1:-}" in
|
||||
-h|--help|"") usage ;;
|
||||
--list) list_variants ;;
|
||||
--down) down_running; exit 0 ;;
|
||||
esac
|
||||
|
||||
VARIANT="$1"
|
||||
shift || true
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
FORCE="${FORCE:-0}"
|
||||
VARIANT=""
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
-h|--help) usage ;;
|
||||
--list) list_variants ;;
|
||||
--down) down_running; exit 0 ;;
|
||||
--no-wait) WAIT=0 ;;
|
||||
*) echo "Unknown flag: $arg"; exit 1 ;;
|
||||
--force) FORCE=1 ;;
|
||||
--*) echo "Unknown flag: $1"; exit 1 ;;
|
||||
*)
|
||||
if [[ -n "$VARIANT" ]]; then
|
||||
echo "ERROR: multiple variants supplied: '${VARIANT}' and '$1'" >&2
|
||||
exit 1
|
||||
fi
|
||||
VARIANT="$1"
|
||||
;;
|
||||
esac
|
||||
shift
|
||||
done
|
||||
|
||||
[[ -n "$VARIANT" ]] || usage
|
||||
|
||||
resolve_ready_url "${VARIANT}"
|
||||
down_running
|
||||
gpu_preflight
|
||||
up_variant "${VARIANT}"
|
||||
[[ $WAIT -eq 1 ]] && wait_ready
|
||||
echo "[switch] done. Try: curl -s ${READY_URL%/v1/models}/v1/models | jq ."
|
||||
|
||||
Executable
+187
@@ -0,0 +1,187 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||
TMP_DIR="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
make_mock_nvidia_smi() {
|
||||
mkdir -p "${TMP_DIR}/bin"
|
||||
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK'
|
||||
#!/usr/bin/env bash
|
||||
case "$*" in
|
||||
*"--query-gpu=index,name,memory.total,compute_cap"*)
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
|
||||
;;
|
||||
*"--query-gpu=index,memory.total"*)
|
||||
printf '%s\n' "${MOCK_GPU_MEM_QUERY:-${MOCK_GPU_QUERY:?}}" \
|
||||
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $3); print $1 ", " $3}'
|
||||
;;
|
||||
*"--query-gpu=index,memory.free,memory.total"*)
|
||||
printf '%s\n' "${MOCK_GPU_FREE_QUERY:?MOCK_GPU_FREE_QUERY not set}"
|
||||
;;
|
||||
"-L")
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?}" \
|
||||
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
|
||||
;;
|
||||
*)
|
||||
echo "unexpected nvidia-smi invocation: $*" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
MOCK
|
||||
chmod +x "${TMP_DIR}/bin/nvidia-smi"
|
||||
export PATH="${TMP_DIR}/bin:${PATH}"
|
||||
}
|
||||
|
||||
make_compose() {
|
||||
local path="$1"
|
||||
local min_vram="$2"
|
||||
local min_gpu="$3"
|
||||
local tp="$4"
|
||||
local sm="${5:-}"
|
||||
|
||||
{
|
||||
echo "# Hardware metadata (test fixture):"
|
||||
echo "# Requires-min-vram-gb: ${min_vram}"
|
||||
echo "# Requires-min-gpu-count: ${min_gpu}"
|
||||
echo "# Tensor-parallel: ${tp}"
|
||||
if [[ -n "$sm" ]]; then
|
||||
echo "# Requires-sm: ${sm}"
|
||||
fi
|
||||
echo "services: {}"
|
||||
} > "$path"
|
||||
}
|
||||
|
||||
assert_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" != *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
run_case() {
|
||||
local compose="$1"
|
||||
local variant="$2"
|
||||
local force="${3:-0}"
|
||||
(
|
||||
unset CLUB3090_GPU CUDA_VISIBLE_DEVICES NVIDIA_VISIBLE_DEVICES FORCE
|
||||
source "${ROOT_DIR}/scripts/preflight.sh"
|
||||
if preflight_compose_hardware "$compose" "$variant" "$force"; then
|
||||
echo "STATUS=ok"
|
||||
echo "NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-}"
|
||||
else
|
||||
rc=$?
|
||||
echo "STATUS=fail:${rc}"
|
||||
exit "$rc"
|
||||
fi
|
||||
) 2>&1
|
||||
}
|
||||
|
||||
expect_failure() {
|
||||
local compose="$1"
|
||||
local variant="$2"
|
||||
local force="${3:-0}"
|
||||
local output
|
||||
|
||||
if output="$(run_case "$compose" "$variant" "$force")"; then
|
||||
echo "ASSERTION FAILED: expected preflight failure for ${variant}" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$output" >&2
|
||||
exit 1
|
||||
fi
|
||||
printf '%s' "$output"
|
||||
}
|
||||
|
||||
make_mock_nvidia_smi
|
||||
|
||||
single_compose="${TMP_DIR}/single.yml"
|
||||
dual_compose="${TMP_DIR}/dual.yml"
|
||||
quad_compose="${TMP_DIR}/quad.yml"
|
||||
gemma_single="${TMP_DIR}/gemma-single.yml"
|
||||
missing_meta="${TMP_DIR}/missing.yml"
|
||||
|
||||
make_compose "$single_compose" 24 1 1
|
||||
make_compose "$dual_compose" 24 2 2
|
||||
make_compose "$quad_compose" 24 4 4
|
||||
make_compose "$gemma_single" 32 1 1 "9.0+"
|
||||
echo "services: {}" > "$missing_meta"
|
||||
|
||||
# 1. Matched 2x3090 + TP=1 compose: pass, deterministic GPU 0 selection.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576\n1, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(run_case "$single_compose" "vllm/default")"
|
||||
assert_contains "$out" "auto-selected GPU 0"
|
||||
assert_contains "$out" "NVIDIA_VISIBLE_DEVICES=0"
|
||||
|
||||
# 2. Matched 2x3090 + TP=2 compose: pass.
|
||||
out="$(run_case "$dual_compose" "vllm/dual")"
|
||||
assert_contains "$out" "TP=2 requires 2 GPU(s)"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 3. Matched 2x3090 + TP=4 compose: hard fail.
|
||||
out="$(expect_failure "$quad_compose" "vllm/dual4")"
|
||||
assert_contains "$out" "requires 4 visible GPU(s)"
|
||||
|
||||
# 4. 1x3090 + TP=2 compose: hard fail.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(expect_failure "$dual_compose" "vllm/dual")"
|
||||
assert_contains "$out" "requires 2 visible GPU(s)"
|
||||
|
||||
# 5. 1x3090 + TP=1, 24 GB floor: pass.
|
||||
out="$(run_case "$single_compose" "vllm/default")"
|
||||
assert_contains "$out" "auto-selected GPU 0"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 6. 16 GB + 24 GB + TP=1: auto-select the 24 GB card.
|
||||
MOCK_GPU_QUERY=$'0, RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 16000, 16384\n1, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(run_case "$single_compose" "vllm/default")"
|
||||
assert_contains "$out" "auto-selected GPU 1"
|
||||
assert_contains "$out" "NVIDIA_VISIBLE_DEVICES=1"
|
||||
|
||||
# 7. 16 GB + 24 GB + TP=2: warn, then proceed for tuned sub-24 GB rigs.
|
||||
out="$(run_case "$dual_compose" "vllm/dual")"
|
||||
assert_contains "$out" "WARN:"
|
||||
assert_contains "$out" "TP=2"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 8. 1x3090 + TP=1 compose with 32 GB + sm_9.0+ floor: hard fail.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(expect_failure "$gemma_single" "vllm/gemma-mtp-tp1")"
|
||||
assert_contains "$out" "requires one GPU with >=32 GB VRAM, sm_9.0+"
|
||||
|
||||
# 9. 1xH100 + TP=1 compose with sm_9.0+ floor: pass.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA H100 80GB HBM3, 81920, 9.0'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 80000, 81920'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(run_case "$gemma_single" "vllm/gemma-mtp-tp1")"
|
||||
assert_contains "$out" "auto-selected GPU 0"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 10. --force skips the hardware gate.
|
||||
out="$(run_case "$gemma_single" "vllm/gemma-mtp-tp1" 1)"
|
||||
assert_contains "$out" "hardware: skipped"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 11. Missing metadata warns and allows.
|
||||
out="$(run_case "$missing_meta" "vllm/local")"
|
||||
assert_contains "$out" "no hardware metadata"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
echo "test-preflight-vram: ok"
|
||||
Executable
+111
@@ -0,0 +1,111 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||
cd "$ROOT_DIR"
|
||||
|
||||
# shellcheck source=../lib/bench-row-formatter.sh
|
||||
source "$ROOT_DIR/scripts/lib/bench-row-formatter.sh"
|
||||
|
||||
assert_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" != *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
assert_columns() {
|
||||
local row="$1"
|
||||
local expected="$2"
|
||||
local cols
|
||||
cols="$(awk -F'|' '{print NF - 2}' <<< "$row")"
|
||||
if [[ "$cols" != "$expected" ]]; then
|
||||
echo "ASSERTION FAILED: expected $expected columns, got $cols" >&2
|
||||
echo "$row" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
fixtures=()
|
||||
while IFS= read -r fixture; do
|
||||
fixtures+=("$fixture")
|
||||
done < <(bench_row_fixtures)
|
||||
|
||||
if [[ "${#fixtures[@]}" -lt 6 ]]; then
|
||||
echo "ASSERTION FAILED: expected at least 6 submit-bench fixtures, found ${#fixtures[@]}" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
for dir in "${fixtures[@]}"; do
|
||||
section="$(bench_row_section "$dir")"
|
||||
row="$(bench_row_format "$dir")"
|
||||
[[ "$row" == \|*\| ]] || {
|
||||
echo "ASSERTION FAILED: row is not a markdown table row for $dir" >&2
|
||||
echo "$row" >&2
|
||||
exit 1
|
||||
}
|
||||
assert_contains "$row" "Report: \`results/rebench/${dir##*/}/REPORT.md\`"
|
||||
if [[ "$section" == "Gemma 4 31B (community-experimental)" ]]; then
|
||||
assert_columns "$row" 10
|
||||
else
|
||||
assert_columns "$row" 8
|
||||
fi
|
||||
done
|
||||
|
||||
tag="qwen-int8-pth-n4-2026-05-10"
|
||||
rm -f "results/rebench/${tag}/BENCHMARKS-row.md" \
|
||||
"results/rebench/${tag}/PR-body.md" \
|
||||
"results/rebench/${tag}/ISSUE-body.md" \
|
||||
"results/rebench/${tag}/auto-submit-mock.log"
|
||||
|
||||
out="$(bash scripts/submit-bench.sh --tag "$tag")"
|
||||
assert_contains "$out" "Generated BENCHMARKS row for section: Dual-card (2× RTX 3090, TP=2)"
|
||||
assert_contains "$out" "Wrote: results/rebench/${tag}/BENCHMARKS-row.md"
|
||||
assert_contains "$out" "1. Issue + maintainer integrates"
|
||||
assert_contains "$out" "2. Direct PR"
|
||||
assert_contains "$out" "3. Manual edit"
|
||||
test -s "results/rebench/${tag}/BENCHMARKS-row.md"
|
||||
|
||||
if out="$(bash scripts/submit-bench.sh --tag does-not-exist 2>&1)"; then
|
||||
echo "ASSERTION FAILED: missing tag unexpectedly succeeded" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "tag dir not found: results/rebench/does-not-exist"
|
||||
|
||||
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit)"
|
||||
assert_contains "$out" "Issue title: [bench] @octocat ${tag}"
|
||||
test -s "results/rebench/${tag}/auto-submit-mock.log"
|
||||
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh issue create --title [bench] @octocat ${tag}"
|
||||
test -s "results/rebench/${tag}/ISSUE-body.md"
|
||||
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "results/rebench/${tag}/REPORT.md"
|
||||
assert_contains "$(cat "results/rebench/${tag}/ISSUE-body.md")" "Proposed BENCHMARKS.md row"
|
||||
|
||||
out="$(GH_MOCK=1 GH_MOCK_USER=octocat bash scripts/submit-bench.sh --tag "$tag" --auto-submit --as-pr)"
|
||||
assert_contains "$out" "PR title: bench(matrix): @octocat ${tag}"
|
||||
test -s "results/rebench/${tag}/PR-body.md"
|
||||
assert_contains "$(cat "results/rebench/${tag}/auto-submit-mock.log")" "gh pr create --title bench(matrix): @octocat ${tag}"
|
||||
assert_contains "$(cat "results/rebench/${tag}/PR-body.md")" "results/rebench/${tag}/REPORT.md"
|
||||
|
||||
tmp_bin="$(mktemp -d)"
|
||||
trap 'rm -rf "$tmp_bin"; rm -f "results/rebench/${tag}/BENCHMARKS-row.md" "results/rebench/${tag}/PR-body.md" "results/rebench/${tag}/ISSUE-body.md" "results/rebench/${tag}/auto-submit-mock.log"' EXIT
|
||||
cat > "${tmp_bin}/gh" <<'MOCK_GH'
|
||||
#!/usr/bin/env bash
|
||||
if [[ "$1" == "auth" && "$2" == "status" ]]; then
|
||||
exit 1
|
||||
fi
|
||||
echo "unexpected gh call: $*" >&2
|
||||
exit 2
|
||||
MOCK_GH
|
||||
chmod +x "${tmp_bin}/gh"
|
||||
|
||||
if out="$(PATH="${tmp_bin}:${PATH}" bash scripts/submit-bench.sh --tag "$tag" --auto-submit 2>&1)"; then
|
||||
echo "ASSERTION FAILED: unauthenticated gh path unexpectedly succeeded" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "not authed with gh. Run: gh auth login"
|
||||
|
||||
echo "test-submit-bench: ok"
|
||||
Reference in New Issue
Block a user