Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
26985527f7 | ||
|
|
642dfba8ef |
+10
-5
@@ -38,11 +38,16 @@
|
||||
# GPU selection
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
# Which GPUs to expose to docker. Single-card composes use just `0`,
|
||||
# dual-card composes use `0,1`. The compose files set this themselves;
|
||||
# override here if your physical layout differs (e.g. you run dual-card
|
||||
# on cards 2,3).
|
||||
# CUDA_VISIBLE_DEVICES=0,1
|
||||
# Which GPU(s) to expose to docker.
|
||||
#
|
||||
# scripts/switch.sh auto-selects the largest eligible card for TP=1 vLLM
|
||||
# composes. Override with CLUB3090_GPU when you know which physical GPU should
|
||||
# run a single-card compose (for example, a 24 GB card beside a 16 GB card).
|
||||
# CLUB3090_GPU=1
|
||||
#
|
||||
# For direct `docker compose ... up` runs, or TP>=2 non-default layouts, set
|
||||
# NVIDIA_VISIBLE_DEVICES explicitly. Use comma-separated physical indices.
|
||||
# NVIDIA_VISIBLE_DEVICES=0,1
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
@@ -16,6 +16,22 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.5.1 — 2026-05-13
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(qwen): PR #35936 overlay — sidecar pattern to resolve Genesis RO-mount conflict ([6617e1e](https://github.com/noonghunna/club-3090/commit/6617e1e090a6f52708aaf83821a92b612c4ad869))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs: clarify MODEL_DIR — second drive / HF cache / Windows-WSL ([1678ca0](https://github.com/noonghunna/club-3090/commit/1678ca0c8ba43ea09fad9073639a48337f2ff163))
|
||||
- docs(upstream): correct stale vllm#40807 row + add #40798/#42215 row ([14ffe45](https://github.com/noonghunna/club-3090/commit/14ffe45667fb0aa292839be05ae3ad0d139d3a04))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.1`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.0...v0.5.1)
|
||||
## v0.5.0 — 2026-05-12
|
||||
|
||||
|
||||
|
||||
@@ -33,6 +33,22 @@ The recipes are written against 3090 specifically but should work on:
|
||||
|
||||
**Won't work:** anything with <20 GB VRAM (3060, 3070, stock 3080, 3080 Ti). The 27B model in INT4 is ~18 GB — KV pool + activations push past 24 GB on smaller cards even with aggressive quantization. **Modded 20 GB 3080s do work** (see row above) — the mod gives them enough headroom for the 27B + TQ K8V4 KV path on TP=2, with `mem-util=0.82` to absorb cudagraph profiling overhead.
|
||||
|
||||
### Mismatched / heterogeneous GPUs
|
||||
|
||||
`scripts/switch.sh` reads hardware metadata from the vLLM compose headers before starting Docker. The preflight checks required GPU count, per-GPU VRAM, tensor parallel size, and any hard SM floor.
|
||||
|
||||
For TP=1 vLLM composes, `switch.sh` auto-selects the largest eligible GPU and exports it through `NVIDIA_VISIBLE_DEVICES`. On a mixed 16 GB + 24 GB rig, `bash scripts/switch.sh vllm/default` should pick the 24 GB card instead of trying to boot on GPU 0 blindly.
|
||||
|
||||
Overrides:
|
||||
|
||||
```bash
|
||||
CLUB3090_GPU=1 bash scripts/switch.sh vllm/default
|
||||
NVIDIA_VISIBLE_DEVICES=2,3 bash scripts/switch.sh vllm/dual
|
||||
bash scripts/switch.sh --force vllm/gemma-mtp-tp1
|
||||
```
|
||||
|
||||
Use `--force` only when you are intentionally testing an unsupported combo. Example: `vllm/gemma-mtp-tp1` is now preflight-blocked on a 24 GB 3090 because the compose is preserved for 32 GB / newer-SM single-card rigs.
|
||||
|
||||
### Note for sub-24 GB cards
|
||||
|
||||
On 20 GB cards (modded 3080) the cudagraph-profiling overhead is a meaningful slice of available VRAM. Drop `--gpu-memory-utilization` to **0.82** (vs shipped 0.95 for 24 GB). vLLM nightly's `gpu_worker.py` reports the equivalent effective KV size in the boot log; tune to keep activation headroom for the ~15K tool-prefill peak (verify-full check 8). Credit: [@troymroberts](https://github.com/troymroberts).
|
||||
|
||||
@@ -58,6 +58,10 @@
|
||||
# - max-num-seqs default 4 (3dluvr's 1 is single-stream-only); override
|
||||
# MAX_NUM_SEQS=1 for max-context single-stream like 3dluvr.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-awq:
|
||||
# Latest nightly that contains PR #41745 (Gemma 4 MTP). AWQ doesn't need
|
||||
@@ -72,6 +76,7 @@ services:
|
||||
- ../../cache/torch_compile_awq:/root/.cache/vllm/torch_compile_cache
|
||||
- ../../cache/triton_awq:/root/.triton/cache
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -89,6 +89,10 @@
|
||||
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
|
||||
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-bf16:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -118,6 +122,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -93,6 +93,10 @@
|
||||
# 1. PR #41703 z-lab DFlash drafter — see ../../patches/vllm-gemma4-dflash/README
|
||||
# 2. PR #40391 per-token-head — see ../../patches/vllm-gemma4-dflash-int8/README
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-dflash-int8:
|
||||
# SAME pin as dual-dflash.yml — DFlash overlay was rebased to e47c98ef.
|
||||
@@ -139,6 +143,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-dflash-int8/v1/worker/kv_cache_shape_utils.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/kv_cache_shape_utils.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -65,6 +65,10 @@
|
||||
# - default (auto = bfloat16) → sidesteps all the above.
|
||||
# Smaller KV pool than fp8 would give but it's the only Ampere-shippable.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-dflash:
|
||||
# Nightly bumped to 2026-05-06 to match rebase target proximity (Codex
|
||||
@@ -104,6 +108,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-dflash/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -48,6 +48,10 @@
|
||||
# - default (auto = bfloat16) → sidesteps all fp8 kernel paths
|
||||
# Smaller KV pool than fp8 but at 32K test ctx that's not the bottleneck.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -63,6 +67,7 @@ services:
|
||||
- ../../cache/triton:/root/.triton/cache
|
||||
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -177,6 +177,11 @@
|
||||
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
|
||||
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
# Requires-sm: 9.0+
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-int8-tq3:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -213,6 +218,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -89,6 +89,10 @@
|
||||
# gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'
|
||||
# gh api repos/vllm-project/vllm/pulls/41991 --jq '.state, .merged_at'
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-int8:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -118,6 +122,7 @@ services:
|
||||
- ../../patches/vllm-gemma4-tool-parser-fixes/tool_parsers/gemma4_tool_parser.py:/usr/local/lib/python3.12/dist-packages/vllm/tool_parsers/gemma4_tool_parser.py:ro
|
||||
# --------------------------------------------------------------------
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -48,6 +48,11 @@
|
||||
# - default (auto = bfloat16) → sidesteps all fp8 kernel paths
|
||||
# Smaller KV pool than fp8 but at 32K test ctx that's not the bottleneck.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 32
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
# Requires-sm: 9.0+
|
||||
services:
|
||||
vllm-gemma-4-31b-mtp-tp1:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -63,6 +68,7 @@ services:
|
||||
- ../../cache/triton:/root/.triton/cache
|
||||
# PR #41745 overlay dropped 2026-05-08: merged upstream + nightly contains it.
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -20,6 +20,10 @@
|
||||
# Run:
|
||||
# docker compose -f dual/bf16.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-bf16:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/bf16.yml — 2026-05-08
|
||||
@@ -59,6 +63,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -48,6 +48,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/carnice-bf16mtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-carnice-bf16mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -82,6 +86,7 @@ services:
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# PCIe-only stack — disable NCCL features that assume NVLink.
|
||||
|
||||
@@ -43,6 +43,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/dflash-noviz.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -77,6 +81,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -67,6 +67,10 @@
|
||||
#
|
||||
# docker compose -f dual/dflash.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -101,6 +105,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -49,6 +49,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/docker-compose.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual:
|
||||
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
|
||||
@@ -89,6 +93,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# PCIe-only stack — disable NCCL features that assume NVLink.
|
||||
|
||||
@@ -20,6 +20,10 @@
|
||||
# Run:
|
||||
# docker compose -f dual/int8.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-int8:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8.yml — 2026-05-08
|
||||
@@ -59,6 +63,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -75,6 +75,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink-dflash-noviz.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink-dflash-noviz:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -109,6 +113,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
|
||||
@@ -71,6 +71,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink-dflash.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -105,6 +109,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
|
||||
@@ -61,6 +61,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink-turbo.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink-turbo:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -102,6 +106,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
|
||||
@@ -57,6 +57,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/nvlink.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-nvlink:
|
||||
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
|
||||
@@ -95,6 +99,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# NVLink bridge present — let NCCL use P2P (don't disable it the way
|
||||
|
||||
@@ -33,6 +33,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/qwopus-bf16mtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwopus-bf16mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -58,6 +62,7 @@ services:
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -61,6 +61,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/tq3-mtp-genesis.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-tq3-mtp-genesis:
|
||||
# Pinned to a Genesis v7.72.2 known-good vLLM nightly. The canonical
|
||||
@@ -101,6 +105,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -59,6 +59,10 @@
|
||||
# Run (only when all 5 upstream fixes are available):
|
||||
# docker compose -f dual/tq3-mtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-int8-tq3:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8-tq3.yml — 2026-05-08
|
||||
@@ -106,6 +110,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- VLLM_ENFORCE_EAGER=1
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
|
||||
@@ -35,6 +35,10 @@
|
||||
# To run:
|
||||
# docker compose -f dual/tq3-nomtp.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-int8-tq3-nomtp:
|
||||
# Same vLLM nightly as gemma-4-31b/vllm/compose/dual/int8-tq3.yml — 2026-05-08
|
||||
@@ -90,6 +94,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -45,6 +45,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f dual/turbo.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 2
|
||||
# Tensor-parallel: 2
|
||||
services:
|
||||
vllm-qwen36-27b-dual-turbo:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -86,6 +90,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -60,6 +60,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f multi4/dflash.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 4
|
||||
# Tensor-parallel: 4
|
||||
services:
|
||||
vllm-qwen36-27b-multi4-dflash:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -94,6 +98,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
- NCCL_CUMEM_ENABLE=0
|
||||
|
||||
@@ -57,6 +57,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f multi4/docker-compose.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 4
|
||||
# Tensor-parallel: 4
|
||||
services:
|
||||
vllm-qwen36-27b-multi4:
|
||||
# Tracking latest nightly intentionally — this stack uses fp8 KV (not
|
||||
@@ -97,6 +101,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
# PCIe-only stack — disable NCCL features that assume NVLink.
|
||||
|
||||
@@ -129,6 +129,10 @@
|
||||
# Then send requests with structured_outputs in extra_body — see
|
||||
# docs/STRUCTURED_COT.md for client examples.
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-bounded-thinking:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -163,6 +167,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
|
||||
@@ -95,6 +95,10 @@
|
||||
# bash ../scripts/setup.sh # ensures Genesis tree at v7.14+ layout
|
||||
# cd compose && docker compose up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -134,6 +138,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
|
||||
@@ -109,6 +109,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/long-text.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-long-text-no-mtp:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -151,6 +155,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
|
||||
@@ -119,6 +119,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/long-text.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-long-text:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -161,6 +165,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
|
||||
@@ -95,6 +95,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/long-vision.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-long-vision:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -129,6 +133,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
|
||||
@@ -31,6 +31,10 @@
|
||||
# Run:
|
||||
# cd compose && docker compose -f single/minimal.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 20
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b-minimal:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -61,6 +65,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# Uncomment the next line to pin to a specific GPU (e.g. GPU 0):
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
|
||||
@@ -33,6 +33,10 @@
|
||||
# cd <repo>/models/qwen3.6-27b/vllm/compose
|
||||
# docker compose -f single/tools-text.yml up -d
|
||||
# ===========================================================================
|
||||
# Hardware metadata (parsed by scripts/preflight.sh):
|
||||
# Requires-min-vram-gb: 24
|
||||
# Requires-min-gpu-count: 1
|
||||
# Tensor-parallel: 1
|
||||
services:
|
||||
vllm-qwen36-27b:
|
||||
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
|
||||
@@ -66,6 +70,7 @@ services:
|
||||
# https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
|
||||
- ../../patches/froggeric-chat-template/chat_template.jinja:/etc/qwen-froggeric-chat-template.jinja:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-all}
|
||||
# - CUDA_VISIBLE_DEVICES=0
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# Tiny parser for hardware metadata stored as compose header comments.
|
||||
#
|
||||
# Expected form:
|
||||
# # Requires-min-vram-gb: 24
|
||||
# # Requires-min-gpu-count: 2
|
||||
# # Tensor-parallel: 2
|
||||
# # Requires-sm: 9.0+
|
||||
#
|
||||
# This intentionally does not parse YAML. These fields are comments so that
|
||||
# older docker compose versions and direct `docker compose -f ... up` flows keep
|
||||
# working unchanged.
|
||||
|
||||
_compose_meta_trim() {
|
||||
local value="$1"
|
||||
value="${value#"${value%%[![:space:]]*}"}"
|
||||
value="${value%"${value##*[![:space:]]}"}"
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
_compose_meta_norm_key() {
|
||||
local key="$1"
|
||||
key="$(_compose_meta_trim "$key")"
|
||||
key="${key//_/-}"
|
||||
key="${key// /-}"
|
||||
printf '%s' "$key" | tr '[:upper:]' '[:lower:]'
|
||||
}
|
||||
|
||||
_compose_meta_wants_key() {
|
||||
local requested="$(_compose_meta_norm_key "$1")"
|
||||
local candidate="$(_compose_meta_norm_key "$2")"
|
||||
|
||||
case "$requested" in
|
||||
min-vram-gb) requested="requires-min-vram-gb" ;;
|
||||
min-gpu-count) requested="requires-min-gpu-count" ;;
|
||||
tp) requested="tensor-parallel" ;;
|
||||
sm) requested="requires-sm" ;;
|
||||
esac
|
||||
|
||||
[[ "$candidate" == "$requested" ]]
|
||||
}
|
||||
|
||||
compose_meta_get() {
|
||||
local compose_file="$1"
|
||||
local field="$2"
|
||||
|
||||
[[ -f "$compose_file" ]] || return 1
|
||||
|
||||
local line key value
|
||||
while IFS= read -r line; do
|
||||
[[ "$line" =~ ^[[:space:]]*# ]] || continue
|
||||
line="${line#*\#}"
|
||||
[[ "$line" == *:* ]] || continue
|
||||
key="${line%%:*}"
|
||||
value="${line#*:}"
|
||||
if _compose_meta_wants_key "$field" "$key"; then
|
||||
_compose_meta_trim "$value"
|
||||
return 0
|
||||
fi
|
||||
done < "$compose_file"
|
||||
|
||||
return 1
|
||||
}
|
||||
+315
-3
@@ -12,6 +12,7 @@
|
||||
# preflight_running — warn if a club-3090 container is already up
|
||||
# preflight_genesis_pin — warn if on-disk Genesis tree differs from setup.sh's pin
|
||||
# preflight_repo_drift — warn if local HEAD is behind origin/master
|
||||
# preflight_compose_hardware— check compose VRAM/GPU-count/SM metadata
|
||||
#
|
||||
# Style: each function prints one or more "[preflight] ..." lines.
|
||||
# Hard failures get a one-line ERROR + a "Fix:" hint.
|
||||
@@ -19,6 +20,12 @@
|
||||
# Avoid double-sourcing.
|
||||
[[ -n "${_PREFLIGHT_LOADED:-}" ]] && return 0
|
||||
_PREFLIGHT_LOADED=1
|
||||
_PREFLIGHT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
if [[ -f "${_PREFLIGHT_DIR}/lib/compose-meta.sh" ]]; then
|
||||
# shellcheck source=lib/compose-meta.sh
|
||||
source "${_PREFLIGHT_DIR}/lib/compose-meta.sh"
|
||||
fi
|
||||
|
||||
preflight_docker() {
|
||||
if ! command -v docker >/dev/null 2>&1; then
|
||||
@@ -72,6 +79,301 @@ preflight_gpu() {
|
||||
return 0
|
||||
}
|
||||
|
||||
_preflight_trim() {
|
||||
local value="$1"
|
||||
value="${value#"${value%%[![:space:]]*}"}"
|
||||
value="${value%"${value##*[![:space:]]}"}"
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
_preflight_csv_token() {
|
||||
local value="$1"
|
||||
value="$(_preflight_trim "$value")"
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
_preflight_selector() {
|
||||
if [[ -n "${CLUB3090_GPU:-}" ]]; then
|
||||
printf '%s' "${CLUB3090_GPU}"
|
||||
elif [[ -n "${NVIDIA_VISIBLE_DEVICES:-}" && "${NVIDIA_VISIBLE_DEVICES}" != "all" && "${NVIDIA_VISIBLE_DEVICES}" != "void" ]]; then
|
||||
printf '%s' "${NVIDIA_VISIBLE_DEVICES}"
|
||||
elif [[ -n "${CUDA_VISIBLE_DEVICES:-}" && "${CUDA_VISIBLE_DEVICES}" != "all" && "${CUDA_VISIBLE_DEVICES}" != "void" ]]; then
|
||||
printf '%s' "${CUDA_VISIBLE_DEVICES}"
|
||||
fi
|
||||
}
|
||||
|
||||
_preflight_selector_is_specific() {
|
||||
local selector="${1:-}"
|
||||
[[ -n "$selector" && "$selector" != "all" && "$selector" != "void" ]]
|
||||
}
|
||||
|
||||
_preflight_selector_allows_index() {
|
||||
local selector="$1"
|
||||
local idx="$2"
|
||||
local token
|
||||
|
||||
if ! _preflight_selector_is_specific "$selector"; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
IFS=',' read -ra _preflight_selector_tokens <<< "$selector"
|
||||
for token in "${_preflight_selector_tokens[@]}"; do
|
||||
token="$(_preflight_trim "$token")"
|
||||
[[ "$token" == "$idx" ]] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
_preflight_selector_first_numeric() {
|
||||
local selector="$1"
|
||||
local token
|
||||
|
||||
IFS=',' read -ra _preflight_selector_tokens <<< "$selector"
|
||||
for token in "${_preflight_selector_tokens[@]}"; do
|
||||
token="$(_preflight_trim "$token")"
|
||||
if [[ "$token" =~ ^[0-9]+$ ]]; then
|
||||
printf '%s' "$token"
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
_preflight_sm_to_int() {
|
||||
local sm="$1"
|
||||
sm="${sm%%+}"
|
||||
sm="${sm//sm_/}"
|
||||
sm="${sm//SM_/}"
|
||||
sm="${sm// /}"
|
||||
[[ -z "$sm" ]] && { echo 0; return; }
|
||||
|
||||
local major minor
|
||||
if [[ "$sm" == *.* ]]; then
|
||||
major="${sm%%.*}"
|
||||
minor="${sm#*.}"
|
||||
else
|
||||
major="$sm"
|
||||
minor="0"
|
||||
fi
|
||||
major="${major//[^0-9]/}"
|
||||
minor="${minor//[^0-9]/}"
|
||||
[[ -z "$major" ]] && major=0
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
if [[ "${#minor}" -eq 1 ]]; then
|
||||
minor=$(( minor * 10 ))
|
||||
else
|
||||
minor="${minor:0:2}"
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
fi
|
||||
echo $(( major * 100 + minor ))
|
||||
}
|
||||
|
||||
_preflight_vram_gb() {
|
||||
local mib="$1"
|
||||
echo $(( (mib + 1023) / 1024 ))
|
||||
}
|
||||
|
||||
_preflight_hardware_suggestions() {
|
||||
local variant="${1:-}"
|
||||
|
||||
echo "[preflight]" >&2
|
||||
echo "[preflight] Suggested next steps:" >&2
|
||||
echo "[preflight] - Pick a compose that matches the detected GPU VRAM/topology." >&2
|
||||
if [[ "$variant" == vllm/gemma-mtp-tp1 ]]; then
|
||||
echo "[preflight] - On 2x 24 GB cards, use: bash scripts/switch.sh vllm/gemma-mtp" >&2
|
||||
fi
|
||||
echo "[preflight] - On a single 24 GB card, start with: bash scripts/switch.sh vllm/default" >&2
|
||||
echo "[preflight] - For maximum compatibility, use: bash scripts/switch.sh llamacpp/default" >&2
|
||||
echo "[preflight] - Explicit bypass: bash scripts/switch.sh --force ${variant:-<variant>}" >&2
|
||||
}
|
||||
|
||||
# preflight_compose_hardware <compose_file> [variant] [force]
|
||||
#
|
||||
# Reads compose header metadata and checks the target host before docker compose
|
||||
# starts. This is intentionally conservative:
|
||||
# - Missing metadata warns and allows the boot.
|
||||
# - TP=1 composes auto-select the largest eligible GPU unless the user set
|
||||
# CLUB3090_GPU, CUDA_VISIBLE_DEVICES, or NVIDIA_VISIBLE_DEVICES.
|
||||
# - TP>=2 composes hard-fail only on insufficient GPU count or hard SM gates;
|
||||
# heterogeneous VRAM below the requested floor warns because advanced users
|
||||
# may be validating sub-24 GB configs with tuned memory-utilization.
|
||||
preflight_compose_hardware() {
|
||||
local compose_file="$1"
|
||||
local variant="${2:-}"
|
||||
local force="${3:-${FORCE:-0}}"
|
||||
|
||||
if [[ "${PREFLIGHT_NO_HARDWARE:-0}" == "1" ]]; then
|
||||
return 0
|
||||
fi
|
||||
if [[ "$force" == "1" || "${FORCE:-0}" == "1" ]]; then
|
||||
echo "[preflight] hardware: skipped (--force/FORCE=1)"
|
||||
return 0
|
||||
fi
|
||||
if [[ ! -f "$compose_file" ]]; then
|
||||
echo "[preflight] ERROR: compose file not found: $compose_file" >&2
|
||||
return 1
|
||||
fi
|
||||
if ! command -v nvidia-smi >/dev/null 2>&1; then
|
||||
echo "[preflight] WARN: nvidia-smi not found; skipping compose hardware metadata check." >&2
|
||||
return 0
|
||||
fi
|
||||
if ! declare -F compose_meta_get >/dev/null 2>&1; then
|
||||
echo "[preflight] WARN: compose metadata parser unavailable; skipping hardware metadata check." >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
local min_vram_gb min_gpu_count tp requires_sm
|
||||
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
|
||||
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
|
||||
tp="$(compose_meta_get "$compose_file" tensor-parallel || true)"
|
||||
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
|
||||
|
||||
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" || -z "$tp" ]]; then
|
||||
echo "[preflight] WARN: compose has no hardware metadata; allowing boot: $compose_file" >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
requires_sm="${requires_sm:-0.0}"
|
||||
local required_sm_int
|
||||
required_sm_int="$(_preflight_sm_to_int "$requires_sm")"
|
||||
|
||||
local gpu_query
|
||||
gpu_query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
if [[ -z "$gpu_query" ]]; then
|
||||
echo "[preflight] WARN: could not query GPU VRAM/SM via nvidia-smi; skipping hardware metadata check." >&2
|
||||
return 0
|
||||
fi
|
||||
|
||||
local selector
|
||||
selector="$(_preflight_selector || true)"
|
||||
|
||||
local total_count=0 selected_count=0 eligible_count=0 selected_below_vram=0 selected_below_sm=0
|
||||
local best_idx="" best_name="" best_mib=0 best_sm=""
|
||||
local first_idx="" first_name="" first_mib=0 first_sm=""
|
||||
local idx name mem_mib sm rest vram_gb sm_int
|
||||
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
name="$(_preflight_csv_token "$name")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
sm="$(_preflight_csv_token "$sm")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
total_count=$(( total_count + 1 ))
|
||||
_preflight_selector_allows_index "$selector" "$idx" || continue
|
||||
|
||||
selected_count=$(( selected_count + 1 ))
|
||||
if [[ -z "$first_idx" ]]; then
|
||||
first_idx="$idx"
|
||||
first_name="$name"
|
||||
first_mib="$mem_mib"
|
||||
first_sm="$sm"
|
||||
fi
|
||||
|
||||
vram_gb="$(_preflight_vram_gb "$mem_mib")"
|
||||
sm_int="$(_preflight_sm_to_int "$sm")"
|
||||
|
||||
if (( vram_gb < min_vram_gb )); then
|
||||
selected_below_vram=1
|
||||
fi
|
||||
if (( sm_int < required_sm_int )); then
|
||||
selected_below_sm=1
|
||||
fi
|
||||
|
||||
if (( vram_gb >= min_vram_gb && sm_int >= required_sm_int )); then
|
||||
eligible_count=$(( eligible_count + 1 ))
|
||||
if (( mem_mib > best_mib )); then
|
||||
best_idx="$idx"
|
||||
best_name="$name"
|
||||
best_mib="$mem_mib"
|
||||
best_sm="$sm"
|
||||
fi
|
||||
fi
|
||||
done <<< "$gpu_query"
|
||||
|
||||
if (( total_count == 0 )); then
|
||||
echo "[preflight] ERROR: no NVIDIA GPUs detected." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
if (( selected_count == 0 )); then
|
||||
echo "[preflight] ERROR: GPU selector '${selector}' did not match any detected GPU index." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
|
||||
local requires_sm_display="${requires_sm%%+}"
|
||||
local sm_label=""
|
||||
if (( required_sm_int > 0 )); then
|
||||
sm_label=", sm_${requires_sm_display}+"
|
||||
fi
|
||||
|
||||
if (( tp <= 1 )); then
|
||||
if _preflight_selector_is_specific "$selector"; then
|
||||
local first_vram_gb first_sm_int
|
||||
first_vram_gb="$(_preflight_vram_gb "$first_mib")"
|
||||
first_sm_int="$(_preflight_sm_to_int "$first_sm")"
|
||||
if (( first_vram_gb < min_vram_gb || first_sm_int < required_sm_int )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires one GPU with >=${min_vram_gb} GB VRAM${sm_label}." >&2
|
||||
echo "[preflight] Explicit selector '${selector}' starts with GPU ${first_idx}: ${first_name}, ${first_vram_gb} GB, sm_${first_sm}." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
export CLUB3090_GPU="${CLUB3090_GPU:-$selector}"
|
||||
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-$selector}"
|
||||
export NVIDIA_VISIBLE_DEVICES="${NVIDIA_VISIBLE_DEVICES:-$selector}"
|
||||
echo "[preflight] hardware: ${variant:-compose} TP=1 requires >=${min_vram_gb} GB${sm_label}; using explicit GPU ${first_idx} (${first_vram_gb} GB, sm_${first_sm})"
|
||||
return 0
|
||||
fi
|
||||
|
||||
if (( eligible_count == 0 )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires one GPU with >=${min_vram_gb} GB VRAM${sm_label}; none found." >&2
|
||||
echo "[preflight] Detected GPUs:" >&2
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
name="$(_preflight_csv_token "$name")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
sm="$(_preflight_csv_token "$sm")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
echo "[preflight] GPU ${idx}: ${name}, $(_preflight_vram_gb "$mem_mib") GB, sm_${sm}" >&2
|
||||
done <<< "$gpu_query"
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
|
||||
export CLUB3090_GPU="$best_idx"
|
||||
export CUDA_VISIBLE_DEVICES="$best_idx"
|
||||
export NVIDIA_VISIBLE_DEVICES="$best_idx"
|
||||
echo "[preflight] hardware: ${variant:-compose} TP=1 requires >=${min_vram_gb} GB${sm_label}; auto-selected GPU ${best_idx} ($(_preflight_vram_gb "$best_mib") GB, sm_${best_sm})"
|
||||
return 0
|
||||
fi
|
||||
|
||||
if (( selected_count < min_gpu_count )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires ${min_gpu_count} visible GPU(s) for TP=${tp}; found ${selected_count}." >&2
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
if (( selected_below_sm == 1 )); then
|
||||
echo "[preflight] ERROR: ${variant:-compose} requires sm_${requires_sm_display}+ on visible GPUs." >&2
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
name="$(_preflight_csv_token "$name")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
sm="$(_preflight_csv_token "$sm")"
|
||||
_preflight_selector_allows_index "$selector" "$idx" || continue
|
||||
echo "[preflight] GPU ${idx}: ${name}, $(_preflight_vram_gb "$mem_mib") GB, sm_${sm}" >&2
|
||||
done <<< "$gpu_query"
|
||||
_preflight_hardware_suggestions "$variant"
|
||||
return 1
|
||||
fi
|
||||
if (( selected_below_vram == 1 )); then
|
||||
echo "[preflight] WARN: ${variant:-compose} requires >=${min_vram_gb} GB per visible GPU for TP=${tp}, but at least one selected GPU is smaller." >&2
|
||||
echo "[preflight] Continuing because TP>=2 sub-24 GB rigs may use tuned gpu-memory-utilization/KV settings." >&2
|
||||
fi
|
||||
|
||||
echo "[preflight] hardware: ${variant:-compose} TP=${tp} requires ${min_gpu_count} GPU(s), >=${min_vram_gb} GB each${sm_label}; ${selected_count} visible GPU(s) detected"
|
||||
return 0
|
||||
}
|
||||
|
||||
preflight_disk() {
|
||||
local path="$1"
|
||||
local need_gb="$2"
|
||||
@@ -412,9 +714,19 @@ preflight_kv_format_hint() {
|
||||
return 0
|
||||
fi
|
||||
|
||||
# Detect smallest VRAM among visible cards (the TP-split ceiling).
|
||||
local min_vram_mib
|
||||
min_vram_mib="$(nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | sort -n | head -1)"
|
||||
# Detect smallest VRAM among selected/visible cards (the TP-split ceiling).
|
||||
local min_vram_mib="" mem_query selector idx mem_mib
|
||||
selector="$(_preflight_selector || true)"
|
||||
mem_query="$(nvidia-smi --query-gpu=index,memory.total --format=csv,noheader,nounits 2>/dev/null || true)"
|
||||
while IFS=',' read -r idx mem_mib; do
|
||||
idx="$(_preflight_csv_token "$idx")"
|
||||
mem_mib="$(_preflight_csv_token "$mem_mib")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
_preflight_selector_allows_index "$selector" "$idx" || continue
|
||||
if [[ -z "$min_vram_mib" || "$mem_mib" -lt "$min_vram_mib" ]]; then
|
||||
min_vram_mib="$mem_mib"
|
||||
fi
|
||||
done <<< "$mem_query"
|
||||
if [[ -z "$min_vram_mib" ]] || [[ "$min_vram_mib" -ge 24000 ]]; then
|
||||
return 0 # 24 GB+ cards — TQ3 is the right pick, no hint needed
|
||||
fi
|
||||
|
||||
+35
-13
@@ -9,6 +9,7 @@
|
||||
# Usage:
|
||||
# bash scripts/switch.sh <variant> # switch + tail until ready
|
||||
# bash scripts/switch.sh <variant> --no-wait # switch and return immediately
|
||||
# bash scripts/switch.sh --force <variant> # skip hardware/free-VRAM preflight
|
||||
# bash scripts/switch.sh --list # show all variants
|
||||
# bash scripts/switch.sh --down # just bring down whatever's up
|
||||
#
|
||||
@@ -42,6 +43,8 @@
|
||||
#
|
||||
# Env overrides (rarely needed):
|
||||
# COMPOSE_BIN Default: "docker compose" (set to e.g. "podman compose" if needed)
|
||||
# CLUB3090_GPU Single-card GPU index override, e.g. "1" on a hetero rig
|
||||
# FORCE Set to 1 to skip hardware/free-VRAM preflight
|
||||
# READY_URL Default: http://localhost:8020/v1/models
|
||||
# READY_TIMEOUT Default: 600 (seconds — longer for cold cudagraph capture)
|
||||
|
||||
@@ -169,15 +172,23 @@ gpu_preflight() {
|
||||
return
|
||||
fi
|
||||
# Free MiB per GPU. Tolerate small overhead (driver, X server) — abort
|
||||
# if any GPU has <80% of its total memory free.
|
||||
# if any selected GPU has <80% of its total memory free.
|
||||
local mem_query
|
||||
mem_query=$(nvidia-smi --query-gpu=index,memory.free,memory.total --format=csv,noheader,nounits 2>/dev/null) || return
|
||||
local selector="${NVIDIA_VISIBLE_DEVICES:-${CUDA_VISIBLE_DEVICES:-}}"
|
||||
local selector_specific=0
|
||||
if [[ -n "$selector" && "$selector" != "all" && "$selector" != "void" ]]; then
|
||||
selector_specific=1
|
||||
fi
|
||||
local bad=0
|
||||
while IFS=',' read -r idx free total; do
|
||||
free=$(echo "$free" | tr -d ' ')
|
||||
total=$(echo "$total" | tr -d ' ')
|
||||
idx=$(echo "$idx" | tr -d ' ')
|
||||
[[ -z "$free" || -z "$total" ]] && continue
|
||||
if [[ "$selector_specific" -eq 1 && ",${selector}," != *",${idx},"* ]]; then
|
||||
continue
|
||||
fi
|
||||
# Require ≥80% free. Compose default gpu-memory-utilization is 0.92.
|
||||
local need=$(( total * 80 / 100 ))
|
||||
if [[ "$free" -lt "$need" ]]; then
|
||||
@@ -241,8 +252,12 @@ up_variant() {
|
||||
preflight_genesis_pin "${ROOT_DIR}" || true
|
||||
preflight_repo_drift "${ROOT_DIR}" || true
|
||||
preflight_compose_deps "${full_dir}/${file}" || exit 1
|
||||
if [[ "$eng" == "vllm" ]]; then
|
||||
preflight_compose_hardware "${full_dir}/${file}" "$v" "${FORCE:-0}" || exit 1
|
||||
fi
|
||||
preflight_kv_format_hint "${full_dir}/${file}" || true
|
||||
fi
|
||||
gpu_preflight
|
||||
|
||||
echo "[switch] bringing up: ${v} (${dir}/${file})"
|
||||
(cd "${full_dir}" && ${COMPOSE_BIN} -f "${file}" up -d)
|
||||
@@ -319,24 +334,31 @@ wait_ready() {
|
||||
|
||||
# --- arg parsing ---
|
||||
WAIT=1
|
||||
case "${1:-}" in
|
||||
-h|--help|"") usage ;;
|
||||
--list) list_variants ;;
|
||||
--down) down_running; exit 0 ;;
|
||||
esac
|
||||
|
||||
VARIANT="$1"
|
||||
shift || true
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
FORCE="${FORCE:-0}"
|
||||
VARIANT=""
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
-h|--help) usage ;;
|
||||
--list) list_variants ;;
|
||||
--down) down_running; exit 0 ;;
|
||||
--no-wait) WAIT=0 ;;
|
||||
*) echo "Unknown flag: $arg"; exit 1 ;;
|
||||
--force) FORCE=1 ;;
|
||||
--*) echo "Unknown flag: $1"; exit 1 ;;
|
||||
*)
|
||||
if [[ -n "$VARIANT" ]]; then
|
||||
echo "ERROR: multiple variants supplied: '${VARIANT}' and '$1'" >&2
|
||||
exit 1
|
||||
fi
|
||||
VARIANT="$1"
|
||||
;;
|
||||
esac
|
||||
shift
|
||||
done
|
||||
|
||||
[[ -n "$VARIANT" ]] || usage
|
||||
|
||||
resolve_ready_url "${VARIANT}"
|
||||
down_running
|
||||
gpu_preflight
|
||||
up_variant "${VARIANT}"
|
||||
[[ $WAIT -eq 1 ]] && wait_ready
|
||||
echo "[switch] done. Try: curl -s ${READY_URL%/v1/models}/v1/models | jq ."
|
||||
|
||||
Executable
+187
@@ -0,0 +1,187 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||
TMP_DIR="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
make_mock_nvidia_smi() {
|
||||
mkdir -p "${TMP_DIR}/bin"
|
||||
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK'
|
||||
#!/usr/bin/env bash
|
||||
case "$*" in
|
||||
*"--query-gpu=index,name,memory.total,compute_cap"*)
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
|
||||
;;
|
||||
*"--query-gpu=index,memory.total"*)
|
||||
printf '%s\n' "${MOCK_GPU_MEM_QUERY:-${MOCK_GPU_QUERY:?}}" \
|
||||
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $3); print $1 ", " $3}'
|
||||
;;
|
||||
*"--query-gpu=index,memory.free,memory.total"*)
|
||||
printf '%s\n' "${MOCK_GPU_FREE_QUERY:?MOCK_GPU_FREE_QUERY not set}"
|
||||
;;
|
||||
"-L")
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?}" \
|
||||
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
|
||||
;;
|
||||
*)
|
||||
echo "unexpected nvidia-smi invocation: $*" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
MOCK
|
||||
chmod +x "${TMP_DIR}/bin/nvidia-smi"
|
||||
export PATH="${TMP_DIR}/bin:${PATH}"
|
||||
}
|
||||
|
||||
make_compose() {
|
||||
local path="$1"
|
||||
local min_vram="$2"
|
||||
local min_gpu="$3"
|
||||
local tp="$4"
|
||||
local sm="${5:-}"
|
||||
|
||||
{
|
||||
echo "# Hardware metadata (test fixture):"
|
||||
echo "# Requires-min-vram-gb: ${min_vram}"
|
||||
echo "# Requires-min-gpu-count: ${min_gpu}"
|
||||
echo "# Tensor-parallel: ${tp}"
|
||||
if [[ -n "$sm" ]]; then
|
||||
echo "# Requires-sm: ${sm}"
|
||||
fi
|
||||
echo "services: {}"
|
||||
} > "$path"
|
||||
}
|
||||
|
||||
assert_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" != *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
run_case() {
|
||||
local compose="$1"
|
||||
local variant="$2"
|
||||
local force="${3:-0}"
|
||||
(
|
||||
unset CLUB3090_GPU CUDA_VISIBLE_DEVICES NVIDIA_VISIBLE_DEVICES FORCE
|
||||
source "${ROOT_DIR}/scripts/preflight.sh"
|
||||
if preflight_compose_hardware "$compose" "$variant" "$force"; then
|
||||
echo "STATUS=ok"
|
||||
echo "NVIDIA_VISIBLE_DEVICES=${NVIDIA_VISIBLE_DEVICES:-}"
|
||||
else
|
||||
rc=$?
|
||||
echo "STATUS=fail:${rc}"
|
||||
exit "$rc"
|
||||
fi
|
||||
) 2>&1
|
||||
}
|
||||
|
||||
expect_failure() {
|
||||
local compose="$1"
|
||||
local variant="$2"
|
||||
local force="${3:-0}"
|
||||
local output
|
||||
|
||||
if output="$(run_case "$compose" "$variant" "$force")"; then
|
||||
echo "ASSERTION FAILED: expected preflight failure for ${variant}" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$output" >&2
|
||||
exit 1
|
||||
fi
|
||||
printf '%s' "$output"
|
||||
}
|
||||
|
||||
make_mock_nvidia_smi
|
||||
|
||||
single_compose="${TMP_DIR}/single.yml"
|
||||
dual_compose="${TMP_DIR}/dual.yml"
|
||||
quad_compose="${TMP_DIR}/quad.yml"
|
||||
gemma_single="${TMP_DIR}/gemma-single.yml"
|
||||
missing_meta="${TMP_DIR}/missing.yml"
|
||||
|
||||
make_compose "$single_compose" 24 1 1
|
||||
make_compose "$dual_compose" 24 2 2
|
||||
make_compose "$quad_compose" 24 4 4
|
||||
make_compose "$gemma_single" 32 1 1 "9.0+"
|
||||
echo "services: {}" > "$missing_meta"
|
||||
|
||||
# 1. Matched 2x3090 + TP=1 compose: pass, deterministic GPU 0 selection.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576\n1, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(run_case "$single_compose" "vllm/default")"
|
||||
assert_contains "$out" "auto-selected GPU 0"
|
||||
assert_contains "$out" "NVIDIA_VISIBLE_DEVICES=0"
|
||||
|
||||
# 2. Matched 2x3090 + TP=2 compose: pass.
|
||||
out="$(run_case "$dual_compose" "vllm/dual")"
|
||||
assert_contains "$out" "TP=2 requires 2 GPU(s)"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 3. Matched 2x3090 + TP=4 compose: hard fail.
|
||||
out="$(expect_failure "$quad_compose" "vllm/dual4")"
|
||||
assert_contains "$out" "requires 4 visible GPU(s)"
|
||||
|
||||
# 4. 1x3090 + TP=2 compose: hard fail.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(expect_failure "$dual_compose" "vllm/dual")"
|
||||
assert_contains "$out" "requires 2 visible GPU(s)"
|
||||
|
||||
# 5. 1x3090 + TP=1, 24 GB floor: pass.
|
||||
out="$(run_case "$single_compose" "vllm/default")"
|
||||
assert_contains "$out" "auto-selected GPU 0"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 6. 16 GB + 24 GB + TP=1: auto-select the 24 GB card.
|
||||
MOCK_GPU_QUERY=$'0, RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 16000, 16384\n1, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(run_case "$single_compose" "vllm/default")"
|
||||
assert_contains "$out" "auto-selected GPU 1"
|
||||
assert_contains "$out" "NVIDIA_VISIBLE_DEVICES=1"
|
||||
|
||||
# 7. 16 GB + 24 GB + TP=2: warn, then proceed for tuned sub-24 GB rigs.
|
||||
out="$(run_case "$dual_compose" "vllm/dual")"
|
||||
assert_contains "$out" "WARN:"
|
||||
assert_contains "$out" "TP=2"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 8. 1x3090 + TP=1 compose with 32 GB + sm_9.0+ floor: hard fail.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 24000, 24576'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(expect_failure "$gemma_single" "vllm/gemma-mtp-tp1")"
|
||||
assert_contains "$out" "requires one GPU with >=32 GB VRAM, sm_9.0+"
|
||||
|
||||
# 9. 1xH100 + TP=1 compose with sm_9.0+ floor: pass.
|
||||
MOCK_GPU_QUERY=$'0, NVIDIA H100 80GB HBM3, 81920, 9.0'
|
||||
MOCK_GPU_FREE_QUERY=$'0, 80000, 81920'
|
||||
export MOCK_GPU_QUERY MOCK_GPU_FREE_QUERY
|
||||
|
||||
out="$(run_case "$gemma_single" "vllm/gemma-mtp-tp1")"
|
||||
assert_contains "$out" "auto-selected GPU 0"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 10. --force skips the hardware gate.
|
||||
out="$(run_case "$gemma_single" "vllm/gemma-mtp-tp1" 1)"
|
||||
assert_contains "$out" "hardware: skipped"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
# 11. Missing metadata warns and allows.
|
||||
out="$(run_case "$missing_meta" "vllm/local")"
|
||||
assert_contains "$out" "no hardware metadata"
|
||||
assert_contains "$out" "STATUS=ok"
|
||||
|
||||
echo "test-preflight-vram: ok"
|
||||
Reference in New Issue
Block a user