Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e05f1969bc | ||
|
|
12a33fbdd1 | ||
|
|
45fa421e93 |
@@ -16,6 +16,16 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.5.4 — 2026-05-13
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(scripts): make submit-bench issue-first ([22bf2e9](https://github.com/noonghunna/club-3090/commit/22bf2e9398c7907aae6b62809bde30e111e4a700))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.4`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.5.3...v0.5.4)
|
||||
## v0.5.3 — 2026-05-13
|
||||
|
||||
|
||||
|
||||
+1
-1
@@ -8,7 +8,7 @@ Thanks for being here. This repo collects working recipes for serving big LLMs o
|
||||
|
||||
### ✅ Yes please
|
||||
|
||||
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
|
||||
- **Numbers from your rig.** Different power caps, different motherboards, different models — we want all of it. Use the [Numbers from your rig](https://github.com/noonghunna/club-3090/issues/new?template=numbers-from-your-rig.yml) issue template (no PR needed). The template asks for `bash scripts/report.sh --full > my-rig.md` — one ~35-min pass captures hardware (incl. power caps + NVLink topology), stack version, verify-full + verify-stress 7/7, **SOAK_MODE=continuous summary (catches Cliff 2b)**, AND the canonical bench numbers. If you're unsure what to run before measuring, start with `bash scripts/setup.sh` and `bash scripts/launch.sh`; both wizards mark hardware-fit. High-signal contributions land in `BENCHMARKS` with attribution. **Not running our Docker composes?** All scripts now work on non-Docker host builds (llama.cpp host server, SGLang, etc.) via `URL=... CONTAINER=none MODEL=... bash scripts/...` — engine is auto-detected, vLLM-specific checks skip cleanly. See [discussion #88](https://github.com/noonghunna/club-3090/discussions/88) for the full host-build contributor flow.
|
||||
- **Power-cap efficiency curves.** `sudo bash scripts/power-cap-sweep.sh --cooling air|water|aio --load-mode decode-concurrent --concurrency auto --bench-runs 3` produces cross-rig efficiency-knee data ([discussion #86](https://github.com/noonghunna/club-3090/discussions/86)). ~15-20 min for a 30-cap sweep on a 3090/4090/5090. Especially valuable on cards we don't have anchors for yet (A5000/A6000, 4080, 5060 Ti / 5080, modded variants). **Keep `--step-size 10` (the default).** Larger step-sizes (e.g. `--step-size 50`) are too coarse for the efficiency knee and only useful for quick smoke tests. See [docs/HARDWARE.md](docs/HARDWARE.md#cross-rig-power-cap-data-anchor-points) for the full canonical command and rationale.
|
||||
- **Bug reports with the data we ask for.** The [bug report template](https://github.com/noonghunna/club-3090/issues/new?template=bug-report.yml) leads with `bash scripts/report.sh > my-rig.md` (add `--verify` to include verify-full output, `--soak` to also run SOAK_MODE=continuous if you suspect a multi-turn agent cliff) — single command captures the rig state we'd otherwise ask for individually (hardware, container state, Genesis patches, KV pool sizing, engine config). With that paste, the first reply is usually a fix or a clear next step instead of "can you send me…".
|
||||
- **Bug reproductions / minimum repros for upstream issues.** vLLM / llama.cpp / Genesis bugs that affect this stack are most useful when they have a one-paragraph reduction. Drop them in an issue or open a draft PR adding a reproducer to `verify-stress.sh`.
|
||||
|
||||
@@ -66,13 +66,15 @@ Bench protocol: 3 warm + 5 measured runs of the canonical narrative + code promp
|
||||
git clone https://github.com/noonghunna/club-3090.git
|
||||
cd club-3090
|
||||
|
||||
# 2. Download + SHA-verify the model (~20 GB; clones Genesis patches too)
|
||||
# (asks you where to put model weights — pick in-repo default, ~/models, or
|
||||
# a custom path on a different drive. To skip the prompt:
|
||||
# `export MODEL_DIR=/mnt/your-drive/models` before running. See FAQ + .env.example.)
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
# 2. Pick/download + SHA-verify the model (interactive hardware-aware picker)
|
||||
# (asks you which model, then where to put model weights — pick in-repo
|
||||
# default, ~/models, or a custom path on a different drive. To skip prompts:
|
||||
# `export MODEL_DIR=/mnt/your-drive/models` and pass the model name. See FAQ.)
|
||||
bash scripts/setup.sh
|
||||
# Or scripted:
|
||||
# bash scripts/setup.sh qwen3.6-27b
|
||||
|
||||
# 3. Pick a config + boot it (interactive wizard — asks engine / cards / workload)
|
||||
# 3. Pick a config + boot it (interactive hardware-aware wizard — asks cards / workload)
|
||||
bash scripts/launch.sh
|
||||
# Or skip the wizard:
|
||||
# bash scripts/launch.sh --variant vllm/default # single-card chat (recommended)
|
||||
|
||||
+7
-1
@@ -130,6 +130,12 @@ If your numbers on the same compose look different from ours by >15%, the most l
|
||||
|
||||
## Setup
|
||||
|
||||
### How do I pick the right model + variant?
|
||||
|
||||
For a first install, run `bash scripts/setup.sh` with no model argument in a normal terminal. It opens a hardware-aware model picker, marks Qwen / Gemma / Both as eligible or not for your detected GPUs, then continues into the existing download flow.
|
||||
|
||||
After setup, run `bash scripts/launch.sh`. Its existing cards + workload wizard now marks compose variants with hardware fit and picks the recommended default for the rig (`vllm/long-text` on one 24 GB card, `vllm/dual` on matched 2× 3090). Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, plus `setup.sh --help` / `launch.sh --help`.
|
||||
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
|
||||
|
||||
Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
|
||||
@@ -140,7 +146,7 @@ Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
```
|
||||
2. **`.env` file at repo root** — picked up automatically on every script run. See [`.env.example`](../.env.example).
|
||||
3. **Interactive prompt** — `bash scripts/setup.sh qwen3.6-27b` with nothing set offers three choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
|
||||
3. **Interactive prompt** — `bash scripts/setup.sh` with nothing set first asks which model to download, then offers three model-dir choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
|
||||
4. **Silent fallback** — `<repo>/models-cache/`. Functional but pollutes the git tree; not recommended.
|
||||
|
||||
Every script that touches model paths reads from the same `MODEL_DIR`. The compose YAMLs' volume mount is `${MODEL_DIR:-...}:/root/.cache/huggingface` — once set, every container reads + writes there.
|
||||
|
||||
+145
-6
@@ -7,12 +7,14 @@
|
||||
# what you want, use `scripts/switch.sh <variant>` directly.
|
||||
#
|
||||
# Usage:
|
||||
# bash scripts/launch.sh # interactive wizard
|
||||
# bash scripts/launch.sh # interactive hardware-aware wizard
|
||||
# bash scripts/launch.sh --variant <name> # skip wizard, boot directly
|
||||
# bash scripts/launch.sh --engine vllm --cards 1 # partial flags, ask the rest
|
||||
# bash scripts/launch.sh --no-verify # skip post-launch verify-full
|
||||
# bash scripts/launch.sh --no-preflight # skip docker/GPU pre-flight
|
||||
#
|
||||
# The wizard marks variants that don't fit the detected GPUs; direct
|
||||
# --variant keeps the power-user path and delegates final gating to switch.sh.
|
||||
# All flags accept the same names as `switch.sh --list` produces.
|
||||
# Examples:
|
||||
# bash scripts/launch.sh --variant vllm/default
|
||||
@@ -87,7 +89,12 @@ choose() {
|
||||
done
|
||||
while true; do
|
||||
local pick
|
||||
read -rp "Choice [1-${#labels[@]}]: " pick
|
||||
if ! read -rp "Choice [1-${#labels[@]}]: " pick; then
|
||||
echo "" >&2
|
||||
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
|
||||
kill -INT $$
|
||||
exit 1
|
||||
fi
|
||||
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
|
||||
echo "${values[$((pick-1))]}"
|
||||
return
|
||||
@@ -96,6 +103,121 @@ choose() {
|
||||
done
|
||||
}
|
||||
|
||||
declare -A LAUNCH_VARIANT_COMPOSE=(
|
||||
[vllm/default]="models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
|
||||
[vllm/long-vision]="models/qwen3.6-27b/vllm/compose/single/long-vision.yml"
|
||||
[vllm/long-text]="models/qwen3.6-27b/vllm/compose/single/long-text.yml"
|
||||
[vllm/long-text-no-mtp]="models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml"
|
||||
[vllm/bounded-thinking]="models/qwen3.6-27b/vllm/compose/single/bounded-thinking.yml"
|
||||
[vllm/tools-text]="models/qwen3.6-27b/vllm/compose/single/tools-text.yml"
|
||||
[vllm/minimal]="models/qwen3.6-27b/vllm/compose/single/minimal.yml"
|
||||
[vllm/dual]="models/qwen3.6-27b/vllm/compose/dual/docker-compose.yml"
|
||||
[vllm/dual4]="models/qwen3.6-27b/vllm/compose/multi4/docker-compose.yml"
|
||||
[vllm/dual4-dflash]="models/qwen3.6-27b/vllm/compose/multi4/dflash.yml"
|
||||
[vllm/dual-turbo]="models/qwen3.6-27b/vllm/compose/dual/turbo.yml"
|
||||
[vllm/dual-dflash]="models/qwen3.6-27b/vllm/compose/dual/dflash.yml"
|
||||
[vllm/dual-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/dflash-noviz.yml"
|
||||
[vllm/dual-nvlink]="models/qwen3.6-27b/vllm/compose/dual/nvlink.yml"
|
||||
[vllm/dual-nvlink-turbo]="models/qwen3.6-27b/vllm/compose/dual/nvlink-turbo.yml"
|
||||
[vllm/dual-nvlink-dflash]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash.yml"
|
||||
[vllm/dual-nvlink-dflash-noviz]="models/qwen3.6-27b/vllm/compose/dual/nvlink-dflash-noviz.yml"
|
||||
[vllm/gemma-mtp]="models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
|
||||
[vllm/gemma-mtp-tp1]="models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
|
||||
[vllm/gemma-dflash]="models/gemma-4-31b/vllm/compose/dual/dflash.yml"
|
||||
)
|
||||
|
||||
variant_hw_status() {
|
||||
local variant="$1"
|
||||
local rel="${LAUNCH_VARIANT_COMPOSE[$variant]:-}"
|
||||
if [[ -z "$rel" ]]; then
|
||||
printf 'ok|fits your rig'
|
||||
return 0
|
||||
fi
|
||||
|
||||
local compose_file="${ROOT_DIR}/${rel}"
|
||||
if [[ ! -f "$compose_file" ]]; then
|
||||
printf 'unknown|compose metadata unavailable'
|
||||
return 2
|
||||
fi
|
||||
compose_hw_compose_status "$compose_file" 2>/dev/null || true
|
||||
}
|
||||
|
||||
choose_variant() {
|
||||
# choose_variant "prompt" "default-variant" "label1" "value1" ...
|
||||
local prompt="$1" default_variant="$2"
|
||||
shift 2
|
||||
|
||||
local i labels=() values=() statuses=() eligible=()
|
||||
while [[ $# -gt 0 ]]; do
|
||||
labels+=("$1")
|
||||
values+=("$2")
|
||||
statuses+=("$(variant_hw_status "$2")")
|
||||
shift 2
|
||||
done
|
||||
|
||||
local default_idx=""
|
||||
for i in "${!values[@]}"; do
|
||||
if [[ "${values[$i]}" == "$default_variant" && "${statuses[$i]}" == ok\|* ]]; then
|
||||
default_idx=$((i + 1))
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [[ -z "$default_idx" ]]; then
|
||||
for i in "${!values[@]}"; do
|
||||
if [[ "${statuses[$i]}" == ok\|* || "${statuses[$i]}" == unknown\|* ]]; then
|
||||
default_idx=$((i + 1))
|
||||
break
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
echo "" >&2
|
||||
echo "$prompt" >&2
|
||||
for i in "${!labels[@]}"; do
|
||||
local status="${statuses[$i]}"
|
||||
local state="${status%%|*}"
|
||||
local reason="${status#*|}"
|
||||
local marker="✓"
|
||||
case "$state" in
|
||||
ok) marker="✓" ;;
|
||||
unknown) marker="?" ;;
|
||||
*) marker="✗" ;;
|
||||
esac
|
||||
if [[ -n "$default_idx" && $((i + 1)) -eq "$default_idx" ]]; then
|
||||
printf " %d) %s %s %s [default]\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
|
||||
else
|
||||
printf " %d) %s %s %s\n" "$((i + 1))" "${labels[$i]}" "$marker" "$reason" >&2
|
||||
fi
|
||||
done
|
||||
|
||||
if [[ -z "$default_idx" ]]; then
|
||||
echo "ERROR: no eligible variants in this menu. Use scripts/switch.sh --force <variant> to attempt anyway." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
while true; do
|
||||
local pick
|
||||
if ! read -rp "Choice [1-${#labels[@]}, default ${default_idx}]: " pick; then
|
||||
echo "" >&2
|
||||
echo " EOF on stdin — wizard needs interactive input. Use --variant <name> to skip." >&2
|
||||
kill -INT $$
|
||||
exit 1
|
||||
fi
|
||||
pick="${pick:-$default_idx}"
|
||||
if [[ "$pick" =~ ^[0-9]+$ ]] && (( pick >= 1 && pick <= ${#labels[@]} )); then
|
||||
local status="${statuses[$((pick - 1))]}"
|
||||
if [[ "$status" == no\|* ]]; then
|
||||
echo " That variant won't run on your detected rig: ${status#*|}" >&2
|
||||
echo " Pick another, or use: bash scripts/switch.sh --force ${values[$((pick - 1))]}" >&2
|
||||
continue
|
||||
fi
|
||||
echo "${values[$((pick - 1))]}"
|
||||
return
|
||||
fi
|
||||
echo " invalid — pick a number 1-${#labels[@]}" >&2
|
||||
done
|
||||
}
|
||||
|
||||
# --- wizard ---
|
||||
# Flow: cards → workload → auto-pick engine. Newcomers can answer "how
|
||||
# many GPUs" and "what do I want to do" but rarely "vLLM or llama.cpp" —
|
||||
@@ -151,8 +273,15 @@ if [[ -z "$VARIANT" ]]; then
|
||||
"[fallback] tools-text 75K FP8 (FP8 KV alternative for accuracy compare)" "vllm/tools-text"
|
||||
"[fallback] minimal 32K (no Genesis, no spec-decode — diagnostic stack)" "vllm/minimal"
|
||||
)
|
||||
VLLM_DUAL_OPTS=(
|
||||
"[2-card] vllm/dual — 262K + vision + 2 streams" "vllm/dual"
|
||||
"[2-card] vllm/dual-turbo — 4 streams @ 262K, TQ3 KV" "vllm/dual-turbo"
|
||||
"[2-card] vllm/dual-dflash — peak code TPS with vision" "vllm/dual-dflash"
|
||||
"[2-card] vllm/gemma-mtp — Gemma 4 dual-card default" "vllm/gemma-mtp"
|
||||
)
|
||||
else
|
||||
VLLM_FALLBACK_OPTS=()
|
||||
VLLM_DUAL_OPTS=()
|
||||
fi
|
||||
if [[ -z "$ENGINE" || "$ENGINE" == "llamacpp" ]]; then
|
||||
LLAMA_FALLBACK_OPTS=(
|
||||
@@ -161,19 +290,29 @@ if [[ -z "$VARIANT" ]]; then
|
||||
else
|
||||
LLAMA_FALLBACK_OPTS=()
|
||||
fi
|
||||
VARIANT=$(choose "What's your main workload?" \
|
||||
if [[ -n "$ENGINE" && "$ENGINE" != "vllm" && "$ENGINE" != "llamacpp" ]]; then
|
||||
echo "ERROR: --engine ${ENGINE} unsupported (expected vllm or llamacpp)." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ "$ENGINE" == "llamacpp" ]]; then
|
||||
DEFAULT_VARIANT="llamacpp/default"
|
||||
else
|
||||
DEFAULT_VARIANT="vllm/long-text"
|
||||
fi
|
||||
VARIANT=$(choose_variant "What's your main workload?" "$DEFAULT_VARIANT" \
|
||||
"${VLLM_OPTS[@]}" "${LLAMA_OPTS[@]}" \
|
||||
"${VLLM_FALLBACK_OPTS[@]}" "${LLAMA_FALLBACK_OPTS[@]}")
|
||||
"${VLLM_FALLBACK_OPTS[@]}" "${VLLM_DUAL_OPTS[@]}" "${LLAMA_FALLBACK_OPTS[@]}")
|
||||
elif [[ "$CARDS" == "2" ]]; then
|
||||
if [[ -n "$ENGINE" && "$ENGINE" != "vllm" ]]; then
|
||||
echo "ERROR: --engine ${ENGINE} not supported on 2× cards (no llama.cpp dual recipe yet)." >&2
|
||||
exit 1
|
||||
fi
|
||||
VARIANT=$(choose "What's your dual-card priority?" \
|
||||
VARIANT=$(choose_variant "What's your dual-card priority?" "vllm/dual" \
|
||||
"Balanced default — 262K + vision + 2 streams (recommended)" "vllm/dual" \
|
||||
"Multi-tenant — 4 concurrent streams @ 262K, TQ3 KV" "vllm/dual-turbo" \
|
||||
"Peak code TPS with vision (185K, DFlash N=5)" "vllm/dual-dflash" \
|
||||
"Peak code TPS no vision (200K, DFlash N=5)" "vllm/dual-dflash-noviz")
|
||||
"Peak code TPS no vision (200K, DFlash N=5)" "vllm/dual-dflash-noviz" \
|
||||
"Gemma 4 dual-card default — 32K + vision + MTP" "vllm/gemma-mtp")
|
||||
else
|
||||
echo "ERROR: --cards ${CARDS} unsupported (expected 1 or 2)." >&2
|
||||
exit 1
|
||||
|
||||
@@ -62,3 +62,229 @@ compose_meta_get() {
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
compose_hw_sm_to_int() {
|
||||
local sm="$1"
|
||||
sm="${sm%%+}"
|
||||
sm="${sm//sm_/}"
|
||||
sm="${sm//SM_/}"
|
||||
sm="${sm// /}"
|
||||
[[ -z "$sm" ]] && { echo 0; return; }
|
||||
|
||||
local major minor
|
||||
if [[ "$sm" == *.* ]]; then
|
||||
major="${sm%%.*}"
|
||||
minor="${sm#*.}"
|
||||
else
|
||||
major="$sm"
|
||||
minor="0"
|
||||
fi
|
||||
major="${major//[^0-9]/}"
|
||||
minor="${minor//[^0-9]/}"
|
||||
[[ -z "$major" ]] && major=0
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
if [[ "${#minor}" -eq 1 ]]; then
|
||||
minor=$(( minor * 10 ))
|
||||
else
|
||||
minor="${minor:0:2}"
|
||||
[[ -z "$minor" ]] && minor=0
|
||||
fi
|
||||
echo $(( major * 100 + minor ))
|
||||
}
|
||||
|
||||
compose_hw_vram_gb() {
|
||||
local mib="$1"
|
||||
echo $(( (mib + 1023) / 1024 ))
|
||||
}
|
||||
|
||||
compose_hw_detect_gpus() {
|
||||
if [[ "${_COMPOSE_HW_GPU_CACHE_SET:-0}" == "1" ]]; then
|
||||
[[ -n "${_COMPOSE_HW_GPU_CACHE:-}" ]] || return 1
|
||||
printf '%s\n' "${_COMPOSE_HW_GPU_CACHE}"
|
||||
return 0
|
||||
fi
|
||||
|
||||
command -v nvidia-smi >/dev/null 2>&1 || return 1
|
||||
|
||||
local query idx name mem_mib sm rest
|
||||
query="$(nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv,noheader,nounits 2>/dev/null)" || return 1
|
||||
[[ -n "$query" ]] || return 1
|
||||
|
||||
local parsed=""
|
||||
while IFS=',' read -r idx name mem_mib sm rest; do
|
||||
idx="$(_compose_meta_trim "$idx")"
|
||||
name="$(_compose_meta_trim "$name")"
|
||||
mem_mib="$(_compose_meta_trim "$mem_mib")"
|
||||
sm="$(_compose_meta_trim "$sm")"
|
||||
[[ -z "$idx" || -z "$mem_mib" ]] && continue
|
||||
parsed+="${idx}"$'\t'"${name}"$'\t'"${mem_mib}"$'\t'"${sm}"$'\n'
|
||||
done <<< "$query"
|
||||
|
||||
parsed="${parsed%$'\n'}"
|
||||
_COMPOSE_HW_GPU_CACHE_SET=1
|
||||
_COMPOSE_HW_GPU_CACHE="$parsed"
|
||||
[[ -n "$parsed" ]] || return 1
|
||||
printf '%s\n' "$parsed"
|
||||
}
|
||||
|
||||
compose_hw_summary() {
|
||||
local gpu_lines
|
||||
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
|
||||
if [[ -z "$gpu_lines" ]]; then
|
||||
printf 'no NVIDIA GPUs detected'
|
||||
return 0
|
||||
fi
|
||||
|
||||
local count=0 first_name="" first_gb="" mixed=0 idx name mem_mib sm
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
local gb
|
||||
gb="$(compose_hw_vram_gb "$mem_mib")"
|
||||
name="${name#NVIDIA }"
|
||||
name="${name#GeForce }"
|
||||
count=$((count + 1))
|
||||
if [[ -z "$first_name" ]]; then
|
||||
first_name="$name"
|
||||
first_gb="$gb"
|
||||
elif [[ "$name" != "$first_name" || "$gb" != "$first_gb" ]]; then
|
||||
mixed=1
|
||||
fi
|
||||
done <<< "$gpu_lines"
|
||||
|
||||
if (( count == 0 )); then
|
||||
printf 'no NVIDIA GPUs detected'
|
||||
elif (( mixed == 0 )); then
|
||||
if (( count == 1 )); then
|
||||
printf '1× %s, %s GB' "$first_name" "$first_gb"
|
||||
else
|
||||
printf '%d× %s, %s GB each' "$count" "$first_name" "$first_gb"
|
||||
fi
|
||||
else
|
||||
local parts=()
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
name="${name#NVIDIA }"
|
||||
name="${name#GeForce }"
|
||||
parts+=("${name}, $(compose_hw_vram_gb "$mem_mib") GB")
|
||||
done <<< "$gpu_lines"
|
||||
local joined=""
|
||||
for part in "${parts[@]}"; do
|
||||
if [[ -z "$joined" ]]; then
|
||||
joined="$part"
|
||||
else
|
||||
joined="${joined} + ${part}"
|
||||
fi
|
||||
done
|
||||
printf '%s' "$joined"
|
||||
fi
|
||||
}
|
||||
|
||||
compose_hw_requirement_text() {
|
||||
local min_vram_gb="$1"
|
||||
local min_gpu_count="$2"
|
||||
local requires_sm="${3:-}"
|
||||
|
||||
local req
|
||||
if [[ "$min_gpu_count" == "1" ]]; then
|
||||
req="${min_vram_gb} GB+"
|
||||
else
|
||||
req="${min_gpu_count}× ${min_vram_gb} GB"
|
||||
fi
|
||||
if [[ -n "$requires_sm" && "$requires_sm" != "0.0" ]]; then
|
||||
req="${req}, sm_${requires_sm%%+}+"
|
||||
fi
|
||||
printf '%s' "$req"
|
||||
}
|
||||
|
||||
compose_hw_compose_status() {
|
||||
local compose_file="$1"
|
||||
local min_vram_gb min_gpu_count requires_sm
|
||||
|
||||
min_vram_gb="$(compose_meta_get "$compose_file" requires-min-vram-gb || true)"
|
||||
min_gpu_count="$(compose_meta_get "$compose_file" requires-min-gpu-count || true)"
|
||||
requires_sm="$(compose_meta_get "$compose_file" requires-sm || true)"
|
||||
|
||||
if [[ -z "$min_vram_gb" || -z "$min_gpu_count" ]]; then
|
||||
printf 'unknown|metadata unavailable'
|
||||
return 2
|
||||
fi
|
||||
|
||||
requires_sm="${requires_sm:-0.0}"
|
||||
local required_sm_int
|
||||
required_sm_int="$(compose_hw_sm_to_int "$requires_sm")"
|
||||
|
||||
local gpu_lines
|
||||
gpu_lines="$(compose_hw_detect_gpus 2>/dev/null || true)"
|
||||
if [[ -z "$gpu_lines" ]]; then
|
||||
printf 'no|no NVIDIA GPUs detected'
|
||||
return 1
|
||||
fi
|
||||
|
||||
local eligible_count=0 idx name mem_mib sm gb sm_int
|
||||
while IFS=$'\t' read -r idx name mem_mib sm; do
|
||||
[[ -z "$idx" ]] && continue
|
||||
gb="$(compose_hw_vram_gb "$mem_mib")"
|
||||
sm_int="$(compose_hw_sm_to_int "$sm")"
|
||||
if (( gb >= min_vram_gb && sm_int >= required_sm_int )); then
|
||||
eligible_count=$((eligible_count + 1))
|
||||
fi
|
||||
done <<< "$gpu_lines"
|
||||
|
||||
if (( eligible_count >= min_gpu_count )); then
|
||||
printf 'ok|fits your rig'
|
||||
return 0
|
||||
fi
|
||||
|
||||
printf 'no|needs %s (your rig: %s)' \
|
||||
"$(compose_hw_requirement_text "$min_vram_gb" "$min_gpu_count" "$requires_sm")" \
|
||||
"$(compose_hw_summary)"
|
||||
return 1
|
||||
}
|
||||
|
||||
compose_hw_compose_eligible() {
|
||||
local status
|
||||
status="$(compose_hw_compose_status "$1" 2>/dev/null || true)"
|
||||
[[ "$status" == ok\|* ]]
|
||||
}
|
||||
|
||||
compose_hw_model_status() {
|
||||
local repo_root="$1"
|
||||
local model="$2"
|
||||
local candidates=()
|
||||
local friendly_need=""
|
||||
|
||||
case "$model" in
|
||||
qwen3.6-27b)
|
||||
candidates=(
|
||||
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/long-text.yml"
|
||||
"${repo_root}/models/qwen3.6-27b/vllm/compose/single/docker-compose.yml"
|
||||
)
|
||||
friendly_need="needs 20 GB+ VRAM (24 GB recommended)"
|
||||
;;
|
||||
gemma-4-31b)
|
||||
candidates=(
|
||||
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/docker-compose.yml"
|
||||
"${repo_root}/models/gemma-4-31b/vllm/compose/dual/int8.yml"
|
||||
"${repo_root}/models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
|
||||
)
|
||||
friendly_need="needs 32 GB+ on single card OR 2× 24 GB"
|
||||
;;
|
||||
*)
|
||||
printf 'no|unknown model: %s' "$model"
|
||||
return 1
|
||||
;;
|
||||
esac
|
||||
|
||||
local file status
|
||||
for file in "${candidates[@]}"; do
|
||||
[[ -f "$file" ]] || continue
|
||||
status="$(compose_hw_compose_status "$file" 2>/dev/null || true)"
|
||||
if [[ "$status" == ok\|* ]]; then
|
||||
printf 'ok|fits your rig'
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
|
||||
printf 'no|%s (your rig: %s)' "$friendly_need" "$(compose_hw_summary)"
|
||||
return 1
|
||||
}
|
||||
|
||||
+96
-7
@@ -2,7 +2,8 @@
|
||||
#
|
||||
# Model-aware one-shot setup for club-3090.
|
||||
#
|
||||
# bash scripts/setup.sh <model-name>
|
||||
# bash scripts/setup.sh # interactive model picker in a TTY
|
||||
# bash scripts/setup.sh <model-name> # scripted/CI positional form
|
||||
#
|
||||
# Currently supported:
|
||||
# qwen3.6-27b → Lorbus/Qwen3.6-27B-int4-AutoRound + Genesis patches
|
||||
@@ -37,15 +38,89 @@
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
# ---------- Model dispatch ----------
|
||||
MODEL_NAME="${1:-}"
|
||||
if [[ -z "${MODEL_NAME}" ]]; then
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
usage() {
|
||||
echo "Usage: $0 <model-name>"
|
||||
echo " $0 # interactive model picker in a TTY"
|
||||
echo ""
|
||||
echo "Run with no model name in a normal terminal to open the hardware-aware"
|
||||
echo "model picker. Use the positional form in scripts/CI to skip prompts."
|
||||
echo ""
|
||||
echo "Supported model names:"
|
||||
echo " qwen3.6-27b"
|
||||
echo " gemma-4-31b"
|
||||
exit 1
|
||||
}
|
||||
|
||||
model_label() {
|
||||
case "$1" in
|
||||
qwen3.6-27b) echo "Qwen 3.6 27B" ;;
|
||||
gemma-4-31b) echo "Gemma 4 31B" ;;
|
||||
*) echo "$1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
model_picker_line() {
|
||||
local idx="$1" model="$2" size="$3" status mark reason
|
||||
status="$(compose_hw_model_status "$ROOT_DIR" "$model" 2>/dev/null || true)"
|
||||
reason="${status#*|}"
|
||||
if [[ "$status" == ok\|* ]]; then
|
||||
mark="✓"
|
||||
else
|
||||
mark="✗"
|
||||
fi
|
||||
printf " %s. %-14s (%s) %s %s\n" "$idx" "$(model_label "$model")" "$size" "$mark" "$reason"
|
||||
}
|
||||
|
||||
pick_model_interactive() {
|
||||
# shellcheck source=lib/compose-meta.sh
|
||||
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
|
||||
|
||||
echo "[setup] Which model to download?" >&2
|
||||
echo "" >&2
|
||||
model_picker_line "1" "qwen3.6-27b" "~14 GB AutoRound INT4" >&2
|
||||
model_picker_line "2" "gemma-4-31b" "~21 GB AutoRound INT4 + drafter" >&2
|
||||
echo " 3. Both (~30 GB total) downloads both model families" >&2
|
||||
echo "" >&2
|
||||
while true; do
|
||||
local pick
|
||||
read -rp "Choice [1-3]: " pick
|
||||
case "$pick" in
|
||||
1) echo "qwen3.6-27b"; return ;;
|
||||
2) echo "gemma-4-31b"; return ;;
|
||||
3) echo "both"; return ;;
|
||||
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
|
||||
esac
|
||||
done
|
||||
}
|
||||
|
||||
# ---------- Model dispatch ----------
|
||||
case "${1:-}" in
|
||||
-h|--help)
|
||||
usage
|
||||
exit 0
|
||||
;;
|
||||
esac
|
||||
|
||||
MODEL_NAME="${1:-}"
|
||||
if [[ -z "${MODEL_NAME}" ]]; then
|
||||
if [[ -t 0 && -t 1 ]]; then
|
||||
MODEL_NAME="$(pick_model_interactive)"
|
||||
else
|
||||
usage
|
||||
echo ""
|
||||
echo "(Interactive picker available in a TTY shell. Use the positional form in scripts/CI.)"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ "${MODEL_NAME}" == "both" ]]; then
|
||||
# Resolve MODEL_DIR once in the parent by reusing the normal prompt below,
|
||||
# then recurse through the positional form for each model.
|
||||
SETUP_BOTH_MODE=1
|
||||
MODEL_NAME="qwen3.6-27b"
|
||||
else
|
||||
SETUP_BOTH_MODE=0
|
||||
fi
|
||||
|
||||
# ALWAYS_DRAFT_REPO + ALWAYS_DRAFT_SUBDIR: a drafter that this model REQUIRES
|
||||
@@ -89,8 +164,6 @@ case "${MODEL_NAME}" in
|
||||
;;
|
||||
esac
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
# ---------- MODEL_DIR resolution ----------
|
||||
# Order of precedence:
|
||||
# 1. MODEL_DIR already exported in the calling shell → use as-is
|
||||
@@ -158,6 +231,18 @@ fi
|
||||
|
||||
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
|
||||
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
|
||||
if [[ "${SETUP_BOTH_MODE:-0}" == "1" ]]; then
|
||||
export MODEL_DIR
|
||||
echo "[setup] downloading both supported models into ${MODEL_DIR}"
|
||||
echo ""
|
||||
bash "$0" qwen3.6-27b
|
||||
echo ""
|
||||
bash "$0" gemma-4-31b
|
||||
echo ""
|
||||
echo "[setup] ✓ Both models downloaded."
|
||||
echo "[setup] Next: bash scripts/launch.sh"
|
||||
exit 0
|
||||
fi
|
||||
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
|
||||
|
||||
cd "${ROOT_DIR}"
|
||||
@@ -445,6 +530,7 @@ echo ""
|
||||
# refactored 2026-05-03 to vendor the two files in-repo, fixing #37.)
|
||||
|
||||
# Per-model "next steps" — different composes / served-model-name / port between models.
|
||||
SETUP_MODEL_DISPLAY="$(model_label "${MODEL_NAME}")"
|
||||
case "${MODEL_NAME}" in
|
||||
qwen3.6-27b)
|
||||
SAMPLE_CONTAINER="vllm-qwen36-27b"
|
||||
@@ -468,6 +554,9 @@ case "${MODEL_NAME}" in
|
||||
;;
|
||||
esac
|
||||
|
||||
echo "[setup] ✓ ${SETUP_MODEL_DISPLAY} downloaded."
|
||||
echo "[setup] Next: bash scripts/launch.sh"
|
||||
echo ""
|
||||
echo "Next — single-card vLLM (default):"
|
||||
if [[ "${MODEL_NAME}" == "gemma-4-31b" ]]; then
|
||||
echo " bash scripts/switch.sh vllm/gemma-mtp"
|
||||
|
||||
Executable
+162
@@ -0,0 +1,162 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||
TMP_DIR="$(mktemp -d)"
|
||||
ORIG_PATH="$PATH"
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
assert_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" != *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
assert_not_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
if [[ "$haystack" == *"$needle"* ]]; then
|
||||
echo "ASSERTION FAILED: expected output not to contain: $needle" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$haystack" >&2
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
make_mock_tools() {
|
||||
mkdir -p "${TMP_DIR}/bin"
|
||||
cat > "${TMP_DIR}/bin/nvidia-smi" <<'MOCK_NVIDIA_SMI'
|
||||
#!/usr/bin/env bash
|
||||
case "$*" in
|
||||
*"--query-gpu=index,name,memory.total,compute_cap"*)
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}"
|
||||
;;
|
||||
"-L")
|
||||
printf '%s\n' "${MOCK_GPU_QUERY:?MOCK_GPU_QUERY not set}" \
|
||||
| awk -F, '{gsub(/^[ \t]+|[ \t]+$/, "", $1); gsub(/^[ \t]+|[ \t]+$/, "", $2); print "GPU " $1 ": " $2}'
|
||||
;;
|
||||
*)
|
||||
echo "unexpected nvidia-smi invocation: $*" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
MOCK_NVIDIA_SMI
|
||||
chmod +x "${TMP_DIR}/bin/nvidia-smi"
|
||||
|
||||
cat > "${TMP_DIR}/switch-mock" <<'MOCK_SWITCH'
|
||||
#!/usr/bin/env bash
|
||||
echo "SWITCHED $*"
|
||||
MOCK_SWITCH
|
||||
chmod +x "${TMP_DIR}/switch-mock"
|
||||
|
||||
export PATH="${TMP_DIR}/bin:${ORIG_PATH}"
|
||||
}
|
||||
|
||||
set_rig() {
|
||||
export MOCK_GPU_QUERY="$1"
|
||||
}
|
||||
|
||||
model_status() {
|
||||
local model="$1"
|
||||
(
|
||||
source "${ROOT_DIR}/scripts/lib/compose-meta.sh"
|
||||
compose_hw_model_status "$ROOT_DIR" "$model"
|
||||
)
|
||||
}
|
||||
|
||||
assert_model_status() {
|
||||
local model="$1"
|
||||
local expected_prefix="$2"
|
||||
local expected_text="${3:-}"
|
||||
local status
|
||||
status="$(model_status "$model" || true)"
|
||||
if [[ "$status" != "${expected_prefix}"* ]]; then
|
||||
echo "ASSERTION FAILED: ${model} status expected prefix '${expected_prefix}', got '${status}'" >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ -n "$expected_text" ]]; then
|
||||
assert_contains "$status" "$expected_text"
|
||||
fi
|
||||
}
|
||||
|
||||
make_mock_tools
|
||||
|
||||
# Matched 2x3090: Qwen and Gemma both have a viable compose.
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "ok|fits your rig"
|
||||
|
||||
# Single 24 GB Ampere: Qwen fits; Gemma needs either 32 GB+ single-card or 2x24 GB.
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
|
||||
assert_contains "$(model_status "gemma-4-31b")" "1× RTX 3090, 24 GB"
|
||||
|
||||
# Single 16 GB: neither shipped model has a viable compose.
|
||||
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9'
|
||||
assert_model_status "qwen3.6-27b" "no|" "needs 20 GB+ VRAM"
|
||||
assert_model_status "gemma-4-31b" "no|" "needs 32 GB+ on single card OR 2× 24 GB"
|
||||
|
||||
# Heterogeneous 16 + 24 GB: Qwen can run on the 24 GB card; Gemma dual cannot.
|
||||
set_rig $'0, NVIDIA RTX 4060 Ti, 16384, 8.9\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "no|" "RTX 4060 Ti, 16 GB + RTX 3090, 24 GB"
|
||||
|
||||
# 32 GB+ modern card: Gemma's single-card compose is eligible.
|
||||
set_rig $'0, NVIDIA GeForce RTX 5090, 32768, 12.0'
|
||||
assert_model_status "qwen3.6-27b" "ok|fits your rig"
|
||||
assert_model_status "gemma-4-31b" "ok|fits your rig"
|
||||
|
||||
# Non-TTY no-arg setup fails fast with usage rather than hanging.
|
||||
if out="$(echo | bash "${ROOT_DIR}/scripts/setup.sh" 2>&1)"; then
|
||||
echo "ASSERTION FAILED: non-TTY no-arg setup unexpectedly succeeded" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
assert_contains "$out" "Usage:"
|
||||
assert_contains "$out" "Interactive picker available in a TTY shell"
|
||||
|
||||
# Positional setup path remains non-interactive and reaches the existing flow.
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
out="$(MODEL_DIR="${TMP_DIR}/models" PREFLIGHT_DISK_GB=0 SKIP_GENESIS=1 SKIP_MODEL=1 bash "${ROOT_DIR}/scripts/setup.sh" qwen3.6-27b 2>&1)"
|
||||
assert_not_contains "$out" "Which model to download?"
|
||||
assert_contains "$out" "[model] SKIP_MODEL=1"
|
||||
|
||||
# The launch wizard's variant-display step marks hardware viability and defaults
|
||||
# to the single-card long-text recommendation on a single 24 GB card.
|
||||
out="$(printf '\n' | SWITCH="${TMP_DIR}/switch-mock" bash "${ROOT_DIR}/scripts/launch.sh" --no-preflight --no-verify --engine vllm --cards 1 2>&1)"
|
||||
assert_contains "$out" "Long ctx, text only — Balanced MTP"
|
||||
assert_contains "$out" "[default]"
|
||||
assert_contains "$out" "vllm/dual"
|
||||
assert_contains "$out" "✗ needs 2× 24 GB"
|
||||
assert_contains "$out" "SWITCHED vllm/long-text"
|
||||
|
||||
# TTY-backed no-arg setup supports the cosmetic but real "Both" choice by
|
||||
# dispatching through the positional path for both model families.
|
||||
if ! command -v script >/dev/null 2>&1; then
|
||||
echo "ASSERTION FAILED: util-linux 'script' is required for TTY picker coverage" >&2
|
||||
exit 1
|
||||
fi
|
||||
set_rig $'0, NVIDIA GeForce RTX 3090, 24576, 8.6\n1, NVIDIA GeForce RTX 3090, 24576, 8.6'
|
||||
export MODEL_DIR="${TMP_DIR}/models"
|
||||
export PREFLIGHT_DISK_GB=0
|
||||
export SKIP_GENESIS=1
|
||||
export SKIP_MODEL=1
|
||||
out="$(printf '3\n' | script -qec "bash '${ROOT_DIR}/scripts/setup.sh'" /dev/null 2>&1)"
|
||||
assert_contains "$out" "[setup] Which model to download?"
|
||||
assert_contains "$out" "Both"
|
||||
assert_contains "$out" "[setup] downloading both supported models"
|
||||
skip_count="$(grep -c "\[model\] SKIP_MODEL=1" <<< "$out" || true)"
|
||||
if [[ "$skip_count" != "2" ]]; then
|
||||
echo "ASSERTION FAILED: expected Both choice to dispatch two model setup runs, got ${skip_count}" >&2
|
||||
echo "--- output ---" >&2
|
||||
echo "$out" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "test-setup-picker: ok"
|
||||
Reference in New Issue
Block a user