PR-B of the compose-quant-hierarchy refactor (follows PR-A #231). launch.sh no longer hardcodes its variant/port/container/kvcalc maps — it derives them from compose_registry.py via a shared emitter, the same single source of truth switch.sh already uses. Adds a topology-aware <engine>/default resolver. - scripts/lib/registry-emit.sh (new): shared emitter — emits switch/launch tables from COMPOSE_REGISTRY, parses container_name from each compose, and exposes registry_default_target() for <engine>[/<topology>]/default. - launch.sh: drop the hardcoded `declare -A LAUNCH_*` maps; populate from the emitter. PRIMARY_MODEL constant (default qwen3.6-27b) backs bare <engine>/default. - switch.sh: drop inline derivation; resolve vllm/default, vllm/dual/default, vllm/multi4/default via the shared emitter. - compose_registry.py: add kvcalc_key to entries; tools/kv-calc.py aliases. - bench-row-formatter.sh: DEFAULTS-aware compose_display (drops the now-dead docker-compose.yml branch). - tests: test-launch-registry-parity.sh + test-default-resolver.sh (new); test-switch-registry-parity refreshed onto the shared emitter. Default resolution is an alias layer over existing registry keys — no keys renamed, no compose content changed. Side effect: the launcher<->registry drift that left vllm/gemma-mtp pointing at a non-existent dual/.../fp8-mtp.yml is gone (now derives the registry's bf16-mtp.yml). Implemented via Codex per the PR-B brief; independently re-validated before commit: 7/7 gates PASS (switch+launch parity, default-resolver, launch-compat, registry-disk, mounts-resolve, kv-calc 22/22); live vllm/default->dual + vllm/dual/default->dual both verify-full 8/8; leak-clean. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
47 lines
2.0 KiB
Bash
Executable File
47 lines
2.0 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# PR-B — <engine>/default resolver uses DEFAULTS + detected topology.
|
|
set -euo pipefail
|
|
|
|
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
|
|
cd "$ROOT_DIR"
|
|
|
|
assert_contains() {
|
|
local haystack="$1" needle="$2"
|
|
if [[ "$haystack" != *"$needle"* ]]; then
|
|
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
|
echo "--- output ---" >&2
|
|
echo "$haystack" >&2
|
|
exit 1
|
|
fi
|
|
}
|
|
|
|
fake_one='0:RTX_3090:24576:8.6'
|
|
fake_two='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6'
|
|
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection vllm/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/default"
|
|
assert_contains "$out" "vllm/default"
|
|
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_two" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection vllm/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/dual"
|
|
assert_contains "$out" "vllm/dual"
|
|
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection vllm/dual/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/dual"
|
|
|
|
if out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection llamacpp/dual/default 2>&1)"; then
|
|
echo "ASSERTION FAILED: bad topology default unexpectedly resolved" >&2
|
|
echo "$out" >&2
|
|
exit 1
|
|
fi
|
|
assert_contains "$out" "cannot resolve default variant 'llamacpp/dual/default'"
|
|
assert_contains "$out" "Available defaults"
|
|
|
|
out="$(NVIDIA_VISIBLE_DEVICES=0,1 FORCE=1 PREFLIGHT_NO_COMPOSE_DEPS=1 COMPOSE_BIN=: READY_TIMEOUT=1 bash scripts/switch.sh --no-wait vllm/default 2>&1 || true)"
|
|
assert_contains "$out" "bringing up: vllm/dual"
|
|
|
|
out="$(NVIDIA_VISIBLE_DEVICES=0 FORCE=1 PREFLIGHT_NO_COMPOSE_DEPS=1 COMPOSE_BIN=: READY_TIMEOUT=1 bash scripts/switch.sh --no-wait vllm/dual/default 2>&1 || true)"
|
|
assert_contains "$out" "bringing up: vllm/dual"
|
|
|
|
echo "test-default-resolver: ok"
|