Release / release (push) Failing after 49s
Catches the failure mode reported by alexpolo1 on Discord 2026-05-11:
switch.sh reports "no club-3090 container running" but the GPU is still
pinned at ~22 GiB from a non-managed process, and the new container
OOMs at boot with a cryptic vLLM ValueError.
Two changes:
1. Widen RUNNING_PATTERN from a hard-coded variant list to `^(vllm-|llama-cpp-)`
so down_running() also catches locally-built and one-off `docker run`
instances under the same image families.
2. Add gpu_preflight() between down_running() and up_variant():
- Queries nvidia-smi for free memory per GPU.
- If any card has <80% free (insufficient for the typical 0.92
gpu-memory-utilization), abort with a diagnostic listing the
holding PIDs from nvidia-smi --query-compute-apps and suggesting
specific cleanup commands.
- FORCE=1 env bypasses the check for users who know what they're doing.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>