Files
club-3090/scripts
noonghunnaandClaude Opus 4.7 4866913a10
Release / release (push) Failing after 49s
fix(switch): GPU memory pre-flight + widen RUNNING_PATTERN
Catches the failure mode reported by alexpolo1 on Discord 2026-05-11:
switch.sh reports "no club-3090 container running" but the GPU is still
pinned at ~22 GiB from a non-managed process, and the new container
OOMs at boot with a cryptic vLLM ValueError.

Two changes:

1. Widen RUNNING_PATTERN from a hard-coded variant list to `^(vllm-|llama-cpp-)`
   so down_running() also catches locally-built and one-off `docker run`
   instances under the same image families.

2. Add gpu_preflight() between down_running() and up_variant():
   - Queries nvidia-smi for free memory per GPU.
   - If any card has <80% free (insufficient for the typical 0.92
     gpu-memory-utilization), abort with a diagnostic listing the
     holding PIDs from nvidia-smi --query-compute-apps and suggesting
     specific cleanup commands.
   - FORCE=1 env bypasses the check for users who know what they're doing.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 11:20:40 +00:00
..