Files
club-3090/scripts/tests/test-gpu-mode-list.sh
47ac38385a Director placement lever: CPU / GPU0 / GPU1 (c3 Settings) (#473)
* Director placement lever: env-driven CPU / GPU0 / GPU1 (backend)

The studio director's GPU/CPU placement is now a single lever,
STUDIO_DIRECTOR_DEVICE (gpu0 | gpu1 | cpu, default gpu0), read from the rig
.env by gpu-mode's start_studio_director and translated into the compose
env (-ngl + CUDA_VISIBLE_DEVICES + device_ids):

- gpu0 (default): -ngl 99, GPU0 — fast craft (~50-100 tok/s), ~4.6 GB,
  coexists with the image lanes. Unchanged from before.
- gpu1: -ngl 99, GPU1 — only when GPU1 has room (NOT during a video render;
  GPU1 is the DisTorch DiT donor).
- cpu: -ngl 0, CUDA_VISIBLE_DEVICES="" — frees ~4.6 GB off GPU0 (lifts the
  single-card Wan window 121→161 frames) at ~single-digit tok/s craft.

Compose now reads ${DIRECTOR_NGL:-99} + ${STUDIO_DIRECTOR_CUDA-0} (no-colon
so an explicit empty value = CPU survives). Default (no override) preserves
current GPU0 behaviour exactly.

Live-validated: CPU mode starts with GPU0 full (gemma12b), adds 0 MiB VRAM
to GPU0, serves on :8090, generates (~5 tok/s CPU). The c3 Settings field
that writes STUDIO_DIRECTOR_DEVICE follows in the next commit.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* c3 Settings: Director placement (CPU / GPU0 / GPU1) + CPU thread cap

Exposes the director-placement lever in the c3 Settings modal so users pick
GPU0 (fast, default) / GPU1 / CPU without hand-editing the .env.

- services.py: director_device() reads STUDIO_DIRECTOR_DEVICE from the repo
  .env (default gpu0, validates the value); set_repo_env_var() upserts a key
  in place (preserves other lines, no duplicates, creates the file if absent).
- app.py: SettingsScreen gains a "Director placement" Select; apply_settings
  persists the choice to the repo .env (the SHARED config gpu-mode reads —
  distinct from c3-settings.json for MODEL_DIR/HF_TOKEN). Applies on the next
  ai-studio start.
- compose: CPU thread cap — -t ${DIRECTOR_THREADS:-8} bounds CPU use so the
  director doesn't starve OWUI's embedder/reranker (also CPU). The ~2.6 GB
  GGUF loads into system RAM (mmap'd; resident in page cache, not run from SSD).
- tests: +6 data-layer (TestDirectorPlacement) + 1 headless apply-settings
  round-trip (persists STUDIO_DIRECTOR_DEVICE, idempotent re-apply). Full
  suite green (728), settings/director subset 13/13.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* tests: fix stale scene names in gpu-mode-list EXPECT

The scene catalog renamed its dispatch keywords to qwen27b / gemma-31b,
but the test's EXPECT spot-check map still referenced the old 27b / gemma
short names — so the JSON-shape assertion had been red on master. Point
EXPECT at the canonical names the catalog now emits.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* studio: run director in chat scene as the catalog-support layer

Bring the uncensored director up in the `chat` scene (honoring the
STUDIO_DIRECTOR_DEVICE placement knob), framing chat as the supporting-
infra home — OWUI + LiteLLM + Qdrant + SearXNG + director — for ad-hoc
Catalog models launched via `switch.sh --owui`.

A CPU-placed director uses no GPU, so it's the always-on path: it survives
scene switches and stays live in OWUI. New _director_evict_if_gpu helper
frees only a GPU-resident director when a dual-card LLM scene claims the
cards; mode_off stops it outright. Also brings mode_gemma_int8 in line with
its dual-card siblings (it was missing the studio teardown entirely).

Docs: requirements.md gains a "Chat scene — the Catalog-support layer"
section + reframes director placement around the unified knob / c3 Setting;
video.md note synced.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* studio: disable director thinking on CPU only (latency)

The uncensored director is an "Aggressive" reasoning fine-tune that emits
a full <think> trace before every answer. On GPU that's cheap and the
trace lands in reasoning_content (content stays clean), so leave it on.
On CPU (~14 tok/s) the trace dominates latency, so gpu-mode now passes
`--jinja --reasoning off` for the cpu placement only — forcing the
template's enable_thinking=false (this fine-tune ignores /no_think and
--reasoning-budget 0, but honors --reasoning off).

Wired via a new DIRECTOR_THINK_ARGS compose param (empty on GPU). Live:
CPU director now answers in one pass, no reasoning trace, craft quality
intact (full cinematic spec, finish=stop).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

* c3: Containers-pane director start honors placement + enumerate sidecars

Two consistency fixes for the Containers tab:

1. Starting studio-director from Containers now honors STUDIO_DIRECTOR_DEVICE
   instead of falling back to the GPU0/thinking-on compose default. New
   director_compose_env() mirrors gpu-mode's start_studio_director translation
   (NGL/CUDA/GPU/THINK_ARGS), injected as an `env K=V …` prefix on the compose
   up cmd (process env wins over --env-file). cpu → -ngl 0 + --reasoning off.

2. The nested studio sidecars (director/gallery/orchestrator/image-shim/
   step-voice/tts) now enumerate when STOPPED, so they're startable rows — not
   only visible while running. New STUDIO_SIDECARS map is the single SoT for
   resolving the container-name → services/studio/<sub>/ project (fixing the
   director↔enhancer name mismatch that previously returned None → docker
   restart, which fails on a fresh install).

+8 tests (director_compose_env cpu/gpu, director resolves to enhancer with the
env prefix, sidecar enumeration). Live: c3 service_start plan starts the director
CPU + no-think (argv -ngl 0 --reasoning off, GPU0 free, clean generation).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-25 09:01:46 +05:00

116 lines
4.6 KiB
Bash
Executable File

#!/usr/bin/env bash
# Serve-cockpit — gpu-mode --list-modes scene-catalog emitter.
#
# `gpu-mode --list-modes` prints a human scene catalog; `--list-modes --json`
# emits the machine-readable [{name,group,description,services,ports,gpus}]
# array the cockpit consumes. This test exercises the new emit WITHOUT touching
# GPUs/docker (the catalog is static data derived from the dispatch case) and
# asserts its SHAPE: valid JSON, the six required keys per row, the contract's
# three groups, services/ports as arrays, and that every dispatch keyword that
# should appear is present in the correct group. It also guards the strictly-
# additive contract: an unknown mode still falls through to usage().
set -euo pipefail
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
cd "$ROOT_DIR"
GPU_MODE="$ROOT_DIR/scripts/gpu-mode.sh"
fail=0
note() { echo "FAIL: $1" >&2; fail=1; }
assert_contains() {
local hay="$1" needle="$2" msg="$3"
[[ "$hay" == *"$needle"* ]] || note "${msg}: output lacks '${needle}'"
}
assert_not_contains() {
local hay="$1" needle="$2" msg="$3"
[[ "$hay" != *"$needle"* ]] || note "${msg}: output unexpectedly contains '${needle}'"
}
# --- 1. --list-modes --json is valid JSON with the contract shape -----------
JSON_OUT="$(bash "$GPU_MODE" --list-modes --json 2>&1)"
JSON_OUT="$JSON_OUT" python3 - <<'PY' || note "JSON shape assertion failed"
import json, os, sys
raw = os.environ["JSON_OUT"]
try:
data = json.loads(raw)
except json.JSONDecodeError as e:
print(f" not valid JSON: {e}", file=sys.stderr)
sys.exit(1)
ok = True
def bad(msg):
global ok
print(f" {msg}", file=sys.stderr)
ok = False
if not isinstance(data, list) or not data:
bad("top-level is not a non-empty array")
sys.exit(0 if ok else 1)
REQUIRED = {"name", "group", "description", "services", "ports", "gpus"}
GROUPS = {"models", "studio", "ops"}
# Every dispatch keyword the catalog must expose, with its contract group.
# (serving → models; chat → ops — browser-chat infra, no GPU model. bigmodel +
# diffusiongemma scenes were removed — bigmodel ≈ off, dgemma → catalog slug. The
# standalone comfyui scene was removed — comfyui runs via ai-studio.)
EXPECT = {
"qwen27b": "models", "gemma-31b": "models", "deckard": "models",
"ai-studio": "studio",
"chat": "ops", "off": "ops", "power-cap": "ops", "prune": "ops", "prune-all": "ops",
}
seen = {}
for i, row in enumerate(data):
if not isinstance(row, dict):
bad(f"row {i} is not an object")
continue
missing = REQUIRED - set(row)
if missing:
bad(f"row {row.get('name', i)} missing keys: {sorted(missing)}")
if row.get("group") not in GROUPS:
bad(f"row {row.get('name', i)} has non-contract group {row.get('group')!r}")
if not isinstance(row.get("services"), list):
bad(f"row {row.get('name', i)} services is not an array")
if not isinstance(row.get("ports"), list):
bad(f"row {row.get('name', i)} ports is not an array")
if not isinstance(row.get("description"), str) or not row["description"]:
bad(f"row {row.get('name', i)} description empty/non-string")
seen[row.get("name")] = row.get("group")
for name, grp in EXPECT.items():
if name not in seen:
bad(f"expected mode {name!r} absent from catalog")
elif seen[name] != grp:
bad(f"mode {name!r} in group {seen[name]!r}, expected {grp!r}")
# Spot-check that a known mode carries real service/port data.
chat = next((r for r in data if r["name"] == "chat"), None)
if chat is None or "litellm" not in chat["services"] or "8080" not in chat["ports"]:
bad("chat row missing expected litellm service / 8080 port")
sys.exit(0 if ok else 1)
PY
# --- 2. plain --list-modes renders the grouped catalog ----------------------
PLAIN_OUT="$(bash "$GPU_MODE" --list-modes 2>&1)"
assert_contains "$PLAIN_OUT" "Scene Catalog" "plain render has catalog header"
assert_contains "$PLAIN_OUT" "[models]" "plain render groups models"
assert_contains "$PLAIN_OUT" "[studio]" "plain render groups studio"
assert_contains "$PLAIN_OUT" "[ops]" "plain render groups ops"
assert_contains "$PLAIN_OUT" "deckard" "plain render lists a models mode"
# --- 3. additive guard: unknown mode still prints usage ---------------------
USAGE_OUT="$(bash "$GPU_MODE" definitely-not-a-mode 2>&1)"
assert_contains "$USAGE_OUT" "GPU Mode Switcher" "unknown mode falls through to usage"
assert_not_contains "$USAGE_OUT" "Scene Catalog" "usage does not leak catalog output"
if [[ "$fail" -ne 0 ]]; then
echo "[gpu-mode-list] FAIL" >&2
exit 1
fi
echo "test-gpu-mode-list: ok"