Make `beellama/dflash` the single-GPU default for Qwen3.6-27B, served via an unofficial multi-arch image, and demote ik-llama to the balanced alt. Decision basis (single-3090, 370 W; bench.sh n=5 + quality-test --full): - Code TPS ~100 vs ik 69 — fastest single-card 3090 code path. - 8-pack 107/150 (71%) think-off vs ik 99/150 (66%); 113/150 think-on (same-session). DFlash is output-lossless, so quality == the Q5_K_S target. - Context ceiling ladder (measured): 130K comfortable, 160K usable (115K-tok prefill OK, 0.8 GB headroom), 200K boots but OOMs on prefill. Ships 102K default; raise to 160K via CTX_SIZE. - ik still wins narrative TPS (63), VRAM, and shipped vision — kept as the balanced alt (`--variant ik-llama/iq4ks-mtp`). Registry: `beellama/dflash` experimental -> caveats + a new DEFAULTS[(qwen3.6-27b, beellama, single)] row. beellama is already #1 in ENGINE_PREFERENCE[single], so the resolver now returns it for Qwen single. Image: unofficial multi-arch ghcr build `beellama-cpp:multiarch-b9459-07ac3ce` (sm_86/89/120 = RTX 3090/4090/5090), built from Anbeeld/beellama.cpp .devops/cuda.Dockerfile with CUDA_DOCKER_ARCH="86;89;120" + GGML_CUDA_FA_ALL_QUANTS=ON. Engine reports `ARCHS = 860,890,1200`; cuobjdump confirms all 3 SASS. CAVEAT: sm_89/sm_120 are compiled but UNVALIDATED — only sm_86/3090 is verified on-rig. Also bumps the gemma beellama compose to the same multi-arch image (stays experimental — not promoted). Docs: README single-card framing, INFERENCE_ENGINES.md, UPSTREAM.md row, BENCHMARKS.md Qwen beellama row. Resolver guard tests flipped to the new default; full suite 38/38. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
65 lines
3.3 KiB
Bash
Executable File
65 lines
3.3 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# PR-B — <engine>/default resolver uses DEFAULTS + detected topology.
|
|
set -euo pipefail
|
|
|
|
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
|
|
cd "$ROOT_DIR"
|
|
|
|
assert_contains() {
|
|
local haystack="$1" needle="$2"
|
|
if [[ "$haystack" != *"$needle"* ]]; then
|
|
echo "ASSERTION FAILED: expected output to contain: $needle" >&2
|
|
echo "--- output ---" >&2
|
|
echo "$haystack" >&2
|
|
exit 1
|
|
fi
|
|
}
|
|
|
|
fake_one='0:RTX_3090:24576:8.6'
|
|
fake_two='0:RTX_3090:24576:8.6,1:RTX_3090:24576:8.6'
|
|
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection vllm/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/default"
|
|
assert_contains "$out" "vllm/default"
|
|
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_two" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection vllm/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/dual"
|
|
assert_contains "$out" "vllm/dual"
|
|
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection vllm/dual/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/dual"
|
|
|
|
if out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection llamacpp/dual/default 2>&1)"; then
|
|
echo "ASSERTION FAILED: bad topology default unexpectedly resolved" >&2
|
|
echo "$out" >&2
|
|
exit 1
|
|
fi
|
|
assert_contains "$out" "cannot resolve default variant 'llamacpp/dual/default'"
|
|
assert_contains "$out" "Available defaults"
|
|
|
|
out="$(NVIDIA_VISIBLE_DEVICES=0,1 FORCE=1 PREFLIGHT_NO_COMPOSE_DEPS=1 COMPOSE_BIN=: READY_TIMEOUT=1 bash scripts/switch.sh --no-wait vllm/default 2>&1 || true)"
|
|
assert_contains "$out" "bringing up: vllm/dual"
|
|
|
|
out="$(NVIDIA_VISIBLE_DEVICES=0 FORCE=1 PREFLIGHT_NO_COMPOSE_DEPS=1 COMPOSE_BIN=: READY_TIMEOUT=1 bash scripts/switch.sh --no-wait vllm/dual/default 2>&1 || true)"
|
|
assert_contains "$out" "bringing up: vllm/dual"
|
|
|
|
# PR-B: `<model>/default` token dispatch through launch.sh (engine-vs-model).
|
|
# Single rig: qwen3.6-27b/default → curated beellama (ENGINE_PREFERENCE #1, caveats);
|
|
# dual → vllm/dual.
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection --variant qwen3.6-27b/default 2>&1)"
|
|
assert_contains "$out" "selected variant: beellama/dflash"
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_two" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection --variant qwen3.6-27b/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/dual"
|
|
# gemma-4-31b/default dual → vllm/gemma-mtp (model token overrides PRIMARY_MODEL).
|
|
out="$(CLUB3090_FAKE_GPUS="$fake_two" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection --variant gemma-4-31b/default 2>&1)"
|
|
assert_contains "$out" "selected variant: vllm/gemma-mtp"
|
|
# Unknown X/default → clear error (neither engine nor model).
|
|
if out="$(CLUB3090_FAKE_GPUS="$fake_one" SWITCH=/bin/echo bash scripts/launch.sh --no-preflight --no-verify --no-projection --variant bogus/default 2>&1)"; then
|
|
echo "ASSERTION FAILED: bogus/default unexpectedly resolved" >&2
|
|
echo "$out" >&2
|
|
exit 1
|
|
fi
|
|
assert_contains "$out" "neither a known engine nor a known model"
|
|
|
|
echo "test-default-resolver: ok"
|