diff --git a/docs/CLIFFS.md b/docs/CLIFFS.md index a690d1aa..ad471b3e 100644 --- a/docs/CLIFFS.md +++ b/docs/CLIFFS.md @@ -417,6 +417,21 @@ On 24 GB / 3090 the per-card budget absorbs TQ3's activation peak and the smalle The general principle: **the variant matrix is per-card-budget × KV-format-tradeoff aware**. Compose defaults are tuned for 24 GB / 3090; users on different VRAM classes may need to override `--kv-cache-dtype` to relocate the activation/pool balance for their hardware. +#### The naming trap — fp8 is *larger* than TQ3 per token + +The KV-format names are misleading. **fp8_e5m2 stores 8 bits per cached element; turboquant_3bit_nc packs 3 bits.** At long single-prompt contexts on memory-tight rigs, this flips the conventional intuition: + +| KV format | Bytes / cached token | Verdict at 180K on 24 GB single-card | +|---|---:|---| +| `turboquant_3bit_nc` (TQ3) | 0.375 | ✅ fits at 180K with mem-util 0.93 | +| `fp8_e5m2` | 1.0 | ❌ OOMs — needs 6.64 GiB KV pool, only 4.36 GiB available | + +**Validated by [@easel #102](https://github.com/noonghunna/club-3090/issues/102#issuecomment-4412264989) on RTX 5090 Laptop**: switching from TQ3 → fp8_e5m2 at 180K context produced `ValueError: 6.64 GiB KV cache is needed, larger than available (4.36 GiB)`. Reverted to TQ3 and the boot succeeded. + +**Pin**: on 24 GB single-card, **TQ3 is the long-context KV; fp8 is for short-context throughput**. The TQ3 activation-peak trade discussed above is real but ~1 GB; the fp8 KV-pool inflation at long-ctx is several GB. At 180K the activation-peak trade is dominated by the pool-size trade. + +**On dual-card rigs** the analysis is the same per-card; TQ3 is still the long-context-on-tight-budget choice. fp8 starts to make sense again when (a) context is short enough that pool size doesn't dominate, or (b) you have generous per-card headroom (32 GB+ cards, or shorter max_model_len giving you VRAM to spare). + ### Why llama.cpp doesn't have Cliff 2 llama.cpp's Qwen3-Next implementation processes DeltaNet/GDN layers with **online state updates** (incremental) rather than materializing the full intermediate. State is updated per-token or per-tile, never as a single multi-GB tensor. Different algorithm, different memory profile. diff --git a/docs/HARDWARE.md b/docs/HARDWARE.md index 00c1619a..fc1687d5 100644 --- a/docs/HARDWARE.md +++ b/docs/HARDWARE.md @@ -273,6 +273,29 @@ sudo nvidia-smi -rgc -i 0 # release graphics-clock lock - This is an **air-cooled 5090 finding** — water-cooled rigs may have different optimal clock pairs (lower thermals → higher sustained boost-clock-vs-power tradeoff) - The freq-cap methodology hasn't been wrapped into `power-cap-sweep.sh` yet — apnar's data is hand-rolled. If you want to run a similar sweep on your 5090, copy his approach until we ship a `freq-cap-sweep.sh` companion +### Laptop GPUs — EC-managed power (no software power-cap) + +On laptop-class Ampere/Ada/Blackwell GPUs (RTX 30/40/50-series Laptop variants), `nvidia-smi -pl ` returns `[N/A]` and software power-cap tools cannot enforce a limit. The power envelope is owned by the **embedded controller (EC)** via the platform firmware (a vendor-specific implementation of NVIDIA's Dynamic Boost / OEM platform-power policy), not exposed to the OS: + +```text +Power: limit=[N/A] (default=95W, max=175W) | current_draw=94W @ load +``` + +The card reports a max TDP in PCI config but the EC enforces the actual operating point based on platform thermals, AC-vs-battery, cooling fans, and OEM-specific tuning. `nvidia-smi` cannot override the EC. + +**Confirmed on**: RTX 5090 Laptop (driver 596.36, EC profile 95W) — [@easel #102 follow-up 2026-05-09](https://github.com/noonghunna/club-3090/issues/102#issuecomment-4412264989). + +**Implications**: +- `scripts/power-cap-sweep.sh` detects the limitation and exits gracefully — it cannot characterize laptop GPUs +- The matrix entry for laptop rigs in HARDWARE.md should read: *software power-cap: N/A (EC-managed)* +- The clock-lock approach (`-lgc` / `-lmc` from the Blackwell desktop section above) is the **only available characterization path** on laptop GPUs — clock-locking does not have the same EC-dependency as the power-cap actuator + +**Practical guidance for laptop owners**: +- Don't try to run our power-cap sweep tools — they'll fail to actuate +- Tune via clock-locking instead, but expect the EC to potentially override your clock locks under thermal pressure +- For sustained throughput, focus on cooling (laptop cooling pad, undervolt via vendor tools, AC power) rather than software caps +- The pre-set EC profiles (Performance / Quiet / Eco modes in vendor tools like NVIDIA App, Lenovo Vantage, ASUS Armoury Crate) are the user-accessible knobs + ### Interpreting "draw plateaued below cap" sweeps If your sweep ends with the high-cap rows showing **actual draw < cap by 5-15%** (e.g. 547W actual at 600W cap on a 5090 with `decode-concurrent N=4`), it usually means one of: diff --git a/scripts/verify-stress.sh b/scripts/verify-stress.sh index 8b923899..ec35b5b8 100755 --- a/scripts/verify-stress.sh +++ b/scripts/verify-stress.sh @@ -49,6 +49,13 @@ # PREFILL_TARGET_CHARS Tool-response prefill payload size in chars # (default: 100000 ≈ 25K tokens; set higher to # push closer to the cliff under investigation). +# STRESS_LONGCTX_TIMEOUT_S Curl timeout for long-context needle checks +# (default: 300, auto-bumped to 600 if container has +# VLLM_ENFORCE_EAGER=1 — eager prefill at 60K+ can +# take 200-290s, see easel #102 follow-up). +# STRESS_TOOL_PREFILL_TIMEOUT_S Curl timeout for tool-prefill OOM check +# (default: 240, auto-bumped to 480 if container has +# VLLM_ENFORCE_EAGER=1). set -euo pipefail @@ -64,6 +71,26 @@ URL="${URL:-http://localhost:8020}" MODEL="${MODEL:-qwen3.6-27b-autoround}" CONTAINER="${CONTAINER:-vllm-qwen36-27b}" +# Detect VLLM_ENFORCE_EAGER=1 in the running container's env. Eager-mode +# prefill at 60K-140K can take 200-290s (vs <60s with CUDA graphs); the +# default curl timeouts below would false-positive as HTTP 000 on rigs that +# need eager mode to fit (typical for WSL2 / laptop GPUs at long ctx, see +# easel #102). When detected, scale the long-ctx + tool-prefill timeouts up. +EAGER_MODE_DETECTED=0 +if command -v docker >/dev/null 2>&1; then + if docker inspect "${CONTAINER}" --format '{{range .Config.Env}}{{println .}}{{end}}' 2>/dev/null \ + | grep -qE '^VLLM_ENFORCE_EAGER=1$'; then + EAGER_MODE_DETECTED=1 + fi +fi +if [[ "${EAGER_MODE_DETECTED}" == "1" ]]; then + STRESS_LONGCTX_TIMEOUT_S="${STRESS_LONGCTX_TIMEOUT_S:-600}" + STRESS_TOOL_PREFILL_TIMEOUT_S="${STRESS_TOOL_PREFILL_TIMEOUT_S:-480}" +else + STRESS_LONGCTX_TIMEOUT_S="${STRESS_LONGCTX_TIMEOUT_S:-300}" + STRESS_TOOL_PREFILL_TIMEOUT_S="${STRESS_TOOL_PREFILL_TIMEOUT_S:-240}" +fi + pass() { printf " \033[32m✓\033[0m %s\n" "$1"; } fail() { printf " \033[31m✗\033[0m %s\n" "$1"; printf " \033[33m→\033[0m %s\n" "$2"; return 1; } skip() { printf " \033[33m⊘\033[0m %s (skipped)\n" "$1"; } @@ -201,7 +228,7 @@ EOF secret="$(cat "$secret_file")" local resp content_raw prompt_tok http_code resp_file resp_file="$(mktemp --suffix=.json)" - http_code="$(curl -s -m 300 -o "${resp_file}" -w '%{http_code}' \ + http_code="$(curl -s -m "${STRESS_LONGCTX_TIMEOUT_S}" -o "${resp_file}" -w '%{http_code}' \ "${URL}/v1/chat/completions" \ -H "Content-Type: application/json" \ --data-binary "@${req_file}")" || http_code="000" @@ -327,7 +354,7 @@ with open(os.environ['REQ_FILE'], 'w') as f: EOF local http_code - http_code="$(curl -s -m 240 -o "${resp_file}" -w '%{http_code}' \ + http_code="$(curl -s -m "${STRESS_TOOL_PREFILL_TIMEOUT_S}" -o "${resp_file}" -w '%{http_code}' \ "${URL}/v1/chat/completions" \ -H "Content-Type: application/json" \ --data-binary "@${req_file}")" || http_code="000"