Commit Graph
11 Commits
Author SHA1 Message Date
noonghunnaandtekgnosis-net 6cafaf80b6 Operational robustness (#281): orphan-safe switch.sh · reboot-surviving vLLM · multi-GPU power sweep (#285)
Re-bases tekgnosis-net's #281/#282 onto master: (1) switch.sh registry-derived VARIANT_CONTAINER closed-world teardown (+--remove-orphans) — fixes beellama/ik-llama/sglang VRAM leak; (2) 29 vLLM composes restart: ${CLUB3090_RESTART:-unless-stopped} (reboot survival, opt-out knob); (3) power-cap-sweep.sh multi-GPU (board-power sum + cap restore). 3 new tests; suite 41/41. Closes #281, supersedes #282.

Co-Authored-By: tekgnosis-net <[email protected]>
2026-05-31 22:13:52 +05:00
noonghunnaandnoonghunna 6291ce5081 feat(eval): expose request-level thinking toggles (#196)
Closes #195.

Add ENABLE_THINKING support to bench.sh and pass --enable-thinking / --thinking-max-tokens through quality-test.sh. Propagate the same env through rebench-full.sh, warn on likely reasoning-on server/request-off mismatches, and document the workflow for reasoning-model evals.

Co-authored-by: noonghunna <[email protected]>
2026-05-23 08:24:58 +05:00
noonghunnaandnoonghunna 3b5b213db7 feat(compose): expose sampling defaults via env (#194)
Closes #193.

Wire TEMP/TEMPERATURE, TOP_P, TOP_K, MIN_P, and REPEAT_PENALTY into tracked llama.cpp, ik_llama, and vLLM compose variants. Keep model-specific compose fallbacks while allowing shell/.env overrides; vLLM uses override-generation-config, and llama.cpp/ik_llama use native sampling flags.

Co-authored-by: noonghunna <[email protected]>
2026-05-23 06:03:33 +05:00
noonghunna 28bd0e8970 docs: fix stale NVLINK_MODE comment, INTERNALS.md cliff status, and dead companion repo link 2026-05-18 19:50:47 +00:00
John Karabudak e00626a50e feat: unify dual-card composes with NVLink auto-detection
As discussed in #98, collapse 8 Qwen dual-card compose files into 4.
NVLink presence is now detected at startup instead of requiring
a separate compose per interconnect. The 4 nvlink-*.yml files become
deprecated stubs that extend the unified compose with NVLINK_MODE=force_on.

Add entrypoint conditionals and env-var substitution so the same compose
selects the correct NCCL settings and --disable-custom-all-reduce flag based
on detected topology. Extend the same pattern to all 7 Gemma dual composes.
2026-05-14 02:36:15 -02:30
noonghunna 26985527f7 Add hardware-aware compose preflight
Release / release (push) Failing after 50s
2026-05-13 18:03:37 +00:00
Erik LaBiancaandClaude Sonnet 4.6 af9fb0cd29 compose: parametrize VLLM_ENFORCE_EAGER, KV_CACHE_DTYPE, P40/P82/PN54 across all variants (#110)
All defaults unchanged — existing users see identical behavior out of the box.
New opt-ins let RTX 5090 / WSL2 / high-L2 rigs tune without forking files.

Changes across 18 compose files + .env.example:

VLLM_ENFORCE_EAGER — hook added to 7 files that lacked it:
  carnice-bf16mtp, dual-dflash, dual-dflash-noviz, dual-nvlink-dflash,
  dual-nvlink-dflash-noviz, dual4-dflash, minimal
  (bounded-thinking, dual, dual-turbo, long-*, tools-text, docker-compose.yml
   already had the hook)

KV_CACHE_DTYPE — parameterised in all 13 variants that hardcoded it:
  turboquant_3bit_nc default: bounded-thinking, docker-compose.yml,
    long-text, long-text-no-mtp, long-vision, dual-turbo, dual-nvlink-turbo
  fp8_e5m2 default: carnice-bf16mtp, dual, dual-nvlink, dual4, minimal, tools-text

GENESIS_ENABLE_P40 + GENESIS_ENABLE_PN54 — opt-in stanzas added to all 8
  Genesis-using variants: bounded-thinking, docker-compose.yml, dual-turbo,
  dual-nvlink-turbo, long-text, long-text-no-mtp, long-vision, tools-text

GENESIS_ENABLE_P82 — promoted from hardcoded 0 → ${GENESIS_ENABLE_P82:-0}
  in 6 spec-decode variants: bounded-thinking, dual-turbo, dual-nvlink-turbo,
  long-text, long-text-no-mtp, long-vision

.env.example additions:
  - Docs for VLLM_ENFORCE_EAGER, KV_CACHE_DTYPE, P40, P82, PN54
  - Validated RTX 5090 Laptop + WSL2 profile block (issue #102):
    PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False,max_split_size_mb:512
    GPU_MEMORY_UTILIZATION=0.94, VLLM_ENFORCE_EAGER=1, GENESIS_ENABLE_P40=1,
    GENESIS_ENABLE_P82=1, SOAK_TIMEOUT_S=3600

Co-authored-by: Claude Sonnet 4.6 <[email protected]>
2026-05-10 05:51:49 +05:00
Erik LaBiancaandClaude Sonnet 4.6 73c31848ea fix: BIND_HOST opt-in + localhost script fixes (#109)
Three related fixes for running benchmarks without IDE agent interference:

1. All 18 vLLM compose files: port binding is now
   ${BIND_HOST:-0.0.0.0}:${PORT:-<n>}:8000
   Setting BIND_HOST=127.0.0.1 in .env restricts the API to localhost,
   preventing IDE agents (Cline, Cursor) from competing for the
   max-num-seqs=1 slot and causing verify-stress HTTP 000 failures.

2. scripts/preflight.sh: port auto-detection regex now matches
   127.0.0.1:<port>->8000/tcp in addition to 0.0.0.0: and [::]:
   Previously all verify-*/bench scripts silently produced no output
   when BIND_HOST=127.0.0.1 was set.

3. scripts/report.sh: SOAK_TIMEOUT_S is now forwarded to soak-test.sh
   from the shell environment. Previously the variable was read from
   compose .env (docker-compose only) and silently ignored by the
   script, always using the 1800s default regardless of what was set.

Docs: .env.example gains BIND_HOST and SOAK_TIMEOUT_S entries.

Co-authored-by: Claude Sonnet 4.6 <[email protected]>
2026-05-10 05:50:08 +05:00
Erik LaBiancaandClaude Opus 4.7 276ab89291 composes: PYTORCH_CUDA_ALLOC_CONF env-override knob + WSL2 boot-crash docs (#84)
A single-card RTX 3090 Ti rig on WSL2 (driver 596.36) hits
RuntimeError: CUDA driver error: device not ready from gptq_marlin_repack
immediately after weight load on the v7.72.2-uplift nightly pin.
Bisect ruled out Genesis, spec-decode, TQ3 KV, async-residual error,
and TDR (registry already extended + Windows rebooted).
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False resolves it. Same
env-var workaround as JusefPol's NVLink boot-crash (PR #31), already
hardcoded in the dual-nvlink*.yml composes.

Replaces the hardcoded PYTORCH_CUDA_ALLOC_CONF line in 14 single-card
and PCIe dual-card composes with a ${PYTORCH_CUDA_ALLOC_CONF:-...}
override (defaults preserved). Pattern matches existing MAX_MODEL_LEN /
GPU_MEMORY_UTILIZATION overrides from #79. The two dual-nvlink*.yml
composes are unchanged — their existing JusefPol-driven default already
has expandable_segments off.

Documentation:
- docs/HARDWARE.md: new "disable PyTorch expandable_segments" subsection
  alongside the TDR fix, with stack trace, what was ruled out, override
  recipe, and a single uncontrolled observation about weight-load time
  (32 sec → 13 sec).
- docs/FAQ.md: WSL2 question now cross-links both the TDR and
  expandable_segments fix subsections.
- .env.example: documents the override under "vLLM tuning knobs".
- CHANGELOG.md (top-level + per-model): dated 2026-05-06 entries.

The exact failing call hasn't been isolated. The cuMemMap virtual-memory
API used by expandable_segments:True is the suspected culprit since
both known occurrences respond to the same workaround, but no specific
call has been proven to return cudaErrorNotReady.

Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-07 05:26:54 +05:00
noonghunnaandClaude Opus 4.7 ebacba1efd fix: address open issues #1, #4, #7
Closes / addresses 3 reported issues + adds requested feature:

#7 vid (PORT not honored, MODEL_DIR vs MODELS_DIR confusion):
  - All 8 vLLM compose files now use "${PORT:-XXXX}:8000" so .env PORT
    flows through. Defaults preserved per-variant (8020 single, 8010-8013
    dual). llama.cpp composes already had this pattern.
  - scripts/switch.sh: load .env early; per-variant default-port table;
    new resolve_ready_url() picks PORT > variant default for the readiness
    probe.
  - scripts/launch.sh: same default-port table; final endpoint URL printed
    to user reflects actual mapped port.
  - .env.example: ⚠ box callout that variable names are CASE-SENSITIVE
    (MODEL_DIR singular, NOT MODELS_DIR plural — silently ignored).
    New PORT section documenting per-variant defaults.

#4 timxx (tools-text.yml fails "Free memory ... less than desired"):
  - docs/FAQ.md: new entry "Container fails to start: Free memory..."
    explaining the vLLM startup check, the two workarounds (free VRAM /
    lower mem-util), and which configs hit it most often (0.97+ mem-util).
  - Compose defaults unchanged (0.97 stays the right design target on
    headless rigs); the FAQ documents the workaround for users with X11.

#1 fabriciomalta (per-config VRAM column):
  - docs/SINGLE_CARD.md: TL;DR table now has VRAM column with mem-util.
  - docs/DUAL_CARD.md: TL;DR table same + footnote explaining per-card
    semantics and which dual configs would/wouldn't fit on 2× 20 GB cards
    (relevant to fabriciomalta's 2× 3080-20GB use case).

#2 tenitram (empty responses) — fixed in master via aab8ff4
(P68/P69 disabled). Closed with reply pointing at the fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-30 16:19:53 +00:00
noonghunnaandClaude Opus 4.7 ec704e4e2e Pin Genesis to exact tested commit + add .env.example + issue templates
- setup.sh: GENESIS_PIN now defaults to commit bf667c7 (Genesis HEAD as of
  2026-04-27, semver "v7.54"). This is the exact tree our published TPS
  numbers were measured against; tagged v7.51-stable was one minor older
  but came up first because the SHA isn't durable. Switch to commit pin
  removes the doc-vs-runtime mismatch. Clone strategy adjusted since
  --branch + --depth 1 doesn't accept SHAs.
- .env.example: documents MODEL_DIR / HF_TOKEN / CUDA_VISIBLE_DEVICES /
  MEM_UTIL / MAX_MODEL_LEN / GENESIS_PIN / SKIP_GENESIS / URL / WARMUPS /
  RUNS with the same defaults the composes ship. Pure opt-in.
- .github/ISSUE_TEMPLATE/: bug-report.yml requires docker logs --tail 100,
  verify-full.sh output, nvidia-smi, GPU config, compose variant, repo
  commit. numbers-from-your-rig.yml structures cross-rig TPS contributions
  with rig spec, bench output, VRAM, max ctx, and notes. config.yml
  routes Q&A to Discussions.
- .gitignore: drop trailing slash on genesis pattern so it also ignores
  local symlinks that some of us point at out-of-tree clones.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-28 13:03:04 +00:00