Closes#195.
Add ENABLE_THINKING support to bench.sh and pass --enable-thinking / --thinking-max-tokens through quality-test.sh. Propagate the same env through rebench-full.sh, warn on likely reasoning-on server/request-off mismatches, and document the workflow for reasoning-model evals.
Co-authored-by: noonghunna <[email protected]>
As discussed in #98, collapse 8 Qwen dual-card compose files into 4.
NVLink presence is now detected at startup instead of requiring
a separate compose per interconnect. The 4 nvlink-*.yml files become
deprecated stubs that extend the unified compose with NVLINK_MODE=force_on.
Add entrypoint conditionals and env-var substitution so the same compose
selects the correct NCCL settings and --disable-custom-all-reduce flag based
on detected topology. Extend the same pattern to all 7 Gemma dual composes.
Three related fixes for running benchmarks without IDE agent interference:
1. All 18 vLLM compose files: port binding is now
${BIND_HOST:-0.0.0.0}:${PORT:-<n>}:8000
Setting BIND_HOST=127.0.0.1 in .env restricts the API to localhost,
preventing IDE agents (Cline, Cursor) from competing for the
max-num-seqs=1 slot and causing verify-stress HTTP 000 failures.
2. scripts/preflight.sh: port auto-detection regex now matches
127.0.0.1:<port>->8000/tcp in addition to 0.0.0.0: and [::]:
Previously all verify-*/bench scripts silently produced no output
when BIND_HOST=127.0.0.1 was set.
3. scripts/report.sh: SOAK_TIMEOUT_S is now forwarded to soak-test.sh
from the shell environment. Previously the variable was read from
compose .env (docker-compose only) and silently ignored by the
script, always using the 1800s default regardless of what was set.
Docs: .env.example gains BIND_HOST and SOAK_TIMEOUT_S entries.
Co-authored-by: Claude Sonnet 4.6 <[email protected]>
A single-card RTX 3090 Ti rig on WSL2 (driver 596.36) hits
RuntimeError: CUDA driver error: device not ready from gptq_marlin_repack
immediately after weight load on the v7.72.2-uplift nightly pin.
Bisect ruled out Genesis, spec-decode, TQ3 KV, async-residual error,
and TDR (registry already extended + Windows rebooted).
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False resolves it. Same
env-var workaround as JusefPol's NVLink boot-crash (PR #31), already
hardcoded in the dual-nvlink*.yml composes.
Replaces the hardcoded PYTORCH_CUDA_ALLOC_CONF line in 14 single-card
and PCIe dual-card composes with a ${PYTORCH_CUDA_ALLOC_CONF:-...}
override (defaults preserved). Pattern matches existing MAX_MODEL_LEN /
GPU_MEMORY_UTILIZATION overrides from #79. The two dual-nvlink*.yml
composes are unchanged — their existing JusefPol-driven default already
has expandable_segments off.
Documentation:
- docs/HARDWARE.md: new "disable PyTorch expandable_segments" subsection
alongside the TDR fix, with stack trace, what was ruled out, override
recipe, and a single uncontrolled observation about weight-load time
(32 sec → 13 sec).
- docs/FAQ.md: WSL2 question now cross-links both the TDR and
expandable_segments fix subsections.
- .env.example: documents the override under "vLLM tuning knobs".
- CHANGELOG.md (top-level + per-model): dated 2026-05-06 entries.
The exact failing call hasn't been isolated. The cuMemMap virtual-memory
API used by expandable_segments:True is the suspected culprit since
both known occurrences respond to the same workaround, but no specific
call has been proven to return cudaErrorNotReady.
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Closes / addresses 3 reported issues + adds requested feature:
#7 vid (PORT not honored, MODEL_DIR vs MODELS_DIR confusion):
- All 8 vLLM compose files now use "${PORT:-XXXX}:8000" so .env PORT
flows through. Defaults preserved per-variant (8020 single, 8010-8013
dual). llama.cpp composes already had this pattern.
- scripts/switch.sh: load .env early; per-variant default-port table;
new resolve_ready_url() picks PORT > variant default for the readiness
probe.
- scripts/launch.sh: same default-port table; final endpoint URL printed
to user reflects actual mapped port.
- .env.example: ⚠ box callout that variable names are CASE-SENSITIVE
(MODEL_DIR singular, NOT MODELS_DIR plural — silently ignored).
New PORT section documenting per-variant defaults.
#4 timxx (tools-text.yml fails "Free memory ... less than desired"):
- docs/FAQ.md: new entry "Container fails to start: Free memory..."
explaining the vLLM startup check, the two workarounds (free VRAM /
lower mem-util), and which configs hit it most often (0.97+ mem-util).
- Compose defaults unchanged (0.97 stays the right design target on
headless rigs); the FAQ documents the workaround for users with X11.
#1 fabriciomalta (per-config VRAM column):
- docs/SINGLE_CARD.md: TL;DR table now has VRAM column with mem-util.
- docs/DUAL_CARD.md: TL;DR table same + footnote explaining per-card
semantics and which dual configs would/wouldn't fit on 2× 20 GB cards
(relevant to fabriciomalta's 2× 3080-20GB use case).
#2 tenitram (empty responses) — fixed in master via aab8ff4
(P68/P69 disabled). Closed with reply pointing at the fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
- setup.sh: GENESIS_PIN now defaults to commit bf667c7 (Genesis HEAD as of
2026-04-27, semver "v7.54"). This is the exact tree our published TPS
numbers were measured against; tagged v7.51-stable was one minor older
but came up first because the SHA isn't durable. Switch to commit pin
removes the doc-vs-runtime mismatch. Clone strategy adjusted since
--branch + --depth 1 doesn't accept SHAs.
- .env.example: documents MODEL_DIR / HF_TOKEN / CUDA_VISIBLE_DEVICES /
MEM_UTIL / MAX_MODEL_LEN / GENESIS_PIN / SKIP_GENESIS / URL / WARMUPS /
RUNS with the same defaults the composes ship. Pure opt-in.
- .github/ISSUE_TEMPLATE/: bug-report.yml requires docker logs --tail 100,
verify-full.sh output, nvidia-smi, GPU config, compose variant, repo
commit. numbers-from-your-rig.yml structures cross-rig TPS contributions
with rig spec, bench output, VRAM, max ctx, and notes. config.yml
routes Q&A to Discussions.
- .gitignore: drop trailing slash on genesis pattern so it also ignores
local symlinks that some of us point at out-of-tree clones.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>