Vendored local overlay for vllm-project/vllm#35936 — fixes empty
tool_calls=[] response when tool_choice="required" is used with the
qwen3_coder parser. Under "required", vLLM forces structured JSON
output regardless of which parser is configured; the qwen3_coder
parser then scans the JSON for its XML <tool_call> sentinel, finds
none, and returns tools_called=False.
Overlay reshapes _parse_tool_calls_from_content() in
vllm/entrypoints/openai/engine/serving.py to try JSON-validate first
when tool_choice="required", and fall back to the configured parser
only when validation fails — covering both the supports_required_and_named
True and False parser paths.
End-to-end validated 2026-05-12:
- curl tool_choice="required" + get_weather tool: tool_calls populated
- curl tool_choice="auto" regression: still works
- MLS-Bench ml-ensemble-boosting with thinking.enabled=false: agent
reaches Step 1 (edit) with no "No action returned" stall
Mounted into the 23 Qwen 3.6-27B composes pinning the post-#41434
nightly-1acd67a79 image. The single compose on the older 01d4d1ad3
pin (dual/tq3-mtp-genesis.yml) is intentionally excluded — its source
tree diverges from the rebase base.
Added UPSTREAM.md row tracking the upstream PR + drop trigger.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
- RDNA 4: include RX 8000 alongside RX 9000 (early roadmap name; AMD
shipped under RX 9000 brand). Future-proofs for both references.
- Intel Xe2 FP8: add explicit maturity caveat — silicon supports it but
full production-stack acceleration in vLLM-SYCL / IPEX-LLM is still
maturing in mid-2026. OpenVINO has the most coverage today.
- AMD CDNA 4 MX*: clarify that NVFP4 is NVIDIA-proprietary (E2M1 + E4M3
per-16-element scale), so AMD targets the OCP-standard block-scaled
FP4/FP6 (MXFP4 / MXFP6) and AMD variants — for cross-vendor 4-bit
portability, MXFP4 is the safer target.
- AMD detection: fix the cross-vendor routing code block. HIP-via-PyTorch
appears as torch.cuda; reliable distinguishers are torch.version.hip
(not None) and torch.cuda.get_device_properties(0).gcnArchName plus
rocm-smi --showproductname. NVIDIA detection now explicitly excludes
torch.version.hip not None to avoid HIP false-positives.
- Per-vendor KV cache option tables added for Intel + AMD — symmetric
with NVIDIA's KV section above. Captures vLLM-ROCm FP8 KV on CDNA 3+
and RDNA 4; OpenVINO/IPEX-LLM coverage on Intel; llama.cpp portable
path everywhere.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Adds docs/DTYPE_MATRIX.md — a reference table mapping NVIDIA GPU
architectures (Pascal → Blackwell DC) to native vs emulated Tensor Core
support for each dtype + quant scheme.
Sections:
- At-a-glance compute-dtype matrix (9 archs × 10 dtypes, ✓/SW/✗ marks)
- Weight-quantization schemes (GPTQ / AWQ / AutoRound / NF4 / SmoothQuant
/ FP8 weights / MXFP8 / MXFP4 / NVFP4 / GGUF / HQQ-AQLM-SqueezeLLM)
with storage-vs-compute paths
- Weight-only vs weight+activation axis (W4A16 vs W8A8 vs W4A8)
- KV-cache dtype support (FP16 / FP8 / INT8 PTH / TQ3 / TQ4 / k8v4)
- Per-arch compose recommendations (which compose to ship for which
GPU class)
- Runtime detection (points at Genesis guards.py)
- Corner cases (Ada FP8 vs Hopper FP8, Blackwell consumer vs DC, NVFP4
vs MXFP4 block-size differences, MX* family overview, FP6)
- References (NVIDIA whitepapers, Marlin, Genesis, BENCHMARKS)
Cross-linked from:
- HARDWARE.md GPU-compat table
- GLOSSARY.md Quantization section
- FAQ.md as a new "What dtype/quant should I pick for my GPU?" Q under
Hardware
This sets the foundation for future per-arch compose optimization —
detecting compute capability at boot and picking the right KV dtype /
weight quant scheme automatically based on what the hardware actually
accelerates rather than what's nominally supported.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Adds the dual-card row for @ygafarov's MiniPC eGPU setup from #120:
3090 via USB4 dock (PCIe 3.0 x4 ≈ 3.94 GB/s) + 5070 Ti via OCuLink
(PCIe 4.0 x4 ≈ 7.88 GB/s) on AMD Ryzen AI MAX+ 395 / Strix Halo. fp8
KV, 200K ctx, MTP n=3 — 65.10 / 85.81 narr/code TPS at 1.00× concurrency.
Notable bits captured in the row:
- First heterogeneous Ampere + Blackwell consumer eGPU dual on the matrix
- 5070 Ti at 91% util but only 125 W (out of 290 W cap) — visibly waiting
on the 3090; vLLM compiles for sm_86 across both cards
- KV pool capped at 200K @ 1.00× by the 5070 Ti's 16 GiB VRAM, not the
3090's 24 GiB
- verify-stress 7/7 including 91K needle recall (Cliff 2 clean)
- Soak ⚠ borderline 360 MiB VRAM growth — same eGPU-bus accretion as his
own #113 single-card row at 240 MiB; not a leak, just x4-PCIe allocator
behavior under prefill
- Slower than his own single-3090 row from #113 (68.86 / 91.70 at 48K),
so the row also serves as the canonical "when is dual worse than single
on an eGPU rig" data point
Reply with diagnosis + experiments queued at #120.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Surfaces the canonical int8_per_token_head anti-scaling trade-off so the
next dual-3090 user who hits this finds the answer via search rather than
having to reason from the head-to-head matrix.
Frames it as:
- not a bug — per-(token, head) scale serializes at concurrency, fp8's
single global scale doesn't
- pick by workload: INT8 PTH for single-stream throughput, fp8 for
aggregate concurrency, Genesis-backed TQ3+MTP if you want both
- includes the diagnostic checklist for matching baseline numbers
(power cap, vLLM nightly, Genesis on/off, MTP n, prompt shape) so users
can self-troubleshoot a cross-rig gap before posting
Triggered by a Discord question on a dual-3090 + NVLink rig that observed
the canonical INT8 anti-scaling pattern (150 TPS single / flat at concurrency)
vs fp8 (70-100 TPS single / 400+ at concurrency).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Adds `docs/TQ3_MTP_GENESIS.md` — user-facing writeup of the 6th matched-config
rebench leg (qwen-tq3-mtp-genesis-2026-05-11). Frames the result as a sweet
spot on the quality-vs-concurrency frontier: 1.22M-token KV pool at 262K ctx
(2× the INT8 PTH baseline), 4.66× concurrency, within ~5pp quality of INT8 PTH
on the 150-scenario quality suite, within noise on aider-polyglot-30 (18/30
vs 19/30 baseline). 7/7 verify-stress, 0/100 silent-empty in soak.
Charts under `docs/assets/tq3-mtp-genesis/`:
- 01-kv-pool-by-config.png — KV pool capacity by leg
- 02-quality-vs-concurrency.png — trade-off frontier
- 03-verdict-matrix.png — per-phase verdict, patch-only vs Genesis
The page documents both paths (patch-only broken / Genesis P67 working) and
points readers at the right compose for their workload: tq3-mtp-genesis.yml
when they want MTP+TQ3, tq3-nomtp.yml when they want max KV without Genesis,
int8.yml as the production-safe baseline, turbo.yml for multi-tenant 4-stream.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The canonical club-3090 pin (1acd67a79... = 0.20.2rc1.dev129) is post-#41434
main and is NOT on Genesis v7.72.2's KNOWN_GOOD_VLLM_PINS allowlist. With
that pin + Genesis loaded, vLLM hits the transformers 5.8.0 cached_file
regression at `maybe_override_with_speculators` and aborts at boot:
OSError: Repo id must be in the form 'repo_name' or 'namespace/repo_name':
'/root/.cache/huggingface/qwen3.6-27b-autoround-int4'
Downgrade this compose only to nightly-01d4d1ad3 (=0.20.2rc1.dev9), which
Genesis explicitly allowlists. Validated end-to-end on dual 3090 (TP=2,
max-num-seqs=2, MTP n=3, TQ3 KV):
bench: 89.2 narr / 119.1 code TPS, CV 1-4% (matched-config-comparable)
verify-stress: 7/7 incl 60K needle PASS (Cliff 2 territory)
quality: 86/150 (57.3%) — within ~5pp of Qwen INT8 PTH baseline
soak: PASS, 0/100 silent_empty, 99.2% TPS retention, 0 MiB growth
aider-30: 18/30 (60.0%) — within noise of Qwen INT8 PTH (19/30)
KV pool: 1.22M tokens / 4.66× concurrency @ 262K (2× the INT8 PTH pool)
spec-decode: AL 3.50 sustained, per-position [0.95, 0.84, 0.75]
Re-evaluate the pin when Sander tags v7.73.x with a refreshed allowlist
(see memory: genesis_v773_memory_rework_pending).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Five-format matrix across TQ3/TQ4/k8v4 + MTP all fail long-context needle
with first-word repetition, regardless of cudagraph state or PR #40914 K+1
dispatch. TQ3 with MTP disabled passes 7/7 verify-stress, so this is an
MTP × TurboQuant multi-query interaction (Sander's P67/P67b territory),
not TQ precision.
PR #40914 actively makes things worse on our stack: dropping it lifts
verify-stress from 3/7 to 5/7. Per vllm.ai/blog/turboquant the upstream
position is that hybrid-attention models (which Qwen3-Next is) are not yet
supported by TurboQuant — matches our matrix.
Updates anchor the conclusion in tree:
- docs/UPSTREAM.md: #40914 reframed as OPEN/negative on this stack;
P67/P67b called out as the only known working multi-query fix.
- docs/FAQ.md: #40914 answer updated with the 2026-05-11 local result.
- models/qwen3.6-27b/INTERNALS.md: P67 → P67/P67b, #40914 status flipped.
- vllm/patches/vllm-pr40914-k1-only/README.md: prepended "negative result"
status section; round-2/3 MTP-skip experiment annotated as ineffective.
- vllm/patches/vllm-pr40914-k1-only/.../turboquant_attn.py: epilogue now
writes through the caller's output buffer (Codex fix; left for the
re-test artifact, not deployed).
- vllm/compose/dual/tq3-mtp.yml: re-tombstoned to point at the upstream
feature-compat gap rather than a 5-PR landing list.
Deployable Genesis-free path remains dual/tq3-nomtp.yml; TQ + MTP path
remains dual/tq3-mtp-genesis.yml.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
CLUB3090_TQ_K1_SKIP_MTP=1 (default, skips K+1 dispatch on mtp.* layers)
was validated 2026-05-11 round 3. Result: same `!`-flood corruption,
same 100% drafter+target agreement on wrong output. Rules out the
drafter as primary corruption source — the target full-attention layers'
K+1 dispatch is also producing wrong attention on this nightly + this
model layout.
Net: PR #40914's K+1 dispatch can't be salvaged with a layer filter
alone. Genesis P64/P65/P66/P68/P69 remains the only full fix; only 1
of those 5 has an upstream PR analog today. Compose stays tombstoned;
patches retained for re-test when upstream catches up.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Round-2 instrumentation (Codex co-diagnosed, 2026-05-11) captured the
PR #40914 K+1 dispatch firing first on `mtp.layers.0.self_attn.attn` —
the MTP drafter's attention layer, NOT a target-model full-attention
verify layer. Acceptance stabilized at AL=4.0 / ~100% but output
degenerated into repeated '!' on every MTP fire from 24-token prompts
on up, meaning drafter and target were agreeing on the same corrupt
synthetic-seq_lens attention path.
This commit adds a layer-name filter that excludes `mtp.*` from K+1
dispatch eligibility by default, preserving the original spec-verify
routing only for target-model attention layers. Env escape hatch
`CLUB3090_TQ_K1_SKIP_MTP=0` reverts to the unfiltered behavior for A/B.
Not yet validated — boot + verify-stress + bench pending.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Compose layout for Qwen 3.6 27B + TQ3 KV now reflects the three real
operating points after today's TQ3+MTP investigation:
- dual/tq3-mtp.yml — TQ3+MTP attempt without Genesis, TOMBSTONED.
Needs 4 of 5 missing upstream PRs (Genesis P64/P65/P66/P68/P69
equivalents); only PR #40914 has a community analog. Re-test when
upstream catches up.
- dual/tq3-nomtp.yml — TQ3 without MTP. Validated working on pure
upstream nightly: 1.73M KV pool, 6.59× concurrency at 262K,
verify-stress 7/7 pass. The deployable Genesis-free TQ3 path today.
- dual/tq3-mtp-genesis.yml — TQ3+MTP via Genesis, matched-config 2-stream
sibling of dual/turbo.yml's 4-stream production compose.
Plus retained research artifacts in patches/:
- vllm-pr40798-rebased/ — partial upstream workspace fix
- vllm-pr40914-k1-only/ — manually-rebased K+1 dispatch (still has
open question per vllm-issue#40880 closure history)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Catches the failure mode reported by alexpolo1 on Discord 2026-05-11:
switch.sh reports "no club-3090 container running" but the GPU is still
pinned at ~22 GiB from a non-managed process, and the new container
OOMs at boot with a cryptic vLLM ValueError.
Two changes:
1. Widen RUNNING_PATTERN from a hard-coded variant list to `^(vllm-|llama-cpp-)`
so down_running() also catches locally-built and one-off `docker run`
instances under the same image families.
2. Add gpu_preflight() between down_running() and up_variant():
- Queries nvidia-smi for free memory per GPU.
- If any card has <80% free (insufficient for the typical 0.92
gpu-memory-utilization), abort with a diagnostic listing the
holding PIDs from nvidia-smi --query-compute-apps and suggesting
specific cleanup commands.
- FORCE=1 env bypasses the check for users who know what they're doing.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Walked the full vllm#41403 5-gate chain for Gemma 4 + turboquant_3bit_nc on
2× RTX 3090 (SM 8.6). Gates 1-5 worked around; gate 6 is hardware-physical
(FA2 head_dim ≤ 256 on Ampere, Gemma 4 global layers use head_dim=512, FA3
support requires Hopper SM 9.0+).
Retains the int8-tq3.yml compose + the vendored #40108 overlay as a
re-test surface for when either (a) TQ backend gains pure-Triton prefill
for head_dim>256 on Ampere, or (b) vLLM gains per-layer attention backend
routing, or (c) we have Hopper hardware. None close.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Today's TQ3 boot probe surfaced that PR #41434 ([Perf][3/n] Eliminate
GPU<->CPU syncs in attention impls — MERGED) supersedes our local
patch_tolist_cudagraph.py — both anchors no longer match because
upstream replaced them with CPU-resident metadata fields (seq_lens_cpu,
query_start_loc_cpu) and a GPU `.tolist()` fallback when those aren't
populated. Patch ran clean, found no anchors, exited as no-op.
Related: PR #39931 (TurboQuant: support hybrid models and uniform
quantization — MERGED) is what unblocks Qwen3-Next + TurboQuant natively,
with no Genesis dependency. Confirmed by booting Qwen
int8-tq3.yml against vLLM nightly 1acd67a7: 1,416,904 tokens KV pool +
5.41× max concurrency @ 262K, no patches needed.
Bumps (19 Qwen composes): nightly-01d4d1ad (2026-05-04) → 1acd67a7 (2026-05-08).
Matches what Gemma already pinned; converges the stack on a single
nightly. Bump is small (4 days) but includes:
- #39931 TurboQuant + hybrid attention
- #41434 GPU<->CPU sync elimination ([Perf][3/n])
- #40092 TurboQuant FA3/FA4 prefill paths
- #41991 Gemma 4 tool parser bugfix (partial — #42006 still open)
- misc other fixes
Excluded from bump (need separate validation):
- models/gemma-4-31b/vllm/compose/dual/dflash.yml
- models/gemma-4-31b/vllm/compose/dual/dflash-int8.yml
Both pinned at nightly-e47c98ef (2026-05-06) with heavy vendored overlays
(vllm-gemma4-dflash/, vllm-pr40391-rebased/) keyed to specific upstream
file paths. Bumping their base nightly without re-validating the overlay
anchors risks silent breakage. To bump, A/B against current pin first.
patch_tolist_cudagraph removed from:
- models/qwen3.6-27b/vllm/compose/dual/int8-tq3.yml (mount + entrypoint)
- models/qwen3.6-27b/vllm/compose/dual/turbo.yml (mount + entrypoint)
- models/qwen3.6-27b/vllm/compose/dual/nvlink-turbo.yml (mount + entrypoint)
- models/qwen3.6-27b/vllm/compose/single/long-text.yml (mount + entrypoint)
- models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml (mount + entrypoint)
- models/gemma-4-31b/vllm/compose/dual/int8-tq3.yml (mount + wrapper entrypoint)
- models/qwen3.6-27b/vllm/patches/patch_tolist_cudagraph.py (file deleted)
- models/gemma-4-31b/vllm/patches/patch_tolist_cudagraph.py (file deleted)
Some composes still have patch_tolist_cudagraph mentioned in comments —
those are historical context about Genesis P78 / PN34 superseding it,
not active code, so left in place.
Gemma int8-tq3.yml: rewrote header as tombstone documenting the upstream
blocker chain. TurboQuant on Gemma 4 needs vllm#40108 (sliding-window
TQ) + vllm#41497 + vllm#40308 to merge. Tracker: vllm-issue#41403
(5-gate blocker stack). Until then, use dual/int8.yml for Gemma.
New today (BF16 baseline composes):
- models/qwen3.6-27b/vllm/compose/dual/bf16.yml (port 8012, 200K + n=4 + max-num-seqs=1)
- models/gemma-4-31b/vllm/compose/dual/bf16.yml (port 8033, 200K + n=4 + max-num-seqs=1)
Both default to 200K context with max-num-seqs=1 — that's BF16's natural
ceiling at TP=2 on 24 GiB cards.
Verified state after this commit:
- 28 composes on nightly-1acd67a7 (canonical)
- 2 composes on nightly-e47c98ef (DFlash family, deferred)
- 0 active patch_tolist_cudagraph mounts or entrypoint invocations
scripts/rebench-report.py:
- Date extraction via regex (handles tag-with-or-without YYYY-MM-DD suffix),
falls back to dir mtime.
- Quality latency parser now reads nested latency.{p50,p95,mean} (was looking
for non-existent flat fields).
- New 'Paste-ready compose Quality: schema line' subsection grepped from
quality-full.log.
- Rig info section in Meta: hostname, GPU models, per-card power cap,
parsed from rig.txt.
- Auto-computed TL;DR section (3-6 bullets covering TPS, KV pool,
verify-stress, quality, soak, aider).
- Phase timings table from timings.json with grand total.
- Reproducer-commands section with paste-ready bash.
- --compare-to <tag-dir> flag for numerical deltas across runs; writes
_internal.json sidecar so future runs can diff against this one.
- REPORT-discuss.md trimmed variant for GitHub Discussion comments.
- Aider per_language already fixed in prior commit; quant-string fix
in discuss header (was rendering 'INT16' from --dtype float16 suffix).
scripts/rebench-full.sh:
- Writes rig.txt (hostname + nvidia-smi -L + power cap) at preamble.
- Writes timings.json incrementally as each phase completes, via
record_timing helper called from run_step.
- quality-full.log is the source for the Quality: one-liner so the
rebench-report grep just works.
Validated against the synthesized Qwen INT8 PTH n=4 tag dir — all 9
sections render correctly.
aider's benchmark.py emits per-exercise results as a dict keyed by
'<lang>/<exercise>' path, nested under verifier_trace.trace. Earlier
parser assumed a flat list of {language, passed} dicts and so produced
no per-language tally.
Also handle the dual nesting (verifier_trace.trace vs flat) and use the
path prefix as the language key. Fall back to dict-aware tally when
top-level pass_rate/passed_count/total_count are None.
Validated against today's Qwen INT8 PTH n=4 aider rerun (quality-2026-05-11
T00-12-30.json) — produces:
Total: 19/30
cpp: 3/5 go: 4/5 java: 4/5 javascript: 4/5 python: 3/5 rust: 1/5
Total per-leg runtime drops ~2-2.5 hr → ~1.75-2 hr. Soak still gives 50 turns
which catches the silent-empty / Cliff 2b patterns just as reliably as 100
turns (those failures cluster in early sessions when they happen). Bump back
to SOAK_SESSIONS=20 for the canonical stability matrix when validating new
compose paths.
Adds qwen3.6-27b/vllm/compose/dual/int8.yml (new) — same vLLM nightly
1acd67a7 as gemma-int8, same int8_per_token_head KV class, MTP n=4
parameterized via SPEC_N_MAX env var. First validation of INT8 PTH KV
on Qwen3-Next DeltaNet hybrid attention on our stack.
Updates gemma-int8.yml to parameterize num_speculative_tokens via SPEC_N_MAX.
BENCHMARKS.md head-to-head section: matched-config addendum showing
Gemma's per-stream advantage shrinks to ~10% decode / ~45% TTFT (vs
+60% in the morning's mismatched run), and Qwen has +30% more KV pool /
concurrency. Original mismatched table preserved.
Side-effect: Qwen int8.yml is a candidate shipping path — +27/+33% TPS
over canonical dual.yml at -2.3 GiB/card VRAM.
Filter out 'chore(changelog): regenerate for vX.Y.Z [skip ci]' commits
that the release.yml workflow auto-creates. They're pure machine
traffic and only added noise to release notes.
CHANGELOG.md and GitHub Release notes now render commit subjects only.
Full commit message body (why/how/validation data) stays in 'git log',
one click away via the SHA link. Keeps both surfaces skim-readable.
v0.3.2 release pages were 100+ lines per commit; this brings them down
to ~1 line per commit.
The v0.3.2 tag was originally pushed at commit 64b0474 but GitHub didn't
emit a CreateEvent (likely dedup after delete + re-push to same SHA), so
the release.yml workflow never fired. Empty commit gives the tag a fresh
SHA that GitHub will process cleanly.
CHANGELOG.md is now auto-generated from commit messages by git-cliff in
the release workflow. Hand-edits below the static header will be wiped on
the next tag.
Workflow (`.github/workflows/release.yml`):
- On tag push (`v[0-9]+.[0-9]+.[0-9]+`):
1. Render GitHub Release body: `git-cliff --latest --strip header`
→ just the per-version section, no SemVer preamble repeat
2. Regenerate full CHANGELOG.md: `git-cliff` (default = all tags)
→ preserves header + all historical sections
3. Commit CHANGELOG.md back to master with `[skip ci]` marker
4. Publish GitHub Release with the latest-only body
Template (`cliff.toml`):
- `[changelog].header` now holds the SemVer preamble + CalVer→SemVer
mapping table (preserved across regens; stripped from GitHub Release
bodies via `--strip header`).
- `body` template now renders the **full commit message** (subject as
bold bullet, body indented below) instead of just the first line.
Rich narrative I write in commit message bodies (tables, validation
numbers, before/after diffs) now flows into both CHANGELOG.md and the
GitHub Release page from the same source.
- Per-release Pin/Diff footer guarded with `{% if version %}` so the
Unreleased section doesn't emit empty links.
CHANGELOG.md replaced with the auto-gen output. Past hand-written tables
and phase breakdowns are replaced by the corresponding commit messages
(those were already rich for commits that mattered — v0.3.1 soak-helper
fix has its Before/After table in the commit body and renders fine).
Going forward: just write rich commit messages and tag. Both surfaces
update automatically. No hand-edit of CHANGELOG.md required.
When the user runs quality-test.sh with a localhost-style URL
(default `http://localhost:8020`, or any `localhost`/`127.x`/`[::1]`
variant), auto-export `BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1` so
benchlocal-cli rewrites the hermes-agent's outbound model endpoint
from `localhost:<port>` to `host.docker.internal:<port>` inside the
Docker sandbox container.
Without this, the hermes-agent inside the sandbox can't reach the
host's vLLM (localhost resolves to the container itself) and every
scenario fails with `"API call failed after 3 retries: Connection
error."` — produced spurious 0/20 grades on this rig prior to the
benchlocal-cli runner.py 9c1566f fix.
Skips the auto-set when:
- User already set the env var (explicit override)
- URL points at a non-loopback host (real LAN IP, k8s service name,
host.docker.internal already) — no rewrite needed
Emits a stderr breadcrumb when the auto-set fires so users can see
what changed.
vLLM nightly (0.20.2rc1.dev9+) emits the qwen3 reasoning parser's output
under `delta.reasoning` (legacy field name), not `delta.reasoning_content`
that soak-helper.py was watching. Result: for any thinking-on response
whose `<think>` block doesn't close within `max_tokens`, soak-helper saw
zero deltas → fell back to the "couldn't measure" path → reported
`ttft_ms == t_ms` and `decode_tps = 0.0`. The model was generating
correctly; the harness just couldn't see the wire output.
Repro request (JDWarner's #107 turn 5): math problem with
`max_tokens=2000` + `chat_template_kwargs.enable_thinking=true`.
Before patch:
status=200 t_ms=22709 ttft_ms=22709 decode_tps=0.0
completion_tokens=2000 content="" reasoning_content=""
After patch (same request, same compose, same model):
status=200 t_ms=22709 ttft_ms=234 decode_tps=88.985
completion_tokens=2000 content="" reasoning_content="Here's a thinking
process:\n\n1. **Understand the User's Problem:**\n..." (3959 chars)
Validation soak (fresh-mode, 20 sessions × 5 turns = 100 turns, qwen3.6-27b
dual.yml):
verdict PASS
silent_empty 0 / 100 (0.0%) ← was ~3-5/40 baseline
p50_decode_tps 90.22
p95_ttft_ms 1389
errors 0
max_growth 0 MiB / 200
Closes the cross-rig "silent-empty turn-5" pattern parked behind the
Cliff 2b investigation — it was a harness measurement bug, not a model
or rig issue.
club-3090 is software (docker composes + system scripts + patch bundles
that downstream rigs run as-is), not just rolling recipes. Switch from
CalVer to SemVer from v0.3.0 onward; past CalVer tags (v2026.05.09,
v2026.05.10) preserved for history.
CHANGELOG.md: convention note + retroactive CalVer→SemVer mapping +
new v0.3.0 (2026-05-10) entry covering 8 commits since v2026.05.10:
- Qwen 3.6 27B thinking OFF default across all 21 composes
- MODEL_DIR UX overhaul (closes#116) — interactive setup prompt,
$MODEL_DIR placeholder everywhere
- power-cap-sweep --include-commit (closes#112)
- BENCHMARKS aider-polyglot-30 row
cliff.toml: drop "snapshot of the rolling stack — not a versioned
API" framing; add SemVer note. Tag pattern v[0-9]+.[0-9]+.[0-9]+
already matches both CalVer and SemVer, so the cliff release
workflow needs no changes.
New "Quality benches — Aider Polyglot 30" section captures pass-rate /
agentic-coding signal alongside the existing TPS rows. First two rows:
- Qwen 3.6 27B (AutoRound INT4) on dual.yml: 20/30 = 66.7%, 19 min wall
- Gemma 4 31B (Intel AutoRound INT4) on dual.yml: 17/30 = 56.7%, 19 min wall
Both run on 2× 3090 PCIe, 230 W cap, threads=2. Qwen edges Gemma by +10pp
despite being smaller; java is the biggest swing (Qwen 4/5 vs Gemma 1/5).
Critical caveat documented: Qwen with thinking ON is unusable for agentic
benches on this hardware — hits the 1500s subprocess cap before any
exercise completes. The new --default-chat-template-kwargs flag in our
vLLM Qwen composes (commit 534d29f) sets enable_thinking=false by default.
Aider-polyglot run via benchlocal-cli's aider-polyglot-30 pack. Cross-rig
re-run path documented in the section.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
@laurimyllari noted in disc #62 that with the project moving fast,
including the club-3090 git commit in sweep output helps correlate
cross-rig sweeps to the script revision they were run against. He
stamped `aa99173` manually; the script should do it for us.
New flag:
--include-commit Stamp the club-3090 git commit (short SHA) in the
report header next to the date. Off by default.
Implementation:
- Captures `git -C "$REPO_ROOT" rev-parse --short HEAD` once at header-build
time (REPO_ROOT was already known to the script).
- Injects into the report header next to **Date:**, e.g.:
**Date:** 2026-05-10T17:55:00Z **club-3090 commit:** `534d29f`
- Suppress (don't stamp "n/a") when run from a non-clone or git is
unreachable. Closes the curl-pipe-from-docs UX hole — `curl ... | bash`
users don't have a clone, so the field just disappears rather than
showing a confusing "n/a".
Off by default per the issue rationale: surprise stamping confuses
contributors running from documentation snippets.
Closes#112.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Follow-up to 29d17ed which used the wrong vLLM flag name (`--chat-template-kwargs`),
causing boot failure: "vllm: error: unrecognized arguments: --chat-template-kwargs
{"enable_thinking": false}".
vLLM's actual flag for setting server-side default chat template kwargs is
`--default-chat-template-kwargs` (with the `default-` prefix). Confirmed by
- vLLM nightly source: vllm/engine/arg_utils.py defines
`default_chat_template_kwargs: dict[str, Any] | None = None` with
json.loads parsing.
- vLLM PR #37739 ("Fix default_chat_template_kwargs handling in Responses API")
references it as already available in the shared render stack.
Behavior unchanged: thinking OFF by default for all 17 vLLM Qwen 27B composes
(plus bounded-thinking unaffected — it intentionally keeps thinking ON).
Per-request override still works via OpenAI extra_body:
{"chat_template_kwargs": {"enable_thinking": true}}
Verified: vllm-qwen36-27b-dual now boots cleanly with the corrected flag.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Previously, all 19 vLLM Qwen composes used --reasoning-parser qwen3 (parses
<think>...</think> output blocks) but did NOT explicitly disable the
thinking template. That meant Qwen3 default — thinking ON — applied
across the board. This burns hidden token budget on internal CoT for
every request, hurts latency, and creates a "Qwen looks much slower
than Gemma" gap on agentic benchmarks (which is exactly what we just
hit on aider-polyglot — Qwen exceeded the 1500s timeout, Gemma
finished in 19 min).
Change:
- All vLLM composes (18 of them) now pass `--chat-template-kwargs
'{"enable_thinking": false}'` after `--reasoning-parser qwen3`.
- llama.cpp single/docker-compose.yml: DISABLE_THINKING default flipped
0 → 1 (thinking now OFF by default; opt back in via DISABLE_THINKING=0).
- llama.cpp single/concurrent.yml: gained `--chat-template-kwargs` flag
with default `{"enable_thinking":false}` (overridable via
CHAT_TEMPLATE_KWARGS env).
NOT changed:
- bounded-thinking.yml — that's the structured-CoT compose where thinking
IS the feature. Reverted my initial blanket change for that one.
- qwopus-bf16mtp.yml — already had enable_thinking=false (preview compose).
Users who want thinking ON can:
- For vLLM: pass `chat_template_kwargs: {enable_thinking: true}` in the
per-request body (works fine).
- For llama.cpp: set `DISABLE_THINKING=0` in compose/.env.
Aligns with Gemma 4's "thinking off by default" (it ships that way upstream)
and removes the Qwen vs Gemma framework-bench skew.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Previously setup.sh silently defaulted MODEL_DIR to <repo>/models-cache,
which meant fresh users got ~14-21 GB of model weights downloaded INTO
their git tree without realizing it. The relative path is also wrong
when run from anywhere except the repo root.
New 4-step resolution order in setup.sh:
1. MODEL_DIR exported in calling shell → use as-is (unchanged)
2. .env at repo root sets MODEL_DIR → source it (NEW)
3. Interactive prompt (only on TTY) → ask user (NEW)
4. Silent fallback to <repo>/models-cache (unchanged for non-TTY)
The interactive prompt only fires when:
- MODEL_DIR is not in the calling env, AND
- .env doesn't already set it, AND
- both stdin AND stdout are TTYs (CI / scripted runs unaffected)
Three options offered:
1. <repo>/models-cache (the old silent default — kept as option)
2. $HOME/models (sensible cross-rig default)
3. custom absolute path
After picking, optionally persists the choice to .env (gitignored) so
re-runs skip the prompt. Existing .env files are updated in-place if
they already set MODEL_DIR; appended-to otherwise.
Closes the UX hole RobH589 hit in club-3090#116 — the relative
../../../../../models-cache default that was resolving wrong when
not run from the compose dir.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Two recipe scripts (single-card-default.sh, single-card-max-ctx.sh) had
the dev rig path /mnt/models/huggingface/... baked into the MODEL_PATH
default + the comment-block instructions for downloading the GGUF.
Updated:
- Comments now show \${MODEL_DIR}/qwen3.6-27b-gguf/... as the placeholder
(matches what the just-fixed README + LLAMA_CPP.md docs say).
- MODEL_PATH default changed from /mnt/models/huggingface/... to
\${MODEL_DIR:-\$HOME/models}/qwen3.6-27b-gguf/... — falls back to
~/models/ if MODEL_DIR isn't set, which is a more reasonable default
for cross-rig users than our /mnt/models/huggingface/ path.
- file-exists check at line 25 still fails loudly with the resolved path
if neither MODEL_DIR nor MODEL_PATH is set correctly.
Follow-up to fbf3431 (de-bind \$MODEL_DIR from rig path in docs).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
User docs were hardcoding the dev rig path (/mnt/models/huggingface/...) as
if it was canonical. It's not — cross-rig users have models at /data/models,
~/models, /mnt/nvme/llms, etc. Setting MODEL_DIR per their setup is the
intended UX (the compose already supports it via env-var default).
Updates:
- models/qwen3.6-27b/llama-cpp/README.md: download examples now use
\$MODEL_DIR/qwen3.6-27b-gguf/ instead of /mnt/models/huggingface/...
- models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml: header
comment uses \$MODEL_DIR/qwen3.6-27b-gguf/ for download examples + says
"MODEL_DIR=/your/models/dir docker compose up -d" instead of our path.
- docs/engines/LLAMA_CPP.md: same treatment + cleaned up Qwen3.5 + DFlash
draft path examples to also use \$MODEL_DIR.
- scripts/preflight.sh: hf download hint shows literal \${MODEL_DIR} so user
knows what to set, plus explicit "set MODEL_DIR first" line. Previously
echoed the resolved relative path (../../../../models-cache) which lands
outside the repo if pwd isn't the compose dir.
Caught by RobH589 in club-3090#116 — they hit the path-resolved-to-root-of-drive
case from the relative-path default. Closes the doc UX side; the compose's
env-var override mechanism was already correct.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
After the GGUF dir move (/mnt/models/gguf/qwen3.6-27b/ → /mnt/models/
huggingface/qwen3.6-27b-gguf/), four refs were not updated and led to
a path-mismatch loop reported in club-3090#116:
- scripts/preflight.sh `hf download` hint pointed at qwen3.6-27b/, but
the compose default expects qwen3.6-27b-gguf/. Same for the mv hint
for mmproj relocation, and the in-container mmproj default at line 329.
- models/qwen3.6-27b/llama-cpp/README.md example command still said
`MODEL_DIR=/mnt/models/gguf` (now /mnt/models/huggingface).
- models/qwen3.6-27b/llama-cpp/compose/single/{docker-compose,concurrent}.yml
header comment said "Q5_K_XL" but the actual default has been Q3_K_XL
for a while (this one predates the reorg — just stale doc).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The dev rig had grown across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`,
`/home/wasif/`) and three repos (single-3090, dual-3090, club-3090). Disk-out
on `/` (97% used) forced a cleanup; rather than just prune, we consolidated
the layout across the whole stack while we were at it. This commit captures
what landed inside this repo.
Services consolidation:
- services/{ollama,openwebui,litellm,qdrant,searxng}/ migrated in from
/opt/ai/compose/<svc>/ (zero functional change — same docker-compose.yml).
- services/litellm/config.yaml rewritten: explicit routes for current
primaries (qwen3.6-27b-autoround → :8010, gemma-4-31b-autoround → :8030).
Removed `* → ollama/*` wildcard.
- services/comfyui/ migrated in (was /opt/ai/compose/comfyui/) — wired into
gpu-mode with full mutex against vLLM/SGLang.
scripts/gpu-mode.sh under git:
- Was a loose /opt/ai/gpu-mode.sh outside any repo. Symlinked at
/usr/local/bin/gpu-mode.
- Five Gemma 4 31B modes added: gemma, gemma-dflash, gemma-int8,
gemma-dflash-int8, gemma-awq.
- One ComfyUI mode (mutex with all LLM serving).
- prune / prune-all subcommands (safe image prune; aggressive variant adds
build cache --keep-storage 5GB + dangling networks).
- gpu-mode status now shows Docker disk + /var/lib/docker + /tmp sizes.
- compose_at() passes --env-file <repo>/.env so MODEL_DIR resolves
regardless of which compose dir gpu-mode cd's into. Fixes the recurring
"MODEL_DIR not set, defaulting to ../../../../../models-cache" warning.
- stderr no longer swallowed by compose_at() (real errors surface).
- Cross-model VRAM mutex: every Qwen mode stop_all_gemma + stop_comfyui
and vice-versa.
scripts/maintenance/ — new hygiene-tools subdir:
- list-image-pins.sh: engine-agnostic pin auditor. Scans every compose's
`image:` line, groups by `<repo>:<tag>`, flags pin-drift (multiple tags
per repo), ranks composes by patch surface.
Pin tracking:
- docs/UPSTREAM.md gains a "Pinned images" section: table of every pinned
image, why each pin exists, retirement candidate criteria.
- docs/NIGHTLY_BUMP_RUNBOOK.md (new): 7-step procedure for bumping pinned
engine images (scope → branch → patch survival → boot → verify-full +
verify-stress → bench delta → land → retire). Engine-specific notes
for vLLM nightly hashes, llama.cpp digest pinning, SGLang variants.
Path updates from the engine + model dir consolidation:
- /opt/ai/vllm-src/ → /opt/ai/engines/vllm/primary/
(in setup.sh, INTERNALS.md, several patch READMEs, docs/HARDWARE.md,
docs/FAQ.md, docs/DUAL_CARD.md, docs/UPSTREAM.md, models/qwen3.6-27b/
CHANGELOG.md)
- /mnt/models/gguf/qwen3.6-27b/ → /mnt/models/huggingface/qwen3.6-27b-gguf/
(in models/qwen3.6-27b/llama-cpp/{compose/single/*.yml, recipes/*.sh,
README.md}, docs/engines/LLAMA_CPP.md)
CHANGELOG.md narrative gap fill:
- 2026-05-10 entry for this reorg.
- 2026-05-09 entry for compose convention formalization (topology
promoted to dir level, profile schema, Status enum + Caveats, cliff
CI swap, Discord launch).
- 2026-05-08 entry for Gemma 4 INT8 PTH unblock + 262K validation.
- 2026-05-07 entry for power-cap-sweep campaign + HARDWARE.md cross-rig
charts + cross-rig benchmark rows.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Lands the Gemma 4 31B work that was in flight last week + Qwopus3.6-27B
preview compose. Captures four distinct serving paths for Gemma 4:
- dual/awq.yml — AWQ-4bit (text-only, simplest path)
- dual/dflash-int8.yml — DFlash + INT8 PTH KV (262K ctx, full pipeline)
- vllm-gemma4-dflash-int8/ — vendored patches stacking DFlash spec-decode
+ INT8 PTH KV across model_executor + v1/spec_decode + v1/attention +
v1/worker (~13 patched files; PR #42102 + #40391-rebased + tool-parser
fixes#42006 + #41991 stacked)
- vllm-gemma4-fp8-ampere/ — earlier Phase 2 attempt before INT8 PTH
reframe (kept for forensics); Ampere has no native FP8 tensor cores
- vllm-perheadkv-hybridpage-fix/ — hybrid-page bug fix surfaced during
Phase 3
- vllm-pr40391-perheadkv/ — PR #40391 vendored at the tree level
(separate from the rebased variant under refs/jianc99-dflash-gemma4)
Plus:
- models/qwen3.6-27b/vllm/compose/dual/qwopus-bf16mtp.yml — preview compose
for Carnice AutoRound Recipe D output (port 8071, NOT production —
see club-3090-todo.md for known gaps + cheap A/Bs).
- AGENTS.md — codify compose naming + profile-schema + experimental-compose
conventions that the new files follow.
- docs/QUALITY_TEST.md — runbook for `quality-test.sh` + the benchlocal-cli
packs it wraps.
CHANGELOG.md narrative entries for these are added separately.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Two findings from @easel's deep cross-rig writeup on issue #102 (5090
Laptop WSL2):
1. **HARDWARE.md** — new "GPU memory budget on WSL2" subsection. WSL2
container CUDA-context consumes ~1.31 GiB before vLLM profiler runs.
`gpu_memory_utilization=0.95` crashes at boot; 0.944 works. Validated
on 2 machines. Formula: `(vram_total - 1.31) / vram_total`.
Recommendation: `GPU_MEMORY_UTILIZATION=0.94` in `.env` on WSL2.
2. **CLIFFS.md** — new section "Cliff 3 — DeltaNet SSM state is not
prefix-cacheable (the prefill cliff)". This is a structural finding
that explains a class of failures we'd been describing without
naming. @easel's warm-cache run (68.7% KV-block hit, turn 10 at
35.6K tokens took 577s — 2.3× the cold-start 254s) is the smoking
gun: prefix cache helps attention but DeltaNet's recurrent state
`h_t = f(h_{t-1}, x_t)` must be recomputed from scratch every turn.
PN32 fixes OOM stability; nothing fixes the O(n) prefill scaling
because the architecture itself is sequential.
Practical ceiling on single-card vLLM (any 24 GB Qwen3-Next config):
sub-30s TTFT only below 5K accumulated tokens. 22-35K = 3-4 min/turn.
~74K = 10+ min client timeout.
Elevates "for single-card agentic Qwen3-Next, use llama.cpp" from
implied to explicit in the docs. Dual-card extends the envelope to
25-30K accumulated; deep sessions (50K+) still need llama.cpp on
either topology.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Three related fixes for running benchmarks without IDE agent interference:
1. All 18 vLLM compose files: port binding is now
${BIND_HOST:-0.0.0.0}:${PORT:-<n>}:8000
Setting BIND_HOST=127.0.0.1 in .env restricts the API to localhost,
preventing IDE agents (Cline, Cursor) from competing for the
max-num-seqs=1 slot and causing verify-stress HTTP 000 failures.
2. scripts/preflight.sh: port auto-detection regex now matches
127.0.0.1:<port>->8000/tcp in addition to 0.0.0.0: and [::]:
Previously all verify-*/bench scripts silently produced no output
when BIND_HOST=127.0.0.1 was set.
3. scripts/report.sh: SOAK_TIMEOUT_S is now forwarded to soak-test.sh
from the shell environment. Previously the variable was read from
compose .env (docker-compose only) and silently ignored by the
script, always using the 1800s default regardless of what was set.
Docs: .env.example gains BIND_HOST and SOAK_TIMEOUT_S entries.
Co-authored-by: Claude Sonnet 4.6 <[email protected]>
Four small fixes addressing low soak-test compliance in cross-rig bench
contributions. Audit of 5 recent #113/#107/#102/#104/#93 showed soak data
IS being run but it's hidden in the main report and the dedicated
template field comes out empty (template said "leave blank if you ran
--full"). Older BENCHMARKS rows often omit soak verdict entirely.
1. **`.github/ISSUE_TEMPLATE/numbers-from-your-rig.yml`**: replace optional
"soak summary" textarea with a required dropdown listing PASS / borderline /
FAIL / Skipped+reason / Not-yet-run. Verdict is now grep-able even when
the data is buried in the main report textarea.
2. **`scripts/soak-test.sh`**: add `--continuous` / `--quick` / `--fresh`
flags + `--help` + cleaner usage docs. Was 5 env vars to invoke
(`SOAK_MODE=continuous SOAK_SESSIONS=5 SOAK_TURNS=5 CONTAINER=... ENDPOINT=...`);
now `bash scripts/soak-test.sh --continuous` does the same with
auto-detect (existing logic preserved + exposed). Env vars still work
for back-compat.
3. **`scripts/report.sh`**: when `--bench` (or partial) ran without
`--soak`/`--full`, append a "⚠ Soak: not included" reminder block to
the report so contributors know what's missing before pasting into
the issue template.
4. **`BENCHMARKS.md`**: Notes-column convention — every row should start
with explicit `Soak: ✓ PASS` / `⚠ borderline` / `✗ FAIL` / `—` so
readers can grep at a glance. Updated 2 recent rows (ygafarov #113,
JDWarner #107) to use the convention. Older rows backfill as the
convention spreads.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
First Strix-Halo-miniPC + oculink-eGPU class on the matrix. Single 3090
over PCIe x4 (oculink) on AMD Ryzen AI MAX+ 395 / 124 GB RAM / CachyOS /
290W cap. Result: 68.86 narr / 91.70 code TPS via vllm/default + TQ3 at
48K — clean MTP AL 3.31 (77% accept), CV 1.6%/2.7%.
Soak FAIL is borderline (240 MiB > 200 MiB threshold, 3 turns >30s) but
100% TPS retention + 0 errors + 0 silent-empty suggests x4-PCIe accretion
+ bus-latency under prefill, not Cliff 2b. Worth flagging as a possible
"eGPU bus class" threshold allowance for soak-test.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>