Commit Graph
339 Commits
Author SHA1 Message Date
github-actions[bot] 722f998ff3 chore(changelog): regenerate for v0.5.0 [skip ci] 2026-05-12 21:23:33 +00:00
noonghunnaandClaude Opus 4.7 84498d47aa feat(qwen): ship froggeric chat-template fixes as default-on
Release / release (push) Failing after 50s
Vendored snapshot of froggeric/Qwen-Fixed-Chat-Templates qwen3.6
template, mounted into all 22 vanilla Qwen 3.6-27B composes via
--chat-template. Replaces the model's default Jinja template with
the community-patched one. Carnice and Qwopus composes intentionally
excluded — they ship bespoke Hermes-JSON templates that must not be
overwritten.

Upstream fixes seven documented bugs in the default Qwen 3.5 / 3.6
templates:
  - empty <think></think> blocks polluting past-turn context
  - </thinking> closing-tag hallucination on Qwen 3.6
  - unclosed <think> before tool_call (mangled output)
  - raise_exception crash when no user query in messages (kills
    agentic loops)
  - "developer" role rejection (blocks modern API clients)
  - |items Jinja filter unsupported in C++ runtimes (llama.cpp,
    LM Studio, MLX)
  - type-aware tojson serialization for tool arguments
Source: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

Surfaced by @troymroberts in club-3090 discussion #121.

First-pass A/B 2026-05-12 vs the matched-config Qwen INT8 PTH n=4
rebench baseline (2026-05-10):

  Pack             Baseline   Froggeric   Δ
  ---------------- --------   ---------   --
  toolcall-15       10/15      10/15       0
  instructfollow    13/15      13/15       0
  structoutput      13/15      13/15       0
  dataextract       15/15      15/15       0
  reasonmath         6/15       6/15       0
  bugfind           11/15      11/15       0
  hermesagent-20     9/20      12/20      +3 (+15pp)
  cli-40            17/40      17/40       0
  TOTAL             94/150     97/150     +3 (+2pp)

hermesagent-20 is the multi-turn agentic pack — exactly where the
empty-think + no-user-query + unclosed-think-before-tool-call fixes
compound. 7 other packs flat = no regression on single-turn flows.

Added a new "Community templates / model assets" section to
docs/UPSTREAM.md tracking the resource + drop trigger (replace when
upstream Qwen pushes equivalent fixes).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
v0.5.0
2026-05-12 21:21:37 +00:00
noonghunnaandClaude Opus 4.7 28b16b5dc9 feat(vllm): add PR #35936 required-tool fallback overlay
Vendored local overlay for vllm-project/vllm#35936 — fixes empty
tool_calls=[] response when tool_choice="required" is used with the
qwen3_coder parser. Under "required", vLLM forces structured JSON
output regardless of which parser is configured; the qwen3_coder
parser then scans the JSON for its XML <tool_call> sentinel, finds
none, and returns tools_called=False.

Overlay reshapes _parse_tool_calls_from_content() in
vllm/entrypoints/openai/engine/serving.py to try JSON-validate first
when tool_choice="required", and fall back to the configured parser
only when validation fails — covering both the supports_required_and_named
True and False parser paths.

End-to-end validated 2026-05-12:
- curl tool_choice="required" + get_weather tool: tool_calls populated
- curl tool_choice="auto" regression: still works
- MLS-Bench ml-ensemble-boosting with thinking.enabled=false: agent
  reaches Step 1 (edit) with no "No action returned" stall

Mounted into the 23 Qwen 3.6-27B composes pinning the post-#41434
nightly-1acd67a79 image. The single compose on the older 01d4d1ad3
pin (dual/tq3-mtp-genesis.yml) is intentionally excluded — its source
tree diverges from the rebase base.

Added UPSTREAM.md row tracking the upstream PR + drop trigger.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 21:20:50 +00:00
noonghunnaandClaude Opus 4.7 62b3b455a9 docs(dtype-matrix): more polish — RDNA naming, FP8 maturity caveats, AMD detection
- RDNA 4: include RX 8000 alongside RX 9000 (early roadmap name; AMD
  shipped under RX 9000 brand). Future-proofs for both references.
- Intel Xe2 FP8: add explicit maturity caveat — silicon supports it but
  full production-stack acceleration in vLLM-SYCL / IPEX-LLM is still
  maturing in mid-2026. OpenVINO has the most coverage today.
- AMD CDNA 4 MX*: clarify that NVFP4 is NVIDIA-proprietary (E2M1 + E4M3
  per-16-element scale), so AMD targets the OCP-standard block-scaled
  FP4/FP6 (MXFP4 / MXFP6) and AMD variants — for cross-vendor 4-bit
  portability, MXFP4 is the safer target.
- AMD detection: fix the cross-vendor routing code block. HIP-via-PyTorch
  appears as torch.cuda; reliable distinguishers are torch.version.hip
  (not None) and torch.cuda.get_device_properties(0).gcnArchName plus
  rocm-smi --showproductname. NVIDIA detection now explicitly excludes
  torch.version.hip not None to avoid HIP false-positives.
- Per-vendor KV cache option tables added for Intel + AMD — symmetric
  with NVIDIA's KV section above. Captures vLLM-ROCm FP8 KV on CDNA 3+
  and RDNA 4; OpenVINO/IPEX-LLM coverage on Intel; llama.cpp portable
  path everywhere.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 12:16:40 +00:00
noonghunnaandClaude Opus 4.7 3d4548c504 docs(dtype-matrix): polish nuances + add Intel and AMD vendor sections
Polish edits (per Grok's feedback):
- "Last verified" date prominently at the top
- Ada FP8 nuance: native TC support but ~½ Hopper's FP8 TFLOPS; lacks
  transformer-engine integration (on-die format conversion, mixed-precision)
- INT4 TC throughput jump Turing → Ampere: ~2× + 2:4 sparsity path
- TF32 inference note: training/mixed-precision convenience, rarely the
  active path for LLM inference
- Blackwell-today note: vLLM nightlies still prefer FP8 paths; NVFP4
  artifacts for 27-31B models scarce as of 2026-05
- MX* family accuracy benefit: per-32-element block scaling recovers
  dynamic range, scores measurably better than plain low-bit at same width
- Emerging KV recipes: block-scaled FP8, NVFP4 KV, TensorRT-LLM W4A8/W4A4

Vendor expansion — Intel + AMD sections:
- Intel XMX matrix: Xe-HPG (Arc A) / Xe-HPC (Flex/Max) / Xe2 (Battlemage,
  Lunar Lake) with INT2/INT4/INT8/FP16/BF16, partial FP8 emerging
- AMD WMMA (RDNA 2/3/4) + MFMA (CDNA 2/3/4) matrices — RDNA 4 brings
  consumer FP8 + 2:4 sparsity; CDNA 4 adds FP6/FP4/MXFP*
- Per-vendor "what this means for LLM serving today" tables
- Cross-vendor routing section with detection logic + tier mapping
- Expanded references — Intel oneAPI/XMX, AMD ROCm/WMMA/MFMA/CK, OCP MX
  spec, llama.cpp portability path

The doc is now genuinely multi-vendor reference material; the NVIDIA half
is still where the actionable compose work happens, but the rest gives
context for cross-rig PRs that want to add compose/intel/ or compose/amd/
paths once cross-rig benchmark data justifies them.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 12:12:45 +00:00
noonghunnaandClaude Opus 4.7 9c6d3cfba1 docs(dtype-matrix): per-arch hardware accelerator matrix for compose optimization
Adds docs/DTYPE_MATRIX.md — a reference table mapping NVIDIA GPU
architectures (Pascal → Blackwell DC) to native vs emulated Tensor Core
support for each dtype + quant scheme.

Sections:
- At-a-glance compute-dtype matrix (9 archs × 10 dtypes, ✓/SW/✗ marks)
- Weight-quantization schemes (GPTQ / AWQ / AutoRound / NF4 / SmoothQuant
  / FP8 weights / MXFP8 / MXFP4 / NVFP4 / GGUF / HQQ-AQLM-SqueezeLLM)
  with storage-vs-compute paths
- Weight-only vs weight+activation axis (W4A16 vs W8A8 vs W4A8)
- KV-cache dtype support (FP16 / FP8 / INT8 PTH / TQ3 / TQ4 / k8v4)
- Per-arch compose recommendations (which compose to ship for which
  GPU class)
- Runtime detection (points at Genesis guards.py)
- Corner cases (Ada FP8 vs Hopper FP8, Blackwell consumer vs DC, NVFP4
  vs MXFP4 block-size differences, MX* family overview, FP6)
- References (NVIDIA whitepapers, Marlin, Genesis, BENCHMARKS)

Cross-linked from:
- HARDWARE.md GPU-compat table
- GLOSSARY.md Quantization section
- FAQ.md as a new "What dtype/quant should I pick for my GPU?" Q under
  Hardware

This sets the foundation for future per-arch compose optimization —
detecting compute capability at boot and picking the right KV dtype /
weight quant scheme automatically based on what the hardware actually
accelerates rather than what's nominally supported.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 12:04:16 +00:00
noonghunnaandClaude Opus 4.7 1770931729 bench(matrix): @ygafarov first heterogeneous Ampere + Blackwell eGPU dual
Adds the dual-card row for @ygafarov's MiniPC eGPU setup from #120:
3090 via USB4 dock (PCIe 3.0 x4 ≈ 3.94 GB/s) + 5070 Ti via OCuLink
(PCIe 4.0 x4 ≈ 7.88 GB/s) on AMD Ryzen AI MAX+ 395 / Strix Halo. fp8
KV, 200K ctx, MTP n=3 — 65.10 / 85.81 narr/code TPS at 1.00× concurrency.

Notable bits captured in the row:
- First heterogeneous Ampere + Blackwell consumer eGPU dual on the matrix
- 5070 Ti at 91% util but only 125 W (out of 290 W cap) — visibly waiting
  on the 3090; vLLM compiles for sm_86 across both cards
- KV pool capped at 200K @ 1.00× by the 5070 Ti's 16 GiB VRAM, not the
  3090's 24 GiB
- verify-stress 7/7 including 91K needle recall (Cliff 2 clean)
- Soak ⚠ borderline 360 MiB VRAM growth — same eGPU-bus accretion as his
  own #113 single-card row at 240 MiB; not a leak, just x4-PCIe allocator
  behavior under prefill
- Slower than his own single-3090 row from #113 (68.86 / 91.70 at 48K),
  so the row also serves as the canonical "when is dual worse than single
  on an eGPU rig" data point

Reply with diagnosis + experiments queued at #120.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 11:17:16 +00:00
noonghunnaandClaude Opus 4.7 df53287b1c docs(faq): add 'INT8 PTH doesn't scale at concurrency — is that a bug?'
Surfaces the canonical int8_per_token_head anti-scaling trade-off so the
next dual-3090 user who hits this finds the answer via search rather than
having to reason from the head-to-head matrix.

Frames it as:
- not a bug — per-(token, head) scale serializes at concurrency, fp8's
  single global scale doesn't
- pick by workload: INT8 PTH for single-stream throughput, fp8 for
  aggregate concurrency, Genesis-backed TQ3+MTP if you want both
- includes the diagnostic checklist for matching baseline numbers
  (power cap, vLLM nightly, Genesis on/off, MTP n, prompt shape) so users
  can self-troubleshoot a cross-rig gap before posting

Triggered by a Discord question on a dual-3090 + NVLink rig that observed
the canonical INT8 anti-scaling pattern (150 TPS single / flat at concurrency)
vs fp8 (70-100 TPS single / 400+ at concurrency).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 09:58:16 +00:00
noonghunnaandClaude Opus 4.7 c2b1c93872 docs(tq3-mtp): writeup + charts for the Genesis-backed TQ3+MTP path
Adds `docs/TQ3_MTP_GENESIS.md` — user-facing writeup of the 6th matched-config
rebench leg (qwen-tq3-mtp-genesis-2026-05-11). Frames the result as a sweet
spot on the quality-vs-concurrency frontier: 1.22M-token KV pool at 262K ctx
(2× the INT8 PTH baseline), 4.66× concurrency, within ~5pp quality of INT8 PTH
on the 150-scenario quality suite, within noise on aider-polyglot-30 (18/30
vs 19/30 baseline). 7/7 verify-stress, 0/100 silent-empty in soak.

Charts under `docs/assets/tq3-mtp-genesis/`:
- 01-kv-pool-by-config.png — KV pool capacity by leg
- 02-quality-vs-concurrency.png — trade-off frontier
- 03-verdict-matrix.png — per-phase verdict, patch-only vs Genesis

The page documents both paths (patch-only broken / Genesis P67 working) and
points readers at the right compose for their workload: tq3-mtp-genesis.yml
when they want MTP+TQ3, tq3-nomtp.yml when they want max KV without Genesis,
int8.yml as the production-safe baseline, turbo.yml for multi-tenant 4-stream.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 20:32:43 +00:00
noonghunnaandClaude Opus 4.7 570fa71240 compose(tq3-mtp-genesis): pin to Genesis v7.72.2 known-good vLLM nightly
The canonical club-3090 pin (1acd67a79... = 0.20.2rc1.dev129) is post-#41434
main and is NOT on Genesis v7.72.2's KNOWN_GOOD_VLLM_PINS allowlist. With
that pin + Genesis loaded, vLLM hits the transformers 5.8.0 cached_file
regression at `maybe_override_with_speculators` and aborts at boot:

  OSError: Repo id must be in the form 'repo_name' or 'namespace/repo_name':
  '/root/.cache/huggingface/qwen3.6-27b-autoround-int4'

Downgrade this compose only to nightly-01d4d1ad3 (=0.20.2rc1.dev9), which
Genesis explicitly allowlists. Validated end-to-end on dual 3090 (TP=2,
max-num-seqs=2, MTP n=3, TQ3 KV):

  bench:         89.2 narr / 119.1 code TPS, CV 1-4% (matched-config-comparable)
  verify-stress: 7/7 incl 60K needle PASS (Cliff 2 territory)
  quality:       86/150 (57.3%) — within ~5pp of Qwen INT8 PTH baseline
  soak:          PASS, 0/100 silent_empty, 99.2% TPS retention, 0 MiB growth
  aider-30:      18/30 (60.0%) — within noise of Qwen INT8 PTH (19/30)
  KV pool:       1.22M tokens / 4.66× concurrency @ 262K (2× the INT8 PTH pool)
  spec-decode:   AL 3.50 sustained, per-position [0.95, 0.84, 0.75]

Re-evaluate the pin when Sander tags v7.73.x with a refreshed allowlist
(see memory: genesis_v773_memory_rework_pending).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 18:19:07 +00:00
noonghunnaandClaude Opus 4.7 9fba03788e docs(qwen-tq3): close round-4 — #40914 not shippable, route to nomtp + Genesis
Five-format matrix across TQ3/TQ4/k8v4 + MTP all fail long-context needle
with first-word repetition, regardless of cudagraph state or PR #40914 K+1
dispatch. TQ3 with MTP disabled passes 7/7 verify-stress, so this is an
MTP × TurboQuant multi-query interaction (Sander's P67/P67b territory),
not TQ precision.

PR #40914 actively makes things worse on our stack: dropping it lifts
verify-stress from 3/7 to 5/7. Per vllm.ai/blog/turboquant the upstream
position is that hybrid-attention models (which Qwen3-Next is) are not yet
supported by TurboQuant — matches our matrix.

Updates anchor the conclusion in tree:

- docs/UPSTREAM.md: #40914 reframed as OPEN/negative on this stack;
  P67/P67b called out as the only known working multi-query fix.
- docs/FAQ.md: #40914 answer updated with the 2026-05-11 local result.
- models/qwen3.6-27b/INTERNALS.md: P67 → P67/P67b, #40914 status flipped.
- vllm/patches/vllm-pr40914-k1-only/README.md: prepended "negative result"
  status section; round-2/3 MTP-skip experiment annotated as ineffective.
- vllm/patches/vllm-pr40914-k1-only/.../turboquant_attn.py: epilogue now
  writes through the caller's output buffer (Codex fix; left for the
  re-test artifact, not deployed).
- vllm/compose/dual/tq3-mtp.yml: re-tombstoned to point at the upstream
  feature-compat gap rather than a 5-PR landing list.

Deployable Genesis-free path remains dual/tq3-nomtp.yml; TQ + MTP path
remains dual/tq3-mtp-genesis.yml.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 16:22:27 +00:00
noonghunnaandClaude Opus 4.7 063d3e943c docs(qwen-tq3): re-tombstone tq3-mtp.yml after round-3 MTP-skip validation
CLUB3090_TQ_K1_SKIP_MTP=1 (default, skips K+1 dispatch on mtp.* layers)
was validated 2026-05-11 round 3. Result: same `!`-flood corruption,
same 100% drafter+target agreement on wrong output. Rules out the
drafter as primary corruption source — the target full-attention layers'
K+1 dispatch is also producing wrong attention on this nightly + this
model layout.

Net: PR #40914's K+1 dispatch can't be salvaged with a layer filter
alone. Genesis P64/P65/P66/P68/P69 remains the only full fix; only 1
of those 5 has an upstream PR analog today. Compose stays tombstoned;
patches retained for re-test when upstream catches up.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 15:06:19 +00:00
noonghunnaandClaude Opus 4.7 6b2a7d553b feat(qwen-tq3): add CLUB3090_TQ_K1_SKIP_MTP layer-filter for PR #40914 K+1 dispatch
Round-2 instrumentation (Codex co-diagnosed, 2026-05-11) captured the
PR #40914 K+1 dispatch firing first on `mtp.layers.0.self_attn.attn` —
the MTP drafter's attention layer, NOT a target-model full-attention
verify layer. Acceptance stabilized at AL=4.0 / ~100% but output
degenerated into repeated '!' on every MTP fire from 24-token prompts
on up, meaning drafter and target were agreeing on the same corrupt
synthetic-seq_lens attention path.

This commit adds a layer-name filter that excludes `mtp.*` from K+1
dispatch eligibility by default, preserving the original spec-verify
routing only for target-model attention layers. Env escape hatch
`CLUB3090_TQ_K1_SKIP_MTP=0` reverts to the unfiltered behavior for A/B.

Not yet validated — boot + verify-stress + bench pending.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 14:46:40 +00:00
noonghunnaandClaude Opus 4.7 6182922225 refactor(qwen): rename int8-tq3 → tq3-* family + add no-MTP + Genesis variants
Compose layout for Qwen 3.6 27B + TQ3 KV now reflects the three real
operating points after today's TQ3+MTP investigation:

- dual/tq3-mtp.yml — TQ3+MTP attempt without Genesis, TOMBSTONED.
  Needs 4 of 5 missing upstream PRs (Genesis P64/P65/P66/P68/P69
  equivalents); only PR #40914 has a community analog. Re-test when
  upstream catches up.
- dual/tq3-nomtp.yml — TQ3 without MTP. Validated working on pure
  upstream nightly: 1.73M KV pool, 6.59× concurrency at 262K,
  verify-stress 7/7 pass. The deployable Genesis-free TQ3 path today.
- dual/tq3-mtp-genesis.yml — TQ3+MTP via Genesis, matched-config 2-stream
  sibling of dual/turbo.yml's 4-stream production compose.

Plus retained research artifacts in patches/:
- vllm-pr40798-rebased/  — partial upstream workspace fix
- vllm-pr40914-k1-only/  — manually-rebased K+1 dispatch (still has
  open question per vllm-issue#40880 closure history)

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 14:46:05 +00:00
github-actions[bot] a3b66c489c chore(changelog): regenerate for v0.4.0 [skip ci] 2026-05-11 11:24:38 +00:00
noonghunnaandClaude Opus 4.7 4866913a10 fix(switch): GPU memory pre-flight + widen RUNNING_PATTERN
Release / release (push) Failing after 49s
Catches the failure mode reported by alexpolo1 on Discord 2026-05-11:
switch.sh reports "no club-3090 container running" but the GPU is still
pinned at ~22 GiB from a non-managed process, and the new container
OOMs at boot with a cryptic vLLM ValueError.

Two changes:

1. Widen RUNNING_PATTERN from a hard-coded variant list to `^(vllm-|llama-cpp-)`
   so down_running() also catches locally-built and one-off `docker run`
   instances under the same image families.

2. Add gpu_preflight() between down_running() and up_variant():
   - Queries nvidia-smi for free memory per GPU.
   - If any card has <80% free (insufficient for the typical 0.92
     gpu-memory-utilization), abort with a diagnostic listing the
     holding PIDs from nvidia-smi --query-compute-apps and suggesting
     specific cleanup commands.
   - FORCE=1 env bypasses the check for users who know what they're doing.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
v0.4.0
2026-05-11 11:20:40 +00:00
noonghunnaandClaude Opus 4.7 f8c706699c docs(gemma-4-31b): document TQ3 Ampere FA2 head_dim wall + vendor #40108 overlay
Walked the full vllm#41403 5-gate chain for Gemma 4 + turboquant_3bit_nc on
2× RTX 3090 (SM 8.6). Gates 1-5 worked around; gate 6 is hardware-physical
(FA2 head_dim ≤ 256 on Ampere, Gemma 4 global layers use head_dim=512, FA3
support requires Hopper SM 9.0+).

Retains the int8-tq3.yml compose + the vendored #40108 overlay as a
re-test surface for when either (a) TQ backend gains pure-Triton prefill
for head_dim>256 on Ampere, or (b) vLLM gains per-layer attention backend
routing, or (c) we have Hopper hardware. None close.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 02:50:04 +00:00
noonghunna 16a1374053 chore(composes): bump Qwen pins → 1acd67a7, drop obsolete patch_tolist_cudagraph
Today's TQ3 boot probe surfaced that PR #41434 ([Perf][3/n] Eliminate
GPU<->CPU syncs in attention impls — MERGED) supersedes our local
patch_tolist_cudagraph.py — both anchors no longer match because
upstream replaced them with CPU-resident metadata fields (seq_lens_cpu,
query_start_loc_cpu) and a GPU `.tolist()` fallback when those aren't
populated. Patch ran clean, found no anchors, exited as no-op.

Related: PR #39931 (TurboQuant: support hybrid models and uniform
quantization — MERGED) is what unblocks Qwen3-Next + TurboQuant natively,
with no Genesis dependency. Confirmed by booting Qwen
int8-tq3.yml against vLLM nightly 1acd67a7: 1,416,904 tokens KV pool +
5.41× max concurrency @ 262K, no patches needed.

Bumps (19 Qwen composes): nightly-01d4d1ad (2026-05-04) → 1acd67a7 (2026-05-08).
Matches what Gemma already pinned; converges the stack on a single
nightly. Bump is small (4 days) but includes:
  - #39931 TurboQuant + hybrid attention
  - #41434 GPU<->CPU sync elimination ([Perf][3/n])
  - #40092 TurboQuant FA3/FA4 prefill paths
  - #41991 Gemma 4 tool parser bugfix (partial — #42006 still open)
  - misc other fixes

Excluded from bump (need separate validation):
  - models/gemma-4-31b/vllm/compose/dual/dflash.yml
  - models/gemma-4-31b/vllm/compose/dual/dflash-int8.yml
    Both pinned at nightly-e47c98ef (2026-05-06) with heavy vendored overlays
    (vllm-gemma4-dflash/, vllm-pr40391-rebased/) keyed to specific upstream
    file paths. Bumping their base nightly without re-validating the overlay
    anchors risks silent breakage. To bump, A/B against current pin first.

patch_tolist_cudagraph removed from:
  - models/qwen3.6-27b/vllm/compose/dual/int8-tq3.yml (mount + entrypoint)
  - models/qwen3.6-27b/vllm/compose/dual/turbo.yml (mount + entrypoint)
  - models/qwen3.6-27b/vllm/compose/dual/nvlink-turbo.yml (mount + entrypoint)
  - models/qwen3.6-27b/vllm/compose/single/long-text.yml (mount + entrypoint)
  - models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml (mount + entrypoint)
  - models/gemma-4-31b/vllm/compose/dual/int8-tq3.yml (mount + wrapper entrypoint)
  - models/qwen3.6-27b/vllm/patches/patch_tolist_cudagraph.py (file deleted)
  - models/gemma-4-31b/vllm/patches/patch_tolist_cudagraph.py (file deleted)

Some composes still have patch_tolist_cudagraph mentioned in comments —
those are historical context about Genesis P78 / PN34 superseding it,
not active code, so left in place.

Gemma int8-tq3.yml: rewrote header as tombstone documenting the upstream
blocker chain. TurboQuant on Gemma 4 needs vllm#40108 (sliding-window
TQ) + vllm#41497 + vllm#40308 to merge. Tracker: vllm-issue#41403
(5-gate blocker stack). Until then, use dual/int8.yml for Gemma.

New today (BF16 baseline composes):
  - models/qwen3.6-27b/vllm/compose/dual/bf16.yml (port 8012, 200K + n=4 + max-num-seqs=1)
  - models/gemma-4-31b/vllm/compose/dual/bf16.yml (port 8033, 200K + n=4 + max-num-seqs=1)

Both default to 200K context with max-num-seqs=1 — that's BF16's natural
ceiling at TP=2 on 24 GiB cards.

Verified state after this commit:
  - 28 composes on nightly-1acd67a7 (canonical)
  - 2 composes on nightly-e47c98ef (DFlash family, deferred)
  - 0 active patch_tolist_cudagraph mounts or entrypoint invocations
2026-05-11 01:56:32 +00:00
noonghunna be7f9aa143 feat(rebench-report): close 9 gaps — TL;DR + rig + timings + reproducer + delta + discuss variant
scripts/rebench-report.py:
  - Date extraction via regex (handles tag-with-or-without YYYY-MM-DD suffix),
    falls back to dir mtime.
  - Quality latency parser now reads nested latency.{p50,p95,mean} (was looking
    for non-existent flat fields).
  - New 'Paste-ready compose Quality: schema line' subsection grepped from
    quality-full.log.
  - Rig info section in Meta: hostname, GPU models, per-card power cap,
    parsed from rig.txt.
  - Auto-computed TL;DR section (3-6 bullets covering TPS, KV pool,
    verify-stress, quality, soak, aider).
  - Phase timings table from timings.json with grand total.
  - Reproducer-commands section with paste-ready bash.
  - --compare-to <tag-dir> flag for numerical deltas across runs; writes
    _internal.json sidecar so future runs can diff against this one.
  - REPORT-discuss.md trimmed variant for GitHub Discussion comments.
  - Aider per_language already fixed in prior commit; quant-string fix
    in discuss header (was rendering 'INT16' from --dtype float16 suffix).

scripts/rebench-full.sh:
  - Writes rig.txt (hostname + nvidia-smi -L + power cap) at preamble.
  - Writes timings.json incrementally as each phase completes, via
    record_timing helper called from run_step.
  - quality-full.log is the source for the Quality: one-liner so the
    rebench-report grep just works.

Validated against the synthesized Qwen INT8 PTH n=4 tag dir — all 9
sections render correctly.
2026-05-11 00:45:19 +00:00
noonghunna 7c4b310cca fix(rebench-report): parse aider upstream_per_exercise as dict (not list)
aider's benchmark.py emits per-exercise results as a dict keyed by
'<lang>/<exercise>' path, nested under verifier_trace.trace. Earlier
parser assumed a flat list of {language, passed} dicts and so produced
no per-language tally.

Also handle the dual nesting (verifier_trace.trace vs flat) and use the
path prefix as the language key. Fall back to dict-aware tally when
top-level pass_rate/passed_count/total_count are None.

Validated against today's Qwen INT8 PTH n=4 aider rerun (quality-2026-05-11
T00-12-30.json) — produces:
  Total: 19/30
  cpp: 3/5  go: 4/5  java: 4/5  javascript: 4/5  python: 3/5  rust: 1/5
2026-05-11 00:35:13 +00:00
noonghunna 18355f41f6 feat(rebench): add REPORT.md synthesizer + container/boot/GPU captures
scripts/rebench-report.py — parses raw artifacts (bench.log, verify-stress.log,
quality-full.json, soak.log, aider-polyglot.json, container-config.json,
vllm-boot.log) and renders a single REPORT.md at the top of the tag dir.

Sections: meta · config · performance · concurrency+VRAM · verify-stress
matrix · quality 8-pack table · failure examples · soak KPIs · aider per-language.

scripts/rebench-full.sh additions:
  - container-config.json (docker inspect snapshot)
  - vllm-boot.log (KV pool size + max concurrency + MTP detect, trimmed)
  - gpu-state-{start,end}.log (nvidia-smi snapshots)
  - final REPORT.md synthesis phase

Standalone re-render: python3 scripts/rebench-report.py results/rebench/<tag>/

Dry-run validated against today's running Qwen INT8 PTH compose — meta/config/
concurrency sections render cleanly with patches-mounted manifest + Genesis
detection working.
2026-05-11 00:28:30 +00:00
noonghunna 340689427f feat(rebench): halve default soak to 10 sessions × 5 turns (~15-20 min)
Total per-leg runtime drops ~2-2.5 hr → ~1.75-2 hr. Soak still gives 50 turns
which catches the silent-empty / Cliff 2b patterns just as reliably as 100
turns (those failures cluster in early sessions when they happen). Bump back
to SOAK_SESSIONS=20 for the canonical stability matrix when validating new
compose paths.
2026-05-11 00:16:52 +00:00
noonghunna 94a2522417 feat(rebench): one-shot canonical 5-step bench orchestrator
scripts/rebench-full.sh — runs bench + verify-stress + quality-full + soak +
aider-polyglot in canonical order, collects artifacts to results/rebench/<tag>/.

Fixes the six recurring mistakes from manual matrix runs:
  1. Wrong cwd — script resolves ROOT_DIR via $BASH_SOURCE and cd's upfront
  2. Forgot --save-json — routes aider through quality-test.sh (auto-saves),
     snapshots JSON into per-tag dir
  3. Forgot MODEL= override — auto-detects served-model-name from /v1/models
  4. Forgot BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 — auto-exports when URL
     matches localhost-style, idempotent re user override
  5. Wrong port — preflight_autodetect_endpoint resolves it
  6. No idempotent resume — --resume skips steps whose artifact files exist

Usage:
  bash scripts/rebench-full.sh                    # auto-tag from MODEL
  bash scripts/rebench-full.sh --tag qwen-int8    # explicit tag
  bash scripts/rebench-full.sh --skip soak,aider  # skip phases (CSV)
  bash scripts/rebench-full.sh --resume           # resume after interrupt

Run twice on different models to assemble a matched-config head-to-head.
2026-05-11 00:15:20 +00:00
noonghunna 755e5199ff bench(head-to-head): matched-config rebench + Qwen INT8 PTH KV compose
Adds qwen3.6-27b/vllm/compose/dual/int8.yml (new) — same vLLM nightly
1acd67a7 as gemma-int8, same int8_per_token_head KV class, MTP n=4
parameterized via SPEC_N_MAX env var. First validation of INT8 PTH KV
on Qwen3-Next DeltaNet hybrid attention on our stack.

Updates gemma-int8.yml to parameterize num_speculative_tokens via SPEC_N_MAX.

BENCHMARKS.md head-to-head section: matched-config addendum showing
Gemma's per-stream advantage shrinks to ~10% decode / ~45% TTFT (vs
+60% in the morning's mismatched run), and Qwen has +30% more KV pool /
concurrency. Original mismatched table preserved.

Side-effect: Qwen int8.yml is a candidate shipping path — +27/+33% TPS
over canonical dual.yml at -2.3 GiB/card VRAM.
2026-05-10 23:09:37 +00:00
noonghunna edda3b3eca docs(benchmarks): Qwen 3.6 27B vs Gemma 4 31B head-to-head on dual 3090
Adds a consolidated head-to-head section under BENCHMARKS.md aggregating
today's TPS bench + 8-pack quality-test --full + aider-polyglot-30 results
for both dual.yml configs on @noonghunna's 2× 3090 rig.

- Config & pin delta table (vLLM nightly 01d4d1ad vs 1acd67a7, fp8 vs BF16
  KV, built-in MTP n=3 vs external 0.5B drafter n=4, 262K vs 32K ctx).
- TPS: Gemma +57% narr / +60% code, ~2× faster TTFT, 0.8 GiB less VRAM/card.
- Quality (150 scenarios): tied at 67% / 68%; per-pack split shows Gemma
  wins agentic + bug-fix (hermes +10 pp, bugfind +13 pp); Qwen wins
  polyglot code editing (aider-polyglot +10 pp, Java the biggest swing).
- 180-scenario combined: Qwen 120/180 (66.7%), Gemma 119/180 (66.1%).
- Reproducer commands inline.
2026-05-10 22:05:52 +00:00
noonghunna a258e496bf chore(cliff): skip auto-regen bot commits in changelog parser
Filter out 'chore(changelog): regenerate for vX.Y.Z [skip ci]' commits
that the release.yml workflow auto-creates. They're pure machine
traffic and only added noise to release notes.
2026-05-10 21:16:35 +00:00
github-actions[bot] 9ec08cf7eb chore(changelog): regenerate for v0.3.3 [skip ci] 2026-05-10 21:14:50 +00:00
noonghunna eeb946b0a7 chore(changelog): subject-only rendering (drop commit body verbosity)
Release / release (push) Failing after 51s
CHANGELOG.md and GitHub Release notes now render commit subjects only.
Full commit message body (why/how/validation data) stays in 'git log',
one click away via the SHA link. Keeps both surfaces skim-readable.

v0.3.2 release pages were 100+ lines per commit; this brings them down
to ~1 line per commit.
v0.3.3
2026-05-10 21:14:34 +00:00
github-actions[bot] 507a36f7f7 chore(changelog): regenerate for v0.3.2 [skip ci] 2026-05-10 21:08:51 +00:00
noonghunna 255c743dff chore: trigger v0.3.2 release workflow (GitHub deduped previous tag push)
Release / release (push) Failing after 48s
The v0.3.2 tag was originally pushed at commit 64b0474 but GitHub didn't
emit a CreateEvent (likely dedup after delete + re-push to same SHA), so
the release.yml workflow never fired. Empty commit gives the tag a fresh
SHA that GitHub will process cleanly.
v0.3.2
2026-05-10 21:08:33 +00:00
noonghunna 64b0474a62 chore(changelog): automate CHANGELOG + release notes from commits via cliff (Option A)
CHANGELOG.md is now auto-generated from commit messages by git-cliff in
the release workflow. Hand-edits below the static header will be wiped on
the next tag.

Workflow (`.github/workflows/release.yml`):
  - On tag push (`v[0-9]+.[0-9]+.[0-9]+`):
    1. Render GitHub Release body: `git-cliff --latest --strip header`
       → just the per-version section, no SemVer preamble repeat
    2. Regenerate full CHANGELOG.md: `git-cliff` (default = all tags)
       → preserves header + all historical sections
    3. Commit CHANGELOG.md back to master with `[skip ci]` marker
    4. Publish GitHub Release with the latest-only body

Template (`cliff.toml`):
  - `[changelog].header` now holds the SemVer preamble + CalVer→SemVer
    mapping table (preserved across regens; stripped from GitHub Release
    bodies via `--strip header`).
  - `body` template now renders the **full commit message** (subject as
    bold bullet, body indented below) instead of just the first line.
    Rich narrative I write in commit message bodies (tables, validation
    numbers, before/after diffs) now flows into both CHANGELOG.md and the
    GitHub Release page from the same source.
  - Per-release Pin/Diff footer guarded with `{% if version %}` so the
    Unreleased section doesn't emit empty links.

CHANGELOG.md replaced with the auto-gen output. Past hand-written tables
and phase breakdowns are replaced by the corresponding commit messages
(those were already rich for commits that mattered — v0.3.1 soak-helper
fix has its Before/After table in the commit body and renders fine).

Going forward: just write rich commit messages and tag. Both surfaces
update automatically. No hand-edit of CHANGELOG.md required.
2026-05-10 20:53:51 +00:00
noonghunna 83bf73d3ec feat(quality-test): auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for localhost URLs
When the user runs quality-test.sh with a localhost-style URL
(default `http://localhost:8020`, or any `localhost`/`127.x`/`[::1]`
variant), auto-export `BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1` so
benchlocal-cli rewrites the hermes-agent's outbound model endpoint
from `localhost:<port>` to `host.docker.internal:<port>` inside the
Docker sandbox container.

Without this, the hermes-agent inside the sandbox can't reach the
host's vLLM (localhost resolves to the container itself) and every
scenario fails with `"API call failed after 3 retries: Connection
error."` — produced spurious 0/20 grades on this rig prior to the
benchlocal-cli runner.py 9c1566f fix.

Skips the auto-set when:
  - User already set the env var (explicit override)
  - URL points at a non-loopback host (real LAN IP, k8s service name,
    host.docker.internal already) — no rewrite needed

Emits a stderr breadcrumb when the auto-set fires so users can see
what changed.
2026-05-10 20:29:49 +00:00
noonghunna 9db8b2603c docs(changelog): v0.3.1 entry for soak-helper delta.reasoning capture
Release / release (push) Failing after 46s
Documents the silent-empty turn-5 root cause (vLLM nightly field-name
shift to delta.reasoning) + validation soak results.
v0.3.1
2026-05-10 19:16:44 +00:00
noonghunna 88eb67aa18 fix(soak-helper): capture delta.reasoning alongside delta.reasoning_content
vLLM nightly (0.20.2rc1.dev9+) emits the qwen3 reasoning parser's output
under `delta.reasoning` (legacy field name), not `delta.reasoning_content`
that soak-helper.py was watching. Result: for any thinking-on response
whose `<think>` block doesn't close within `max_tokens`, soak-helper saw
zero deltas → fell back to the "couldn't measure" path → reported
`ttft_ms == t_ms` and `decode_tps = 0.0`. The model was generating
correctly; the harness just couldn't see the wire output.

Repro request (JDWarner's #107 turn 5): math problem with
`max_tokens=2000` + `chat_template_kwargs.enable_thinking=true`.

Before patch:
  status=200  t_ms=22709  ttft_ms=22709  decode_tps=0.0
  completion_tokens=2000  content=""  reasoning_content=""

After patch (same request, same compose, same model):
  status=200  t_ms=22709  ttft_ms=234  decode_tps=88.985
  completion_tokens=2000  content=""  reasoning_content="Here's a thinking
  process:\n\n1. **Understand the User's Problem:**\n..."  (3959 chars)

Validation soak (fresh-mode, 20 sessions × 5 turns = 100 turns, qwen3.6-27b
dual.yml):
  verdict        PASS
  silent_empty   0 / 100 (0.0%)   ← was ~3-5/40 baseline
  p50_decode_tps 90.22
  p95_ttft_ms    1389
  errors         0
  max_growth     0 MiB / 200

Closes the cross-rig "silent-empty turn-5" pattern parked behind the
Cliff 2b investigation — it was a harness measurement bug, not a model
or rig issue.
2026-05-10 19:15:57 +00:00
noonghunna 7080f1f89b release: SemVer adoption + v0.3.0 changelog entry
Release / release (push) Failing after 1m25s
club-3090 is software (docker composes + system scripts + patch bundles
that downstream rigs run as-is), not just rolling recipes. Switch from
CalVer to SemVer from v0.3.0 onward; past CalVer tags (v2026.05.09,
v2026.05.10) preserved for history.

CHANGELOG.md: convention note + retroactive CalVer→SemVer mapping +
new v0.3.0 (2026-05-10) entry covering 8 commits since v2026.05.10:
- Qwen 3.6 27B thinking OFF default across all 21 composes
- MODEL_DIR UX overhaul (closes #116) — interactive setup prompt,
  $MODEL_DIR placeholder everywhere
- power-cap-sweep --include-commit (closes #112)
- BENCHMARKS aider-polyglot-30 row

cliff.toml: drop "snapshot of the rolling stack — not a versioned
API" framing; add SemVer note. Tag pattern v[0-9]+.[0-9]+.[0-9]+
already matches both CalVer and SemVer, so the cliff release
workflow needs no changes.
v0.3.0
2026-05-10 18:29:06 +00:00
noonghunnaandClaude Opus 4.7 e08988e614 docs(benchmarks): aider-polyglot-30 — Qwen 27B 20/30 (66.7%) > Gemma 4 31B 17/30 (56.7%)
New "Quality benches — Aider Polyglot 30" section captures pass-rate /
agentic-coding signal alongside the existing TPS rows. First two rows:

- Qwen 3.6 27B (AutoRound INT4) on dual.yml: 20/30 = 66.7%, 19 min wall
- Gemma 4 31B (Intel AutoRound INT4) on dual.yml: 17/30 = 56.7%, 19 min wall

Both run on 2× 3090 PCIe, 230 W cap, threads=2. Qwen edges Gemma by +10pp
despite being smaller; java is the biggest swing (Qwen 4/5 vs Gemma 1/5).

Critical caveat documented: Qwen with thinking ON is unusable for agentic
benches on this hardware — hits the 1500s subprocess cap before any
exercise completes. The new --default-chat-template-kwargs flag in our
vLLM Qwen composes (commit 534d29f) sets enable_thinking=false by default.

Aider-polyglot run via benchlocal-cli's aider-polyglot-30 pack. Cross-rig
re-run path documented in the section.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 18:04:58 +00:00
noonghunnaandClaude Opus 4.7 7d91ac75e0 feat(power-cap-sweep): --include-commit flag stamps club-3090 git SHA in report header (closes #112)
@laurimyllari noted in disc #62 that with the project moving fast,
including the club-3090 git commit in sweep output helps correlate
cross-rig sweeps to the script revision they were run against. He
stamped `aa99173` manually; the script should do it for us.

New flag:

  --include-commit   Stamp the club-3090 git commit (short SHA) in the
                     report header next to the date. Off by default.

Implementation:
- Captures `git -C "$REPO_ROOT" rev-parse --short HEAD` once at header-build
  time (REPO_ROOT was already known to the script).
- Injects into the report header next to **Date:**, e.g.:

    **Date:** 2026-05-10T17:55:00Z &nbsp; **club-3090 commit:** `534d29f`

- Suppress (don't stamp "n/a") when run from a non-clone or git is
  unreachable. Closes the curl-pipe-from-docs UX hole — `curl ... | bash`
  users don't have a clone, so the field just disappears rather than
  showing a confusing "n/a".

Off by default per the issue rationale: surprise stamping confuses
contributors running from documentation snippets.

Closes #112.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:57:20 +00:00
noonghunnaandClaude Opus 4.7 534d29f1b1 fix(qwen3.6-27b): use --default-chat-template-kwargs (not --chat-template-kwargs)
Follow-up to 29d17ed which used the wrong vLLM flag name (`--chat-template-kwargs`),
causing boot failure: "vllm: error: unrecognized arguments: --chat-template-kwargs
{"enable_thinking": false}".

vLLM's actual flag for setting server-side default chat template kwargs is
`--default-chat-template-kwargs` (with the `default-` prefix). Confirmed by
- vLLM nightly source: vllm/engine/arg_utils.py defines
  `default_chat_template_kwargs: dict[str, Any] | None = None` with
  json.loads parsing.
- vLLM PR #37739 ("Fix default_chat_template_kwargs handling in Responses API")
  references it as already available in the shared render stack.

Behavior unchanged: thinking OFF by default for all 17 vLLM Qwen 27B composes
(plus bounded-thinking unaffected — it intentionally keeps thinking ON).
Per-request override still works via OpenAI extra_body:
  {"chat_template_kwargs": {"enable_thinking": true}}

Verified: vllm-qwen36-27b-dual now boots cleanly with the corrected flag.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:44:50 +00:00
noonghunnaandClaude Opus 4.7 29d17ed82d feat(qwen3.6-27b): thinking OFF by default across all 21 composes
Previously, all 19 vLLM Qwen composes used --reasoning-parser qwen3 (parses
<think>...</think> output blocks) but did NOT explicitly disable the
thinking template. That meant Qwen3 default — thinking ON — applied
across the board. This burns hidden token budget on internal CoT for
every request, hurts latency, and creates a "Qwen looks much slower
than Gemma" gap on agentic benchmarks (which is exactly what we just
hit on aider-polyglot — Qwen exceeded the 1500s timeout, Gemma
finished in 19 min).

Change:
- All vLLM composes (18 of them) now pass `--chat-template-kwargs
  '{"enable_thinking": false}'` after `--reasoning-parser qwen3`.
- llama.cpp single/docker-compose.yml: DISABLE_THINKING default flipped
  0 → 1 (thinking now OFF by default; opt back in via DISABLE_THINKING=0).
- llama.cpp single/concurrent.yml: gained `--chat-template-kwargs` flag
  with default `{"enable_thinking":false}` (overridable via
  CHAT_TEMPLATE_KWARGS env).

NOT changed:
- bounded-thinking.yml — that's the structured-CoT compose where thinking
  IS the feature. Reverted my initial blanket change for that one.
- qwopus-bf16mtp.yml — already had enable_thinking=false (preview compose).

Users who want thinking ON can:
- For vLLM: pass `chat_template_kwargs: {enable_thinking: true}` in the
  per-request body (works fine).
- For llama.cpp: set `DISABLE_THINKING=0` in compose/.env.

Aligns with Gemma 4's "thinking off by default" (it ships that way upstream)
and removes the Qwen vs Gemma framework-bench skew.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:25:45 +00:00
noonghunnaandClaude Opus 4.7 3909c2d6b8 feat(setup): interactive MODEL_DIR prompt for fresh TTY users
Previously setup.sh silently defaulted MODEL_DIR to <repo>/models-cache,
which meant fresh users got ~14-21 GB of model weights downloaded INTO
their git tree without realizing it. The relative path is also wrong
when run from anywhere except the repo root.

New 4-step resolution order in setup.sh:
  1. MODEL_DIR exported in calling shell  → use as-is (unchanged)
  2. .env at repo root sets MODEL_DIR     → source it (NEW)
  3. Interactive prompt (only on TTY)     → ask user (NEW)
  4. Silent fallback to <repo>/models-cache (unchanged for non-TTY)

The interactive prompt only fires when:
  - MODEL_DIR is not in the calling env, AND
  - .env doesn't already set it, AND
  - both stdin AND stdout are TTYs (CI / scripted runs unaffected)

Three options offered:
  1. <repo>/models-cache   (the old silent default — kept as option)
  2. $HOME/models           (sensible cross-rig default)
  3. custom absolute path

After picking, optionally persists the choice to .env (gitignored) so
re-runs skip the prompt. Existing .env files are updated in-place if
they already set MODEL_DIR; appended-to otherwise.

Closes the UX hole RobH589 hit in club-3090#116 — the relative
../../../../../models-cache default that was resolving wrong when
not run from the compose dir.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:16:55 +00:00
noonghunnaandClaude Opus 4.7 cc3a717524 docs(recipes): use \$MODEL_DIR placeholder + sensible cross-rig default
Two recipe scripts (single-card-default.sh, single-card-max-ctx.sh) had
the dev rig path /mnt/models/huggingface/... baked into the MODEL_PATH
default + the comment-block instructions for downloading the GGUF.

Updated:
- Comments now show \${MODEL_DIR}/qwen3.6-27b-gguf/... as the placeholder
  (matches what the just-fixed README + LLAMA_CPP.md docs say).
- MODEL_PATH default changed from /mnt/models/huggingface/... to
  \${MODEL_DIR:-\$HOME/models}/qwen3.6-27b-gguf/... — falls back to
  ~/models/ if MODEL_DIR isn't set, which is a more reasonable default
  for cross-rig users than our /mnt/models/huggingface/ path.
- file-exists check at line 25 still fails loudly with the resolved path
  if neither MODEL_DIR nor MODEL_PATH is set correctly.

Follow-up to fbf3431 (de-bind \$MODEL_DIR from rig path in docs).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:14:21 +00:00
noonghunnaandClaude Opus 4.7 fbf343129c docs: use \$MODEL_DIR placeholder, not the dev rig's /mnt/models/huggingface/
User docs were hardcoding the dev rig path (/mnt/models/huggingface/...) as
if it was canonical. It's not — cross-rig users have models at /data/models,
~/models, /mnt/nvme/llms, etc. Setting MODEL_DIR per their setup is the
intended UX (the compose already supports it via env-var default).

Updates:
- models/qwen3.6-27b/llama-cpp/README.md: download examples now use
  \$MODEL_DIR/qwen3.6-27b-gguf/ instead of /mnt/models/huggingface/...
- models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml: header
  comment uses \$MODEL_DIR/qwen3.6-27b-gguf/ for download examples + says
  "MODEL_DIR=/your/models/dir docker compose up -d" instead of our path.
- docs/engines/LLAMA_CPP.md: same treatment + cleaned up Qwen3.5 + DFlash
  draft path examples to also use \$MODEL_DIR.
- scripts/preflight.sh: hf download hint shows literal \${MODEL_DIR} so user
  knows what to set, plus explicit "set MODEL_DIR first" line. Previously
  echoed the resolved relative path (../../../../models-cache) which lands
  outside the repo if pwd isn't the compose dir.

Caught by RobH589 in club-3090#116 — they hit the path-resolved-to-root-of-drive
case from the relative-path default. Closes the doc UX side; the compose's
env-var override mechanism was already correct.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:13:29 +00:00
noonghunnaandClaude Opus 4.7 cf7f1959fd fix: 4 stale refs missed in 2026-05-10 reorg push (caught by RobH589 #116)
After the GGUF dir move (/mnt/models/gguf/qwen3.6-27b/ → /mnt/models/
huggingface/qwen3.6-27b-gguf/), four refs were not updated and led to
a path-mismatch loop reported in club-3090#116:

- scripts/preflight.sh `hf download` hint pointed at qwen3.6-27b/, but
  the compose default expects qwen3.6-27b-gguf/. Same for the mv hint
  for mmproj relocation, and the in-container mmproj default at line 329.
- models/qwen3.6-27b/llama-cpp/README.md example command still said
  `MODEL_DIR=/mnt/models/gguf` (now /mnt/models/huggingface).
- models/qwen3.6-27b/llama-cpp/compose/single/{docker-compose,concurrent}.yml
  header comment said "Q5_K_XL" but the actual default has been Q3_K_XL
  for a while (this one predates the reorg — just stale doc).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:06:20 +00:00
noonghunnaandClaude Opus 4.7 00366a58d7 reorg: services/ consolidation + gpu-mode under git + ComfyUI + pin tracker + path updates
Release / release (push) Failing after 53s
The dev rig had grown across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`,
`/home/wasif/`) and three repos (single-3090, dual-3090, club-3090). Disk-out
on `/` (97% used) forced a cleanup; rather than just prune, we consolidated
the layout across the whole stack while we were at it. This commit captures
what landed inside this repo.

Services consolidation:
- services/{ollama,openwebui,litellm,qdrant,searxng}/ migrated in from
  /opt/ai/compose/<svc>/ (zero functional change — same docker-compose.yml).
- services/litellm/config.yaml rewritten: explicit routes for current
  primaries (qwen3.6-27b-autoround → :8010, gemma-4-31b-autoround → :8030).
  Removed `* → ollama/*` wildcard.
- services/comfyui/ migrated in (was /opt/ai/compose/comfyui/) — wired into
  gpu-mode with full mutex against vLLM/SGLang.

scripts/gpu-mode.sh under git:
- Was a loose /opt/ai/gpu-mode.sh outside any repo. Symlinked at
  /usr/local/bin/gpu-mode.
- Five Gemma 4 31B modes added: gemma, gemma-dflash, gemma-int8,
  gemma-dflash-int8, gemma-awq.
- One ComfyUI mode (mutex with all LLM serving).
- prune / prune-all subcommands (safe image prune; aggressive variant adds
  build cache --keep-storage 5GB + dangling networks).
- gpu-mode status now shows Docker disk + /var/lib/docker + /tmp sizes.
- compose_at() passes --env-file <repo>/.env so MODEL_DIR resolves
  regardless of which compose dir gpu-mode cd's into. Fixes the recurring
  "MODEL_DIR not set, defaulting to ../../../../../models-cache" warning.
- stderr no longer swallowed by compose_at() (real errors surface).
- Cross-model VRAM mutex: every Qwen mode stop_all_gemma + stop_comfyui
  and vice-versa.

scripts/maintenance/ — new hygiene-tools subdir:
- list-image-pins.sh: engine-agnostic pin auditor. Scans every compose's
  `image:` line, groups by `<repo>:<tag>`, flags pin-drift (multiple tags
  per repo), ranks composes by patch surface.

Pin tracking:
- docs/UPSTREAM.md gains a "Pinned images" section: table of every pinned
  image, why each pin exists, retirement candidate criteria.
- docs/NIGHTLY_BUMP_RUNBOOK.md (new): 7-step procedure for bumping pinned
  engine images (scope → branch → patch survival → boot → verify-full +
  verify-stress → bench delta → land → retire). Engine-specific notes
  for vLLM nightly hashes, llama.cpp digest pinning, SGLang variants.

Path updates from the engine + model dir consolidation:
- /opt/ai/vllm-src/ → /opt/ai/engines/vllm/primary/
  (in setup.sh, INTERNALS.md, several patch READMEs, docs/HARDWARE.md,
   docs/FAQ.md, docs/DUAL_CARD.md, docs/UPSTREAM.md, models/qwen3.6-27b/
   CHANGELOG.md)
- /mnt/models/gguf/qwen3.6-27b/ → /mnt/models/huggingface/qwen3.6-27b-gguf/
  (in models/qwen3.6-27b/llama-cpp/{compose/single/*.yml, recipes/*.sh,
   README.md}, docs/engines/LLAMA_CPP.md)

CHANGELOG.md narrative gap fill:
- 2026-05-10 entry for this reorg.
- 2026-05-09 entry for compose convention formalization (topology
  promoted to dir level, profile schema, Status enum + Caveats, cliff
  CI swap, Discord launch).
- 2026-05-08 entry for Gemma 4 INT8 PTH unblock + 262K validation.
- 2026-05-07 entry for power-cap-sweep campaign + HARDWARE.md cross-rig
  charts + cross-rig benchmark rows.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
v2026.05.10
2026-05-10 16:57:03 +00:00
noonghunnaandClaude Opus 4.7 403b16f303 feat(gemma-4-31b): INT8 PTH KV unblocks 262K + AWQ + DFlash compose family
Lands the Gemma 4 31B work that was in flight last week + Qwopus3.6-27B
preview compose. Captures four distinct serving paths for Gemma 4:

- dual/awq.yml — AWQ-4bit (text-only, simplest path)
- dual/dflash-int8.yml — DFlash + INT8 PTH KV (262K ctx, full pipeline)
- vllm-gemma4-dflash-int8/ — vendored patches stacking DFlash spec-decode
  + INT8 PTH KV across model_executor + v1/spec_decode + v1/attention +
  v1/worker (~13 patched files; PR #42102 + #40391-rebased + tool-parser
  fixes #42006 + #41991 stacked)
- vllm-gemma4-fp8-ampere/ — earlier Phase 2 attempt before INT8 PTH
  reframe (kept for forensics); Ampere has no native FP8 tensor cores
- vllm-perheadkv-hybridpage-fix/ — hybrid-page bug fix surfaced during
  Phase 3
- vllm-pr40391-perheadkv/ — PR #40391 vendored at the tree level
  (separate from the rebased variant under refs/jianc99-dflash-gemma4)

Plus:
- models/qwen3.6-27b/vllm/compose/dual/qwopus-bf16mtp.yml — preview compose
  for Carnice AutoRound Recipe D output (port 8071, NOT production —
  see club-3090-todo.md for known gaps + cheap A/Bs).
- AGENTS.md — codify compose naming + profile-schema + experimental-compose
  conventions that the new files follow.
- docs/QUALITY_TEST.md — runbook for `quality-test.sh` + the benchlocal-cli
  packs it wraps.

CHANGELOG.md narrative entries for these are added separately.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 16:56:17 +00:00
noonghunnaandClaude Opus 4.7 6e12700f9a docs: WSL2 budget formula + Cliff 3 (DeltaNet SSM-state non-cacheable)
Two findings from @easel's deep cross-rig writeup on issue #102 (5090
Laptop WSL2):

1. **HARDWARE.md** — new "GPU memory budget on WSL2" subsection. WSL2
   container CUDA-context consumes ~1.31 GiB before vLLM profiler runs.
   `gpu_memory_utilization=0.95` crashes at boot; 0.944 works. Validated
   on 2 machines. Formula: `(vram_total - 1.31) / vram_total`.
   Recommendation: `GPU_MEMORY_UTILIZATION=0.94` in `.env` on WSL2.

2. **CLIFFS.md** — new section "Cliff 3 — DeltaNet SSM state is not
   prefix-cacheable (the prefill cliff)". This is a structural finding
   that explains a class of failures we'd been describing without
   naming. @easel's warm-cache run (68.7% KV-block hit, turn 10 at
   35.6K tokens took 577s — 2.3× the cold-start 254s) is the smoking
   gun: prefix cache helps attention but DeltaNet's recurrent state
   `h_t = f(h_{t-1}, x_t)` must be recomputed from scratch every turn.
   PN32 fixes OOM stability; nothing fixes the O(n) prefill scaling
   because the architecture itself is sequential.

   Practical ceiling on single-card vLLM (any 24 GB Qwen3-Next config):
   sub-30s TTFT only below 5K accumulated tokens. 22-35K = 3-4 min/turn.
   ~74K = 10+ min client timeout.

   Elevates "for single-card agentic Qwen3-Next, use llama.cpp" from
   implied to explicit in the docs. Dual-card extends the envelope to
   25-30K accumulated; deep sessions (50K+) still need llama.cpp on
   either topology.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 01:10:05 +00:00
Erik LaBiancaandClaude Sonnet 4.6 af9fb0cd29 compose: parametrize VLLM_ENFORCE_EAGER, KV_CACHE_DTYPE, P40/P82/PN54 across all variants (#110)
All defaults unchanged — existing users see identical behavior out of the box.
New opt-ins let RTX 5090 / WSL2 / high-L2 rigs tune without forking files.

Changes across 18 compose files + .env.example:

VLLM_ENFORCE_EAGER — hook added to 7 files that lacked it:
  carnice-bf16mtp, dual-dflash, dual-dflash-noviz, dual-nvlink-dflash,
  dual-nvlink-dflash-noviz, dual4-dflash, minimal
  (bounded-thinking, dual, dual-turbo, long-*, tools-text, docker-compose.yml
   already had the hook)

KV_CACHE_DTYPE — parameterised in all 13 variants that hardcoded it:
  turboquant_3bit_nc default: bounded-thinking, docker-compose.yml,
    long-text, long-text-no-mtp, long-vision, dual-turbo, dual-nvlink-turbo
  fp8_e5m2 default: carnice-bf16mtp, dual, dual-nvlink, dual4, minimal, tools-text

GENESIS_ENABLE_P40 + GENESIS_ENABLE_PN54 — opt-in stanzas added to all 8
  Genesis-using variants: bounded-thinking, docker-compose.yml, dual-turbo,
  dual-nvlink-turbo, long-text, long-text-no-mtp, long-vision, tools-text

GENESIS_ENABLE_P82 — promoted from hardcoded 0 → ${GENESIS_ENABLE_P82:-0}
  in 6 spec-decode variants: bounded-thinking, dual-turbo, dual-nvlink-turbo,
  long-text, long-text-no-mtp, long-vision

.env.example additions:
  - Docs for VLLM_ENFORCE_EAGER, KV_CACHE_DTYPE, P40, P82, PN54
  - Validated RTX 5090 Laptop + WSL2 profile block (issue #102):
    PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False,max_split_size_mb:512
    GPU_MEMORY_UTILIZATION=0.94, VLLM_ENFORCE_EAGER=1, GENESIS_ENABLE_P40=1,
    GENESIS_ENABLE_P82=1, SOAK_TIMEOUT_S=3600

Co-authored-by: Claude Sonnet 4.6 <[email protected]>
2026-05-10 05:51:49 +05:00
Erik LaBiancaandClaude Sonnet 4.6 73c31848ea fix: BIND_HOST opt-in + localhost script fixes (#109)
Three related fixes for running benchmarks without IDE agent interference:

1. All 18 vLLM compose files: port binding is now
   ${BIND_HOST:-0.0.0.0}:${PORT:-<n>}:8000
   Setting BIND_HOST=127.0.0.1 in .env restricts the API to localhost,
   preventing IDE agents (Cline, Cursor) from competing for the
   max-num-seqs=1 slot and causing verify-stress HTTP 000 failures.

2. scripts/preflight.sh: port auto-detection regex now matches
   127.0.0.1:<port>->8000/tcp in addition to 0.0.0.0: and [::]:
   Previously all verify-*/bench scripts silently produced no output
   when BIND_HOST=127.0.0.1 was set.

3. scripts/report.sh: SOAK_TIMEOUT_S is now forwarded to soak-test.sh
   from the shell environment. Previously the variable was read from
   compose .env (docker-compose only) and silently ignored by the
   script, always using the 1800s default regardless of what was set.

Docs: .env.example gains BIND_HOST and SOAK_TIMEOUT_S entries.

Co-authored-by: Claude Sonnet 4.6 <[email protected]>
2026-05-10 05:50:08 +05:00
noonghunnaandClaude Opus 4.7 c298b60f76 encourage-soak: template dropdown + script ergonomics + report reminder + Notes convention
Four small fixes addressing low soak-test compliance in cross-rig bench
contributions. Audit of 5 recent #113/#107/#102/#104/#93 showed soak data
IS being run but it's hidden in the main report and the dedicated
template field comes out empty (template said "leave blank if you ran
--full"). Older BENCHMARKS rows often omit soak verdict entirely.

1. **`.github/ISSUE_TEMPLATE/numbers-from-your-rig.yml`**: replace optional
   "soak summary" textarea with a required dropdown listing PASS / borderline /
   FAIL / Skipped+reason / Not-yet-run. Verdict is now grep-able even when
   the data is buried in the main report textarea.

2. **`scripts/soak-test.sh`**: add `--continuous` / `--quick` / `--fresh`
   flags + `--help` + cleaner usage docs. Was 5 env vars to invoke
   (`SOAK_MODE=continuous SOAK_SESSIONS=5 SOAK_TURNS=5 CONTAINER=... ENDPOINT=...`);
   now `bash scripts/soak-test.sh --continuous` does the same with
   auto-detect (existing logic preserved + exposed). Env vars still work
   for back-compat.

3. **`scripts/report.sh`**: when `--bench` (or partial) ran without
   `--soak`/`--full`, append a "⚠ Soak: not included" reminder block to
   the report so contributors know what's missing before pasting into
   the issue template.

4. **`BENCHMARKS.md`**: Notes-column convention — every row should start
   with explicit `Soak: ✓ PASS` / `⚠ borderline` / `✗ FAIL` / `—` so
   readers can grep at a glance. Updated 2 recent rows (ygafarov #113,
   JDWarner #107) to use the convention. Older rows backfill as the
   convention spreads.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 00:44:19 +00:00
noonghunnaandClaude Opus 4.7 a589058761 BENCHMARKS: add @ygafarov Strix-Halo + oculink-eGPU x4-PCIe single-3090 row (#113)
First Strix-Halo-miniPC + oculink-eGPU class on the matrix. Single 3090
over PCIe x4 (oculink) on AMD Ryzen AI MAX+ 395 / 124 GB RAM / CachyOS /
290W cap. Result: 68.86 narr / 91.70 code TPS via vllm/default + TQ3 at
48K — clean MTP AL 3.31 (77% accept), CV 1.6%/2.7%.

Soak FAIL is borderline (240 MiB > 200 MiB threshold, 3 turns >30s) but
100% TPS retention + 0 errors + 0 silent-empty suggests x4-PCIe accretion
+ bus-latency under prefill, not Cliff 2b. Worth flagging as a possible
"eGPU bus class" threshold allowance for soak-test.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 00:32:32 +00:00