11 Commits
Author SHA1 Message Date
noonghunna 9db8b2603c docs(changelog): v0.3.1 entry for soak-helper delta.reasoning capture
Release / release (push) Failing after 46s
Documents the silent-empty turn-5 root cause (vLLM nightly field-name
shift to delta.reasoning) + validation soak results.
2026-05-10 19:16:44 +00:00
noonghunna 88eb67aa18 fix(soak-helper): capture delta.reasoning alongside delta.reasoning_content
vLLM nightly (0.20.2rc1.dev9+) emits the qwen3 reasoning parser's output
under `delta.reasoning` (legacy field name), not `delta.reasoning_content`
that soak-helper.py was watching. Result: for any thinking-on response
whose `<think>` block doesn't close within `max_tokens`, soak-helper saw
zero deltas → fell back to the "couldn't measure" path → reported
`ttft_ms == t_ms` and `decode_tps = 0.0`. The model was generating
correctly; the harness just couldn't see the wire output.

Repro request (JDWarner's #107 turn 5): math problem with
`max_tokens=2000` + `chat_template_kwargs.enable_thinking=true`.

Before patch:
  status=200  t_ms=22709  ttft_ms=22709  decode_tps=0.0
  completion_tokens=2000  content=""  reasoning_content=""

After patch (same request, same compose, same model):
  status=200  t_ms=22709  ttft_ms=234  decode_tps=88.985
  completion_tokens=2000  content=""  reasoning_content="Here's a thinking
  process:\n\n1. **Understand the User's Problem:**\n..."  (3959 chars)

Validation soak (fresh-mode, 20 sessions × 5 turns = 100 turns, qwen3.6-27b
dual.yml):
  verdict        PASS
  silent_empty   0 / 100 (0.0%)   ← was ~3-5/40 baseline
  p50_decode_tps 90.22
  p95_ttft_ms    1389
  errors         0
  max_growth     0 MiB / 200

Closes the cross-rig "silent-empty turn-5" pattern parked behind the
Cliff 2b investigation — it was a harness measurement bug, not a model
or rig issue.
2026-05-10 19:15:57 +00:00
noonghunna 7080f1f89b release: SemVer adoption + v0.3.0 changelog entry
Release / release (push) Failing after 1m25s
club-3090 is software (docker composes + system scripts + patch bundles
that downstream rigs run as-is), not just rolling recipes. Switch from
CalVer to SemVer from v0.3.0 onward; past CalVer tags (v2026.05.09,
v2026.05.10) preserved for history.

CHANGELOG.md: convention note + retroactive CalVer→SemVer mapping +
new v0.3.0 (2026-05-10) entry covering 8 commits since v2026.05.10:
- Qwen 3.6 27B thinking OFF default across all 21 composes
- MODEL_DIR UX overhaul (closes #116) — interactive setup prompt,
  $MODEL_DIR placeholder everywhere
- power-cap-sweep --include-commit (closes #112)
- BENCHMARKS aider-polyglot-30 row

cliff.toml: drop "snapshot of the rolling stack — not a versioned
API" framing; add SemVer note. Tag pattern v[0-9]+.[0-9]+.[0-9]+
already matches both CalVer and SemVer, so the cliff release
workflow needs no changes.
2026-05-10 18:29:06 +00:00
noonghunnaandClaude Opus 4.7 e08988e614 docs(benchmarks): aider-polyglot-30 — Qwen 27B 20/30 (66.7%) > Gemma 4 31B 17/30 (56.7%)
New "Quality benches — Aider Polyglot 30" section captures pass-rate /
agentic-coding signal alongside the existing TPS rows. First two rows:

- Qwen 3.6 27B (AutoRound INT4) on dual.yml: 20/30 = 66.7%, 19 min wall
- Gemma 4 31B (Intel AutoRound INT4) on dual.yml: 17/30 = 56.7%, 19 min wall

Both run on 2× 3090 PCIe, 230 W cap, threads=2. Qwen edges Gemma by +10pp
despite being smaller; java is the biggest swing (Qwen 4/5 vs Gemma 1/5).

Critical caveat documented: Qwen with thinking ON is unusable for agentic
benches on this hardware — hits the 1500s subprocess cap before any
exercise completes. The new --default-chat-template-kwargs flag in our
vLLM Qwen composes (commit 534d29f) sets enable_thinking=false by default.

Aider-polyglot run via benchlocal-cli's aider-polyglot-30 pack. Cross-rig
re-run path documented in the section.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 18:04:58 +00:00
noonghunnaandClaude Opus 4.7 7d91ac75e0 feat(power-cap-sweep): --include-commit flag stamps club-3090 git SHA in report header (closes #112)
@laurimyllari noted in disc #62 that with the project moving fast,
including the club-3090 git commit in sweep output helps correlate
cross-rig sweeps to the script revision they were run against. He
stamped `aa99173` manually; the script should do it for us.

New flag:

  --include-commit   Stamp the club-3090 git commit (short SHA) in the
                     report header next to the date. Off by default.

Implementation:
- Captures `git -C "$REPO_ROOT" rev-parse --short HEAD` once at header-build
  time (REPO_ROOT was already known to the script).
- Injects into the report header next to **Date:**, e.g.:

    **Date:** 2026-05-10T17:55:00Z &nbsp; **club-3090 commit:** `534d29f`

- Suppress (don't stamp "n/a") when run from a non-clone or git is
  unreachable. Closes the curl-pipe-from-docs UX hole — `curl ... | bash`
  users don't have a clone, so the field just disappears rather than
  showing a confusing "n/a".

Off by default per the issue rationale: surprise stamping confuses
contributors running from documentation snippets.

Closes #112.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:57:20 +00:00
noonghunnaandClaude Opus 4.7 534d29f1b1 fix(qwen3.6-27b): use --default-chat-template-kwargs (not --chat-template-kwargs)
Follow-up to 29d17ed which used the wrong vLLM flag name (`--chat-template-kwargs`),
causing boot failure: "vllm: error: unrecognized arguments: --chat-template-kwargs
{"enable_thinking": false}".

vLLM's actual flag for setting server-side default chat template kwargs is
`--default-chat-template-kwargs` (with the `default-` prefix). Confirmed by
- vLLM nightly source: vllm/engine/arg_utils.py defines
  `default_chat_template_kwargs: dict[str, Any] | None = None` with
  json.loads parsing.
- vLLM PR #37739 ("Fix default_chat_template_kwargs handling in Responses API")
  references it as already available in the shared render stack.

Behavior unchanged: thinking OFF by default for all 17 vLLM Qwen 27B composes
(plus bounded-thinking unaffected — it intentionally keeps thinking ON).
Per-request override still works via OpenAI extra_body:
  {"chat_template_kwargs": {"enable_thinking": true}}

Verified: vllm-qwen36-27b-dual now boots cleanly with the corrected flag.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:44:50 +00:00
noonghunnaandClaude Opus 4.7 29d17ed82d feat(qwen3.6-27b): thinking OFF by default across all 21 composes
Previously, all 19 vLLM Qwen composes used --reasoning-parser qwen3 (parses
<think>...</think> output blocks) but did NOT explicitly disable the
thinking template. That meant Qwen3 default — thinking ON — applied
across the board. This burns hidden token budget on internal CoT for
every request, hurts latency, and creates a "Qwen looks much slower
than Gemma" gap on agentic benchmarks (which is exactly what we just
hit on aider-polyglot — Qwen exceeded the 1500s timeout, Gemma
finished in 19 min).

Change:
- All vLLM composes (18 of them) now pass `--chat-template-kwargs
  '{"enable_thinking": false}'` after `--reasoning-parser qwen3`.
- llama.cpp single/docker-compose.yml: DISABLE_THINKING default flipped
  0 → 1 (thinking now OFF by default; opt back in via DISABLE_THINKING=0).
- llama.cpp single/concurrent.yml: gained `--chat-template-kwargs` flag
  with default `{"enable_thinking":false}` (overridable via
  CHAT_TEMPLATE_KWARGS env).

NOT changed:
- bounded-thinking.yml — that's the structured-CoT compose where thinking
  IS the feature. Reverted my initial blanket change for that one.
- qwopus-bf16mtp.yml — already had enable_thinking=false (preview compose).

Users who want thinking ON can:
- For vLLM: pass `chat_template_kwargs: {enable_thinking: true}` in the
  per-request body (works fine).
- For llama.cpp: set `DISABLE_THINKING=0` in compose/.env.

Aligns with Gemma 4's "thinking off by default" (it ships that way upstream)
and removes the Qwen vs Gemma framework-bench skew.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:25:45 +00:00
noonghunnaandClaude Opus 4.7 3909c2d6b8 feat(setup): interactive MODEL_DIR prompt for fresh TTY users
Previously setup.sh silently defaulted MODEL_DIR to <repo>/models-cache,
which meant fresh users got ~14-21 GB of model weights downloaded INTO
their git tree without realizing it. The relative path is also wrong
when run from anywhere except the repo root.

New 4-step resolution order in setup.sh:
  1. MODEL_DIR exported in calling shell  → use as-is (unchanged)
  2. .env at repo root sets MODEL_DIR     → source it (NEW)
  3. Interactive prompt (only on TTY)     → ask user (NEW)
  4. Silent fallback to <repo>/models-cache (unchanged for non-TTY)

The interactive prompt only fires when:
  - MODEL_DIR is not in the calling env, AND
  - .env doesn't already set it, AND
  - both stdin AND stdout are TTYs (CI / scripted runs unaffected)

Three options offered:
  1. <repo>/models-cache   (the old silent default — kept as option)
  2. $HOME/models           (sensible cross-rig default)
  3. custom absolute path

After picking, optionally persists the choice to .env (gitignored) so
re-runs skip the prompt. Existing .env files are updated in-place if
they already set MODEL_DIR; appended-to otherwise.

Closes the UX hole RobH589 hit in club-3090#116 — the relative
../../../../../models-cache default that was resolving wrong when
not run from the compose dir.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:16:55 +00:00
noonghunnaandClaude Opus 4.7 cc3a717524 docs(recipes): use \$MODEL_DIR placeholder + sensible cross-rig default
Two recipe scripts (single-card-default.sh, single-card-max-ctx.sh) had
the dev rig path /mnt/models/huggingface/... baked into the MODEL_PATH
default + the comment-block instructions for downloading the GGUF.

Updated:
- Comments now show \${MODEL_DIR}/qwen3.6-27b-gguf/... as the placeholder
  (matches what the just-fixed README + LLAMA_CPP.md docs say).
- MODEL_PATH default changed from /mnt/models/huggingface/... to
  \${MODEL_DIR:-\$HOME/models}/qwen3.6-27b-gguf/... — falls back to
  ~/models/ if MODEL_DIR isn't set, which is a more reasonable default
  for cross-rig users than our /mnt/models/huggingface/ path.
- file-exists check at line 25 still fails loudly with the resolved path
  if neither MODEL_DIR nor MODEL_PATH is set correctly.

Follow-up to fbf3431 (de-bind \$MODEL_DIR from rig path in docs).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:14:21 +00:00
noonghunnaandClaude Opus 4.7 fbf343129c docs: use \$MODEL_DIR placeholder, not the dev rig's /mnt/models/huggingface/
User docs were hardcoding the dev rig path (/mnt/models/huggingface/...) as
if it was canonical. It's not — cross-rig users have models at /data/models,
~/models, /mnt/nvme/llms, etc. Setting MODEL_DIR per their setup is the
intended UX (the compose already supports it via env-var default).

Updates:
- models/qwen3.6-27b/llama-cpp/README.md: download examples now use
  \$MODEL_DIR/qwen3.6-27b-gguf/ instead of /mnt/models/huggingface/...
- models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml: header
  comment uses \$MODEL_DIR/qwen3.6-27b-gguf/ for download examples + says
  "MODEL_DIR=/your/models/dir docker compose up -d" instead of our path.
- docs/engines/LLAMA_CPP.md: same treatment + cleaned up Qwen3.5 + DFlash
  draft path examples to also use \$MODEL_DIR.
- scripts/preflight.sh: hf download hint shows literal \${MODEL_DIR} so user
  knows what to set, plus explicit "set MODEL_DIR first" line. Previously
  echoed the resolved relative path (../../../../models-cache) which lands
  outside the repo if pwd isn't the compose dir.

Caught by RobH589 in club-3090#116 — they hit the path-resolved-to-root-of-drive
case from the relative-path default. Closes the doc UX side; the compose's
env-var override mechanism was already correct.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:13:29 +00:00
noonghunnaandClaude Opus 4.7 cf7f1959fd fix: 4 stale refs missed in 2026-05-10 reorg push (caught by RobH589 #116)
After the GGUF dir move (/mnt/models/gguf/qwen3.6-27b/ → /mnt/models/
huggingface/qwen3.6-27b-gguf/), four refs were not updated and led to
a path-mismatch loop reported in club-3090#116:

- scripts/preflight.sh `hf download` hint pointed at qwen3.6-27b/, but
  the compose default expects qwen3.6-27b-gguf/. Same for the mv hint
  for mmproj relocation, and the in-container mmproj default at line 329.
- models/qwen3.6-27b/llama-cpp/README.md example command still said
  `MODEL_DIR=/mnt/models/gguf` (now /mnt/models/huggingface).
- models/qwen3.6-27b/llama-cpp/compose/single/{docker-compose,concurrent}.yml
  header comment said "Q5_K_XL" but the actual default has been Q3_K_XL
  for a while (this one predates the reorg — just stale doc).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:06:20 +00:00
30 changed files with 257 additions and 36 deletions
+21
View File
@@ -244,3 +244,24 @@ Cross-rig data on Google's official Gemma 4 MTP "assistant" drafter (released 20
None close the **-13% narr / -11% code gap to 3dluvr's anchor**. Remaining gap likely rig-specific (3dluvr's EPYC 7J13 / different PCIe topology / 275W cap / etc) rather than tunable via flags. Cross-rig productionizable settings: stick with default `--dtype bfloat16`, default cudagraph, default scheduling. Custom override `cudagraph_capture_sizes [9]` worth it ONLY if you serve >95% n=8-MTP-single-stream code traffic (e.g. dedicated coding-agent endpoint) where the +2.5% code lift exceeds the -3% narrative loss. |
| `dual.yml`-shape forced TP=1 | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | 32K | **159.67 / 215.10** (decode 160.71 / 217.30) | 27.5 GB | 2026-05-07 | **First single-5090 Gemma 4 MTP data point.** First non-OOM single-card Gemma 4 result on the matrix — the 32 GB Blackwell envelope clears the 24 GB Ampere boot OOM. CV 1.9%/1.8%, peak 426 W. **+46% narr / +51% code over @noonghunna's 2× 3090 TP=2 baseline (109/142)** — single-card 5090 beats dual-3090 on Gemma 4. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
| `dual-dflash.yml`-shape forced TP=1 (mem-util 0.96, max-model-len 12000) | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | **12K** | **150.40 / 261.06** (decode 151.16 / 264.62) | 28.8 GB | 2026-05-07 | **First single-5090 Gemma 4 DFlash data point.** Trade vs MTP row above: ~6% narr loss, **+21% code lift** (215→261). 1st-warmup TTFT outlier (73 s) suggests cudagraph warmup taking longer on first request; subsequent warmups stable at <40 ms. CV 3.6%/2.8%, peak 440 W. **Required mem-util 0.96 + max-model-len 12K** to fit BF16 weights + DFlash N=5 drafter on 32 GB — DFlash drafter footprint pushes out ctx ceiling vs MTP's 32K. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
---
## Quality benches — Aider Polyglot 30
Pass rate on a curated 30-exercise subset of [aider-polyglot-benchmark](https://github.com/Aider-AI/polyglot-benchmark) (5 per language across cpp/go/java/javascript/python/rust, mix of easy/medium/hard). Tests **edit-format reliability** AND **algorithmic correctness** — does the model emit diffs aider can apply, AND do the resulting tests pass.
Run via [`benchlocal-cli`](https://github.com/noonghunna/benchlocal-cli) `aider-polyglot-30` pack. Different from the TPS rows above — this is a quality / agentic-coding signal, not a throughput measurement.
| Model | Compose | Rig | Pass / Total | % | Wall (real) | Wall (sum-dur) | Tokens (P+C) | Date | Notes |
|---|---|---|---:|---:|---:|---:|---:|---|---|
| **Qwen 3.6 27B** (AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | **20 / 30** | **66.7%** | 19.0 min | 34.0 min | 436K + 111K = 547K | 2026-05-10 | **`enable_thinking=false`** (server-side `--default-chat-template-kwargs '{"enable_thinking": false}'` + per-request `extra_body` belt). With thinking ON: 0/30 (1500s timeout exceeded before any exercise completed — Qwen burns the token budget on hidden CoT). Per-language: cpp 3/5 · go 4/5 · **java 4/5** · js 4/5 · python 2/5 · rust 3/5. `threads=2`. |
| **Gemma 4 31B** (Intel AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | 17 / 30 | 56.7% | 19.2 min | 19.2 min | 380K + 72K = 452K | 2026-05-10 | Default thinking off (Gemma 4's chat template requires explicit `enable_thinking=true` to enable). Per-language: cpp 2/5 · **go 4/5** · java 1/5 · **js 4/5** · python 3/5 · rust 3/5. `threads=2`. |
**Notable**:
- Qwen 3.6 27B beats Gemma 4 31B by **+10 pp** despite 4 GB fewer parameters. Java is the biggest swing (4/5 vs 1/5 — `affine-cipher` specifically tripped Gemma).
- Gemma is meaningfully faster wall-clock (sum-of-exercise-durations 19 vs 34 min), suggesting Qwen produces longer per-turn answers but they convert to passes more reliably.
- Qwen with thinking ON is unusable for this kind of bench on club-3090 hardware: hits the 1500s subprocess timeout cap (now bumped to 2700s) before completing any exercises — the hidden CoT eats the per-exercise token budget. **Set `enable_thinking=false` for any agentic / multi-turn workload.**
**To re-run cross-rig**: `bash scripts/quality-test.sh --pack aider-polyglot-30 --enable-sandboxed-packs` against your endpoint. See [docs/QUALITY_TEST.md](docs/QUALITY_TEST.md) for the harness setup.
+75
View File
@@ -2,8 +2,83 @@
Changes that span the entire stack — engine version pins, script behavior, repo structure. Per-model dated history lives in `models/<name>/CHANGELOG.md`.
## Versioning
club-3090 ships docker composes, system scripts, and patch bundles that downstream rigs run as-is — it's effectively software, not just rolling recipes. **SemVer** in `0.x` from `v0.3.0` onward; treat any minor bump as potentially breaking until `1.0`.
Past CalVer tags (preserved for history):
| CalVer tag | SemVer equivalent | Date |
|---|---|---|
| `v2026.05.09` | (≈ v0.1.0) | 2026-05-09 — first tagged release |
| `v2026.05.10` | (≈ v0.2.0) | 2026-05-10 — stack reorg + Gemma 4 INT8 PTH unblock |
CHANGELOG entries below 2026-05-10 use date-prose headings (pre-SemVer convention). New entries from `v0.3.0` use `## v0.X.Y (YYYY-MM-DD) — title`.
## v0.3.1 (2026-05-10) — soak-helper captures `delta.reasoning` (closes cross-rig silent-empty turn-5 mystery)
One-line patch behind a meaningful finding: vLLM nightly (`0.20.2rc1.dev9+`) emits the qwen3 reasoning parser's output under `delta.reasoning` (legacy field name), not `delta.reasoning_content` that `soak-helper.py` was watching. Result: any thinking-on response whose `<think>` block didn't close within `max_tokens` showed up as a silent-empty turn (`ttft_ms == t_ms`, `decode_tps = 0.0`, empty `content` + empty `reasoning_content`) — the model was generating correctly, the harness just couldn't see the wire.
This closes the cross-rig "silent-empty turn 5" pattern that was parked behind the Cliff 2b investigation. It was a harness measurement bug, not a model/rig issue.
Validation on JDWarner's exact repro request (#107 turn 5, math problem + `enable_thinking=true` + `max_tokens=2000`):
| | Before | After |
|---|---|---|
| `ttft_ms` | 22709 (== `t_ms`) | **234** |
| `decode_tps` | **0.0** | **88.985** |
| `reasoning_content` chars | 0 | 3959 |
| `completion_tokens` | 2000 | 2000 (unchanged) |
Validation soak (`qwen3.6-27b dual.yml`, fresh-mode, 20 sessions × 5 turns):
```
verdict PASS
silent_empty 0 / 100 (0.0%) ← was ~3-5/40 baseline
p50_decode_tps 90.22
p95_ttft_ms 1389
errors 0
max_growth_mib 0 / 200
```
**Patch** ([88eb67a](https://github.com/noonghunna/club-3090/commit/88eb67a)): `soak-helper.py` accumulates `delta.reasoning_content || delta.reasoning` (covers both vLLM streaming-output shapes, current and legacy). Three-line change inside `cmd_run`'s SSE chunk loop.
## v0.3.0 (2026-05-10) — Qwen thinking-off default, MODEL_DIR UX overhaul, docs cleanup, aider-polyglot bench rows
Eight commits since `v2026.05.10`. Headline: Qwen 3.6 27B now defaults `enable_thinking=false` server-side via `--default-chat-template-kwargs` across all 21 vLLM Qwen composes (closes the recurring "Qwen looks much slower than Gemma" agentic-bench skew). MODEL_DIR docs no longer bake the dev-rig path. New `--include-commit` flag on `power-cap-sweep.sh`.
**Compose change — Qwen 3.6 27B thinking OFF by default** ([534d29f](https://github.com/noonghunna/club-3090/commit/534d29f), [29d17ed](https://github.com/noonghunna/club-3090/commit/29d17ed)):
- All 17 vLLM Qwen 27B composes (single + dual + multi4 + nvlink variants + carnice + qwopus) gain `--default-chat-template-kwargs '{"enable_thinking": false}'` after `--reasoning-parser qwen3`. Server-side default — no per-request kwarg needed.
- llama.cpp single composes: `DISABLE_THINKING` default flipped 0 → 1; `concurrent.yml` gains `--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}`.
- `bounded-thinking.yml` intentionally keeps thinking ON (its whole purpose is constrained CoT).
- Per-request override still works for users who want thinking on: `chat_template_kwargs: {enable_thinking: true}` in OpenAI request body.
- Empirical impact: aider-polyglot-30 bench went from 0/30 (1500s timeout, hidden CoT eats budget) → 20/30 (66.7%) on the same hardware once thinking defaulted off.
- Note: 29d17ed used the wrong flag name (`--chat-template-kwargs`); 534d29f corrected to `--default-chat-template-kwargs` after vLLM nightly source confirmed the actual CLI option (referenced in upstream PR #37739).
**MODEL_DIR UX overhaul** (closes club-3090#116 — RobH589's "downloads landing in root of drive" report):
- `cf7f195` — fixed 4 stale refs missed in the 2026-05-10 reorg push (preflight.sh hf-download hint, README.md example, compose header Q5_K_XL→Q3_K_XL stale comment).
- `fbf3431` — replaced the dev-rig path `/mnt/models/huggingface/...` with `$MODEL_DIR/...` placeholders across `models/qwen3.6-27b/llama-cpp/README.md`, `docs/engines/LLAMA_CPP.md`, the compose header, and the `preflight.sh` hint. Cross-rig users no longer see our path baked into instructions.
- `cc3a717` — recipe scripts (`single-card-default.sh`, `single-card-max-ctx.sh`) default `MODEL_PATH` to `${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/...`.
- `3909c2d` — `setup.sh` now **prompts interactively** when `MODEL_DIR` isn't set and stdin is a TTY: pick `<repo>/models-cache`, `$HOME/models`, or custom path; optionally persist to `.env` (gitignored). CI / scripted runs (no TTY) get the silent fallback unchanged.
**`power-cap-sweep.sh` — `--include-commit` flag** ([7d91ac7](https://github.com/noonghunna/club-3090/commit/7d91ac7), closes [#112](https://github.com/noonghunna/club-3090/issues/112)):
- Off by default. When set, stamps the club-3090 git short SHA in the report header next to the date. Useful for cross-rig sweep correlation per @laurimyllari's request.
- Suppresses the field entirely (no "n/a") when run from a non-clone — closes the curl-pipe-from-docs UX hole.
**BENCHMARKS.md — Aider Polyglot 30 quality bench** ([e08988e](https://github.com/noonghunna/club-3090/commit/e08988e)):
- New "Quality benches — Aider Polyglot 30" section with cross-model rows: Qwen 3.6 27B `dual.yml` 20/30 (66.7%) vs Gemma 4 31B `dual.yml` 17/30 (56.7%).
- Documents the thinking-on/thinking-off Qwen failure mode + per-language breakdown.
- Run via [`benchlocal-cli`](https://github.com/noonghunna/benchlocal-cli)'s `aider-polyglot-30` pack.
**cliff.toml + CHANGELOG meta**: switched from CalVer rolling-stack framing to SemVer software framing. Past CalVer tags preserved for history.
---
## 2026-05-10 — Stack reorg: services consolidation, gpu-mode under git, ComfyUI in services/, pin tracker 🧹
(This entry covers `v2026.05.10` — see SemVer mapping above.)
The rig had grown organically across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`, `/home/wasif/`), three repos (single-3090, dual-3090, club-3090), and ad-hoc paths under `/opt/ai/vllm-src/`, `/opt/ai/vendor/`, `/home/wasif/lucebox-hub/`, `/home/wasif/llama.cpp*`. Disk-out on `/` (97% used) forced a clean-up; rather than just prune, we consolidated the layout while we were at it.
**What landed in this repo:**
+1 -1
View File
@@ -34,7 +34,7 @@ body = """
## Pinning to this release
This is a snapshot of the rolling stack — not a versioned API. To pin to this exact state:
**Versioning:** SemVer in `0.x` — treat any minor bump as potentially breaking until `1.0`. Past CalVer tags (`v2026.05.09`, `v2026.05.10`) are preserved for history; SemVer takes over from `v0.3.0` onward.
```bash
git checkout {{ version }}
+10 -10
View File
@@ -86,7 +86,7 @@ This is exactly why our launch frame is **two routes, not one** ([README](../../
```bash
# Use hf CLI (pip install 'huggingface-hub[hf_transfer]')
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
```
Confirm size matches the HuggingFace listing. If a `sha256` is published, verify it.
@@ -108,7 +108,7 @@ For a sane mid-context default (65K, plenty for chat + light agent work):
```bash
/opt/llama.cpp/build/bin/llama-server \
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-c 65536 \
--host 0.0.0.0 --port 8020 \
-ngl 999 \
@@ -131,7 +131,7 @@ Recipe (community-reported, validated by multiple users on r/LocalLLaMA):
```bash
/opt/llama.cpp/build/bin/llama-server \
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-ngl 99 \
-c 262144 \
-np 1 \
@@ -154,12 +154,12 @@ Sustained throughput at 262K with this config is typically **35-45 tok/s** on a
Download the `mmproj` model:
```bash
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
```
Add to launch:
```bash
--mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
--mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
```
### 5. Tool calls (limited)
@@ -185,12 +185,12 @@ cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
# Download draft model (~500 MB)
hf download z-lab/Qwen3.6-27B-DFlash --local-dir /mnt/models/huggingface/z-lab/Qwen3.6-27B-DFlash/
hf download z-lab/Qwen3.6-27B-DFlash --local-dir $MODEL_DIR/z-lab/Qwen3.6-27B-DFlash/
# Launch
/opt/lucebox-hub/build/bin/llama-server \
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
--draft /mnt/models/gguf/qwen3.6-27b-dflash/dflash-N5.gguf \
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
--draft $MODEL_DIR/qwen3.6-27b-dflash-gguf/dflash-N5.gguf \
--draft-max 5 \
--draft-min 1 \
-c 65536 \
@@ -210,8 +210,8 @@ If you have two GPUs (e.g. 2× 3090), lucebox-hub now supports a heterogeneous-s
```bash
# Target on GPU 0, DFlash draft on GPU 1
/opt/lucebox-hub/build/bin/llama-server \
-m /mnt/models/gguf/qwen3.5-27b/Qwen3.5-27B-Q4_K_M.gguf \
--draft /mnt/models/gguf/qwen3.5-27b-dflash/dflash-N5.gguf \
-m $MODEL_DIR/qwen3.5-27b-gguf/Qwen3.5-27B-Q4_K_M.gguf \
--draft $MODEL_DIR/qwen3.5-27b-dflash-gguf/dflash-N5.gguf \
--target-gpu 0 --draft-gpu 1 \
--draft-max 16 --draft-min 1 \
-c 262144 \
+4 -4
View File
@@ -30,7 +30,7 @@ Showcase: full **262K context** on one 3090 with vision + q4_0 KV.
```bash
cd models/qwen3.6-27b/llama-cpp/compose
MODEL_DIR=/mnt/models/gguf docker compose up -d
MODEL_DIR=/your/models/dir docker compose up -d
```
Memory budget: 14.5 GB (Q3_K_XL) + 4.5 GB KV @ 262K + 0.8 GB mmproj ≈ 20 GB / 24 GB.
@@ -66,7 +66,7 @@ The Q3_K_XL number at 262K is **lower than community-reported 35-45 tok/s** ([Re
```bash
# 1. Get a GGUF quant (recommended: Unsloth's Q4_K_M)
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
# 2. Build llama.cpp with CUDA support
git clone https://github.com/ggerganov/llama.cpp /opt/llama.cpp
@@ -99,9 +99,9 @@ GGUFs of this model are at [unsloth/Qwen3.6-27B-GGUF](https://huggingface.co/uns
## Vision (mmproj)
```bash
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
# Add to launch: --mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
# Add to launch: --mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
```
Vision works via the mmproj model. Sample text+image queries are OpenAI-compat.
@@ -1,6 +1,6 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
# Engine: llama.cpp (NOT vLLM)
# Topology: Single 3090 (TP=1)
# Drafter: none
@@ -66,6 +66,7 @@ services:
--cont-batching
--jinja
--reasoning-format ${REASONING_FORMAT:-none}
--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}
deploy:
resources:
reservations:
@@ -1,6 +1,6 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
# Engine: llama.cpp (NOT vLLM — different engine, different memory model)
# Topology: Single 3090 (TP=1)
# Drafter: none (vanilla llama.cpp; MTP via PR #22673 not adopted yet)
@@ -43,15 +43,15 @@
# 1. Get the GGUF + mmproj:
# hf download unsloth/Qwen3.6-27B-GGUF \
# --include "Qwen3.6-27B-UD-Q3_K_XL.gguf" "mmproj-F16.gguf" \
# --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/unsloth-q3kxl
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-q3kxl
# (mmproj sometimes ships at the repo root rather than under unsloth-q3kxl;
# the MMPROJ env var below points at the canonical sidecar location.)
# 2. From this directory:
# MODEL_DIR=/mnt/models/huggingface docker compose up -d
# MODEL_DIR=/your/models/dir docker compose up -d
# 3. curl http://localhost:8020/v1/models → should list the model.
#
# Override defaults via .env or shell:
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /mnt/models/huggingface on this stack)
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /path/to/your/models on your stack)
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf)
# MMPROJ_FILE path under /models (default: qwen3.6-27b-gguf/mmproj-F16.gguf)
# CTX_SIZE total KV pool (default: 262144)
@@ -60,7 +60,7 @@
# Override to `auto` to get separate `reasoning_content` field
# (Qwen3.6 thinking trace) — useful for clients that render
# reasoning_content (most don't). Issue: club-3090#97.
# DISABLE_THINKING set to 1 in .env to add `--chat-template-kwargs '{"enable_thinking":false}'`
# DISABLE_THINKING default: 1 (thinking OFF). Set to 0 to opt INTO thinking
# which forces empty <think></think> blocks in responses. Useful for
# clients (e.g. opencode) that display <think> content as the response.
# Tradeoff: applies to ALL clients on this server — Hermes/agents that
@@ -92,7 +92,7 @@ services:
- -c
- |
set -e
# DISABLE_THINKING=1 in compose/.env appends --chat-template-kwargs to disable
# DISABLE_THINKING=1 (default) appends --chat-template-kwargs to disable
# Qwen3 thinking server-side. Forces the chat template to insert empty
# <think></think> blocks → output goes straight to the response. Useful for
# clients (e.g. opencode) that display <think> content as the response.
@@ -100,7 +100,7 @@ services:
# use thinking lose reasoning capability. See docs/HARDWARE.md and disc club-3090#97.
# Note: $$VAR is YAML-escape for $VAR (compose passes literal $ to bash).
EXTRA_ARGS=()
if [ "$${DISABLE_THINKING:-0}" = "1" ]; then
if [ "$${DISABLE_THINKING:-1}" = "1" ]; then
EXTRA_ARGS+=("--chat-template-kwargs" '{"enable_thinking":false}')
echo "[entrypoint] DISABLE_THINKING=1 — chat template will produce empty <think></think>"
fi
@@ -4,15 +4,15 @@
#
# Prereqs:
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/)
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir ${MODEL_DIR}/qwen3.6-27b-gguf/)
#
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX
set -euo pipefail
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
PORT="${PORT:-8020}"
CTX="${CTX:-65536}"
@@ -18,14 +18,14 @@
#
# Prereqs:
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
#
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX, KV_TYPE
set -euo pipefail
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
PORT="${PORT:-8020}"
CTX="${CTX:-262144}"
KV_TYPE="${KV_TYPE:-q4_0}"
@@ -127,6 +127,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_xml
@@ -118,6 +118,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -141,6 +141,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -130,6 +130,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -154,6 +154,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -149,6 +149,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -257,6 +257,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -139,6 +139,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -237,6 +237,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -134,6 +134,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -138,6 +138,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -247,6 +247,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -357,6 +357,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -374,6 +374,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -274,6 +274,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -97,6 +97,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -157,6 +157,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
+17 -1
View File
@@ -22,6 +22,7 @@
# sudo bash scripts/power-cap-sweep.sh --load-mode decode-concurrent --concurrency 8 --bench-runs 3
# sudo bash scripts/power-cap-sweep.sh --load-mode prefill-heavy
# sudo bash scripts/power-cap-sweep.sh --no-reset # leave at last cap (you reset manually)
# sudo bash scripts/power-cap-sweep.sh --include-commit # stamp club-3090 git short SHA in report header
#
# Load modes:
# decode-single:
@@ -157,6 +158,10 @@ PREFILL_FILLER_REPEATS=""
PREFILL_PROMPT_TOKENS=""
DECODE_CONCURRENT_RUN_SECONDS=""
CALIBRATION_NOTE=""
INCLUDE_COMMIT=0 # --include-commit stamps the club-3090 git short SHA in
# the report header. Off by default — `curl ... | bash`
# users have no clone, and stamping "n/a" is confusing
# (better to suppress the field entirely there).
while [ $# -gt 0 ]; do
case "$1" in
@@ -171,6 +176,7 @@ while [ $# -gt 0 ]; do
--load-target) LOAD_TARGET="$2"; shift 2 ;;
--concurrency-stretch) CONCURRENCY_STRETCH="$2"; shift 2 ;;
--target-cap-seconds) TARGET_CAP_SECONDS="$2"; shift 2 ;;
--include-commit) INCLUDE_COMMIT=1; shift ;;
--no-reset) RESET=0; shift ;;
-h|--help)
sed -n '1,/^set -euo/p' "$0" | grep '^#' | sed 's/^# \?//'
@@ -1120,7 +1126,17 @@ RESULTS_FILE=/tmp/power-cap-summary.md
echo "**Model:** \`${MODEL}\` &nbsp; **Engine:** \`${CONTAINER}\` &nbsp; **Endpoint:** ${URL}"
echo "**Load mode:** \`${LOAD_MODE}\`$([ "$LOAD_MODE" = "decode-single" ] && echo " (${TARGET_CAP_SECONDS}s × 2 timed streams)")$([ "$LOAD_MODE" = "decode-concurrent" ] && echo " (concurrency=${CONCURRENCY}, ${DECODE_CONCURRENT_RUN_SECONDS}s/run × ${BENCH_RUNS} runs × 2 timed batches)")$([ "$LOAD_MODE" = "prefill-heavy" ] && echo " (target-prefill=${TARGET_PREFILL_SECONDS}s, filler_repeats=${PREFILL_FILLER_REPEATS})")$([ "$LOAD_MODE" != "decode-single" ] && echo " (bench-runs=${BENCH_RUNS})")"
[ -n "$CALIBRATION_NOTE" ] && echo "**Calibration:** ${CALIBRATION_NOTE}"
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z"
# --include-commit: stamp club-3090 git short SHA next to the date if requested.
# Suppress entirely (rather than show "n/a") when run from a non-clone or
# when git isn't reachable — closes the curl-pipe-from-docs UX hole.
COMMIT_FRAGMENT=""
if [ "${INCLUDE_COMMIT:-0}" = "1" ]; then
COMMIT_SHA=$(git -C "$REPO_ROOT" rev-parse --short HEAD 2>/dev/null || true)
if [ -n "$COMMIT_SHA" ]; then
COMMIT_FRAGMENT=" &nbsp; **club-3090 commit:** \`${COMMIT_SHA}\`"
fi
fi
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z${COMMIT_FRAGMENT}"
echo ""
if [ "$COOLING" = "unspecified" ]; then
echo "> ⚠️ Cooling class not specified at run time. Add **air / water / AIO** when posting"
+5 -4
View File
@@ -325,8 +325,8 @@ preflight_compose_deps() {
mmproj_in_container="${mmproj_in_container//\$\{MMPROJ_FILE:-/}"
mmproj_in_container="${mmproj_in_container%\}}"
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b/mmproj-F16.gguf"
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b-gguf/mmproj-F16.gguf"
if [[ -n "${GGUF_FILE:-}" ]]; then gguf_in_container="$GGUF_FILE"; fi
if [[ -n "${MMPROJ_FILE:-}" ]]; then mmproj_in_container="$MMPROJ_FILE"; fi
@@ -375,9 +375,10 @@ preflight_compose_deps() {
if [[ $hint_gguf -eq 1 ]]; then
echo "[preflight] hf download unsloth/Qwen3.6-27B-GGUF \\" >&2
echo "[preflight] Qwen3.6-27B-UD-Q3_K_XL.gguf mmproj-F16.gguf \\" >&2
echo "[preflight] --local-dir ${model_dir}/qwen3.6-27b/unsloth-q3kxl" >&2
echo "[preflight] --local-dir \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl" >&2
echo "[preflight] # (set MODEL_DIR first: export MODEL_DIR=\${MODEL_DIR:-/path/to/your/models})" >&2
echo "[preflight] # mmproj lands at unsloth-q3kxl/ — move it up so the default --mmproj path resolves:" >&2
echo "[preflight] # mv ${model_dir}/qwen3.6-27b/unsloth-q3kxl/mmproj-F16.gguf ${model_dir}/qwen3.6-27b/" >&2
echo "[preflight] # mv \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl/mmproj-F16.gguf \${MODEL_DIR}/qwen3.6-27b-gguf/" >&2
echo "[preflight] (~16 GB total. setup.sh today only fetches the vLLM AutoRound weights;" >&2
echo "[preflight] GGUF must be fetched separately for any llamacpp/* variant.)" >&2
fi
+67
View File
@@ -90,6 +90,73 @@ case "${MODEL_NAME}" in
esac
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
# ---------- MODEL_DIR resolution ----------
# Order of precedence:
# 1. MODEL_DIR already exported in the calling shell → use as-is
# 2. .env at repo root sets MODEL_DIR → source it
# 3. Interactive prompt (only if stdin is a TTY) → ask user
# 4. Silent fallback to <repo>/models-cache → in-repo default
#
# The prompt only fires for fresh users on a TTY who haven't set anything.
# CI / scripted runs (no TTY) get the silent fallback, preserving prior behavior.
# Step 2: source repo-root .env if present (lets a saved choice persist)
if [[ -z "${MODEL_DIR:-}" && -f "${ROOT_DIR}/.env" ]]; then
# shellcheck source=/dev/null
set -a; source "${ROOT_DIR}/.env"; set +a
fi
# Step 3: prompt if still unset + interactive
if [[ -z "${MODEL_DIR:-}" && -t 0 && -t 1 ]]; then
echo ""
echo "Where should I put model weights?"
echo " Models are large (Qwen3.6-27B AutoRound: ~14 GB; Gemma 4 31B: ~21 GB)."
echo " This dir lives outside the git tree — pick a location with sufficient free space."
echo ""
echo " 1) ${ROOT_DIR}/models-cache (in-repo, default — pollutes git tree)"
echo " 2) ${HOME}/models (recommended for cross-rig — outside repo)"
echo " 3) custom path"
echo ""
while true; do
read -rp "Choice [1-3] (or set MODEL_DIR env var to skip): " pick
case "${pick}" in
1) MODEL_DIR="${ROOT_DIR}/models-cache"; break ;;
2) MODEL_DIR="${HOME}/models"; break ;;
3)
read -rp " Enter absolute path: " custom
if [[ "${custom}" =~ ^/ ]]; then
MODEL_DIR="${custom}"; break
else
echo " ! must be an absolute path (start with /)" >&2
fi
;;
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
esac
done
echo ""
# Offer to persist the choice so future runs skip the prompt
read -rp "Save MODEL_DIR=${MODEL_DIR} to .env so we skip this next time? [Y/n]: " save
if [[ "${save:-y}" =~ ^[Yy]$ || -z "${save:-}" ]]; then
if [[ -f "${ROOT_DIR}/.env" ]]; then
# Update existing .env (replace MODEL_DIR= line if present, else append)
if grep -qE "^MODEL_DIR=" "${ROOT_DIR}/.env"; then
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=${MODEL_DIR}|" "${ROOT_DIR}/.env"
else
echo "MODEL_DIR=${MODEL_DIR}" >> "${ROOT_DIR}/.env"
fi
else
echo "MODEL_DIR=${MODEL_DIR}" > "${ROOT_DIR}/.env"
fi
echo " → saved. (.env is gitignored.)"
else
echo " → not saved. Set MODEL_DIR=... when re-running, or you'll get this prompt again."
fi
echo ""
fi
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
+9 -3
View File
@@ -557,15 +557,21 @@ def cmd_run(endpoint, req_path, timeout_s, metrics_path):
choices = chunk.get("choices") or []
if choices:
delta = choices[0].get("delta") or {}
if ttft is None and (delta.get("content") or delta.get("reasoning_content") or delta.get("tool_calls")):
# vLLM emits reasoning under either `delta.reasoning_content`
# (older qwen3 reasoner path) or `delta.reasoning` (current
# nightly as of vllm-0.20.2rc1+; legacy field name). Watch
# both so the soak harness doesn't go silent when the
# underlying field name shifts under us.
reasoning_delta = delta.get("reasoning_content") or delta.get("reasoning")
if ttft is None and (delta.get("content") or reasoning_delta or delta.get("tool_calls")):
ttft = time.time() - t0
# Accumulate streamed parts. vLLM splits content/reasoning
# across many small deltas; tool_calls stream as indexed
# objects whose fields (name, arguments) arrive in pieces.
if delta.get("content"):
content_parts.append(delta["content"])
if delta.get("reasoning_content"):
reasoning_parts.append(delta["reasoning_content"])
if reasoning_delta:
reasoning_parts.append(reasoning_delta)
for tc in (delta.get("tool_calls") or []):
idx = tc.get("index", 0)
slot = tool_calls_acc.setdefault(idx, {"id": "", "type": "function", "name": "", "args": ""})