14 Commits
Author SHA1 Message Date
noonghunna 255c743dff chore: trigger v0.3.2 release workflow (GitHub deduped previous tag push)
Release / release (push) Failing after 48s
The v0.3.2 tag was originally pushed at commit 64b0474 but GitHub didn't
emit a CreateEvent (likely dedup after delete + re-push to same SHA), so
the release.yml workflow never fired. Empty commit gives the tag a fresh
SHA that GitHub will process cleanly.
2026-05-10 21:08:33 +00:00
noonghunna 64b0474a62 chore(changelog): automate CHANGELOG + release notes from commits via cliff (Option A)
CHANGELOG.md is now auto-generated from commit messages by git-cliff in
the release workflow. Hand-edits below the static header will be wiped on
the next tag.

Workflow (`.github/workflows/release.yml`):
  - On tag push (`v[0-9]+.[0-9]+.[0-9]+`):
    1. Render GitHub Release body: `git-cliff --latest --strip header`
       → just the per-version section, no SemVer preamble repeat
    2. Regenerate full CHANGELOG.md: `git-cliff` (default = all tags)
       → preserves header + all historical sections
    3. Commit CHANGELOG.md back to master with `[skip ci]` marker
    4. Publish GitHub Release with the latest-only body

Template (`cliff.toml`):
  - `[changelog].header` now holds the SemVer preamble + CalVer→SemVer
    mapping table (preserved across regens; stripped from GitHub Release
    bodies via `--strip header`).
  - `body` template now renders the **full commit message** (subject as
    bold bullet, body indented below) instead of just the first line.
    Rich narrative I write in commit message bodies (tables, validation
    numbers, before/after diffs) now flows into both CHANGELOG.md and the
    GitHub Release page from the same source.
  - Per-release Pin/Diff footer guarded with `{% if version %}` so the
    Unreleased section doesn't emit empty links.

CHANGELOG.md replaced with the auto-gen output. Past hand-written tables
and phase breakdowns are replaced by the corresponding commit messages
(those were already rich for commits that mattered — v0.3.1 soak-helper
fix has its Before/After table in the commit body and renders fine).

Going forward: just write rich commit messages and tag. Both surfaces
update automatically. No hand-edit of CHANGELOG.md required.
2026-05-10 20:53:51 +00:00
noonghunna 83bf73d3ec feat(quality-test): auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for localhost URLs
When the user runs quality-test.sh with a localhost-style URL
(default `http://localhost:8020`, or any `localhost`/`127.x`/`[::1]`
variant), auto-export `BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1` so
benchlocal-cli rewrites the hermes-agent's outbound model endpoint
from `localhost:<port>` to `host.docker.internal:<port>` inside the
Docker sandbox container.

Without this, the hermes-agent inside the sandbox can't reach the
host's vLLM (localhost resolves to the container itself) and every
scenario fails with `"API call failed after 3 retries: Connection
error."` — produced spurious 0/20 grades on this rig prior to the
benchlocal-cli runner.py 9c1566f fix.

Skips the auto-set when:
  - User already set the env var (explicit override)
  - URL points at a non-loopback host (real LAN IP, k8s service name,
    host.docker.internal already) — no rewrite needed

Emits a stderr breadcrumb when the auto-set fires so users can see
what changed.
2026-05-10 20:29:49 +00:00
noonghunna 9db8b2603c docs(changelog): v0.3.1 entry for soak-helper delta.reasoning capture
Release / release (push) Failing after 46s
Documents the silent-empty turn-5 root cause (vLLM nightly field-name
shift to delta.reasoning) + validation soak results.
2026-05-10 19:16:44 +00:00
noonghunna 88eb67aa18 fix(soak-helper): capture delta.reasoning alongside delta.reasoning_content
vLLM nightly (0.20.2rc1.dev9+) emits the qwen3 reasoning parser's output
under `delta.reasoning` (legacy field name), not `delta.reasoning_content`
that soak-helper.py was watching. Result: for any thinking-on response
whose `<think>` block doesn't close within `max_tokens`, soak-helper saw
zero deltas → fell back to the "couldn't measure" path → reported
`ttft_ms == t_ms` and `decode_tps = 0.0`. The model was generating
correctly; the harness just couldn't see the wire output.

Repro request (JDWarner's #107 turn 5): math problem with
`max_tokens=2000` + `chat_template_kwargs.enable_thinking=true`.

Before patch:
  status=200  t_ms=22709  ttft_ms=22709  decode_tps=0.0
  completion_tokens=2000  content=""  reasoning_content=""

After patch (same request, same compose, same model):
  status=200  t_ms=22709  ttft_ms=234  decode_tps=88.985
  completion_tokens=2000  content=""  reasoning_content="Here's a thinking
  process:\n\n1. **Understand the User's Problem:**\n..."  (3959 chars)

Validation soak (fresh-mode, 20 sessions × 5 turns = 100 turns, qwen3.6-27b
dual.yml):
  verdict        PASS
  silent_empty   0 / 100 (0.0%)   ← was ~3-5/40 baseline
  p50_decode_tps 90.22
  p95_ttft_ms    1389
  errors         0
  max_growth     0 MiB / 200

Closes the cross-rig "silent-empty turn-5" pattern parked behind the
Cliff 2b investigation — it was a harness measurement bug, not a model
or rig issue.
2026-05-10 19:15:57 +00:00
noonghunna 7080f1f89b release: SemVer adoption + v0.3.0 changelog entry
Release / release (push) Failing after 1m25s
club-3090 is software (docker composes + system scripts + patch bundles
that downstream rigs run as-is), not just rolling recipes. Switch from
CalVer to SemVer from v0.3.0 onward; past CalVer tags (v2026.05.09,
v2026.05.10) preserved for history.

CHANGELOG.md: convention note + retroactive CalVer→SemVer mapping +
new v0.3.0 (2026-05-10) entry covering 8 commits since v2026.05.10:
- Qwen 3.6 27B thinking OFF default across all 21 composes
- MODEL_DIR UX overhaul (closes #116) — interactive setup prompt,
  $MODEL_DIR placeholder everywhere
- power-cap-sweep --include-commit (closes #112)
- BENCHMARKS aider-polyglot-30 row

cliff.toml: drop "snapshot of the rolling stack — not a versioned
API" framing; add SemVer note. Tag pattern v[0-9]+.[0-9]+.[0-9]+
already matches both CalVer and SemVer, so the cliff release
workflow needs no changes.
2026-05-10 18:29:06 +00:00
noonghunnaandClaude Opus 4.7 e08988e614 docs(benchmarks): aider-polyglot-30 — Qwen 27B 20/30 (66.7%) > Gemma 4 31B 17/30 (56.7%)
New "Quality benches — Aider Polyglot 30" section captures pass-rate /
agentic-coding signal alongside the existing TPS rows. First two rows:

- Qwen 3.6 27B (AutoRound INT4) on dual.yml: 20/30 = 66.7%, 19 min wall
- Gemma 4 31B (Intel AutoRound INT4) on dual.yml: 17/30 = 56.7%, 19 min wall

Both run on 2× 3090 PCIe, 230 W cap, threads=2. Qwen edges Gemma by +10pp
despite being smaller; java is the biggest swing (Qwen 4/5 vs Gemma 1/5).

Critical caveat documented: Qwen with thinking ON is unusable for agentic
benches on this hardware — hits the 1500s subprocess cap before any
exercise completes. The new --default-chat-template-kwargs flag in our
vLLM Qwen composes (commit 534d29f) sets enable_thinking=false by default.

Aider-polyglot run via benchlocal-cli's aider-polyglot-30 pack. Cross-rig
re-run path documented in the section.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 18:04:58 +00:00
noonghunnaandClaude Opus 4.7 7d91ac75e0 feat(power-cap-sweep): --include-commit flag stamps club-3090 git SHA in report header (closes #112)
@laurimyllari noted in disc #62 that with the project moving fast,
including the club-3090 git commit in sweep output helps correlate
cross-rig sweeps to the script revision they were run against. He
stamped `aa99173` manually; the script should do it for us.

New flag:

  --include-commit   Stamp the club-3090 git commit (short SHA) in the
                     report header next to the date. Off by default.

Implementation:
- Captures `git -C "$REPO_ROOT" rev-parse --short HEAD` once at header-build
  time (REPO_ROOT was already known to the script).
- Injects into the report header next to **Date:**, e.g.:

    **Date:** 2026-05-10T17:55:00Z &nbsp; **club-3090 commit:** `534d29f`

- Suppress (don't stamp "n/a") when run from a non-clone or git is
  unreachable. Closes the curl-pipe-from-docs UX hole — `curl ... | bash`
  users don't have a clone, so the field just disappears rather than
  showing a confusing "n/a".

Off by default per the issue rationale: surprise stamping confuses
contributors running from documentation snippets.

Closes #112.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:57:20 +00:00
noonghunnaandClaude Opus 4.7 534d29f1b1 fix(qwen3.6-27b): use --default-chat-template-kwargs (not --chat-template-kwargs)
Follow-up to 29d17ed which used the wrong vLLM flag name (`--chat-template-kwargs`),
causing boot failure: "vllm: error: unrecognized arguments: --chat-template-kwargs
{"enable_thinking": false}".

vLLM's actual flag for setting server-side default chat template kwargs is
`--default-chat-template-kwargs` (with the `default-` prefix). Confirmed by
- vLLM nightly source: vllm/engine/arg_utils.py defines
  `default_chat_template_kwargs: dict[str, Any] | None = None` with
  json.loads parsing.
- vLLM PR #37739 ("Fix default_chat_template_kwargs handling in Responses API")
  references it as already available in the shared render stack.

Behavior unchanged: thinking OFF by default for all 17 vLLM Qwen 27B composes
(plus bounded-thinking unaffected — it intentionally keeps thinking ON).
Per-request override still works via OpenAI extra_body:
  {"chat_template_kwargs": {"enable_thinking": true}}

Verified: vllm-qwen36-27b-dual now boots cleanly with the corrected flag.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:44:50 +00:00
noonghunnaandClaude Opus 4.7 29d17ed82d feat(qwen3.6-27b): thinking OFF by default across all 21 composes
Previously, all 19 vLLM Qwen composes used --reasoning-parser qwen3 (parses
<think>...</think> output blocks) but did NOT explicitly disable the
thinking template. That meant Qwen3 default — thinking ON — applied
across the board. This burns hidden token budget on internal CoT for
every request, hurts latency, and creates a "Qwen looks much slower
than Gemma" gap on agentic benchmarks (which is exactly what we just
hit on aider-polyglot — Qwen exceeded the 1500s timeout, Gemma
finished in 19 min).

Change:
- All vLLM composes (18 of them) now pass `--chat-template-kwargs
  '{"enable_thinking": false}'` after `--reasoning-parser qwen3`.
- llama.cpp single/docker-compose.yml: DISABLE_THINKING default flipped
  0 → 1 (thinking now OFF by default; opt back in via DISABLE_THINKING=0).
- llama.cpp single/concurrent.yml: gained `--chat-template-kwargs` flag
  with default `{"enable_thinking":false}` (overridable via
  CHAT_TEMPLATE_KWARGS env).

NOT changed:
- bounded-thinking.yml — that's the structured-CoT compose where thinking
  IS the feature. Reverted my initial blanket change for that one.
- qwopus-bf16mtp.yml — already had enable_thinking=false (preview compose).

Users who want thinking ON can:
- For vLLM: pass `chat_template_kwargs: {enable_thinking: true}` in the
  per-request body (works fine).
- For llama.cpp: set `DISABLE_THINKING=0` in compose/.env.

Aligns with Gemma 4's "thinking off by default" (it ships that way upstream)
and removes the Qwen vs Gemma framework-bench skew.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:25:45 +00:00
noonghunnaandClaude Opus 4.7 3909c2d6b8 feat(setup): interactive MODEL_DIR prompt for fresh TTY users
Previously setup.sh silently defaulted MODEL_DIR to <repo>/models-cache,
which meant fresh users got ~14-21 GB of model weights downloaded INTO
their git tree without realizing it. The relative path is also wrong
when run from anywhere except the repo root.

New 4-step resolution order in setup.sh:
  1. MODEL_DIR exported in calling shell  → use as-is (unchanged)
  2. .env at repo root sets MODEL_DIR     → source it (NEW)
  3. Interactive prompt (only on TTY)     → ask user (NEW)
  4. Silent fallback to <repo>/models-cache (unchanged for non-TTY)

The interactive prompt only fires when:
  - MODEL_DIR is not in the calling env, AND
  - .env doesn't already set it, AND
  - both stdin AND stdout are TTYs (CI / scripted runs unaffected)

Three options offered:
  1. <repo>/models-cache   (the old silent default — kept as option)
  2. $HOME/models           (sensible cross-rig default)
  3. custom absolute path

After picking, optionally persists the choice to .env (gitignored) so
re-runs skip the prompt. Existing .env files are updated in-place if
they already set MODEL_DIR; appended-to otherwise.

Closes the UX hole RobH589 hit in club-3090#116 — the relative
../../../../../models-cache default that was resolving wrong when
not run from the compose dir.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:16:55 +00:00
noonghunnaandClaude Opus 4.7 cc3a717524 docs(recipes): use \$MODEL_DIR placeholder + sensible cross-rig default
Two recipe scripts (single-card-default.sh, single-card-max-ctx.sh) had
the dev rig path /mnt/models/huggingface/... baked into the MODEL_PATH
default + the comment-block instructions for downloading the GGUF.

Updated:
- Comments now show \${MODEL_DIR}/qwen3.6-27b-gguf/... as the placeholder
  (matches what the just-fixed README + LLAMA_CPP.md docs say).
- MODEL_PATH default changed from /mnt/models/huggingface/... to
  \${MODEL_DIR:-\$HOME/models}/qwen3.6-27b-gguf/... — falls back to
  ~/models/ if MODEL_DIR isn't set, which is a more reasonable default
  for cross-rig users than our /mnt/models/huggingface/ path.
- file-exists check at line 25 still fails loudly with the resolved path
  if neither MODEL_DIR nor MODEL_PATH is set correctly.

Follow-up to fbf3431 (de-bind \$MODEL_DIR from rig path in docs).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:14:21 +00:00
noonghunnaandClaude Opus 4.7 fbf343129c docs: use \$MODEL_DIR placeholder, not the dev rig's /mnt/models/huggingface/
User docs were hardcoding the dev rig path (/mnt/models/huggingface/...) as
if it was canonical. It's not — cross-rig users have models at /data/models,
~/models, /mnt/nvme/llms, etc. Setting MODEL_DIR per their setup is the
intended UX (the compose already supports it via env-var default).

Updates:
- models/qwen3.6-27b/llama-cpp/README.md: download examples now use
  \$MODEL_DIR/qwen3.6-27b-gguf/ instead of /mnt/models/huggingface/...
- models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml: header
  comment uses \$MODEL_DIR/qwen3.6-27b-gguf/ for download examples + says
  "MODEL_DIR=/your/models/dir docker compose up -d" instead of our path.
- docs/engines/LLAMA_CPP.md: same treatment + cleaned up Qwen3.5 + DFlash
  draft path examples to also use \$MODEL_DIR.
- scripts/preflight.sh: hf download hint shows literal \${MODEL_DIR} so user
  knows what to set, plus explicit "set MODEL_DIR first" line. Previously
  echoed the resolved relative path (../../../../models-cache) which lands
  outside the repo if pwd isn't the compose dir.

Caught by RobH589 in club-3090#116 — they hit the path-resolved-to-root-of-drive
case from the relative-path default. Closes the doc UX side; the compose's
env-var override mechanism was already correct.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:13:29 +00:00
noonghunnaandClaude Opus 4.7 cf7f1959fd fix: 4 stale refs missed in 2026-05-10 reorg push (caught by RobH589 #116)
After the GGUF dir move (/mnt/models/gguf/qwen3.6-27b/ → /mnt/models/
huggingface/qwen3.6-27b-gguf/), four refs were not updated and led to
a path-mismatch loop reported in club-3090#116:

- scripts/preflight.sh `hf download` hint pointed at qwen3.6-27b/, but
  the compose default expects qwen3.6-27b-gguf/. Same for the mv hint
  for mmproj relocation, and the in-container mmproj default at line 329.
- models/qwen3.6-27b/llama-cpp/README.md example command still said
  `MODEL_DIR=/mnt/models/gguf` (now /mnt/models/huggingface).
- models/qwen3.6-27b/llama-cpp/compose/single/{docker-compose,concurrent}.yml
  header comment said "Q5_K_XL" but the actual default has been Q3_K_XL
  for a while (this one predates the reorg — just stale doc).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:06:20 +00:00
32 changed files with 9023 additions and 471 deletions
+40 -3
View File
@@ -16,17 +16,54 @@ jobs:
uses: actions/checkout@v4
with:
fetch-depth: 0 # full history needed for git-cliff
token: ${{ secrets.GITHUB_TOKEN }}
- name: Generate release notes
id: cliff
# ---- (1) Generate the GitHub Release body (latest only, no header) ----
# `--strip header` drops the static SemVer preamble so the release page
# shows just the per-version section, while CHANGELOG.md (below) keeps
# the header at the top of the file.
- name: Generate release notes (latest only)
id: cliff-release
uses: orhun/git-cliff-action@v4
with:
config: cliff.toml
args: --latest --github-repo noonghunna/club-3090
args: --latest --strip header --github-repo noonghunna/club-3090
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
OUTPUT: RELEASE_NOTES.md
# ---- (2) Regenerate full CHANGELOG.md (all tags, with header) ----
- name: Regenerate CHANGELOG.md
uses: orhun/git-cliff-action@v4
with:
config: cliff.toml
args: --github-repo noonghunna/club-3090
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
OUTPUT: CHANGELOG.md
# ---- (3) Commit the regenerated CHANGELOG.md back to master ----
# The tag's commit doesn't include this auto-regen — but the GitHub
# Release page is correct (step 1), and master's CHANGELOG.md catches
# up ~1 min after tag push. Skipped if nothing changed.
- name: Commit CHANGELOG.md back to master
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git fetch origin master
git checkout -B master origin/master
# Re-run cliff against master HEAD so the regen reflects what master
# actually contains (the tag may not yet be on master if pushed from
# a feature branch; uncommon but handled).
git add CHANGELOG.md
if git diff --staged --quiet; then
echo "CHANGELOG.md unchanged — skipping commit."
else
git commit -m "chore(changelog): regenerate for ${{ github.ref_name }} [skip ci]"
git push origin master
fi
# ---- (4) Publish GitHub Release ----
- name: Create GitHub Release
uses: softprops/action-gh-release@v2
with:
+21
View File
@@ -244,3 +244,24 @@ Cross-rig data on Google's official Gemma 4 MTP "assistant" drafter (released 20
None close the **-13% narr / -11% code gap to 3dluvr's anchor**. Remaining gap likely rig-specific (3dluvr's EPYC 7J13 / different PCIe topology / 275W cap / etc) rather than tunable via flags. Cross-rig productionizable settings: stick with default `--dtype bfloat16`, default cudagraph, default scheduling. Custom override `cudagraph_capture_sizes [9]` worth it ONLY if you serve >95% n=8-MTP-single-stream code traffic (e.g. dedicated coding-agent endpoint) where the +2.5% code lift exceeds the -3% narrative loss. |
| `dual.yml`-shape forced TP=1 | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | 32K | **159.67 / 215.10** (decode 160.71 / 217.30) | 27.5 GB | 2026-05-07 | **First single-5090 Gemma 4 MTP data point.** First non-OOM single-card Gemma 4 result on the matrix — the 32 GB Blackwell envelope clears the 24 GB Ampere boot OOM. CV 1.9%/1.8%, peak 426 W. **+46% narr / +51% code over @noonghunna's 2× 3090 TP=2 baseline (109/142)** — single-card 5090 beats dual-3090 on Gemma 4. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
| `dual-dflash.yml`-shape forced TP=1 (mem-util 0.96, max-model-len 12000) | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | **12K** | **150.40 / 261.06** (decode 151.16 / 264.62) | 28.8 GB | 2026-05-07 | **First single-5090 Gemma 4 DFlash data point.** Trade vs MTP row above: ~6% narr loss, **+21% code lift** (215→261). 1st-warmup TTFT outlier (73 s) suggests cudagraph warmup taking longer on first request; subsequent warmups stable at <40 ms. CV 3.6%/2.8%, peak 440 W. **Required mem-util 0.96 + max-model-len 12K** to fit BF16 weights + DFlash N=5 drafter on 32 GB — DFlash drafter footprint pushes out ctx ceiling vs MTP's 32K. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
---
## Quality benches — Aider Polyglot 30
Pass rate on a curated 30-exercise subset of [aider-polyglot-benchmark](https://github.com/Aider-AI/polyglot-benchmark) (5 per language across cpp/go/java/javascript/python/rust, mix of easy/medium/hard). Tests **edit-format reliability** AND **algorithmic correctness** — does the model emit diffs aider can apply, AND do the resulting tests pass.
Run via [`benchlocal-cli`](https://github.com/noonghunna/benchlocal-cli) `aider-polyglot-30` pack. Different from the TPS rows above — this is a quality / agentic-coding signal, not a throughput measurement.
| Model | Compose | Rig | Pass / Total | % | Wall (real) | Wall (sum-dur) | Tokens (P+C) | Date | Notes |
|---|---|---|---:|---:|---:|---:|---:|---|---|
| **Qwen 3.6 27B** (AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | **20 / 30** | **66.7%** | 19.0 min | 34.0 min | 436K + 111K = 547K | 2026-05-10 | **`enable_thinking=false`** (server-side `--default-chat-template-kwargs '{"enable_thinking": false}'` + per-request `extra_body` belt). With thinking ON: 0/30 (1500s timeout exceeded before any exercise completed — Qwen burns the token budget on hidden CoT). Per-language: cpp 3/5 · go 4/5 · **java 4/5** · js 4/5 · python 2/5 · rust 3/5. `threads=2`. |
| **Gemma 4 31B** (Intel AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | 17 / 30 | 56.7% | 19.2 min | 19.2 min | 380K + 72K = 452K | 2026-05-10 | Default thinking off (Gemma 4's chat template requires explicit `enable_thinking=true` to enable). Per-language: cpp 2/5 · **go 4/5** · java 1/5 · **js 4/5** · python 3/5 · rust 3/5. `threads=2`. |
**Notable**:
- Qwen 3.6 27B beats Gemma 4 31B by **+10 pp** despite 4 GB fewer parameters. Java is the biggest swing (4/5 vs 1/5 — `affine-cipher` specifically tripped Gemma).
- Gemma is meaningfully faster wall-clock (sum-of-exercise-durations 19 vs 34 min), suggesting Qwen produces longer per-turn answers but they convert to passes more reliably.
- Qwen with thinking ON is unusable for this kind of bench on club-3090 hardware: hits the 1500s subprocess timeout cap (now bumped to 2700s) before completing any exercises — the hidden CoT eats the per-exercise token budget. **Set `enable_thinking=false` for any agentic / multi-turn workload.**
**To re-run cross-rig**: `bash scripts/quality-test.sh --pack aider-polyglot-30 --enable-sandboxed-packs` against your endpoint. See [docs/QUALITY_TEST.md](docs/QUALITY_TEST.md) for the harness setup.
+8747 -413
View File
File diff suppressed because it is too large Load Diff
+42 -20
View File
@@ -8,15 +8,41 @@
# → release published.
[changelog]
# Header rendered once at top of every release body.
# Header rendered once at top of CHANGELOG.md (preserved across full regens).
# Use `--strip header` on `--latest` for GitHub Release bodies so they don't
# duplicate this intro on every release page.
header = """
# Changelog
Auto-generated from commit messages by [git-cliff](https://git-cliff.org/).
Update flow: write rich commit message bodies → tag → CI regenerates this file
and the GitHub Release notes from the same source. Don't hand-edit below the
header — your changes will be overwritten on the next tag.
**Versioning:** SemVer in `0.x` — treat any minor bump as potentially breaking
until `1.0`. Past CalVer tags (`v2026.05.09`, `v2026.05.10`) are preserved for
history; SemVer takes over from `v0.3.0` onward.
| CalVer tag | SemVer equivalent | Date |
|---|---|---|
| `v2026.05.09` | (≈ v0.1.0) | 2026-05-09 — first tagged release |
| `v2026.05.10` | (≈ v0.2.0) | 2026-05-10 — stack reorg + Gemma 4 INT8 PTH unblock |
---
"""
# Body template — rendered per release (we use --latest so only one).
# Tera templating syntax. Each commit shows as a bullet with PR link if present.
# Body template — rendered per release. For `--latest` (GitHub Release) only
# the most recent block renders; for full regen of CHANGELOG.md, every tagged
# block renders in reverse-chronological order.
#
# Commit message body (everything after subject + blank line) renders below
# the subject bullet so rich narrative (tables, validation data, before/after
# numbers) ends up in both CHANGELOG.md and the GitHub Release page from the
# same source. Tera templating; see https://keats.github.io/tera/docs/.
body = """
{% if version %}\
## What's in {{ version }}
## {{ version }}{% if timestamp %} — {{ timestamp | date(format="%Y-%m-%d") }}{% endif %}
{% else %}\
## Unreleased
@@ -26,25 +52,21 @@ body = """
### {{ group }}
{% for commit in commits %}\
- {{ commit.message | split(pat="\\n") | first | trim }}{% if commit.github.pr_number %} ([#{{ commit.github.pr_number }}](https://github.com/noonghunna/club-3090/pull/{{ commit.github.pr_number }}) by @{{ commit.github.username }}){% else %} ([{{ commit.id | truncate(length=7, end="") }}](https://github.com/noonghunna/club-3090/commit/{{ commit.id }})){% endif %}
{% endfor %}
{% endfor %}
- **{{ commit.message | split(pat="\\n") | first | trim }}**{% if commit.github.pr_number %} ([#{{ commit.github.pr_number }}](https://github.com/noonghunna/club-3090/pull/{{ commit.github.pr_number }}) by @{{ commit.github.username }}){% else %} ([{{ commit.id | truncate(length=7, end="") }}](https://github.com/noonghunna/club-3090/commit/{{ commit.id }})){% endif %}
{% set body_lines = commit.message | split(pat="\\n") %}\
{% if body_lines | length > 1 %}\
{% set body = body_lines | slice(start=1) | join(sep="\\n") | trim %}\
{% if body %}
---
{{ body | replace(from="\\n", to="\\n ") }}
## Pinning to this release
This is a snapshot of the rolling stack — not a versioned API. To pin to this exact state:
```bash
git checkout {{ version }}
```
When posting cross-rig benchmark numbers ([disc #86](https://github.com/noonghunna/club-3090/discussions/86)), please include this version tag (or commit SHA) so others can reproduce against the same script revision.
{% if previous.version %}\
**Full diff:** [{{ previous.version }}...{{ version }}](https://github.com/noonghunna/club-3090/compare/{{ previous.version }}...{{ version }})
{% endif %}\
{% endif %}\
{% endfor %}
{% endfor %}
{% if version %}[Pin: `git checkout {{ version }}`]{% if previous.version %} · [Full diff](https://github.com/noonghunna/club-3090/compare/{{ previous.version }}...{{ version }}){% endif %}
{% endif %}
"""
footer = ""
+10 -10
View File
@@ -86,7 +86,7 @@ This is exactly why our launch frame is **two routes, not one** ([README](../../
```bash
# Use hf CLI (pip install 'huggingface-hub[hf_transfer]')
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
```
Confirm size matches the HuggingFace listing. If a `sha256` is published, verify it.
@@ -108,7 +108,7 @@ For a sane mid-context default (65K, plenty for chat + light agent work):
```bash
/opt/llama.cpp/build/bin/llama-server \
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-c 65536 \
--host 0.0.0.0 --port 8020 \
-ngl 999 \
@@ -131,7 +131,7 @@ Recipe (community-reported, validated by multiple users on r/LocalLLaMA):
```bash
/opt/llama.cpp/build/bin/llama-server \
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
-ngl 99 \
-c 262144 \
-np 1 \
@@ -154,12 +154,12 @@ Sustained throughput at 262K with this config is typically **35-45 tok/s** on a
Download the `mmproj` model:
```bash
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
```
Add to launch:
```bash
--mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
--mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
```
### 5. Tool calls (limited)
@@ -185,12 +185,12 @@ cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
# Download draft model (~500 MB)
hf download z-lab/Qwen3.6-27B-DFlash --local-dir /mnt/models/huggingface/z-lab/Qwen3.6-27B-DFlash/
hf download z-lab/Qwen3.6-27B-DFlash --local-dir $MODEL_DIR/z-lab/Qwen3.6-27B-DFlash/
# Launch
/opt/lucebox-hub/build/bin/llama-server \
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
--draft /mnt/models/gguf/qwen3.6-27b-dflash/dflash-N5.gguf \
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
--draft $MODEL_DIR/qwen3.6-27b-dflash-gguf/dflash-N5.gguf \
--draft-max 5 \
--draft-min 1 \
-c 65536 \
@@ -210,8 +210,8 @@ If you have two GPUs (e.g. 2× 3090), lucebox-hub now supports a heterogeneous-s
```bash
# Target on GPU 0, DFlash draft on GPU 1
/opt/lucebox-hub/build/bin/llama-server \
-m /mnt/models/gguf/qwen3.5-27b/Qwen3.5-27B-Q4_K_M.gguf \
--draft /mnt/models/gguf/qwen3.5-27b-dflash/dflash-N5.gguf \
-m $MODEL_DIR/qwen3.5-27b-gguf/Qwen3.5-27B-Q4_K_M.gguf \
--draft $MODEL_DIR/qwen3.5-27b-dflash-gguf/dflash-N5.gguf \
--target-gpu 0 --draft-gpu 1 \
--draft-max 16 --draft-min 1 \
-c 262144 \
+4 -4
View File
@@ -30,7 +30,7 @@ Showcase: full **262K context** on one 3090 with vision + q4_0 KV.
```bash
cd models/qwen3.6-27b/llama-cpp/compose
MODEL_DIR=/mnt/models/gguf docker compose up -d
MODEL_DIR=/your/models/dir docker compose up -d
```
Memory budget: 14.5 GB (Q3_K_XL) + 4.5 GB KV @ 262K + 0.8 GB mmproj ≈ 20 GB / 24 GB.
@@ -66,7 +66,7 @@ The Q3_K_XL number at 262K is **lower than community-reported 35-45 tok/s** ([Re
```bash
# 1. Get a GGUF quant (recommended: Unsloth's Q4_K_M)
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
# 2. Build llama.cpp with CUDA support
git clone https://github.com/ggerganov/llama.cpp /opt/llama.cpp
@@ -99,9 +99,9 @@ GGUFs of this model are at [unsloth/Qwen3.6-27B-GGUF](https://huggingface.co/uns
## Vision (mmproj)
```bash
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
# Add to launch: --mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
# Add to launch: --mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
```
Vision works via the mmproj model. Sample text+image queries are OpenAI-compat.
@@ -1,6 +1,6 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
# Engine: llama.cpp (NOT vLLM)
# Topology: Single 3090 (TP=1)
# Drafter: none
@@ -66,6 +66,7 @@ services:
--cont-batching
--jinja
--reasoning-format ${REASONING_FORMAT:-none}
--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}
deploy:
resources:
reservations:
@@ -1,6 +1,6 @@
# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
# Engine: llama.cpp (NOT vLLM — different engine, different memory model)
# Topology: Single 3090 (TP=1)
# Drafter: none (vanilla llama.cpp; MTP via PR #22673 not adopted yet)
@@ -43,15 +43,15 @@
# 1. Get the GGUF + mmproj:
# hf download unsloth/Qwen3.6-27B-GGUF \
# --include "Qwen3.6-27B-UD-Q3_K_XL.gguf" "mmproj-F16.gguf" \
# --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/unsloth-q3kxl
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-q3kxl
# (mmproj sometimes ships at the repo root rather than under unsloth-q3kxl;
# the MMPROJ env var below points at the canonical sidecar location.)
# 2. From this directory:
# MODEL_DIR=/mnt/models/huggingface docker compose up -d
# MODEL_DIR=/your/models/dir docker compose up -d
# 3. curl http://localhost:8020/v1/models → should list the model.
#
# Override defaults via .env or shell:
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /mnt/models/huggingface on this stack)
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /path/to/your/models on your stack)
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf)
# MMPROJ_FILE path under /models (default: qwen3.6-27b-gguf/mmproj-F16.gguf)
# CTX_SIZE total KV pool (default: 262144)
@@ -60,7 +60,7 @@
# Override to `auto` to get separate `reasoning_content` field
# (Qwen3.6 thinking trace) — useful for clients that render
# reasoning_content (most don't). Issue: club-3090#97.
# DISABLE_THINKING set to 1 in .env to add `--chat-template-kwargs '{"enable_thinking":false}'`
# DISABLE_THINKING default: 1 (thinking OFF). Set to 0 to opt INTO thinking
# which forces empty <think></think> blocks in responses. Useful for
# clients (e.g. opencode) that display <think> content as the response.
# Tradeoff: applies to ALL clients on this server — Hermes/agents that
@@ -92,7 +92,7 @@ services:
- -c
- |
set -e
# DISABLE_THINKING=1 in compose/.env appends --chat-template-kwargs to disable
# DISABLE_THINKING=1 (default) appends --chat-template-kwargs to disable
# Qwen3 thinking server-side. Forces the chat template to insert empty
# <think></think> blocks → output goes straight to the response. Useful for
# clients (e.g. opencode) that display <think> content as the response.
@@ -100,7 +100,7 @@ services:
# use thinking lose reasoning capability. See docs/HARDWARE.md and disc club-3090#97.
# Note: $$VAR is YAML-escape for $VAR (compose passes literal $ to bash).
EXTRA_ARGS=()
if [ "$${DISABLE_THINKING:-0}" = "1" ]; then
if [ "$${DISABLE_THINKING:-1}" = "1" ]; then
EXTRA_ARGS+=("--chat-template-kwargs" '{"enable_thinking":false}')
echo "[entrypoint] DISABLE_THINKING=1 — chat template will produce empty <think></think>"
fi
@@ -4,15 +4,15 @@
#
# Prereqs:
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/)
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir ${MODEL_DIR}/qwen3.6-27b-gguf/)
#
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX
set -euo pipefail
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
PORT="${PORT:-8020}"
CTX="${CTX:-65536}"
@@ -18,14 +18,14 @@
#
# Prereqs:
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
#
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX, KV_TYPE
set -euo pipefail
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
PORT="${PORT:-8020}"
CTX="${CTX:-262144}"
KV_TYPE="${KV_TYPE:-q4_0}"
@@ -127,6 +127,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_xml
@@ -118,6 +118,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -141,6 +141,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -130,6 +130,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -154,6 +154,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -149,6 +149,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -257,6 +257,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -139,6 +139,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -237,6 +237,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -134,6 +134,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -138,6 +138,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -247,6 +247,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -357,6 +357,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -374,6 +374,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -274,6 +274,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -97,6 +97,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
@@ -157,6 +157,8 @@ services:
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": false}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
+17 -1
View File
@@ -22,6 +22,7 @@
# sudo bash scripts/power-cap-sweep.sh --load-mode decode-concurrent --concurrency 8 --bench-runs 3
# sudo bash scripts/power-cap-sweep.sh --load-mode prefill-heavy
# sudo bash scripts/power-cap-sweep.sh --no-reset # leave at last cap (you reset manually)
# sudo bash scripts/power-cap-sweep.sh --include-commit # stamp club-3090 git short SHA in report header
#
# Load modes:
# decode-single:
@@ -157,6 +158,10 @@ PREFILL_FILLER_REPEATS=""
PREFILL_PROMPT_TOKENS=""
DECODE_CONCURRENT_RUN_SECONDS=""
CALIBRATION_NOTE=""
INCLUDE_COMMIT=0 # --include-commit stamps the club-3090 git short SHA in
# the report header. Off by default — `curl ... | bash`
# users have no clone, and stamping "n/a" is confusing
# (better to suppress the field entirely there).
while [ $# -gt 0 ]; do
case "$1" in
@@ -171,6 +176,7 @@ while [ $# -gt 0 ]; do
--load-target) LOAD_TARGET="$2"; shift 2 ;;
--concurrency-stretch) CONCURRENCY_STRETCH="$2"; shift 2 ;;
--target-cap-seconds) TARGET_CAP_SECONDS="$2"; shift 2 ;;
--include-commit) INCLUDE_COMMIT=1; shift ;;
--no-reset) RESET=0; shift ;;
-h|--help)
sed -n '1,/^set -euo/p' "$0" | grep '^#' | sed 's/^# \?//'
@@ -1120,7 +1126,17 @@ RESULTS_FILE=/tmp/power-cap-summary.md
echo "**Model:** \`${MODEL}\` &nbsp; **Engine:** \`${CONTAINER}\` &nbsp; **Endpoint:** ${URL}"
echo "**Load mode:** \`${LOAD_MODE}\`$([ "$LOAD_MODE" = "decode-single" ] && echo " (${TARGET_CAP_SECONDS}s × 2 timed streams)")$([ "$LOAD_MODE" = "decode-concurrent" ] && echo " (concurrency=${CONCURRENCY}, ${DECODE_CONCURRENT_RUN_SECONDS}s/run × ${BENCH_RUNS} runs × 2 timed batches)")$([ "$LOAD_MODE" = "prefill-heavy" ] && echo " (target-prefill=${TARGET_PREFILL_SECONDS}s, filler_repeats=${PREFILL_FILLER_REPEATS})")$([ "$LOAD_MODE" != "decode-single" ] && echo " (bench-runs=${BENCH_RUNS})")"
[ -n "$CALIBRATION_NOTE" ] && echo "**Calibration:** ${CALIBRATION_NOTE}"
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z"
# --include-commit: stamp club-3090 git short SHA next to the date if requested.
# Suppress entirely (rather than show "n/a") when run from a non-clone or
# when git isn't reachable — closes the curl-pipe-from-docs UX hole.
COMMIT_FRAGMENT=""
if [ "${INCLUDE_COMMIT:-0}" = "1" ]; then
COMMIT_SHA=$(git -C "$REPO_ROOT" rev-parse --short HEAD 2>/dev/null || true)
if [ -n "$COMMIT_SHA" ]; then
COMMIT_FRAGMENT=" &nbsp; **club-3090 commit:** \`${COMMIT_SHA}\`"
fi
fi
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z${COMMIT_FRAGMENT}"
echo ""
if [ "$COOLING" = "unspecified" ]; then
echo "> ⚠️ Cooling class not specified at run time. Add **air / water / AIO** when posting"
+5 -4
View File
@@ -325,8 +325,8 @@ preflight_compose_deps() {
mmproj_in_container="${mmproj_in_container//\$\{MMPROJ_FILE:-/}"
mmproj_in_container="${mmproj_in_container%\}}"
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b/mmproj-F16.gguf"
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b-gguf/mmproj-F16.gguf"
if [[ -n "${GGUF_FILE:-}" ]]; then gguf_in_container="$GGUF_FILE"; fi
if [[ -n "${MMPROJ_FILE:-}" ]]; then mmproj_in_container="$MMPROJ_FILE"; fi
@@ -375,9 +375,10 @@ preflight_compose_deps() {
if [[ $hint_gguf -eq 1 ]]; then
echo "[preflight] hf download unsloth/Qwen3.6-27B-GGUF \\" >&2
echo "[preflight] Qwen3.6-27B-UD-Q3_K_XL.gguf mmproj-F16.gguf \\" >&2
echo "[preflight] --local-dir ${model_dir}/qwen3.6-27b/unsloth-q3kxl" >&2
echo "[preflight] --local-dir \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl" >&2
echo "[preflight] # (set MODEL_DIR first: export MODEL_DIR=\${MODEL_DIR:-/path/to/your/models})" >&2
echo "[preflight] # mmproj lands at unsloth-q3kxl/ — move it up so the default --mmproj path resolves:" >&2
echo "[preflight] # mv ${model_dir}/qwen3.6-27b/unsloth-q3kxl/mmproj-F16.gguf ${model_dir}/qwen3.6-27b/" >&2
echo "[preflight] # mv \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl/mmproj-F16.gguf \${MODEL_DIR}/qwen3.6-27b-gguf/" >&2
echo "[preflight] (~16 GB total. setup.sh today only fetches the vLLM AutoRound weights;" >&2
echo "[preflight] GGUF must be fetched separately for any llamacpp/* variant.)" >&2
fi
+13
View File
@@ -177,6 +177,19 @@ if [[ -n "$DETECTED_MODEL" && "$DETECTED_MODEL" != "$MODEL" ]]; then
MODEL="$DETECTED_MODEL"
fi
# hermesagent-20 runs its agent inside a Docker sandbox container. Localhost-style
# URLs (localhost/127.x/[::1]) inside the container resolve to the container itself,
# not the host's vLLM. Auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 so benchlocal-cli
# (a) adds --add-host=host.docker.internal:host-gateway to the sandbox container, and
# (b) rewrites the model endpoint URL to use host.docker.internal:<port> for the
# hermes-agent's outbound API calls. Skip if already set (user override) or if URL
# already uses host.docker.internal / a non-loopback host (real LAN IP, k8s service).
if [[ -z "${BENCHLOCAL_HERMES_RESOLVE_LOCALHOST:-}" ]] \
&& [[ "$URL" =~ ^https?://(localhost|127\.|\[::1\]) ]]; then
export BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1
echo "[quality-test] localhost URL detected — auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for hermes sandbox endpoint rewrite" >&2
fi
# ---- run benchlocal-cli ------------------------------------------------------
RESULTS_DIR="${ROOT_DIR}/results/quality"
+67
View File
@@ -90,6 +90,73 @@ case "${MODEL_NAME}" in
esac
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
# ---------- MODEL_DIR resolution ----------
# Order of precedence:
# 1. MODEL_DIR already exported in the calling shell → use as-is
# 2. .env at repo root sets MODEL_DIR → source it
# 3. Interactive prompt (only if stdin is a TTY) → ask user
# 4. Silent fallback to <repo>/models-cache → in-repo default
#
# The prompt only fires for fresh users on a TTY who haven't set anything.
# CI / scripted runs (no TTY) get the silent fallback, preserving prior behavior.
# Step 2: source repo-root .env if present (lets a saved choice persist)
if [[ -z "${MODEL_DIR:-}" && -f "${ROOT_DIR}/.env" ]]; then
# shellcheck source=/dev/null
set -a; source "${ROOT_DIR}/.env"; set +a
fi
# Step 3: prompt if still unset + interactive
if [[ -z "${MODEL_DIR:-}" && -t 0 && -t 1 ]]; then
echo ""
echo "Where should I put model weights?"
echo " Models are large (Qwen3.6-27B AutoRound: ~14 GB; Gemma 4 31B: ~21 GB)."
echo " This dir lives outside the git tree — pick a location with sufficient free space."
echo ""
echo " 1) ${ROOT_DIR}/models-cache (in-repo, default — pollutes git tree)"
echo " 2) ${HOME}/models (recommended for cross-rig — outside repo)"
echo " 3) custom path"
echo ""
while true; do
read -rp "Choice [1-3] (or set MODEL_DIR env var to skip): " pick
case "${pick}" in
1) MODEL_DIR="${ROOT_DIR}/models-cache"; break ;;
2) MODEL_DIR="${HOME}/models"; break ;;
3)
read -rp " Enter absolute path: " custom
if [[ "${custom}" =~ ^/ ]]; then
MODEL_DIR="${custom}"; break
else
echo " ! must be an absolute path (start with /)" >&2
fi
;;
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
esac
done
echo ""
# Offer to persist the choice so future runs skip the prompt
read -rp "Save MODEL_DIR=${MODEL_DIR} to .env so we skip this next time? [Y/n]: " save
if [[ "${save:-y}" =~ ^[Yy]$ || -z "${save:-}" ]]; then
if [[ -f "${ROOT_DIR}/.env" ]]; then
# Update existing .env (replace MODEL_DIR= line if present, else append)
if grep -qE "^MODEL_DIR=" "${ROOT_DIR}/.env"; then
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=${MODEL_DIR}|" "${ROOT_DIR}/.env"
else
echo "MODEL_DIR=${MODEL_DIR}" >> "${ROOT_DIR}/.env"
fi
else
echo "MODEL_DIR=${MODEL_DIR}" > "${ROOT_DIR}/.env"
fi
echo " → saved. (.env is gitignored.)"
else
echo " → not saved. Set MODEL_DIR=... when re-running, or you'll get this prompt again."
fi
echo ""
fi
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
+9 -3
View File
@@ -557,15 +557,21 @@ def cmd_run(endpoint, req_path, timeout_s, metrics_path):
choices = chunk.get("choices") or []
if choices:
delta = choices[0].get("delta") or {}
if ttft is None and (delta.get("content") or delta.get("reasoning_content") or delta.get("tool_calls")):
# vLLM emits reasoning under either `delta.reasoning_content`
# (older qwen3 reasoner path) or `delta.reasoning` (current
# nightly as of vllm-0.20.2rc1+; legacy field name). Watch
# both so the soak harness doesn't go silent when the
# underlying field name shifts under us.
reasoning_delta = delta.get("reasoning_content") or delta.get("reasoning")
if ttft is None and (delta.get("content") or reasoning_delta or delta.get("tool_calls")):
ttft = time.time() - t0
# Accumulate streamed parts. vLLM splits content/reasoning
# across many small deltas; tool_calls stream as indexed
# objects whose fields (name, arguments) arrive in pieces.
if delta.get("content"):
content_parts.append(delta["content"])
if delta.get("reasoning_content"):
reasoning_parts.append(delta["reasoning_content"])
if reasoning_delta:
reasoning_parts.append(reasoning_delta)
for tc in (delta.get("tool_calls") or []):
idx = tc.get("index", 0)
slot = tool_calls_acc.setdefault(idx, {"id": "", "type": "function", "name": "", "args": ""})