5 Commits
Author SHA1 Message Date
noonghunna 255c743dff chore: trigger v0.3.2 release workflow (GitHub deduped previous tag push)
Release / release (push) Failing after 48s
The v0.3.2 tag was originally pushed at commit 64b0474 but GitHub didn't
emit a CreateEvent (likely dedup after delete + re-push to same SHA), so
the release.yml workflow never fired. Empty commit gives the tag a fresh
SHA that GitHub will process cleanly.
2026-05-10 21:08:33 +00:00
noonghunna 64b0474a62 chore(changelog): automate CHANGELOG + release notes from commits via cliff (Option A)
CHANGELOG.md is now auto-generated from commit messages by git-cliff in
the release workflow. Hand-edits below the static header will be wiped on
the next tag.

Workflow (`.github/workflows/release.yml`):
  - On tag push (`v[0-9]+.[0-9]+.[0-9]+`):
    1. Render GitHub Release body: `git-cliff --latest --strip header`
       → just the per-version section, no SemVer preamble repeat
    2. Regenerate full CHANGELOG.md: `git-cliff` (default = all tags)
       → preserves header + all historical sections
    3. Commit CHANGELOG.md back to master with `[skip ci]` marker
    4. Publish GitHub Release with the latest-only body

Template (`cliff.toml`):
  - `[changelog].header` now holds the SemVer preamble + CalVer→SemVer
    mapping table (preserved across regens; stripped from GitHub Release
    bodies via `--strip header`).
  - `body` template now renders the **full commit message** (subject as
    bold bullet, body indented below) instead of just the first line.
    Rich narrative I write in commit message bodies (tables, validation
    numbers, before/after diffs) now flows into both CHANGELOG.md and the
    GitHub Release page from the same source.
  - Per-release Pin/Diff footer guarded with `{% if version %}` so the
    Unreleased section doesn't emit empty links.

CHANGELOG.md replaced with the auto-gen output. Past hand-written tables
and phase breakdowns are replaced by the corresponding commit messages
(those were already rich for commits that mattered — v0.3.1 soak-helper
fix has its Before/After table in the commit body and renders fine).

Going forward: just write rich commit messages and tag. Both surfaces
update automatically. No hand-edit of CHANGELOG.md required.
2026-05-10 20:53:51 +00:00
noonghunna 83bf73d3ec feat(quality-test): auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for localhost URLs
When the user runs quality-test.sh with a localhost-style URL
(default `http://localhost:8020`, or any `localhost`/`127.x`/`[::1]`
variant), auto-export `BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1` so
benchlocal-cli rewrites the hermes-agent's outbound model endpoint
from `localhost:<port>` to `host.docker.internal:<port>` inside the
Docker sandbox container.

Without this, the hermes-agent inside the sandbox can't reach the
host's vLLM (localhost resolves to the container itself) and every
scenario fails with `"API call failed after 3 retries: Connection
error."` — produced spurious 0/20 grades on this rig prior to the
benchlocal-cli runner.py 9c1566f fix.

Skips the auto-set when:
  - User already set the env var (explicit override)
  - URL points at a non-loopback host (real LAN IP, k8s service name,
    host.docker.internal already) — no rewrite needed

Emits a stderr breadcrumb when the auto-set fires so users can see
what changed.
2026-05-10 20:29:49 +00:00
noonghunna 9db8b2603c docs(changelog): v0.3.1 entry for soak-helper delta.reasoning capture
Release / release (push) Failing after 46s
Documents the silent-empty turn-5 root cause (vLLM nightly field-name
shift to delta.reasoning) + validation soak results.
2026-05-10 19:16:44 +00:00
noonghunna 88eb67aa18 fix(soak-helper): capture delta.reasoning alongside delta.reasoning_content
vLLM nightly (0.20.2rc1.dev9+) emits the qwen3 reasoning parser's output
under `delta.reasoning` (legacy field name), not `delta.reasoning_content`
that soak-helper.py was watching. Result: for any thinking-on response
whose `<think>` block doesn't close within `max_tokens`, soak-helper saw
zero deltas → fell back to the "couldn't measure" path → reported
`ttft_ms == t_ms` and `decode_tps = 0.0`. The model was generating
correctly; the harness just couldn't see the wire output.

Repro request (JDWarner's #107 turn 5): math problem with
`max_tokens=2000` + `chat_template_kwargs.enable_thinking=true`.

Before patch:
  status=200  t_ms=22709  ttft_ms=22709  decode_tps=0.0
  completion_tokens=2000  content=""  reasoning_content=""

After patch (same request, same compose, same model):
  status=200  t_ms=22709  ttft_ms=234  decode_tps=88.985
  completion_tokens=2000  content=""  reasoning_content="Here's a thinking
  process:\n\n1. **Understand the User's Problem:**\n..."  (3959 chars)

Validation soak (fresh-mode, 20 sessions × 5 turns = 100 turns, qwen3.6-27b
dual.yml):
  verdict        PASS
  silent_empty   0 / 100 (0.0%)   ← was ~3-5/40 baseline
  p50_decode_tps 90.22
  p95_ttft_ms    1389
  errors         0
  max_growth     0 MiB / 200

Closes the cross-rig "silent-empty turn-5" pattern parked behind the
Cliff 2b investigation — it was a harness measurement bug, not a model
or rig issue.
2026-05-10 19:15:57 +00:00
5 changed files with 8807 additions and 442 deletions
+40 -3
View File
@@ -16,17 +16,54 @@ jobs:
uses: actions/checkout@v4
with:
fetch-depth: 0 # full history needed for git-cliff
token: ${{ secrets.GITHUB_TOKEN }}
- name: Generate release notes
id: cliff
# ---- (1) Generate the GitHub Release body (latest only, no header) ----
# `--strip header` drops the static SemVer preamble so the release page
# shows just the per-version section, while CHANGELOG.md (below) keeps
# the header at the top of the file.
- name: Generate release notes (latest only)
id: cliff-release
uses: orhun/git-cliff-action@v4
with:
config: cliff.toml
args: --latest --github-repo noonghunna/club-3090
args: --latest --strip header --github-repo noonghunna/club-3090
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
OUTPUT: RELEASE_NOTES.md
# ---- (2) Regenerate full CHANGELOG.md (all tags, with header) ----
- name: Regenerate CHANGELOG.md
uses: orhun/git-cliff-action@v4
with:
config: cliff.toml
args: --github-repo noonghunna/club-3090
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
OUTPUT: CHANGELOG.md
# ---- (3) Commit the regenerated CHANGELOG.md back to master ----
# The tag's commit doesn't include this auto-regen — but the GitHub
# Release page is correct (step 1), and master's CHANGELOG.md catches
# up ~1 min after tag push. Skipped if nothing changed.
- name: Commit CHANGELOG.md back to master
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git fetch origin master
git checkout -B master origin/master
# Re-run cliff against master HEAD so the regen reflects what master
# actually contains (the tag may not yet be on master if pushed from
# a feature branch; uncommon but handled).
git add CHANGELOG.md
if git diff --staged --quiet; then
echo "CHANGELOG.md unchanged — skipping commit."
else
git commit -m "chore(changelog): regenerate for ${{ github.ref_name }} [skip ci]"
git push origin master
fi
# ---- (4) Publish GitHub Release ----
- name: Create GitHub Release
uses: softprops/action-gh-release@v2
with:
+8703 -416
View File
File diff suppressed because it is too large Load Diff
+42 -20
View File
@@ -8,15 +8,41 @@
# → release published.
[changelog]
# Header rendered once at top of every release body.
# Header rendered once at top of CHANGELOG.md (preserved across full regens).
# Use `--strip header` on `--latest` for GitHub Release bodies so they don't
# duplicate this intro on every release page.
header = """
# Changelog
Auto-generated from commit messages by [git-cliff](https://git-cliff.org/).
Update flow: write rich commit message bodies → tag → CI regenerates this file
and the GitHub Release notes from the same source. Don't hand-edit below the
header — your changes will be overwritten on the next tag.
**Versioning:** SemVer in `0.x` — treat any minor bump as potentially breaking
until `1.0`. Past CalVer tags (`v2026.05.09`, `v2026.05.10`) are preserved for
history; SemVer takes over from `v0.3.0` onward.
| CalVer tag | SemVer equivalent | Date |
|---|---|---|
| `v2026.05.09` | (≈ v0.1.0) | 2026-05-09 — first tagged release |
| `v2026.05.10` | (≈ v0.2.0) | 2026-05-10 — stack reorg + Gemma 4 INT8 PTH unblock |
---
"""
# Body template — rendered per release (we use --latest so only one).
# Tera templating syntax. Each commit shows as a bullet with PR link if present.
# Body template — rendered per release. For `--latest` (GitHub Release) only
# the most recent block renders; for full regen of CHANGELOG.md, every tagged
# block renders in reverse-chronological order.
#
# Commit message body (everything after subject + blank line) renders below
# the subject bullet so rich narrative (tables, validation data, before/after
# numbers) ends up in both CHANGELOG.md and the GitHub Release page from the
# same source. Tera templating; see https://keats.github.io/tera/docs/.
body = """
{% if version %}\
## What's in {{ version }}
## {{ version }}{% if timestamp %} — {{ timestamp | date(format="%Y-%m-%d") }}{% endif %}
{% else %}\
## Unreleased
@@ -26,25 +52,21 @@ body = """
### {{ group }}
{% for commit in commits %}\
- {{ commit.message | split(pat="\\n") | first | trim }}{% if commit.github.pr_number %} ([#{{ commit.github.pr_number }}](https://github.com/noonghunna/club-3090/pull/{{ commit.github.pr_number }}) by @{{ commit.github.username }}){% else %} ([{{ commit.id | truncate(length=7, end="") }}](https://github.com/noonghunna/club-3090/commit/{{ commit.id }})){% endif %}
{% endfor %}
{% endfor %}
- **{{ commit.message | split(pat="\\n") | first | trim }}**{% if commit.github.pr_number %} ([#{{ commit.github.pr_number }}](https://github.com/noonghunna/club-3090/pull/{{ commit.github.pr_number }}) by @{{ commit.github.username }}){% else %} ([{{ commit.id | truncate(length=7, end="") }}](https://github.com/noonghunna/club-3090/commit/{{ commit.id }})){% endif %}
{% set body_lines = commit.message | split(pat="\\n") %}\
{% if body_lines | length > 1 %}\
{% set body = body_lines | slice(start=1) | join(sep="\\n") | trim %}\
{% if body %}
---
{{ body | replace(from="\\n", to="\\n ") }}
## Pinning to this release
**Versioning:** SemVer in `0.x` — treat any minor bump as potentially breaking until `1.0`. Past CalVer tags (`v2026.05.09`, `v2026.05.10`) are preserved for history; SemVer takes over from `v0.3.0` onward.
```bash
git checkout {{ version }}
```
When posting cross-rig benchmark numbers ([disc #86](https://github.com/noonghunna/club-3090/discussions/86)), please include this version tag (or commit SHA) so others can reproduce against the same script revision.
{% if previous.version %}\
**Full diff:** [{{ previous.version }}...{{ version }}](https://github.com/noonghunna/club-3090/compare/{{ previous.version }}...{{ version }})
{% endif %}\
{% endif %}\
{% endfor %}
{% endfor %}
{% if version %}[Pin: `git checkout {{ version }}`]{% if previous.version %} · [Full diff](https://github.com/noonghunna/club-3090/compare/{{ previous.version }}...{{ version }}){% endif %}
{% endif %}
"""
footer = ""
+13
View File
@@ -177,6 +177,19 @@ if [[ -n "$DETECTED_MODEL" && "$DETECTED_MODEL" != "$MODEL" ]]; then
MODEL="$DETECTED_MODEL"
fi
# hermesagent-20 runs its agent inside a Docker sandbox container. Localhost-style
# URLs (localhost/127.x/[::1]) inside the container resolve to the container itself,
# not the host's vLLM. Auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 so benchlocal-cli
# (a) adds --add-host=host.docker.internal:host-gateway to the sandbox container, and
# (b) rewrites the model endpoint URL to use host.docker.internal:<port> for the
# hermes-agent's outbound API calls. Skip if already set (user override) or if URL
# already uses host.docker.internal / a non-loopback host (real LAN IP, k8s service).
if [[ -z "${BENCHLOCAL_HERMES_RESOLVE_LOCALHOST:-}" ]] \
&& [[ "$URL" =~ ^https?://(localhost|127\.|\[::1\]) ]]; then
export BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1
echo "[quality-test] localhost URL detected — auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for hermes sandbox endpoint rewrite" >&2
fi
# ---- run benchlocal-cli ------------------------------------------------------
RESULTS_DIR="${ROOT_DIR}/results/quality"
+9 -3
View File
@@ -557,15 +557,21 @@ def cmd_run(endpoint, req_path, timeout_s, metrics_path):
choices = chunk.get("choices") or []
if choices:
delta = choices[0].get("delta") or {}
if ttft is None and (delta.get("content") or delta.get("reasoning_content") or delta.get("tool_calls")):
# vLLM emits reasoning under either `delta.reasoning_content`
# (older qwen3 reasoner path) or `delta.reasoning` (current
# nightly as of vllm-0.20.2rc1+; legacy field name). Watch
# both so the soak harness doesn't go silent when the
# underlying field name shifts under us.
reasoning_delta = delta.get("reasoning_content") or delta.get("reasoning")
if ttft is None and (delta.get("content") or reasoning_delta or delta.get("tool_calls")):
ttft = time.time() - t0
# Accumulate streamed parts. vLLM splits content/reasoning
# across many small deltas; tool_calls stream as indexed
# objects whose fields (name, arguments) arrive in pieces.
if delta.get("content"):
content_parts.append(delta["content"])
if delta.get("reasoning_content"):
reasoning_parts.append(delta["reasoning_content"])
if reasoning_delta:
reasoning_parts.append(reasoning_delta)
for tc in (delta.get("tool_calls") or []):
idx = tc.get("index", 0)
slot = tool_calls_acc.setdefault(idx, {"id": "", "type": "function", "name": "", "args": ""})