2 Commits
Author SHA1 Message Date
noonghunna 9db8b2603c docs(changelog): v0.3.1 entry for soak-helper delta.reasoning capture
Release / release (push) Failing after 46s
Documents the silent-empty turn-5 root cause (vLLM nightly field-name
shift to delta.reasoning) + validation soak results.
2026-05-10 19:16:44 +00:00
noonghunna 88eb67aa18 fix(soak-helper): capture delta.reasoning alongside delta.reasoning_content
vLLM nightly (0.20.2rc1.dev9+) emits the qwen3 reasoning parser's output
under `delta.reasoning` (legacy field name), not `delta.reasoning_content`
that soak-helper.py was watching. Result: for any thinking-on response
whose `<think>` block doesn't close within `max_tokens`, soak-helper saw
zero deltas → fell back to the "couldn't measure" path → reported
`ttft_ms == t_ms` and `decode_tps = 0.0`. The model was generating
correctly; the harness just couldn't see the wire output.

Repro request (JDWarner's #107 turn 5): math problem with
`max_tokens=2000` + `chat_template_kwargs.enable_thinking=true`.

Before patch:
  status=200  t_ms=22709  ttft_ms=22709  decode_tps=0.0
  completion_tokens=2000  content=""  reasoning_content=""

After patch (same request, same compose, same model):
  status=200  t_ms=22709  ttft_ms=234  decode_tps=88.985
  completion_tokens=2000  content=""  reasoning_content="Here's a thinking
  process:\n\n1. **Understand the User's Problem:**\n..."  (3959 chars)

Validation soak (fresh-mode, 20 sessions × 5 turns = 100 turns, qwen3.6-27b
dual.yml):
  verdict        PASS
  silent_empty   0 / 100 (0.0%)   ← was ~3-5/40 baseline
  p50_decode_tps 90.22
  p95_ttft_ms    1389
  errors         0
  max_growth     0 MiB / 200

Closes the cross-rig "silent-empty turn-5" pattern parked behind the
Cliff 2b investigation — it was a harness measurement bug, not a model
or rig issue.
2026-05-10 19:15:57 +00:00
2 changed files with 37 additions and 3 deletions
+28
View File
@@ -15,6 +15,34 @@ Past CalVer tags (preserved for history):
CHANGELOG entries below 2026-05-10 use date-prose headings (pre-SemVer convention). New entries from `v0.3.0` use `## v0.X.Y (YYYY-MM-DD) — title`.
## v0.3.1 (2026-05-10) — soak-helper captures `delta.reasoning` (closes cross-rig silent-empty turn-5 mystery)
One-line patch behind a meaningful finding: vLLM nightly (`0.20.2rc1.dev9+`) emits the qwen3 reasoning parser's output under `delta.reasoning` (legacy field name), not `delta.reasoning_content` that `soak-helper.py` was watching. Result: any thinking-on response whose `<think>` block didn't close within `max_tokens` showed up as a silent-empty turn (`ttft_ms == t_ms`, `decode_tps = 0.0`, empty `content` + empty `reasoning_content`) — the model was generating correctly, the harness just couldn't see the wire.
This closes the cross-rig "silent-empty turn 5" pattern that was parked behind the Cliff 2b investigation. It was a harness measurement bug, not a model/rig issue.
Validation on JDWarner's exact repro request (#107 turn 5, math problem + `enable_thinking=true` + `max_tokens=2000`):
| | Before | After |
|---|---|---|
| `ttft_ms` | 22709 (== `t_ms`) | **234** |
| `decode_tps` | **0.0** | **88.985** |
| `reasoning_content` chars | 0 | 3959 |
| `completion_tokens` | 2000 | 2000 (unchanged) |
Validation soak (`qwen3.6-27b dual.yml`, fresh-mode, 20 sessions × 5 turns):
```
verdict PASS
silent_empty 0 / 100 (0.0%) ← was ~3-5/40 baseline
p50_decode_tps 90.22
p95_ttft_ms 1389
errors 0
max_growth_mib 0 / 200
```
**Patch** ([88eb67a](https://github.com/noonghunna/club-3090/commit/88eb67a)): `soak-helper.py` accumulates `delta.reasoning_content || delta.reasoning` (covers both vLLM streaming-output shapes, current and legacy). Three-line change inside `cmd_run`'s SSE chunk loop.
## v0.3.0 (2026-05-10) — Qwen thinking-off default, MODEL_DIR UX overhaul, docs cleanup, aider-polyglot bench rows
Eight commits since `v2026.05.10`. Headline: Qwen 3.6 27B now defaults `enable_thinking=false` server-side via `--default-chat-template-kwargs` across all 21 vLLM Qwen composes (closes the recurring "Qwen looks much slower than Gemma" agentic-bench skew). MODEL_DIR docs no longer bake the dev-rig path. New `--include-commit` flag on `power-cap-sweep.sh`.
+9 -3
View File
@@ -557,15 +557,21 @@ def cmd_run(endpoint, req_path, timeout_s, metrics_path):
choices = chunk.get("choices") or []
if choices:
delta = choices[0].get("delta") or {}
if ttft is None and (delta.get("content") or delta.get("reasoning_content") or delta.get("tool_calls")):
# vLLM emits reasoning under either `delta.reasoning_content`
# (older qwen3 reasoner path) or `delta.reasoning` (current
# nightly as of vllm-0.20.2rc1+; legacy field name). Watch
# both so the soak harness doesn't go silent when the
# underlying field name shifts under us.
reasoning_delta = delta.get("reasoning_content") or delta.get("reasoning")
if ttft is None and (delta.get("content") or reasoning_delta or delta.get("tool_calls")):
ttft = time.time() - t0
# Accumulate streamed parts. vLLM splits content/reasoning
# across many small deltas; tool_calls stream as indexed
# objects whose fields (name, arguments) arrive in pieces.
if delta.get("content"):
content_parts.append(delta["content"])
if delta.get("reasoning_content"):
reasoning_parts.append(delta["reasoning_content"])
if reasoning_delta:
reasoning_parts.append(reasoning_delta)
for tc in (delta.get("tool_calls") or []):
idx = tc.get("index", 0)
slot = tool_calls_acc.setdefault(idx, {"id": "", "type": "function", "name": "", "args": ""})