The wrapper exposed --enable-thinking (force thinking on for every pack)
but had no force-off counterpart, so a clean all-off arm of a reasoning
A/B wasn't reachable through the wrapper — only the mixed per-pack
default. Add --no-thinking / NO_THINKING=1, forwarded to benchlocal-cli
--no-thinking, mutually exclusive with --enable-thinking. Keeps the
wrapper's localhost hermes-resolve + timeout sizing on both arms.
Validated live driving the DiffusionGemma 8-pack off/on A/B.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
The 'Failure breakdown:' block printed at the end of every run already shows
failure_mode + full detail per failed scenario — the quickest read. Add a
'Diagnosing failures' section (breakdown first, then JSON / benchlocal-cli
inspect for full trace / older runs / filter / diff) and a quality-test.sh
footer pointing at it. Fix the QUALITY_TEST.md sample, which showed a grouped
format + invented failure modes (missing_field/wrong_value) that the tool
never emits, to the real flat per-line format + real failure_mode vocabulary.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The benchlocal-cli `--progress` flag emits per-scenario `[N/M]` lines
to stderr — critical for the long-running modes (--full ~30-40 min,
--reasoning ~similar, --pack aider-polyglot-30 ~25-30 min) which
otherwise go dark for the whole duration. The wrapper never forwarded
the flag, so users (including agents) ran blind unless they bypassed
the wrapper.
Make --progress on by default, --no-progress as the opt-out for CI /
log-volume-sensitive contexts. Mirror PROGRESS=0/1 env var. Document
the convention in AGENTS.md so agents see it as policy, not just
private/personal memory.
Convention surfaced after wasting 25 min on a blind --full run during
autonomous validation: the wrapper logged its own progress line and
nothing else for the entire 30+ min window.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
The wrapper was always passing `--timeout-per-case 60` to benchlocal-cli,
which under benchlocal-cli #41 (per-pack `timeout_per_case_default`
metadata) would override the new agentic-pack defaults back to 60s and
defeat the fix — cli-40 / hermesagent-20 / aider-polyglot-30 would
still hit budget timeouts.
Now: only pass `--timeout-per-case` to benchlocal-cli when the user
explicitly sets `TIMEOUT_PER_CASE=...` env or `--timeout-per-case N`
CLI flag. When unset, benchlocal-cli applies its per-pack metadata
defaults (60s deterministic / 300s cli-40+hermes / 1800s aider).
Backward compat: against an OLD benchlocal-cli (pre-#41), the CLI flag's
own default is 60.0 — so unsetting the wrapper override gives the same
behavior as before. Against the NEW benchlocal-cli, the per-pack default
applies. Either way the wrapper does the right thing.
Pairs with benchlocal-cli PR https://github.com/noonghunna/benchlocal-cli/pull/43.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Closes#195.
Add ENABLE_THINKING support to bench.sh and pass --enable-thinking / --thinking-max-tokens through quality-test.sh. Propagate the same env through rebench-full.sh, warn on likely reasoning-on server/request-off mismatches, and document the workflow for reasoning-model evals.
Co-authored-by: noonghunna <[email protected]>
Companion to benchlocal-cli #22: quality-test.sh gains --sampling-from-server
flag (also settable via SAMPLING_FROM_SERVER=1 env), forwarded to benchlocal-cli.
rebench-full.sh propagates the env var to both quality-test invocations.
Use case: when the compose encodes the model's recommended sampling
(e.g. Qwopus temp=0.8), run eval with the serve-side defaults instead
of the hardcoded temp=0 baseline.
bash scripts/quality-test.sh --sampling-from-server
SAMPLING_FROM_SERVER=1 bash scripts/rebench-full.sh
rebench-full now passes --sandbox-log-dir "$OUT_DIR" to its quality-full
(step 3) and aider-polyglot-30 (step 5) quality-test.sh calls, and
quality-test.sh forwards it to benchlocal-cli. Sandboxed packs (bugfind-15,
hermesagent-20, cli-40, aider-polyglot-30) now write sandbox-<pack>.log into
results/rebench/<tag>/ before container teardown — previously those container
logs were discarded on cleanup, so a surprising agentic result (e.g. the
aider 15/30 batch) had no inspectable trace. Also exposed standalone via
`quality-test.sh --sandbox-log-dir DIR` / SANDBOX_LOG_DIR env.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Fixes the model-override footgun on llama-swap / multi-model endpoints
(reported by @ampersandru in disc #152). quality-test.sh unconditionally
overrode MODEL with the first id from /v1/models whenever they differed —
fine for single-model composes (the wrong-name → HTTP 404 fix), but on
llama-swap /v1/models returns the first registered model (often the wrong
one), so the whole run got routed at the wrong model. @ampersandru was
manually `sed`-ing out the override line every run as a workaround.
New behaviour:
- MODEL unset → auto-detect from /v1/models (unchanged; keeps the
single-model footgun fix)
- MODEL set → respect it verbatim, never override; warn once if the
endpoint disagrees ("using YOUR value")
- New --model NAME → equivalent to MODEL env, sets the explicit flag
Tracks whether MODEL was explicitly provided (env or --model) before the
default is applied, and gates the auto-detect on that.
Tested via a mock /v1/models endpoint returning a mismatched model:
A. MODEL unset → auto-detects the served id ✓
B. --model 'Qwen3.6-ik' → respects it, warns on mismatch ✓
C. MODEL=… env → same respect behaviour ✓
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Surfaced by ampersandru's external eval in #152 — running aider-polyglot-30
on a single 3090 at 250W power cap hit the previous default timeout
mid-batch (29/30 exercises completed in <45 min, the 30th got killed
during a long edit). benchlocal-cli already accepts --timeout-per-case
but quality-test.sh wasn't exposing it as a flag, only as an env var.
Changes:
* quality-test.sh: add --timeout-per-case N CLI flag with validation;
overrides the TIMEOUT_PER_CASE env var when both are set. Update
usage + examples.
* rebench-full.sh: pass --timeout-per-case 3600 (1 hour) by default for
the aider-polyglot-30 step. Override via AIDER_TIMEOUT_PER_CASE env.
Bumped from benchlocal-cli's lower default to give power-capped /
single-card rigs enough budget to finish the full 30-exercise batch.
Both fixes are scoped to club-3090. The deeper benchlocal-cli UX issues
(default per-case timeout too low; "0/1 fail" CLI headline hiding
partial-results JSON when agent_runner_timeout fires mid-batch) are
filed separately upstream.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
When the user runs quality-test.sh with a localhost-style URL
(default `http://localhost:8020`, or any `localhost`/`127.x`/`[::1]`
variant), auto-export `BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1` so
benchlocal-cli rewrites the hermes-agent's outbound model endpoint
from `localhost:<port>` to `host.docker.internal:<port>` inside the
Docker sandbox container.
Without this, the hermes-agent inside the sandbox can't reach the
host's vLLM (localhost resolves to the container itself) and every
scenario fails with `"API call failed after 3 retries: Connection
error."` — produced spurious 0/20 grades on this rig prior to the
benchlocal-cli runner.py 9c1566f fix.
Skips the auto-set when:
- User already set the env var (explicit override)
- URL points at a non-loopback host (real LAN IP, k8s service name,
host.docker.internal already) — no rewrite needed
Emits a stderr breadcrumb when the auto-set fires so users can see
what changed.