Commit Graph
14 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.8 3878b24849 quality-test: add --no-thinking (symmetric force-off for reasoning A/B)
The wrapper exposed --enable-thinking (force thinking on for every pack)
but had no force-off counterpart, so a clean all-off arm of a reasoning
A/B wasn't reachable through the wrapper — only the mixed per-pack
default. Add --no-thinking / NO_THINKING=1, forwarded to benchlocal-cli
--no-thinking, mutually exclusive with --enable-thinking. Keeps the
wrapper's localhost hermes-resolve + timeout sizing on both arms.

Validated live driving the DiffusionGemma 8-pack off/on A/B.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-11 12:16:41 +05:00
noonghunnaandClaude Opus 4.8 77c1cd6c01 docs(quality-test): document failure-reason reading; fix drifted Failure-breakdown sample
The 'Failure breakdown:' block printed at the end of every run already shows
failure_mode + full detail per failed scenario — the quickest read. Add a
'Diagnosing failures' section (breakdown first, then JSON / benchlocal-cli
inspect for full trace / older runs / filter / diff) and a quality-test.sh
footer pointing at it. Fix the QUALITY_TEST.md sample, which showed a grouped
format + invented failure modes (missing_field/wrong_value) that the tool
never emits, to the real flat per-line format + real failure_mode vocabulary.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-05-29 01:40:37 +00:00
f1fe9205f0 quality-test.sh: forward --progress to benchlocal-cli, default ON (#248)
The benchlocal-cli `--progress` flag emits per-scenario `[N/M]` lines
to stderr — critical for the long-running modes (--full ~30-40 min,
--reasoning ~similar, --pack aider-polyglot-30 ~25-30 min) which
otherwise go dark for the whole duration. The wrapper never forwarded
the flag, so users (including agents) ran blind unless they bypassed
the wrapper.

Make --progress on by default, --no-progress as the opt-out for CI /
log-volume-sensitive contexts. Mirror PROGRESS=0/1 env var. Document
the convention in AGENTS.md so agents see it as policy, not just
private/personal memory.

Convention surfaced after wasting 25 min on a blind --full run during
autonomous validation: the wrapper logged its own progress line and
nothing else for the entire 30+ min window.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-28 20:38:57 +05:00
8b1d00d110 quality-test.sh: stop hardcoding --timeout-per-case 60 (benchlocal-cli #41) (#245)
The wrapper was always passing `--timeout-per-case 60` to benchlocal-cli,
which under benchlocal-cli #41 (per-pack `timeout_per_case_default`
metadata) would override the new agentic-pack defaults back to 60s and
defeat the fix — cli-40 / hermesagent-20 / aider-polyglot-30 would
still hit budget timeouts.

Now: only pass `--timeout-per-case` to benchlocal-cli when the user
explicitly sets `TIMEOUT_PER_CASE=...` env or `--timeout-per-case N`
CLI flag. When unset, benchlocal-cli applies its per-pack metadata
defaults (60s deterministic / 300s cli-40+hermes / 1800s aider).

Backward compat: against an OLD benchlocal-cli (pre-#41), the CLI flag's
own default is 60.0 — so unsetting the wrapper override gives the same
behavior as before. Against the NEW benchlocal-cli, the per-pack default
applies. Either way the wrapper does the right thing.

Pairs with benchlocal-cli PR https://github.com/noonghunna/benchlocal-cli/pull/43.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-28 15:57:26 +05:00
noonghunna 605f1df52e Document reasoning quality suite 2026-05-24 12:00:11 +00:00
noonghunna caf6fc2b72 Expose benchlocal reasoning suite 2026-05-24 08:30:07 +00:00
noonghunnaandnoonghunna 6291ce5081 feat(eval): expose request-level thinking toggles (#196)
Closes #195.

Add ENABLE_THINKING support to bench.sh and pass --enable-thinking / --thinking-max-tokens through quality-test.sh. Propagate the same env through rebench-full.sh, warn on likely reasoning-on server/request-off mismatches, and document the workflow for reasoning-model evals.

Co-authored-by: noonghunna <[email protected]>
2026-05-23 08:24:58 +05:00
noonghunna dd1f070626 feat(scripts): pass --sampling-from-server through quality-test.sh + rebench-full.sh
Companion to benchlocal-cli #22: quality-test.sh gains --sampling-from-server
flag (also settable via SAMPLING_FROM_SERVER=1 env), forwarded to benchlocal-cli.
rebench-full.sh propagates the env var to both quality-test invocations.

Use case: when the compose encodes the model's recommended sampling
(e.g. Qwopus temp=0.8), run eval with the serve-side defaults instead
of the hardcoded temp=0 baseline.

bash scripts/quality-test.sh --sampling-from-server
SAMPLING_FROM_SERVER=1 bash scripts/rebench-full.sh
2026-05-23 01:23:10 +00:00
82e3e50284 fix(rebench): always capture sandboxed-pack logs to the per-tag results dir (#179)
rebench-full now passes --sandbox-log-dir "$OUT_DIR" to its quality-full
(step 3) and aider-polyglot-30 (step 5) quality-test.sh calls, and
quality-test.sh forwards it to benchlocal-cli. Sandboxed packs (bugfind-15,
hermesagent-20, cli-40, aider-polyglot-30) now write sandbox-<pack>.log into
results/rebench/<tag>/ before container teardown — previously those container
logs were discarded on cleanup, so a surprising agentic result (e.g. the
aider 15/30 batch) had no inspectable trace. Also exposed standalone via
`quality-test.sh --sandbox-log-dir DIR` / SANDBOX_LOG_DIR env.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-22 20:50:51 +05:00
f2c4b111f5 quality-test: respect explicit MODEL/--model, don't clobber from /v1/models (#177)
Fixes the model-override footgun on llama-swap / multi-model endpoints
(reported by @ampersandru in disc #152). quality-test.sh unconditionally
overrode MODEL with the first id from /v1/models whenever they differed —
fine for single-model composes (the wrong-name → HTTP 404 fix), but on
llama-swap /v1/models returns the first registered model (often the wrong
one), so the whole run got routed at the wrong model. @ampersandru was
manually `sed`-ing out the override line every run as a workaround.

New behaviour:
- MODEL unset      → auto-detect from /v1/models (unchanged; keeps the
                     single-model footgun fix)
- MODEL set        → respect it verbatim, never override; warn once if the
                     endpoint disagrees ("using YOUR value")
- New --model NAME → equivalent to MODEL env, sets the explicit flag

Tracks whether MODEL was explicitly provided (env or --model) before the
default is applied, and gates the auto-detect on that.

Tested via a mock /v1/models endpoint returning a mismatched model:
  A. MODEL unset           → auto-detects the served id          ✓
  B. --model 'Qwen3.6-ik'  → respects it, warns on mismatch      ✓
  C. MODEL=… env           → same respect behaviour              ✓

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-21 08:27:43 +05:00
31948858c5 quality-test: expose --timeout-per-case + bump aider-polyglot-30 to 3600s (#175)
Surfaced by ampersandru's external eval in #152 — running aider-polyglot-30
on a single 3090 at 250W power cap hit the previous default timeout
mid-batch (29/30 exercises completed in <45 min, the 30th got killed
during a long edit). benchlocal-cli already accepts --timeout-per-case
but quality-test.sh wasn't exposing it as a flag, only as an env var.

Changes:

* quality-test.sh: add --timeout-per-case N CLI flag with validation;
  overrides the TIMEOUT_PER_CASE env var when both are set. Update
  usage + examples.
* rebench-full.sh: pass --timeout-per-case 3600 (1 hour) by default for
  the aider-polyglot-30 step. Override via AIDER_TIMEOUT_PER_CASE env.
  Bumped from benchlocal-cli's lower default to give power-capped /
  single-card rigs enough budget to finish the full 30-exercise batch.

Both fixes are scoped to club-3090. The deeper benchlocal-cli UX issues
(default per-case timeout too low; "0/1 fail" CLI headline hiding
partial-results JSON when agent_runner_timeout fires mid-batch) are
filed separately upstream.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-21 05:23:33 +05:00
noonghunna 83bf73d3ec feat(quality-test): auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for localhost URLs
When the user runs quality-test.sh with a localhost-style URL
(default `http://localhost:8020`, or any `localhost`/`127.x`/`[::1]`
variant), auto-export `BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1` so
benchlocal-cli rewrites the hermes-agent's outbound model endpoint
from `localhost:<port>` to `host.docker.internal:<port>` inside the
Docker sandbox container.

Without this, the hermes-agent inside the sandbox can't reach the
host's vLLM (localhost resolves to the container itself) and every
scenario fails with `"API call failed after 3 retries: Connection
error."` — produced spurious 0/20 grades on this rig prior to the
benchlocal-cli runner.py 9c1566f fix.

Skips the auto-set when:
  - User already set the env var (explicit override)
  - URL points at a non-loopback host (real LAN IP, k8s service name,
    host.docker.internal already) — no rewrite needed

Emits a stderr breadcrumb when the auto-set fires so users can see
what changed.
2026-05-10 20:29:49 +00:00
noonghunna 7020d965bd quality-test.sh: --sandboxed-only passthrough 2026-05-10 00:09:41 +00:00
noonghunnaandClaude Opus 4.7 1be02d2271 quality-test.sh: --help, --pack passthrough, align with benchlocal-cli v0.5
- Add comprehensive --help with mode descriptions + examples.
- Add --pack PACK_ID passthrough (run a single named pack, overrides mode).
- Add --no-sandboxed opt-out for --full (mirrors benchlocal-cli flag).
- Add --list-packs convenience flag.
- Drop ENABLE_SANDBOXED env var — no longer needed since benchlocal-cli's
  --full now defaults to sandboxed packs.
- Refresh mode descriptions: medium=5 packs (incl reasonmath now), full=8
  packs requiring Docker.

Pairs with benchlocal-cli v0.5.0 (commit eb7ddb0).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-09 22:23:22 +00:00