Files
club-3090/tools/grammar-eval
noonghunnaandClaude Opus 4.7 d82e89807a chore(results): commit grammar bench evidence + gitignore investigation artifacts (#217)
The results/ directory accumulated 21 MB of intermediate investigation
output during today's cliff/grammar work. Most of that (residency probes,
soak iteration matrices) lived its useful life in issue threads and the
gitignored docs/diagnostics/ memos, and isn't worth eternal evidence.

What's evidence (committed): 1.5 MB of raw bench output backing published
numbers in STRUCTURED_COT.md / BENCHMARKS.md / CHANGELOG:

- results/grammar-ab-20260503-224235/ — Phase 2 30-problem subset bench
  (was the basis for the n=30 PROMPT_TERSE-wins finding that Phase 3
  later disproved at scale)
- results/grammar-full-20260504-003118-gpu0/ — Phase 3 HE+ shard 1 (82 rows)
- results/grammar-full-20260504-003118-gpu1/ — Phase 3 HE+ shard 2 (82 rows)
- results/grammar-full-lcb-20260504-021003/ — Phase 3 LCB v6 (50 rows + shard
  metadata; this is the post-bug-fix re-run after the harness LCB issue
  surfaced and was patched)
- tools/grammar-eval/codex-check-scratchpad.gbnf — fourth grammar candidate
  exploring DeepSeek + STATE/CHECK pre-VERDICT scaffolding (untested in
  Phase 3; available for future bench)

What's investigation artifact (now gitignored): residency-* (Codex Cliff 2b
investigation probes, ~5 MB) + soak-* (~16 MB of cliff probe iterations,
final validation matrix already in #41 + cliff2_accumulated_ctx_finding.md
memory).

.gitignore pattern: `results/*` with explicit negations for `grammar-ab-*`,
`grammar-full-*`, and the existing `v0.20-migration` directory. Future
grammar bench runs auto-track when committed; future residency/soak runs
auto-ignore. The pattern works because git's directory-ignore precedence
applies to immediate children, not the parent itself, so negations
re-include named subdirs.

Total committed: 522 lines of bench output (jsonl + json + summary md).
The data is small (per-row generations cap at ~4096 tokens × 5 conditions
× 214 problems = ~4 MB raw, much of which compresses well in jsonl).

Held back for user decision: tools/residency-instrument/ (Codex's Cliff 2b
instrumentation harness — useful tool, no docs yet, may want a README
before publication).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-04 13:24:06 +00:00
..

Grammar evaluation harness

Tools to A/B-test alternative bounded-thinking grammars against the currently-shipped andthattoo/structured-cot GOAL/APPROACH/EDGE grammar (in docs/STRUCTURED_COT.md). Specifically targeting Holiday_Purpose_3166's tagline grammar from r/LocalLLaMA.

The hypothesis: the K (keywords, 1-5 free tokens) and R (result-keywords, 1-5 free tokens) fields in Holiday's grammar give the model a pressure-relief valve that GOAL/APPROACH/EDGE's rigid 3-line shape lacks. We have 6 documented HE+ regression cases (HE/97, 101, 108, 129, 137, 151) where FSM under-thinks and FREE wins; if Holiday's grammar rescues even some of those without losing too much compression, it's a Pareto improvement worth shipping.

Files

File Purpose
holiday-tagline.gbnf Translated grammar, xgrammar-compatible
TRANSLATION.md Translation decisions + Phase-1 smoke results
smoke-test.py Phase-1 — verify the grammar parses + applies on vLLM
subset-bench.py Phase-2 — 30-prompt HE+ A/B vs current grammar + FREE + PROMPT_TERSE

How to run

Phase 1 — smoke (~3 min):

python3 tools/grammar-eval/smoke-test.py --base-url http://localhost:8020/v1
# or with auto-boot:
python3 tools/grammar-eval/smoke-test.py --boot

5 prompts × tagline-grammar. Pass criterion: all 5 return successfully + match the BODY_RE regex. Validated 2026-05-03 PM — see TRANSLATION.md.

Phase 2 — 30-prompt subset bench (~30-60 min):

/tmp/structured-cot-venv/bin/python tools/grammar-eval/subset-bench.py \
  --base-url http://localhost:8020/v1

Subset = 6 known FSM-regress problems + 4 known FREE-regress + 20 random HE+. Compares: FREE vs current GOAL/APPROACH/EDGE vs Holiday tagline vs PROMPT_TERSE. Outputs results/grammar-ab-<timestamp>/{results.jsonl, summary.md}.

Headline question phase 2 answers: of the 6 known HE+ FSM regressions, how many does Holiday's grammar rescue? 0 = drop the experiment. 3+ = phase 3 (full HE+ 164 + LCB v6 50) is worth running.

Phase 3 — full bench (only if phase 2 shows signal):

Re-run the same harness against the full HumanEval+ 164 + LiveCodeBench v6 50. ~6-8 hours of compute. Updates the table in docs/STRUCTURED_COT.md.

Background

The brief that produced these tools is at docs/diagnostics/grammar-eval-codex-brief.md (gitignored). The companion LLM-agnostic design exercise prompt (for refining beyond Holiday's design) is at docs/diagnostics/grammar-design-llm-prompt.md (gitignored).