The results/ directory accumulated 21 MB of intermediate investigation output during today's cliff/grammar work. Most of that (residency probes, soak iteration matrices) lived its useful life in issue threads and the gitignored docs/diagnostics/ memos, and isn't worth eternal evidence. What's evidence (committed): 1.5 MB of raw bench output backing published numbers in STRUCTURED_COT.md / BENCHMARKS.md / CHANGELOG: - results/grammar-ab-20260503-224235/ — Phase 2 30-problem subset bench (was the basis for the n=30 PROMPT_TERSE-wins finding that Phase 3 later disproved at scale) - results/grammar-full-20260504-003118-gpu0/ — Phase 3 HE+ shard 1 (82 rows) - results/grammar-full-20260504-003118-gpu1/ — Phase 3 HE+ shard 2 (82 rows) - results/grammar-full-lcb-20260504-021003/ — Phase 3 LCB v6 (50 rows + shard metadata; this is the post-bug-fix re-run after the harness LCB issue surfaced and was patched) - tools/grammar-eval/codex-check-scratchpad.gbnf — fourth grammar candidate exploring DeepSeek + STATE/CHECK pre-VERDICT scaffolding (untested in Phase 3; available for future bench) What's investigation artifact (now gitignored): residency-* (Codex Cliff 2b investigation probes, ~5 MB) + soak-* (~16 MB of cliff probe iterations, final validation matrix already in #41 + cliff2_accumulated_ctx_finding.md memory). .gitignore pattern: `results/*` with explicit negations for `grammar-ab-*`, `grammar-full-*`, and the existing `v0.20-migration` directory. Future grammar bench runs auto-track when committed; future residency/soak runs auto-ignore. The pattern works because git's directory-ignore precedence applies to immediate children, not the parent itself, so negations re-include named subdirs. Total committed: 522 lines of bench output (jsonl + json + summary md). The data is small (per-row generations cap at ~4096 tokens × 5 conditions × 214 problems = ~4 MB raw, much of which compresses well in jsonl). Held back for user decision: tools/residency-instrument/ (Codex's Cliff 2b instrumentation harness — useful tool, no docs yet, may want a README before publication). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Grammar evaluation harness
Tools to A/B-test alternative bounded-thinking grammars against the currently-shipped andthattoo/structured-cot GOAL/APPROACH/EDGE grammar (in docs/STRUCTURED_COT.md). Specifically targeting Holiday_Purpose_3166's tagline grammar from r/LocalLLaMA.
The hypothesis: the K (keywords, 1-5 free tokens) and R (result-keywords, 1-5 free tokens) fields in Holiday's grammar give the model a pressure-relief valve that GOAL/APPROACH/EDGE's rigid 3-line shape lacks. We have 6 documented HE+ regression cases (HE/97, 101, 108, 129, 137, 151) where FSM under-thinks and FREE wins; if Holiday's grammar rescues even some of those without losing too much compression, it's a Pareto improvement worth shipping.
Files
| File | Purpose |
|---|---|
holiday-tagline.gbnf |
Translated grammar, xgrammar-compatible |
TRANSLATION.md |
Translation decisions + Phase-1 smoke results |
smoke-test.py |
Phase-1 — verify the grammar parses + applies on vLLM |
subset-bench.py |
Phase-2 — 30-prompt HE+ A/B vs current grammar + FREE + PROMPT_TERSE |
How to run
Phase 1 — smoke (~3 min):
python3 tools/grammar-eval/smoke-test.py --base-url http://localhost:8020/v1
# or with auto-boot:
python3 tools/grammar-eval/smoke-test.py --boot
5 prompts × tagline-grammar. Pass criterion: all 5 return successfully + match the BODY_RE regex. Validated 2026-05-03 PM — see TRANSLATION.md.
Phase 2 — 30-prompt subset bench (~30-60 min):
/tmp/structured-cot-venv/bin/python tools/grammar-eval/subset-bench.py \
--base-url http://localhost:8020/v1
Subset = 6 known FSM-regress problems + 4 known FREE-regress + 20 random HE+. Compares: FREE vs current GOAL/APPROACH/EDGE vs Holiday tagline vs PROMPT_TERSE. Outputs results/grammar-ab-<timestamp>/{results.jsonl, summary.md}.
Headline question phase 2 answers: of the 6 known HE+ FSM regressions, how many does Holiday's grammar rescue? 0 = drop the experiment. 3+ = phase 3 (full HE+ 164 + LCB v6 50) is worth running.
Phase 3 — full bench (only if phase 2 shows signal):
Re-run the same harness against the full HumanEval+ 164 + LiveCodeBench v6 50. ~6-8 hours of compute. Updates the table in docs/STRUCTURED_COT.md.
Background
The brief that produced these tools is at docs/diagnostics/grammar-eval-codex-brief.md (gitignored). The companion LLM-agnostic design exercise prompt (for refining beyond Holiday's design) is at docs/diagnostics/grammar-design-llm-prompt.md (gitignored).