Files
club-3090/tools/grammar-eval/README.md
T
7be8ecc9e0 feat(grammar-eval): land harness for Holiday tagline grammar A/B
Codex-implemented harness from docs/diagnostics/grammar-eval-codex-brief.md
(gitignored). Sets up the A/B test of Holiday_Purpose_3166's tagline
grammar (Reddit r/LocalLLaMA 1sx7w55) against our shipped
andthattoo/structured-cot GOAL/APPROACH/EDGE grammar.

The hypothesis: Holiday's K/R free-token-list fields are a pressure-relief
valve that GOAL/APPROACH/EDGE's rigid 3-line shape lacks. We have 6
documented HE+ regression cases (HE/97, 101, 108, 129, 137, 151) where
FSM under-thinks and FREE wins; if Holiday's grammar rescues some
without losing too much compression, it's a Pareto improvement worth
shipping.

Files:
  - tools/grammar-eval/holiday-tagline.gbnf — translated grammar
    (Holiday's bounded-repetition `[A-Za-z][A-Za-z0-9_.!-]{0,18}` rewritten
    as explicit nullable tail rules for xgrammar compatibility; opening
    `<think>\n` removed since Qwen3.6 chat template prefixes it; closing
    `</think>\n\n` preserved verbatim).
  - tools/grammar-eval/TRANSLATION.md — translation decisions + Phase-1
    smoke results (5/5 tagline prompts PASS shape regex against
    vllm/bounded-thinking; FREE failures are existing max_tokens trap, not
    grammar issues).
  - tools/grammar-eval/smoke-test.py — Phase-1 runner. Validates grammar
    parses + applies on vLLM/xgrammar before committing compute to bench.
  - tools/grammar-eval/subset-bench.py — Phase-2 30-prompt HE+ A/B harness
    (6 FSM-regress + 4 FREE-regress + 20 random; runs FREE / current /
    Holiday / PROMPT_TERSE; outputs results.jsonl + summary.md).
  - tools/grammar-eval/README.md — overview + run instructions.

docs/STRUCTURED_COT.md — extended "FSM-regress cases are real" section
with note on this active experiment, scope, and gating decision (≥3 of
6 rescues = phase 3; 0-2 rescues = drop).

Phase-1 smoke validation cross-rig (this rig, 2026-05-03 PM):
  vllm/bounded-thinking + Qwen3.6-27B AutoRound INT4 + vLLM
  0.20.1rc1.dev16 + Genesis 2db18df/v7.69. All 5 tagline prompts return
  successfully + match BODY_RE regex. Translation is valid xgrammar
  syntax; ready for phase-2 bench whenever compute is available.

Phase 2/3 deferred — bench takes ~30-60 min (subset) / ~6-8h (full HE+ +
LCB v6). User asked to ship harness now and run bench later.

Co-Authored-By: Holiday_Purpose_3166 <[email protected]>
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-03 16:01:36 +00:00

2.6 KiB
Raw Blame History

Grammar evaluation harness

Tools to A/B-test alternative bounded-thinking grammars against the currently-shipped andthattoo/structured-cot GOAL/APPROACH/EDGE grammar (in docs/STRUCTURED_COT.md). Specifically targeting Holiday_Purpose_3166's tagline grammar from r/LocalLLaMA.

The hypothesis: the K (keywords, 1-5 free tokens) and R (result-keywords, 1-5 free tokens) fields in Holiday's grammar give the model a pressure-relief valve that GOAL/APPROACH/EDGE's rigid 3-line shape lacks. We have 6 documented HE+ regression cases (HE/97, 101, 108, 129, 137, 151) where FSM under-thinks and FREE wins; if Holiday's grammar rescues even some of those without losing too much compression, it's a Pareto improvement worth shipping.

Files

File Purpose
holiday-tagline.gbnf Translated grammar, xgrammar-compatible
TRANSLATION.md Translation decisions + Phase-1 smoke results
smoke-test.py Phase-1 — verify the grammar parses + applies on vLLM
subset-bench.py Phase-2 — 30-prompt HE+ A/B vs current grammar + FREE + PROMPT_TERSE

How to run

Phase 1 — smoke (~3 min):

python3 tools/grammar-eval/smoke-test.py --base-url http://localhost:8020/v1
# or with auto-boot:
python3 tools/grammar-eval/smoke-test.py --boot

5 prompts × tagline-grammar. Pass criterion: all 5 return successfully + match the BODY_RE regex. Validated 2026-05-03 PM — see TRANSLATION.md.

Phase 2 — 30-prompt subset bench (~30-60 min):

/tmp/structured-cot-venv/bin/python tools/grammar-eval/subset-bench.py \
  --base-url http://localhost:8020/v1

Subset = 6 known FSM-regress problems + 4 known FREE-regress + 20 random HE+. Compares: FREE vs current GOAL/APPROACH/EDGE vs Holiday tagline vs PROMPT_TERSE. Outputs results/grammar-ab-<timestamp>/{results.jsonl, summary.md}.

Headline question phase 2 answers: of the 6 known HE+ FSM regressions, how many does Holiday's grammar rescue? 0 = drop the experiment. 3+ = phase 3 (full HE+ 164 + LCB v6 50) is worth running.

Phase 3 — full bench (only if phase 2 shows signal):

Re-run the same harness against the full HumanEval+ 164 + LiveCodeBench v6 50. ~6-8 hours of compute. Updates the table in docs/STRUCTURED_COT.md.

Background

The brief that produced these tools is at docs/diagnostics/grammar-eval-codex-brief.md (gitignored). The companion LLM-agnostic design exercise prompt (for refining beyond Holiday's design) is at docs/diagnostics/grammar-design-llm-prompt.md (gitignored).