Files
club-3090/tools/grammar-eval/TRANSLATION.md
T
7be8ecc9e0 feat(grammar-eval): land harness for Holiday tagline grammar A/B
Codex-implemented harness from docs/diagnostics/grammar-eval-codex-brief.md
(gitignored). Sets up the A/B test of Holiday_Purpose_3166's tagline
grammar (Reddit r/LocalLLaMA 1sx7w55) against our shipped
andthattoo/structured-cot GOAL/APPROACH/EDGE grammar.

The hypothesis: Holiday's K/R free-token-list fields are a pressure-relief
valve that GOAL/APPROACH/EDGE's rigid 3-line shape lacks. We have 6
documented HE+ regression cases (HE/97, 101, 108, 129, 137, 151) where
FSM under-thinks and FREE wins; if Holiday's grammar rescues some
without losing too much compression, it's a Pareto improvement worth
shipping.

Files:
  - tools/grammar-eval/holiday-tagline.gbnf — translated grammar
    (Holiday's bounded-repetition `[A-Za-z][A-Za-z0-9_.!-]{0,18}` rewritten
    as explicit nullable tail rules for xgrammar compatibility; opening
    `<think>\n` removed since Qwen3.6 chat template prefixes it; closing
    `</think>\n\n` preserved verbatim).
  - tools/grammar-eval/TRANSLATION.md — translation decisions + Phase-1
    smoke results (5/5 tagline prompts PASS shape regex against
    vllm/bounded-thinking; FREE failures are existing max_tokens trap, not
    grammar issues).
  - tools/grammar-eval/smoke-test.py — Phase-1 runner. Validates grammar
    parses + applies on vLLM/xgrammar before committing compute to bench.
  - tools/grammar-eval/subset-bench.py — Phase-2 30-prompt HE+ A/B harness
    (6 FSM-regress + 4 FREE-regress + 20 random; runs FREE / current /
    Holiday / PROMPT_TERSE; outputs results.jsonl + summary.md).
  - tools/grammar-eval/README.md — overview + run instructions.

docs/STRUCTURED_COT.md — extended "FSM-regress cases are real" section
with note on this active experiment, scope, and gating decision (≥3 of
6 rescues = phase 3; 0-2 rescues = drop).

Phase-1 smoke validation cross-rig (this rig, 2026-05-03 PM):
  vllm/bounded-thinking + Qwen3.6-27B AutoRound INT4 + vLLM
  0.20.1rc1.dev16 + Genesis 2db18df/v7.69. All 5 tagline prompts return
  successfully + match BODY_RE regex. Translation is valid xgrammar
  syntax; ready for phase-2 bench whenever compute is available.

Phase 2/3 deferred — bench takes ~30-60 min (subset) / ~6-8h (full HE+ +
LCB v6). User asked to ship harness now and run bench later.

Co-Authored-By: Holiday_Purpose_3166 <[email protected]>
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-03 16:01:36 +00:00

4.0 KiB

Holiday tagline grammar translation for vLLM/xgrammar

This translates Holiday_Purpose_3166's Reddit GBNF grammar into the form used by club-3090's bounded-thinking vLLM stack.

The translated grammar lives at:

tools/grammar-eval/holiday-tagline.gbnf

Translated grammar

root ::= "Q=" q "\n" "M=" m "\n" "K=" toks "\n" "R=" toks "\n" "V=" v "\n" "</think>\n\n" out

q ::= "solve" | "prove" | "route" | "debug" | "patch" | "code" | "calc" | "compare" | "explain"
m ::= "case" | "enum" | "check" | "derive" | "edit" | "test" | "trace" | "rank"
v ::= "ok" | "fail" | "done" | "blocked" | "candidate" | "verify"

toks ::= tok | tok "," tok | tok "," tok "," tok | tok "," tok "," tok "," tok | tok "," tok "," tok "," tok "," tok
tok ::= alpha tail18
alpha ::= [A-Za-z]
tok_char ::= [A-Za-z0-9_.!-]

tail18 ::= "" | tok_char tail17
tail17 ::= "" | tok_char tail16
tail16 ::= "" | tok_char tail15
tail15 ::= "" | tok_char tail14
tail14 ::= "" | tok_char tail13
tail13 ::= "" | tok_char tail12
tail12 ::= "" | tok_char tail11
tail11 ::= "" | tok_char tail10
tail10 ::= "" | tok_char tail9
tail9 ::= "" | tok_char tail8
tail8 ::= "" | tok_char tail7
tail7 ::= "" | tok_char tail6
tail6 ::= "" | tok_char tail5
tail5 ::= "" | tok_char tail4
tail4 ::= "" | tok_char tail3
tail3 ::= "" | tok_char tail2
tail2 ::= "" | tok_char tail1
tail1 ::= "" | tok_char tail0
tail0 ::= ""

out ::= [\x09\x0A\x0D\x20-\x7E]+

Translation decisions

  • Opening <think>\n removed: Qwen3.6's chat template already prefixes the assistant reasoning channel with <think>\n. The grammar therefore starts at Q=, matching the existing fsm_grammar_no_open.gbnf port.

  • Closing </think>\n\n preserved verbatim: the grammar still forces the model to terminate the reasoning block and then emit normal answer text. vLLM's Qwen reasoning parser may expose only the body as reasoning_content; the smoke test normalizes this before regex validation.

  • tok bounded repetition unrolled: Holiday's original tok ::= [A-Za-z][A-Za-z0-9_.!-]{0,18} was rewritten as alpha tail18 with explicit nullable tail rules. This preserves the same 1-19 character token length without relying on {0,18} support.

  • Keyword lists preserved: q, m, v, and the 1-5 comma-separated toks alternatives are unchanged except for formatting.

  • ASCII answer region preserved: out ::= [\x09\x0A\x0D\x20-\x7E]+ matches the current structured-CoT grammar's permissive code/output region. It allows tabs, newlines, carriage returns, spaces, and printable ASCII.

Validation status

The local host environment does not have xgrammar installed, so static parser validation must happen through vLLM. Use:

python3 tools/grammar-eval/smoke-test.py --boot

The smoke test sends the grammar through structured_outputs: {"grammar": ...} and verifies that generated reasoning matches:

^Q=\w+\nM=\w+\nK=[\w,.!-]+\nR=[\w,.!-]+\nV=\w+$

Any 4xx/5xx response or shape mismatch should be treated as a translation failure before running the 30-prompt subset bench.

Phase-1 smoke results — 2026-05-03 PM

Ran against vllm/bounded-thinking on Qwen3.6-27B AutoRound INT4, vLLM 0.20.1rc1.dev16+g7a1eb8ac2, Genesis 2db18df/v7.69, RTX 3090. All five tagline-grammar prompts passed with shape-matching output:

[grammar-smoke] PASS simple-math/tagline: ok
[grammar-smoke] PASS simple-code/tagline: ok
[grammar-smoke] PASS simple-debug/tagline: ok
[grammar-smoke] PASS fallback-hard-1/tagline: ok
[grammar-smoke] PASS fallback-hard-2/tagline: ok

Three FREE-baseline prompts (run alongside for sanity comparison) failed with empty answer content — that's the existing max_tokens=4096 truncation trap documented in docs/STRUCTURED_COT.md "Honest caveats" section, not a grammar issue. Tagline grammar's structure caps think tokens, dodges the trap.

Translation status: validated. Phase 2 (30-prompt HumanEval+ subset bench) is unblocked.