The compose headers were stale (long-text said 218K + 0.985, long-vision
192K + 0.98, bounded-thinking 218K) and didn't reference v7.64 or the
new patch stack. Updated all three to reflect the shipped 185K + 0.975
(text) / 140K + 0.95 (vision) configs and link to docs/CLIFFS.md "Update
2026-05-01 PM" for the bisection rationale.
Each compose header now also documents the recommended client max_tokens
default:
- long-text + long-vision (FREE thinking): 8192 (16384 for hard
reasoning / competition problems). 4096 was the trap that bit our
LCB v6 baseline mid-think.
- bounded-thinking (FSM grammar caps think): 4096 is sufficient — the
grammar bounds think to ~150-300 structured tokens.
docs/EXAMPLES.md gets a new top-of-doc "max_tokens defaults" table so
copy-paste users land on the right number without reading the bench
forensics. Two existing examples bumped: math reasoning 400 → 2048 (easy
math but FREE thinking can run that), Quicksort code 800 → 4096 (code
gen with FREE thinking traps at 800).
Smoke-test "Capital of France" examples kept at max_tokens=200 — that's
the documented intentional headroom for thinking + short answer.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>