Files
club-3090/scripts
5b3793b85e feat(kv-calc): opt-in --kv-breakdown architecture cache planning layer (#213)
Additive + reporting-only. Default text and --json output are byte-identical
unless --kv-breakdown is passed; predict()/raw_verdict()/--calibration remain
the calibrated fit authority. Adds --gpus N (informational), a conflict-guarded
--sequences alias, and opt-in compressed-KV / indexer-cache / draft-KV estimate
buckets scoped to single-node 1-8 GPU home/workstation planning.

Implemented via Codex collab; independently validated on-rig by Claude:
- default output IDENTICAL (HEAD vs change) across text / --json / --solve-max-ctx
- --calibration 22/22, both kv-calc test scripts pass, py_compile clean, diff --check clean
- 4 error cases exit 2 (sequences/max-num-seqs conflict, --gpus bounds,
  indexer/compressed missing head-dim)

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-24 08:57:05 +05:00
..
…
…
…