The producer could write a `boot-fit-measured` record with null decode
TPS when bench.sh output drifted (summary block absent/unparseable) or a
metric was missing. A measured record with null TPS is worse than no
record for optimizer calibration: it looks like real data.
For a measured result_class:
- raise MeasuredRecordError if no parseable bench summary block was found
(decode_TPS mean= absent => output drift), or if the parse produced no
decode TPS. This matches the module's existing fail-loud posture
(KeyError on an unknown registry tag).
- a malformed/absent `=== GPU state ===` line is a SOFT gap (VRAM is a
fingerprint extension, not the core measured TPS): surface it in a new
top-level `parse_warnings` list instead of raising, so the gap is
explicit and never a silent null.
Non-measured classes (predicted/derived) impose no decode-TPS
requirement; genuinely-optional optimizer fields stay null. The CLI
catches MeasuredRecordError (exit 2, clean message) and echoes warnings
to stderr. Extends the test with: measured + no decode summary => fail
loud; measured + bad/absent GPU line => parse_warnings; non-measured +
no decode => no raise; and a happy-path regression asserting empty
parse_warnings.
Addresses the Codex review Medium finding on #249.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>