Bug 1 — get_vram_free_mb inflated on multi-GPU hosts:
Summed ALL GPUs → reported 24507 MB when GPU0 truly had 381 free
(added idle GPU1's 24 GB). Margin gate was defeated on any host with
more GPUs than the model uses.
FIX: reads Docker HostConfig.DeviceRequests[0].DeviceIDs to identify
the model's GPU(s), passes nvidia-smi -i <those>. Falls back to
CUDA_VISIBLE_DEVICES / NVIDIA_VISIBLE_DEVICES env, then all-GPUs with
a warning. Live-validated: 353 MB (GPU0 only) vs old 24479 MB.
Bug 2 — filler→token ratio overshoots ~18x:
The rung used scale = target_tokens / 3.5, treating scale as chars
when it's actually block-repetition count. The 95K rung produced
1.77M tokens (674% of n_ctx) → HTTP 400 on a healthy engine.
FIX: calibration probe before the ladder sends scale=100, reads back
prompt_tokens, computes the real tok/scale_unit ratio. Live-validated:
95K rung now produces 94788 tokens (0.2% error, 36% of n_ctx).
Bug 3 — no-op ladder reported PASS:
A ladder that tested nothing (all rungs HTTP 400) said 'All stress
checks passed.' A rung with target < n_ctx returning 400 is a sizing
error, not a clean skip.
FIX: distinguishes sizing errors (target < n_ctx + HTTP 400) from
legitimate engine rejections (target > n_ctx + HTTP 400). Sizing
errors FAIL the probe. 'All rungs skipped' now FAILs with a clear
diagnostic instead of silently passing.
Tests: 37 pass (9 new — GPU-scoped VRAM query for Bug 1, dual/single/
scoped GPU variants, streaming helper timing extraction with real
llama.cpp response shape).
All three bugs were caught by running the ladder against the live
262K compose on the dual-3090 dev rig. The unit tests alone didn't
catch them — mocks fed the wrong shapes (top-level n_ctx, sum-all-GPUs,
clean-skip semantics). Live validation is now part of the workflow.