Files
club-3090/scripts/tests
noonghunna b84249c805 fix(verify-stress): three live-caught bugs in ceiling ladder (#199)
Bug 1 — get_vram_free_mb inflated on multi-GPU hosts:
  Summed ALL GPUs → reported 24507 MB when GPU0 truly had 381 free
  (added idle GPU1's 24 GB). Margin gate was defeated on any host with
  more GPUs than the model uses.
  FIX: reads Docker HostConfig.DeviceRequests[0].DeviceIDs to identify
  the model's GPU(s), passes nvidia-smi -i <those>. Falls back to
  CUDA_VISIBLE_DEVICES / NVIDIA_VISIBLE_DEVICES env, then all-GPUs with
  a warning. Live-validated: 353 MB (GPU0 only) vs old 24479 MB.

Bug 2 — filler→token ratio overshoots ~18x:
  The rung used scale = target_tokens / 3.5, treating scale as chars
  when it's actually block-repetition count. The 95K rung produced
  1.77M tokens (674% of n_ctx) → HTTP 400 on a healthy engine.
  FIX: calibration probe before the ladder sends scale=100, reads back
  prompt_tokens, computes the real tok/scale_unit ratio. Live-validated:
  95K rung now produces 94788 tokens (0.2% error, 36% of n_ctx).

Bug 3 — no-op ladder reported PASS:
  A ladder that tested nothing (all rungs HTTP 400) said 'All stress
  checks passed.' A rung with target < n_ctx returning 400 is a sizing
  error, not a clean skip.
  FIX: distinguishes sizing errors (target < n_ctx + HTTP 400) from
  legitimate engine rejections (target > n_ctx + HTTP 400). Sizing
  errors FAIL the probe. 'All rungs skipped' now FAILs with a clear
  diagnostic instead of silently passing.

Tests: 37 pass (9 new — GPU-scoped VRAM query for Bug 1, dual/single/
scoped GPU variants, streaming helper timing extraction with real
llama.cpp response shape).

All three bugs were caught by running the ladder against the live
262K compose on the dual-3090 dev rig. The unit tests alone didn't
catch them — mocks fed the wrong shapes (top-level n_ctx, sum-all-GPUs,
clean-skip semantics). Live validation is now part of the workflow.
2026-05-23 04:42:18 +00:00
..