Headline finding from @apnar's 4th + 5th sweeps in one day (disc #86): - Decode N=8 (Gemma 4 + MTP): tops at 551W actual draw (vs 547W at N=4) → confirms decode is memory-bandwidth bound, not concurrency-limited - Prefill-heavy (Qwen3.6 long-text): hits 599.98W actual draw at 600W cap → 99.997% cap-respect, full TDP saturation Per-workload-class power ceiling on the 5090: - Decode workloads: ~547-551W max (memory bandwidth limit, regardless of cap) - Prefill workloads: ~600W (compute-bound, scales with TDP) Both workload classes share efficiency knee at 400W cap (67% of stock TDP) — that "60-85% of stock TDP" cross-rig pattern holds across workloads. But absolute throughput behavior differs: decode at 600W gives only 4% more TPS than at 400W (bandwidth ceiling); prefill at 600W gives 19% more TPS (compute actually uses the watts). Practical implication for 5090 deployments: - Chat/IDE-agent: cap at 400W (huge efficiency win, ~5% TPS loss) - RAG/long-context: leave at stock 600W (compute-bound, watts buy throughput) - Mixed: 400W if chat-dominant; pure prefill loads suffer New chart: docs/img/power-cap-5090-qwen36-prefill.png shows the curve with the "599.98W actual at 600W cap" callout. HARDWARE.md additions: - 2 new BENCHMARKS rows for prefill-heavy 400W (sweet spot) + 600W (saturation) - Inline embed of the prefill chart after the decode (Gemma 4 + MTP) chart - New "Per-workload-class power ceilings" subsection with the cross-workload bottleneck table Note on apnar's contribution: he ran 5 sweeps total today (50W resolution v1, 50W after calibration fix v2, 10W canonical anchor v3, decode N=8 headroom probe v4, prefill-heavy compute-saturation test v5) — all on the old token-bounded bench architecture before today's time-bounded redesign shipped (his sweeps at 13-14 UTC, my time-bounded redesign at 18:15 UTC). The cross-validation depth is rare — most cross-rig contributions stop at one sweep.
252 KiB
1956x922px
252 KiB
1956x922px