Same 3090 (GPU 0, air-cooled), same engine (mainline llama.cpp), same
Q4-class quant (Q4_K_XL). Only the model changes: dense Qwen3.6-27B vs
A3B 35B (3B active per token).
Findings:
1. MoE shifts decode sweet spot 80W lower:
- Dense decode: 290W → 0.111 TPS/W, SM 1380 MHz at sweet spot
- A3B decode: 210W → 0.546 TPS/W, SM 1290 MHz at sweet spot
Each token only activates 3B of 35B params on MoE → less compute per
token → bandwidth-bound knee fires at lower power.
2. Prefill sweet spot is workload-determined, NOT model-determined:
- Dense prefill: 250W → 3.633 TPS/W
- A3B prefill: 250W → 9.865 TPS/W
Both converge to same cap because prefill is compute-bound regardless
of MoE routing.
3. Boost-clock plateau depends on workload AND model:
- Dense decode: PLATEAU at 340-370W (SM 1560 MHz lock)
- A3B decode: NO PLATEAU (SM climbs smoothly 1875→1890→1890→1905)
- Dense prefill: PLATEAU at 330-370W (SM 1605-1620)
- A3B prefill: PLATEAU at 340-370W (SM 1680-1710)
Plateau auto-detection correctly flagged dense decode but not A3B
decode — confirming firmware operating-point selection responds to
compute pressure, not just to cap value.
Adds:
- docs/img/power-cap-3090-a3b-decode.py + .png
- docs/img/power-cap-3090-a3b-prefill.py + .png
- HARDWARE.md cross-rig table rows for both A3B sweeps
- HARDWARE.md "Same hardware, MoE workload" subsection with comparison
table + practical recommendation (A3B users → 210W cap, vs 290W for
dense Qwen)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>