Surfaces the canonical int8_per_token_head anti-scaling trade-off so the
next dual-3090 user who hits this finds the answer via search rather than
having to reason from the head-to-head matrix.
Frames it as:
- not a bug — per-(token, head) scale serializes at concurrency, fp8's
single global scale doesn't
- pick by workload: INT8 PTH for single-stream throughput, fp8 for
aggregate concurrency, Genesis-backed TQ3+MTP if you want both
- includes the diagnostic checklist for matching baseline numbers
(power cap, vLLM nightly, Genesis on/off, MTP n, prompt shape) so users
can self-troubleshoot a cross-rig gap before posting
Triggered by a Discord question on a dual-3090 + NVLink rig that observed
the canonical INT8 anti-scaling pattern (150 TPS single / flat at concurrency)
vs fp8 (70-100 TPS single / 400+ at concurrency).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>