Files
club-3090/docs
noonghunnaandClaude Opus 4.7 df53287b1c docs(faq): add 'INT8 PTH doesn't scale at concurrency — is that a bug?'
Surfaces the canonical int8_per_token_head anti-scaling trade-off so the
next dual-3090 user who hits this finds the answer via search rather than
having to reason from the head-to-head matrix.

Frames it as:
- not a bug — per-(token, head) scale serializes at concurrency, fp8's
  single global scale doesn't
- pick by workload: INT8 PTH for single-stream throughput, fp8 for
  aggregate concurrency, Genesis-backed TQ3+MTP if you want both
- includes the diagnostic checklist for matching baseline numbers
  (power cap, vLLM nightly, Genesis on/off, MTP n, prompt shape) so users
  can self-troubleshoot a cross-rig gap before posting

Triggered by a Discord question on a dual-3090 + NVLink rig that observed
the canonical INT8 anti-scaling pattern (150 TPS single / flat at concurrency)
vs fp8 (70-100 TPS single / 400+ at concurrency).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 09:58:16 +00:00
..