Files
club-3090/models/gemma-4-31b
noonghunnaandClaude Opus 4.7 f8c706699c docs(gemma-4-31b): document TQ3 Ampere FA2 head_dim wall + vendor #40108 overlay
Walked the full vllm#41403 5-gate chain for Gemma 4 + turboquant_3bit_nc on
2× RTX 3090 (SM 8.6). Gates 1-5 worked around; gate 6 is hardware-physical
(FA2 head_dim ≤ 256 on Ampere, Gemma 4 global layers use head_dim=512, FA3
support requires Hopper SM 9.0+).

Retains the int8-tq3.yml compose + the vendored #40108 overlay as a
re-test surface for when either (a) TQ backend gains pure-Triton prefill
for head_dim>256 on Ampere, or (b) vLLM gains per-layer attention backend
routing, or (c) we have Hopper hardware. None close.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-11 02:50:04 +00:00
..