Files
club-3090/models/gemma-4-31b
noonghunna 5ea64bc539 Re-gate beellama/gemma-dflash on the official pin; fix rebench-full n=3
The single-card Gemma-4 default was the catalog's only default with no
baseline row (wave-2 PRIORITY gap). Full gate on Anbeeld's OFFICIAL
v0.3.2-preview digest (the current beellama-local pin), rebench tag
gemma-dflash-regate:

- decode 44.91 narr / 79.76 code (n=5, CV 2.1%/5.6%), TTFT 115 ms
- NIAH ladder clean to 117,513 tok (91% of 128K); ceiling margin 929 MB
  (< 1024 MB bar -> caveat stays, but up from ~177 MB pre-v0.3.2)
- 8-pack current harness: 108/150 off / 113/150 on (+5 in-band wash,
  think-off default re-confirmed)
- soak PASS (p50 59.8, 100.6% retention, 0 growth, 0/100 silent-empty)

Row seeded via catalog-baseline.sh (first non-vLLM engine through the
producer flow); BENCHMARKS row + compose header (Quality + margin
numbers) updated to match.

Tooling fallout the gate exposed:
- rebench-full.sh hardcoded RUNS:-3/WARMUPS:-1, silently under-running
  the bench protocol on every orchestrated gate. The A1 "n=3 env leak"
  attribution was wrong -- it was this line (baselines comment
  corrected). Fixed to protocol 5/3; n=3 had flattered gemma-dflash
  code TPS by +8%.
- catalog-baseline.sh counted prefill-probe run-lines into its bench-n
  gate (read n=8 for an n=5 log); now reads the summary headers.

Full scripts gate 64/64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-07-04 19:28:03 +00:00
..

Gemma 4 31B — on 2× RTX 3090

Run Gemma 4 31B — with vision and tool calling — on 2× RTX 3090s, on stock vLLM v0.24.0 (overlay-free).

⚠️ Single-card boot OOMs on 24 GB Ampere regardless of KV format. Needs ≥32 GB single-card (validated on RTX 5090 by @apnar).


Deployment

See docs/DUAL_CARD.md for workload-driven config picks. TL;DR:

Config Max ctx Decode TPS Best for
vllm/gemma-31b-dual (default) 224K ~59 General-purpose, vision + tools — stock vLLM v0.24.0, overlay-free

Run via:

bash scripts/launch.sh --variant vllm/gemma-31b-dual     # bf16 @224K, v0.24.0, overlay-free

v0.24.0 consolidation (2026-07-02): the 31b is now a single overlay-free bf16 dual slug on vllm-stable. The v0.22.0 composes (gemma-int8-mtp = 262K int8-PTH + PR #40391, gemma-bf16-mtp = 131K, gemma-mtp-tp1, gemma-31b-qat-w4a16-dual) are deprecated (switch.sh --list --all). MTP is off — Gemma-4 MTP × tool-calling is broken on v0.24.0 (vLLM #39043 / #42006). The 262K int8-PTH path returns overlay-free when PR #40391 merges upstream (on v0.24.0 int8-PTH allocates 262K but silently craters recall past ~32K, so bf16 @224K is the honest default).


Models

Key details

Aspect Notes
Quants Intel AutoRound INT4
KV bfloat16 @224K on stock v0.24.0 (overlay-free). The int8-PTH + PR #40391 262K path is deprecated (returns free when #40391 merges)
Drafter none — MTP disabled on v0.24.0 (Gemma-4 MTP × tools broken, vLLM #39043 / #42006)
Vision Yes
Tools --tool-call-parser gemma4
NVLink Auto-detected via NVLINK_MODE env var

Upstream tracker

  • vLLM PR #41745 — Gemma 4 MTP support (merged)
  • vLLM PR #40391 — INT8 PTH KV page-align (OPEN/unmerged). The deprecated vllm/gemma-int8-mtp (v0.22.0) vendors it for 262K; on v0.24.0 int8-PTH craters recall without it, so the default is bf16 (vllm/gemma-31b-dual). 262K int8-PTH returns when this merges.
  • Discussion #67 — first Ampere consumer cross-rig data