Files
club-3090/scripts/lib
noonghunnaandClaude Opus 4.7 6dc9a0dce1 feat(gemma-4-26b-a4b): AWQ + MTP n=4 — +12% narr / +49% code over no-MTP baseline
New compose models/gemma-4-26b-a4b/vllm/compose/dual/awq-mtp.yml adds
external MTP drafter wiring via google/gemma-4-26B-A4B-it-assistant
(832 MB BF16, already downloaded). COMPOSE_REGISTRY entry
vllm/gemma-a4b-awq-mtp at port 8043.

Live boot + bench 2026-05-15, dual 3090 PCIe, 230 W cap:

  awq.yml        (no MTP) → 138.88 / 138.67 wall TPS (139 / 139 decode)
  awq-mtp.yml    (n=4)    → 155.05 / 207.02 wall TPS (157 / 211 decode)
                            +12% narr / +49% code

MTP metrics:
  AL: 3.04 narrative, 3.79 code
  Per-pos accept (narr): 0.77 / 0.55 / 0.40 / 0.29 (50.9% avg)
  Per-pos accept (code): 73.5% avg
  Both GPUs at 98% util / 357 W and 305 W (symmetric)

Cross-MoE finding (opposite direction from Qwen 35B-A3B MTP, see
preview-mtp.yml row in BENCHMARKS): Gemma's external assistant
drafter is a small dense model (~0.5 B params, Gemma4AssistantForCausalLM)
that BYPASSES the MoE expert routing entirely on the draft pass.
The Qwen built-in MTP head runs through the model's MoE forward, so
each draft step pays MoE routing + inter-GPU sync overhead — that's
why Qwen MTP n=3 was 50% SLOWER but Gemma MTP n=4 is 49% FASTER.

Practical rule for MoE models on Ampere: prefer external drafters
over built-in MTP heads where both are available. Documented in
learnings/gemma-4-26b-a4b.md + cross-ref in qwen3.6-35b-a3b.md.

Recommendation: awq-mtp.yml becomes the recommended default for
Gemma 26B-A4B going forward; awq.yml stays as the no-drafter A/B
reference.

PR #40886 overlay still applied via the same patches/install.sh
bind-mount pattern as awq.yml.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
..