Per request: turn on the external MTP drafter for BOTH topologies and
ladder each compose to its empirically-validated max context.
- single/awq/base.yml -> single/awq/mtp.yml (rename per ADDING_MODELS.md:
drafter=MTP named, bf16 KV is the gemma4-on-Ampere default so NOT named).
Adds --speculative-config (gemma-26b-it-assistant n=4) + prefix-caching +
chunked-prefill (parity with the dual). Max ctx 8K -> 16K: KV pool 17,490
tok at 0.92 util on 1x 3090 with the drafter loaded (17,408 would leave
<100 tok headroom). Registry drafter None -> gemma-26b-it-assistant.
- dual/awq/mtp.yml: max ctx 32K -> 262,144 (model max). KV pool 806,821 tok
at 262144/0.92 on 2x 3090 — sliding_window=1024 keeps 25/30 layers' KV
cheap, so the model max fits with 3x headroom.
Both LIVE-validated on the rig 2026-06-06 (stock v0.22.0): single tp=1 +
MTP boots (SpeculativeConfig method=mtp n=4), coherent, clean gemma4
tool-call; dual tp=2 + MTP @262K boots, coherent, clean tool-call.
Registry max_ctx 16384/262144 + single compose_path -> mtp.yml; profile_runtime
mirrors (single gains the speculative_config slot, dual maxlen 262144).
test-generate-from-profile mk_einput now clears the base slug's drafter (the
clean-derived fixtures need a no-drafter seed; the drafter-refuse fixture sets
one explicitly) so it's stable now that gemma-26ba4b-single carries an MTP drafter.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>