Faster single-card sibling to the beellama 256K compose: mainline ggml-org serves
gemma4 ~9% faster (51 vs 47 TPS on the same Q8 weights). Same --override-kv
gemma4.context_length=int:262144 lever lifts the n_ctx_train cap (verified on stock
ggml-org: n_ctx=262144, coherent, 16 GB @ 256K). Uses q8_0 KV (FA-OK without
FA_ALL_QUANTS, half of f16 so 256K fits). KV-quant is TPS-neutral here (beellama
q8_0 47.7 == q5_0/q4_1 47.0). No-spec; MTP sibling stays gated on llama.cpp #23398.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>