Files
club-3090/models
noonghunnaandClaude Opus 4.8 8c1f811dd2 gemma-4-12b mainline llama.cpp: single-card 256K (override-kv, q8_0 KV)
Faster single-card sibling to the beellama 256K compose: mainline ggml-org serves
gemma4 ~9% faster (51 vs 47 TPS on the same Q8 weights). Same --override-kv
gemma4.context_length=int:262144 lever lifts the n_ctx_train cap (verified on stock
ggml-org: n_ctx=262144, coherent, 16 GB @ 256K). Uses q8_0 KV (FA-OK without
FA_ALL_QUANTS, half of f16 so 256K fits). KV-quant is TPS-neutral here (beellama
q8_0 47.7 == q5_0/q4_1 47.0). No-spec; MTP sibling stays gated on llama.cpp #23398.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-04 03:18:11 +00:00
..