Files
club-3090/scripts/lib
5da4eab12e Add llamacpp/qwen27b-pi-reasoning-single (Qwen3.6-27B Pi-style coding agent, mainline llama.cpp + MTP) (#425)
bytkim/Qwen3.6-27B-MTP-pi-reasoning Q4_K_M GGUF (embedded MTP head) — a "Pi-style"
reasoning-supervised CODING / terminal-agent fine-tune — on MAINLINE llama.cpp
(server-cuda-b9246, PR #22673), single 3090, q4_0/q4_0 KV + MTP, reasoning-ON.

- New compose: models/qwen3.6-27b/llama-cpp/compose/single/pi-reasoning-q4km/mtp.yml
- Registry slug llamacpp/qwen27b-pi-reasoning-single (experimental, port 8063).
- Weights entry pi-reasoning-q4km; drafter qwen-mtp-builtin (spec_method mtp).

CONFIG FOLLOWS THE MODEL CARD: temp 1.0 / top-p 0.95 / top-k 0 / min-p 0 (NOT the
stack's 0.6/20), reasoning ON, q4_0/q4_0 KV, --jinja -ngl 99 -fa. Card recommends
MTP n=3; on-rig A/B found n=2 marginally faster (within noise) — kept n=2,
MTP_DRAFT_N_MAX=3 matches the card. presence-penalty 1.5 is a documented knob for
the card's DIRECT/instruct (REASONING=off) mode.

CONTEXT (measured 2026-06-17, GPU0/GPU1): default 200K-alloc fills ~188K usable with
correct needle recall (22.7 GB / ~1.8 GB free; ~23 t/s decode at ~188K depth). Do NOT
alloc 262K — the FA scratch grows with the allocation, so 262K OOMs at ~176K (LESS
usable than 200K); full 262K usable is beellama-only. Author tested only 128K, so
128-188K is engine-proven but past the card's validated window (CTX_SIZE=131072 for
strict compliance).

BENCH (canonical bench.sh n=3, thinking-off, short-prompt): narrative 28.5 wall / 28.7
decode, code 32.9 / 33.4, PP 743 tok/s — ~45% below base llamacpp/default (50/59) on
identical engine/KV/MTP, i.e. this fine-tune's embedded MTP head is weaker. Engine A/B:
mainline ~25% faster than a beellama q4_0/q4_1 path → mainline chosen. Stays
experimental (--force): verify-stress / soak / quality ladder pending.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-18 03:56:28 +05:00
..