Files
noonghunna 2abe025513 Add llamacpp/hauhaucs-35ba3b-dual uncensored MTP compose (🧪) (#410)
Wires morikomorizz/Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP (Q6_K_P GGUF
with an embedded nextn MTP head) as a dual-card mainline llama.cpp b9570
compose: -ts 0.55,0.45, q8_0 KV, MTP n=3, 262K, reasoning-on by default.

The MTP head loads clean on mainline ("speculative decoding context
initialized") — the prior HauhauCS-MTP ret=-3 was an ik-llama/older-build
issue, not the model arch. The -ts 0.55,0.45 split rebalances the MTP draft
card (even 1,1 skews ~3 GB at 262K).

Validated 2026-06-14: verify-stress 8/8 (NIAH ceiling ladder -> 240K),
bench.sh n=3 @262K (narr 113.4 / code ~150 decode TPS, CV<1%), soak fresh
20x5 PASS (0 growth, 0/100 silent-empty, p50 162.4, 99.6% retention),
8-pack think-OFF 103/150 / think-ON 105/150 (wash). n=3 vs n=1 @262K =
-9% prose / +10% code (code-leaning default by request; n=1 prose-best via
MTP_DRAFT_N_MAX=1).

Status 🧪 Experimental: community GGUF (digest-unpinned) + uncensored.
No DEFAULTS row — opt-in only. Guard suite green.

Co-authored-by: noonghunna <10742901+noonghunna@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 06:28:08 +05:00
..