Files
club-3090/scripts
7df24a46fb Add ik-llama apex-fit-q8q5 variant for Qwen3.6-35B-A3B (#242) (#243)
* Add ik-llama apex-fit-q8q5 variant for Qwen3.6-35B-A3B (#242)

Captures @laurimyllari's `--fit` + asymmetric q8_0(K)/q5_0(V) KV config
from discussion #241 as a single-card ik-llama variant on the APEX
I-Compact GGUF (registry tag `ik-llama/apex-fit-q8q5`, port 8057).

Measured 1× 3090 + 370 W, n=5:

  q4/q4 mtp.yml baseline (first real run — APEX weights weren't on disk
    before today): 96.47 / 144.01 wall TPS narr/code (CV 5.4% / 4.5%)
    20.46 GB VRAM
  q8/q5 fit-mtp.yml (this variant):           103.25 / 149.12 wall TPS
    (CV 3.0% / 1.5%) — +7% narr / +4% code at tighter CV, +0.6 GB VRAM
    The Anbeeld K-high/V-low asymmetric-KV pattern materialised here.

Gates:
  verify-full        8/8 PASS
  verify-stress      8/8 PASS incl. 180K NIAH (91% of n_ctx 196608)
  bench              above
  soak-continuous    PASS (0 errors, 0/25 silent_empty, 0 VRAM growth,
                     100% TPS retention, p50 decode 223 TPS)
  deterministic q    76/90 = 84% on PR #38 verifiers
                     toolcall 14/15 · instructfollow 15/15 ·
                     structoutput 14/15 · dataextract 11/15 ·
                     reasonmath 12/15 · bugfind 10/15

MoE × MTP sub-question (from #242 body) — answered: ik-llama built-in
MTP on the 35B-A3B MoE does NOT pay the vLLM −51% / −35% penalty
(cf. `qwen3.6-35b-a3b/dual/preview-mtp.yml` BENCHMARKS row). MTP context
ready at n_ctx=196608, decode bursts 270+ TPS in soak. The MoE×MTP
penalty is vLLM-scheduler-specific, not architectural.

Status set to ⚠️ Production w/ caveats because sandbox-pack quality
(hermesagent-20 0/20, aider-polyglot-30 0/30, cli-40 timeouts) hit
benchlocal-cli sandbox infrastructure issues 2026-05-28 — hermes
`tool_events=0` after the agent runs (suspected PR #38 thinking-on
sampler interaction); aider fails at git checkout
`fatal: path 'aider/__init__.py' does not exist in 'f46766c'` before
any LLM call. Neither is attributable to the model or the compose.
Pre-PR-#38 cross-rig reference from @laurimyllari's 4090 in #241:
hermes 9-12/20, aider 14-18/30 — the model class is capable; the
sandbox state on this rig needs separate work.

Catalog gates: test-compose-registry-disk count bumped 55→56 per the
documented model-add workflow. test-profiles-compat, test-model-
weights-registry, test-switch-registry-parity, test-launch-registry-
parity all PASS. Two inherited reds (test-compose-mounts-resolve and
test-patch-attribution) unchanged from master baseline — they affect
qwen3.6-27b vLLM composes, not this addition.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

* fit-mtp.yml: tighten --fit/--no-mmap/--cache-ram rationale (PR #243 review)

@laurimyllari clarified on PR #243 that for the I-Compact GGUF
(~17 GB on 24 GB card) both `--fit` and `--no-mmap` are largely inert
since the model fits in VRAM — the reason to keep them as defaults
is forward-compat: swapping GGUF_FILE for a bigger quant
(UD-Q8_K_XL, APEX Quality, etc.) "just works" with reasonable
partial-offload performance, without re-tuning the compose.

Replaced the two "Why" blocks in the compose docstring to lead with
his framing. Also flagged `--cache-ram 4096` (lower than ik-llama's
8192 default) with the same suspected forward-compat rationale,
pending his confirmation on the PR.

No flag changes; docstring-only.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

* Promote fit-mtp.yml ✅ Production + update sandbox-pack quality

Updates the compose Status + Quality line + BENCHMARKS row after live
validation of the sandbox-pack fixes (benchlocal-cli #42/#43/#44 +
club-3090 #245, all merged today):

  hermesagent-20:    0/20 → 11/20 (55%)  via #42 deterministic sampler
  aider-polyglot-30: 0/30 → 12/30 (40%)  via #44 git-checkout from AIDER_DIR
  cli-40:            11/40 → 12/40 (30%, ±1 noise)  via #43 budget fix
                     (helps aider/hermes wall-clock; cli-40 "timeouts"
                     turn out to be sandbox-internal agent-give-up, not
                     wall-clock budget — see follow-up benchlocal-cli issue)

All gates clean: verify-full 8/8, verify-stress 8/8 (incl. 180K NIAH),
bench n=5 (+7% narr / +4% code vs q4/q4 mtp.yml), soak-continuous PASS
(0 errors, 0 silent_empty, 0 VRAM growth, 100% retention). Status:
🧪/⚠️ → ✅ Production. Caveats trimmed to the 3090 power-sensitivity
note.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-28 16:31:09 +05:00
..
…
…
…