ik-llama/iq4ks-mtp is the fastest single-card path (~18-20% faster decode +
leanest VRAM, #184) but was absent from the README quick-start and the registry
ctx values had drifted from the 200K shipped default. Keeping llamacpp/default =
mainline (the simplest / clean-upstream-image pick) — surfacing ik, not renaming.
- README quick-start: add `ik-llama/iq4ks-mtp` (fastest single-card) alongside
the llamacpp/* variants.
- docs/SINGLE_CARD.md: one-line "simplest (llamacpp/default) vs fastest
(ik-llama/iq4ks-mtp)" steer atop the config table (ik rows were already
present + ⭐-marked).
- compose_registry.py: fix stale wizard-projection max_ctx to the 200K default —
llamacpp/default + llamacpp/mtp were 131072 (too low), ik-llama/iq4ks-mtp was
262144 (boots-not-fills); all → 200000. Vision entries (49152 / 163840) already
correct. Reworded the "262K via -ub 512" comment (that was the boots≠fills
false ceiling). NOTE: launch.sh doesn't pass max_ctx to runtime, so this is
wizard-projection accuracy only — runtime ctx still comes from the compose
CTX_SIZE=200000 default.
Validated: registry imports; test-launch-compat, test-switch-registry-parity
(45 composes, parity), test-profiles-compat all pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>