Files
club-3090/models/qwen3.6-27b/llama-cpp/compose/single/unsloth-q4km/mtp.yml
noonghunna 9821c94efb refactor(compose): insert <quant> layer + make the registry the single source of truth
Restructure all 47 composes to
models/<model>/<engine>/compose/<topology>/<quant-slug>/<serving>.yml and centralize
defaults as registry pointers. Move + path-rewire only — no compose runtime-config
changes (plus the +1 ../ depth bump each moved file requires, and descriptive names
for the former docker-compose.yml defaults).

- <quant-slug> = 1:1 with a weights artifact == compose_registry weights_variant ==
  weights.py key (provider-quant form). Fixes the prior coarse `gguf` and the
  bf16/int8-files-mislabeled-as-autoround_int4 weights_variant.
- default = DEFAULTS[(model,engine,topology)] pointer in compose_registry.py; no
  default.yml / docker-compose.yml. KV-dtype variants (bf16/int8) stay as files
  under autoround-int4/.
- +1 ../ on every relative mount in moved composes (cache/patches/scripts/models-cache).
- Rewired: compose_registry (47 compose_path + weights_variant + DEFAULTS), launch.sh,
  gpu-mode.sh, bench-row-formatter, profile *.yml, generate_compose.py, weights.py
  (+ back-compat ALIASES), AGENTS.md, ADDING_MODELS.md, BENCHMARKS + doc path refs.
- New guard tests: test-compose-registry-disk.sh, test-compose-mounts-resolve.sh.
- Variant keys (vllm/dual, ik-llama/iq4ks-mtp, ...) unchanged — launch/switch UX preserved.

Validation: docker compose config 47/47 · parity/launch-compat/registry-disk/
mounts-resolve/profiles-compat/model-weights-registry PASS · kv-calc 22/22 · leak-clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 15:58:08 +00:00

129 lines
7.0 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# ===========================================================================
# Profile (at-a-glance):
# Model: Qwen3.6-27B (Unsloth MTP-enabled Q4_K_M GGUF)
# Engine: llama.cpp (local MTP-enabled build, build 9235)
# Topology: Single 3090 (TP=1)
# Drafter: MTP n=2 (--spec-type draft-mtp, sweet spot — see BENCHMARKS.md)
# KV: q4_0 K + q4_0 V (densest mainline, Ampere-fast)
# Vision: no (mmproj NOT mounted — for vision, see mtp-vision.yml)
# Template: native (GGUF-embedded) — froggeric A/B'd here, regressed 8-pack 102→95
# Max ctx: 200000 default (@ -ub 512, fills ~183K w/ ~1.1 GB margin); 262144 boots but
# walls ~125K (FA scratch at fill — see CTX_SIZE note); 131072 = faster prefill @ -ub 1024
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
# Status: ✅ Production
# Engine-profile: llama-cpp-local
# Best for: IDE agents, opencode, Hermes, long-multi-turn agentic — the
# speed + ctx workhorse. ~51 narr / ~60 code TPS, 7/7 verify-stress
# (incl. 60K + 91K needle), 102/150 quality (68%) on the 8-pack.
# ---------------------------------------------------------------------------
# Qwen3.6-27B on llama.cpp — single 3090, MTP, 200K ctx, no vision.
#
# This is one of TWO named single-card profiles for this model (collapsed from
# three on 2026-05-22 — the old vanilla Q3_K_XL `docker-compose.yml` was retired
# once Q4_K_M was shown to cover long context too; see the CTX_SIZE note below):
# - mtp.yml → MTP n=2 + no vision (THIS file). `llamacpp/default`
# is now an ALIAS for this profile.
# - mtp-vision.yml → MTP n=2 + vision (multimodal)
#
# Why no vision here: mmproj F16 costs ~0.8 GB. Without it the MTP-safe ctx
# ceiling jumps from ~49K to 131K (sweep-verified 2026-05-19 on build 9235).
# If you don't need image input, this is the better MTP profile.
#
# Template: native (GGUF-embedded) + `--reasoning off` + `--reasoning-format
# deepseek` — the stack-wide thinking-off lever. We A/B'd froggeric v19 here
# (it loads + honours --reasoning off fine on b9246 — the old "froggeric
# suppresses --reasoning off" lore is stale), but it REGRESSED the 8-pack
# 102 → 95 on mainline llama.cpp. froggeric helps the vLLM/Qwen3-Next path,
# not this engine — so native is the default. Closes opencode hang (#97).
#
# Cliff-survival: -ub 1024 (lowered from the older 2048 default). The 25K
# tool-prefill check in verify-stress fails at -ub 2048 on tight single-card
# headroom; -ub 1024 halves the per-pass activation peak and the boundary
# matrix goes 5/7 → 7/7. The "Cliff 2 single-prompt at 5060K is
# architectural" narrative is partially superseded by this config: at
# -ub 1024 + 131K + MTP n=2 + thinking-off, verify-stress recalls needles
# cleanly at 58K and 91K.
#
# VRAM budget on 24 GB (Q4_K_M):
# weights (Q4_K_M): ~17.0 GB
# KV at 131K (q4_0 K+V): ~5.0 GB
# MTP draft head + overhead: ~0.5 GB
# total: ~22.5 GB
# headroom: ~1.6 GB for prompt + activation peaks
#
# Quick start:
# 1. Get the MTP-enabled GGUF:
# hf download unsloth/Qwen3.6-27B-MTP-GGUF Qwen3.6-27B-Q4_K_M.gguf \
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-mtp-q4km
# 2. From this directory:
# MODEL_DIR=/your/models/dir docker compose -f mtp.yml up -d
# 3. curl http://localhost:8020/v1/models → should list the model.
#
# Override defaults via .env or shell:
# MODEL_DIR host dir to mount as /models (default: ../../../../../../models-cache)
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf)
# CTX_SIZE KV pool size (default: 200000 @ -ub 512 — fills ~183K with
# ~1.1 GB free, the max-safe single-card value. 262144 boots ~23 GB,
# MTP active, passes verify-stress(91K) + soak — but only *fills* to
# ~125K before OOMing on the FA scratch at high fill: a false ceiling.
# See docs/CLIFFS.md (2026-05-23). Faster prefill: CTX_SIZE=131072 UBATCH_SIZE=1024.)
# BATCH_SIZE llama.cpp -b (default: 4096 — on mainline -b does NOT drive VRAM)
# UBATCH_SIZE llama.cpp -ub (default: 512 — funds the 200K default; raise to 1024 at 131K)
# KV_TYPE K and V quant type (default: q4_0)
# NP parallel slots (default: 1 — see ⚠ below)
# MTP_DRAFT_N_MAX MTP draft tokens (default: 2 — sweet spot per BENCHMARKS.md)
# REASONING thinking gate (default: off — stack-wide policy)
# REASONING_FORMAT reasoning routing (default: deepseek — hygiene)
# PORT host port (default: 8020)
# CUDA_VISIBLE_DEVICES which GPU to use (default: 0)
services:
llama-cpp-qwen36-27b-mtp:
# Pinned to b9246 (validated 2026-05-20). The rolling `:server-cuda` tag
# regressed at b9282 (broken lib packaging: `libllama-common.so.0: cannot
# open shared object file` → crash loop) — see #187. Bump via env when a
# newer build is validated: LLAMACPP_IMAGE=ghcr.io/ggml-org/llama.cpp:server-cuda-bXXXX
image: ${LLAMACPP_IMAGE:-ghcr.io/ggml-org/llama.cpp:server-cuda-b9246}
container_name: "${ESTATE_CONTAINER:-llama-cpp-qwen36-27b}"
restart: unless-stopped
ports:
- "${ESTATE_PORT:-${PORT:-8020}}:8080"
volumes:
- "${MODEL_DIR:-../../../../../../models-cache}:/models:ro"
# ⚠ -np 1 is intentional on a single 24 GB card — do NOT raise it to
# "parallelize." One GPU is compute-bound: extra slots divide its
# throughput, they don't multiply it. At -np 4 each slot fell to
# ~14 tok/s here — slow enough to trip agentic clients' per-request
# timeouts (aider ran 1/30) — and -np>1 also auto-disables MTP and
# can OOM the spec-context buffer. On a higher-throughput card (e.g.
# 5090) or multi-GPU the trade may flip — re-validate before raising.
command: >-
--host 0.0.0.0
--port 8080
-m /models/${GGUF_FILE:-qwen3.6-27b-gguf/unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf}
-c ${CTX_SIZE:-200000}
-b ${BATCH_SIZE:-4096}
-ub ${UBATCH_SIZE:-512}
-ngl 99
-fa on
--cache-type-k ${KV_TYPE:-q4_0}
--cache-type-v ${KV_TYPE:-q4_0}
-np ${NP:-1}
--spec-type draft-mtp
--spec-draft-n-max ${MTP_DRAFT_N_MAX:-2}
--jinja
--reasoning ${REASONING:-off}
--reasoning-format ${REASONING_FORMAT:-deepseek}
--temp ${TEMP:-${TEMPERATURE:-0.6}}
--top-p ${TOP_P:-0.95}
--top-k ${TOP_K:-20}
--min-p ${MIN_P:-0.0}
--repeat-penalty ${REPEAT_PENALTY:-1.0}
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["${ESTATE_GPUS:-${CUDA_VISIBLE_DEVICES:-0}}"]
capabilities: [compute, utility]