bytkim/Qwen3.6-27B-MTP-pi-reasoning Q4_K_M GGUF (embedded MTP head) — a "Pi-style" reasoning-supervised CODING / terminal-agent fine-tune — on MAINLINE llama.cpp (server-cuda-b9246, PR #22673), single 3090, q4_0/q4_0 KV + MTP, reasoning-ON. - New compose: models/qwen3.6-27b/llama-cpp/compose/single/pi-reasoning-q4km/mtp.yml - Registry slug llamacpp/qwen27b-pi-reasoning-single (experimental, port 8063). - Weights entry pi-reasoning-q4km; drafter qwen-mtp-builtin (spec_method mtp). CONFIG FOLLOWS THE MODEL CARD: temp 1.0 / top-p 0.95 / top-k 0 / min-p 0 (NOT the stack's 0.6/20), reasoning ON, q4_0/q4_0 KV, --jinja -ngl 99 -fa. Card recommends MTP n=3; on-rig A/B found n=2 marginally faster (within noise) — kept n=2, MTP_DRAFT_N_MAX=3 matches the card. presence-penalty 1.5 is a documented knob for the card's DIRECT/instruct (REASONING=off) mode. CONTEXT (measured 2026-06-17, GPU0/GPU1): default 200K-alloc fills ~188K usable with correct needle recall (22.7 GB / ~1.8 GB free; ~23 t/s decode at ~188K depth). Do NOT alloc 262K — the FA scratch grows with the allocation, so 262K OOMs at ~176K (LESS usable than 200K); full 262K usable is beellama-only. Author tested only 128K, so 128-188K is engine-proven but past the card's validated window (CTX_SIZE=131072 for strict compliance). BENCH (canonical bench.sh n=3, thinking-off, short-prompt): narrative 28.5 wall / 28.7 decode, code 32.9 / 33.4, PP 743 tok/s — ~45% below base llamacpp/default (50/59) on identical engine/KV/MTP, i.e. this fine-tune's embedded MTP head is weaker. Engine A/B: mainline ~25% faster than a beellama q4_0/q4_1 path → mainline chosen. Stays experimental (--force): verify-stress / soak / quality ladder pending. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 <[email protected]>
80 lines
3.2 KiB
Bash
Executable File
80 lines
3.2 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
set -euo pipefail
|
|
|
|
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
|
|
cd "$ROOT_DIR"
|
|
export PYTHONPATH="$ROOT_DIR${PYTHONPATH:+:$PYTHONPATH}"
|
|
|
|
python3 - <<'PY'
|
|
from pathlib import Path
|
|
|
|
from scripts.lib.profiles.compat import load_profiles
|
|
from scripts.lib.profiles.compose_registry import COMPOSE_REGISTRY, DEFAULTS
|
|
|
|
root = Path.cwd()
|
|
profiles = load_profiles()
|
|
registry_paths = {Path(entry["compose_path"]) for entry in COMPOSE_REGISTRY.values()}
|
|
disk_paths = set(Path("models").glob("*/*/compose/*/*/*.yml"))
|
|
|
|
failures = []
|
|
|
|
def check(cond, msg):
|
|
if cond:
|
|
print(f"PASS: {msg}")
|
|
else:
|
|
print(f"FAIL: {msg}")
|
|
failures.append(msg)
|
|
|
|
check(len(COMPOSE_REGISTRY) == 53, f"registry has 53 entries (got {len(COMPOSE_REGISTRY)})")
|
|
check(len(disk_paths) == 54, f"disk has 54 compose files (got {len(disk_paths)})")
|
|
check(registry_paths <= disk_paths, "all registry compose_path values exist on disk")
|
|
parked_disk_only = disk_paths - registry_paths
|
|
# Disk-only (non-registry) composes allowed: parked SGLang archives, plus the experimental
|
|
# vLLM-Omni Qwen3-Omni compose (intentionally NOT registry-wired — custom-engine, direct
|
|
# `docker compose`-only deploy; see models/qwen3-omni-30b-a3b/vllm-omni/README.md).
|
|
def _allowed_disk_only(path):
|
|
return (
|
|
"/sglang/compose/" in f"/{path.as_posix()}"
|
|
or path == Path("models/qwen3-omni-30b-a3b/vllm-omni/compose/dual/autoround-int4/omni.yml")
|
|
)
|
|
check(
|
|
all(_allowed_disk_only(path) for path in parked_disk_only),
|
|
"only parked SGLang archives + the non-registry vLLM-Omni compose are disk-only",
|
|
)
|
|
if parked_disk_only:
|
|
print("INFO: disk-only parked composes: " + ", ".join(str(p) for p in sorted(parked_disk_only)))
|
|
|
|
for name, entry in sorted(COMPOSE_REGISTRY.items()):
|
|
path = Path(entry["compose_path"])
|
|
parts = path.parts
|
|
check(path.exists(), f"{name}: compose_path exists")
|
|
check(path.name not in {"docker-compose.yml", "default.yml"}, f"{name}: filename is descriptive")
|
|
try:
|
|
idx = parts.index("compose")
|
|
topology, quant_slug, filename = parts[idx + 1:idx + 4]
|
|
except (ValueError, IndexError):
|
|
check(False, f"{name}: path follows compose/<topology>/<quant>/<file>.yml")
|
|
continue
|
|
check(topology in {"single", "dual", "multi4"}, f"{name}: topology segment valid")
|
|
check(filename.endswith(".yml"), f"{name}: compose filename is .yml")
|
|
check(quant_slug == entry["weights_variant"], f"{name}: quant slug matches weights_variant")
|
|
model = profiles.models[entry["model"]]
|
|
check(entry["weights_variant"] in model.weights, f"{name}: weights_variant exists in ModelProfile")
|
|
|
|
for key, name in sorted(DEFAULTS.items()):
|
|
model, _engine, topology = key
|
|
entry = COMPOSE_REGISTRY.get(name)
|
|
check(entry is not None, f"DEFAULTS{key}: target exists")
|
|
if entry is None:
|
|
continue
|
|
path_parts = Path(entry["compose_path"]).parts
|
|
idx = path_parts.index("compose")
|
|
check(entry["model"] == model, f"DEFAULTS{key}: model matches target")
|
|
check(path_parts[idx + 1] == topology, f"DEFAULTS{key}: topology matches target path")
|
|
|
|
if failures:
|
|
raise SystemExit(f"{len(failures)} registry/disk checks failed")
|
|
PY
|
|
|
|
echo "test-compose-registry-disk: ok"
|