Files
club-3090/scripts/lib/profiles/engines/llama-cpp-mainline.yml
T
noonghunnaandClaude Opus 4.8 93acbf979f Promote Deckard-40B to ✅ Production; fix arch + slug naming
The arch-confirm blocker named in the 🧪 status note is resolved: the GGUF
header reports general.architecture=qwen35 (standard GQA, 97 layers), NOT the
Qwen3-Next/DeltaNet hybrid the original scaffold guessed. Corrected the profile
to family=qwen35-dense and dropped the fictional GDN/linear-attention fields.

Fixes found while correcting the profile:
- Restored the required `attention_k_eq_v: false` field (compat.py reads it with
  bracket notation — dropping it broke load_profiles for the whole catalog).
- Fixed a duplicate-key bug in the weights block: a second `hf_repo:` (DavidAU
  base, no MTP head) shadowed PiehSoft's MTP GGUF under yaml.safe_load
  last-key-wins — setup.sh would have fetched the wrong, head-less repo.
- Added qwen35-dense to llama-cpp-local supported_model_families (the C10 gate).

Renamed the weights slug mtp-q6k → piehsoft-q6k to follow the provider-quant
idiom used by every other GGUF compose (ubergarm-iq4ks, unsloth-q4km,
byteshape-iq4xs) and to drop the redundant mtp in mtp-q6k/mtp.yml. The serving
filename stays mtp.yml (KV isn't encoded in GGUF-engine filenames). Renamed the
compose dir + on-disk weights dir + all path references in lockstep.

Validation (2× 3090, server-cuda-b9570): verify-full 8/8, verify-stress 8/8
(ceiling ladder filled to 120K / 91% of n_ctx, 0 MiB VRAM growth, needle recall
clean through 90K), 8-pack 105/150 with MTP off==on (spec-dec lossless),
soak-continuous PASS. Guard suite 42/0.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-10 00:38:15 +00:00

30 lines
1.1 KiB
YAML

schema_version: 1
id: llama-cpp-local
display_name: llama.cpp local CUDA server (MTP-enabled)
type: llama.cpp
stability: stable
install:
method: docker_image
spec: ghcr.io/ggml-org/llama.cpp:server-cuda
min_sm: 6.0
supported_model_families:
- qwen3-next-hybrid
- qwen3-next-moe
- qwen35-dense # Deckard-40B (standard GQA, arch=qwen35; MTP via embedded nextn head)
- gemma4-unified
features: {}
supported_kv_formats:
- q4_0
- q5_0
- q8_0
- k8v4
supported_drafters:
- mtp
supported_weight_formats:
- gguf
required_overlays: []
vendored_overlays: []
required_genesis: false
notes: "Cliff-immune single-card fallback for Qwen GGUF serving. Uses upstream `ghcr.io/ggml-org/llama.cpp`, pinned to build `server-cuda-b9246` (NOT the rolling `:server-cuda` tag — it regressed at b9282 with a broken-lib crash loop, #187). MTP PR #22673 merged upstream + draft-eagle3 + draft-mtp spec types available natively. Bump via env once a newer build is validated: `LLAMACPP_IMAGE=ghcr.io/.../server-cuda-bXXXX`. No club-3090 patches on llama.cpp; we follow upstream as-is."