The arch-confirm blocker named in the 🧪 status note is resolved: the GGUF header reports general.architecture=qwen35 (standard GQA, 97 layers), NOT the Qwen3-Next/DeltaNet hybrid the original scaffold guessed. Corrected the profile to family=qwen35-dense and dropped the fictional GDN/linear-attention fields. Fixes found while correcting the profile: - Restored the required `attention_k_eq_v: false` field (compat.py reads it with bracket notation — dropping it broke load_profiles for the whole catalog). - Fixed a duplicate-key bug in the weights block: a second `hf_repo:` (DavidAU base, no MTP head) shadowed PiehSoft's MTP GGUF under yaml.safe_load last-key-wins — setup.sh would have fetched the wrong, head-less repo. - Added qwen35-dense to llama-cpp-local supported_model_families (the C10 gate). Renamed the weights slug mtp-q6k → piehsoft-q6k to follow the provider-quant idiom used by every other GGUF compose (ubergarm-iq4ks, unsloth-q4km, byteshape-iq4xs) and to drop the redundant mtp in mtp-q6k/mtp.yml. The serving filename stays mtp.yml (KV isn't encoded in GGUF-engine filenames). Renamed the compose dir + on-disk weights dir + all path references in lockstep. Validation (2× 3090, server-cuda-b9570): verify-full 8/8, verify-stress 8/8 (ceiling ladder filled to 120K / 91% of n_ctx, 0 MiB VRAM growth, needle recall clean through 90K), 8-pack 105/150 with MTP off==on (spec-dec lossless), soak-continuous PASS. Guard suite 42/0. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
30 lines
1.1 KiB
YAML
30 lines
1.1 KiB
YAML
schema_version: 1
|
|
id: llama-cpp-local
|
|
display_name: llama.cpp local CUDA server (MTP-enabled)
|
|
type: llama.cpp
|
|
stability: stable
|
|
install:
|
|
method: docker_image
|
|
spec: ghcr.io/ggml-org/llama.cpp:server-cuda
|
|
min_sm: 6.0
|
|
supported_model_families:
|
|
- qwen3-next-hybrid
|
|
- qwen3-next-moe
|
|
- qwen35-dense # Deckard-40B (standard GQA, arch=qwen35; MTP via embedded nextn head)
|
|
- gemma4-unified
|
|
features: {}
|
|
supported_kv_formats:
|
|
- q4_0
|
|
- q5_0
|
|
- q8_0
|
|
- k8v4
|
|
supported_drafters:
|
|
- mtp
|
|
supported_weight_formats:
|
|
- gguf
|
|
required_overlays: []
|
|
vendored_overlays: []
|
|
required_genesis: false
|
|
notes: "Cliff-immune single-card fallback for Qwen GGUF serving. Uses upstream `ghcr.io/ggml-org/llama.cpp`, pinned to build `server-cuda-b9246` (NOT the rolling `:server-cuda` tag — it regressed at b9282 with a broken-lib crash loop, #187). MTP PR #22673 merged upstream + draft-eagle3 + draft-mtp spec types available natively. Bump via env once a newer build is validated: `LLAMACPP_IMAGE=ghcr.io/.../server-cuda-bXXXX`. No club-3090 patches on llama.cpp; we follow upstream as-is."
|
|
|