Dense 40B uncensored Qwen3.6 community merge, dual-3090 llama.cpp with
embedded MTP head (Q6_K GGUF, 31 GB). First dual llama.cpp compose in
the catalog; first 'category' field in the registry (uncensored).
Changes:
- Compose: models/qwen3.6-40b-deckard/llama-cpp/compose/dual/mtp-q6k/mtp.yml
Status: 🧪 Unverified (quality /150 + soak pending)
Config: -ngl 99 -ts 1,1 -fa on --spec-type draft-mtp --spec-draft-n-max 2
--cache-type-k q8_0 --cache-type-v q8_0 -c 131072
- Registry: llamacpp/deckard40B-dual-mtp (category=uncensored, experimental)
- DEFAULTS: (qwen3.6-40b-deckard, llamacpp, dual) → slug
- Model profile: scripts/lib/profiles/models/qwen3.6-40b-deckard.yml
- Drafter: qwen-mtp-builtin model_compat extended with qwen3.6-40b-deckard
- LiteLLM route: deckard-40b → :8199
- launch.sh: suggest_default_variant case for Deckard
- BENCHMARKS.md: new section with validated MTP n=2 numbers
- Test fixtures: registry count 43→44, disk 44→49, models 5→6
Wrinkle decisions (see PR description):
1. Dual llama.cpp: uses 'count: all' + '-ts 1,1' (tensor-split layer-split),
matching the validated serving config. No launcher changes needed — the
compose's deploy section handles GPU reservation directly.
2. Engine pin: uses rolling server-cuda tag (matching engine profile spec),
not the b9246 pin other composes use. The validated build (2026-06-09
digest 1c4ff61a) is newer than b9246 and has draft-mtp working. No engine
pin bump for other models.
NOT DONE (live validation — GPUs busy with quality eval):
- Boot/bench/soak/quality NOT run — maintainer to validate post-eval
- Status stays 🧪 until full gate passes
From maintainer feedback on the P1 bundle:
- scripts/setup-image-studio.sh: pre-run plan + confirm prompt (--yes / CI=1 /
non-TTY auto-yes to never hang), --help/usage banner, and a proper "Get started"
block — create your admin account (first sign-up = admin; no creds pre-set),
pick the gemma-4-12b chat model, then 🖼️ to generate. States the fresh-vs-existing
volume wiring caveat.
- services/litellm/config.yaml: add a gemma-4-12b route (-> :8069), so the
image-studio chat brain is reachable through the gateway too (it's the one route
live in image-studio mode; the big-model routes are GPU-mutex with ComfyUI).
Open WebUI still points direct to :8069 by default for a clean picker.
- docs/IMAGE_STUDIO.md: architecture section + ASCII diagram (front-end -> chat /
image; the 2-GPU split; LiteLLM gateway), explicit first-run + how-to-generate-an-
image-in-chat steps, chat-routing explanation, and pin/v0.9.6/secret accuracy fixes.
Live-validated: gemma-4-12b responds through LiteLLM :4000; setup --help + bash -n clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Hardening from the live cutover validation:
- ComfyUI pinned to a known-good commit (COMFYUI_REF default cb9f6394… — has
Ideogram-4) instead of floating HEAD, so an upstream change can't silently break
users. Set COMFYUI_REF=HEAD to float. Entrypoint is now mounted into the container
so pin/bootstrap edits apply on `up` without a 30-min image rebuild (kept +x — a
non-exec mounted entrypoint = OCI "permission denied", caught in validation).
- OWUI WEBUI_SECRET_KEY no longer hardcoded to "change-this-secret-key" (a shared,
forgeable secret); now empty → OWUI generates+persists a strong per-deployment key
in the data volume. Override via host env only to share sessions across replicas.
Live-validated on the committed compose: end-to-end gen OWUI :8080 → ComfyUI →
Ideogram-4 (81 s); boot log confirms "pinned to cb9f6394" + GPU0 pin; OWUI on the
persisted key file.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The gateway config still routed to retired pre-club-3090 endpoints (:8000
vLLM-patched, :8001/:8002 SGLang, :8003 longctx/ngram, :8004 luce-dflash) —
none of which gpu-mode brings up anymore — and had no route to the shipped
35B-A3B dual.
Now 3 live primaries only:
- qwen3.6-27b-autoround -> :8010 (gpu-mode 27b)
- qwen3.6-35b-a3b-autoround -> :8051 (launch.sh vllm/qwen-35b-a3b-dual; served-model-name matched)
- gemma-4-31b-autoround -> :8030 (gpu-mode gemma)
Gemma-4-26B-A4B left unrouted on purpose: its registry-default compose
(vllm/gemma-a4b, autoround-int4-mixed) is SM86-blocked (Marlin K-dim); the
working AWQ variant isn't a gpu-mode primary. Noted inline.
Takes effect on next `docker restart litellm` / gpu-mode switch. YAML
validated; no repo scripts/tests reference the removed model_names.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
The dev rig had grown across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`,
`/home/wasif/`) and three repos (single-3090, dual-3090, club-3090). Disk-out
on `/` (97% used) forced a cleanup; rather than just prune, we consolidated
the layout across the whole stack while we were at it. This commit captures
what landed inside this repo.
Services consolidation:
- services/{ollama,openwebui,litellm,qdrant,searxng}/ migrated in from
/opt/ai/compose/<svc>/ (zero functional change — same docker-compose.yml).
- services/litellm/config.yaml rewritten: explicit routes for current
primaries (qwen3.6-27b-autoround → :8010, gemma-4-31b-autoround → :8030).
Removed `* → ollama/*` wildcard.
- services/comfyui/ migrated in (was /opt/ai/compose/comfyui/) — wired into
gpu-mode with full mutex against vLLM/SGLang.
scripts/gpu-mode.sh under git:
- Was a loose /opt/ai/gpu-mode.sh outside any repo. Symlinked at
/usr/local/bin/gpu-mode.
- Five Gemma 4 31B modes added: gemma, gemma-dflash, gemma-int8,
gemma-dflash-int8, gemma-awq.
- One ComfyUI mode (mutex with all LLM serving).
- prune / prune-all subcommands (safe image prune; aggressive variant adds
build cache --keep-storage 5GB + dangling networks).
- gpu-mode status now shows Docker disk + /var/lib/docker + /tmp sizes.
- compose_at() passes --env-file <repo>/.env so MODEL_DIR resolves
regardless of which compose dir gpu-mode cd's into. Fixes the recurring
"MODEL_DIR not set, defaulting to ../../../../../models-cache" warning.
- stderr no longer swallowed by compose_at() (real errors surface).
- Cross-model VRAM mutex: every Qwen mode stop_all_gemma + stop_comfyui
and vice-versa.
scripts/maintenance/ — new hygiene-tools subdir:
- list-image-pins.sh: engine-agnostic pin auditor. Scans every compose's
`image:` line, groups by `<repo>:<tag>`, flags pin-drift (multiple tags
per repo), ranks composes by patch surface.
Pin tracking:
- docs/UPSTREAM.md gains a "Pinned images" section: table of every pinned
image, why each pin exists, retirement candidate criteria.
- docs/NIGHTLY_BUMP_RUNBOOK.md (new): 7-step procedure for bumping pinned
engine images (scope → branch → patch survival → boot → verify-full +
verify-stress → bench delta → land → retire). Engine-specific notes
for vLLM nightly hashes, llama.cpp digest pinning, SGLang variants.
Path updates from the engine + model dir consolidation:
- /opt/ai/vllm-src/ → /opt/ai/engines/vllm/primary/
(in setup.sh, INTERNALS.md, several patch READMEs, docs/HARDWARE.md,
docs/FAQ.md, docs/DUAL_CARD.md, docs/UPSTREAM.md, models/qwen3.6-27b/
CHANGELOG.md)
- /mnt/models/gguf/qwen3.6-27b/ → /mnt/models/huggingface/qwen3.6-27b-gguf/
(in models/qwen3.6-27b/llama-cpp/{compose/single/*.yml, recipes/*.sh,
README.md}, docs/engines/LLAMA_CPP.md)
CHANGELOG.md narrative gap fill:
- 2026-05-10 entry for this reorg.
- 2026-05-09 entry for compose convention formalization (topology
promoted to dir level, profile schema, Status enum + Caveats, cliff
CI swap, Discord launch).
- 2026-05-08 entry for Gemma 4 INT8 PTH unblock + 262K validation.
- 2026-05-07 entry for power-cap-sweep campaign + HARDWARE.md cross-rig
charts + cross-rig benchmark rows.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>