Commit Graph
6 Commits
Author SHA1 Message Date
noonghunna 5dafaa61fe feat: add Qwen3.6-40B-Deckard to catalog (llamacpp/deckard40B-dual-mtp)
Dense 40B uncensored Qwen3.6 community merge, dual-3090 llama.cpp with
embedded MTP head (Q6_K GGUF, 31 GB). First dual llama.cpp compose in
the catalog; first 'category' field in the registry (uncensored).

Changes:
- Compose: models/qwen3.6-40b-deckard/llama-cpp/compose/dual/mtp-q6k/mtp.yml
  Status: 🧪 Unverified (quality /150 + soak pending)
  Config: -ngl 99 -ts 1,1 -fa on --spec-type draft-mtp --spec-draft-n-max 2
          --cache-type-k q8_0 --cache-type-v q8_0 -c 131072
- Registry: llamacpp/deckard40B-dual-mtp (category=uncensored, experimental)
- DEFAULTS: (qwen3.6-40b-deckard, llamacpp, dual) → slug
- Model profile: scripts/lib/profiles/models/qwen3.6-40b-deckard.yml
- Drafter: qwen-mtp-builtin model_compat extended with qwen3.6-40b-deckard
- LiteLLM route: deckard-40b → :8199
- launch.sh: suggest_default_variant case for Deckard
- BENCHMARKS.md: new section with validated MTP n=2 numbers
- Test fixtures: registry count 43→44, disk 44→49, models 5→6

Wrinkle decisions (see PR description):
1. Dual llama.cpp: uses 'count: all' + '-ts 1,1' (tensor-split layer-split),
   matching the validated serving config. No launcher changes needed — the
   compose's deploy section handles GPU reservation directly.
2. Engine pin: uses rolling server-cuda tag (matching engine profile spec),
   not the b9246 pin other composes use. The validated build (2026-06-09
   digest 1c4ff61a) is newer than b9246 and has draft-mtp working. No engine
   pin bump for other models.

NOT DONE (live validation — GPUs busy with quality eval):
- Boot/bench/soak/quality NOT run — maintainer to validate post-eval
- Status stays 🧪 until full gate passes
2026-06-09 21:59:13 +00:00
noonghunnaandClaude Opus 4.8 18902fa495 image-studio P1 follow-up: setup UX, LiteLLM route, architecture docs
From maintainer feedback on the P1 bundle:

- scripts/setup-image-studio.sh: pre-run plan + confirm prompt (--yes / CI=1 /
  non-TTY auto-yes to never hang), --help/usage banner, and a proper "Get started"
  block — create your admin account (first sign-up = admin; no creds pre-set),
  pick the gemma-4-12b chat model, then 🖼️ to generate. States the fresh-vs-existing
  volume wiring caveat.
- services/litellm/config.yaml: add a gemma-4-12b route (-> :8069), so the
  image-studio chat brain is reachable through the gateway too (it's the one route
  live in image-studio mode; the big-model routes are GPU-mutex with ComfyUI).
  Open WebUI still points direct to :8069 by default for a clean picker.
- docs/IMAGE_STUDIO.md: architecture section + ASCII diagram (front-end -> chat /
  image; the 2-GPU split; LiteLLM gateway), explicit first-run + how-to-generate-an-
  image-in-chat steps, chat-routing explanation, and pin/v0.9.6/secret accuracy fixes.

Live-validated: gemma-4-12b responds through LiteLLM :4000; setup --help + bash -n clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-09 05:13:31 +00:00
noonghunnaandClaude Opus 4.8 092b7c7a46 image-studio P1: pin ComfyUI commit + drop hardcoded OWUI secret
Hardening from the live cutover validation:
- ComfyUI pinned to a known-good commit (COMFYUI_REF default cb9f6394… — has
  Ideogram-4) instead of floating HEAD, so an upstream change can't silently break
  users. Set COMFYUI_REF=HEAD to float. Entrypoint is now mounted into the container
  so pin/bootstrap edits apply on `up` without a 30-min image rebuild (kept +x — a
  non-exec mounted entrypoint = OCI "permission denied", caught in validation).
- OWUI WEBUI_SECRET_KEY no longer hardcoded to "change-this-secret-key" (a shared,
  forgeable secret); now empty → OWUI generates+persists a strong per-deployment key
  in the data volume. Override via host env only to share sessions across replicas.

Live-validated on the committed compose: end-to-end gen OWUI :8080 → ComfyUI →
Ideogram-4 (81 s); boot log confirms "pinned to cb9f6394" + GPU0 pin; OWUI on the
persisted key file.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-09 04:22:19 +00:00
noonghunnaandClaude Opus 4.8 9981b28698 Add image-studio bundle P1: Ideogram-4 + gemma-12b chat + gpu-mode mode
Wires the committed services/comfyui + services/openwebui scaffold into a
seamless image-gen + chat bundle (coexisting on a 2-GPU box):

- services/comfyui/download_ideogram4.sh: fetch the Ideogram-4 fp8 set
  (2 transformers + Qwen3-VL-8B enc + flux2 VAE) into the ComfyUI models tree
- services/comfyui: opt-in COMFYUI_CUDA_VISIBLE_DEVICES GPU pin (entrypoint guard;
  empty default = all GPUs, preserves current behavior)
- services/openwebui: pin :v0.9.6; image-gen wired via imagegen.env (env_file,
  PersistentConfig — fresh-volume); default chat -> gemma-4-12b :8069 (LiteLLM alt)
- scripts/gpu-mode.sh: `image-studio` mode — ComfyUI/Ideogram-4 on GPU0 +
  gemma-4-12b chat on GPU1 (compose_at_env passes env through sudo); :8069 status
- scripts/setup-image-studio.sh: one-shot build + download + bring-up

Config-validated (docker compose config, bash -n, gpu-mode usage); test suite
40/42 (2 failures pre-existing, unrelated). Docs (IMAGE_STUDIO.md + deltas) and
live cutover validation to follow on this branch.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-09 03:15:51 +00:00
fb07edc686 chore(litellm): prune dead routes, add 35B-A3B dual, drop wildcard (#263)
The gateway config still routed to retired pre-club-3090 endpoints (:8000
vLLM-patched, :8001/:8002 SGLang, :8003 longctx/ngram, :8004 luce-dflash) —
none of which gpu-mode brings up anymore — and had no route to the shipped
35B-A3B dual.

Now 3 live primaries only:
- qwen3.6-27b-autoround     -> :8010  (gpu-mode 27b)
- qwen3.6-35b-a3b-autoround -> :8051  (launch.sh vllm/qwen-35b-a3b-dual; served-model-name matched)
- gemma-4-31b-autoround     -> :8030  (gpu-mode gemma)

Gemma-4-26B-A4B left unrouted on purpose: its registry-default compose
(vllm/gemma-a4b, autoround-int4-mixed) is SM86-blocked (Marlin K-dim); the
working AWQ variant isn't a gpu-mode primary. Noted inline.

Takes effect on next `docker restart litellm` / gpu-mode switch. YAML
validated; no repo scripts/tests reference the removed model_names.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-05-30 18:03:46 +05:00
noonghunnaandClaude Opus 4.7 00366a58d7 reorg: services/ consolidation + gpu-mode under git + ComfyUI + pin tracker + path updates
Release / release (push) Failing after 53s
The dev rig had grown across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`,
`/home/wasif/`) and three repos (single-3090, dual-3090, club-3090). Disk-out
on `/` (97% used) forced a cleanup; rather than just prune, we consolidated
the layout across the whole stack while we were at it. This commit captures
what landed inside this repo.

Services consolidation:
- services/{ollama,openwebui,litellm,qdrant,searxng}/ migrated in from
  /opt/ai/compose/<svc>/ (zero functional change — same docker-compose.yml).
- services/litellm/config.yaml rewritten: explicit routes for current
  primaries (qwen3.6-27b-autoround → :8010, gemma-4-31b-autoround → :8030).
  Removed `* → ollama/*` wildcard.
- services/comfyui/ migrated in (was /opt/ai/compose/comfyui/) — wired into
  gpu-mode with full mutex against vLLM/SGLang.

scripts/gpu-mode.sh under git:
- Was a loose /opt/ai/gpu-mode.sh outside any repo. Symlinked at
  /usr/local/bin/gpu-mode.
- Five Gemma 4 31B modes added: gemma, gemma-dflash, gemma-int8,
  gemma-dflash-int8, gemma-awq.
- One ComfyUI mode (mutex with all LLM serving).
- prune / prune-all subcommands (safe image prune; aggressive variant adds
  build cache --keep-storage 5GB + dangling networks).
- gpu-mode status now shows Docker disk + /var/lib/docker + /tmp sizes.
- compose_at() passes --env-file <repo>/.env so MODEL_DIR resolves
  regardless of which compose dir gpu-mode cd's into. Fixes the recurring
  "MODEL_DIR not set, defaulting to ../../../../../models-cache" warning.
- stderr no longer swallowed by compose_at() (real errors surface).
- Cross-model VRAM mutex: every Qwen mode stop_all_gemma + stop_comfyui
  and vice-versa.

scripts/maintenance/ — new hygiene-tools subdir:
- list-image-pins.sh: engine-agnostic pin auditor. Scans every compose's
  `image:` line, groups by `<repo>:<tag>`, flags pin-drift (multiple tags
  per repo), ranks composes by patch surface.

Pin tracking:
- docs/UPSTREAM.md gains a "Pinned images" section: table of every pinned
  image, why each pin exists, retirement candidate criteria.
- docs/NIGHTLY_BUMP_RUNBOOK.md (new): 7-step procedure for bumping pinned
  engine images (scope → branch → patch survival → boot → verify-full +
  verify-stress → bench delta → land → retire). Engine-specific notes
  for vLLM nightly hashes, llama.cpp digest pinning, SGLang variants.

Path updates from the engine + model dir consolidation:
- /opt/ai/vllm-src/ → /opt/ai/engines/vllm/primary/
  (in setup.sh, INTERNALS.md, several patch READMEs, docs/HARDWARE.md,
   docs/FAQ.md, docs/DUAL_CARD.md, docs/UPSTREAM.md, models/qwen3.6-27b/
   CHANGELOG.md)
- /mnt/models/gguf/qwen3.6-27b/ → /mnt/models/huggingface/qwen3.6-27b-gguf/
  (in models/qwen3.6-27b/llama-cpp/{compose/single/*.yml, recipes/*.sh,
   README.md}, docs/engines/LLAMA_CPP.md)

CHANGELOG.md narrative gap fill:
- 2026-05-10 entry for this reorg.
- 2026-05-09 entry for compose convention formalization (topology
  promoted to dir level, profile schema, Status enum + Caveats, cliff
  CI swap, Discord launch).
- 2026-05-08 entry for Gemma 4 INT8 PTH unblock + 262K validation.
- 2026-05-07 entry for power-cap-sweep campaign + HARDWARE.md cross-rig
  charts + cross-rig benchmark rows.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 16:57:03 +00:00