Commit Graph
171 Commits
Author SHA1 Message Date
cc1e0a2251 Add Ornith-1.0-35B experimental slug (ik-llama/ornith35b-dual) (#479)
DeepReinforce agentic-coding RL fine-tune of Qwen3.6-35B-A3B (same qwen35moe arch, no MTP head, text-only). ik_llama dual 3090, Q8_0, full 262K, drafter-free ngram (opt-in). Gate PASS: ~108.8/108.6 TPS, verify-stress 8/8, soak PASS, 8-pack 105/105. Ties the base on the 8-pack (110), EDGES it on aider (15/30 vs 12-13) -> coding-leaning lane. Catalog suite 58/58 green.


Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-26 08:54:51 +05:00
6cdefab1f3 Add Ornith-1.0-9B experimental slug (ik-llama/ornith9b-single) (#477)
DeepReinforce agentic-coding RL fine-tune — Qwen3-Next dense-FFN HYBRID (arch=qwen35: 8 full-attn + 24 GDN/DeltaNet layers, NON-MoE), single 3090, Q4_K_M + q8_0 KV, full 262K (KV only 4.25 GiB — just 8/32 layers carry full-attn GQA KV; 13.4 GiB total). Drafter-free ngram self-spec on ik_llama (works despite the DeltaNet hybrid). Full gate PASS: bench ~102 TPS, verify-stress 8/8, soak PASS, 8-pack 91/150 off / 95 on. 🧪 niche only — gemma-4-12b beats it on quality (105) + speed; pick for the lean 13.4 GiB footprint. Catalog suite 58/58 green.


Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-26 06:01:56 +05:00
9a24463374 Make the OWUI model picker scene-accurate (studio lanes + LLMs) (#470)
The OWUI picker listed models that weren't actually serving: the studio
image/video/audio lanes (always shown via the studio pipe) and every
LLM catalog model (always shown via the always-up LiteLLM :4000
gateway). Both error until you switch to the matching gpu-mode scene.
Now the picker reflects what's live.

Studio lanes — gate the pipe's pipes() on backend liveness:
- build_studio_pipe.py (source template) + regenerated studio_pipe.py:
  add a _LANES catalog tagged by backend (ComfyUI :8188 for all media
  lanes, the on-demand voice service :8193 for the voice lane), a cheap
  reachability probe (_alive, short timeout) with an 8s TTL cache, and
  filter pipes() to only the lanes whose backend answers. New valve
  hide_unavailable_lanes (default true) toggles it off. When the
  ai-studio scene is down, the Studio group drops out of the picker
  instead of listing dead lanes.

LLMs — point OWUI at each backend directly instead of the gateway:
- services/openwebui/docker-compose.yml: seed one connection per backend
  (director :8090 + :8010/:8051/:8032/:8038/:8199), drop :4000. OWUI
  hides models from an unreachable connection, so each model shows only
  while its scene serves. Served names match the catalog (IDs unchanged).
  LiteLLM stays up for other clients; it just leaves OWUI's picker.
- scripts/lib/owui-unregister.sh: new symmetric inverse of
  owui-register.sh (remove a connection by port; forged-JWT config API,
  idempotent, no-op if OWUI down / not present).
- scripts/setup-ai-studio.sh: idempotent step that registers the 6
  per-backend connections + drops a stale :4000, so existing installs
  converge (fresh installs get it from the compose env seed).

Live-validated on the rig (gemma12b scene): picker LLMs = only
gemma-4-12b-int8 (the four down models hidden, :4000 gone), Studio group
hidden (ComfyUI + voice down). Probe verified both directions (live
endpoint -> shown even on a 404 path; dead port -> hidden).


Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-25 05:22:05 +05:00
noonghunnaandClaude Opus 4.8 2693d937f8 Bump ComfyUI pin to v0.26.0 + queue Krea 2 download
v0.26.0 (comfyanonymous/ComfyUI #14589) adds native LOCAL Krea2 support —
the earlier "cloud-only, dropped" verdict was pin-specific (cb9f6394
predated it; only the cloud Krea2ImageNode existed → local load failed
"Could not detect model type").

- entrypoint.sh: COMFYUI_REF default cb9f6394 -> f6c162dd (v0.26.0); reword
  the pin comment (Krea2 + Qwen3-VL text-gen; re-validate ALL lanes on bump).
- docker-compose.yml: pin comment v0.26.0 + Krea2.
- download_krea.sh (new): fetches the 3 Comfy-Org/Krea-2 assets (turbo fp8
  DiT + Qwen3-VL-4B encoder + Qwen-Image VAE) into the ComfyUI models tree.
- download_studio_models.sh + studio-models.tsv: add Krea as an image-lane
  model (roster <-> manifest kept mirrored).

NOT yet merge-ready: the v0.26.0 pin is a load-bearing bump — the custom
HiDream-O1 node, ComfyUI-GGUF, and DisTorch multi-GPU all ride it. On-rig
re-validation of every studio lane (Krea renders + the existing 9 still
work) is the merge gate. The OWUI lane graph (krea2.json workflow +
studio_pipe entry) is authored at that validation step (the Krea sampler /
Qwen3-VL encode graph differs from Z-Image — not clonable blind).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-06-24 15:40:00 +00:00
noonghunnaandClaude Opus 4.8 5fb8fa1c7f studio: download scripts + manifest rows for Z-Image + Wan2.2
Wire the two new lanes into the install/preflight surfaces so a fresh rig fetches them
and c3 / gpu-mode know they're expected:

- download_zimage.sh — Z-Image-Turbo fp8 + Qwen3-4B encoder + flux ae VAE (~12 GB).
- download_wan.sh — Wan2.2-Rapid Mega NSFW Q8 GGUF + umt5 encoder + Wan 2.1 VAE (~25 GB).
- download_studio_models.sh ("grab all") calls both (idempotent — skips what's present).
- studio-models.tsv: +image/Z-Image, +video/Wan2.2 rows (the shared manifest read by both
  c3's StudioModel loader and gpu-mode preflight). Roster note records Krea2 as dropped.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-06-24 00:16:22 +00:00
noonghunnaandClaude Opus 4.8 3fc26857f2 studio: consolidate image/video into one ai-studio gpu-mode scene
Replace the separate `image-studio` / `video-studio` / `comfyui` scenes with
a single `ai-studio` scene: ComfyUI holds both GPUs always, and image / video /
audio are picked as *lanes* in OWUI rather than gpu-mode switches. Removes the
gemma-12b chat brain from the studio bundle (kept on disk, just unwired) and
points OWUI's DEFAULT_MODELS at the qwen director.

- gpu-mode.sh: `mode_ai_studio` replaces the two studio modes; drop the gemma
  start/stop + the GPU0-pin split; preflight now reads the shared manifest.
- scripts/lib/studio-models.tsv: single source-of-truth model manifest (modality,
  label, root, rel_path, size, installer) consumed by both gpu-mode preflight and
  the c3 cockpit, so the two can't drift.
- download_{director,video,audio,ace_step,stable_audio,studio}_models.sh: idempotent
  fetchers; download_studio_models.sh orchestrates "grab everything".
- setup-video-studio.sh + repointed setup-image-studio.sh → `gpu-mode ai-studio`.
- litellm: drop the gemma-4-12b :8069 route; comfyui compose/entrypoint GPU-pin
  comments corrected (no scene sets the pin now).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
2026-06-23 23:43:12 +00:00
noonghunnaandClaude Opus 4.8 639525bd85 c3 Download: registry-driven companion artifacts (DFlash draft / mmproj)
A catalog slug's compose can mount a SEPARATE weight artifact beyond its core
weights — a DFlash draft model or an mmproj vision projector, from a different
subdir.  The cockpit Download fetched only the core, so those 8 slugs (5 DFlash
+ 3 GGUF-vision) would download, read "present", offer Start, then fail to boot
for the missing companion.

The registry is the single source of truth: the cockpit gets every catalog slug
from it, so the slug's required artifacts come from it too.

- compose_registry.py: new per-slug `weights_companions` field — the extra
  weight-variant keys a slug needs beyond `weights_variant`.  Set on the 5 DFlash
  slugs (-> anbeeld-dflash-iq4xs) + 3 vision slugs (-> gguf_mmproj_f16).  Per-slug,
  so a text slug sharing a vision GGUF variant (e.g. llamacpp/default on
  unsloth-q4km) does NOT over-fetch the ~1GB mmproj.
- registry-emit --json: emit `weights_companions` + `drafter` + `vision` (vision
  derived from the vision-coding workload).  drafter/vision also enable a future
  catalog badge.  test-registry-json contract updated.
- setup.sh: read WEIGHT_EXTRA_KEYS (a <model>:<variant> list) and fetch those
  alongside the core via the existing companion-download loop + per-file SHA
  verify.  setup.sh stays the weights-layer puller (model profiles); it does NOT
  depend on the slug registry, so the historical pre-registry "download + test a
  new model before it has a slug" path is untouched.
- cockpit: run_weights_download passes entry.weights_companions as
  WEIGHT_EXTRA_KEYS (model-qualified); weights_state is companion-aware (a slug
  is PARTIAL, Start gated + Download offered, until BOTH core and companions are
  on disk); CatalogEntry gains weights_companions / drafter / vision (attached to
  the row like `source`, no shared-core schema change).

+5 tests (row-attach, companion env injection, companion-aware state, setup.sh
WEIGHT_EXTRA_KEYS, registry-json contract).  Suite: tui-core 66, serve-cockpit
665 (+1 skip), repo gates 56/0.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-22 13:41:27 +00:00
noonghunnaandClaude Opus 4.8 eb81c5f0af profiles: wire hf_repo for 5 verified weight variants
Fill in the HuggingFace source repo for five weight variants that had no direct
download recipe (were "manual"), each verified live against the HF API:
  - gemma-4-31b:bf16          → google/gemma-4-31B-it
  - qwen3.6-27b:awq           → cyankiwi/Qwen3.6-27B-AWQ-INT4
  - qwen3.6-35b-a3b:gptq_int4 → palmfuture/Qwen3.6-35B-A3B-GPTQ-Int4
  - qwen3.6-35b-a3b:dflash    → z-lab/Qwen3.6-35B-A3B-DFlash
  - qwen3.6-35b-a3b:dflash_gguf → abhinand/Qwen3.6-35B-A3B-DFlash-GGUF

These are now fetchable via `WEIGHT_KEY=<model>:<variant> bash scripts/setup.sh
<model>` (and, once a compose adopts them, one-click in the c3 Download UX).
Weights-registry tests green.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-22 02:06:14 +00:00
noonghunnaandClaude Opus 4.8 1730e0405b weights.py: add list --json (download-state source for c3)
Batch-emit the static weights metadata for every (model, variant) with a
local_subdir — subdir / hf_repo / size_gb / verify_glob / status / kind — as a
single JSON array.  The serve-cockpit TUI shells this once to learn where each
slug's weights live + how to fetch them, then stats the dirs itself against its
own (user-configurable) model dir for the download-vs-serve state.  Purely
additive: a new subcommand, no FS access, `entry`/`lookup` unchanged.

Foundation for the upcoming Download UX (Download-vs-Start, listing progress).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-22 01:23:25 +00:00
noonghunnaandClaude Opus 4.8 bb2c30ebca c3 cockpit UX batch 3: honesty + serving actions (fit-vs-live-VRAM, real running config, targeted stop)
From the live-rig UX audit:

- A6 Catalog "fits" respects LIVE free-VRAM, not an empty card: a row whose vram_est
  exceeds live per-GPU free-VRAM (from the estate poll) is downgraded ●→⚠/✗ with a
  reason; with no live data the column reads "(vs empty card)". Never fabricates a
  free number; the live-serving model's own row is exempt (its VRAM is counted as
  used, so it is provably fitting).
- A7 Serving panel shows the REAL running config: probes /v1/models max_model_len
  (running ctx) + docker inspect image, and badges "config differs from catalog slug
  <slug>" only when the probe diverges from the slug's CONFIGURED ctx (registry
  max_ctx, exact int) — not the kv-calc fit ceiling. Registry-derived fields stay
  labelled "(per catalog slug)".
- A4 Targeted serving verbs: [k] stop / [b] restart resolve the serving container by
  matched_slug and route through the existing confirm/reconcile gate (vs [o] stop-ALL
  which kills co-resident services); [n] switch jumps to Run·Catalog.
- #4/N7 Doctor: [y] re-runs the (read-only) doctor read; the full-validation report is
  surfaced from Doctor; clear issues offer the obvious next action.
- A9 ③ Gate ladder shows each step's last outcome (·/⟳/✓/✗) from the run's exit.

Threads a new configured_ctx field through the registry --json contract
(registry-emit.sh → VariantRow → CatalogEntry) so the A7 badge compares the
configured ctx; unifies the K-label convention on registry-emit's ÷1000 so the
running-ctx and catalog labels match. All registry-parity gates pass.

Verified by an implement→gate→repair→verify→critic workflow; folded the findings: the
serving model's own row no longer false-"✗ won't fit"; the divergence badge compares
configured vs probed (not the fit ceiling) so an honest 262K serve on the 295K-fit
qwen doesn't false-badge; the ladder glyph awaits run completion before resolving
⟳→✓/✗ (was stuck at ⟳).

Suite green: cockpit 444; registry-json / switch- / launch-parity / compose-disk gates pass.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-21 00:58:44 +00:00
Will Hampson d03abf43d3 Fix estate GPU visibility overrides (#447)
Generate a per-instance Compose override for estate boots so CUDA_VISIBLE_DEVICES and NVIDIA_VISIBLE_DEVICES are set inside each container. This prevents parallel vLLM estate instances from both binding to physical GPUs 0/1 when Docker exposes all reserved devices.\n\nValidated with test-diagnose-estate, test-estate-json, test-profiles-compat, test-launch-compat, and a live 4x3090 estate boot.
2026-06-21 04:38:46 +05:00
noonghunnaandClaude Opus 4.8 d09a063317 Phase 2b: add --json/CLI contracts to 7 stack scripts (additive)
Data-layer contracts the cockpit (and any jq user) consumes — all strictly
additive (existing human output byte-identical), full guard suite green (54/54):
- registry-emit.sh --json : {variants,defaults,profiles{engines,models,hardware,drafters}}
- tools/kv-calc.py --fit <slug|model> --card <gpu> --json : structured fit verdict
- gpu-mode.sh --list-modes [--json] : scene catalog (serving/studio/ops)
- estate_cli.py report-state/diagnose --json : structured estate read
- pull.sh --profile-like --dry-run --json : structured swap_path (not a message blob)
- health.sh CONTAINER= : Doctor probes any engine container (was qwen36-27b-hardcoded)
- switch.sh --explain <slug> [--json] : joined registry/engine/model/hw/drafter + fit + bench

Built + adversarially reviewed via workflow. The review caught a real
switch<->kv-calc seam defect (switch fed hyphenated 'rtx-3090', kv-calc matched
only 'rtx3090' -> fit silently 'unavailable' on the 3090 rig); fixed: kv-calc
accepts hyphenated hardware-profile ids + hyphen-strip fallback; switch surfaces
kv-calc's structured verdict regardless of RC and renders its real keys; both
tests now exercise the seam. No shared-module edits.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-18 08:43:00 +00:00
noonghunnaandClaude Opus 4.8 79173b12d5 Record measured pi-reasoning rebench-full results + BENCHMARKS row
Full rebench-full (--with-8pack-thinking=both) completed 2026-06-18 for
llamacpp/qwen27b-pi-reasoning-single. Replaces the earlier inferred "≈ base
50/59 @370W" note (+ the PENDING-ladder caveats) with the measured numbers, and
adds the BENCHMARKS.md single-card row.

- Bench @370W: 47.9/55.3 decode narr/code (n=5, CV<2%); @230W cap 28.5/32.9.
- verify-stress 8/8 (NIAH->183K), 8-pack 104/150 off / 106/150 on, soak PASS
  (0 MiB growth, 0/100 silent-empty, p50 54.5, 102% retention).
- MTP head NOT weaker than base: matched-power A/B dead-even (73% vs 72% accept);
  47.9/55.3 @370W ~on par with base 50.3/58.9. Power-sensitive (mainline -42%
  370->230W) — the earlier "slow" 28/33 was the rig's 230W cap, not the model.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-18 02:23:00 +00:00
noonghunnaandClaude Opus 4.8 78c4038850 Correct pi-reasoning bench framing: 230W power-cap artifact, MTP head ≡ base
The merged llamacpp/qwen27b-pi-reasoning-single docs claimed its embedded MTP
head was "~45% below base / weaker" — that was a power-cap measurement artifact,
not a model property. A matched-power A/B (230W, same engine/KV/n=2/prompt) shows
the head performs IDENTICALLY to the base Qwen3.6-27B MTP head: 73% vs 72% draft
acceptance, 41.2 vs 41.2 t/s. The base llamacpp/default 50.3/58.9 figure I
compared against is a 370W BENCHMARKS number; mainline is -42% from 370->230W, so
the comparison was only valid at matched power (cf. the ik-llama/iq4ks-mtp
BENCHMARKS row, which documents exactly this trap). The 28.5/33.4 decode numbers
are correct AS 230W measurements; at 370W this config matches base's 50/59.
Corrected the status_note, compose header, and weights manual_note.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-17 23:46:31 +00:00
5da4eab12e Add llamacpp/qwen27b-pi-reasoning-single (Qwen3.6-27B Pi-style coding agent, mainline llama.cpp + MTP) (#425)
bytkim/Qwen3.6-27B-MTP-pi-reasoning Q4_K_M GGUF (embedded MTP head) — a "Pi-style"
reasoning-supervised CODING / terminal-agent fine-tune — on MAINLINE llama.cpp
(server-cuda-b9246, PR #22673), single 3090, q4_0/q4_0 KV + MTP, reasoning-ON.

- New compose: models/qwen3.6-27b/llama-cpp/compose/single/pi-reasoning-q4km/mtp.yml
- Registry slug llamacpp/qwen27b-pi-reasoning-single (experimental, port 8063).
- Weights entry pi-reasoning-q4km; drafter qwen-mtp-builtin (spec_method mtp).

CONFIG FOLLOWS THE MODEL CARD: temp 1.0 / top-p 0.95 / top-k 0 / min-p 0 (NOT the
stack's 0.6/20), reasoning ON, q4_0/q4_0 KV, --jinja -ngl 99 -fa. Card recommends
MTP n=3; on-rig A/B found n=2 marginally faster (within noise) — kept n=2,
MTP_DRAFT_N_MAX=3 matches the card. presence-penalty 1.5 is a documented knob for
the card's DIRECT/instruct (REASONING=off) mode.

CONTEXT (measured 2026-06-17, GPU0/GPU1): default 200K-alloc fills ~188K usable with
correct needle recall (22.7 GB / ~1.8 GB free; ~23 t/s decode at ~188K depth). Do NOT
alloc 262K — the FA scratch grows with the allocation, so 262K OOMs at ~176K (LESS
usable than 200K); full 262K usable is beellama-only. Author tested only 128K, so
128-188K is engine-proven but past the card's validated window (CTX_SIZE=131072 for
strict compliance).

BENCH (canonical bench.sh n=3, thinking-off, short-prompt): narrative 28.5 wall / 28.7
decode, code 32.9 / 33.4, PP 743 tok/s — ~45% below base llamacpp/default (50/59) on
identical engine/KV/MTP, i.e. this fine-tune's embedded MTP head is weaker. Engine A/B:
mainline ~25% faster than a beellama q4_0/q4_1 path → mainline chosen. Stays
experimental (--force): verify-stress / soak / quality ladder pending.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-18 03:56:28 +05:00
5cedde3572 Pull gate: actionable messages for uncurated derives of curated models (#424)
Three related [C0]/eligibility pull-gate fixes for the curated-swap surface —
an uncurated derive (abliterated / fine-tune) of a model we already serve.

1. Wrapper-arch alias. `pull.sh --profile-like` was false-aborting at [C0] with
   "no arch_patches matrix row for 'Qwen3_5ForConditionalGeneration'". That is
   the OUTER multimodal wrapper class the weights report; the patch matrix is
   keyed on the inner canonical `Qwen3NextForCausalLM`. arch_patches.yml is a
   closed key-set, so the alias lives in the editable arch_model_xref.
   - profile_runtime.yml: `config_architectures: [Qwen3_5ForConditionalGeneration]`
     on the Qwen3NextForCausalLM xref entry.
   - generate_compose.py: `resolve_arch_from_config()` maps a config.json
     architectures[0] string -> (canonical_arch, arch_row) via that alias.
   - gates.py [C0]: resolve the wrapper arch via the alias before declaring
     NO_ARCH_ROW. The hybrid now reports ENGINE_SUPPORTED.

2. GGUF axis. `supported_weight_formats` was declared on every engine but never
   enforced (only `kv_format` was). The deriver blocks GGUF on the derive path,
   but the curated registry / curated-swap path had no such guard. gates.py [C0]
   now rejects a `gguf` weight_format on an engine whose supported_weight_formats
   lacks `gguf` (structural axis; matches the `gguf` token only, so a derive's
   raw dtype spelling bf16/float16 is never false-rejected).

3. Won't-fit size advisory. The eligibility no-fit-model abort for a hybrid/MoE
   derive now appends (a) an actionable NOTE pointing at the curated-swap path +
   docs/BRING_YOUR_OWN.md, and (b) a coarse weights-only VRAM verdict: when the
   raw weights exceed the detected topology's total VRAM they won't fit at ANY
   KV, so say so concretely (the huihui abliterated bf16 ~54 GB vs 2×24 GB case)
   instead of a generic stop. `_weights_oversize_advisory()` is pure/total —
   empty when it fits / size unknown / headless.

Docs + tests:
   - BRING_YOUR_OWN.md: new section C — "Swap a curated model for a fine-tune /
     abliterated variant -> reuse its compose" (artifact↔engine + quant + MTP
     caveats, worked example).
   - test-pullgate-gates.sh: wrapper-arch ALIAS [C0] case + resolve_arch_from_config()
     unit + GGUF-on-vLLM runtime-incompatible + no-false-positive control +
     _weights_oversize_advisory() unit (oversize / fits / headless / malformed).

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-18 01:10:01 +05:00
ff4bbbc09d LMCache: in-repo L2 default + LMCACHE_L2=1 toggle, preflight disk-check, corrected sizing (#422)
* LMCache: in-repo L2 default + LMCACHE_L2=1 toggle, preflight disk-check, corrected sizing

- LMCACHE_L2=1 enables the L2 disk tier at a gitignored repo-root lmcache-kv/ (zero
  config); LMCACHE_L2_ADAPTER still fully overrides. L2 stays OFF by default.
- preflight_lmcache_ram now also soft-warns on low L2 disk space at the L2 host dir.
- CORRECTED capacity (measured > estimated): the LMCache offload cache is
  ~131 KB/token MEASURED (lmcache_mp_l1_memory_usage_bytes 4.93 GB + 4.6 GB on L2
  disk, 36,808-token session) — ~7x the GPU's 18.9 int8-PTH rate, which only sets
  max servable context. Prior docs used 18.9 for capacity and overstated it ~7x:
  --l1-size-gb 30 holds ~4 x 50K sessions (not ~33), <1 full 262K.
- Added a RAM/disk-vs-context sizing table to INTERNALS.md.
- Measured L2 rehydrate: 4.8s vs 43s cold re-prefill (~9x), cross-restart persistence
  confirmed; use the fs adapter (not nixl_store — this image's NIXL is broken).

Refs #133.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

* LMCache: document fp8 serde is broken on this model (no reduced-size L2 yet)

Measured 2026-06-17: LMCache's fp8 serde shape-errors on this model's KV
('[2,16,1584,260]' invalid — wrong layout for the hybrid GDN/head_dim-256/int8-PTH
source), 60 serialize fails, L2 stays empty. CacheGen-MP is 'coming soon'. So L2 is
uncompressed (~131 KB/token); the fp8-serde toggle was NOT wired (it doesn't work).
Re-test trigger noted. Refs #133.

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-18 00:06:54 +05:00
3531fd3551 Add opt-in LMCache KV-offload compose (vllm/qwen-27b-dual-lmcache, incubating) (#421)
Layers an LMCache tiered persistent prefix-KV cache (MP/HMA connector) onto the
dual-max fidelity profile (FP8 + int8-PTH KV + MTP n=3 @262K). Controlled A/B
shows ZERO decode penalty (74 narr / 94 code TPS == without LMCache, MTP intact
~83% accept); reuses cached prefix KV across long / multi-session workloads
(cold->warm TTFT ~7-8x) instead of re-prefilling.

- New vllm-lmcache engine profile: lmcache/vllm-openai DIGEST-pinned (the tag is
  mutable — bundled vLLM 0.22.1-dev then 0.23.1-dev under the same tag).
- preflight_lmcache_ram guard: rejects --l1-size-gb over-allocation even under
  --force (the l1=100-on-94GB host-OOM that forced a reboot, #133).
- env-gated L2 disk tier hook (LMCACHE_L2_ADAPTER, off by default).
- INTERNALS.md RAM/disk sizing section; registry entry + count bumps; full
  guard suite green.

Incubating (hidden, --force to launch): runs LMCache's third-party image with a
newer vLLM than our v0.22.0 pin; L2 latency unmeasured on-rig; 38 GB on-demand.

Refs #133.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-17 20:46:38 +05:00
noonghunnaandClaude Opus 4.8 8b3aa13f54 docs(vibethinker-3b): add measured one-shot-coding scores (HE+ 97%, LCB 83%)
Code packs run 2026-06-16 (tool-free — sandbox executes the model's code, no
tool-calls): humaneval-plus-30 29/30 (97%), lcb-v6-30 25/30 (83%, 2 losses =
token_limit on the hardest). Substantiates the 'strong one-shot/competition
coding' characterization with our own numbers (matches the authors' ~96% LeetCode).
Folded into the llamacpp compose Quality field + registry status_note.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-16 21:08:40 +00:00
noonghunnaandClaude Opus 4.8 d3cacfad0c feat(vibethinker-3b): add llamacpp/vibethinker-3b-single (Q8, incubating)
The performance-max VibeThinker path — prithivMLmods Q8_0 GGUF on mainline
llama.cpp, single 3090, q8_0 KV, full 131K, -b 4096 -ub 2048. A sibling to
the vLLM bf16 compose (#418), and the better one on every axis.

Why it's better than the vLLM sibling:
- ~166 TPS decode (vs ~110 vLLM bf16), ~6.2 GB at full 131K (vs ~9.8 GB),
  prefill ~6,630 tok/s (-b 4096/-ub 2048 = +20%), boot ~9s.
- Q8_0 is near-lossless → quality INTACT, where vLLM's fp8 weight-quant broke
  this quant-sensitive 3B (non-terminating empty output). KV quant is fine
  (storage-only); fp8 *weights* are the problem.

Validated 2026-06-16 (tool-free packs, temp 0.6, thinking-on):
  gsm-symbolic-30 100% · instructfollow-15 100% · reasonmath-15 80% ·
  structoutput-15 80% · dataextract-15 40%.
  dataextract is a genuine extraction ceiling — full temp sweep {0:27%,
  0.6:40%, 1.0:27%} can't lift it (value mismatches persist). A reasoning
  specialist, not an extractor. No tool-calling. verify-full 5/9 by design.

Wiring: registry entry + dense-family added to llama-cpp-local engine +
prithivmlmods-q8 weights variant + disk/registry counts (51/52). No MTP head
in the GGUF (plain qwen2 conversion) → no spec-dec. Guard suite 47/47.

Status: 🐣 incubating (hidden from switch.sh --list; --force to launch).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-16 19:53:05 +00:00
noonghunnaandClaude Opus 4.8 35a69650fe feat(catalog): add 🐣 Incubating status tier + VibeThinker-3B (incubating)
Introduce a new pre-experimental status tier, `incubating`, for composes that
work but aren't ready for the actionable list — niche specialists or models
that fail the standard functional gate by design. Incubating composes are
HIDDEN from `switch.sh --list` by default (revealed by `--list --all`) and
launch-gated behind `--force` (non-functional), so a half-validated model is
catalogued and discoverable without cluttering the recommended set.

Tier wiring (reuses the existing status plumbing — no new emit column):
- compose_registry.py: `incubating` in STATUS_VALUES + `🐣` in COMPOSE_STATUS_EMOJI
  (stays out of FUNCTIONAL_STATUSES → --force-gated, never auto-defaulted)
- switch.sh: both --list filters skip incubating unless --all, with a
  `(+N incubating hidden — --all)` header note
- AGENTS.md: Status enum row + Caveats-required references
- docs/ADDING_MODELS.md: new rule — NEW MODELS START at 🐣 Incubating, promote
  up the enum (🐣 → 🧪 → ⚠️/✅) as they earn the actionable list

First occupant — VibeThinker-3B (WeiboAI, Qwen2 dense reasoning fine-tune):
- `vllm/vibethinker-3b-single` — bf16 weights + fp8_e5m2 KV, single 3090,
  full 131072 ctx, mem_util 0.40 (~9.8 GB, single-concurrency sized),
  --reasoning-parser qwen3, no tool-calling
- Live-validated 2026-06-16: serves clean correct reasoning/code (~110 TPS),
  qwen3 parser splits <think>. fp8 WEIGHTS rejected (break the quant-sensitive
  3B: non-terminating empty output). Always-reasoning + no-tools → fails
  verify-full's fixed-small-budget checks (5/9) by design → incubating, not gate-passing.

Guard suite 47/47 green (nex-n2-mini parked aside; it's an unrelated local experiment).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-16 15:03:06 +00:00
noonghunnaandClaude Opus 4.8 186ab9b2b6 feat(qwen3.6-27b): add beellama/carnice-v2-dual-q8-mtp dual compose (#403)
Dual-GPU Carnice-V2-27B Q8_0 + embedded MTP head on beellama v0.3.2-preview
(layer-split -ts 0.55,0.45, q8_0/q8_0 KV, 262K). The dual / quality-max
follow-through requested in discussion #403.

Validated via rebench-full (2026-06-16, 2x 3090 PCIe):
- bench n=5: narr 40.7 / code 44.0 decode TPS, TTFT ~79 ms, PP 1197 t/s
- verify-full all-pass; verify-stress 8/8 (NIAH ladder -> 240K)
- soak fresh 20x5 PASS (0 growth, 0/100 silent-empty, p50 42.2, 100% retention)
- 8-pack think-OFF 103/150 / think-ON 105/150 (wash; in-band vs qwopus-coder)

Key decisions (measured A/Bs, captured in compose header + learnings):
- q8_0 KV over the requested kvarn6: +17% prefill (1003 vs 860 t/s; escapes
  KVarN software-compression compute, q4=q8=1004 so it's the path not the
  bit-width), higher fidelity, reference-aligned, fits 262K on dual. KVarN's
  compression only pays off on a tight single card.
- MTP-only; DFlash ruled out (only base-27B drafters exist -> ~10% accept on
  the fine-tune; no Carnice-matched drafter).
- n=2 = +13% validated opt-in (DRAFT_N_MAX=2); n=1 default.
- -b/-ub/--no-mmap A/B'd flat -> KV type was the only prefill lever.

Status: experimental (beellama v0.3.2 is a rolling pre-release; #455 un-park gate).

Catalog wiring: registry entry + qwen3.6-27b.yml carnice-v2-q8 weights variant
+ disk-count bump. Full guard suite green (47/47).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-16 11:57:35 +00:00
8c3c0cb40f Pin hauhaucs-35ba3b-dual weights to morikomorizz commit 49a080d (#319) (#413)
First use of the new `revision:` lever (#408). Pins the
morikomorizz/Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP fetch to commit
49a080db7406cd98f94f0f2a18539bbcc1520444 — the exact bytes the BENCHMARKS
row and #411 Results Card were measured against.

Verified: repo last-modified 2026-06-10 predates our 2026-06-14 download,
and the HF x-linked-etag at this sha equals our local file's sha256
(76a0d4c2...). weights.py round-trips revision: -> WEIGHT_REVISION; guard
suite green. Guards against a silent upstream re-quant (the #316 class).

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-15 06:50:20 +05:00
DeuceBucket 6be52294c8 feat(weights): optional revision: pin in weights-fetch schema (#319) (#408)
adds an optional `revision:` key per weights variant. weights.py emits
WEIGHT_REVISION; setup.sh threads it into `hf download --revision` and
pins the post-download sha-verify etag lookup to the same revision, so a
stale pin can't false-fail against a newer HEAD. preflight's manual hint
mirrors the flag. unset = track HEAD, so behavior is unchanged for every
current entry (nothing sets revision: today).

this is the weights half of #316: upstream quant repos re-quant
silently, and we had no lever to pin the bytes a BENCHMARKS row was
measured against. engine images already pin; weights didn't. mechanism
only, no real entry is pinned in this PR (that's a per-entry,
rig-validated call that's yours to make).

refs #319, #316
2026-06-15 06:35:43 +05:00
2abe025513 Add llamacpp/hauhaucs-35ba3b-dual uncensored MTP compose (🧪) (#410)
Wires morikomorizz/Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP (Q6_K_P GGUF
with an embedded nextn MTP head) as a dual-card mainline llama.cpp b9570
compose: -ts 0.55,0.45, q8_0 KV, MTP n=3, 262K, reasoning-on by default.

The MTP head loads clean on mainline ("speculative decoding context
initialized") — the prior HauhauCS-MTP ret=-3 was an ik-llama/older-build
issue, not the model arch. The -ts 0.55,0.45 split rebalances the MTP draft
card (even 1,1 skews ~3 GB at 262K).

Validated 2026-06-14: verify-stress 8/8 (NIAH ceiling ladder -> 240K),
bench.sh n=3 @262K (narr 113.4 / code ~150 decode TPS, CV<1%), soak fresh
20x5 PASS (0 growth, 0/100 silent-empty, p50 162.4, 99.6% retention),
8-pack think-OFF 103/150 / think-ON 105/150 (wash). n=3 vs n=1 @262K =
-9% prose / +10% code (code-leaning default by request; n=1 prose-best via
MTP_DRAFT_N_MAX=1).

Status 🧪 Experimental: community GGUF (digest-unpinned) + uncensored.
No DEFAULTS row — opt-in only. Guard suite green.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-15 06:28:08 +05:00
87a6dbb924 Add Carnice-V2-27B beellama single-card compose (beellama/carnice-v2-single-q5km-mtp) (#406)
* Add Carnice-V2-27B beellama single-card compose (beellama/carnice-v2-single-q5km-mtp)

stuchapin Carnice-V2-27B Q5_K_M GGUF (kai-os/Carnice-V2-27b, a Hermes-style agentic
SFT of Qwen3.6-27B) with embedded MTP head, on beellama v0.3.2-preview KVarN — single
3090, kvarn4/kvarn4 KV, MTP n=1, reasoning-on default.

Validated 2026-06-14 (rebench-full, reasoning-on): engine-compat PASS (beellama loads
the PR#22673-fused GGUF — the card's "mainline fails to load" does not apply),
verify-stress 8/8 (NIAH clean to 150K), soak PASS (0-growth, 0/100 silent-empty,
100.5% retention), bench 46.8/50.5 TPS narr/code, MTP accept ~94%. 8-pack reasoning-on
110/150 — beats sibling beellama/qwopus-coder 103/150 (edge is agentic/instruct).

- compose: models/qwen3.6-27b/beellama/compose/single/carnice-v2-q5km/mtp-kvarn4.yml
- drafter profile: scripts/lib/profiles/drafters/carnice-mtp-gguf.yml (n_default/n_max=1
  — mtp_num_hidden_layers=1; author warns n=3 is wrong)
- weights_variant carnice-v2-q5km + compose_registry entry (port 8068, kvcalc SKIP)
- bump test count assertions (registry 47, disk 48, drafters 11)

Stays experimental (beellama v0.3.2 is a rolling pre-release, the #455 engine gate),
not for any quality reason. n=1 is the card-faithful default; n=2 is a documented
+12%-TPS opt-in (DRAFT_N_MAX=2, does not crash on our single-card kvarn4) pending a
dedicated soak.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

* setup.sh: add WEIGHTS=carnice-v2 fetch knob for beellama/carnice-v2-single-q5km-mtp

UX parity with WEIGHTS=qwopus-coder — maps to qwen3.6-27b:carnice-v2-q5km so
`WEIGHTS=carnice-v2 bash scripts/setup.sh qwen3.6-27b` fetches the GGUF.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-14 17:04:47 +05:00
520f5b909b catalog: add beellama/qwopus-coder (Qwopus3.6-27B-Coder, KVarN-4) — first KVarN compose (#391)
* catalog: add beellama/qwopus-coder (Qwopus3.6-27B-Coder, KVarN-4) — first KVarN compose

Wires the Qwopus3.6-27B-Coder Q5_K_M GGUF (Jackrong, embedded MTP head) into
the curated catalog on beellama + the new KVarN-4 KV format. Single 3090,
kvarn4/kvarn4 @ 160K + MTP (230K no-MTP env opt-in). 🧪 experimental (--force),
gated on the v0.3.2-preview KVarN engine build (digest-pinned, merged earlier).

- engines/beellama-local.yml + hardware/rtx-3090.yml: register `kvarn4` KV format
- drafters/qwopus-mtp-gguf.yml: embedded-MTP-GGUF drafter (Jackrong)
- models/qwen3.6-27b.yml: qwopus-coder-mtp-q5km weights entry (Jackrong GGUF)
- compose_registry.py: `beellama/qwopus-coder` slug (kvcalc_key SKIP, port 8067)
- compose mtp.yml: git-tracked; image fallback → the KVarN digest (the v0.3.0
  multiarch fallback rejects kvarn*)
- setup.sh: `WEIGHTS=qwopus-coder` fetch knob
- bump stale count asserts (drafters 9→10, registry 45→46, disk 46→47)

Validated 2026-06-12 (sm_86): embedded MTP loads, verify-full all-pass, NIAH
@72K = q5_0/q4_1 control, bench ~46/58 TPS (KVarN decode-neutral), 8-pack
104/103 ≈ q5_0/q4_1 102/107 (quality-neutral; disc #329). scripts/tests/*.sh
green in a clean tree (the lone test-compose-registry-disk red is pre-existing
untracked nex-n2-mini WIP, #473 — not in this commit). pending soak.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

* qwopus-coder: soak PASS + launcher-path validated — drop pending-soak note

Launcher path (switch.sh beellama/qwopus-coder --force → KVarN digest injected →
kvarn4 boots → serves :8067) + verify-full all-pass + soak-continuous PASS
(0 MiB growth, 0/25 silent-empty, 100% TPS retention). Stays 🧪 (pre-release
engine); full verify-stress NIAH ladder is the only gate left for ⚠️ promotion.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-13 03:27:58 +05:00
ce71d4376a beellama: bump pin v0.3.0 → v0.3.2-preview (commit-pinned) + validate (#389)
* beellama: bump pin v0.3.0 → v0.3.2-preview (commit-pinned) + validate

Maintainer chose the v0.3.2 preview over the v0.3.1 stable for the newer
build (adds experimental KVarN KV-compression). v0.3.2 is a rolling
pre-release — Anbeeld replaces its moving Docker tags with newer branch
builds — so we pin the COMMIT-suffixed tag for an immutable pin:
  install.spec → ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.3.2-317c65e27e1e

Validated on-rig (single 3090, q5ks-dflash): boots on the preview image,
verify-full all-pass (Paris / tool_calls / streaming / thinking), prose
coherent, DFlash spec-dec active (acceptance ~0.12-0.17 on short tasks,
not collapsed). beellama has no vendored patches, so nothing to rebase.

Composes STAY 🧪 experimental: preview ≠ stable. The first stable tag now
exists (v0.3.1, server-cuda-v0.3.1, non-prerelease — Qwen3 MTP post-norm +
CUDA KV-quant fixes); repoint install.spec there to un-park (#455) once it
passes the full gate. Multiarch fallback (compose-literal, sm_120 direct-
compose) left at v0.3.0 — rebuild at the chosen tag is a separate follow-up.

scripts/tests/*.sh green (the 2 reds are pre-existing untracked-experimental-
compose artifacts: qwopus-coder + nex-n2-mini, unrelated to this pin).

Co-Authored-By: Claude Opus 4.8 <[email protected]>

* beellama: repoint pin to KVarN build BY DIGEST (the -317c65 tag lacks KVarN)

The commit-suffixed preview tag server-cuda-preview-v0.3.2-317c65e27e1e
PREDATES the KVarN merge — its --cache-type-k rejects kvarn* (only
turbo/TCQ). KVarN is only in the latest rolling server-cuda-preview-v0.3.2
build (commit 98caf25), which has no immutable commit-suffixed tag, so we
pin its DIGEST (immutable + KVarN), same as the vLLM :gemma digest pin:
  install.spec → ghcr.io/anbeeld/beellama.cpp@sha256:858e7cfbfb0d5d…

Measured (single 3090, Qwen3.6-27B Q5_K_S, -np 1): KVarN lifts the single-
request ceiling from ~196K (q5_0/q4_1; 262K OOMs) to the full 262K —
kvarn4 (≈q5_0 quality) fits 262K tight (~1GB free), kvarn2 ~3GB free.
Recall-at-depth NIAH + Qwopus-coder re-validate on the digest build pending.

Co-Authored-By: Claude Opus 4.8 <[email protected]>

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-06-13 02:56:45 +05:00
noonghunnaandClaude Opus 4.8 eae0e36f81 setup: fix fp8 weights nits after #369 (typo, help string, manual_note)
Follow-up cleanup to the merged #369:
- error message said "WEIGHTS=ft8" → "WEIGHTS=fp8"
- the unrecognized-WEIGHTS help string now lists 'fp8'
- the fp8 entry's manual_note said "no direct pull recipe wired" — now
  that hf_repo is set, document the WEIGHTS=fp8 setup.sh path (hf-download.sh
  kept as the stall-resistant manual option for the 29 GB layer-split repo)

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-12 11:08:36 +00:00
hlo-world 64821ac4a4 chore: make setup script work with WEIGHTS=FP8 and qwen3.6-27b (#369) 2026-06-12 16:07:02 +05:00
noonghunnaandClaude Opus 4.8 0cfd099b37 dgemma: adopt official vllm/vllm-openai:gemma image (drop 120-file sideload)
vLLM publishes an OFFICIAL `vllm/vllm-openai:gemma` image (pushed 2026-06-10, a
stock build of the dgemma branch commit 74b5964f) with DiffusionGemma baked in —
`DiffusionGemmaForBlockDiffusion` registers natively, transformers 5.10.2. So we
pin that image (BY DIGEST, purge-resistant) and drop the bespoke sideload from
PR #358 (stock nightly + 123-file branch overlay + install_script).

3 fixes are NOT upstream (vLLM tests H100/B200 + TP=1) so they're not in :gemma —
the compose now bind-mounts them (site_package_overlay) from the new lean dir
models/diffusiongemma-26b-a4b/vllm/patches/gemma-image-fixes/:
  - marlin.py + marlin_utils_fp8.py — sm_86 fp8 Marlin sub-tile-K pad. :gemma
    clean dies in warmup ("Invalid thread config ... num_bits=8 ...
    max_shared_mem=101376", K=352/1056) without it.
  - diffusion_gemma.py — TP-vocab soft-embed + dtype fix (TP=2; their recipe is TP=1).

Changes:
  - base.yml: image -> :gemma@sha256:9c719fc0...; default `vllm serve` entrypoint
    + 3 file mounts (was: overlay-dir mount + install_script bash entrypoint).
    Status 🧪 experimental (was upstream-gated; supersedes PR #359 too).
  - engine vllm-diffusion-gemma: install.spec -> :gemma@digest; vendored_overlays
    -> the 3 fixes (delivery site_package_overlay).
  - registry status -> experimental; note rewritten.
  - patches.yml: dgemma-gemma-image-fixes (site_package_overlay, 3 overlay_files).
  - diagnose_profile_cli OVERLAY_PATH_HINTS + docs/UPSTREAM.md #45163 row.
  - DELETE the 123-file dgemma-overlay/ + the regeneration Dockerfile.

Validated live on 2x RTX 3090 (2026-06-11): :gemma clean dies on the Marlin wall;
:gemma + 3 mounts (via `docker compose -f base.yml up`) boots, serves coherent
output, 262K, ~177/180 TPS typical / ~1100 peak, 23.1 GB/card. verify-full: gen +
tool-call + reasoning + output-quality pass (streaming-SSE "1 chunk" is the
expected block-diffusion artifact). Full gate green (test-compose-registry-disk
local-only red = untracked nex-n2-mini WIP; CI-clean validated: disk 46 / reg 45).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-11 05:31:36 +00:00
noonghunnaandClaude Opus 4.8 af1e909b04 Wire DiffusionGemma into the catalog via stock-nightly sideload
Make `vllm/diffusiongemma-dual` usable through setup/pull/switch/launch, and
pivot the engine delivery from a baked local image to a SIDELOADED overlay on a
stock, pullable nightly (so non-rig users can actually run it).

Delivery (no baked image):
- Engine `vllm-diffusion-gemma` pins a STOCK `vllm/vllm-openai:nightly-2c9c07c8…`
  and sideloads the vendored overlay at boot via install_script (entrypoint).
- `models/diffusiongemma-26b-a4b/vllm/patches/dgemma-overlay/` is the vendored
  payload: the FULL stock-vs-`dgemma`-branch `vllm/` delta (123 .py) — a lean
  PR-#45163-only overlay version-skews (`build_attn_metadata() got an unexpected
  keyword argument 'causal'`: load-bearing dgemma changes live outside the PR
  diff) — plus Codex's 3 fixes (marlin K-pad ×2 + diffusion_gemma TP-vocab/dtype).
  install.sh cp's it over the installed vllm pkg, fail-loud arch assert.
- base.yml rewired: stock image + overlay mount + install_script entrypoint;
  drops the 3 standalone marlin-k-pad mounts (folded into the overlay) and the
  baked-image dependency. The dgemma-pr45163 Dockerfile is now the overlay-
  regeneration helper, not the runtime engine.

Catalog wiring:
- Model profile `diffusiongemma-26b-a4b` (gemma4-swa-moe backbone + block
  diffusion; text; valid_tp [1,2]; kv_calc_supported=false → kvcalc_key SKIP).
- Engine profile `vllm-diffusion-gemma` (stock install + vendored_overlays).
- compose_registry `vllm/diffusiongemma-dual` (port 8042, status upstream-gated).
  NO DEFAULTS row — upstream-gated is non-functional, reachable only by slug.
- patches.yml `dgemma-vllm45163-sideload` (install_script + drift_guard) and the
  diagnose_profile_cli overlay-path hint.
- docs/UPSTREAM.md row for vllm#45163 (re-pin + drop triggers).

Output-length fix carried in base.yml: lift the model's 256-tok max_new_tokens
default to 16384 via --override-generation-config (keeps the diffusion denoising
params; fixes OWUI truncation + next-turn echo).

Tests: bump the count fixtures for +1 model / +1 engine / +1 compose / +1 registry
entry (test-profiles-compat 6→7 models, 11→12 engines; test-compose-registry-disk
44→45 registry, 45→46 disk). Full gate green except test-compose-registry-disk
local-only redness from untracked nex-n2-mini WIP in the shared tree (CI-clean
validated green: tracked-only disk=46, registry=45, disk-only set all-allowed).

Validated live on 2x RTX 3090: stock nightly + sideloaded overlay boots, serves
coherent output, 262K, via `docker compose -f base.yml up` (the switch/launch path).
Upstream-gated until #45163 merges into a pinnable engine.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-11 04:00:30 +00:00
noonghunnaandClaude Opus 4.8 93acbf979f Promote Deckard-40B to ✅ Production; fix arch + slug naming
The arch-confirm blocker named in the 🧪 status note is resolved: the GGUF
header reports general.architecture=qwen35 (standard GQA, 97 layers), NOT the
Qwen3-Next/DeltaNet hybrid the original scaffold guessed. Corrected the profile
to family=qwen35-dense and dropped the fictional GDN/linear-attention fields.

Fixes found while correcting the profile:
- Restored the required `attention_k_eq_v: false` field (compat.py reads it with
  bracket notation — dropping it broke load_profiles for the whole catalog).
- Fixed a duplicate-key bug in the weights block: a second `hf_repo:` (DavidAU
  base, no MTP head) shadowed PiehSoft's MTP GGUF under yaml.safe_load
  last-key-wins — setup.sh would have fetched the wrong, head-less repo.
- Added qwen35-dense to llama-cpp-local supported_model_families (the C10 gate).

Renamed the weights slug mtp-q6k → piehsoft-q6k to follow the provider-quant
idiom used by every other GGUF compose (ubergarm-iq4ks, unsloth-q4km,
byteshape-iq4xs) and to drop the redundant mtp in mtp-q6k/mtp.yml. The serving
filename stays mtp.yml (KV isn't encoded in GGUF-engine filenames). Renamed the
compose dir + on-disk weights dir + all path references in lockstep.

Validation (2× 3090, server-cuda-b9570): verify-full 8/8, verify-stress 8/8
(ceiling ladder filled to 120K / 91% of n_ctx, 0 MiB VRAM growth, needle recall
clean through 90K), 8-pack 105/150 with MTP off==on (spec-dec lossless),
soak-continuous PASS. Guard suite 42/0.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-10 00:38:15 +00:00
noonghunnaandClaude Opus 4.8 111c72b0ea switch.sh --owui: auto-register a launched model in Open WebUI
New opt-in flag: after a catalog model is up + ready, switch.sh --owui upserts
an OpenAI connection in Open WebUI (host.docker.internal:<port>) so it appears
in the chat picker — no manual Admin->Connections step. scripts/lib/owui-register.sh
is conditional (no-op if OWUI not running), idempotent (skips if already present),
and drives OWUI's admin config API via a token forged from the container secret.
Validated end-to-end on Deckard-40B (:8199).

Guard suite 41/42; the 1 failure (test-measurement-record) is pre-existing
bench-output-format drift, unrelated to this change (no switch.sh/owui reference).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-10 00:07:19 +00:00
noonghunnaandClaude Opus 4.8 d6725faf02 Deckard-40B: fix provenance — GGUF is PiehSoft's, wire hf_repo fetch
The MTP-injected GGUF is NOT a local build (the scaffold's comment was wrong) —
it's published at PiehSoft/Qwen3.6-40B-Deckard-MTP-Q6_K. Wire hf_repo + files so
setup.sh/launch.sh auto-fetch it, and correct the compose/profile comments.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-09 23:04:14 +00:00
noonghunnaandClaude Opus 4.8 20e1d6f362 Deckard-40B: record soak PASS + final 105/150 (gates all green)
Soak-continuous PASS (0 MiB growth, 0/25 silent-empty, 25 turns), 8-pack
105/150 with MTP off==on (spec-dec lossless), verify-full 8/8. Updates the
compose caveats, registry status_note, and BENCHMARKS row from 'pending' to
the measured results. Stays 🧪 Experimental pending profile config.json arch
confirm before ✅.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-09 22:43:57 +00:00
noonghunna 5dafaa61fe feat: add Qwen3.6-40B-Deckard to catalog (llamacpp/deckard40B-dual-mtp)
Dense 40B uncensored Qwen3.6 community merge, dual-3090 llama.cpp with
embedded MTP head (Q6_K GGUF, 31 GB). First dual llama.cpp compose in
the catalog; first 'category' field in the registry (uncensored).

Changes:
- Compose: models/qwen3.6-40b-deckard/llama-cpp/compose/dual/mtp-q6k/mtp.yml
  Status: 🧪 Unverified (quality /150 + soak pending)
  Config: -ngl 99 -ts 1,1 -fa on --spec-type draft-mtp --spec-draft-n-max 2
          --cache-type-k q8_0 --cache-type-v q8_0 -c 131072
- Registry: llamacpp/deckard40B-dual-mtp (category=uncensored, experimental)
- DEFAULTS: (qwen3.6-40b-deckard, llamacpp, dual) → slug
- Model profile: scripts/lib/profiles/models/qwen3.6-40b-deckard.yml
- Drafter: qwen-mtp-builtin model_compat extended with qwen3.6-40b-deckard
- LiteLLM route: deckard-40b → :8199
- launch.sh: suggest_default_variant case for Deckard
- BENCHMARKS.md: new section with validated MTP n=2 numbers
- Test fixtures: registry count 43→44, disk 44→49, models 5→6

Wrinkle decisions (see PR description):
1. Dual llama.cpp: uses 'count: all' + '-ts 1,1' (tensor-split layer-split),
   matching the validated serving config. No launcher changes needed — the
   compose's deploy section handles GPU reservation directly.
2. Engine pin: uses rolling server-cuda tag (matching engine profile spec),
   not the b9246 pin other composes use. The validated build (2026-06-09
   digest 1c4ff61a) is newer than b9246 and has draft-mtp working. No engine
   pin bump for other models.

NOT DONE (live validation — GPUs busy with quality eval):
- Boot/bench/soak/quality NOT run — maintainer to validate post-eval
- Status stays 🧪 until full gate passes
2026-06-09 21:59:13 +00:00
74b30abfe3 Add vllm/qwen-27b-dual-balanced + record 3-way 8-pack A/B (tie) (#343)
* Add vllm/qwen-27b-dual-balanced + record 3-way 8-pack A/B (tie)

New "balanced" dual tier: cyankiwi AWQ-BF16-INT4 (int4 group-32 + BF16 mtp,
compressed-tensors) + int8-PTH KV + MTP n=3, TP=2 @262K. Live-validated
2026-06-07 (Marlin WNA16, KV pool 370K tok / 1.41x — the LARGEST of the dual
family, since int4 weights free more VRAM than fp8's 8-bit; ~67 TPS decode).

3-way 8-pack A/B (--full, same harness, same day 2026-06-07):
  fast (autoround+fp8KV) 109 · balanced (awq+int8KV) 105 · max (fp8+int8KV) 110
  → a TIE (deterministic packs 64/64/65; spread within +/-5-7 noise).

The short-context 8-pack does NOT separate the quants. So:
- fast (vllm/dual) stays the default;
- balanced/max are framed by SERVING characteristics (KV-pool headroom,
  int8-PTH fidelity), NOT behavioral quality — headers/docs say so explicitly.
- The dimension where int8-PTH should win (long-ctx NIAH recall) is an open
  follow-up, not claimed here.

Also corrects the stale fast-tier 129/150 baseline: that was the 2026-05-09
harness; today's benchlocal-cli (verifier fixes since) scores the same fast
config at 109/150. Annotated as not-comparable across all four compose headers.

Wiring: awq-bf16-int4 weights_variant (compressed-tensors, cyankiwi repo) +
registry entry (vllm-stable, kv int8_per_token_head, port 8016, SKIP kvcalc,
experimental) + dual-max status_note updated with the A/B + DUAL_CARD rows +
A/B footnote. test-compose-registry-disk 42->43 / 43->44. Full gate green 43/0;
compat C4/C10/C14 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* Correct balanced KV-pool claim: fast has the largest pool, not balanced

Measured v0.22.0 @262K, TP=2: fast 622K/2.37× > balanced 370K/1.41× > max
295K/1.13×. My initial framing claimed balanced had "the largest KV pool of
the dual family" — wrong: I only compared it to max and forgot fast. Fast's
17.5 GB autoround weights are far lighter than balanced's 27 GB AWQ, so fast
has the biggest pool by a wide margin.

This reframes balanced honestly: it is DOMINATED by the fast tier — slower
(~67 vs ~89 code TPS), smaller KV pool, and tied/below on the 8-pack. Its
only possible edge is int8-PTH KV fidelity > fast's fp8_e5m2 (same size — a
fidelity bet, not a memory one), which is UNPROVEN (the short-ctx 8-pack is
blind to it). Keep only if the long-ctx NIAH A/B (#470) proves it; else
deprecate. Fixed across the compose header, registry status_note, and
DUAL_CARD (rows + footnote); fast row now cites its 622K/2.37× pool.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

---------

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-08 02:25:29 +05:00
d77d4acfac Add Qwen3.6-27B fast/max tiers across dual + multi4 (4 slugs) (#340)
A symmetric 4-slug family for the Qwen vLLM path:
  *-fast = AutoRound INT4 + fp8_e5m2 KV  (peak TPS, the proven path)
  *-max  = official FP8     + int8-PTH KV (higher fidelity @ 262K)

Slugs:
  vllm/qwen-27b-dual-fast   alias of vllm/dual (AutoRound INT4, TP=2) — production
  vllm/qwen-27b-dual-max    FP8 + int8-PTH, TP=2 — 🧪 live-validated 2026-06-07
  vllm/qwen-27b-multi-fast  AutoRound INT4, TP=4 — 🧪 cross-rig (dual is the proxy)
  vllm/qwen-27b-multi-max   FP8 + int8-PTH, TP=4 — 🧪 cross-rig (dual-max is the proxy)

dual-max validated live on this 2-card rig: FP8 weights load via MarlinFP8
W8A16 on Ampere sm_86, int8_per_token_head KV keeps the full 262K (KV pool
295K tok / 1.13x concurrency), MTP n=3 active, coherent. The multi4 configs
are byte-identical to their dual sibling apart from TP + gpu-count, so they
ship Experimental until a real >=4x 3090 host validates them (this dev rig
has 2 cards). No multi4 DEFAULT added — the resolver degrades to vllm/dual
(honest; no model carries a multi default), and the slugs stay reachable by name.

Catalog wiring:
- models/qwen3.6-27b.yml: new fp8 weights_variant (official e4m3 release,
  layer-split safetensors + embedded mtp.safetensors + vision tower)
- vllm-stable.yml: + fp8 weight format (Marlin W8A16 on Ampere) and
  + int8_per_token_head KV — NATIVE on stock v0.22.0 for uniform-head-dim
  models (qwen3-next); features flag stays false (no overlay, unlike Gemma's
  #40391-provided int8-PTH on vllm-gemma-stable)
- compose_registry.py: 4 entries; test-compose-registry-disk counts 38->42 / 40->43

Full gate green (42/42); compat C4/C5/C10/C14 pass for all 4 slugs.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-07 19:34:26 +05:00
noonghunnaandClaude Opus 4.8 31dc2c4756 Set 31B w4a16 default MTP n=4->3 (n-swept optimum) + update A/B record
n-sweep @370W (dual TP=2): n2 72.9/86.1 . n3 74.0/87.7 . n4 71.6/87.8 wall-TPS
(all within ~3% - the QAT-int4's fast acceptance decay caps the spec benefit).
n=3 is the best balance (top narrative, code tied with n=4, ~20% less drafting:
43% draft-accept vs n=4's 37%), so SPEC_N_MAX default 4->3. The w4a16 compose had
inherited n=4 from the autoround clone, but autoround's slower decay (88/73/59/48%
per-pos) earns n=4 while the QAT-int4's faster decay (64/39/25% at n=3) favors a
lower n - optimal MTP n is quant-specific. Updates the compose default + Quality
field, registry status_note, BENCHMARKS row. Quality unchanged (n = speed only).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-07 02:08:11 +00:00
noonghunnaandClaude Opus 4.8 97b678e405 Record 31B w4a16 8-pack A/B: 109/150 vs autoround 105 (comparable quality, weaker MTP)
vllm/gemma-31b-qat-w4a16-dual A/B vs the autoround-int4 default (gemma-int8-mtp),
same int8-PTH KV + assistant MTP n=4 (byte-identical speculative-config):
- 8-pack 109/150 vs 105 (+4, within +/-5-7 noise = quality tie; real instructfollow
  edge IF 15-vs-8 offset by hermes/cli/TC/RM).
- WEAKER spec-decode: MTP accept-len ~2.4 vs autoround's ~3.9 (per-position
  acceptance lower at every position) -> ~30% lower TPS (71.6/87.8 @370W <
  autoround 106/139 @230W despite MORE power; AL is the clean power-indep signal).
- Boots clean on stock vllm-gemma-stable (no #44494 workaround - tower arch).
Comparable quality but slower -> autoround-int4 stays the default; stays Experimental.
Updates the compose Quality field + registry status_note + a BENCHMARKS row.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-07 01:44:36 +00:00
noonghunna 579a717201 Experimental Gemma-4 QAT W4A16 vLLM composes + kv-calc int4 fix (#339)
Two 🧪 experimental QAT-W4A16 vLLM slugs (vllm/gemma-12b-qat-w4a16-single, vllm/gemma-31b-qat-w4a16-dual) + the standalone kv-calc gemma-12B int4 weight-pricing fix. 12B needs the gemma4-unified-vision-unquant sitecustomize workaround (vLLM #44494, both bugs, self-contained); 31B boots clean on stock vllm-gemma-stable. Full suite 42/0.
2026-06-07 05:26:16 +05:00
noonghunnaandClaude Opus 4.8 3a5ece7bb5 gemma-26ba4b-single: promote INT8-PTH single → ⚠️ Production-w/-caveats (gate PASS)
rebench-full (gemma-26ba4b-int8r, 1x 3090 @370W) PASSED every gate:
  verify-full ✓ · bench 168 narr / 217 code TPS (MTP AL 3.06-3.79) ·
  verify-stress NIAH-clean → 161K (91% of 176K) ·
  quality 109/150 think-ON (98/150 off) — on par with gemma-4-31B's 107/150 ·
  soak 20x5 PASS, 0 MiB growth.

Flip Status 🧪 Experimental → ⚠️ Production-w/-caveats in the int8.yml header
(+ Quality line) and compose_registry (status experimental→caveats + gate-result
note). Add the BENCHMARKS.md row. Caveats: needs the #40391 overlay; 176K @
mem_util 0.94 (262K only without the MTP drafter — drafter weights + cudagraph
capture cost ~86K tok, 0.96 OOMs the capture tail); think-OFF agentic/extraction
softer (cli-40 30% / DataExtract 60%, recover to 52%/73% with thinking). Gate 42/42.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-06 14:17:50 +00:00
noonghunnaandClaude Opus 4.8 e2efb757a9 Repoint vllm/gemma-26ba4b-single to INT8-PTH long-ctx (#465)
Repoint the single-card Gemma-4-26B-A4B slug in place from the bf16/16K
AWQ+MTP path to an INT8-per-token-head long-context path. Same weights
(cyankiwi AWQ-4bit, Marlin WNA16 MoE), same external MTP drafter (n=4),
same gemma4 tool-call; only the KV format changes (bf16 -> int8_per_token_head)
via the vendored vLLM PR #40391 hybrid-SWA KV page-size overlay. The dual
(vllm/gemma-26ba4b-dual) is untouched.

INT8-PTH (1 byte/token) lifts the single-card ceiling from bf16 16K to a
176K default. Live-validated on 1x RTX 3090 (2026-06-06): #40391 applies
cleanly, int8_per_token_head KV inits, Marlin WNA16 MoE backend, MTP
SpeculativeConfig active; KV pool 183,357 tok >= 176K; coherent generation
(post cudagraph-warmup) + clean gemma4 tool-call. mem_util 0.94, not 0.96:
0.96 passed the upfront KV check but OOM'd in the drafter's later cudagraph
capture (240 MiB free, needed 256) on a single 24 GB card.

- engine vllm-gemma-stable: + gemma4-swa-moe family, + awq/compressed-tensors
  weight formats
- patches.yml: + gemma-a4b-vllm-pr40391-rebased (model-scoped copy of the
  patch dir; reaches the compose, passes test-patch-attribution)
- compose single/awq/mtp.yml -> single/awq/int8.yml (full rewrite to the
  #40391 overlay + entrypoint + int8 KV; Status: Experimental)
- registry + profile_runtime: engine/kv/max_ctx/mem_util/compose_path/entrypoint
- test-generate-from-profile: repoint the clean-derived-seed fixture off this
  slug (now overlay-carrying) to vllm/gemma-26ba4b-dual + a tp= override

Status stays Experimental; rebench-full + soak are the #464 follow-up. Full
gate suite green (test-submit-bench pre-existing/environmental, fails on
baseline too).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-06 11:30:27 +00:00
noonghunnaandClaude Opus 4.8 95e47fedc9 kv-calc: per-sequence KV-pool floor (capped at 1 GB), fix KV-light false-FAIL
The C12 budget verdict used a flat absolute floor (`MIN_KV_GB = 0.05 if
qwen3-next-moe else 1.0`) and FAILed any config leaving <1 GB for the growing
KV pool. That false-FAILs KV-light models that demonstrably boot: gemma-4-26b
-a4b single-card AWQ + MTP projects 0.78 GB growing pool, yet the live boot
holds 17,490 tok >= 16,384 max_ctx and serves (sliding_window=1024 makes 25/30
layers' KV trivially cheap).

vLLM's real pre-check is token-capacity based: it boots iff the capped pool
holds >= ONE max_model_len sequence, NOT iff KV >= 1 GB. Growing KV scales
linearly with max_num_seqs, so one sequence's KV is
`kv_pool_requested_gb / max_num_seqs`. Use that as the floor — which also
generalizes (and removes) the old qwen3-next-moe 0.05 special-case.

CAP the floor at the legacy 1 GB: for dense/long-KV configs one sequence needs
many GB and a hard per-seq threshold would false-FAIL measured-working configs
sitting inside the estimator's +-1.5 GB band (e.g. gemma-dual-int8 @262K:
10.71 GB avail vs 10.84 GB/seq, 1.2% short but boots). Net: relax the floor
ONLY for KV-light models, leave dense on its exact prior >=1 GB behavior.

The fix lives in the estimator (where the gap is), NOT in disabling
kv_calc_supported for a config we actively ship the drafter on. Design
bounced with Codex (per-sequence floor); capped-at-1-GB refinement added here
after the pure per-seq form regressed the gemma-4-31b dense duals.

Verdicts after: gemma-26ba4b-single FAIL -> TIGHT (correct); dense duals
unchanged; --calibration still 7/7; full gate 42/42.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-06 01:41:10 +00:00
noonghunnaandClaude Opus 4.8 95e1448217 Enable MTP on gemma-4-26b-a4b single + ladder both to max ctx (#326)
Per request: turn on the external MTP drafter for BOTH topologies and
ladder each compose to its empirically-validated max context.

- single/awq/base.yml -> single/awq/mtp.yml (rename per ADDING_MODELS.md:
  drafter=MTP named, bf16 KV is the gemma4-on-Ampere default so NOT named).
  Adds --speculative-config (gemma-26b-it-assistant n=4) + prefix-caching +
  chunked-prefill (parity with the dual). Max ctx 8K -> 16K: KV pool 17,490
  tok at 0.92 util on 1x 3090 with the drafter loaded (17,408 would leave
  <100 tok headroom). Registry drafter None -> gemma-26b-it-assistant.
- dual/awq/mtp.yml: max ctx 32K -> 262,144 (model max). KV pool 806,821 tok
  at 262144/0.92 on 2x 3090 — sliding_window=1024 keeps 25/30 layers' KV
  cheap, so the model max fits with 3x headroom.

Both LIVE-validated on the rig 2026-06-06 (stock v0.22.0): single tp=1 +
MTP boots (SpeculativeConfig method=mtp n=4), coherent, clean gemma4
tool-call; dual tp=2 + MTP @262K boots, coherent, clean tool-call.

Registry max_ctx 16384/262144 + single compose_path -> mtp.yml; profile_runtime
mirrors (single gains the speculative_config slot, dual maxlen 262144).
test-generate-from-profile mk_einput now clears the base slug's drafter (the
clean-derived fixtures need a no-drafter seed; the drafter-refuse fixture sets
one explicitly) so it's stable now that gemma-26ba4b-single carries an MTP drafter.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-06 01:40:53 +00:00
noonghunnaandClaude Opus 4.8 48eaf8283c Wire gemma-4-26b-a4b AWQ on stock vLLM v0.22.0; retire AutoRound (#326)
The gemma-4-26b-a4b slugs were the last users of the purged
`vllm-nightly-clean` engine. Migrate them to AWQ-4bit on stock
`vllm-stable` v0.22.0, retiring the Ampere-dead AutoRound path.

Why AutoRound is dead on Ampere: AutoRound int4-MIXED packs to
`uint8b128`, which has no W4A16 kernel on sm_86 (Cutlass/Machete are
Hopper-only; Triton W4A16 takes only uint4b8/uint4). The cyankiwi
AWQ-4bit (pure uint4, compressed-tensors) resolves to Marlin WNA16 MoE
on sm_86 and serves cleanly.

Slug changes (registry 38 -> 36):
  - REMOVE vllm/gemma-a4b, vllm/gemma-a4b-awq (AutoRound/AWQ-bf16 duals)
  - RENAME vllm/gemma-a4b-single    -> vllm/gemma-26ba4b-single (AWQ tp=1, :8040)
  - RENAME vllm/gemma-a4b-awq-mtp   -> vllm/gemma-26ba4b-dual   (AWQ tp=2 + ext MTP, :8041)
Both on vllm-stable, kv bf16 (kv_cache_dtype=auto), kvcalc_key=SKIP
(experimental; rebench-full/soak pending before Production).

PR #40886 (AWQ compressed-tensors MoE key remap) is native in v0.22.0 —
the vendored overlay + install entrypoint are dropped; patch marked
deprecated_on 2026-06-06. vllm-nightly-clean deprecated (final
purged-nightly engine, now zero registry users); gemma4-swa-moe added
to vllm-stable supported_model_families; arch_patches engine_pin
flipped to [email protected].0 loads:true.

Boot-validated on 2x 3090 (2026-06-06): single tp=1 coherent + clean
gemma4 tool-call; dual tp=2 + external MTP (num_spec_tokens=4) coherent
+ clean tool-call, draft layers wired. 3 AutoRound/AWQ-bf16 composes
moved to _archive/ (kept for a future Hopper engine).

Gate: 42/42 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-06 00:46:13 +00:00
noonghunnaandClaude Opus 4.8 cc3e906f8c chore(327): archive the Qwen3.6-27B DFlash path + deprecate vllm-nightly-dflash
#327 audit found all 3 DFlash slugs rode the purged nightly-e47c98ef image (the
same dead SHA as the deprecated vllm-nightly-full), and vllm/dual4-dflash was
mislabeled `production` while (a) on a dead image, (b) TP=4 (unrunnable on this
2-card rig), (c) DFlash is blocked on Qwen3-Next per the repo hardware truths
(DeltaNet rollback). Same treatment as the #254 Genesis tail:

- Archived dual-dflash, dual-dflash-noviz, dual4-dflash -> compose/_archive/
  (+ manifest); removed their 3 registry entries (41 -> 38).
- DEFAULTS[(qwen,vllm,multi4)] removed -- no functional multi4 vLLM slug remains,
  so the resolver degrades multi4 -> dual. The 4-card vLLM tier is now empty
  (recoverable from git if demand returns).
- vllm-nightly-dflash engine -> stability: deprecated; its qwen3-next-hybrid arch
  pin flipped loads:true -> false (zero registry users; not offered for derive).
- Cascade: calibration (-3 rows), kv-calc alias map + override, patches.yml
  (-18 dflash refs), test fixtures (registry-disk 38/40, calibration 9/9;
  launch-compat/list-topology-filter/model-default-resolver/pullgate-gates/pull
  adapted to the now-empty multi4 + nightly-slug categories).

Completes the nightly-engine cleanup: vllm-nightly-mtp/full/dflash all deprecated
with zero registry users; only the gated vllm-nightly-clean (#326) remains.
42/42 gate green. Refs #327.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-05 15:26:29 +00:00
noonghunnaandClaude Opus 4.8 baac1acafd chore(254): deprecate the now-unused Genesis nightly engines + patches [Phase 3+4]
Phase 3 — engine profiles (zero registry users after the archival):
- vllm-nightly-mtp + vllm-nightly-full -> stability: deprecated + DEPRECATED notes
  (both purged nightlies, 404 on Docker Hub). nightly-mtp retained as the
  genesis_equipped test anchor (required_genesis:true).
- arch_patches.yml: flipped the qwen3-next-hybrid vllm-nightly-mtp pin
  loads:true -> false (deriver won't offer the dead engine; [email protected].0
  stays the loads:true primary).
- generate-compose.sh usage example + docs/UPSTREAM.md engine-pin rows updated.

Phase 4 — patches.yml: stamped 55 patches deprecated_on:2026-06-05 (kept on disk):
- the ~46 genesis-p*/pn* env-gated patches (via the &genesis_env_patch anchor)
- 9 dead overlays: sglang x2, pr40798/pr40914 (negative-result), gemma-pr41800
  (merged upstream), gemma4-fp8-ampere (Ampere-dead), perheadkv-hybridpage-fix +
  pr40391-perheadkv (superseded by pr40391-rebased), carnice-chat-template
  (carnice compose archived).
- Left 9 ACTIVE: overlays still mounted by functional composes + the gated
  gemma-a4b #326 patch.

42/42 gate green. Refs #254.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-05 14:44:21 +00:00
noonghunnaandClaude Opus 4.8 b5ef9e697f chore(254): archive 15 Genesis/nightly composes + migrate the cascade [WIP: 5 test fixtures]
Archive (per #254, user-approved "drop registry"):
- 15 Genesis/nightly vLLM composes -> models/qwen3.6-27b/vllm/compose/_archive/
  + revival manifest (README.md): the 9 vllm-nightly-mtp slugs, dual-int8
  (nightly-full), tools-text + dual-bf16 (nightly-clean deprecated), and the 3
  Genesis-stack slugs dual4/carnice/qwopus.
- Removed their 15 COMPOSE_REGISTRY entries (56 -> 41).

Cascade migrated to keep the catalog coherent:
- calibration/qwen3.6-27b.yml: dropped 5 archived-slug rows.
- patches.yml: dropped 69 archived-slug refs from compose lists.
- tools/kv-calc.py: pruned the legacy alias map + dual-turbo/dual4 overrides
  (--calibration now 13/13).
- DEFAULTS (qwen,vllm,multi4): vllm/dual4 -> vllm/dual4-dflash.
- compose-meta.sh: compose_hw_model_status qwen candidates repointed off the two
  archived composes -> minimal.yml (fixes a real setup.sh "doesn't fit" bug).
- test fixtures: registry-disk counts (41/43), kv-generic-dense + submit-pull
  calibration counts (13/13), model-default-resolver (dual4-dflash + dual-dflash pin).

WIP: 5 test files still anchor on archived slugs (handed to Codex to finish to
42/42 green). Refs #254.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-06-05 13:37:48 +00:00