DeepReinforce agentic-coding RL fine-tune of Qwen3.6-35B-A3B (same qwen35moe arch, no MTP head, text-only). ik_llama dual 3090, Q8_0, full 262K, drafter-free ngram (opt-in). Gate PASS: ~108.8/108.6 TPS, verify-stress 8/8, soak PASS, 8-pack 105/105. Ties the base on the 8-pack (110), EDGES it on aider (15/30 vs 12-13) -> coding-leaning lane. Catalog suite 58/58 green.
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
The OWUI picker listed models that weren't actually serving: the studio
image/video/audio lanes (always shown via the studio pipe) and every
LLM catalog model (always shown via the always-up LiteLLM :4000
gateway). Both error until you switch to the matching gpu-mode scene.
Now the picker reflects what's live.
Studio lanes — gate the pipe's pipes() on backend liveness:
- build_studio_pipe.py (source template) + regenerated studio_pipe.py:
add a _LANES catalog tagged by backend (ComfyUI :8188 for all media
lanes, the on-demand voice service :8193 for the voice lane), a cheap
reachability probe (_alive, short timeout) with an 8s TTL cache, and
filter pipes() to only the lanes whose backend answers. New valve
hide_unavailable_lanes (default true) toggles it off. When the
ai-studio scene is down, the Studio group drops out of the picker
instead of listing dead lanes.
LLMs — point OWUI at each backend directly instead of the gateway:
- services/openwebui/docker-compose.yml: seed one connection per backend
(director :8090 + :8010/:8051/:8032/:8038/:8199), drop :4000. OWUI
hides models from an unreachable connection, so each model shows only
while its scene serves. Served names match the catalog (IDs unchanged).
LiteLLM stays up for other clients; it just leaves OWUI's picker.
- scripts/lib/owui-unregister.sh: new symmetric inverse of
owui-register.sh (remove a connection by port; forged-JWT config API,
idempotent, no-op if OWUI down / not present).
- scripts/setup-ai-studio.sh: idempotent step that registers the 6
per-backend connections + drops a stale :4000, so existing installs
converge (fresh installs get it from the compose env seed).
Live-validated on the rig (gemma12b scene): picker LLMs = only
gemma-4-12b-int8 (the four down models hidden, :4000 gone), Studio group
hidden (ComfyUI + voice down). Probe verified both directions (live
endpoint -> shown even on a 404 path; dead port -> hidden).
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
v0.26.0 (comfyanonymous/ComfyUI #14589) adds native LOCAL Krea2 support —
the earlier "cloud-only, dropped" verdict was pin-specific (cb9f6394
predated it; only the cloud Krea2ImageNode existed → local load failed
"Could not detect model type").
- entrypoint.sh: COMFYUI_REF default cb9f6394 -> f6c162dd (v0.26.0); reword
the pin comment (Krea2 + Qwen3-VL text-gen; re-validate ALL lanes on bump).
- docker-compose.yml: pin comment v0.26.0 + Krea2.
- download_krea.sh (new): fetches the 3 Comfy-Org/Krea-2 assets (turbo fp8
DiT + Qwen3-VL-4B encoder + Qwen-Image VAE) into the ComfyUI models tree.
- download_studio_models.sh + studio-models.tsv: add Krea as an image-lane
model (roster <-> manifest kept mirrored).
NOT yet merge-ready: the v0.26.0 pin is a load-bearing bump — the custom
HiDream-O1 node, ComfyUI-GGUF, and DisTorch multi-GPU all ride it. On-rig
re-validation of every studio lane (Krea renders + the existing 9 still
work) is the merge gate. The OWUI lane graph (krea2.json workflow +
studio_pipe entry) is authored at that validation step (the Krea sampler /
Qwen3-VL encode graph differs from Z-Image — not clonable blind).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Wire the two new lanes into the install/preflight surfaces so a fresh rig fetches them
and c3 / gpu-mode know they're expected:
- download_zimage.sh — Z-Image-Turbo fp8 + Qwen3-4B encoder + flux ae VAE (~12 GB).
- download_wan.sh — Wan2.2-Rapid Mega NSFW Q8 GGUF + umt5 encoder + Wan 2.1 VAE (~25 GB).
- download_studio_models.sh ("grab all") calls both (idempotent — skips what's present).
- studio-models.tsv: +image/Z-Image, +video/Wan2.2 rows (the shared manifest read by both
c3's StudioModel loader and gpu-mode preflight). Roster note records Krea2 as dropped.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
Replace the separate `image-studio` / `video-studio` / `comfyui` scenes with
a single `ai-studio` scene: ComfyUI holds both GPUs always, and image / video /
audio are picked as *lanes* in OWUI rather than gpu-mode switches. Removes the
gemma-12b chat brain from the studio bundle (kept on disk, just unwired) and
points OWUI's DEFAULT_MODELS at the qwen director.
- gpu-mode.sh: `mode_ai_studio` replaces the two studio modes; drop the gemma
start/stop + the GPU0-pin split; preflight now reads the shared manifest.
- scripts/lib/studio-models.tsv: single source-of-truth model manifest (modality,
label, root, rel_path, size, installer) consumed by both gpu-mode preflight and
the c3 cockpit, so the two can't drift.
- download_{director,video,audio,ace_step,stable_audio,studio}_models.sh: idempotent
fetchers; download_studio_models.sh orchestrates "grab everything".
- setup-video-studio.sh + repointed setup-image-studio.sh → `gpu-mode ai-studio`.
- litellm: drop the gemma-4-12b :8069 route; comfyui compose/entrypoint GPU-pin
comments corrected (no scene sets the pin now).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
A catalog slug's compose can mount a SEPARATE weight artifact beyond its core
weights — a DFlash draft model or an mmproj vision projector, from a different
subdir. The cockpit Download fetched only the core, so those 8 slugs (5 DFlash
+ 3 GGUF-vision) would download, read "present", offer Start, then fail to boot
for the missing companion.
The registry is the single source of truth: the cockpit gets every catalog slug
from it, so the slug's required artifacts come from it too.
- compose_registry.py: new per-slug `weights_companions` field — the extra
weight-variant keys a slug needs beyond `weights_variant`. Set on the 5 DFlash
slugs (-> anbeeld-dflash-iq4xs) + 3 vision slugs (-> gguf_mmproj_f16). Per-slug,
so a text slug sharing a vision GGUF variant (e.g. llamacpp/default on
unsloth-q4km) does NOT over-fetch the ~1GB mmproj.
- registry-emit --json: emit `weights_companions` + `drafter` + `vision` (vision
derived from the vision-coding workload). drafter/vision also enable a future
catalog badge. test-registry-json contract updated.
- setup.sh: read WEIGHT_EXTRA_KEYS (a <model>:<variant> list) and fetch those
alongside the core via the existing companion-download loop + per-file SHA
verify. setup.sh stays the weights-layer puller (model profiles); it does NOT
depend on the slug registry, so the historical pre-registry "download + test a
new model before it has a slug" path is untouched.
- cockpit: run_weights_download passes entry.weights_companions as
WEIGHT_EXTRA_KEYS (model-qualified); weights_state is companion-aware (a slug
is PARTIAL, Start gated + Download offered, until BOTH core and companions are
on disk); CatalogEntry gains weights_companions / drafter / vision (attached to
the row like `source`, no shared-core schema change).
+5 tests (row-attach, companion env injection, companion-aware state, setup.sh
WEIGHT_EXTRA_KEYS, registry-json contract). Suite: tui-core 66, serve-cockpit
665 (+1 skip), repo gates 56/0.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Fill in the HuggingFace source repo for five weight variants that had no direct
download recipe (were "manual"), each verified live against the HF API:
- gemma-4-31b:bf16 → google/gemma-4-31B-it
- qwen3.6-27b:awq → cyankiwi/Qwen3.6-27B-AWQ-INT4
- qwen3.6-35b-a3b:gptq_int4 → palmfuture/Qwen3.6-35B-A3B-GPTQ-Int4
- qwen3.6-35b-a3b:dflash → z-lab/Qwen3.6-35B-A3B-DFlash
- qwen3.6-35b-a3b:dflash_gguf → abhinand/Qwen3.6-35B-A3B-DFlash-GGUF
These are now fetchable via `WEIGHT_KEY=<model>:<variant> bash scripts/setup.sh
<model>` (and, once a compose adopts them, one-click in the c3 Download UX).
Weights-registry tests green.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Batch-emit the static weights metadata for every (model, variant) with a
local_subdir — subdir / hf_repo / size_gb / verify_glob / status / kind — as a
single JSON array. The serve-cockpit TUI shells this once to learn where each
slug's weights live + how to fetch them, then stats the dirs itself against its
own (user-configurable) model dir for the download-vs-serve state. Purely
additive: a new subcommand, no FS access, `entry`/`lookup` unchanged.
Foundation for the upcoming Download UX (Download-vs-Start, listing progress).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
From the live-rig UX audit:
- A6 Catalog "fits" respects LIVE free-VRAM, not an empty card: a row whose vram_est
exceeds live per-GPU free-VRAM (from the estate poll) is downgraded ●→⚠/✗ with a
reason; with no live data the column reads "(vs empty card)". Never fabricates a
free number; the live-serving model's own row is exempt (its VRAM is counted as
used, so it is provably fitting).
- A7 Serving panel shows the REAL running config: probes /v1/models max_model_len
(running ctx) + docker inspect image, and badges "config differs from catalog slug
<slug>" only when the probe diverges from the slug's CONFIGURED ctx (registry
max_ctx, exact int) — not the kv-calc fit ceiling. Registry-derived fields stay
labelled "(per catalog slug)".
- A4 Targeted serving verbs: [k] stop / [b] restart resolve the serving container by
matched_slug and route through the existing confirm/reconcile gate (vs [o] stop-ALL
which kills co-resident services); [n] switch jumps to Run·Catalog.
- #4/N7 Doctor: [y] re-runs the (read-only) doctor read; the full-validation report is
surfaced from Doctor; clear issues offer the obvious next action.
- A9 ③ Gate ladder shows each step's last outcome (·/⟳/✓/✗) from the run's exit.
Threads a new configured_ctx field through the registry --json contract
(registry-emit.sh → VariantRow → CatalogEntry) so the A7 badge compares the
configured ctx; unifies the K-label convention on registry-emit's ÷1000 so the
running-ctx and catalog labels match. All registry-parity gates pass.
Verified by an implement→gate→repair→verify→critic workflow; folded the findings: the
serving model's own row no longer false-"✗ won't fit"; the divergence badge compares
configured vs probed (not the fit ceiling) so an honest 262K serve on the 295K-fit
qwen doesn't false-badge; the ladder glyph awaits run completion before resolving
⟳→✓/✗ (was stuck at ⟳).
Suite green: cockpit 444; registry-json / switch- / launch-parity / compose-disk gates pass.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Generate a per-instance Compose override for estate boots so CUDA_VISIBLE_DEVICES and NVIDIA_VISIBLE_DEVICES are set inside each container. This prevents parallel vLLM estate instances from both binding to physical GPUs 0/1 when Docker exposes all reserved devices.\n\nValidated with test-diagnose-estate, test-estate-json, test-profiles-compat, test-launch-compat, and a live 4x3090 estate boot.
Data-layer contracts the cockpit (and any jq user) consumes — all strictly
additive (existing human output byte-identical), full guard suite green (54/54):
- registry-emit.sh --json : {variants,defaults,profiles{engines,models,hardware,drafters}}
- tools/kv-calc.py --fit <slug|model> --card <gpu> --json : structured fit verdict
- gpu-mode.sh --list-modes [--json] : scene catalog (serving/studio/ops)
- estate_cli.py report-state/diagnose --json : structured estate read
- pull.sh --profile-like --dry-run --json : structured swap_path (not a message blob)
- health.sh CONTAINER= : Doctor probes any engine container (was qwen36-27b-hardcoded)
- switch.sh --explain <slug> [--json] : joined registry/engine/model/hw/drafter + fit + bench
Built + adversarially reviewed via workflow. The review caught a real
switch<->kv-calc seam defect (switch fed hyphenated 'rtx-3090', kv-calc matched
only 'rtx3090' -> fit silently 'unavailable' on the 3090 rig); fixed: kv-calc
accepts hyphenated hardware-profile ids + hyphen-strip fallback; switch surfaces
kv-calc's structured verdict regardless of RC and renders its real keys; both
tests now exercise the seam. No shared-module edits.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Full rebench-full (--with-8pack-thinking=both) completed 2026-06-18 for
llamacpp/qwen27b-pi-reasoning-single. Replaces the earlier inferred "≈ base
50/59 @370W" note (+ the PENDING-ladder caveats) with the measured numbers, and
adds the BENCHMARKS.md single-card row.
- Bench @370W: 47.9/55.3 decode narr/code (n=5, CV<2%); @230W cap 28.5/32.9.
- verify-stress 8/8 (NIAH->183K), 8-pack 104/150 off / 106/150 on, soak PASS
(0 MiB growth, 0/100 silent-empty, p50 54.5, 102% retention).
- MTP head NOT weaker than base: matched-power A/B dead-even (73% vs 72% accept);
47.9/55.3 @370W ~on par with base 50.3/58.9. Power-sensitive (mainline -42%
370->230W) — the earlier "slow" 28/33 was the rig's 230W cap, not the model.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
The merged llamacpp/qwen27b-pi-reasoning-single docs claimed its embedded MTP
head was "~45% below base / weaker" — that was a power-cap measurement artifact,
not a model property. A matched-power A/B (230W, same engine/KV/n=2/prompt) shows
the head performs IDENTICALLY to the base Qwen3.6-27B MTP head: 73% vs 72% draft
acceptance, 41.2 vs 41.2 t/s. The base llamacpp/default 50.3/58.9 figure I
compared against is a 370W BENCHMARKS number; mainline is -42% from 370->230W, so
the comparison was only valid at matched power (cf. the ik-llama/iq4ks-mtp
BENCHMARKS row, which documents exactly this trap). The 28.5/33.4 decode numbers
are correct AS 230W measurements; at 370W this config matches base's 50/59.
Corrected the status_note, compose header, and weights manual_note.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Three related [C0]/eligibility pull-gate fixes for the curated-swap surface —
an uncurated derive (abliterated / fine-tune) of a model we already serve.
1. Wrapper-arch alias. `pull.sh --profile-like` was false-aborting at [C0] with
"no arch_patches matrix row for 'Qwen3_5ForConditionalGeneration'". That is
the OUTER multimodal wrapper class the weights report; the patch matrix is
keyed on the inner canonical `Qwen3NextForCausalLM`. arch_patches.yml is a
closed key-set, so the alias lives in the editable arch_model_xref.
- profile_runtime.yml: `config_architectures: [Qwen3_5ForConditionalGeneration]`
on the Qwen3NextForCausalLM xref entry.
- generate_compose.py: `resolve_arch_from_config()` maps a config.json
architectures[0] string -> (canonical_arch, arch_row) via that alias.
- gates.py [C0]: resolve the wrapper arch via the alias before declaring
NO_ARCH_ROW. The hybrid now reports ENGINE_SUPPORTED.
2. GGUF axis. `supported_weight_formats` was declared on every engine but never
enforced (only `kv_format` was). The deriver blocks GGUF on the derive path,
but the curated registry / curated-swap path had no such guard. gates.py [C0]
now rejects a `gguf` weight_format on an engine whose supported_weight_formats
lacks `gguf` (structural axis; matches the `gguf` token only, so a derive's
raw dtype spelling bf16/float16 is never false-rejected).
3. Won't-fit size advisory. The eligibility no-fit-model abort for a hybrid/MoE
derive now appends (a) an actionable NOTE pointing at the curated-swap path +
docs/BRING_YOUR_OWN.md, and (b) a coarse weights-only VRAM verdict: when the
raw weights exceed the detected topology's total VRAM they won't fit at ANY
KV, so say so concretely (the huihui abliterated bf16 ~54 GB vs 2×24 GB case)
instead of a generic stop. `_weights_oversize_advisory()` is pure/total —
empty when it fits / size unknown / headless.
Docs + tests:
- BRING_YOUR_OWN.md: new section C — "Swap a curated model for a fine-tune /
abliterated variant -> reuse its compose" (artifact↔engine + quant + MTP
caveats, worked example).
- test-pullgate-gates.sh: wrapper-arch ALIAS [C0] case + resolve_arch_from_config()
unit + GGUF-on-vLLM runtime-incompatible + no-false-positive control +
_weights_oversize_advisory() unit (oversize / fits / headless / malformed).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
* LMCache: in-repo L2 default + LMCACHE_L2=1 toggle, preflight disk-check, corrected sizing
- LMCACHE_L2=1 enables the L2 disk tier at a gitignored repo-root lmcache-kv/ (zero
config); LMCACHE_L2_ADAPTER still fully overrides. L2 stays OFF by default.
- preflight_lmcache_ram now also soft-warns on low L2 disk space at the L2 host dir.
- CORRECTED capacity (measured > estimated): the LMCache offload cache is
~131 KB/token MEASURED (lmcache_mp_l1_memory_usage_bytes 4.93 GB + 4.6 GB on L2
disk, 36,808-token session) — ~7x the GPU's 18.9 int8-PTH rate, which only sets
max servable context. Prior docs used 18.9 for capacity and overstated it ~7x:
--l1-size-gb 30 holds ~4 x 50K sessions (not ~33), <1 full 262K.
- Added a RAM/disk-vs-context sizing table to INTERNALS.md.
- Measured L2 rehydrate: 4.8s vs 43s cold re-prefill (~9x), cross-restart persistence
confirmed; use the fs adapter (not nixl_store — this image's NIXL is broken).
Refs #133.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
* LMCache: document fp8 serde is broken on this model (no reduced-size L2 yet)
Measured 2026-06-17: LMCache's fp8 serde shape-errors on this model's KV
('[2,16,1584,260]' invalid — wrong layout for the hybrid GDN/head_dim-256/int8-PTH
source), 60 serialize fails, L2 stays empty. CacheGen-MP is 'coming soon'. So L2 is
uncompressed (~131 KB/token); the fp8-serde toggle was NOT wired (it doesn't work).
Re-test trigger noted. Refs #133.
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Layers an LMCache tiered persistent prefix-KV cache (MP/HMA connector) onto the
dual-max fidelity profile (FP8 + int8-PTH KV + MTP n=3 @262K). Controlled A/B
shows ZERO decode penalty (74 narr / 94 code TPS == without LMCache, MTP intact
~83% accept); reuses cached prefix KV across long / multi-session workloads
(cold->warm TTFT ~7-8x) instead of re-prefilling.
- New vllm-lmcache engine profile: lmcache/vllm-openai DIGEST-pinned (the tag is
mutable — bundled vLLM 0.22.1-dev then 0.23.1-dev under the same tag).
- preflight_lmcache_ram guard: rejects --l1-size-gb over-allocation even under
--force (the l1=100-on-94GB host-OOM that forced a reboot, #133).
- env-gated L2 disk tier hook (LMCACHE_L2_ADAPTER, off by default).
- INTERNALS.md RAM/disk sizing section; registry entry + count bumps; full
guard suite green.
Incubating (hidden, --force to launch): runs LMCache's third-party image with a
newer vLLM than our v0.22.0 pin; L2 latency unmeasured on-rig; 38 GB on-demand.
Refs #133.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Code packs run 2026-06-16 (tool-free — sandbox executes the model's code, no
tool-calls): humaneval-plus-30 29/30 (97%), lcb-v6-30 25/30 (83%, 2 losses =
token_limit on the hardest). Substantiates the 'strong one-shot/competition
coding' characterization with our own numbers (matches the authors' ~96% LeetCode).
Folded into the llamacpp compose Quality field + registry status_note.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The performance-max VibeThinker path — prithivMLmods Q8_0 GGUF on mainline
llama.cpp, single 3090, q8_0 KV, full 131K, -b 4096 -ub 2048. A sibling to
the vLLM bf16 compose (#418), and the better one on every axis.
Why it's better than the vLLM sibling:
- ~166 TPS decode (vs ~110 vLLM bf16), ~6.2 GB at full 131K (vs ~9.8 GB),
prefill ~6,630 tok/s (-b 4096/-ub 2048 = +20%), boot ~9s.
- Q8_0 is near-lossless → quality INTACT, where vLLM's fp8 weight-quant broke
this quant-sensitive 3B (non-terminating empty output). KV quant is fine
(storage-only); fp8 *weights* are the problem.
Validated 2026-06-16 (tool-free packs, temp 0.6, thinking-on):
gsm-symbolic-30 100% · instructfollow-15 100% · reasonmath-15 80% ·
structoutput-15 80% · dataextract-15 40%.
dataextract is a genuine extraction ceiling — full temp sweep {0:27%,
0.6:40%, 1.0:27%} can't lift it (value mismatches persist). A reasoning
specialist, not an extractor. No tool-calling. verify-full 5/9 by design.
Wiring: registry entry + dense-family added to llama-cpp-local engine +
prithivmlmods-q8 weights variant + disk/registry counts (51/52). No MTP head
in the GGUF (plain qwen2 conversion) → no spec-dec. Guard suite 47/47.
Status: 🐣 incubating (hidden from switch.sh --list; --force to launch).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Introduce a new pre-experimental status tier, `incubating`, for composes that
work but aren't ready for the actionable list — niche specialists or models
that fail the standard functional gate by design. Incubating composes are
HIDDEN from `switch.sh --list` by default (revealed by `--list --all`) and
launch-gated behind `--force` (non-functional), so a half-validated model is
catalogued and discoverable without cluttering the recommended set.
Tier wiring (reuses the existing status plumbing — no new emit column):
- compose_registry.py: `incubating` in STATUS_VALUES + `🐣` in COMPOSE_STATUS_EMOJI
(stays out of FUNCTIONAL_STATUSES → --force-gated, never auto-defaulted)
- switch.sh: both --list filters skip incubating unless --all, with a
`(+N incubating hidden — --all)` header note
- AGENTS.md: Status enum row + Caveats-required references
- docs/ADDING_MODELS.md: new rule — NEW MODELS START at 🐣 Incubating, promote
up the enum (🐣 → 🧪 → ⚠️/✅) as they earn the actionable list
First occupant — VibeThinker-3B (WeiboAI, Qwen2 dense reasoning fine-tune):
- `vllm/vibethinker-3b-single` — bf16 weights + fp8_e5m2 KV, single 3090,
full 131072 ctx, mem_util 0.40 (~9.8 GB, single-concurrency sized),
--reasoning-parser qwen3, no tool-calling
- Live-validated 2026-06-16: serves clean correct reasoning/code (~110 TPS),
qwen3 parser splits <think>. fp8 WEIGHTS rejected (break the quant-sensitive
3B: non-terminating empty output). Always-reasoning + no-tools → fails
verify-full's fixed-small-budget checks (5/9) by design → incubating, not gate-passing.
Guard suite 47/47 green (nex-n2-mini parked aside; it's an unrelated local experiment).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
First use of the new `revision:` lever (#408). Pins the
morikomorizz/Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP fetch to commit
49a080db7406cd98f94f0f2a18539bbcc1520444 — the exact bytes the BENCHMARKS
row and #411 Results Card were measured against.
Verified: repo last-modified 2026-06-10 predates our 2026-06-14 download,
and the HF x-linked-etag at this sha equals our local file's sha256
(76a0d4c2...). weights.py round-trips revision: -> WEIGHT_REVISION; guard
suite green. Guards against a silent upstream re-quant (the #316 class).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
adds an optional `revision:` key per weights variant. weights.py emits
WEIGHT_REVISION; setup.sh threads it into `hf download --revision` and
pins the post-download sha-verify etag lookup to the same revision, so a
stale pin can't false-fail against a newer HEAD. preflight's manual hint
mirrors the flag. unset = track HEAD, so behavior is unchanged for every
current entry (nothing sets revision: today).
this is the weights half of #316: upstream quant repos re-quant
silently, and we had no lever to pin the bytes a BENCHMARKS row was
measured against. engine images already pin; weights didn't. mechanism
only, no real entry is pinned in this PR (that's a per-entry,
rig-validated call that's yours to make).
refs #319, #316
* beellama: bump pin v0.3.0 → v0.3.2-preview (commit-pinned) + validate
Maintainer chose the v0.3.2 preview over the v0.3.1 stable for the newer
build (adds experimental KVarN KV-compression). v0.3.2 is a rolling
pre-release — Anbeeld replaces its moving Docker tags with newer branch
builds — so we pin the COMMIT-suffixed tag for an immutable pin:
install.spec → ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.3.2-317c65e27e1e
Validated on-rig (single 3090, q5ks-dflash): boots on the preview image,
verify-full all-pass (Paris / tool_calls / streaming / thinking), prose
coherent, DFlash spec-dec active (acceptance ~0.12-0.17 on short tasks,
not collapsed). beellama has no vendored patches, so nothing to rebase.
Composes STAY 🧪 experimental: preview ≠ stable. The first stable tag now
exists (v0.3.1, server-cuda-v0.3.1, non-prerelease — Qwen3 MTP post-norm +
CUDA KV-quant fixes); repoint install.spec there to un-park (#455) once it
passes the full gate. Multiarch fallback (compose-literal, sm_120 direct-
compose) left at v0.3.0 — rebuild at the chosen tag is a separate follow-up.
scripts/tests/*.sh green (the 2 reds are pre-existing untracked-experimental-
compose artifacts: qwopus-coder + nex-n2-mini, unrelated to this pin).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
* beellama: repoint pin to KVarN build BY DIGEST (the -317c65 tag lacks KVarN)
The commit-suffixed preview tag server-cuda-preview-v0.3.2-317c65e27e1e
PREDATES the KVarN merge — its --cache-type-k rejects kvarn* (only
turbo/TCQ). KVarN is only in the latest rolling server-cuda-preview-v0.3.2
build (commit 98caf25), which has no immutable commit-suffixed tag, so we
pin its DIGEST (immutable + KVarN), same as the vLLM :gemma digest pin:
install.spec → ghcr.io/anbeeld/beellama.cpp@sha256:858e7cfbfb0d5d…
Measured (single 3090, Qwen3.6-27B Q5_K_S, -np 1): KVarN lifts the single-
request ceiling from ~196K (q5_0/q4_1; 262K OOMs) to the full 262K —
kvarn4 (≈q5_0 quality) fits 262K tight (~1GB free), kvarn2 ~3GB free.
Recall-at-depth NIAH + Qwopus-coder re-validate on the digest build pending.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Follow-up cleanup to the merged #369:
- error message said "WEIGHTS=ft8" → "WEIGHTS=fp8"
- the unrecognized-WEIGHTS help string now lists 'fp8'
- the fp8 entry's manual_note said "no direct pull recipe wired" — now
that hf_repo is set, document the WEIGHTS=fp8 setup.sh path (hf-download.sh
kept as the stall-resistant manual option for the 29 GB layer-split repo)
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Make `vllm/diffusiongemma-dual` usable through setup/pull/switch/launch, and
pivot the engine delivery from a baked local image to a SIDELOADED overlay on a
stock, pullable nightly (so non-rig users can actually run it).
Delivery (no baked image):
- Engine `vllm-diffusion-gemma` pins a STOCK `vllm/vllm-openai:nightly-2c9c07c8…`
and sideloads the vendored overlay at boot via install_script (entrypoint).
- `models/diffusiongemma-26b-a4b/vllm/patches/dgemma-overlay/` is the vendored
payload: the FULL stock-vs-`dgemma`-branch `vllm/` delta (123 .py) — a lean
PR-#45163-only overlay version-skews (`build_attn_metadata() got an unexpected
keyword argument 'causal'`: load-bearing dgemma changes live outside the PR
diff) — plus Codex's 3 fixes (marlin K-pad ×2 + diffusion_gemma TP-vocab/dtype).
install.sh cp's it over the installed vllm pkg, fail-loud arch assert.
- base.yml rewired: stock image + overlay mount + install_script entrypoint;
drops the 3 standalone marlin-k-pad mounts (folded into the overlay) and the
baked-image dependency. The dgemma-pr45163 Dockerfile is now the overlay-
regeneration helper, not the runtime engine.
Catalog wiring:
- Model profile `diffusiongemma-26b-a4b` (gemma4-swa-moe backbone + block
diffusion; text; valid_tp [1,2]; kv_calc_supported=false → kvcalc_key SKIP).
- Engine profile `vllm-diffusion-gemma` (stock install + vendored_overlays).
- compose_registry `vllm/diffusiongemma-dual` (port 8042, status upstream-gated).
NO DEFAULTS row — upstream-gated is non-functional, reachable only by slug.
- patches.yml `dgemma-vllm45163-sideload` (install_script + drift_guard) and the
diagnose_profile_cli overlay-path hint.
- docs/UPSTREAM.md row for vllm#45163 (re-pin + drop triggers).
Output-length fix carried in base.yml: lift the model's 256-tok max_new_tokens
default to 16384 via --override-generation-config (keeps the diffusion denoising
params; fixes OWUI truncation + next-turn echo).
Tests: bump the count fixtures for +1 model / +1 engine / +1 compose / +1 registry
entry (test-profiles-compat 6→7 models, 11→12 engines; test-compose-registry-disk
44→45 registry, 45→46 disk). Full gate green except test-compose-registry-disk
local-only redness from untracked nex-n2-mini WIP in the shared tree (CI-clean
validated green: tracked-only disk=46, registry=45, disk-only set all-allowed).
Validated live on 2x RTX 3090: stock nightly + sideloaded overlay boots, serves
coherent output, 262K, via `docker compose -f base.yml up` (the switch/launch path).
Upstream-gated until #45163 merges into a pinnable engine.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
The arch-confirm blocker named in the 🧪 status note is resolved: the GGUF
header reports general.architecture=qwen35 (standard GQA, 97 layers), NOT the
Qwen3-Next/DeltaNet hybrid the original scaffold guessed. Corrected the profile
to family=qwen35-dense and dropped the fictional GDN/linear-attention fields.
Fixes found while correcting the profile:
- Restored the required `attention_k_eq_v: false` field (compat.py reads it with
bracket notation — dropping it broke load_profiles for the whole catalog).
- Fixed a duplicate-key bug in the weights block: a second `hf_repo:` (DavidAU
base, no MTP head) shadowed PiehSoft's MTP GGUF under yaml.safe_load
last-key-wins — setup.sh would have fetched the wrong, head-less repo.
- Added qwen35-dense to llama-cpp-local supported_model_families (the C10 gate).
Renamed the weights slug mtp-q6k → piehsoft-q6k to follow the provider-quant
idiom used by every other GGUF compose (ubergarm-iq4ks, unsloth-q4km,
byteshape-iq4xs) and to drop the redundant mtp in mtp-q6k/mtp.yml. The serving
filename stays mtp.yml (KV isn't encoded in GGUF-engine filenames). Renamed the
compose dir + on-disk weights dir + all path references in lockstep.
Validation (2× 3090, server-cuda-b9570): verify-full 8/8, verify-stress 8/8
(ceiling ladder filled to 120K / 91% of n_ctx, 0 MiB VRAM growth, needle recall
clean through 90K), 8-pack 105/150 with MTP off==on (spec-dec lossless),
soak-continuous PASS. Guard suite 42/0.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
New opt-in flag: after a catalog model is up + ready, switch.sh --owui upserts
an OpenAI connection in Open WebUI (host.docker.internal:<port>) so it appears
in the chat picker — no manual Admin->Connections step. scripts/lib/owui-register.sh
is conditional (no-op if OWUI not running), idempotent (skips if already present),
and drives OWUI's admin config API via a token forged from the container secret.
Validated end-to-end on Deckard-40B (:8199).
Guard suite 41/42; the 1 failure (test-measurement-record) is pre-existing
bench-output-format drift, unrelated to this change (no switch.sh/owui reference).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The MTP-injected GGUF is NOT a local build (the scaffold's comment was wrong) —
it's published at PiehSoft/Qwen3.6-40B-Deckard-MTP-Q6_K. Wire hf_repo + files so
setup.sh/launch.sh auto-fetch it, and correct the compose/profile comments.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Dense 40B uncensored Qwen3.6 community merge, dual-3090 llama.cpp with
embedded MTP head (Q6_K GGUF, 31 GB). First dual llama.cpp compose in
the catalog; first 'category' field in the registry (uncensored).
Changes:
- Compose: models/qwen3.6-40b-deckard/llama-cpp/compose/dual/mtp-q6k/mtp.yml
Status: 🧪 Unverified (quality /150 + soak pending)
Config: -ngl 99 -ts 1,1 -fa on --spec-type draft-mtp --spec-draft-n-max 2
--cache-type-k q8_0 --cache-type-v q8_0 -c 131072
- Registry: llamacpp/deckard40B-dual-mtp (category=uncensored, experimental)
- DEFAULTS: (qwen3.6-40b-deckard, llamacpp, dual) → slug
- Model profile: scripts/lib/profiles/models/qwen3.6-40b-deckard.yml
- Drafter: qwen-mtp-builtin model_compat extended with qwen3.6-40b-deckard
- LiteLLM route: deckard-40b → :8199
- launch.sh: suggest_default_variant case for Deckard
- BENCHMARKS.md: new section with validated MTP n=2 numbers
- Test fixtures: registry count 43→44, disk 44→49, models 5→6
Wrinkle decisions (see PR description):
1. Dual llama.cpp: uses 'count: all' + '-ts 1,1' (tensor-split layer-split),
matching the validated serving config. No launcher changes needed — the
compose's deploy section handles GPU reservation directly.
2. Engine pin: uses rolling server-cuda tag (matching engine profile spec),
not the b9246 pin other composes use. The validated build (2026-06-09
digest 1c4ff61a) is newer than b9246 and has draft-mtp working. No engine
pin bump for other models.
NOT DONE (live validation — GPUs busy with quality eval):
- Boot/bench/soak/quality NOT run — maintainer to validate post-eval
- Status stays 🧪 until full gate passes
* Add vllm/qwen-27b-dual-balanced + record 3-way 8-pack A/B (tie)
New "balanced" dual tier: cyankiwi AWQ-BF16-INT4 (int4 group-32 + BF16 mtp,
compressed-tensors) + int8-PTH KV + MTP n=3, TP=2 @262K. Live-validated
2026-06-07 (Marlin WNA16, KV pool 370K tok / 1.41x — the LARGEST of the dual
family, since int4 weights free more VRAM than fp8's 8-bit; ~67 TPS decode).
3-way 8-pack A/B (--full, same harness, same day 2026-06-07):
fast (autoround+fp8KV) 109 · balanced (awq+int8KV) 105 · max (fp8+int8KV) 110
→ a TIE (deterministic packs 64/64/65; spread within +/-5-7 noise).
The short-context 8-pack does NOT separate the quants. So:
- fast (vllm/dual) stays the default;
- balanced/max are framed by SERVING characteristics (KV-pool headroom,
int8-PTH fidelity), NOT behavioral quality — headers/docs say so explicitly.
- The dimension where int8-PTH should win (long-ctx NIAH recall) is an open
follow-up, not claimed here.
Also corrects the stale fast-tier 129/150 baseline: that was the 2026-05-09
harness; today's benchlocal-cli (verifier fixes since) scores the same fast
config at 109/150. Annotated as not-comparable across all four compose headers.
Wiring: awq-bf16-int4 weights_variant (compressed-tensors, cyankiwi repo) +
registry entry (vllm-stable, kv int8_per_token_head, port 8016, SKIP kvcalc,
experimental) + dual-max status_note updated with the A/B + DUAL_CARD rows +
A/B footnote. test-compose-registry-disk 42->43 / 43->44. Full gate green 43/0;
compat C4/C10/C14 pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* Correct balanced KV-pool claim: fast has the largest pool, not balanced
Measured v0.22.0 @262K, TP=2: fast 622K/2.37× > balanced 370K/1.41× > max
295K/1.13×. My initial framing claimed balanced had "the largest KV pool of
the dual family" — wrong: I only compared it to max and forgot fast. Fast's
17.5 GB autoround weights are far lighter than balanced's 27 GB AWQ, so fast
has the biggest pool by a wide margin.
This reframes balanced honestly: it is DOMINATED by the fast tier — slower
(~67 vs ~89 code TPS), smaller KV pool, and tied/below on the 8-pack. Its
only possible edge is int8-PTH KV fidelity > fast's fp8_e5m2 (same size — a
fidelity bet, not a memory one), which is UNPROVEN (the short-ctx 8-pack is
blind to it). Keep only if the long-ctx NIAH A/B (#470) proves it; else
deprecate. Fixed across the compose header, registry status_note, and
DUAL_CARD (rows + footnote); fast row now cites its 622K/2.37× pool.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
A symmetric 4-slug family for the Qwen vLLM path:
*-fast = AutoRound INT4 + fp8_e5m2 KV (peak TPS, the proven path)
*-max = official FP8 + int8-PTH KV (higher fidelity @ 262K)
Slugs:
vllm/qwen-27b-dual-fast alias of vllm/dual (AutoRound INT4, TP=2) — production
vllm/qwen-27b-dual-max FP8 + int8-PTH, TP=2 — 🧪 live-validated 2026-06-07
vllm/qwen-27b-multi-fast AutoRound INT4, TP=4 — 🧪 cross-rig (dual is the proxy)
vllm/qwen-27b-multi-max FP8 + int8-PTH, TP=4 — 🧪 cross-rig (dual-max is the proxy)
dual-max validated live on this 2-card rig: FP8 weights load via MarlinFP8
W8A16 on Ampere sm_86, int8_per_token_head KV keeps the full 262K (KV pool
295K tok / 1.13x concurrency), MTP n=3 active, coherent. The multi4 configs
are byte-identical to their dual sibling apart from TP + gpu-count, so they
ship Experimental until a real >=4x 3090 host validates them (this dev rig
has 2 cards). No multi4 DEFAULT added — the resolver degrades to vllm/dual
(honest; no model carries a multi default), and the slugs stay reachable by name.
Catalog wiring:
- models/qwen3.6-27b.yml: new fp8 weights_variant (official e4m3 release,
layer-split safetensors + embedded mtp.safetensors + vision tower)
- vllm-stable.yml: + fp8 weight format (Marlin W8A16 on Ampere) and
+ int8_per_token_head KV — NATIVE on stock v0.22.0 for uniform-head-dim
models (qwen3-next); features flag stays false (no overlay, unlike Gemma's
#40391-provided int8-PTH on vllm-gemma-stable)
- compose_registry.py: 4 entries; test-compose-registry-disk counts 38->42 / 40->43
Full gate green (42/42); compat C4/C5/C10/C14 pass for all 4 slugs.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
n-sweep @370W (dual TP=2): n2 72.9/86.1 . n3 74.0/87.7 . n4 71.6/87.8 wall-TPS
(all within ~3% - the QAT-int4's fast acceptance decay caps the spec benefit).
n=3 is the best balance (top narrative, code tied with n=4, ~20% less drafting:
43% draft-accept vs n=4's 37%), so SPEC_N_MAX default 4->3. The w4a16 compose had
inherited n=4 from the autoround clone, but autoround's slower decay (88/73/59/48%
per-pos) earns n=4 while the QAT-int4's faster decay (64/39/25% at n=3) favors a
lower n - optimal MTP n is quant-specific. Updates the compose default + Quality
field, registry status_note, BENCHMARKS row. Quality unchanged (n = speed only).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
vllm/gemma-31b-qat-w4a16-dual A/B vs the autoround-int4 default (gemma-int8-mtp),
same int8-PTH KV + assistant MTP n=4 (byte-identical speculative-config):
- 8-pack 109/150 vs 105 (+4, within +/-5-7 noise = quality tie; real instructfollow
edge IF 15-vs-8 offset by hermes/cli/TC/RM).
- WEAKER spec-decode: MTP accept-len ~2.4 vs autoround's ~3.9 (per-position
acceptance lower at every position) -> ~30% lower TPS (71.6/87.8 @370W <
autoround 106/139 @230W despite MORE power; AL is the clean power-indep signal).
- Boots clean on stock vllm-gemma-stable (no #44494 workaround - tower arch).
Comparable quality but slower -> autoround-int4 stays the default; stays Experimental.
Updates the compose Quality field + registry status_note + a BENCHMARKS row.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Repoint the single-card Gemma-4-26B-A4B slug in place from the bf16/16K
AWQ+MTP path to an INT8-per-token-head long-context path. Same weights
(cyankiwi AWQ-4bit, Marlin WNA16 MoE), same external MTP drafter (n=4),
same gemma4 tool-call; only the KV format changes (bf16 -> int8_per_token_head)
via the vendored vLLM PR #40391 hybrid-SWA KV page-size overlay. The dual
(vllm/gemma-26ba4b-dual) is untouched.
INT8-PTH (1 byte/token) lifts the single-card ceiling from bf16 16K to a
176K default. Live-validated on 1x RTX 3090 (2026-06-06): #40391 applies
cleanly, int8_per_token_head KV inits, Marlin WNA16 MoE backend, MTP
SpeculativeConfig active; KV pool 183,357 tok >= 176K; coherent generation
(post cudagraph-warmup) + clean gemma4 tool-call. mem_util 0.94, not 0.96:
0.96 passed the upfront KV check but OOM'd in the drafter's later cudagraph
capture (240 MiB free, needed 256) on a single 24 GB card.
- engine vllm-gemma-stable: + gemma4-swa-moe family, + awq/compressed-tensors
weight formats
- patches.yml: + gemma-a4b-vllm-pr40391-rebased (model-scoped copy of the
patch dir; reaches the compose, passes test-patch-attribution)
- compose single/awq/mtp.yml -> single/awq/int8.yml (full rewrite to the
#40391 overlay + entrypoint + int8 KV; Status: Experimental)
- registry + profile_runtime: engine/kv/max_ctx/mem_util/compose_path/entrypoint
- test-generate-from-profile: repoint the clean-derived-seed fixture off this
slug (now overlay-carrying) to vllm/gemma-26ba4b-dual + a tp= override
Status stays Experimental; rebench-full + soak are the #464 follow-up. Full
gate suite green (test-submit-bench pre-existing/environmental, fails on
baseline too).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The C12 budget verdict used a flat absolute floor (`MIN_KV_GB = 0.05 if
qwen3-next-moe else 1.0`) and FAILed any config leaving <1 GB for the growing
KV pool. That false-FAILs KV-light models that demonstrably boot: gemma-4-26b
-a4b single-card AWQ + MTP projects 0.78 GB growing pool, yet the live boot
holds 17,490 tok >= 16,384 max_ctx and serves (sliding_window=1024 makes 25/30
layers' KV trivially cheap).
vLLM's real pre-check is token-capacity based: it boots iff the capped pool
holds >= ONE max_model_len sequence, NOT iff KV >= 1 GB. Growing KV scales
linearly with max_num_seqs, so one sequence's KV is
`kv_pool_requested_gb / max_num_seqs`. Use that as the floor — which also
generalizes (and removes) the old qwen3-next-moe 0.05 special-case.
CAP the floor at the legacy 1 GB: for dense/long-KV configs one sequence needs
many GB and a hard per-seq threshold would false-FAIL measured-working configs
sitting inside the estimator's +-1.5 GB band (e.g. gemma-dual-int8 @262K:
10.71 GB avail vs 10.84 GB/seq, 1.2% short but boots). Net: relax the floor
ONLY for KV-light models, leave dense on its exact prior >=1 GB behavior.
The fix lives in the estimator (where the gap is), NOT in disabling
kv_calc_supported for a config we actively ship the drafter on. Design
bounced with Codex (per-sequence floor); capped-at-1-GB refinement added here
after the pure per-seq form regressed the gemma-4-31b dense duals.
Verdicts after: gemma-26ba4b-single FAIL -> TIGHT (correct); dense duals
unchanged; --calibration still 7/7; full gate 42/42.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Per request: turn on the external MTP drafter for BOTH topologies and
ladder each compose to its empirically-validated max context.
- single/awq/base.yml -> single/awq/mtp.yml (rename per ADDING_MODELS.md:
drafter=MTP named, bf16 KV is the gemma4-on-Ampere default so NOT named).
Adds --speculative-config (gemma-26b-it-assistant n=4) + prefix-caching +
chunked-prefill (parity with the dual). Max ctx 8K -> 16K: KV pool 17,490
tok at 0.92 util on 1x 3090 with the drafter loaded (17,408 would leave
<100 tok headroom). Registry drafter None -> gemma-26b-it-assistant.
- dual/awq/mtp.yml: max ctx 32K -> 262,144 (model max). KV pool 806,821 tok
at 262144/0.92 on 2x 3090 — sliding_window=1024 keeps 25/30 layers' KV
cheap, so the model max fits with 3x headroom.
Both LIVE-validated on the rig 2026-06-06 (stock v0.22.0): single tp=1 +
MTP boots (SpeculativeConfig method=mtp n=4), coherent, clean gemma4
tool-call; dual tp=2 + MTP @262K boots, coherent, clean tool-call.
Registry max_ctx 16384/262144 + single compose_path -> mtp.yml; profile_runtime
mirrors (single gains the speculative_config slot, dual maxlen 262144).
test-generate-from-profile mk_einput now clears the base slug's drafter (the
clean-derived fixtures need a no-drafter seed; the drafter-refuse fixture sets
one explicitly) so it's stable now that gemma-26ba4b-single carries an MTP drafter.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The gemma-4-26b-a4b slugs were the last users of the purged
`vllm-nightly-clean` engine. Migrate them to AWQ-4bit on stock
`vllm-stable` v0.22.0, retiring the Ampere-dead AutoRound path.
Why AutoRound is dead on Ampere: AutoRound int4-MIXED packs to
`uint8b128`, which has no W4A16 kernel on sm_86 (Cutlass/Machete are
Hopper-only; Triton W4A16 takes only uint4b8/uint4). The cyankiwi
AWQ-4bit (pure uint4, compressed-tensors) resolves to Marlin WNA16 MoE
on sm_86 and serves cleanly.
Slug changes (registry 38 -> 36):
- REMOVE vllm/gemma-a4b, vllm/gemma-a4b-awq (AutoRound/AWQ-bf16 duals)
- RENAME vllm/gemma-a4b-single -> vllm/gemma-26ba4b-single (AWQ tp=1, :8040)
- RENAME vllm/gemma-a4b-awq-mtp -> vllm/gemma-26ba4b-dual (AWQ tp=2 + ext MTP, :8041)
Both on vllm-stable, kv bf16 (kv_cache_dtype=auto), kvcalc_key=SKIP
(experimental; rebench-full/soak pending before Production).
PR #40886 (AWQ compressed-tensors MoE key remap) is native in v0.22.0 —
the vendored overlay + install entrypoint are dropped; patch marked
deprecated_on 2026-06-06. vllm-nightly-clean deprecated (final
purged-nightly engine, now zero registry users); gemma4-swa-moe added
to vllm-stable supported_model_families; arch_patches engine_pin
flipped to [email protected].0 loads:true.
Boot-validated on 2x 3090 (2026-06-06): single tp=1 coherent + clean
gemma4 tool-call; dual tp=2 + external MTP (num_spec_tokens=4) coherent
+ clean tool-call, draft layers wired. 3 AutoRound/AWQ-bf16 composes
moved to _archive/ (kept for a future Hopper engine).
Gate: 42/42 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
#327 audit found all 3 DFlash slugs rode the purged nightly-e47c98ef image (the
same dead SHA as the deprecated vllm-nightly-full), and vllm/dual4-dflash was
mislabeled `production` while (a) on a dead image, (b) TP=4 (unrunnable on this
2-card rig), (c) DFlash is blocked on Qwen3-Next per the repo hardware truths
(DeltaNet rollback). Same treatment as the #254 Genesis tail:
- Archived dual-dflash, dual-dflash-noviz, dual4-dflash -> compose/_archive/
(+ manifest); removed their 3 registry entries (41 -> 38).
- DEFAULTS[(qwen,vllm,multi4)] removed -- no functional multi4 vLLM slug remains,
so the resolver degrades multi4 -> dual. The 4-card vLLM tier is now empty
(recoverable from git if demand returns).
- vllm-nightly-dflash engine -> stability: deprecated; its qwen3-next-hybrid arch
pin flipped loads:true -> false (zero registry users; not offered for derive).
- Cascade: calibration (-3 rows), kv-calc alias map + override, patches.yml
(-18 dflash refs), test fixtures (registry-disk 38/40, calibration 9/9;
launch-compat/list-topology-filter/model-default-resolver/pullgate-gates/pull
adapted to the now-empty multi4 + nightly-slug categories).
Completes the nightly-engine cleanup: vllm-nightly-mtp/full/dflash all deprecated
with zero registry users; only the gated vllm-nightly-clean (#326) remains.
42/42 gate green. Refs #327.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>