A lazydocker-style terminal UI for the stack's test scripts (verify / bench /
verify-stress-NIAH / quality / soak / rebench-full): auto-detects the running
model + endpoint, injects MODEL=/URL=, streams live progress, and browses the
results/ history. Built by Qwen Max from docs/club3090-test-tui-prompt.md.
- scripts/c3t — launcher that derives the repo root from its own location and
runs the tool from tools/test-console's own venv (no global installs).
- tools/test-console/ — Python + Textual package, pinned pyproject + uv.lock,
and an 83-test offline pytest suite (parsers vs fixtures, mocked docker /
/v1/models detection, BENCH_MOCK bench parse) — all green, no GPU required.
Fixed before merge: __main__.py and a test fixture hardcoded the rig path
/opt/ai/github/club-3090 — now derived from __file__ (parents[3]; override via
C3T_REPO_ROOT) so it works from any clone, not just this rig. .gitignore
excludes .venv / caches.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Full rebench-full (--with-8pack-thinking=both) completed 2026-06-18 for
llamacpp/qwen27b-pi-reasoning-single. Replaces the earlier inferred "≈ base
50/59 @370W" note (+ the PENDING-ladder caveats) with the measured numbers, and
adds the BENCHMARKS.md single-card row.
- Bench @370W: 47.9/55.3 decode narr/code (n=5, CV<2%); @230W cap 28.5/32.9.
- verify-stress 8/8 (NIAH->183K), 8-pack 104/150 off / 106/150 on, soak PASS
(0 MiB growth, 0/100 silent-empty, p50 54.5, 102% retention).
- MTP head NOT weaker than base: matched-power A/B dead-even (73% vs 72% accept);
47.9/55.3 @370W ~on par with base 50.3/58.9. Power-sensitive (mainline -42%
370->230W) — the earlier "slow" 28/33 was the rig's 230W cap, not the model.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
benchlocal-cli already supports --max-tokens (overrides the per-pack ~1024
completion budget for BOTH arms, #28), but quality-test.sh only forwarded
--thinking-max-tokens. So a VERBOSE model that self-truncates the deterministic
packs (finish_reason=length before its final ANSWER:/</solution> line) could not
be benched at a higher budget through our wrapper.
Surfaced on llamacpp/qwen27b-pi-reasoning (a reasoning fine-tune that reasons in
visible content even thinking-off): ~5-10 reasonmath/bugfind/cli misses in the
thinking-OFF 8-pack were truncations, not wrong answers (e.g. RM-05 had 4/5
checkpoints right but ran out of tokens). The thinking-ON pass (16384 budget)
did not truncate — so the gap is purely the deterministic budget, which had no
knob in our wrapper until now.
- quality-test.sh: add --max-tokens N + MAX_TOKENS env passthrough (mirrors
--thinking-max-tokens exactly: same int validation, help, ENV doc, forward + echo).
- rebench-full.sh: plumb MAX_TOKENS into BOTH 8-pack passes (off + on) + doc it.
- test-quality-thinking.sh: assert env + flag forms forward --max-tokens and that
a non-integer is rejected (mock-benchlocal harness, hermetic).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
The merged llamacpp/qwen27b-pi-reasoning-single docs claimed its embedded MTP
head was "~45% below base / weaker" — that was a power-cap measurement artifact,
not a model property. A matched-power A/B (230W, same engine/KV/n=2/prompt) shows
the head performs IDENTICALLY to the base Qwen3.6-27B MTP head: 73% vs 72% draft
acceptance, 41.2 vs 41.2 t/s. The base llamacpp/default 50.3/58.9 figure I
compared against is a 370W BENCHMARKS number; mainline is -42% from 370->230W, so
the comparison was only valid at matched power (cf. the ik-llama/iq4ks-mtp
BENCHMARKS row, which documents exactly this trap). The 28.5/33.4 decode numbers
are correct AS 230W measurements; at 370W this config matches base's 50/59.
Corrected the status_note, compose header, and weights manual_note.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Three related [C0]/eligibility pull-gate fixes for the curated-swap surface —
an uncurated derive (abliterated / fine-tune) of a model we already serve.
1. Wrapper-arch alias. `pull.sh --profile-like` was false-aborting at [C0] with
"no arch_patches matrix row for 'Qwen3_5ForConditionalGeneration'". That is
the OUTER multimodal wrapper class the weights report; the patch matrix is
keyed on the inner canonical `Qwen3NextForCausalLM`. arch_patches.yml is a
closed key-set, so the alias lives in the editable arch_model_xref.
- profile_runtime.yml: `config_architectures: [Qwen3_5ForConditionalGeneration]`
on the Qwen3NextForCausalLM xref entry.
- generate_compose.py: `resolve_arch_from_config()` maps a config.json
architectures[0] string -> (canonical_arch, arch_row) via that alias.
- gates.py [C0]: resolve the wrapper arch via the alias before declaring
NO_ARCH_ROW. The hybrid now reports ENGINE_SUPPORTED.
2. GGUF axis. `supported_weight_formats` was declared on every engine but never
enforced (only `kv_format` was). The deriver blocks GGUF on the derive path,
but the curated registry / curated-swap path had no such guard. gates.py [C0]
now rejects a `gguf` weight_format on an engine whose supported_weight_formats
lacks `gguf` (structural axis; matches the `gguf` token only, so a derive's
raw dtype spelling bf16/float16 is never false-rejected).
3. Won't-fit size advisory. The eligibility no-fit-model abort for a hybrid/MoE
derive now appends (a) an actionable NOTE pointing at the curated-swap path +
docs/BRING_YOUR_OWN.md, and (b) a coarse weights-only VRAM verdict: when the
raw weights exceed the detected topology's total VRAM they won't fit at ANY
KV, so say so concretely (the huihui abliterated bf16 ~54 GB vs 2×24 GB case)
instead of a generic stop. `_weights_oversize_advisory()` is pure/total —
empty when it fits / size unknown / headless.
Docs + tests:
- BRING_YOUR_OWN.md: new section C — "Swap a curated model for a fine-tune /
abliterated variant -> reuse its compose" (artifact↔engine + quant + MTP
caveats, worked example).
- test-pullgate-gates.sh: wrapper-arch ALIAS [C0] case + resolve_arch_from_config()
unit + GGUF-on-vLLM runtime-incompatible + no-false-positive control +
_weights_oversize_advisory() unit (oversize / fits / headless / malformed).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
* LMCache: in-repo L2 default + LMCACHE_L2=1 toggle, preflight disk-check, corrected sizing
- LMCACHE_L2=1 enables the L2 disk tier at a gitignored repo-root lmcache-kv/ (zero
config); LMCACHE_L2_ADAPTER still fully overrides. L2 stays OFF by default.
- preflight_lmcache_ram now also soft-warns on low L2 disk space at the L2 host dir.
- CORRECTED capacity (measured > estimated): the LMCache offload cache is
~131 KB/token MEASURED (lmcache_mp_l1_memory_usage_bytes 4.93 GB + 4.6 GB on L2
disk, 36,808-token session) — ~7x the GPU's 18.9 int8-PTH rate, which only sets
max servable context. Prior docs used 18.9 for capacity and overstated it ~7x:
--l1-size-gb 30 holds ~4 x 50K sessions (not ~33), <1 full 262K.
- Added a RAM/disk-vs-context sizing table to INTERNALS.md.
- Measured L2 rehydrate: 4.8s vs 43s cold re-prefill (~9x), cross-restart persistence
confirmed; use the fs adapter (not nixl_store — this image's NIXL is broken).
Refs #133.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
* LMCache: document fp8 serde is broken on this model (no reduced-size L2 yet)
Measured 2026-06-17: LMCache's fp8 serde shape-errors on this model's KV
('[2,16,1584,260]' invalid — wrong layout for the hybrid GDN/head_dim-256/int8-PTH
source), 60 serialize fails, L2 stays empty. CacheGen-MP is 'coming soon'. So L2 is
uncompressed (~131 KB/token); the fp8-serde toggle was NOT wired (it doesn't work).
Re-test trigger noted. Refs #133.
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Measured the L2 disk tier on-rig (#133 follow-up): rehydrate 4.8s vs 43s cold
re-prefill (~9x) for a 37K session, cross-restart persistence confirmed
(0 L1 / 46 L2 retained keys post-restart). Two corrections to the shipped docs:
- L2 disk footprint ~125 KB/token measured (~33 GB per 262K session), not the
4.72 GB GPU-KV figure (L2 stores ~7x lower-density).
- Use the fs adapter, not nixl_store — this image's NIXL backend is broken.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Layers an LMCache tiered persistent prefix-KV cache (MP/HMA connector) onto the
dual-max fidelity profile (FP8 + int8-PTH KV + MTP n=3 @262K). Controlled A/B
shows ZERO decode penalty (74 narr / 94 code TPS == without LMCache, MTP intact
~83% accept); reuses cached prefix KV across long / multi-session workloads
(cold->warm TTFT ~7-8x) instead of re-prefilling.
- New vllm-lmcache engine profile: lmcache/vllm-openai DIGEST-pinned (the tag is
mutable — bundled vLLM 0.22.1-dev then 0.23.1-dev under the same tag).
- preflight_lmcache_ram guard: rejects --l1-size-gb over-allocation even under
--force (the l1=100-on-94GB host-OOM that forced a reboot, #133).
- env-gated L2 disk tier hook (LMCACHE_L2_ADAPTER, off by default).
- INTERNALS.md RAM/disk sizing section; registry entry + count bumps; full
guard suite green.
Incubating (hidden, --force to launch): runs LMCache's third-party image with a
newer vLLM than our v0.22.0 pin; L2 latency unmeasured on-rig; 38 GB on-demand.
Refs #133.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Code packs run 2026-06-16 (tool-free — sandbox executes the model's code, no
tool-calls): humaneval-plus-30 29/30 (97%), lcb-v6-30 25/30 (83%, 2 losses =
token_limit on the hardest). Substantiates the 'strong one-shot/competition
coding' characterization with our own numbers (matches the authors' ~96% LeetCode).
Folded into the llamacpp compose Quality field + registry status_note.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The performance-max VibeThinker path — prithivMLmods Q8_0 GGUF on mainline
llama.cpp, single 3090, q8_0 KV, full 131K, -b 4096 -ub 2048. A sibling to
the vLLM bf16 compose (#418), and the better one on every axis.
Why it's better than the vLLM sibling:
- ~166 TPS decode (vs ~110 vLLM bf16), ~6.2 GB at full 131K (vs ~9.8 GB),
prefill ~6,630 tok/s (-b 4096/-ub 2048 = +20%), boot ~9s.
- Q8_0 is near-lossless → quality INTACT, where vLLM's fp8 weight-quant broke
this quant-sensitive 3B (non-terminating empty output). KV quant is fine
(storage-only); fp8 *weights* are the problem.
Validated 2026-06-16 (tool-free packs, temp 0.6, thinking-on):
gsm-symbolic-30 100% · instructfollow-15 100% · reasonmath-15 80% ·
structoutput-15 80% · dataextract-15 40%.
dataextract is a genuine extraction ceiling — full temp sweep {0:27%,
0.6:40%, 1.0:27%} can't lift it (value mismatches persist). A reasoning
specialist, not an extractor. No tool-calling. verify-full 5/9 by design.
Wiring: registry entry + dense-family added to llama-cpp-local engine +
prithivmlmods-q8 weights variant + disk/registry counts (51/52). No MTP head
in the GGUF (plain qwen2 conversion) → no spec-dec. Guard suite 47/47.
Status: 🐣 incubating (hidden from switch.sh --list; --force to launch).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Introduce a new pre-experimental status tier, `incubating`, for composes that
work but aren't ready for the actionable list — niche specialists or models
that fail the standard functional gate by design. Incubating composes are
HIDDEN from `switch.sh --list` by default (revealed by `--list --all`) and
launch-gated behind `--force` (non-functional), so a half-validated model is
catalogued and discoverable without cluttering the recommended set.
Tier wiring (reuses the existing status plumbing — no new emit column):
- compose_registry.py: `incubating` in STATUS_VALUES + `🐣` in COMPOSE_STATUS_EMOJI
(stays out of FUNCTIONAL_STATUSES → --force-gated, never auto-defaulted)
- switch.sh: both --list filters skip incubating unless --all, with a
`(+N incubating hidden — --all)` header note
- AGENTS.md: Status enum row + Caveats-required references
- docs/ADDING_MODELS.md: new rule — NEW MODELS START at 🐣 Incubating, promote
up the enum (🐣 → 🧪 → ⚠️/✅) as they earn the actionable list
First occupant — VibeThinker-3B (WeiboAI, Qwen2 dense reasoning fine-tune):
- `vllm/vibethinker-3b-single` — bf16 weights + fp8_e5m2 KV, single 3090,
full 131072 ctx, mem_util 0.40 (~9.8 GB, single-concurrency sized),
--reasoning-parser qwen3, no tool-calling
- Live-validated 2026-06-16: serves clean correct reasoning/code (~110 TPS),
qwen3 parser splits <think>. fp8 WEIGHTS rejected (break the quant-sensitive
3B: non-terminating empty output). Always-reasoning + no-tools → fails
verify-full's fixed-small-budget checks (5/9) by design → incubating, not gate-passing.
Guard suite 47/47 green (nex-n2-mini parked aside; it's an unrelated local experiment).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Distills the #350 Deckard-40B post into a reusable 'we shipped X' skeleton
that wraps a Results Card (intro+credits / Results Card / getting it / run it
/ what'd help / credits). References RESULTS_CARD.md for the measurement panel
rather than duplicating it; wired into the docs/README.md reference index.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
First cross-rig confirmation of the shipped vllm/diffusiongemma-dual compose
(@steamEngineer, 2x 3090 NVLink). Decode 294/370 — notably above the PCIe
reference (~177/180), plausibly the NVLink TP=2 all-reduce, though block-
diffusion variance is high. NIAH usable to 184K, verify-stress 8/8; the lone
verify-full fail is the SSE-chunk check (false-fail for a block-diffusion model).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
soak-test.sh's auto_container() had drifted into a hardcoded model-name
allowlist (qwen36-27b / qwen36-35b-a3b / gemma-4-31b + beellama gemma
variants). Any shipped compose outside that list silently failed auto-detect
with "no running club-3090 container found" — #405 hit it on the shipped
diffusiongemma-26b-a4b dual compose (soak "failed to launch" through no
fault of the contributor's rig).
Re-syncs with the canonical preflight.sh::preflight_autodetect_endpoint
(which this function's own comment claims to mirror): detect by ENGINE-INTERNAL
port mapping (vLLM 8000 / llama.cpp 8080 / sglang 30000), model-agnostic, then
prefer a recognised engine-family prefix. Same bug class as the #310 preflight
fix — now any compose is found regardless of model.
Verified: unit-tested the pipeline against the #405 container
(vllm-diffusiongemma-26b-a4b-fp8-tp2 + its 8020->8000 mapping) plus qwen/
llama-cpp/sglang variants — all detected; a port-80 studio container is
correctly excluded. bash -n clean; full guard suite green.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
First use of the new `revision:` lever (#408). Pins the
morikomorizz/Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP fetch to commit
49a080db7406cd98f94f0f2a18539bbcc1520444 — the exact bytes the BENCHMARKS
row and #411 Results Card were measured against.
Verified: repo last-modified 2026-06-10 predates our 2026-06-14 download,
and the HF x-linked-etag at this sha equals our local file's sha256
(76a0d4c2...). weights.py round-trips revision: -> WEIGHT_REVISION; guard
suite green. Guards against a silent upstream re-quant (the #316 class).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
adds an optional `revision:` key per weights variant. weights.py emits
WEIGHT_REVISION; setup.sh threads it into `hf download --revision` and
pins the post-download sha-verify etag lookup to the same revision, so a
stale pin can't false-fail against a newer HEAD. preflight's manual hint
mirrors the flag. unset = track HEAD, so behavior is unchanged for every
current entry (nothing sets revision: today).
this is the weights half of #316: upstream quant repos re-quant
silently, and we had no lever to pin the bytes a BENCHMARKS row was
measured against. engine images already pin; weights didn't. mechanism
only, no real entry is pinned in this PR (that's a per-entry,
rig-validated call that's yours to make).
refs #319, #316
adds scripts/tests/test-hf-repos-resolve.sh: an opt-in, network-gated
check that every hf_repo / hf_repos entry in the weights registry
actually resolves on HF, not just string-matches the profile. catches
the renamed / never-existed repo class from #316 that the string-only
guard in test-model-weights-registry.sh can't see.
gated behind CLUB3090_CHECK_HF_REPOS=1 and self-skips otherwise, so the
default `for t in scripts/tests/*.sh` sweep stays offline-safe and green.
follows redirects like `hf download` does, so a rename or recase resolves
and is flagged (not failed); only a real 404 fails. gated 401/403 counts
as exists; network / rate-limit codes are reported but not counted, to
avoid spurious reds.
refs #320, #316
#145 / vLLM #39056 are STREAMING-only tool-call bugs, and our gates were
structurally blind to them: benchlocal-cli runs every pack non-streaming
(benchlocal-cli#68, closed wrong-layer), and verify-full had check_tools
(non-streaming) and check_streaming (no tools) as SEPARATE checks — the bug
lives in their untested intersection (tools × streaming × thinking-on). That's
how the MTP-streaming drop (#39598, un-mitigated since Genesis retired) shipped
silently into the dual composes.
- scripts/verify-full.sh: new [6/9] check_streaming_tools — stream:true + tools
+ tool_choice=auto + enable_thinking=true, reassembles the SSE deltas, and
asserts delta.tool_calls (get_weather) + finish_reason=tool_calls + no
<tool_call> leak into delta.content. Uses tool_choice=auto (the clean path on
v0.22.0); hard-fails on the leak signature (the #145/#39056 class). NB:
tool_choice=required + MTP is a known-open drop (#39598) — the check notes it
but uses auto so it's a green gate on shipped composes. Gated by SKIP_TOOLS.
Renumbered the suite 8 -> 9 checks. Live-validated: vllm/dual all 9 pass.
- scripts/stream-toolcall-probe.py: the deep-sweep instrument behind the #400
investigation (thinking-on, multi-prompt, auto+required, SSE reassembly,
PASS/DROP/OTHER classification, exit-1 on any DROP). It's what proved
qwen3_coder == qwen3_xml and isolated the #39598 MTP-gating. Run it twice
(control vs candidate) for a streaming A/B.
- scripts/tests/test-stream-toolcall-probe.sh: offline mock-SSE gate (clean +
#145-drop modes) — no GPU needed.
Full suite green (test-compose-registry-disk fails pre-existing on an unrelated
untracked model dir).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
deckard-40b loops at default sampling (reported by milano @ 262K ctx). The
compose shipped --repeat-penalty 1.0 (no penalty) and llama.cpp's default
repeat-last-n 64 → degenerate loops (a 10-word phrase repeating up to ~46x in a
2500-token gen, intermittent at temp 0.6). milano's own rep 1.05 + presence 0.05
only takes that to ~18 ("helps then loops again").
Change (validated quality-neutral by a same-session symmetric 8-pack A/B):
- --repeat-penalty 1.0 -> 1.1 (env REPEAT_PENALTY)
- add --repeat-last-n 256 (env REPEAT_LAST_N; was llama.cpp default 64)
→ together these cut looping ~5x (max 46 -> 9 in a per-request sweep) and are
gentle on code (unlike DRY).
- wire DRY sampler env knobs, DEFAULT OFF (--dry-multiplier 0.0 / --dry-base
1.75 / --dry-allowed-length 2): DRY is the strongest loop-breaker but
over-suppresses legitimate repetition in CODE, so it's opt-in (DRY_MULTIPLIER=0.8)
for severe long-ctx loops. Header env-docs updated.
Symmetric 8-pack A/B (same harness, 2026-06-13, think-off, MTP n=2):
OLD rep1.0 99/150 vs NEW rep1.1 100/150 (Δ +1 = noise; per-pack deltas
bidirectional). NB: OLD measured 99 today vs the #350 historical 105 — that
6-pt gap is harness drift since 2026-06-10, NOT this change (the change is +1).
So: anti-loop win at zero quality cost.
Reasoning is OFF by default on this compose, so --reasoning-budget (a common
suggestion) is a no-op here unless the user enables reasoning.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
The streaming sweep for the qwen3_xml PR (#400) surfaced the real picture:
- #39056 auto+thinking+streaming is fixed natively on v0.22.0; the qwen3_coder→
qwen3_xml swap is a no-op (byte-identical A/B, both parsers) → PR #400 closed.
- The residual #145 failure is tool_choice=required + thinking + streaming, and
it's MTP-gated: confirmed 2026-06-13 (scripts/stream-toolcall-probe.py) MTP n=3
dropped 13/20, no-MTP clean 0/20, parser-independent. That's #39598 (MTP
streaming early-return) resurfacing — its Genesis P64 mitigation was retired
with Genesis (#182/#254), so the drop shipped silently back into the dual
composes (benchlocal is non-streaming; verify-full's streaming check has no
tools, so neither caught it).
Updated both rows; mitigation = tool_choice=auto (clean w/ MTP) or no-MTP;
durable fix is upstream (#45413 is parser-side, likely doesn't cover the
spec-decode-streaming path).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
The rolling ghcr.io/ikawrakow/ik-llama-cpp:cu13-server tag moved (2026-05-23
-> 2026-06-10, digest 5f914f1c) and the new build REJECTS the legacy
speculative-decode flags: "legacy speculative option '--multi-token-prediction'
is disabled; use --spec-type mtp:n_max=1,p_min=0.0". Every ik-llama MTP compose
crash-restart-loops on a fresh pull (reported by furfix, Discord — single 3090
WSL). The binary's own --help confirms `-mtp`/`--draft-*`/`--spec-stage` are all
rejected in favour of the unified `--spec-type SPEC[:k=v,...]`.
Flag migration (13 composes):
- 11 MTP composes: --multi-token-prediction + --draft-max N + --draft-p-min P
-> --spec-type mtp:n_max=N,p_min=P (per-compose N preserved).
- 2 two-stage composes: the ngram-mod + mtp cascade used the rejected
--spec-stage with legacy keys -> --spec-type ngram-mod:...,ngram_size_n=... +
--spec-type mtp:...,p_min=... (spec-ngram-size-n -> ngram_size_n,
draft-p-min -> p_min, the canonical keys).
Digest-pin (decided to stop the recurrence — this is the 2nd flag churn from
the moving tag; ik-llama has no engine profile, so the per-compose image:
default is the only pin point): :cu13-server -> @sha256:5f914f1c... in all 13
ik composes + the preflight test fixture + 3 docs (IK_LLAMA / SINGLE_CARD /
INFERENCE_ENGINES). VLLM-style central injection isn't available for ik-llama;
a follow-up could add an engine profile so future pins are one-line.
Live-validated on the pinned image (single 3090):
- ik-llama/iq4ks-mtp: healthy, 0 restarts, no legacy error, MTP context ready,
served a completion, draft acceptance 0.62 (26/42).
- ik-llama/iq4ks-two-stage: healthy, 0 restarts, both stages load
(ngram_mod n=16 + MTP context ready), speculative decoding initialized.
Full test suite green (test-compose-registry-disk pre-existing on an unrelated
untracked model dir). The 3 untracked nex-n2-mini ik composes also carry the
rolling tag but are out of scope here (in-progress catalog work).
Co-authored-by: noonghunna <[email protected]>
setup.sh hard-coded GENESIS_PIN=7b9fd319 and the qwen3.6-27b case set
NEEDS_GENESIS=1, so the default `bash setup.sh qwen3.6-27b` cloned + checked
out Sandermage's Genesis tree at a stale pin on the most common setup path —
the original #182 harm being that the stale pin silently reverts the PN59
Cliff-2b fix.
That harm is already neutralized in practice (the live qwen3.6-27b autoround
composes run overlay-free on stock vllm/vllm-openai:v0.22.0 with `Genesis: none`
and no _genesis mount), but the clone still fired — wasted work, and a latent
landmine. It also contradicted the model profile, which already declares
qwen3.6-27b.requires_genesis: false.
Genesis's last unique justification on this stack was turboquant_3bit_nc KV;
that's now covered by beellama KVarN, and we're not reviving Genesis near-term.
- NEEDS_GENESIS now defaults to 0 and is env-overridable (`${NEEDS_GENESIS:-0}`).
- The qwen3.6-27b dispatch no longer flips it to 1 (matches requires_genesis:false).
- Opt-in preserved for reviving an archived TQ3 compose:
`NEEDS_GENESIS=1 bash scripts/setup.sh qwen3.6-27b` still clones the tree.
- Replaced the stale v7.69-dev narrative comment with the current opt-in state.
No live config needs Genesis (registry has zero non-deprecated slugs on a
genesis_equipped engine), so nothing changes for any shipped path. preflight's
genesis-pin drift check self-skips when the tree isn't cloned. Full test suite
green (test-compose-registry-disk fails pre-existing on an unrelated untracked
model dir; this change touches only setup.sh).
Closes#182.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Phase 1 (offline): the wrapper, the corpus home, and the test gate. Phase 2
(GPU) curates the actual baselines against the live RECOMMENDED_DEFAULT_MODELS
configs.
- scripts/quality-baseline.sh: thin wrapper over quality-test.sh --full that
captures (--capture -> --save-json) or diffs (default -> --previous-result)
an n>=3 aggregate per (registry-slug, thinking-mode). no-thinking is
canonical (temp-0); enable-thinking is the reasoning-on companion. --dry-run
prints the resolved command; extra args pass through to benchlocal-cli.
- scripts/quality-test.sh: forward --repeat / --previous-result and honor a
--save-json path override, so the wrapper's blessed layout works.
- results/baselines/: committed corpus home (whitelisted in .gitignore) + a
README documenting the convention, usage, and an empty index table.
- scripts/tests/test-quality-baseline.sh: offline gate (--dry-run) — asserts
command/path resolution per mode + the required-slug / valid-mode /
positive-repeat / missing-baseline guards.
- docs/QUALITY_TEST.md: regression-baseline subsection pointing at the corpus.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
bench-agentic.sh drove the ramp with tool_choice='required' and RAISED if a
turn returned no parseable tool call → the main loop caught it and broke,
aborting at the first miss. Reachable depth was bounded by tool-call
reliability at depth, not the TURNS/context budget — so on flaky parsers the
ramp stopped below the ~35K DeltaNet degrade zone the producer exists to
characterize (field-observed abort at turn 11 / ~22K).
Fix: on a no-parseable-tool-call turn, synthesize a tool call so the prompt
keeps growing by the same fixed ~chars (the fixture tool_result is injected
regardless and dominates the growth), flag tool_call_missed, and continue.
Genuine transport errors (HTTP/timeout) still propagate and stop the ramp.
Per-turn rows mark misses; a summary line reports "tool-call misses: N/M".
The optional non-tool RAMP_MODE the issue floats is largely subsumed — the
synthesize-on-miss path already lets the ramp reach configured depth on any
engine regardless of tool-call reliability.
Test: scripts/tests/test-bench-agentic-ramp.sh — mock SSE endpoint; asserts the
ramp reaches turn 3 under 100% tool-call misses (counted) AND the success path
is unchanged (no false misses). Offline, no GPU.
Caveat: guarantees reaching configured TURNS; whether the 15-turn fixture
exceeds 35K is separate (extend the fixture if not). Sibling bench-agentic fix
#498 stays separate.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
rebench-full ran the 8-pack quality eval (both think-OFF and think-ON,
~1.5-2.5 hr) by DEFAULT — the longest phase, dominating wall time even when
you only need "does this boot / serve / recall / soak". Flip it: the 8-pack
is now opt-in.
(omit) fast structural gates only (verify + bench + stress + soak)
--with-8pack-thinking 8-pack, reasoning OFF (--full --no-thinking)
--with-8pack-thinking=off same
--with-8pack-thinking=on 8-pack, reasoning ON (--full --enable-thinking)
--with-8pack-thinking=both both passes (the production-promotion gate)
The =off pass now FORCES --no-thinking (all 8 packs think-OFF) for a clean
with/without-reasoning A/B — previously bare `--full` used pack-default MIXED
thinking, which wasn't a true "reasoning off". quality-test.sh runs exactly one
mode per call, so "both" == two invocations; the think-ON pass fires ONLY for
=on/=both, never accidentally.
Naming realigned to benchlocal-cli #65's pack-set-vs-thinking-mode vocabulary
(the issue's --with-quality=off,on -> --with-8pack-thinking[=off|on|both];
"thinking" is explicit in the name so off/on can't be misread as "skip the
pack"). ENABLE_THINKING env no longer drives the 8-pack (only bench.sh).
Promotion-gate call sites updated to --with-8pack-thinking=both: ADDING_MODELS.md
Step 8, QUALITY_TEST.md, TQ3_MTP_GENESIS.md.
Validated: bash -n clean; bad-value guard exits 2 early; --help updated;
test-quality-thinking green (quality-test.sh untouched).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
deucebucket's Qwen3.6-35B-A3B-Cerebellum-v3 (mainline llama.cpp single-3090,
ablation-informed mixed-precision GGUF, 11.96 GB, ~15.1 GB peak @131K, 147.9/146.2
TPS, det 63/75). Landed as a credited author-reported data point rather than a
catalog compose (PR #393 closed): its headline value is the 16 GB fit, which we
have no 16 GB card to validate or support; on 24 GB it's context/quality-dominated
by the ik-llama apex-fit/byteshape siblings. README llama.cpp ❌→✅ correction
(559a3fe) was the durable fix from the same report.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
vLLM merged #45295 (mgoin, 2026-06-12) — general marlin_padded_nk tile-pad
across all dense Marlin paths (WNA16/AWQ/GPTQ + FP4/FP8), subsuming our
per-case W4A16 sub-tile-n pad. Closed our PR #40361 in its favor. #45295 is
NOT in v0.22.0 or the just-released v0.23.0 (branched before the merge), so
the vendored marlin-pad overlay drop + stable pin bump are queued together
for the first release that includes it (v0.24.0 / v0.23.x backport).
- #40361 row → 🟣 closed-superseded; #40354 → 🟢 resolved-upstream-by-#45295
- dgemma #45163 drop-trigger: K-pad half landed via #45295 (pending :gemma rebuild)
- new follow-up row: v0.22.0→v0.23.0+ stable pin bump (wait+combine w/ #45295)
Co-Authored-By: Claude Opus 4.8 <[email protected]>
Was '**v0.8.2 surfaces (current):**' — stale at v0.8.7. The four pull-gate
bullets are still the current surface, so reword to '**Pull-gate surfaces
(current):**' (no pinned version → can't go stale on the next release).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
35B-A3B row (#390 follow-up): drop NEW v0.7.3 tag (we're at v0.8.7; matches
the other established production rows that carry no tag), add the byteshape-iq4xs
single-card path (113/129 TPS @262K, 110/150 8-pack, PR #293) alongside apex-fit.
Also drop the same stale NEW v0.7.3 tag from the Gemma 4 26B-A4B row.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
vLLM dual was promoted preview→✅ Production on 2026-05-30 (vllm/qwen-35b-a3b-dual,
v0.22.0 stable, 262K + vision, 178/174 TPS) per BENCHMARKS.md. README still showed
'Preview (vLLM dual)' / 'vLLM ✅ (preview)' / '182/177 at 16K (no MTP/TQ3/Genesis)'.
Corrected Status + Engines + Highlights to the production numbers.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
deucebucket (#390) ran qwen3.6-35b-a3b on mainline llama.cpp single-3090
(141/144 TPS, 15.1GB @131K). The README support table marked llama.cpp ❌
for this MoE, conflating 'no shipped catalog compose' with 'unsupported' —
contradicted by our own docs/HARDWARE.md mainline power-cap curves on this
exact model and the llama-cpp-mainline engine profile (qwen3-next-moe).
Note added that ik_llama is the shipped single-card path.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
* beellama: bump pin v0.3.0 → v0.3.2-preview (commit-pinned) + validate
Maintainer chose the v0.3.2 preview over the v0.3.1 stable for the newer
build (adds experimental KVarN KV-compression). v0.3.2 is a rolling
pre-release — Anbeeld replaces its moving Docker tags with newer branch
builds — so we pin the COMMIT-suffixed tag for an immutable pin:
install.spec → ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.3.2-317c65e27e1e
Validated on-rig (single 3090, q5ks-dflash): boots on the preview image,
verify-full all-pass (Paris / tool_calls / streaming / thinking), prose
coherent, DFlash spec-dec active (acceptance ~0.12-0.17 on short tasks,
not collapsed). beellama has no vendored patches, so nothing to rebase.
Composes STAY 🧪 experimental: preview ≠ stable. The first stable tag now
exists (v0.3.1, server-cuda-v0.3.1, non-prerelease — Qwen3 MTP post-norm +
CUDA KV-quant fixes); repoint install.spec there to un-park (#455) once it
passes the full gate. Multiarch fallback (compose-literal, sm_120 direct-
compose) left at v0.3.0 — rebuild at the chosen tag is a separate follow-up.
scripts/tests/*.sh green (the 2 reds are pre-existing untracked-experimental-
compose artifacts: qwopus-coder + nex-n2-mini, unrelated to this pin).
Co-Authored-By: Claude Opus 4.8 <[email protected]>
* beellama: repoint pin to KVarN build BY DIGEST (the -317c65 tag lacks KVarN)
The commit-suffixed preview tag server-cuda-preview-v0.3.2-317c65e27e1e
PREDATES the KVarN merge — its --cache-type-k rejects kvarn* (only
turbo/TCQ). KVarN is only in the latest rolling server-cuda-preview-v0.3.2
build (commit 98caf25), which has no immutable commit-suffixed tag, so we
pin its DIGEST (immutable + KVarN), same as the vLLM :gemma digest pin:
install.spec → ghcr.io/anbeeld/beellama.cpp@sha256:858e7cfbfb0d5d…
Measured (single 3090, Qwen3.6-27B Q5_K_S, -np 1): KVarN lifts the single-
request ceiling from ~196K (q5_0/q4_1; 262K OOMs) to the full 262K —
kvarn4 (≈q5_0 quality) fits 262K tight (~1GB free), kvarn2 ~3GB free.
Recall-at-depth NIAH + Qwopus-coder re-validate on the digest build pending.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Follow-up cleanup to the merged #369:
- error message said "WEIGHTS=ft8" → "WEIGHTS=fp8"
- the unrecognized-WEIGHTS help string now lists 'fp8'
- the fp8 entry's manual_note said "no direct pull recipe wired" — now
that hf_repo is set, document the WEIGHTS=fp8 setup.sh path (hf-download.sh
kept as the stall-resistant manual option for the 29 GB layer-split repo)
Co-Authored-By: Claude Opus 4.8 <[email protected]>
report.sh autodetects the running container + URL + engine but not the
served model name, so verify.sh / verify-full.sh / verify-stress.sh /
bench.sh fell back to a hardcoded MODEL=qwen3.6-27b-autoround. Against a
non-qwen vLLM endpoint (e.g. vllm/gemma-26ba4b-single serving
gemma-4-26b-a4b-awq) every request 404'd with "The model
`qwen3.6-27b-autoround` does not exist", failing every check (#372).
llama.cpp ignores the request's model field, so the same wrong default
silently "worked" there (#371) — which masked the bug.
Add a shared preflight_autodetect_model helper that resolves the served
name from the endpoint's /v1/models (first id) when MODEL is unset, and
call it in the four affected scripts before their qwen fallback. This
mirrors what soak-test.sh / bench-agentic.sh / quality-test.sh already
do. An explicit MODEL= still wins (important for llama-swap/multi-model
endpoints); the qwen literal stays as a last resort if detection no-ops.
Validated: unit (resolve / respect-explicit / unreachable-fallback),
end-to-end against a mock /v1/models, and the full scripts/tests suite
stays green (the one pre-existing test-compose-registry-disk failure is
unrelated — untracked nex-n2-mini WIP composes).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
OWUI runs title/tags/follow-up/autocomplete tasks against the selected
model; when that's a Studio generation pipe it rendered the task prompts
as images. Pin TASK_MODEL/TASK_MODEL_EXTERNAL to qwen3.5-4b-uncensored
(:8090, always-on) so tasks run on a chat model — follow-ups stay ON, no
junk renders. (Set live in the OWUI DB 2026-06-12; this seeds fresh volumes.)
Co-Authored-By: Claude Opus 4.8 <[email protected]>
OWUI runs its title/tags/follow-up/autocomplete TASK prompts against the
currently-selected model = the Studio generation pipe, so each chat turn
the pipe was rendering '### Task: Suggest follow-up questions...' as an
(invariably blocked) image — the mysterious extra image. Guard: if the
last user message is an OWUI task prompt (### Task:/### Chat History:/
autocompletion/etc.), return '' without generating. v0.13.2 -> 0.13.3.
(Also recommend setting OWUI's Task Model to a chat model.)
Co-Authored-By: Claude Opus 4.8 <[email protected]>
OWUI's renderer showed the hidden <!--SPEC:base64--> HTML comment (the
refine-context marker) as literal text at the end of every reply. Removed
it from all 5 lanes; _prior_spec now recovers the refine context from the
VISIBLE '**Prompt used:**' / '**Style:**' line already in each reply — so
follow-ups ('make it night') still evolve the previous prompt, with no
leaked marker. v0.13.1 -> 0.13.2.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
OWUI stores the pipe code in its DB (not from studio_pipe.py on disk), so
editing build_studio_pipe.py + regenerating wasn't enough — the installed
function went stale (e.g. v0.11.0 lingering after several lane additions).
This helper rebuilds the pipe, writes it into the OWUI 'studio' function
row, and restarts OWUI to reload it (--no-reload to skip the restart).
README documents the update flow + the stale-function trap.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
"hi" into a Studio lane used to be crafted into a random prompt and
rendered (a 3-4 min HiDream image, or a video-lane OOM) — because the pipe
crafted + rendered on ANY non-empty message. Now the director itself gates:
its system prompt instructs it to reply with `CHAT: <friendly one-liner>`
when the message isn't a real generation request (greeting / small-talk /
question / too vague), and the pipe returns that reply WITHOUT calling the
renderer. A real request ("a red fox in the snow") still crafts + renders.
- _enhance: appends a per-lane CHAT-gating rule to the director system prompt.
- _chat_gate(crafted): returns the friendly reply if the director emitted
"CHAT: ...", else None.
- All crafting lanes (image/chroma/hidream · music · sfx · video) check the
gate right after the craft step, before any render.
Validated on the live director (qwen3.5-4b-uncensored, enable_thinking:False):
"hi" -> "CHAT: Hello! Please describe an image you'd like me to generate…"
"thanks!" -> CHAT reply ; "a red fox in the snow" -> full crafted prompt.
v0.13.0 -> 0.13.1.
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>