6bc91b5cc50ee3a0c2c8a2e8b09029bb4d1b0c02
105
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f2b9dc2de4 |
Retract beellama v0.3.0 prose-DFlash "regression" — net-positive on tok/s
A careful tok/s re-test (bench.sh narrative n=5, MEASURED no-spec controls, three v0.3.0 images: efe856397 / e0663be / 63abcd3) shows DFlash prose is net-POSITIVE everywhere — Qwen single Q5 +27% (45.7 vs 35.9 no-spec), Qwen dual Q8 +52% (35.6 vs 23.4), Gemma single Q4 +28-31% (44.6 vs 34.8). The earlier "v0.3.0-wide prose-acceptance regression (~0.07 AR / net-negative)" was a DOUBLE error: (1) over-reading the noisy/prompt-dependent acceptance-rate diagnostic (the same efe856397 image we logged at ~0.07 now reads ~0.32 AR at the same tok/s — Anbeeld's #288 AR caution was right), and (2) a wrong no-spec baseline (we'd used ~37; the real dual-Q8 no-spec is 23.4). The new adaptive-DM HEAD 63abcd3 is neutral (tok/s flat). Build-arch ruled out (a 3090 runs identical sm_86 SASS from a fat or single-arch binary). - compose_registry.py: 5 beellama status_notes (measured slugs assert net-positive; un-rebenched gemma-duals retract the claim without overclaiming). - 5 beellama compose Caveats: same correction. - docs/UPSTREAM.md row 38: regression clause retracted; prose-recovery half of the promotion gate dropped. - BENCHMARKS.md: added the v0.3.0 DFlash-vs-no-spec A/B note (no-spec baselines). - (learnings/qwen3.6-27b.md + gemma-4-31b.md got dated append-only retractions; reported to Anbeeld at discussion #288.) Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
e7bf8f5ce1 |
beellama: refresh stale b9459 image refs to v0.3.0
The 3 beellama composes still defaulting to the old self-hosted multiarch-b9459-07ac3ce (the 2 singles + the gemma-q4ks dual) and the two single-card registry status_notes predate the #296 BEELLAMA_IMAGE centralization. b9459 is the pre-v0.3.0 build — the gemma-q4ks dual even documented that its default "STILL has the [multi-GPU DFlash] bug". A direct `docker compose up` (no launcher) would pull that stale image. - image: defaults → ghcr.io/noonghunna/beellama-cpp:multiarch-v0.3.0-efe856397 (matches the q8kxl duals; broad sm_86/89/120 compat for direct-compose). - Caveats / Pull-the-image / status_note prose → launchers inject Anbeeld's official server-cuda-v0.3.0 (sm_86/89); 5090/sm_120 uses the multiarch build. - Drop the obsolete "no official Docker yet (v0.3.0 WIP)" claim; note the v0.3.0 prose-DFlash acceptance regression (code unaffected) instead. - Preserve the two historical b9459 references (the 2026-05-30 benchmark attribution + the old-build prose-accept comparison). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
08d4dda6f3 |
feat(35b-a3b): byteshape IQ4_XS ik-llama single-card preset — validated (#299)
Intake of @Rhonstin's #293 byteshape IQ4_XS 35B-A3B preset into the curated catalog + first-party validation on 1x 3090: verify-full/stress 8/8 (NIAH→240K), bench n=5 (narrative 113/code 129 wall TPS), 8-pack 110/150 (≈ author's 111/150), soak-continuous PASS. Registry ik-llama/byteshape-iq4xs-mtp, ⚠️ Production w/ caveats. Intake fixes vs #293: image cu13, port 8058. Credit @Rhonstin. Gate 41/41. |
||
|
|
27b9fe5e45 |
feat(beellama): release v0.3.0 Q8_K_XL dual composes (experimental) + centralize image pin (#296)
Ships 3 dual-card beellama v0.3.0 Q8_K_XL composes 🧪 experimental for community testing (#288): qwen-mtp-dual :8064, qwen-dflash-dual :8065 (262K), gemma-q8-dflash-dual :8066 (192K); flips the q4ks gemma dual ⏸️→🧪. Centralizes BEELLAMA_IMAGE injection (engine install.spec, one-line bump) at Anbeeld's official server-cuda-v0.3.0 tag + adds scripts/beellama-pin-bump.sh (manual, no auto-push). Catalog: q8kxl weights variants, 3 registry entries, beellama-local +mtp_gguf, unsloth-mtp-gguf engine_type fix. Also a stale test-model-default-resolver fix + UPSTREAM.md row 38. Gate 41/41. |
||
|
|
8e7f5a87a5 |
feat(gemma): gemma duals → vLLM v0.22.0 (rebase #40391 lean + re-instate #42006) (#287)
Both gemma duals on immutable v0.22.0. int8-mtp = #40391 (rebased, lean diff-apply) + #42006; bf16-mtp = #42006. #40391 rebased onto v0.22.0 (old full-module copies ImportError'd) + re-delivered as install_script diff (−13K lines). #42006 streaming-multi-tool fix re-instated on both (live-repro'd: streamed multi-tool dropped non-last args). Dropped #41800/#41991 (in stock v0.22.0); fixed stale #41800 patches.yml entry. int8 default 98K→262K. Engine-profile-injection gotcha documented in CLAUDE.md. Validated on real v0.22.0 (docker-inspect): int8 pool 447K@262K bench 95.7/125.8, bf16 pool 195K@131K, streaming multi-tool keeps all args on both. Suite 41/41. |
||
|
|
b327f0311a |
refactor(gemma): rename dual slugs + ladder gemma-bf16-mtp to 131K (#286)
Rename vllm/gemma-int8→gemma-int8-mtp + vllm/gemma-mtp→gemma-bf16-mtp (both carry MTP n=4; only KV format differs); kv-calc alias fix. Ladder gemma-bf16-mtp default 32K→131K (measured BF16 KV pool 196,527 tok @ 0.95; verify-stress 8/8, NIAH→120K, bench 118.8/154.3, zero TPS regression). Codify the dual-card max-context-priority rule in DUAL_CARD.md. Suite 41/41. |
||
|
|
6cafaf80b6 |
Operational robustness (#281): orphan-safe switch.sh · reboot-surviving vLLM · multi-GPU power sweep (#285)
Re-bases tekgnosis-net's #281/#282 onto master: (1) switch.sh registry-derived VARIANT_CONTAINER closed-world teardown (+--remove-orphans) — fixes beellama/ik-llama/sglang VRAM leak; (2) 29 vLLM composes restart: ${CLUB3090_RESTART:-unless-stopped} (reboot survival, opt-out knob); (3) power-cap-sweep.sh multi-GPU (board-power sum + cap restore). 3 new tests; suite 41/41. Closes #281, supersedes #282.
Co-Authored-By: tekgnosis-net <[email protected]>
|
||
|
|
611c430f0a |
beellama Gemma-4 ctx: single 128K (caveats) + dual 262K parked/upstream-gated (#284)
Single beellama/gemma-dflash 128K (caveats; 140K OOMs under load, CTX_SIZE=102400 agent-safe). New dual beellama/gemma-dflash-dual parked upstream-gated (262K works but DFlash multi-GPU broken in 07ac3ce; fixes on v0.3.0 dev branch, no release). switch.sh --list hides upstream-gated (parked) like deprecated with its own note. UPSTREAM row tracks Anbeeld #39. Suite 38/38. |
||
|
|
6274a341c2 |
switch.sh --list: show max context per slug (+ registry<->compose drift) (#283)
Adds the max context size to switch.sh --list. registry-emit threads a ctx label through the variant TSV (before status_note); switch.sh renders it rightmost. Production -> bare ctx; caveats/NA -> folded into the health paren (comma-separated). Rounded to nearest K (163840->164K; 32768 reads 33K). Shows registry vs compose default as a single value when they match, 'validated/compose' (e.g. 164K/200K) when they drift -- the only 3 drifts are experimental lanes; all production/caveats match. Suite 38/38. |
||
|
|
4d47d77fce |
Prune dual vLLM composes: qwen-27b -> one config; gemma-31b default -> gemma-int8 (#279)
Settle the per-model dual vLLM set. qwen3.6-27b dual -> ONE (vllm/dual fp8 262K vision MTP); deprecate dual-dflash/dflash-noviz/tq3-nomtp/bf16/int8. gemma-4-31b dual -> keep TWO: DEFAULTS moved gemma-mtp -> gemma-int8 (full 262K + vision + 4 streams; rides v0.21.0+#40391 overlay), gemma-mtp kept as the stable v0.22.0 32K fallback. Registry entries kept (deprecated, not deleted) so patches.yml + hardware-gating tests stay valid. Out of scope: carnice/qwopus fine-tunes, gemma-26b-a4b, multi4. Suite 38/38. |
||
|
|
4ff14090c9 |
Gemma vLLM -> v0.22.0: bump dual gemma-mtp, deprecate gemma-mtp-tp1 (#278)
Bump vllm/gemma-mtp (dual, bf16+MTP) to stable v0.22.0 (validated 5/5 coherent). Deprecate vllm/gemma-mtp-tp1: fp8 KV is hardware-impossible for Gemma 4 on Ampere sm_86 (fp8e4nv kernel unsupported; fp8_e5m2 rejected by gemma4 attention allowlist; nvfp4 Blackwell-only), live-confirmed on v0.22.0. Remove the dead DEFAULTS[(gemma-4-31b,vllm,single)] row; repoint launch/setup/preflight single-card Gemma -> beellama/gemma-dflash. Also fixes a pre-existing unterminated NEXT_STEPS_NOTE string in setup.sh's gemma-4-31b arm. Suite 38/38. |
||
|
|
bca55e54f2 |
Deprecate Genesis vLLM composes; vLLM single default → vllm/minimal (#276)
* Deprecate Genesis vLLM composes; repoint vLLM single default to vllm/minimal Mark all 8 Genesis-loading qwen3.6-27b vLLM composes deprecated (registry status + compose-header Status): vllm/default, vllm/long-text, vllm/long-text-no-mtp, vllm/long-vision, vllm/bounded-thinking, vllm/tools-text, vllm/dual-turbo, vllm/dual-tq3-mtp-genesis. Genesis is on hold pending Sandermage's next stable release; the stack is moving to stable vLLM + beellama single-card. Repoint DEFAULTS[(qwen3.6-27b, vllm, single)]: vllm/default -> vllm/minimal (the Genesis-free fp8-KV config). The `vllm/default` token now resolves single -> vllm/minimal, dual -> vllm/dual. The model single-card default stays beellama/dflash. Resolver guard tests updated (vllm/default single -> vllm/minimal); SINGLE_CARD.md example commands redirected off the deprecated slugs. Suite 38/38. Follow-ups: (1) bump vllm/minimal + vllm/dual to v0.22.0 (Genesis-free -> stable); (2) SINGLE_CARD.md narrative reframe (single-card vLLM deprecated, beellama default). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> * switch.sh --list: hide deprecated by default, reveal with --all With 9 Genesis composes now deprecated, --list was cluttered. Hide deprecated variants from the default --list (tally + display loops); --all reveals them alongside other-topology variants. Footer shows "(+N deprecated hidden --all)" so they stay discoverable, never silently dropped. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> --------- Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
662122b05d |
Promote beellama/gemma-dflash to single-card Gemma-4-31B default (#272)
Gemma single-card had NO functional default (vLLM single is upstream-gated at head_dim=512; ik-llama walls ~24K; stock llama.cpp ~12 TPS). beellama is the only viable fast single-card Gemma-4 path, so promote it to fill the gap. Registry: beellama/gemma-dflash experimental -> caveats + new DEFAULTS[(gemma-4-31b, beellama, single)]. Resolver now returns it for Gemma single (beellama is #1 in ENGINE_PREFERENCE[single]); Gemma dual unchanged. Validated #441: 131-200K ctx, 47/88 TPS, 8-pack 109/150 think-off / 114/150 think-on (reasoning net-positive on Gemma). Reuses the multi-arch image from #271 (sm_86/89/120; sm_89/120 compiled-not-validated). Docs: compose header (caveats + Quality + ctx ceiling), README Gemma row, BENCHMARKS row marker + intro, UPSTREAM.md row. Resolver guard test flipped from a degradation assertion to beellama/gemma-dflash; full suite 38/38. Re-test trigger: re-point the default to the no-fork mainline path when llama.cpp#23398 (Gemma-4 MTP) merges (docs/UPSTREAM.md). Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
2b671d0552 |
Promote beellama/dflash to single-card Qwen3.6-27B default (#271)
Make `beellama/dflash` the single-GPU default for Qwen3.6-27B, served via an unofficial multi-arch image, and demote ik-llama to the balanced alt. Decision basis (single-3090, 370 W; bench.sh n=5 + quality-test --full): - Code TPS ~100 vs ik 69 — fastest single-card 3090 code path. - 8-pack 107/150 (71%) think-off vs ik 99/150 (66%); 113/150 think-on (same-session). DFlash is output-lossless, so quality == the Q5_K_S target. - Context ceiling ladder (measured): 130K comfortable, 160K usable (115K-tok prefill OK, 0.8 GB headroom), 200K boots but OOMs on prefill. Ships 102K default; raise to 160K via CTX_SIZE. - ik still wins narrative TPS (63), VRAM, and shipped vision — kept as the balanced alt (`--variant ik-llama/iq4ks-mtp`). Registry: `beellama/dflash` experimental -> caveats + a new DEFAULTS[(qwen3.6-27b, beellama, single)] row. beellama is already #1 in ENGINE_PREFERENCE[single], so the resolver now returns it for Qwen single. Image: unofficial multi-arch ghcr build `beellama-cpp:multiarch-b9459-07ac3ce` (sm_86/89/120 = RTX 3090/4090/5090), built from Anbeeld/beellama.cpp .devops/cuda.Dockerfile with CUDA_DOCKER_ARCH="86;89;120" + GGML_CUDA_FA_ALL_QUANTS=ON. Engine reports `ARCHS = 860,890,1200`; cuobjdump confirms all 3 SASS. CAVEAT: sm_89/sm_120 are compiled but UNVALIDATED — only sm_86/3090 is verified on-rig. Also bumps the gemma beellama compose to the same multi-arch image (stays experimental — not promoted). Docs: README single-card framing, INFERENCE_ENGINES.md, UPSTREAM.md row, BENCHMARKS.md Qwen beellama row. Resolver guard tests flipped to the new default; full suite 38/38. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
d2e63b06b6 |
feat(beellama): add beellama.cpp DFlash as first-class compose engine (#268)
Onboards Anbeeld's llama.cpp fork (beellama.cpp) as a registry engine
with two single-card DFlash composes, both experimental pending a
published Docker image:
- beellama/dflash Qwen3.6-27B Q5_K_S + Anbeeld DFlash-IQ4_XS,
q5_0/q4_1 KV, 102K ctx, port 8060
- beellama/gemma-dflash Gemma-4-31B Q4_K_S + Anbeeld DFlash-IQ4_XS,
q5_0/q4_1 KV, 102K ctx, port 8061
Composes reference the locally-built ${BEELLAMA_IMAGE:-beellama-cpp:local}
(entrypoint /app/llama-server) — upstream ships no pullable image, so
they ship as Status: Experimental and surface as (NA: experimental) in
--list (launch is --force-gated). Configs follow the fork's quickstart
docs (Qwen "Precision combo", Gemma "Balanced combo"), deviating only on
the stack-wide --reasoning off default.
Catalog wiring: engine profile beellama-local.yml; two DFlash GGUF
drafters (anbeeld-qwen-dflash, anbeeld-gemma-dflash); compose_registry.py
entries with kvcalc_key="SKIP" (llama.cpp-family, like ik-llama);
weights-map entries (target + draft GGUFs) on both models; q4_1 added to
rtx-3090 supported KV; INFERENCE_ENGINES + UPSTREAM notes; test counts
bumped (engines 8->9, drafters 6->8, registry 45->47, disk 46->48).
Default-resolver invariant preserved: beellama is #1 in
ENGINE_PREFERENCE[single], but its (NA) status + no DEFAULTS row means
the resolver skips it, so qwen3.6-27b single still resolves to
ik-llama/iq4ks-mtp (asserted by test-model-default-resolver.sh). It
auto-promotes only when a published image lands and the composes flip to
production. Full catalog suite green (only the pre-existing
test-submit-bench.sh fixture failure remains).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
|
||
|
|
26eac83f76 |
feat: model-default resolver + user-pinnable defaults (PR-B) (#266)
* feat(switch): model-default resolver + user-pinnable defaults
Add a two-layer default scheme on top of the existing <engine>/default map.
<engine>/default stays the maintainer's recommendation (read-only to users);
the new <model>/default is the user's preference — their .env pin if set,
else a curated pick for the detected topology.
compose_registry.py: two maintainer knobs next to DEFAULTS —
RECOMMENDED_DEFAULT_MODELS (short opt-in shortlist, not an exhaustive
ranking; new models are NOT auto-added) and ENGINE_PREFERENCE (per-topology
engine order). Plus resolver helpers (curated_default_target,
community_default_target stub, model_default_pin_key, engine_set/model_set,
slug_topology, model_of_slug).
registry-emit.sh: shared resolver model_default_target(root, model, topology)
— the single injection point for both launchers. Precedence ladder: --variant
(caller) -> .env pin -> community seam (None today) -> curated ENGINE_PREFERENCE
walk (skips non-functional (NA) slugs) -> degradation (notice + nearest-lower
topology, else a clear "pick explicitly" message; never crashes). Plus
x_default_dispatch: X/default with X in engine-set -> engine rec; X in
model-set -> model default; else error (engines + model-ids are disjoint).
switch.sh: <model>/default token; --set-default <slug> / --clear-default
<model> (round-trip the .env pin CLUB3090_DEFAULT_<MODELID>); a Defaults view
appended to --list (also standalone via --defaults) marking user-pin vs
curated. PR-A's grouping/markers/counts and the --force gate preserved.
launch.sh: bare invocation -> first installed shortlist model -> its
<model>/default (no full wizard); a pinned fast-path ("Launch your default
<slug>? [Y/n]"); a post-boot offer ("Make <slug> your default? [y/N]"). Any
narrowing flag keeps the explicit wizard path. Also load CLUB3090_DEFAULT_*
pin keys from .env even when MODEL_DIR is exported in the shell (the existing
.env loader is MODEL_DIR-gated, which would otherwise hide the pin).
Pin validation is warn + fall back, never blocking: unknown slug / wrong
model / topology-mismatch / (NA) status -> notice + curated default.
Tests: new test-model-default-resolver.sh (curated walk, (NA) skip,
degradation, X/default dispatch, pin override + all validation paths,
community seam skipped, .env round-trip); test-default-resolver.sh extended
with <model>/default launch.sh dispatch. Full suite green (only the
pre-existing test-submit-bench.sh fixture failure remains); 45 entries
unchanged; kv-calc calibration 17/17.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* docs: document the model-default resolver + user pins
Ship the user/contributor docs alongside the resolver code (repo convention —
update existing docs, no new top-level doc).
- README: single-card realign (ik-llama = fastest blessed single default,
llama.cpp = cliff-immune alternative — matches ENGINE_PREFERENCE single
order); add "pin your default" + <model>/default to Quick start.
- CLAUDE.md (= AGENTS.md): document RECOMMENDED_DEFAULT_MODELS +
ENGINE_PREFERENCE + the shared resolver as maintainer knobs next to DEFAULTS.
- FAQ: extend "switch to a different model" with <model>/default; new "How do
I set my own default config?" Q (two-layer model, --set-default/--clear-
default/--defaults, .env key, warn+fallback validation).
- SINGLE_CARD / DUAL_CARD: the resolver + the per-topology engine order; pin
hint.
- GETTING_STARTED: first-run uses <model>/default + the in-flow "set as
default?" prompt.
- ADDING_MODELS: a new model resolves via ENGINE_PREFERENCE; add a DEFAULTS row
per engine; RECOMMENDED_DEFAULT_MODELS is not auto-grown.
- UPSTREAM: beellama Docker-image row (gates beellama onboarding; the resolver
skips it today and it auto-promotes to single default on catalog).
- models/qwen3.6-27b/CHANGELOG: dated PR-B entry.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
|
||
|
|
1182d6b2c3 |
feat(registry): slug health/availability flag (#265)
* feat(registry): add slug health/availability flag Add a lifecycle `status` to every registry slug so `switch.sh --list`, launch, and switch are no longer blind to a compose's health. Previously status lived only in compose-header comments, which drifted: the Genesis dual compose declared "Working (with Genesis)" while its pin is parked and it won't boot clean — a user could boot a broken slug unknowingly. - compose_registry.py: `_entry()` gains keyword-only `status` (default "production") + `status_note`, validated against the enum (production/caveats/experimental/preview/upstream-gated/deprecated). Add `compose_header_status()` mapping a compose's profile-schema `Status:` emoji to that enum. - Sweep every compose `Status:` header to a canonical enum value and re-flag the non-functional slugs: all *genesis* + gemma-4-31b single fp8 -> upstream-gated; carnice -> caveats; qwopus + qwen-a3b-preview -> preview; bf16/int8 A/Bs, llamacpp bounded-thinking, PRISM/APEX eval lanes, gemma-4-26b-a4b onboarding -> experimental; tq3-mtp -> deprecated. - registry-emit.sh emits `status` + `status_note` as the last two VARIANT fields; both loaders + the parity tests read the extended field list. - switch.sh --list: status marker (caveats -> "(caveats)", the NA set -> "(NA: <word>)"); model/topology grouping preserved. Launch/switch gate: production launches, caveats launches with a notice, NA warns + requires --force. launch.sh surfaces the flag before delegating to switch.sh. - New drift-guard test test-compose-status-drift.sh: registry status in enum, compose header maps to enum, and the two agree. 45 entries unchanged; kv-calc calibration 17/17; full suite green (only the pre-existing test-submit-bench fixture failure remains). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> * feat(switch): add model/variant counts to --list Header line shows supported-model count + total variants with the health split (N production · N caveats · N NA); each model group shows its variant count. Widen the marker column so the (NA: …)/(caveats) markers align. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> --------- Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
5a43d4c475 |
chore(vllm): remove redundant nvlink-* dual composes (NVLink auto-detected) (#257)
The four dual/autoround-int4/nvlink-*.yml composes (nvlink-fp8-mtp, nvlink-turbo, nvlink-dflash, nvlink-dflash-noviz) and their vllm/dual-nvlink* launch slugs are removed. They were thin Docker Compose `extends:` stubs whose ONLY override was NVLINK_MODE=force_on — fully redundant since every base dual compose now auto-detects NVLink at boot via detect_nvlink.sh (NVLINK_MODE=auto): it probes nvidia-smi and flips on NCCL_P2P_LEVEL=NVL + custom-all-reduce when a bridge is present, else NCCL_P2P_DISABLE=1 + --disable-custom-all-reduce. NVLink rigs get the identical fast path from the base dual compose with no separate slug; force it explicitly with `NVLINK_MODE=force_on scripts/switch.sh vllm/dual` if auto-detect ever misses. Scope (prune only; no engine-image change — the v0.22.0 stable-engine consolidation is a separate follow-up): - delete 4 nvlink-*.yml composes - compose_registry.py: drop the 4 _entry blocks (50 -> 46 entries) - test-compose-registry-disk.sh: count guards 50/51 -> 46/47 - profile_runtime.yml + patches.yml: drop the 4 nvlink blocks/slug refs - test-profiles-compat.sh: retire the C13/E3 NVLink-required-compose scenarios (the estate NVLink-gating code is now dormant; cleanup tracked as a follow-up, code retained) - docs: switch.sh help, AGENTS.md, HARDWARE.md, UPSTREAM.md, patches/README.md, fp8-mtp/dflash/dflash-noviz compose headers, per-model CHANGELOG removal entry; add a docs/README.md pointer to `switch.sh --list` as the authoritative registry-derived compose x slug matrix Historical NVLink bench rows (JusefPol #31, danbedford #74/#92/#96) are preserved in BENCHMARKS.md. Full test suite green (the lone test-submit-bench failure is a worktree-isolation artifact — it needs gitignored results/rebench/ fixtures absent from a fresh worktree; passes on the working tree). Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
a49944162b |
chore(35b-a3b): drop the built-in-MTP A/B compose (vllm/qwen-a3b-preview-mtp) (#262)
fp8-mtp.yml was only ever an A/B reference -- MTP is net-negative on this MoE at vLLM TP=2 (n2 ~= -45% / n3 ~= -51% wall TPS; the built-in-MTP draft forward pays inter-GPU sync the acceptance can't amortize). It served no production purpose, so remove it and its wiring. Removed: - models/qwen3.6-35b-a3b/vllm/compose/dual/autoround-int4/fp8-mtp.yml - compose_registry.py entry `vllm/qwen-a3b-preview-mtp` (port 8052) - its profile_runtime.yml capture block + calibration anchor (-> 17 anchors) - the kv-calc COMPOSE_ALIAS_TEXT token Test/doc updates for the new counts: - test-kv-generic-dense + test-submit-pull: calibration contract 18/18 -> 17/17 - test-compose-registry-disk: registry 50->49, disk 51->50 composes - test-report-calib: container ref -> live vllm-qwen36-35b-a3b-dual - BENCHMARKS.md + KV_MATH.md notes The MTP-net-negative finding stays recorded in the BENCHMARKS production row and learnings/qwen3.6-35b-a3b.md. Full tests/ suite green (sole red is the pre-existing fixture-dependent test-submit-bench, identical on clean master). Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
46b162ce65 |
feat(qwen3.6-35b-a3b): promote dual → 262K + vision Production (vllm/qwen-35b-a3b-dual) (#259)
Promote the 35B-A3B MoE dual compose out of preview after full live validation on 2× RTX 3090 (2026-05-30): max-ctx probe (262K fits, 7.81× concurrency, 2.05M- token KV pool), bench (178/174 wall · 182 decode TPS, CV <0.5%), verify-stress (NIAH-clean to 240K), soak-continuous (0 growth / 0 err / 0 silent / 100% retention), quality --full (det 79/90 = 88%, aider 13/30 — ≥ ik-llama ref), and a live vision smoke (read a test image correctly @ 262K). - compose: preview.yml → fp8.yml (serving-stack filename per layout convention), preview-mtp.yml → fp8-mtp.yml; fp8.yml = 262K + vision-on + no-MTP, vLLM v0.22.0 stable (no overlays), ✅ Production header. - registry: slug vllm/qwen-a3b-preview → vllm/qwen-35b-a3b-dual, max_ctx 16384 → 262144, compose_path, DEFAULTS, kvcalc alias (kv-calc.py). - profile_runtime capture re-synced to fp8.yml; test-pull slug updated. - BENCHMARKS row + ADDING_MODELS pointer. - MTP intentionally OFF: built-in MTP is net-negative on this MoE under vLLM TP=2 at BOTH n=2 (≈−45%) and n=3 (≈−51%) — the draft forward pays inter-GPU sync the acceptance can't amortize (ik-llama avoids this single-card). fp8-mtp.yml kept as the A/B reference only. GATED: the registry-wide test-diagnose-profile fails for this entry because kv-calc's qwen-MoE branch does NOT divide weights by TP (a deliberate hack: inflated weights ≈ the vLLM-filled peak, which coincidentally passes at 16K but starves the 262K KV in the fit-check). The config is empirically validated; this is a kv-calc MoE-VRAM-model limitation. Lands AFTER the kv-calc fix PR (separate) that distinguishes diagnose fit-need from the calibration filled-peak for cheap-KV MoE. Calibration anchor left untouched here (the re-baseline belongs to that PR). Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
45686808cb |
Prune Gemma vLLM variants (9->3) and pin survivors to v0.21.0
Every gemma-4-31b vLLM compose pinned a purged Docker Hub nightly (e47c98ef / bf610c2f), un-bootable for fresh users (#250, #167 class). Prune the 9-variant set to 3 and repoint survivors onto the immutable stable release tag vllm/vllm-openai:v0.21.0 (never :latest). Survivors: vllm/gemma-mtp (bf16 dual, default), vllm/gemma-int8 (int8-PTH dual), vllm/gemma-mtp-tp1 (fp8 single). Dropped: gemma-dflash, gemma-dflash-int8, gemma-int8-tq3, gemma-bf16, gemma-awq; the 262K path folds into gemma-int8 via a CTX env override. Gemma is decoupled from the shared Qwen nightly profiles via a new vllm-gemma-stable engine profile; Qwen nightly profiles are untouched. Diagnose decoupling: the pruned gemma dflash patch dirs doubled as the diagnose tool's cross-model disk-source proxies for two Qwen overlays (vllm-pr41703-dflash, vllm-pr42102-dflash-kv-quant). Those capabilities are image-baked in the pinned nightly, not mountable files, so mark them image_baked and have diagnose skip the disk-source check for image-baked overlays. Runtime-neutral (no Qwen compose mounts them). switch.sh launcher: handle the new VLLM_IMAGE engine-pin export (the twin loop in launch.sh had it; switch.sh did not -> "unexpected engine pin export"), and guard the VLLM_NIGHTLY_SHA echo for the image-only stable profile (was unbound under set -u). Validated on 2x RTX 3090 (Ampere, v0.21.0): full scripts/tests/*.sh suite 35/0; live boots of gemma-mtp (bf16) and gemma-int8 serve coherent output, qwen vllm/dual regression-clean. gemma-mtp-tp1 (fp8_e4m3) is correctly SM-gated (required_sm=9.0) -- confirmed Ampere-incompatible (Triton "fp8e4nv not supported on sm_86"), a documented 32 GB+/Hopper variant; beellama remains the single-card gemma path on Ampere. Refs #451, #250, #167. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
a289eccd5a |
Fail loud on hollow measured records in measurement_record (#251)
The producer could write a `boot-fit-measured` record with null decode TPS when bench.sh output drifted (summary block absent/unparseable) or a metric was missing. A measured record with null TPS is worse than no record for optimizer calibration: it looks like real data. For a measured result_class: - raise MeasuredRecordError if no parseable bench summary block was found (decode_TPS mean= absent => output drift), or if the parse produced no decode TPS. This matches the module's existing fail-loud posture (KeyError on an unknown registry tag). - a malformed/absent `=== GPU state ===` line is a SOFT gap (VRAM is a fingerprint extension, not the core measured TPS): surface it in a new top-level `parse_warnings` list instead of raising, so the gap is explicit and never a silent null. Non-measured classes (predicted/derived) impose no decode-TPS requirement; genuinely-optional optimizer fields stay null. The CLI catches MeasuredRecordError (exit 2, clean message) and echoes warnings to stderr. Extends the test with: measured + no decode summary => fail loud; measured + bad/absent GPU line => parse_warnings; non-measured + no decode => no raise; and a happy-path regression asserting empty parse_warnings. Addresses the Codex review Medium finding on #249. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
3cdd276756 |
Add measurement-record producer for bench runs (#249)
Pure additive capture: parse a completed scripts/bench.sh stdout (no GPU, no live model) + a compose_registry tag into ONE measurement-record JSON, written to a per-rig gitignored corpus (results/measurement-records/). Conforms to the optimizer design's FROZEN measurement-record field names verbatim; optimizer-only fields (objective/confidence_tier/margin_applied) emitted null, never fabricated. Producer-proposed additions (a context-depth TPS ladder + power_cap_w fingerprint) namespaced under measured_extensions and flagged as Lock-criteria #6 candidates. No consumer, no lookup, no decision logic — that half is gated behind a separate design unlock. bench.sh left untouched (load-bearing); emitter runs standalone on saved bench output. Test runs GPU-free via bench.sh BENCH_MOCK output. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
7df24a46fb |
Add ik-llama apex-fit-q8q5 variant for Qwen3.6-35B-A3B (#242) (#243)
* Add ik-llama apex-fit-q8q5 variant for Qwen3.6-35B-A3B (#242) Captures @laurimyllari's `--fit` + asymmetric q8_0(K)/q5_0(V) KV config from discussion #241 as a single-card ik-llama variant on the APEX I-Compact GGUF (registry tag `ik-llama/apex-fit-q8q5`, port 8057). Measured 1× 3090 + 370 W, n=5: q4/q4 mtp.yml baseline (first real run — APEX weights weren't on disk before today): 96.47 / 144.01 wall TPS narr/code (CV 5.4% / 4.5%) 20.46 GB VRAM q8/q5 fit-mtp.yml (this variant): 103.25 / 149.12 wall TPS (CV 3.0% / 1.5%) — +7% narr / +4% code at tighter CV, +0.6 GB VRAM The Anbeeld K-high/V-low asymmetric-KV pattern materialised here. Gates: verify-full 8/8 PASS verify-stress 8/8 PASS incl. 180K NIAH (91% of n_ctx 196608) bench above soak-continuous PASS (0 errors, 0/25 silent_empty, 0 VRAM growth, 100% TPS retention, p50 decode 223 TPS) deterministic q 76/90 = 84% on PR #38 verifiers toolcall 14/15 · instructfollow 15/15 · structoutput 14/15 · dataextract 11/15 · reasonmath 12/15 · bugfind 10/15 MoE × MTP sub-question (from #242 body) — answered: ik-llama built-in MTP on the 35B-A3B MoE does NOT pay the vLLM −51% / −35% penalty (cf. `qwen3.6-35b-a3b/dual/preview-mtp.yml` BENCHMARKS row). MTP context ready at n_ctx=196608, decode bursts 270+ TPS in soak. The MoE×MTP penalty is vLLM-scheduler-specific, not architectural. Status set to ⚠️ Production w/ caveats because sandbox-pack quality (hermesagent-20 0/20, aider-polyglot-30 0/30, cli-40 timeouts) hit benchlocal-cli sandbox infrastructure issues 2026-05-28 — hermes `tool_events=0` after the agent runs (suspected PR #38 thinking-on sampler interaction); aider fails at git checkout `fatal: path 'aider/__init__.py' does not exist in 'f46766c'` before any LLM call. Neither is attributable to the model or the compose. Pre-PR-#38 cross-rig reference from @laurimyllari's 4090 in #241: hermes 9-12/20, aider 14-18/30 — the model class is capable; the sandbox state on this rig needs separate work. Catalog gates: test-compose-registry-disk count bumped 55→56 per the documented model-add workflow. test-profiles-compat, test-model- weights-registry, test-switch-registry-parity, test-launch-registry- parity all PASS. Two inherited reds (test-compose-mounts-resolve and test-patch-attribution) unchanged from master baseline — they affect qwen3.6-27b vLLM composes, not this addition. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * fit-mtp.yml: tighten --fit/--no-mmap/--cache-ram rationale (PR #243 review) @laurimyllari clarified on PR #243 that for the I-Compact GGUF (~17 GB on 24 GB card) both `--fit` and `--no-mmap` are largely inert since the model fits in VRAM — the reason to keep them as defaults is forward-compat: swapping GGUF_FILE for a bigger quant (UD-Q8_K_XL, APEX Quality, etc.) "just works" with reasonable partial-offload performance, without re-tuning the compose. Replaced the two "Why" blocks in the compose docstring to lead with his framing. Also flagged `--cache-ram 4096` (lower than ik-llama's 8192 default) with the same suspected forward-compat rationale, pending his confirmation on the PR. No flag changes; docstring-only. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * Promote fit-mtp.yml ✅ Production + update sandbox-pack quality Updates the compose Status + Quality line + BENCHMARKS row after live validation of the sandbox-pack fixes (benchlocal-cli #42/#43/#44 + club-3090 #245, all merged today): hermesagent-20: 0/20 → 11/20 (55%) via #42 deterministic sampler aider-polyglot-30: 0/30 → 12/30 (40%) via #44 git-checkout from AIDER_DIR cli-40: 11/40 → 12/40 (30%, ±1 noise) via #43 budget fix (helps aider/hermes wall-clock; cli-40 "timeouts" turn out to be sandbox-internal agent-give-up, not wall-clock budget — see follow-up benchlocal-cli issue) All gates clean: verify-full 8/8, verify-stress 8/8 (incl. 180K NIAH), bench n=5 (+7% narr / +4% code vs q4/q4 mtp.yml), soak-continuous PASS (0 errors, 0 silent_empty, 0 VRAM growth, 100% retention). Status: 🧪/⚠️ → ✅ Production. Caveats trimmed to the 3090 power-sensitivity note. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> --------- Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
aa2a965bb9 |
fix(profiles): cover the new ik-llama PRISM/APEX presets in the compat catalog
Follow-up to #235 — the PRISM/APEX presets exercised (model, engine, KV) combos the profile catalog didn't know about, reddening test-profiles-compat and test-diagnose-profile (caught by the pre-tag FULL-suite run, which the narrower #235 gate subset had missed). - engines/llama-cpp-mainline.yml: supported_model_families += qwen3-next-moe (the 35b-a3b APEX MoE; llama.cpp serves the MoE GGUF — the family list was just incomplete). Fixes constraint C10 for the 3 APEX entries. - hardware/*.yml (all 9): supported_kv_formats += q5_0, q8_0 — the llama.cpp KV quants the engine already lists but no hardware profile did, so the q8_0 presets (prism-pro-dq-dual-vision, apex-mtp-compact-long, apex-mtp-quality-dual) failed C5 "kv not supported by hardware". q5_0/q8_0 are software KV quants that work on any CUDA card (laurimyllari runs q8_0 on a 4090). - patches.yml: register the APEX chat-template (apex-qwen-chat-template) with a symmetric-protocol drift_guard — resolves the orphan #235 introduced into test-patch-attribution. Validation: test-profiles-compat + test-diagnose-profile PASS. APEX patch- attribution orphan resolved (only the 2 PRE-EXISTING sglang orphans remain = the v0.8.5 baseline). Full suite 29/34; the 5 remaining fails are all pre-existing-at-v0.8.5 or test-isolation, NONE from the v0.8.5..master range: generate-compose (pre-existing), setup-picker (mock 5090 rig), submit-bench (0 fixtures), loop-input (test-hwdetect .pull-captures cross-contamination — passes in isolation). Leak-clean. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
458c473908 |
feat(ik-llama): PRISM-PRO-DQ + APEX-MTP presets, conformed to the <quant>/ layout
Ports VykosX's #223 preset families onto the post-refactor convention (PR-A #231 layout + PR-B #234 registry-derived launchers). - 8 ik-llama presets under models/<model>/ik-llama/compose/<topology>/<quant>/: - qwen3.6-27b: ex0bit-prism-pro-dq/{mtp,long,two-stage} (single) + {mtp,mtp-vision} (dual) - qwen3.6-35b-a3b: mudler-apex-compact/{mtp,long} (single) + mudler-apex-quality/mtp (dual) [one repo, two quant files — the canonical two-slug case] - +1 ../ depth on all relocated mounts (models-cache + APEX chat-template). - weights variants added to models/*.yml: ex0bit-prism-pro-dq, mudler-apex-compact, mudler-apex-quality. - 8 compose_registry.py entries (weights_variant=slug, kvcalc_key=SKIP); launchable via the registry — no launch.sh edits needed (PR-B derives them). - catalog-size guard 47 -> 55. Dropped from #223 as obsolete/out-of-convention: the launch.sh hardcoded-array edits (superseded by PR-B) and the root UPSTREAM_CHANGES.md. Validation: 7/7 gates PASS (registry-disk incl. quant-slug<->weights, mounts-resolve, model-weights-registry, switch+launch parity, default-resolver, launch-compat) + kv-calc 22/22. Structurally validated; live boot is weights-gated (community GGUFs not on our rig) — ships community-experimental, crediting VykosX's reported numbers. Co-Authored-By: VykosX <[email protected]> Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
a0520e2060 |
refactor(launch): derive launcher tables from the registry + <engine>/default resolver
PR-B of the compose-quant-hierarchy refactor (follows PR-A #231). launch.sh no longer hardcodes its variant/port/container/kvcalc maps — it derives them from compose_registry.py via a shared emitter, the same single source of truth switch.sh already uses. Adds a topology-aware <engine>/default resolver. - scripts/lib/registry-emit.sh (new): shared emitter — emits switch/launch tables from COMPOSE_REGISTRY, parses container_name from each compose, and exposes registry_default_target() for <engine>[/<topology>]/default. - launch.sh: drop the hardcoded `declare -A LAUNCH_*` maps; populate from the emitter. PRIMARY_MODEL constant (default qwen3.6-27b) backs bare <engine>/default. - switch.sh: drop inline derivation; resolve vllm/default, vllm/dual/default, vllm/multi4/default via the shared emitter. - compose_registry.py: add kvcalc_key to entries; tools/kv-calc.py aliases. - bench-row-formatter.sh: DEFAULTS-aware compose_display (drops the now-dead docker-compose.yml branch). - tests: test-launch-registry-parity.sh + test-default-resolver.sh (new); test-switch-registry-parity refreshed onto the shared emitter. Default resolution is an alias layer over existing registry keys — no keys renamed, no compose content changed. Side effect: the launcher<->registry drift that left vllm/gemma-mtp pointing at a non-existent dual/.../fp8-mtp.yml is gone (now derives the registry's bf16-mtp.yml). Implemented via Codex per the PR-B brief; independently re-validated before commit: 7/7 gates PASS (switch+launch parity, default-resolver, launch-compat, registry-disk, mounts-resolve, kv-calc 22/22); live vllm/default->dual + vllm/dual/default->dual both verify-full 8/8; leak-clean. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
eaa7a8c1ad |
docs+scripts: finish <quant>/ path migration across full repo sweep
Repo-wide follow-up to the compose quant-layer move (
|
||
|
|
9821c94efb |
refactor(compose): insert <quant> layer + make the registry the single source of truth
Restructure all 47 composes to models/<model>/<engine>/compose/<topology>/<quant-slug>/<serving>.yml and centralize defaults as registry pointers. Move + path-rewire only — no compose runtime-config changes (plus the +1 ../ depth bump each moved file requires, and descriptive names for the former docker-compose.yml defaults). - <quant-slug> = 1:1 with a weights artifact == compose_registry weights_variant == weights.py key (provider-quant form). Fixes the prior coarse `gguf` and the bf16/int8-files-mislabeled-as-autoround_int4 weights_variant. - default = DEFAULTS[(model,engine,topology)] pointer in compose_registry.py; no default.yml / docker-compose.yml. KV-dtype variants (bf16/int8) stay as files under autoround-int4/. - +1 ../ on every relative mount in moved composes (cache/patches/scripts/models-cache). - Rewired: compose_registry (47 compose_path + weights_variant + DEFAULTS), launch.sh, gpu-mode.sh, bench-row-formatter, profile *.yml, generate_compose.py, weights.py (+ back-compat ALIASES), AGENTS.md, ADDING_MODELS.md, BENCHMARKS + doc path refs. - New guard tests: test-compose-registry-disk.sh, test-compose-mounts-resolve.sh. - Variant keys (vllm/dual, ik-llama/iq4ks-mtp, ...) unchanged — launch/switch UX preserved. Validation: docker compose config 47/47 · parity/launch-compat/registry-disk/ mounts-resolve/profiles-compat/model-weights-registry PASS · kv-calc 22/22 · leak-clean. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
b116750ef1 |
fix(registry+bench): sync vision defaults to the 2026-05-25 re-tune (#438)
Follow-up to PR #227 (#437). The vision re-tune changed the compose defaults but left the registry / switch.sh / BENCHMARKS describing stale values. - compose_registry: llamacpp/mtp-vision max_ctx 49152 -> 150000 (matches the re-tuned compose default; ik-llama/iq4ks-mtp-vision already correct at 163840). switch.sh derives variants from the registry, so this also fixes its dynamic list. - switch.sh usage comments: both vision variants now show the 1M-px default + "full-res = override, lower ctx" note (were "49K" / bare "160K"). - BENCHMARKS: marked the superseded 2026-05-20 49K row; added two re-tuned vision rows (llama 150K@1M-px fills 138K/561MB; ik 160K@1M-px fills 147K/503MB, graduated EVAL #402 -> Production). Decode TPS carried from siblings (ctx-independent, not re-benched); rows focus on the measured ctx-ceiling + VRAM + verify-stress [8/8]. ik vision was ALREADY registry-wired (parity test green); this is the data sync. Refs #438. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
5f37ae6d4d | Add profile-backed model weight fetch registry | ||
|
|
0b4e694ad3 |
feat(llama.cpp): Structured-CoT bounded-thinking compose + grammar-dialect fix (#214)
* Add llama.cpp bounded-thinking compose * fix(structured-cot/llamacpp): ship no-underscore grammar variant + record on-rig validation On-rig validation found the xgrammar deepseek-scratchpad.gbnf does NOT parse under llama.cpp's GBNF — underscored rule names (line_char/line_chars) trip "parse: error parsing grammar: expecting newline or end at _char", and llama-server silently falls back to UNCONSTRAINED generation (FREE==FSM token counts — silent no-fire). Fix: deepseek-scratchpad.llamacpp.gbnf — identical language, rule names with no underscores (llama.cpp GBNF is [a-zA-Z0-9-] only). Validated 2026-05-24 on 1x 3090: parses, fires (reasoning_content = PLAN/NOTE/VERDICT), coexists with --spec-type draft-mtp (draft accept 0.66-0.88). FREE->FSM 2113->1009 tok (2.1x) and 934->331 (2.8x) @ temp 0.6. Repointed the llama.cpp compose + READMEs + STRUCTURED_COT.md llama.cpp examples to the variant; left the vLLM grammar + refs untouched. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * docs(structured-cot): redact pre-existing /home/wasif absolute paths Pre-commit leak scan caught host-local /home/wasif/structured-cot/runs/... paths in the 'What's saved on disk' section — pre-existing (2026-04-30 doc), not introduced by this branch. Redacted to host-local relative form per no-internal-paths-in-public-docs. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> --------- Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
7f733618e6 |
docs+registry: surface ik-llama on the single-card front door + fix stale max_ctx
ik-llama/iq4ks-mtp is the fastest single-card path (~18-20% faster decode + leanest VRAM, #184) but was absent from the README quick-start and the registry ctx values had drifted from the 200K shipped default. Keeping llamacpp/default = mainline (the simplest / clean-upstream-image pick) — surfacing ik, not renaming. - README quick-start: add `ik-llama/iq4ks-mtp` (fastest single-card) alongside the llamacpp/* variants. - docs/SINGLE_CARD.md: one-line "simplest (llamacpp/default) vs fastest (ik-llama/iq4ks-mtp)" steer atop the config table (ik rows were already present + ⭐-marked). - compose_registry.py: fix stale wizard-projection max_ctx to the 200K default — llamacpp/default + llamacpp/mtp were 131072 (too low), ik-llama/iq4ks-mtp was 262144 (boots-not-fills); all → 200000. Vision entries (49152 / 163840) already correct. Reworded the "262K via -ub 512" comment (that was the boots≠fills false ceiling). NOTE: launch.sh doesn't pass max_ctx to runtime, so this is wizard-projection accuracy only — runtime ctx still comes from the compose CTX_SIZE=200000 default. Validated: registry imports; test-launch-compat, test-switch-registry-parity (45 composes, parity), test-profiles-compat all pass. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
477873af93 |
fix(#168): scope report.sh kv-calc calibration to the running model
`report.sh --full` always dumped kv-calc's full catalog-wide calibration
matrix (4 models × every compose), regardless of what the reporter runs, and
the vLLM-only skip didn't cover ik_llama. Three fixes:
(a) Scope to the running model. Resolve the active container -> kv-calc model
id and filter `--calibration` output to that model's `== <id> ==` section.
A Qwen single-card reporter no longer gets Gemma/MoE rows.
(b) Fix the ggml-engine skip. The skip only matched `llama-cpp-*`; `ik-llama-*`
fell through to "unknown" and ran the (inapplicable) calibration anyway.
Both ggml engines now emit the skip note.
(c) Opt-in full matrix. `--full-calibration` (or REPORT_FULL_CALIBRATION=1)
restores the catalog-wide matrix for maintainer triage. Unknown/unresolved
model also falls back to the full matrix.
Logic is factored into a pure, side-effect-free lib (scripts/lib/report_calib.sh:
calib_engine_for_container / calib_model_for_container / calib_filter_model_section)
so it's unit-testable. New test scripts/tests/test-report-calib.sh covers the
engine map (incl. the ik_llama regression), the model map (all 4 models), and
the section filter (keeps banner + target section + Overall, drops others;
empty id = passthrough).
Validated: bash -n; test-report-calib ok; live `kv-calc --calibration` scoped to
qwen3.6-27b keeps only that section + the Overall line.
Closes #168.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
eba4870893 |
feat(ik-llama): wire iq4ks-mtp + iq4ks-mtp-vision into launch.sh + switch.sh (#189)
The ik_llama composes shipped (#180) but were never registered, so they were raw-`docker compose`-only — awkward now that ik IQ4_KS is the path we point VRAM-tight / WSL single-card users to (its ~0.5-0.8 GB leaner footprint is the one place ik's edge pays rent). Register the two stable variants in compose_registry.py (→ switch.sh derives them) and add them to launch.sh's variant maps (compose/model/engine/kvcalc/order/port/container) so the wizard offers them and `switch.sh ik-llama/iq4ks-mtp` works like the others. The experimental two-stage compose stays raw-compose-only until benched. Verified: registry imports (45 entries), ik present in all 7 launch maps, bash -n clean, composes config-clean. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
ab3c763845 |
fix(llamacpp): pin image to server-cuda-b9246 (rolling tag broke at b9282) (#188)
The upstream rolling `ghcr.io/ggml-org/llama.cpp:server-cuda` tag regressed at b9282 — `llama-server` crash-loops with `libllama-common.so.0: cannot open shared object file` (broken lib packaging). Since our composes defaulted to the rolling tag, every fresh pull of llamacpp/default + llamacpp/mtp-vision hit the crash loop and the endpoint never came up. Reported in #187. Pin both single-card composes to the validated build `server-cuda-b9246` (2026-05-20 — the build all current BENCHMARKS rows were measured on). Override to follow a newer build via LLAMACPP_IMAGE once validated. README + engine profile notes updated to match. Closes #187. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
28cff3f559 |
refactor(llamacpp): collapse single-card composes 3→2 (default = mtp alias) (#181)
* refactor(llamacpp): collapse single-card composes 3→2 (default = mtp alias) Mainline Q4_K_M reaches the full 262K with -ub 512 (validated: boot + serve + verify-stress 7/7 + soak PASS), so the old vanilla Q3_K_XL docker-compose.yml no longer earns a separate profile. Collapse the llama.cpp single-card tree to two compose files: - `llamacpp/default` is now an ALIAS for `llamacpp/mtp` (mtp.yml) — launch.sh + compose_registry repointed. All references (estate-CLI default, tests, docs) keep resolving; the variant name survives. - Retired models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml. - mtp.yml: document the -ub 512 -> 262K recipe in the header. - mtp-vision.yml: align with the ik vision profile — default 160K (was 49K) via -ub 512, and pin --image-min-tokens 1024 / --image-max-tokens 4096 (mainline previously left image bounds at the model default). Validated: full-res 2048^2 image @ ~22.2 GB / 24 (~2.4 GB headroom). - switch.sh / README / SINGLE_CARD.md: descriptions + the dead docker-compose.yml link repointed. Note: on mainline llama.cpp the VRAM lever is -ub (not -b — mainline honors -ub independently, unlike ik_llama which forces n_ubatch=n_batch). Docs follow-up (deferred): SINGLE_CARD.md prose still describes `default` as the old Q3_K_XL path in places; mtp.yml's "froggeric suppresses --reasoning off" note is now disproven (froggeric v19 + --reasoning off works on mainline b9246). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * docs(engines): ik_llama coverage + MTP-merged-on-mainline corrections - Correct stale "llama.cpp MTP via community PR / not merged" across INFERENCE_ENGINES.md, LLAMA_CPP.md, engines/README.md — PR #22673 merged on mainline 2026-05-16 (the b9246 image we bench is post-merge). - Add ik_llama.cpp coverage that was missing: engine list / repo tree / supported-models table (README), workload picker (SINGLE_CARD), choose table (engines/README), model engines line (qwen3.6-27b/README). - Document new ik_llama features: two-stage ngram+MTP (PR #1789), -vhad, --recurrent-ckpt-mode, -ctk-first (IK_LLAMA.md, INFERENCE_ENGINES.md). - IK_LLAMA gotcha: native template won the 8-pack A/B (103 vs 99, toolcall tied) so the ik composes default native; froggeric stays vLLM-only. - New compose iq4ks-two-stage.yml (ngram-mod + MTP via --spec-stage). Boots + two-stage spec-dec engages on cu13-server (validated 2026-05-22); perf bench pending. Review fixes on top of the drafted changes: repo-tree └→├, a 131K↔262K self-contradiction in llama-cpp/README, a now-stale "mainline MTP still open PR" bullet, and softened/attributed the unverified IQ5_KS PPL figures. External PR links verified resolving; no internal paths leaked. Docs drafted via the Qwen coding agent; reviewed, corrected, and the new compose boot-validated by Claude. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * refactor(composes): promote native template default + bump llama-cpp context ceilings llama-cpp/mtp.yml: default ctx 131K → 262K (full native), -ub 512 llama-cpp/mtp-vision.yml: default ctx 49K → 160K, -b/-ub 1024 ik-llama/mtp{,-vision}.yml: native template (froggeric v19 A/B'd, native won 8-pack 103 vs 99), +'-vhad' (V-cache Hadamard), +'--recurrent-ckpt- mode auto' (hybrid-aware MTP rejection), removed ipc:host (not needed) Retire stale artifacts: ik-llama/patches/froggeric-chat-template/ — native won, no longer mounted llama-cpp/recipes/single-card-*.sh — host-binary scripts superseded by compose files (canonical path since v0.8.0) 8-pack ship-gate: 99/150 (vs prior native 103/150) — within noise, 6/8 packs within ±1. Delta is cli-40 variance, not a flag effect (-vhad and --recurrent-ckpt-mode are not output-quality levers). * docs(composes): guard -np 1 with hardware-conditional rationale Add inline ⚠ comment to all 5 single-card composes explaining why -np 1 is intentional on a single 24 GB card (compute-bound, not memory-bound — extra slots divide throughput, don't multiply it). Also expose NP as an env var on the two mainline composes (was hardcoded; ik composes already had ${NP:-1}). The 5090/multi-GPU carve-out is explicit: 'on a higher-throughput card or multi-GPU the trade may flip — re-validate before raising.' Prevents future agents from blindly parallelizing slots on Ampere. * fix(composes): move -np guard comment out of folded scalar The ⚠ -np 1 rationale comment was inside 'command: >-' in all 5 single-card composes. YAML folded scalars treat # lines as content, not comments — docker compose config failed on all 5 files. Move the comment block to YAML-level (between volumes: and command:). All 5 now pass 'docker compose -f <file> config'. Lesson learned: always validate compose edits with 'docker compose -f <file> config >/dev/null' before committing. * docs(FAQ): expand WSL2 section with GPU overhead guidance The existing FAQ entry only mentioned TDR and expandable_segments gotchas. Expanded to explain the ~1.3 GiB invisible WSL2 overhead, how it affects each engine path differently (dual-card: noise, single-card vLLM: one env var, single-card llama.cpp: lower ctx), and provide a concrete VRAM budget table showing which composes OOM on WSL2 vs which fit at defaults. Key finding: mainline llama.cpp mtp.yml at 262K leaves only ~1.5 GB headroom on Linux — WSL2 eats that to ~0.2 GB (instant OOM). Fix is CTX_SIZE=131072 (drops to 20 GB, ~2.7 GB WSL2 headroom). ik_llama composes (IQ4_KS, smaller weights) fit at defaults on WSL2. No compose file changes — this is FAQ docs only. The llama.cpp composes already expose CTX_SIZE as an env var override. --------- Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
4a53edab43 |
llama-cpp: switch to rolling :server-cuda tag (no patches → no pin needed)
Follow-up on
|
||
|
|
c3e7c7ed80 |
llama-cpp: replace orphan llama-cpp:local with upstream pinned image (#170)
v0.8.3 shipped composes (llamacpp/default, llamacpp/mtp, llamacpp/mtp-vision) all reference `image: llama-cpp:local`, a custom image that exists ONLY on the maintainer's rig. There is no Dockerfile, no build script, and no setup.sh hook to produce it for users. Anyone running `bash scripts/switch.sh llamacpp/mtp` on a fresh clone hits "image llama-cpp:local not found" and dies at boot. The custom image was a v0.8.3-dev artifact from when MTP PR #22673 was bleeding edge. The official upstream `ghcr.io/ggml-org/llama.cpp:server-cuda` now has it merged (build b9246 = commit 871b0b70f, 2026-05-20) — pinning to b9246 reproduces the v0.8.3 numbers (50.25 narr / 58.04 code on single 3090, vs shipped 51.28/59.72). Surfaced by @zemaphore in discussion #170. README.md was also lying: claimed "both use the official ghcr.io image, no custom build needed" while composes referenced llama-cpp:local. Override the pin via `LLAMACPP_IMAGE=ghcr.io/.../server-cuda-bXXXX` env if you want to follow upstream master. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
ed1507122c |
llama.cpp single: thinking-off policy alignment + MTP profile family
Closes the dataextract 0/15 quality regression caused by mounting the vLLM-only froggeric Jinja template on the llama.cpp engine, which silently suppressed --reasoning off. Native template + --reasoning off restores it. Stack-wide Qwen3.6 thinking-off-by-default policy now shipped on the llama.cpp path (matches 22/24 vLLM composes). Adds three named single-card profiles: - llamacpp/default (262K vanilla, unchanged) — cliff-immune fallback - llamacpp/mtp (NEW: 131K + MTP n=2) — single-card workhorse, ~60 TPS code - llamacpp/mtp-vision (NEW: 49K + MTP + vision) — first stack profile combining MTP + vision on build 9235 (the older "strip mmproj when MTP" rule was obsolete; sweep-verified MTP + vision coexist) Retires single/concurrent.yml — single-card concurrency is anti-value (per-slot ctx ~48K, worse than every other profile). Concurrency belongs on dual. Fixes a latent user-facing bug: engines/llama-cpp-mainline.yml had supported_drafters: [draft-mtp] (the llama.cpp CLI flag value) while the compat-layer C7 gate compares against drafter spec_method (mtp). Net effect pre-fix: llamacpp/mtp would have been silently filtered out of launch.sh candidate lists. Fixed [draft-mtp] -> [mtp]. Lowers -ub default 2048 -> 1024 on the llama.cpp single composes. The per-pass activation peak halves; verify-stress goes 5/7 -> 7/7 including the 60K + 91K needle rungs previously treated as architectural Cliff 2 territory. Cliff 2 single-prompt at 50-60K was config-driven on llama.cpp, not architectural; CLIFFS.md note added (vLLM Cliff 2 narrative unchanged — different kernel-level failure). Migrates 6 stale call sites for the engine-id rename (llama-cpp-mainline -> llama-cpp-local) that earlier work lagged. Measured impact, Config A (llamacpp/mtp, 131K, MTP n=2, ub=1024): - bench: 51.28 narr / 59.72 code decode TPS (n=3, CV 1.9%/0.5%) - verify-stress: 7/7 PASS (incl. 60K + 91K needle recall) - quality 8-pack: 102/150 (68%) — beats every Qwen vLLM dual config in discussion #119 by 6-16 pp - aider-polyglot-30: 17/30 (56.7%) — matches Qwen vLLM bf16 dual exactly (17/30) on half the hardware - per-GPU code TPS (59.72) ~equal to vLLM dual configs (60-63), confirming the engine-side per-card rate is identical and vLLM dual's aggregate advantage is purely from the second card Config B (llamacpp/mtp-vision, 49K, MTP + vision): - multimodal vision probe passed - bench: 56.52 narr / 66.17 code decode TPS (n=5) - verify-stress: 7/7 PASS Tests green: test-launch-compat, test-profiles-compat, test-switch-registry-parity, test-pullgate-gates, test-patch-attribution. Leak-grep clean. YAML lint clean. Performance charts regenerated to include the two new entries. Doc alignment: SINGLE_CARD, CLIFFS, FAQ, EXAMPLES, README, llama-cpp README, COMPOSE_GENERATOR, PULL_GATE all reflect the new family. CHANGELOG.md + UPSTREAM.md left intact per append-only history rule. Followups (queued, not blocking): #400 MTP retune A/B (n=4 + p-min=0 on the fixed config), #403 soak-test container glob (unlocks Cliff 2b validation on llama.cpp), #405 benchlocal-cli port-offset for parallel rebench, #401 + #404 ik_llama + Tom's TurboQuant trackers. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
26949d7fa8 |
fix(pull): v0.8.2 STEP V5 — recommend must not label a fits-clean model "DOES NOT FIT"
_render_recommendation keyed solely on res.ok, so a confirm→proceed / override-accepted terminal (raw_verdict=fits-clean, res.ok=False because the run needs an explicit --yes/--force-download) fell into the generic "DOES NOT FIT / BLOCKED" branch — dishonest by imprecision (the model fits; only acceptance is pending; CONTRACT-4 is "honest recommendation — fits?"). Add a presentation-only needs-acceptance classification: a fits-clean acceptance terminal now renders "FITS (estimated) — NOT YET ACCEPTED" with the acceptance-gate guidance (not the failure on-ramp); genuine hard-blocks still render DOES NOT FIT / BLOCKED. Pure presentation — still derived only from res, no decision logic. Caught by the V5 on-rig gate (microsoft/phi-2 --dry-run). test-pull.sh rec(2a) updated to assert the honest rendering; suite 25/25, kv-calc N/N. |
||
|
|
c5b5e9b27e |
feat(pull): v0.8.2 STEP V5 — recommend UX + report-a-failed-pull doc + §9-reconciliation
CONTRACT-4: add `--recommend` — an honest aggregated recommendation that is PURE presentation/aggregation over the SHIPPED run_pull verdict. Every line is read straight off the real PullResult (ok/confidence/raw_verdict/ terminal/stratum/abort_reason/notices/emitted); it introduces no decision logic and does not change the exit code. Carries the §7 boot-fit≠runtime caveat + soak-continuous pointer ONLY when the gate itself marked the run boot-fit-satisfied (echoed from res.notices, never re-derived), states which gate decided, is vLLM-only by construction, and never implies a non-emitted artifact (the compose line appears only when res.emitted). CONTRACT-1 user doc: docs/PULL.md gains a "Report a failed pull" section documenting the SHIPPED V1/V2 on-ramp (capture-on-hard-block → surfaced pointer → scripts/pull.sh --submit-last / --submit <dir>, consent prompt, gh + gh-less). Every documented command/flag/output string was verified verbatim against the live shipped CLI on this branch (docs-fidelity RED- LINE). Leak-clean: only repo-relative .pull-captures/<slug>/<ts> forms, no absolute paths. §9-reconciliation: the release headline AND the readiness ledger in docs/PULL.md now state explicitly that GGUF is deferred to a §9 cross- engine design-unlock proposal, and that v0.8.2's scope is the failure on-ramp + registry-expansion + whichllm-hw-detect + recommend — same location/pattern v0.8.0 used for its §9-headline reconciliation. Zero decision-logic change: gates.py / deriver.py / capture.py / loop_input.py / classifier.py / dedup.py / submit_pull.py / hwdetect.py / failure_fingerprints.yml / arch_patches.yml all byte-unchanged. pull.py is a pure addition (zero removed lines): a new _render_recommendation() function + a --recommend flag + one presentation-only call site. test-pull.sh adds the CONTRACT-4 V5 section asserting the recommendation TRACKS a real differing verdict — four genuinely-different real outcomes (fit+emitted / confirm→proceed-blocked / estimated-lower-bound-fit / hard-block) render four pairwise-different blocks, each matching its own real res; rig-independent leak assertion (str(root) absent), not a substring allowlist. Full shipped suite 25/25 green in the CI condition; kv-calc --calibration 11/11 unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
39177282b7 |
feat(pull): v0.8.2 STEP V4 — optional whichllm hw-detect subprocess (CONTRACT-3, hw-detect-only)
CONTRACT-3 §8: an OPTIONAL, bounded subprocess that augments hardware ENUMERATION for the eval path where nvidia-smi does not apply (AMD ROCm / Apple / other-vendor). Strictly detect-only; never feeds kv-calc (kv-calc stays the sole fit authority); no new hard dependency. New isolated leaf module scripts/lib/profiles/hwdetect.py: - detect_non_nvidia_hw()/detect_non_nvidia_sm(): bounded `whichllm list --json` subprocess, defensively parsed into a structured HwDetectResult; maps a recognised non-NVIDIA device class to an SM-equivalent for the [C0] SM gate ONLY. - Every non-delivery path (tool absent / failed / timeout / unparseable / NVIDIA-only / unrecognised) degrades to None and NEVER raises out. Additive consume-point wiring in run_pull (the eval path): a new optional `hwdetect_fn` kwarg, consulted ONLY inside the existing `if hardware_sm is None:` stratum-3 block — i.e. only when nvidia-smi already returned nothing. The NVIDIA majority never enters the seam, so that path is byte-identical whether the augment is absent OR present-but-degrading. On a recognised non-NVIDIA device the eval path gets an SM-equivalent (the [C0] SM gate runs instead of the blind-refuse `hardware-sm-undetermined` terminal) plus an additive notice/diagnostic; no shipped decision field is mutated and kv-calc is not consulted. BOTH RED-LINE halves covered + proven by scripts/tests/test-hwdetect.sh: (a) safety — optional/no-hard-dep, graceful degrade, NVIDIA-path byte-identity, never feeds kv-calc; (b) delivery — a simulated (explicit, deterministic) non-nvidia env yields a structured enumeration the eval path observably consumes (outcome moves OFF the degrade terminal). Rig-independent leak assertion (str(abs_dir) not in shared). All 9 shipped v0.8.0/V1/V2/V3 decision modules byte-unchanged (dedup.py:262 FInput.dedup_hash(_EffProxy()) idiom fenced/untouched). Full scripts/tests/test-*.sh suite green in the CI condition (25/25); kv-calc --calibration unchanged (11/11). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
d78b9a9496 |
fix(pull): v0.8.2 STEP V3 — deliver CONTRACT-2's engine-supported broadening (TRC two-class)
The first expansion marked ALL added arches requires_trust_remote_code: unverified, so a registry-recognised model only moved no-arch-row -> needs-trust-remote-code-ack — a lateral relabel, NOT the "materially more models pass [C0] engine-supported" CONTRACT-2 requires (net newly-passing: zero; caught on-rig via microsoft/phi-2). Two-class TRC posture: long-standing native vLLM built-in classes (no remote code — a documented upstream constraint) carry requires_trust_remote_code:"false" with a documented-constraint evidence anchor; arch families with genuine remote-code lineage (Phi3SmallForCausalLM, InternLM2ForCausalLM) stay unverified/fail-closed. Zero-false-pass preserved: gates.py's has_auto_map is an INDEPENDENT OR-term, so a repo shipping auto_map still hard-blocks needs-trc-ack regardless of the row flag — "false" removes only the arch-row-level over-refusal, never the per-repo trust boundary. On-rig (2026-05-18): microsoft/phi-2 (no auto_map) -> engine-supported clean; PhiForCausalLM+auto_map -> needs-trc-ack; absent arch -> no-arch-row; Phi3Small/InternLM2 -> needs-trc-ack. Suite 24/24, kv-calc 22/22. test-pullgate-gates updated to the two-class invariant. |
||
|
|
999c93fe8c |
feat(pull): v0.8.2 STEP V3 — arch-registry expansion + chat-template attribution/drift_guard
CONTRACT-2 (§10-R4) arch-family registry expansion: +13 safetensors arch rows in arch_patches.yml (PhiForCausalLM — the microsoft/phi-2 STEP V1 on-rig no-arch-row anchor — Phi3Small, Gemma/Gemma3/Gemma3-CG, Starcoder2, Cohere, InternLM2, Mixtral/Qwen2Moe/Qwen3Moe MoE, Qwen2-VL). Additive data only, zero [C0]/decision-logic change. Zero false-pass by construction: each follows the established estimated-lower-bound/unverified-TRC precedent so [C0] still resolves needs-trust-remote-code-ack (fail-closed, bypassable ONLY by --trust-remote-code) — the expansion drops only the --experimental-arch requirement, never auto-passes; an arch still absent still hard-blocks no-arch-row. test-pullgate-gates.sh proves both, plus the #146-shape worked acceptance case (a hand-added awq_bf16_int4 weights variant the expanded flag schema/parity machinery absorbs cleanly). CONTRACT-2b-i chat-template attribution + behavioral drift_guard: new `chat_template` delivery class (VALID_DELIVERY_MECHANISM); froggeric (22 composes — 18 direct + 4 nvlink* via REAL Docker Compose extends: merge) and carnice (mount-only) brought under load_bearing_when + a behavioral drift_guard whose check encodes the self-contained symmetric restart+settle protocol (identical docker restart both arms, /v1/models healthy, 60s settle, >=3 bench runs/arm, grand-mean same-segment compare, flag only a 3/3 deterministic regression). Effective coverage uses REAL merge semantics: docker compose config (preferred) or a deterministic offline extends: merge applying the same rules (additive sequence merge; `!reset` removal) — never the unsound single-base text concat. .jinja artifact discovery catches an orphan vendored template. test-patch-attribution.sh adds the class checks + an H4 fixture asserting a `!reset` child AND a stopped-extending child both lose coverage (the false-negative is the dangerous direction). Generator emit kept in lock-step with reaches(). Documented as PATCH_POLICY.md §3.1. Rig-independent leak assertions added (str(abs_dir) not in shared; repo-relative-only — never a /opt|/home substring allowlist). RED-LINE: gates.py/pull.py/deriver.py/capture.py/loop_input.py/ classifier.py/dedup.py/submit_pull.py/kv-calc.py/failure_fingerprints.yml byte-unchanged; no shipped compose changed; patch_attribution.py c0_state/ is_artifact/compose_text/service_body byte-identical (additive only). Full test-*.sh suite green in the CI condition; kv-calc --calibration N/N. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
52451ca0b0 |
fix(pull): v0.8.2 STEP V2 — gh-less issue body must not carry the absolute capture path
The gh-less paste fallback embedded the absolute bundle dir into the
PUBLIC issue body ("full redacted bundle at `/abs/.../.pull-captures/...`"),
violating the acceptance that nothing the on-ramp tells a user to share
contains an unredacted absolute path. Render a repo-relative
`.pull-captures/<slug>/<ts>` pointer instead. Strengthen the gh-less
leak assertion to a rig-independent check (the absolute bundle dir must
not appear; only the relative pointer may) — the prior /opt|/home check
passed under a tmp sandbox dir and missed this.
|
||
|
|
e1cdcb53c7 |
feat(pull): v0.8.2 STEP V2 — surface pointer + --submit-last/--submit (gh + gh-less, consented, F5 reuse)
CONTRACT-1.2: pull prints the honest one-line on-ramp pointer whenever a gate bundle was emitted for the run, keyed on the V1-recorded capture dir — explicitly NOT gated on the exit code (the bypassable no-arch-row C0 advisory path exits 0 yet emits the #1 §10-R9 bundle). Gate path stays I/O-free: a single stdout line, no network/prompt/auto-send. It does not classify (suppression is loop-side at submit). CONTRACT-1.3: scripts/pull.sh --submit-last / --submit <dir> is a distinct top-level verb parsed before the slug/--profile-like requirement. --submit-last re-reads the V1 shared .last marker at submit (the race defense — surfaces the CURRENT bundle, never a silent wrong-bundle). Re-shows bundle identity + the exact already-redacted payload, requires an explicit y before any network, then reuses the shipped F5 dedup.submit (effective_dedup_hash, bounded loop:dedup-<hash> labels, +1-or-open, collision-safe verify, suppression/review-queue) — not reimplemented. gh-less fallback runs post-F2 classification, gated on should_file: should_file=True -> a prefilled public issues/new URL with the loop:dedup-<hash> label and the deterministic title template; review- queued (unknown / correct-refusal) -> the local _review-queue spool path and the no-public-issue line, with NO public issues/new URL. Never raises; degrades to the local spool + printed paste-path. Console is never a submission source — only the redacted artifact is emitted. New scripts/tests/test-submit-pull.sh (mocked gh, zero network): the .last-marker race re-read, the bundle-emitted-but-exit-0 surfacing, F5-reuse, the gh-less should_file branch with no public URL for review- queued, gate-path I/O-free, and leak-hygiene. Full shipped suite green in the CI condition; kv-calc --calibration unchanged at 22/22; safetensors decision path byte-unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
20f1557d29 |
feat(pull): v0.8.2 STEP V1 — capture-on-hard-block pt1-gate emitter + BaseCaptureBundle protocol lift
CONTRACT-1.1 capture-on-hard-block (additive only; zero v0.8.0 decision-logic
change — the safetensors/GGUF paths are byte-unchanged):
- capture.py: new SEPARATE emit_gate_capture() (the emit_override_capture
byte-preserving precedent — NOT invoked by emit_capture()) writing a
pt1-gate.json + schema:2 manifest.json (outcome:hard-block, exact shipped
abort_reason, failure_class:null) per the per-abort-stratum key table
(model/arch/quant null pre-deriver; topology best-effort/nullable, capture-
only resolve; post-C0 always null). New shared write_last_marker() helper
(atomic tmp+os.replace) called from BOTH emit_capture() and the gate
emitter (centralization mandate — gate-only is the commonest failure).
- pull.py: pass-through capture on the 7 terminal hard-block return paths
(deriver / profile-like / hardware-sm-undetermined / C0 / C2a / no-fit-model
/ C1) — emits a bundle before the existing `return res`; the decision is
byte-unchanged; injectable gate_capture_fn; never raises.
- loop_input.py: BaseCaptureBundle typing.Protocol (Optional[dict] pt2-5);
FInput satisfies it by construction (verified: no isinstance(finput,FInput)
anywhere in F2/F5 — pure static retype, schema==1 byte-identical incl.
dedup_hash); new FInputGate + read_gate_bundle() (schema==2; validates ONLY
the always-present row + outcome==hard-block + failure_class is None — does
NOT reuse the 22-key validator); FInputGate.dedup_tuple() uses .get(k,None)
(behaviour-neutral schema-1, crash-safe schema-2, deterministic null-topo).
- classifier.py / dedup.py: F2+F5 parameter annotations retyped FInput ->
BaseCaptureBundle. The dedup.py FInput.dedup_hash(_EffProxy()) unbound-class
idiom is FENCED (unchanged — not tidied). Additive gate_abort_reason
_match_condition kind (reads pt1_gate.abort_reason; bool like sibling kinds;
no enum/routing change).
- failure_fingerprints.yml: seeded gate_abort_reason rules keyed on the
verified shipped abort_reason strings — only engine-support-unknown/
no-arch-row -> kernel-unsupported (public-filed); runtime-incompatible /
disk-short / hard-block / catch-alls -> unknown (review-queued, not filed).
- tests: extended test-{pull,pullemit-capture,loop-input,classifier,dedup}.sh
with the V1 RED-LINE proofs (emit_capture() still writes ONLY pt1-4; a
schema==1 bundle yields byte-identical FInput / ClassificationResult /
dedup_hash / effective_dedup_hash pre/post the protocol lift; the dedup
fence holds) + gate-emitter / read_gate_bundle / gate_abort_reason routing
/ shared .last marker coverage; all 22 shipped test-*.sh green in the CI
condition (gitignored .pull-captures absent), kv-calc --calibration 11/11.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
344ab87dd3 |
fix(deriver): correct stale "GGUF not supported until v0.8.1" message — now misleading post-v0.8.1-ship
deriver.py:343 and :743 told users GGUF/.bin is "not supported until
v0.8.1". v0.8.1 has shipped (the fix/docs-fidelity stack) and GGUF was
deliberately de-scoped from the v0.8.2 feature work too (cross-engine
serving = a deferred §2/§9 design-unlock, not a near-term version). The
strings actively mislead users on master ("wait for v0.8.1" — which
exists and won't add it). Re-anchored both to accurate, version-free
wording: GGUF/.bin not supported — this path is vLLM + safetensors only.
String-only; zero decision-logic change. Surfaced by the v0.8.2 brief
r1 review (Major finding). Same docs-fidelity class as the v0.8.1 stack.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
b4b20ff7b6 |
fix(patch-attribution): register vendored gemma-4-31b pr41800 overlay (follow-up to #153/#154)
PR #154 vendored models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/ to fix #153, but did not add its patch-attribution entry. test-patch-attribution flags the install.sh as an orphan artifact (lacks patches.yml entry) → rc=1. The repo has no PR-CI so the merge didn't catch it; master is red on this test. Caught by the v0.8.1 pre-tag gate (full suite in CI condition) — exactly the v0.8.0-lesson failure class that per-step verification misses. Fix: add `gemma-vllm-pr41800-truncate-prompt-tokens` to patches.yml mirroring the canonical `qwen-vllm-pr41800-truncate-prompt-tokens` entry (model=gemma-4-31b, the 6 gemma dual compose registry ids that wire it, same delivery_spec/drift_guard/ upstream block), and list it in the Gemma4ForConditionalGeneration arch required_patches for modeling consistency with the qwen arch. Engine-level overlay, no behavior change. Full scripts/tests suite + kv-calc calibration 13/13 GREEN. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |