10 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 46bb271737 docs(examples): correct "thinking on by default" — shipped composes set enable_thinking=false (#372)
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m17s
Release / release (push) Failing after 50s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
EXAMPLES.md asserted thinking is on by default (lines 22/43/228), but
every shipped Qwen3.6 compose sets
--default-chat-template-kwargs '{"enable_thinking": false}'
(bounded-thinking.yml is the only exception). Same docs-vs-shipped
class as the v0.8.0 docs-fidelity gaps. Corrected the 3 inaccurate
spots, added a canonical "thinking is OFF by default + how to enable
per-request + bounded-thinking exception" note under the max_tokens
table. Consistent with the disc #151 public answer and the
enable_thinking-default rationale; does not pre-judge the #150
froggeric re-eval. Doc-only.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 22:28:30 +00:00
noonghunna 760a2da46a Merge pull request #155 from noonghunna/fix/153-patch-attribution
fix(patch-attribution): register vendored gemma-4-31b pr41800 overlay (#153 follow-up)
2026-05-18 03:27:54 +05:00
noonghunnaandClaude Opus 4.7 b4b20ff7b6 fix(patch-attribution): register vendored gemma-4-31b pr41800 overlay (follow-up to #153/#154)
PR #154 vendored models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/
to fix #153, but did not add its patch-attribution entry. test-patch-attribution
flags the install.sh as an orphan artifact (lacks patches.yml entry) → rc=1.
The repo has no PR-CI so the merge didn't catch it; master is red on this test.

Caught by the v0.8.1 pre-tag gate (full suite in CI condition) — exactly the
v0.8.0-lesson failure class that per-step verification misses.

Fix: add `gemma-vllm-pr41800-truncate-prompt-tokens` to patches.yml mirroring
the canonical `qwen-vllm-pr41800-truncate-prompt-tokens` entry (model=gemma-4-31b,
the 6 gemma dual compose registry ids that wire it, same delivery_spec/drift_guard/
upstream block), and list it in the Gemma4ForConditionalGeneration arch
required_patches for modeling consistency with the qwen arch. Engine-level overlay,
no behavior change. Full scripts/tests suite + kv-calc calibration 13/13 GREEN.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 22:03:05 +00:00
noonghunna 6bfb8912c7 Merge pull request #154 from noonghunna/hotfix/153-gemma-pr41800-vendor
fix(gemma-4-31b): vendor missing vllm-pr41800 overlay into model tree (closes #153)
2026-05-18 02:39:41 +05:00
noonghunnaandClaude Opus 4.7 9c7919253d fix(gemma-4-31b): vendor missing vllm-pr41800 overlay into the model tree (closes #153)
5 Gemma 4 31b dual composes (int8, awq, int8-tq3, dflash, dflash-int8)
bind-mount `../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh`
but commit 1d7aad1 ("vendor ... across all pre-fix engines", #139) wired
the mounts without copying the overlay dir into the gemma-4-31b tree —
only the qwen3.6-27b tree got it. Per the per-model-tree `../../patches/`
convention every other Gemma patch mount already follows, the relative
path resolves into the gemma tree, so the mount source was missing and
`docker compose up` failed on all 5.

All 5 composes route through pre-`d5b31c95` engines (vllm-nightly-full
`e47c98ef` / vllm-nightly-dflash `e47c98ef`/`01d4d1ad`) that genuinely
need the kwarg fix, so dropping the mount is NOT correct — the fix is to
vendor the dir. install.sh is engine-level / model-agnostic, copied
byte-identical from the canonical qwen3.6-27b copy. README's compose
list re-scoped to the Gemma tree + a canonical-source pointer added so
the co-located doc isn't misleading.

Verified: all 5 composes now resolve the bind-mount path and parse via
`docker compose config`; install.sh diff-identical to canonical.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 21:31:33 +00:00
noonghunnaandClaude Opus 4.7 b0774f953b docs(hardware): newer-driver 3090 caps long-text.yml at MAX_MODEL_LEN=105000 (#149)
@sethbrasile's controlled 9-run matrix on a headless 3090 + driver
595.71.05 / CUDA 13.2 shows the same env-override pattern as the 4090
display-overhead case: the newer driver's vLLM activation-profile
reserve measured ~2.87 GiB vs ~1.5 GiB on the bare-metal reference rig,
shrinking the KV pool and capping long-text.yml at MAX_MODEL_LEN=105000
(vs 180K default). Added as the 3090 sibling anchor next to the
@laurimyllari 4090 -> 90000 data point so newer-driver 3090 users start
from the right number. Tuning-data contribution, not a bug
(corroborates the known Cliff 2a-under-v7.72.2 / genesis#22 picture).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 21:14:58 +00:00
noonghunnaandClaude Opus 4.7 820eb3845c fix(pull): argparse usage errors exit 64, not 2 — distinguishable from honest hard-stop (#370)
v0.8.0 docs-fidelity finding #1. `pull.py` defined `_EXIT_USAGE=64` and
docs/pull.sh-header promised "64 = usage", but argparse's default
`error()` hard-exits `2` — colliding with `_EXIT_ABORT` (honest gate
hard-stop). A typo and a legitimate gate-block were indistinguishable to
callers/automation (both `2`).

Fix: a contained `_UsageExit64Parser(argparse.ArgumentParser)` overriding
`error()` to exit `_EXIT_USAGE` (64). `--help` is unaffected (goes through
`exit()`, still 0). Verified: no-args / missing-required / unknown-flag
-> 64; --help -> 0; honest hard-stop -> 2 (distinct again); full v0.8.0
suite + kv-calc 22/22 zero regression. Regression-locked by a new
CLI-contract block in test-pull.sh (the pure truth-table can't cover the
argv/exit boundary). docs/PULL.md exit-code table updated to the fixed
contract, with a note that the v0.8.0 *tag* still exits 2 (this lands on
master post-v0.8.0, ships with the next release — not a separate patch
tag, per the maintainer call).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 20:58:10 +00:00
noonghunnaandClaude Opus 4.7 78a7dee247 docs: fix v0.8.0 docs-fidelity gaps (trc-ack first-run heads-up, exit-code honesty, GGUF message claim)
From the v0.8.0 docs-fidelity test (#369) — align docs with shipped CLI:

- PULL.md Quickstart + FAQ: first-run heads-up that common archs
  (Qwen2ForCausalLM &c) hard-block at needs-trust-remote-code-ack even
  with --dry-run; add --trust-remote-code (after vetting the code) to
  clear it. (Was a silent new-user wall.)
- PULL.md exit-codes: documented honestly — argparse usage/arg errors
  exit 2 (shared with honest hard-stop); 64 is reserved, arg-parser
  errors do not currently reach it (tracked CLI follow-up, #370).
- FAQ GGUF claim: "clear message" → accurate "aborts as
  unsupported-format (generic message; clearer GGUF message is a
  tracked v0.8.1 follow-up), not a crash".

Additive, leak-clean, links resolve, curated path untouched. Docs-only
(triggers no CI). Follows the (b) cross-link pass afe56f7.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 20:39:33 +00:00
noonghunnaandClaude Opus 4.7 afe56f763f docs: cross-link the v0.8.0 universal pull flow from the existing user guides
Post-release additive pass — the pre-existing high-traffic guides didn't
know `pull` exists. All edits additive, curated path untouched (same
discipline as the README migration nudge):

- FAQ.md: new Q "model not in the supported list — can I still run it?";
  GGUF Q gets a v0.8.0 note (safetensors-only eval, GGUF→llama.cpp stays
  curated/manual, cross-engine deferred); launch.sh answer points
  non-catalog models at `pull`.
- SINGLE_CARD / DUAL_CARD / MULTI_CARD: one blockquote cross-link each to
  docs/PULL.md ("not in the configs / any HF safetensors repo — both
  paths work").
- ADDING_MODELS.md: reframed catalog-onboarding vs just-run-a-model
  (`pull`); the doc is the heavier calibration-catalog promotion task,
  not a prerequisite for serving.
- GLOSSARY.md: new "Universal pull (v0.8.0)" table (pull, dry-run,
  confidence tier, boot-fit≠runtime, calibration backbone).

Leak-clean; all links resolve on master; CommonMark structure verified
(blockquotes/headings blank-line separated). Docs-only — triggers no CI
(only tags do); lands as post-v0.8.0 polish on master.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-17 17:02:19 +00:00
github-actions[bot] b720840d2e chore(changelog): regenerate for v0.8.0 [skip ci] 2026-05-17 12:32:39 +00:00
16 changed files with 409 additions and 11 deletions
+61
View File
@@ -16,6 +16,67 @@ history; SemVer takes over from `v0.3.0` onward.
---
## v0.8.0 — 2026-05-17
### ⚠️ Cliffs, gotchas, regressions
- v0.8.0 Pull-Gate P4-fix: price Tier-1 curated via curated-exact kv-calc spec, not generic-dense (+ non-mocked regression test) ([087a8ea](https://github.com/noonghunna/club-3090/commit/087a8ea6929144ca64b58d499161087111a290ab))
### 🐛 Bug fixes
- fix(verify-full): warm engine before scored checks (closes #352) ([c595496](https://github.com/noonghunna/club-3090/commit/c5954964c5ffa68a8997866296ec83aa349d9219))
### 📝 Documentation
- docs(tq3-mtp): add missing 04-gemma-vs-qwen.png chart ([da9ef5e](https://github.com/noonghunna/club-3090/commit/da9ef5eb108ed01aa436444e4708ba329e8e7662))
- docs: UPSTREAM Gemma4 TurboQuant row — exact config.py:101 mechanism + fix-PR set ([fd8695f](https://github.com/noonghunna/club-3090/commit/fd8695f1ab64082be8818bd72a72ea2d3f710cca))
- docs+hygiene: track Gemma4 native-TurboQuant upstream blocker; gitignore new MoE cache dirs ([f812715](https://github.com/noonghunna/club-3090/commit/f812715d3e8afce1f79969fb76e5ed600a337122))
- docs(KERNEL_MATRIX): add Kernel Selection Philosophy section ([287c766](https://github.com/noonghunna/club-3090/commit/287c76617d5a3432f04425e876889ed6e43b7891))
### 🧹 Other
- Merge PR #147: v0.8.0 — Universal pull (evaluate & serve any safetensors HF model) ([#147](https://github.com/noonghunna/club-3090/pull/147) by @noonghunna)
- v0.8.0 [docs] PULL.md Quickstart (command-first, top-of-doc) + ARCHITECTURE one-liner: stage names are internal, users run one command ([bef766d](https://github.com/noonghunna/club-3090/commit/bef766d4ae2c5a03958c5da3d79bfa6a92ab01dd))
- v0.8.0 [review] pre-tag fixes: scrub internal-path leaks from shipped source + make .pull-captures-corpus tests CI-safe (skip-when-absent) ([49d9bb4](https://github.com/noonghunna/club-3090/commit/49d9bb4313a97d3cacd6cafdaebc5c2903fc4bfb))
- v0.8.0 [docs] ARCHITECTURE.md: add the universal pull→gate→emit→loop pipeline to the mental model + scripts tree (current-state, was stale for v0.8.0) ([1cdda19](https://github.com/noonghunna/club-3090/commit/1cdda195a77558a5cdef166fb7d184538b773140))
- v0.8.0 [UX] §7 two doc tracks: docs/PULL.md (user front-door) + docs/README.md (track spine) + README migration nudge ([a0b3b5c](https://github.com/noonghunna/club-3090/commit/a0b3b5c8b448a028f6037efeb610d4a53da18881))
- v0.8.0 [F] F8-fix: widen §6.1 Tier-1 OOM signature + pt3.actual regexes to real vLLM v0.21.0+ KV-cache-too-large phrasing — on-rig F8 caught classic-torch-only regexes miss the common KV-prediction failure ([f92624d](https://github.com/noonghunna/club-3090/commit/f92624d9a78c18f5b92257a553abdf07e8071dcf))
- v0.8.0 [F] F7: docs/LOOP.md contributor doc (Loop phase, grounded in shipped F1–F6) + CONTRACT-5(i) risk note ([a8b30d6](https://github.com/noonghunna/club-3090/commit/a8b30d67072fe13bad7810c7ef751fc45c369c2f))
- v0.8.0 [F] F6: CONTRACT-5 mandatory content-hash kv_calc_version (G2) + G1 topo-verify + L2 fixture sync ([1ac0481](https://github.com/noonghunna/club-3090/commit/1ac048189d02ccf709942b1c066737f5d9c29e8e))
- v0.8.0 [F] F5: §6.3 canonical-tuple-hash dedup + bounded label scheme + collision-safe submit path (CONTRACT-4) ([5de7224](https://github.com/noonghunna/club-3090/commit/5de7224a731bfefb70cda038cdd8bf3d0d48915d))
- v0.8.0 [F] F4: §6.2 inbound-trust pipeline raw→candidate→validated→Tier-1 + CONTRACT-3a derived-deferral (CONTRACT-3) ([d758f08](https://github.com/noonghunna/club-3090/commit/d758f08fce86d81d80370ad037ce12e2edc54339))
- v0.8.0 [F] F3: G6-A 3-part additive [E] touch (pt1.predicted_b_breakdown, pt3.failure_log_excerpt+actual, container-log capture) + §6.1 Tier-1 (CONTRACT-2) ([b100979](https://github.com/noonghunna/club-3090/commit/b1009793f26f38370ba2a013ed93d2b6044e0b59))
- v0.8.0 [F] F2: §6.1 Tier-2 semantic-fingerprint classifier + Appendix A seed DB (CONTRACT-2 Tier-2) ([9f80d29](https://github.com/noonghunna/club-3090/commit/9f80d29fbe63cd9302221a7d182693b5f803ca82))
- v0.8.0 [F] F1: FInput capture-bundle reader + schema-1 validation + key-normalization (CONTRACT-1) ([1491cbc](https://github.com/noonghunna/club-3090/commit/1491cbc7afaa48df924e2ea27446b7dc3b782c42))
- v0.8.0 [E] E-outcome-fix: honest 3-state manifest outcome (partial-success != failed) — §6.2 partial is a capability-scoped success ([71148d6](https://github.com/noonghunna/club-3090/commit/71148d6054b4616ada57654b7c405cb0c0d50cc6))
- v0.8.0 [E] E3/E4-fix: boot lifecycle as context manager (server stays up for smoke+capture, teardown on ctx-exit) — on-rig E5 caught teardown-in-finally-before-smoke ([f7c405a](https://github.com/noonghunna/club-3090/commit/f7c405a06d4c8c2ef349533d5d9e6ee59455a073))
- v0.8.0 [E] E3-fix: smoke probes the real served-model-name (not literal "derived") + capture failure detail — on-rig E5 caught red-smoke-on-healthy-boot ([16a1e4d](https://github.com/noonghunna/club-3090/commit/16a1e4d944992682ac5c8db9b8f7482730e3b9e6))
- v0.8.0 [E] E2-fix-2: verify *.safetensors against HF API lfs.sha256 (not Xet-redirect-fragile HEAD x-linked-etag) — on-rig E5 caught false no-etag ([3ae74bf](https://github.com/noonghunna/club-3090/commit/3ae74bfdcfc9ddc4378003c04180d934282b8482))
- v0.8.0 [E] E2-fix: download via hf CLI subprocess (not huggingface_hub lib-import) — on-rig E5 caught ModuleNotFoundError ([806a298](https://github.com/noonghunna/club-3090/commit/806a2985226cd1c6ab9f726882e5a610ba2da69f))
- v0.8.0 [E] E5(docs): docs/PULL_EMIT_DERIVED.md (+ private ledger/recon-checklist updates) ([d134d5a](https://github.com/noonghunna/club-3090/commit/d134d5a5fcc2b509e7e259430954c192d18bea44))
- v0.8.0 [E] E4: post-[C1] derived-[E] orchestration + trigger semantics + override force-capture (pt5) ([2ed18aa](https://github.com/noonghunna/club-3090/commit/2ed18aad3fd1f648d4a143d91e9e86d97b785f08))
- v0.8.0 [E] E3: derived boot (HF_HOME mount) + 4 §6 capture emitters + manifest + derived smoke floor ([f327887](https://github.com/noonghunna/club-3090/commit/f327887c3926b8c119afd1d4883f64f40e1b83a6))
- v0.8.0 [E] E2: HF download stage (download_set allowlist + x-linked-etag SHA, no-etag fail-closed, atomic staging) ([7a2ec86](https://github.com/noonghunna/club-3090/commit/7a2ec8664052044677c1724cc46ea78a7cd2988e))
- v0.8.0 [E] E1: generate_from_profile + derived-vllm template + EInput + CONTRACT-5 gate ([411c84f](https://github.com/noonghunna/club-3090/commit/411c84fd8ae692acd99824754b292f387ecf9486))
- v0.8.0 Pull-Gate P5: docs/PULL_GATE.md (two-path model, 6-stratum taxonomy, §4.1 [C1], hardware-SM) ([2582438](https://github.com/noonghunna/club-3090/commit/2582438d0202b0082464cca0bbb1dbee46d2725c))
- v0.8.0 Pull-Gate P4: stratum-5 + [C1] §4.1 total fn + stratum-6 [D] dry-run + pull orchestrator + exhaustive test-pull.sh ([adf7a3b](https://github.com/noonghunna/club-3090/commit/adf7a3bf13e424e96ffb510d318faaf016ebf975))
- v0.8.0 Pull-Gate P3: stratum-2 precondition + [C0] engine-support/runtime/hardware gate + [C2a] disk ([4a1d385](https://github.com/noonghunna/club-3090/commit/4a1d3857f25de46c60ff8c0b7f566733f32a2a6f))
- v0.8.0 Pull-Gate P2: transformers deriver + ModelProfile/confidence + variant-scoped hf_repos schema ([818b79c](https://github.com/noonghunna/club-3090/commit/818b79ccb523aca5034fefe4abddebf5c697dab2))
- v0.8.0 Pull-Gate P1: kv-calc generic-dense family + eligibility predicate + raw_verdict adapter ([1bafcfe](https://github.com/noonghunna/club-3090/commit/1bafcfe4a008c8d4e89763d872a99618e3853112))
- v0.8.0: doc generated composes are not relocatable (run with --project-directory) ([a2fc05e](https://github.com/noonghunna/club-3090/commit/a2fc05e1c873ff4a6a3bc97de7df1685b41e5006))
- v0.8.0 STEP 5: COMPOSE_GENERATOR.md + PATCH_POLICY.md (#141 contributor contract) ([9546f99](https://github.com/noonghunna/club-3090/commit/9546f99303bfcd907618d7778cf36237c33b858e))
- v0.8.0 STEP 3+4: compose generator + 5-triple golden-parity test (#141) ([6d7a043](https://github.com/noonghunna/club-3090/commit/6d7a043d907aedab88cc10c8b97bea75a2c4a81d))
- v0.8.0 STEP 2: extract patch_attribution.py (sound body-only reaches(), test imports it) ([60f3983](https://github.com/noonghunna/club-3090/commit/60f39832834051a294deb445bfc62d5ca90d0486))
- v0.8.0 Phase A-prime: enrich patch/profile data for #141 generator (compose_service_template, genesis_equipped, delivery metadata, drift_guards, drafter/model_slug/trc fold-ins) ([9f23736](https://github.com/noonghunna/club-3090/commit/9f23736f014dbbea74b56a4287836019d0315fb9))
- Add v0.8 Phase A patch attribution data ([91a9622](https://github.com/noonghunna/club-3090/commit/91a9622619ba2b1360676a95e412f30d3e975d7c))
[Pin: `git checkout v0.8.0`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.7.4...v0.8.0)
## v0.7.4 — 2026-05-15
+3 -1
View File
@@ -1,6 +1,8 @@
# Adding a model to the club-3090 stack
End-to-end workflow for onboarding a new model into the v0.7.0 profile catalog + serving infrastructure. Pairs with [KV_MATH.md](KV_MATH.md) (math reference) and [ARCHITECTURE.md](ARCHITECTURE.md) (current stack state).
End-to-end workflow for onboarding a new model into the **curated profile catalog** + serving infrastructure. Pairs with [KV_MATH.md](KV_MATH.md) (math reference) and [ARCHITECTURE.md](ARCHITECTURE.md) (current stack state).
> **Just want to run a model, not add it to the catalog?** As of v0.8.0 you don't need this workflow — `scripts/pull.sh <org/Model> --profile-like vllm/minimal` evaluates *any* safetensors HF repo against the KV math and boots it if it passes (see [docs/PULL.md](PULL.md)). This page is for the heavier task of promoting a model into the **measured/calibration catalog** (real benchmarks, validated composes, calibration anchors, per-model gotchas) — the high-confidence backbone, not a prerequisite for serving.
## When to add a new model vs a new quant of an existing one
+2
View File
@@ -2,6 +2,8 @@
You have **2× RTX 3090s**. This page is the front door for picking a config and knowing what dual-card unlocks vs single. Model-specific deep dives (quants, Genesis, engine internals) live in the model directory — links at the bottom.
> **Model not in the configs below / want any HF safetensors repo?** → [`docs/PULL.md`](PULL.md): `scripts/pull.sh` evaluates any model against the KV math (honest, no download) and boots it if it passes. The curated configs on this page are the measured path; both work.
**NVLink auto-detection** (since 2026-05-14): the dual-card composes now auto-detect whether an NVLink bridge is present. If you have one, you get the NVLink-optimized path automatically. If not, PCIe mode is used. Override with `NVLINK_MODE=force_on|force_off` in your `.env`. See the "NVLink auto-detection" section below.
> **Have 3+ GPUs?** See [`MULTI_CARD.md`](MULTI_CARD.md) — derivation of TP=4 / TP=8 configs from `dual.yml`, valid TP values for Qwen3.6-27B (1, 2, 4, 5, 8, 10), and what scales vs what doesn't.
+6 -4
View File
@@ -19,13 +19,15 @@ Qwen3.6-27B is a thinking model. The `<think>...</think>` block before the answe
| Scenario | `max_tokens` |
|---|---|
| **FREE thinking on (default long-text / long-vision composes)** | **8192** minimum. 16384 for hard reasoning / competition-grade problems. |
| **FREE thinking (enabled per-request — see note below)** | **8192** minimum. 16384 for hard reasoning / competition-grade problems. |
| **FSM bounded thinking (`bounded-thinking.yml`)** | **4096** is comfortable. The recommended DeepSeek scratchpad grammar uses ~500-1000 think tokens; the andthattoo G/A/E grammar uses ~150. Either fits well below 4096. |
| **`enable_thinking: False`** | Set as tight as the answer needs (50-200 typically). |
| **Tool-using agents (multi-turn)** | 1024-2048 per turn. If a middle turn needs >2K to think, your prompt structure probably needs work. |
The smoke-test examples below use `max_tokens: 200` because they ask short questions where thinking + answer fits comfortably. Real workloads should follow the table above.
> **Thinking is OFF by default on the shipped composes.** Every Qwen3.6 compose sets `--default-chat-template-kwargs '{"enable_thinking": false}'`, so the model answers directly with no `<think>` block unless you opt in. Enable it per-request with `chat_template_kwargs: {"enable_thinking": true}` (no restart) and budget `max_tokens` per the table. The one exception is `bounded-thinking.yml`, which keeps thinking on but bounds its cost via a structured-CoT grammar (see [`docs/STRUCTURED_COT.md`](STRUCTURED_COT.md)).
---
## Quick curl sanity test
@@ -40,7 +42,7 @@ curl -sf http://localhost:8020/v1/chat/completions \
}' | jq -r '.choices[0].message.content'
```
Expected response: a sentence containing `Paris`. The `max_tokens: 200` headroom is intentional — Qwen3.6 thinks before answering by default, so even simple questions burn ~50–150 tokens inside `<think>...</think>` before reaching the answer. Set tighter (`max_tokens: 30`) only if you also pass `chat_template_kwargs: {"enable_thinking": false}` to skip the think block — that's what `verify-full.sh` does internally.
Expected response: a sentence containing `Paris`. The shipped composes set `enable_thinking: false` by default, so the model answers directly with no `<think>` block — `max_tokens: 200` is comfortable slack. If you enable thinking per-request (`chat_template_kwargs: {"enable_thinking": true}`), raise `max_tokens` substantially (see the table above) — the model then emits a `<think>...</think>` block first even for simple questions. `verify-full.sh` passes `enable_thinking: false` explicitly.
---
@@ -225,8 +227,8 @@ const client = new OpenAI({
const resp = await client.chat.completions.create({
model: "qwen3.6-27b-autoround",
messages: [{ role: "user", content: "Quicksort in Rust, please." }],
// FREE thinking is on by default. 4096 covers easy code-gen think+answer;
// 8192 is the safe default for harder coding problems. 800 traps mid-think.
// Shipped composes default enable_thinking:false → this answers with no <think> block; 4096 is generous.
// To get reasoning, add extra_body chat_template_kwargs {"enable_thinking": true} and budget 8192+ (800 traps mid-think).
max_tokens: 4096,
temperature: 0.6,
top_p: 0.95,
+6 -2
View File
@@ -64,9 +64,13 @@ Use LM Studio if you prefer a GUI and don't need the engineering. Use this repo
We tried EAGLE — it's blocked on Qwen3-Next (the family Qwen3.5/3.6 belong to) by DeltaNet hybrid attention's lack of KV rollback support in vLLM/SGLang. MTP works because it's a different protocol (multi-token prediction at draft-head level, not a separate draft model). See [INTERNALS.md "Speculative decoding"](../models/qwen3.6-27b/INTERNALS.md) for the full forensic chain. **Re-test triggers:** if vllm#39931 lands or DeltaNet rollback support arrives upstream, EAGLE becomes viable again.
### The model I want isn't in the supported list — can I still run it?
Yes, if it's a **safetensors** repo. As of v0.8.0, `scripts/pull.sh <org/Model> --profile-like vllm/minimal --dry-run` evaluates *any* safetensors HF repo against this stack's KV math — no download — and tells you honestly whether it fits and at what confidence. Drop `--dry-run` (add `--yes`) and, if it passes the gates, it downloads, generates a minimal compose, and boots it. Non-fits stop with a precise reason, not a crash. Full guide: [docs/PULL.md](PULL.md). One heads-up: many common archs (e.g. `Qwen2ForCausalLM`) stop at `needs-trust-remote-code-ack` on the first try even with `--dry-run` — add `--trust-remote-code` (after checking what code the repo runs) to clear it. Limits: safetensors + vLLM only; GGUF / `.bin` repos abort at derive as `unsupported-format` (not a crash) — see next Q.
### Why not GGUF on vLLM for this model?
Multiple gates blocked. Qwen3.6-27B GGUF on vLLM hits a chain of "fixed but-not-quite" issues — multimodal config routing, ParallelLMHead skip, the `Qwen35TensorProcessor._reverse_reorder_v_heads` weight loader producing garbage output on the 27B layout (transformers PR #45283 only validated on 0.8B). Tracked in [INTERNALS.md](../models/qwen3.6-27b/INTERNALS.md#qwen36-27b-gguf-on-vllm). Use llama.cpp for GGUF on this model.
Multiple gates blocked. Qwen3.6-27B GGUF on vLLM hits a chain of "fixed but-not-quite" issues — multimodal config routing, ParallelLMHead skip, the `Qwen35TensorProcessor._reverse_reorder_v_heads` weight loader producing garbage output on the 27B layout (transformers PR #45283 only validated on 0.8B). Tracked in [INTERNALS.md](../models/qwen3.6-27b/INTERNALS.md#qwen36-27b-gguf-on-vllm). Use llama.cpp for GGUF on this model. **Note (v0.8.0):** `pull` evaluates *safetensors* repos only — GGUF→llama.cpp is **not** served via `pull` (it stays the curated/manual path; cross-engine generation is deliberately deferred). A GGUF/`.bin` repo aborts cleanly at the deriver stage as `unsupported-format` (the message is generic — it does not yet say "GGUF, use llama.cpp"; a clearer message is a tracked v0.8.1 follow-up), not a crash.
### Why AutoRound INT4 not GPTQ / AWQ?
@@ -134,7 +138,7 @@ If your numbers on the same compose look different from ours by >15%, the most l
For a first install, run `bash scripts/setup.sh` with no model argument in a normal terminal. It opens a hardware-aware model picker, marks Qwen / Gemma / Both as eligible or not for your detected GPUs, then continues into the existing download flow.
After setup, run `bash scripts/launch.sh`. The wizard asks which model (filtered to what you've downloaded), then which GPU(s) to use, auto-picks TP for homogeneous sets (PP for heterogeneous), filters variants by hardware fit, shows a per-card VRAM projection from `tools/kv-calc.py` for the suggested default, then boots and runs `verify-full.sh`. Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, partial flags like `bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1` (skips prompts), `--tp 4 --pp 2` to override parallelism, plus `setup.sh --help` / `launch.sh --help` for the full flag list.
After setup, run `bash scripts/launch.sh`. The wizard asks which model (filtered to what you've downloaded), then which GPU(s) to use, auto-picks TP for homogeneous sets (PP for heterogeneous), filters variants by hardware fit, shows a per-card VRAM projection from `tools/kv-calc.py` for the suggested default, then boots and runs `verify-full.sh`. Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, partial flags like `bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1` (skips prompts), `--tp 4 --pp 2` to override parallelism, plus `setup.sh --help` / `launch.sh --help` for the full flag list. This wizard covers the **curated catalog**; for a model *not* in the catalog (any safetensors HF repo), use `scripts/pull.sh` instead — see [docs/PULL.md](PULL.md).
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
+10
View File
@@ -6,6 +6,16 @@ Plain-language definitions for terms used throughout the docs. Roughly grouped b
---
## Universal `pull` (v0.8.0)
| Term | What it means |
|---|---|
| **`pull`** | `scripts/pull.sh <hf-repo> --profile-like <key>` — evaluates *any* safetensors HF repo against this stack's KV math and, if it passes the gates, downloads + generates a minimal compose + boots it. The model-agnostic front door; the curated catalog still works unchanged. See [PULL.md](PULL.md). |
| **dry-run** | `pull … --dry-run` — evaluate only: never downloads, never boots, just prints the verdict. |
| **confidence tier** | How much the fit verdict is trusted: `exact` (a measured/curated profile) vs `estimated-lower-bound` (derived from the repo's own config — a floor, likely under-modeled). Always shown with the verdict. |
| **boot-fit ≠ runtime-stability** | A "fits" verdict is a *boot-time* allocation check. It's necessary-not-sufficient: a config that boots clean can still degrade/OOM under sustained accumulated-context agent workloads (see [CLIFFS.md](CLIFFS.md)). Validate with soak-continuous before relying on it. |
| **calibration backbone** | The curated catalog's role under v0.8.0 — the measured anchor set the KV math is calibrated against (vs. being the only supported models). |
## Throughput / latency
| Term | What it means |
+1 -1
View File
@@ -53,7 +53,7 @@ Use `--force` only when you are intentionally testing an unsupported combo. Exam
On 20 GB cards (modded 3080) the cudagraph-profiling overhead is a meaningful slice of available VRAM. Drop `--gpu-memory-utilization` to **0.82** (vs shipped 0.95 for 24 GB). vLLM nightly's `gpu_worker.py` reports the equivalent effective KV size in the boot log; tune to keep activation headroom for the ~15K tool-prefill peak (verify-full check 8). Credit: [@troymroberts](https://github.com/troymroberts).
**4090s with attached display — env-override the compose defaults.** Some 4090 rigs land at ~23.5 GB usable VRAM with X server + driver overhead, vs the headless 3090s the composes are calibrated for. Boot may fail with `No available memory for the cache blocks` at default `max-model-len`. Cross-rig data: @laurimyllari's 4090 single-card on `long-text.yml` needed `MAX_MODEL_LEN=90000` (down from 180K default) to fit cleanly ([disc #62](../../../noonghunna/club-3090/discussions/62) / [issue #71](../../../noonghunna/club-3090/issues/71)). Pattern:
**4090s with attached display — env-override the compose defaults.** Some 4090 rigs land at ~23.5 GB usable VRAM with X server + driver overhead, vs the headless 3090s the composes are calibrated for. Boot may fail with `No available memory for the cache blocks` at default `max-model-len`. Cross-rig data: @laurimyllari's 4090 single-card on `long-text.yml` needed `MAX_MODEL_LEN=90000` (down from 180K default) to fit cleanly ([disc #62](../../../noonghunna/club-3090/discussions/62) / [issue #71](../../../noonghunna/club-3090/issues/71)). **Newer driver shrinks the budget the same way even on a headless 3090:** @sethbrasile's controlled 9-run matrix on a headless 3090 with driver 595.71.05 / CUDA 13.2 capped `long-text.yml` at `MAX_MODEL_LEN=105000` — the newer driver's activation-profile reserve measured ~2.87 GiB vs ~1.5 GiB on the bare-metal reference rig, shrinking the KV pool by the difference ([issue #149](../../../noonghunna/club-3090/issues/149)). On a newer-driver 3090, start at `MAX_MODEL_LEN=105000` rather than the 180K default. Pattern:
```bash
MAX_MODEL_LEN=90000 bash scripts/switch.sh vllm/long-text
+2
View File
@@ -7,6 +7,8 @@ This page explains what scales (and what doesn't) when going beyond TP=2,
the constraints to know, and how to derive your own compose when `multi4.yml`
isn't your topology.
> **Model not in the configs here / want any HF safetensors repo?** → [`docs/PULL.md`](PULL.md): `scripts/pull.sh` evaluates any model against the KV math (honest, no download) and boots it if it passes. The recipes on this page are the measured/derivation path; both work.
> **Validation note:** the maintainer rig is **2× RTX 3090 PCIe**, but
> Whamp's 4× RTX 3090 PCIe rig validated the TP=4 fp8/MTP baseline in
> [discussion #26](https://github.com/noonghunna/club-3090/discussions/26)
+6 -2
View File
@@ -37,6 +37,8 @@ What you'll see — exactly one of:
| `hard-block` | `2` | Honest stop with a precise reason (unsupported engine/arch, won't-fit, disk, needs `--trust-remote-code`). Nothing downloaded. |
| `override-accepted` | `0` | You explicitly accepted a non-pass path (e.g. `--force-download`); proceeds with the caveat recorded. |
> **First-run heads-up:** many common models (anything `Qwen2ForCausalLM` — Qwen2.5 & a large family, plus other custom-code archs) hard-block at `[C0] needs-trust-remote-code-ack` on the *very first* try — **even with `--dry-run`**. That's the gate working, not a failure. After you've checked what code the repo would run, add **`--trust-remote-code`** to that same command to clear it. See [`--trust-remote-code` — a security decision](#--trust-remote-code--a-security-decision) below.
It is **honest about confidence and never silently passes.** A "fits"
verdict is a *boot-time* check — read [Boot-fit ≠ runtime-stability](#boot-fit--runtime-stability--read-this)
before relying on it for sustained agent workloads. Full detail below.
@@ -116,8 +118,10 @@ scripts/pull.sh some-org/Some-Llama-7B --profile-like vllm/minimal --dry-run
|---|---|
| `0` | Download-eligible / clean verdict. |
| `3` | Needs a flag — a `confirm→proceed` or advisory terminal that is not yet satisfied (re-run with the named flag). |
| `2` | Honest hard-stop (a gate aborted, or `hard-block`). |
| `64` | Usage error. |
| `2` | Honest hard-stop — a gate aborted, or a `hard-block` terminal. |
| `64` | Usage error — missing/unknown argument (distinct from `2`, so a typo is distinguishable from an honest gate-block). |
> *Note: the `64` usage-vs-`2` hard-stop split is a post-`v0.8.0` fix — present on `master`/the next release; the `v0.8.0` release tag still exits `2` for argument errors.*
---
+2
View File
@@ -2,6 +2,8 @@
You have **one RTX 3090 (24 GB VRAM)**. This page is the front door for picking a config and knowing what to expect. The model-specific deep dives (quants, Genesis patches, engine internals) live elsewhere — links at the bottom.
> **Model not in the configs below / want any HF safetensors repo?** → [`docs/PULL.md`](PULL.md): `scripts/pull.sh` evaluates any model against the KV math (honest, no download) and boots it if it passes. The curated configs on this page are the measured path; both work.
---
## ⚠️ Critical — read first if you're running an agentic coding client
@@ -0,0 +1,103 @@
# vLLM PR #41800 overlay — `truncate_prompt_tokens` kwarg on `get_max_tokens`
## What this fixes
Agentic clients (opencode, codex-cli, and similar IDE/agent runtimes) send `truncate_prompt_tokens` on chat-completion requests. Pre-[vLLM PR #41800](https://github.com/vllm-project/vllm/pull/41800), `vllm.entrypoints.utils.get_max_tokens()` doesn't accept that kwarg — and the kwarg propagates from the request handler down into the function call — so requests fail with:
```
HTTP 400: {"error":{"message":"get_max_tokens() got an unexpected keyword argument 'truncate_prompt_tokens'",...}}
```
The fix is upstream PR #41800 (merged 2026-05-06 at commit `d5b31c95`). It adds the kwarg to the function signature and a small body block that clamps `input_length` to `min(input_length, truncate_prompt_tokens or max_model_len)` before the existing length check.
## When this overlay is needed
This overlay is needed on engines pinned to vLLM SHAs that **predate `d5b31c95`**:
| Engine | Pinned SHA | Pre-fix? |
|---|---|---|
| `vllm-nightly-mtp` | `01d4d1ad` (2026-05-04) | ✅ needs overlay |
| `vllm-nightly-dflash` | `e47c98ef` (~2026-05-05) | ✅ needs overlay (20 commits behind d5b31c95) |
| `vllm-nightly-full` | `e47c98ef` | ✅ needs overlay |
| `vllm-nightly-clean` | `bf610c2f` (2026-05-15) | ❌ already includes fix |
If a compose routes through `vllm-nightly-clean`, the overlay is unnecessary — the function signature already accepts the kwarg upstream.
## How the overlay works
`install.sh` is a Python anchor-based in-place patcher. It does two surgical edits to the in-container `/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/utils.py`:
1. **Signature**: adds `truncate_prompt_tokens: int | None = None,` to `get_max_tokens`'s signature, anchored to the existing `override_max_tokens: int | None = None,` line.
2. **Body**: inserts a 6-line truncation-aware `input_length` adjustment block before the existing `if max_model_len < input_length:` check, anchored to that line.
Each insertion carries a sentinel comment (`# PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)`) so re-running the install on an already-patched file is a no-op. Post-patch the file is AST-validated before write.
Why anchor-based and not full-file replacement: the PR diff is +14 / -0 across a 200-line file — replacing the full file would shadow other upstream changes in `utils.py`. Anchor-based insertion is drift-resistant to unrelated upstream movement.
## Canonical source
This is a vendored, byte-identical mirror of the canonical overlay at
`models/qwen3.6-27b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/`.
`install.sh` is model-agnostic (it patches engine-level `vllm/entrypoints/utils.py`),
so it is copied verbatim per the per-model-tree `../../patches/` convention.
Keep `install.sh` identical to the qwen3.6-27b copy; track upstream/drop state in `docs/UPSTREAM.md` only.
## Composes that wire this overlay in (Gemma 4 31b tree)
* `models/gemma-4-31b/vllm/compose/dual/int8.yml`
* `models/gemma-4-31b/vllm/compose/dual/awq.yml`
* `models/gemma-4-31b/vllm/compose/dual/int8-tq3.yml`
* `models/gemma-4-31b/vllm/compose/dual/dflash.yml`
* `models/gemma-4-31b/vllm/compose/dual/dflash-int8.yml`
## How to add this overlay to another affected compose
In any compose that routes through `vllm-nightly-mtp` / `vllm-nightly-dflash` / `vllm-nightly-full`, add:
1. **Volume mount** in the `volumes:` block:
```yaml
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
```
2. **Install line** in the `entrypoint:` bash script, before `exec vllm serve`:
```bash
bash /etc/club3090/install-pr41800.sh
```
Run `bash install.sh` (the file in this directory) standalone to test against a transient vLLM container before wiring into a compose. See the smoke test in the next section.
## Smoke test
```bash
docker run --rm --entrypoint /bin/bash \
-v $(pwd)/install.sh:/install.sh:ro \
vllm/vllm-openai:nightly-01d4d1ad375dc5854779c593eee093bcebb0cada \
-c '
python3 -c "from vllm.entrypoints.utils import get_max_tokens; import inspect; print(inspect.signature(get_max_tokens))"
bash /install.sh
python3 -c "from vllm.entrypoints.utils import get_max_tokens; import inspect; print(inspect.signature(get_max_tokens))"
'
```
Expected: signature lacks `truncate_prompt_tokens` BEFORE install, has it AFTER. Verified on `01d4d1ad` (2026-05-15).
## When to drop this overlay
When **both** are true:
1. PR #41800 has merged upstream (it has — 2026-05-06 at `d5b31c95`)
2. The engine's pinned nightly SHA bumps past `d5b31c95`
For the Genesis-anchored engines, the bump happens with Sander's next Genesis release cycle (v7.73.x). For `vllm-nightly-dflash` and `vllm-nightly-full`, the bump happens when their respective overlays (PR #41703 DFlash, PR #42102 INT8 PTH KV) are re-validated against a newer nightly.
Track in `docs/UPSTREAM.md`.
## Source
- vLLM PR #41800: https://github.com/vllm-project/vllm/pull/41800
- Merged commit: `d5b31c95`
- Tracking issue: noonghunna/club-3090#139
- Triggered by: noonghunna/club-3090#138 (SEVENID's opencode boot failure)
- Patch summary: +7 lines in `vllm/entrypoints/utils.py` (the actual fix) + 5 call-site forward-compat additions in other files (we skip those — the signature fix alone unblocks all known TypeError reports)
@@ -0,0 +1,141 @@
#!/usr/bin/env bash
# Install vLLM PR #41800 — `truncate_prompt_tokens` kwarg on get_max_tokens.
#
# WHY THIS OVERLAY EXISTS:
# opencode (and other agentic clients like codex-cli) send `truncate_prompt_tokens`
# on chat-completion requests. Pre-#41800, vLLM's `get_max_tokens()` doesn't
# accept that kwarg — and somewhere upstream of the function the kwarg gets
# unpacked into the call — so requests fail with:
# HTTP 400: get_max_tokens() got an unexpected keyword argument 'truncate_prompt_tokens'
#
# PR: https://github.com/vllm-project/vllm/pull/41800
# Merged: 2026-05-06 at commit d5b31c95
# Affected pins on master:
# - vllm-nightly-mtp (01d4d1ad, 2026-05-04) — pre-fix
# - vllm-nightly-dflash (e47c98ef) — pre-fix
# - vllm-nightly-full (e47c98ef) — pre-fix
# (vllm-nightly-clean at bf610c2f is POST-fix; doesn't need the overlay)
#
# Tracking issue: #139 (noonghunna/club-3090)
# Triggered by: #138 — SEVENID's opencode boot failure on dual-dflash-noviz.
#
# WHY A PYTHON ANCHOR-BASED PATCHER:
# The PR is +7 lines in `vllm/entrypoints/utils.py` (the actual fix) plus a
# handful of forward-compat call-site additions in 5 other files. The
# function-signature change in utils.py is the ONLY thing required to fix
# the TypeError — once `get_max_tokens` accepts the kwarg, requests stop
# crashing. The call-site changes are nice-to-have semantic completeness
# (actually applying the truncation), so we patch those too via anchors.
#
# Idempotent: each anchor checks for a sentinel marker before inserting.
set -euo pipefail
# Container's vLLM install path. Override via env if vLLM moves.
SITE_PACKAGES="${CLUB3090_PR41800_SITE_PACKAGES:-/usr/local/lib/python3.12/dist-packages}"
UTILS_PY="$SITE_PACKAGES/vllm/entrypoints/utils.py"
if [ ! -f "$UTILS_PY" ]; then
echo "[club3090/pr41800] ERROR: $UTILS_PY not found; aborting overlay install" >&2
exit 1
fi
python3 - <<'PY'
import os
import re
import sys
site_packages = os.environ.get(
"CLUB3090_PR41800_SITE_PACKAGES",
"/usr/local/lib/python3.12/dist-packages",
)
utils_py = f"{site_packages}/vllm/entrypoints/utils.py"
SENTINEL = "# PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)"
# Function signature update: add `truncate_prompt_tokens: int | None = None,`
# as a kwarg on `get_max_tokens`. Anchor on the existing line that closes
# the signature (`override_max_tokens: int | None = None,` line right before `) -> int:`).
SIGNATURE_ANCHOR_RE = re.compile(
r'^(?P<indent>[ \t]+)override_max_tokens: int \| None = None,\n(?P<close>[ \t]*\) -> int:)',
re.MULTILINE,
)
SIGNATURE_INSERT = ''' override_max_tokens: int | None = None,
truncate_prompt_tokens: int | None = None, # PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)
) -> int:'''
# Body update: insert truncation-aware input_length adjustment BEFORE the
# `if max_model_len < input_length:` check. Anchor on that line.
BODY_ANCHOR_RE = re.compile(
r'^(?P<indent>[ \t]+)if max_model_len < input_length:',
re.MULTILINE,
)
BODY_INSERT = ''' # PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)
if truncate_prompt_tokens is not None:
limit = truncate_prompt_tokens
input_length = min(
input_length,
max_model_len if limit == -1 else limit,
)
if max_model_len < input_length:'''
with open(utils_py, "r", encoding="utf-8") as f:
src = f.read()
if SENTINEL in src:
print(f"[club3090/pr41800] {utils_py}: sentinel present, patch already applied; no-op", file=sys.stderr)
sys.exit(0)
# Upstream-fix detection: if the function signature already accepts the kwarg
# (i.e. the engine pinned a post-#41800 nightly), the overlay is unnecessary
# and should no-op gracefully so composes that mount it on a post-fix image
# (e.g. via vllm-nightly-clean) still boot cleanly.
UPSTREAM_RE = re.compile(
r'def get_max_tokens\([^)]*truncate_prompt_tokens\b',
re.DOTALL,
)
if UPSTREAM_RE.search(src):
print(f"[club3090/pr41800] {utils_py}: upstream get_max_tokens() already accepts truncate_prompt_tokens; no-op", file=sys.stderr)
sys.exit(0)
# Apply signature patch first (so the function accepts the kwarg)
m = SIGNATURE_ANCHOR_RE.search(src)
if not m:
print(
f"[club3090/pr41800] ERROR: signature anchor "
f"'override_max_tokens: int | None = None, ... ) -> int:' not found in {utils_py}. "
f"vLLM nightly may have changed entrypoints/utils.py — overlay needs re-anchoring.",
file=sys.stderr,
)
sys.exit(1)
src = SIGNATURE_ANCHOR_RE.sub(SIGNATURE_INSERT, src, count=1)
# Apply body patch
m = BODY_ANCHOR_RE.search(src)
if not m:
print(
f"[club3090/pr41800] ERROR: body anchor 'if max_model_len < input_length:' not found in {utils_py} "
f"after signature patch. vLLM nightly diverged unexpectedly — overlay needs re-anchoring.",
file=sys.stderr,
)
sys.exit(1)
src = BODY_ANCHOR_RE.sub(BODY_INSERT, src, count=1)
with open(utils_py, "w", encoding="utf-8") as f:
f.write(src)
# Quick validity check
import ast
try:
ast.parse(src)
except SyntaxError as e:
print(f"[club3090/pr41800] ERROR: post-patch utils.py is not valid Python: {e}", file=sys.stderr)
sys.exit(1)
print(f"[club3090/pr41800] {utils_py}: signature + body patches applied (truncate_prompt_tokens kwarg)", file=sys.stderr)
PY
echo "[club3090/pr41800] install complete" >&2
+1
View File
@@ -71,6 +71,7 @@ arches:
- gemma-vllm-gemma4-tool-parser-fixes
- gemma-vllm-gemma4-dflash
- gemma-vllm-gemma4-dflash-int8
- gemma-vllm-pr41800-truncate-prompt-tokens
valid_tp:
tp_divisors: [1, 2, 4, 8, 16]
marlin_alignment_required: false
+38
View File
@@ -279,6 +279,44 @@ patches:
drop_when: "each affected engine profile pins a vLLM nightly after d5b31c95"
status: verified
- id: gemma-vllm-pr41800-truncate-prompt-tokens
model: [gemma-4-31b]
files:
- models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens
load_bearing_when:
- composes:
- vllm/gemma-int8
- vllm/gemma-int8-262k
- vllm/gemma-awq
- vllm/gemma-int8-tq3
- vllm/gemma-dflash
- vllm/gemma-dflash-int8
reason: "Pre-d5b31c95 engine pins reject agent clients that send truncate_prompt_tokens."
evidence: "docs/UPSTREAM.md#41800 row; models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/README.md"
delivery: # DEPRECATED/READ-ONLY (test-only) — see header
dockerfile_bake: false
entrypoint_invoke: true
genesis: false
delivery_gaps: []
delivery_mechanism: install_script
delivery_spec:
script: models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/install.sh
mounted_at: /etc/club3090/install-pr41800.sh
invoke: "bash /etc/club3090/install-pr41800.sh"
invoked_before: vllm-serve
wired_at: [volumes, entrypoint]
drift_guard:
kind: behavioral
check: "agent client sending truncate_prompt_tokens does not get HTTP 400 on the selected nightly (no-op if nightly post-d5b31c95)"
on_fail: capability-degraded
capability: truncate-prompt-tokens-kwarg
foundational: false
upstream:
ref: vllm-project/vllm#41800
status: merged
drop_when: "each affected engine profile pins a vLLM nightly after d5b31c95"
status: verified
- id: gemma-vllm-gemma4-dflash
model: [gemma-4-31b]
files:
+13 -1
View File
@@ -1316,7 +1316,19 @@ _EXIT_USAGE = 64
def main(argv: list[str]) -> int:
import argparse
ap = argparse.ArgumentParser(
class _UsageExit64Parser(argparse.ArgumentParser):
"""argparse's default `error()` hard-exits 2 — which collides with
`_EXIT_ABORT` (honest gate hard-stop), so a typo and a legitimate
block are indistinguishable to callers/automation. Override to exit
`_EXIT_USAGE` (64) on argument/usage errors, restoring the
documented contract. `--help` is unaffected (it goes through
`exit()`, not `error()`, and still returns 0)."""
def error(self, message):
self.print_usage(sys.stderr)
self.exit(_EXIT_USAGE, f"{self.prog}: error: {message}\n")
ap = _UsageExit64Parser(
prog="pull.sh",
description="v0.8.0 Pull-Gate — derive an HF repo, gate it through "
"the locked 6-stratum taxonomy, and (Path A, curated+emittable) "
+14
View File
@@ -1335,4 +1335,18 @@ print("\nSUMMARY: all Pull-Gate P4 truth-table assertions passed "
"CONTRACT-5 reject).")
PY
# ---------------------------------------------------------------------------
# CLI-contract: exit-code boundary (the pure truth-table above can't cover
# argv parsing / process exit). #370 regression lock: argparse usage errors
# MUST exit 64 (not 2 — argparse default), so a typo is distinguishable
# from an honest gate hard-stop (2); --help stays 0; hard-stop stays 2.
_clifail=0
_ec(){ bash scripts/pull.sh "$@" >/dev/null 2>&1; echo $?; }
[ "$(_ec)" = 64 ] || { echo "FAIL: no-args -> 64 (#370)" >&2; _clifail=1; }
[ "$(_ec Qwen/Qwen2.5-0.5B-Instruct)" = 64 ] || { echo "FAIL: missing required --profile-like -> 64 (#370)" >&2; _clifail=1; }
[ "$(_ec --nope x)" = 64 ] || { echo "FAIL: unknown flag -> 64 (#370)" >&2; _clifail=1; }
[ "$(_ec --help)" = 0 ] || { echo "FAIL: --help -> 0 (#370 must not regress help)" >&2; _clifail=1; }
[ "$(_ec definitely/nonexistent-xyz123 --profile-like vllm/minimal --dry-run)" = 2 ] || { echo "FAIL: honest hard-stop -> 2 (must stay 2, not 64) (#370)" >&2; _clifail=1; }
[ "$_clifail" = 0 ] && echo "PASS: CLI exit-code contract (#370): usage=64, --help=0, hard-stop=2" || { echo "1+ CLI-contract assertion(s) failed." >&2; exit 1; }
echo "test-pull.sh OK"