Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
46bb271737 | ||
|
|
760a2da46a | ||
|
|
b4b20ff7b6 | ||
|
|
6bfb8912c7 | ||
|
|
9c7919253d | ||
|
|
b0774f953b | ||
|
|
820eb3845c | ||
|
|
78a7dee247 | ||
|
|
afe56f763f | ||
|
|
b720840d2e |
@@ -16,6 +16,67 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.8.0 — 2026-05-17
|
||||
|
||||
|
||||
### ⚠️ Cliffs, gotchas, regressions
|
||||
|
||||
- v0.8.0 Pull-Gate P4-fix: price Tier-1 curated via curated-exact kv-calc spec, not generic-dense (+ non-mocked regression test) ([087a8ea](https://github.com/noonghunna/club-3090/commit/087a8ea6929144ca64b58d499161087111a290ab))
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(verify-full): warm engine before scored checks (closes #352) ([c595496](https://github.com/noonghunna/club-3090/commit/c5954964c5ffa68a8997866296ec83aa349d9219))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs(tq3-mtp): add missing 04-gemma-vs-qwen.png chart ([da9ef5e](https://github.com/noonghunna/club-3090/commit/da9ef5eb108ed01aa436444e4708ba329e8e7662))
|
||||
- docs: UPSTREAM Gemma4 TurboQuant row — exact config.py:101 mechanism + fix-PR set ([fd8695f](https://github.com/noonghunna/club-3090/commit/fd8695f1ab64082be8818bd72a72ea2d3f710cca))
|
||||
- docs+hygiene: track Gemma4 native-TurboQuant upstream blocker; gitignore new MoE cache dirs ([f812715](https://github.com/noonghunna/club-3090/commit/f812715d3e8afce1f79969fb76e5ed600a337122))
|
||||
- docs(KERNEL_MATRIX): add Kernel Selection Philosophy section ([287c766](https://github.com/noonghunna/club-3090/commit/287c76617d5a3432f04425e876889ed6e43b7891))
|
||||
|
||||
|
||||
### 🧹 Other
|
||||
|
||||
- Merge PR #147: v0.8.0 — Universal pull (evaluate & serve any safetensors HF model) ([#147](https://github.com/noonghunna/club-3090/pull/147) by @noonghunna)
|
||||
- v0.8.0 [docs] PULL.md Quickstart (command-first, top-of-doc) + ARCHITECTURE one-liner: stage names are internal, users run one command ([bef766d](https://github.com/noonghunna/club-3090/commit/bef766d4ae2c5a03958c5da3d79bfa6a92ab01dd))
|
||||
- v0.8.0 [review] pre-tag fixes: scrub internal-path leaks from shipped source + make .pull-captures-corpus tests CI-safe (skip-when-absent) ([49d9bb4](https://github.com/noonghunna/club-3090/commit/49d9bb4313a97d3cacd6cafdaebc5c2903fc4bfb))
|
||||
- v0.8.0 [docs] ARCHITECTURE.md: add the universal pull→gate→emit→loop pipeline to the mental model + scripts tree (current-state, was stale for v0.8.0) ([1cdda19](https://github.com/noonghunna/club-3090/commit/1cdda195a77558a5cdef166fb7d184538b773140))
|
||||
- v0.8.0 [UX] §7 two doc tracks: docs/PULL.md (user front-door) + docs/README.md (track spine) + README migration nudge ([a0b3b5c](https://github.com/noonghunna/club-3090/commit/a0b3b5c8b448a028f6037efeb610d4a53da18881))
|
||||
- v0.8.0 [F] F8-fix: widen §6.1 Tier-1 OOM signature + pt3.actual regexes to real vLLM v0.21.0+ KV-cache-too-large phrasing — on-rig F8 caught classic-torch-only regexes miss the common KV-prediction failure ([f92624d](https://github.com/noonghunna/club-3090/commit/f92624d9a78c18f5b92257a553abdf07e8071dcf))
|
||||
- v0.8.0 [F] F7: docs/LOOP.md contributor doc (Loop phase, grounded in shipped F1–F6) + CONTRACT-5(i) risk note ([a8b30d6](https://github.com/noonghunna/club-3090/commit/a8b30d67072fe13bad7810c7ef751fc45c369c2f))
|
||||
- v0.8.0 [F] F6: CONTRACT-5 mandatory content-hash kv_calc_version (G2) + G1 topo-verify + L2 fixture sync ([1ac0481](https://github.com/noonghunna/club-3090/commit/1ac048189d02ccf709942b1c066737f5d9c29e8e))
|
||||
- v0.8.0 [F] F5: §6.3 canonical-tuple-hash dedup + bounded label scheme + collision-safe submit path (CONTRACT-4) ([5de7224](https://github.com/noonghunna/club-3090/commit/5de7224a731bfefb70cda038cdd8bf3d0d48915d))
|
||||
- v0.8.0 [F] F4: §6.2 inbound-trust pipeline raw→candidate→validated→Tier-1 + CONTRACT-3a derived-deferral (CONTRACT-3) ([d758f08](https://github.com/noonghunna/club-3090/commit/d758f08fce86d81d80370ad037ce12e2edc54339))
|
||||
- v0.8.0 [F] F3: G6-A 3-part additive [E] touch (pt1.predicted_b_breakdown, pt3.failure_log_excerpt+actual, container-log capture) + §6.1 Tier-1 (CONTRACT-2) ([b100979](https://github.com/noonghunna/club-3090/commit/b1009793f26f38370ba2a013ed93d2b6044e0b59))
|
||||
- v0.8.0 [F] F2: §6.1 Tier-2 semantic-fingerprint classifier + Appendix A seed DB (CONTRACT-2 Tier-2) ([9f80d29](https://github.com/noonghunna/club-3090/commit/9f80d29fbe63cd9302221a7d182693b5f803ca82))
|
||||
- v0.8.0 [F] F1: FInput capture-bundle reader + schema-1 validation + key-normalization (CONTRACT-1) ([1491cbc](https://github.com/noonghunna/club-3090/commit/1491cbc7afaa48df924e2ea27446b7dc3b782c42))
|
||||
- v0.8.0 [E] E-outcome-fix: honest 3-state manifest outcome (partial-success != failed) — §6.2 partial is a capability-scoped success ([71148d6](https://github.com/noonghunna/club-3090/commit/71148d6054b4616ada57654b7c405cb0c0d50cc6))
|
||||
- v0.8.0 [E] E3/E4-fix: boot lifecycle as context manager (server stays up for smoke+capture, teardown on ctx-exit) — on-rig E5 caught teardown-in-finally-before-smoke ([f7c405a](https://github.com/noonghunna/club-3090/commit/f7c405a06d4c8c2ef349533d5d9e6ee59455a073))
|
||||
- v0.8.0 [E] E3-fix: smoke probes the real served-model-name (not literal "derived") + capture failure detail — on-rig E5 caught red-smoke-on-healthy-boot ([16a1e4d](https://github.com/noonghunna/club-3090/commit/16a1e4d944992682ac5c8db9b8f7482730e3b9e6))
|
||||
- v0.8.0 [E] E2-fix-2: verify *.safetensors against HF API lfs.sha256 (not Xet-redirect-fragile HEAD x-linked-etag) — on-rig E5 caught false no-etag ([3ae74bf](https://github.com/noonghunna/club-3090/commit/3ae74bfdcfc9ddc4378003c04180d934282b8482))
|
||||
- v0.8.0 [E] E2-fix: download via hf CLI subprocess (not huggingface_hub lib-import) — on-rig E5 caught ModuleNotFoundError ([806a298](https://github.com/noonghunna/club-3090/commit/806a2985226cd1c6ab9f726882e5a610ba2da69f))
|
||||
- v0.8.0 [E] E5(docs): docs/PULL_EMIT_DERIVED.md (+ private ledger/recon-checklist updates) ([d134d5a](https://github.com/noonghunna/club-3090/commit/d134d5a5fcc2b509e7e259430954c192d18bea44))
|
||||
- v0.8.0 [E] E4: post-[C1] derived-[E] orchestration + trigger semantics + override force-capture (pt5) ([2ed18aa](https://github.com/noonghunna/club-3090/commit/2ed18aad3fd1f648d4a143d91e9e86d97b785f08))
|
||||
- v0.8.0 [E] E3: derived boot (HF_HOME mount) + 4 §6 capture emitters + manifest + derived smoke floor ([f327887](https://github.com/noonghunna/club-3090/commit/f327887c3926b8c119afd1d4883f64f40e1b83a6))
|
||||
- v0.8.0 [E] E2: HF download stage (download_set allowlist + x-linked-etag SHA, no-etag fail-closed, atomic staging) ([7a2ec86](https://github.com/noonghunna/club-3090/commit/7a2ec8664052044677c1724cc46ea78a7cd2988e))
|
||||
- v0.8.0 [E] E1: generate_from_profile + derived-vllm template + EInput + CONTRACT-5 gate ([411c84f](https://github.com/noonghunna/club-3090/commit/411c84fd8ae692acd99824754b292f387ecf9486))
|
||||
- v0.8.0 Pull-Gate P5: docs/PULL_GATE.md (two-path model, 6-stratum taxonomy, §4.1 [C1], hardware-SM) ([2582438](https://github.com/noonghunna/club-3090/commit/2582438d0202b0082464cca0bbb1dbee46d2725c))
|
||||
- v0.8.0 Pull-Gate P4: stratum-5 + [C1] §4.1 total fn + stratum-6 [D] dry-run + pull orchestrator + exhaustive test-pull.sh ([adf7a3b](https://github.com/noonghunna/club-3090/commit/adf7a3bf13e424e96ffb510d318faaf016ebf975))
|
||||
- v0.8.0 Pull-Gate P3: stratum-2 precondition + [C0] engine-support/runtime/hardware gate + [C2a] disk ([4a1d385](https://github.com/noonghunna/club-3090/commit/4a1d3857f25de46c60ff8c0b7f566733f32a2a6f))
|
||||
- v0.8.0 Pull-Gate P2: transformers deriver + ModelProfile/confidence + variant-scoped hf_repos schema ([818b79c](https://github.com/noonghunna/club-3090/commit/818b79ccb523aca5034fefe4abddebf5c697dab2))
|
||||
- v0.8.0 Pull-Gate P1: kv-calc generic-dense family + eligibility predicate + raw_verdict adapter ([1bafcfe](https://github.com/noonghunna/club-3090/commit/1bafcfe4a008c8d4e89763d872a99618e3853112))
|
||||
- v0.8.0: doc generated composes are not relocatable (run with --project-directory) ([a2fc05e](https://github.com/noonghunna/club-3090/commit/a2fc05e1c873ff4a6a3bc97de7df1685b41e5006))
|
||||
- v0.8.0 STEP 5: COMPOSE_GENERATOR.md + PATCH_POLICY.md (#141 contributor contract) ([9546f99](https://github.com/noonghunna/club-3090/commit/9546f99303bfcd907618d7778cf36237c33b858e))
|
||||
- v0.8.0 STEP 3+4: compose generator + 5-triple golden-parity test (#141) ([6d7a043](https://github.com/noonghunna/club-3090/commit/6d7a043d907aedab88cc10c8b97bea75a2c4a81d))
|
||||
- v0.8.0 STEP 2: extract patch_attribution.py (sound body-only reaches(), test imports it) ([60f3983](https://github.com/noonghunna/club-3090/commit/60f39832834051a294deb445bfc62d5ca90d0486))
|
||||
- v0.8.0 Phase A-prime: enrich patch/profile data for #141 generator (compose_service_template, genesis_equipped, delivery metadata, drift_guards, drafter/model_slug/trc fold-ins) ([9f23736](https://github.com/noonghunna/club-3090/commit/9f23736f014dbbea74b56a4287836019d0315fb9))
|
||||
- Add v0.8 Phase A patch attribution data ([91a9622](https://github.com/noonghunna/club-3090/commit/91a9622619ba2b1360676a95e412f30d3e975d7c))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.8.0`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.7.4...v0.8.0)
|
||||
## v0.7.4 — 2026-05-15
|
||||
|
||||
|
||||
|
||||
@@ -1,6 +1,8 @@
|
||||
# Adding a model to the club-3090 stack
|
||||
|
||||
End-to-end workflow for onboarding a new model into the v0.7.0 profile catalog + serving infrastructure. Pairs with [KV_MATH.md](KV_MATH.md) (math reference) and [ARCHITECTURE.md](ARCHITECTURE.md) (current stack state).
|
||||
End-to-end workflow for onboarding a new model into the **curated profile catalog** + serving infrastructure. Pairs with [KV_MATH.md](KV_MATH.md) (math reference) and [ARCHITECTURE.md](ARCHITECTURE.md) (current stack state).
|
||||
|
||||
> **Just want to run a model, not add it to the catalog?** As of v0.8.0 you don't need this workflow — `scripts/pull.sh <org/Model> --profile-like vllm/minimal` evaluates *any* safetensors HF repo against the KV math and boots it if it passes (see [docs/PULL.md](PULL.md)). This page is for the heavier task of promoting a model into the **measured/calibration catalog** (real benchmarks, validated composes, calibration anchors, per-model gotchas) — the high-confidence backbone, not a prerequisite for serving.
|
||||
|
||||
## When to add a new model vs a new quant of an existing one
|
||||
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
You have **2× RTX 3090s**. This page is the front door for picking a config and knowing what dual-card unlocks vs single. Model-specific deep dives (quants, Genesis, engine internals) live in the model directory — links at the bottom.
|
||||
|
||||
> **Model not in the configs below / want any HF safetensors repo?** → [`docs/PULL.md`](PULL.md): `scripts/pull.sh` evaluates any model against the KV math (honest, no download) and boots it if it passes. The curated configs on this page are the measured path; both work.
|
||||
|
||||
**NVLink auto-detection** (since 2026-05-14): the dual-card composes now auto-detect whether an NVLink bridge is present. If you have one, you get the NVLink-optimized path automatically. If not, PCIe mode is used. Override with `NVLINK_MODE=force_on|force_off` in your `.env`. See the "NVLink auto-detection" section below.
|
||||
|
||||
> **Have 3+ GPUs?** See [`MULTI_CARD.md`](MULTI_CARD.md) — derivation of TP=4 / TP=8 configs from `dual.yml`, valid TP values for Qwen3.6-27B (1, 2, 4, 5, 8, 10), and what scales vs what doesn't.
|
||||
|
||||
+6
-4
@@ -19,13 +19,15 @@ Qwen3.6-27B is a thinking model. The `<think>...</think>` block before the answe
|
||||
|
||||
| Scenario | `max_tokens` |
|
||||
|---|---|
|
||||
| **FREE thinking on (default long-text / long-vision composes)** | **8192** minimum. 16384 for hard reasoning / competition-grade problems. |
|
||||
| **FREE thinking (enabled per-request — see note below)** | **8192** minimum. 16384 for hard reasoning / competition-grade problems. |
|
||||
| **FSM bounded thinking (`bounded-thinking.yml`)** | **4096** is comfortable. The recommended DeepSeek scratchpad grammar uses ~500-1000 think tokens; the andthattoo G/A/E grammar uses ~150. Either fits well below 4096. |
|
||||
| **`enable_thinking: False`** | Set as tight as the answer needs (50-200 typically). |
|
||||
| **Tool-using agents (multi-turn)** | 1024-2048 per turn. If a middle turn needs >2K to think, your prompt structure probably needs work. |
|
||||
|
||||
The smoke-test examples below use `max_tokens: 200` because they ask short questions where thinking + answer fits comfortably. Real workloads should follow the table above.
|
||||
|
||||
> **Thinking is OFF by default on the shipped composes.** Every Qwen3.6 compose sets `--default-chat-template-kwargs '{"enable_thinking": false}'`, so the model answers directly with no `<think>` block unless you opt in. Enable it per-request with `chat_template_kwargs: {"enable_thinking": true}` (no restart) and budget `max_tokens` per the table. The one exception is `bounded-thinking.yml`, which keeps thinking on but bounds its cost via a structured-CoT grammar (see [`docs/STRUCTURED_COT.md`](STRUCTURED_COT.md)).
|
||||
|
||||
---
|
||||
|
||||
## Quick curl sanity test
|
||||
@@ -40,7 +42,7 @@ curl -sf http://localhost:8020/v1/chat/completions \
|
||||
}' | jq -r '.choices[0].message.content'
|
||||
```
|
||||
|
||||
Expected response: a sentence containing `Paris`. The `max_tokens: 200` headroom is intentional — Qwen3.6 thinks before answering by default, so even simple questions burn ~50–150 tokens inside `<think>...</think>` before reaching the answer. Set tighter (`max_tokens: 30`) only if you also pass `chat_template_kwargs: {"enable_thinking": false}` to skip the think block — that's what `verify-full.sh` does internally.
|
||||
Expected response: a sentence containing `Paris`. The shipped composes set `enable_thinking: false` by default, so the model answers directly with no `<think>` block — `max_tokens: 200` is comfortable slack. If you enable thinking per-request (`chat_template_kwargs: {"enable_thinking": true}`), raise `max_tokens` substantially (see the table above) — the model then emits a `<think>...</think>` block first even for simple questions. `verify-full.sh` passes `enable_thinking: false` explicitly.
|
||||
|
||||
---
|
||||
|
||||
@@ -225,8 +227,8 @@ const client = new OpenAI({
|
||||
const resp = await client.chat.completions.create({
|
||||
model: "qwen3.6-27b-autoround",
|
||||
messages: [{ role: "user", content: "Quicksort in Rust, please." }],
|
||||
// FREE thinking is on by default. 4096 covers easy code-gen think+answer;
|
||||
// 8192 is the safe default for harder coding problems. 800 traps mid-think.
|
||||
// Shipped composes default enable_thinking:false → this answers with no <think> block; 4096 is generous.
|
||||
// To get reasoning, add extra_body chat_template_kwargs {"enable_thinking": true} and budget 8192+ (800 traps mid-think).
|
||||
max_tokens: 4096,
|
||||
temperature: 0.6,
|
||||
top_p: 0.95,
|
||||
|
||||
+6
-2
@@ -64,9 +64,13 @@ Use LM Studio if you prefer a GUI and don't need the engineering. Use this repo
|
||||
|
||||
We tried EAGLE — it's blocked on Qwen3-Next (the family Qwen3.5/3.6 belong to) by DeltaNet hybrid attention's lack of KV rollback support in vLLM/SGLang. MTP works because it's a different protocol (multi-token prediction at draft-head level, not a separate draft model). See [INTERNALS.md "Speculative decoding"](../models/qwen3.6-27b/INTERNALS.md) for the full forensic chain. **Re-test triggers:** if vllm#39931 lands or DeltaNet rollback support arrives upstream, EAGLE becomes viable again.
|
||||
|
||||
### The model I want isn't in the supported list — can I still run it?
|
||||
|
||||
Yes, if it's a **safetensors** repo. As of v0.8.0, `scripts/pull.sh <org/Model> --profile-like vllm/minimal --dry-run` evaluates *any* safetensors HF repo against this stack's KV math — no download — and tells you honestly whether it fits and at what confidence. Drop `--dry-run` (add `--yes`) and, if it passes the gates, it downloads, generates a minimal compose, and boots it. Non-fits stop with a precise reason, not a crash. Full guide: [docs/PULL.md](PULL.md). One heads-up: many common archs (e.g. `Qwen2ForCausalLM`) stop at `needs-trust-remote-code-ack` on the first try even with `--dry-run` — add `--trust-remote-code` (after checking what code the repo runs) to clear it. Limits: safetensors + vLLM only; GGUF / `.bin` repos abort at derive as `unsupported-format` (not a crash) — see next Q.
|
||||
|
||||
### Why not GGUF on vLLM for this model?
|
||||
|
||||
Multiple gates blocked. Qwen3.6-27B GGUF on vLLM hits a chain of "fixed but-not-quite" issues — multimodal config routing, ParallelLMHead skip, the `Qwen35TensorProcessor._reverse_reorder_v_heads` weight loader producing garbage output on the 27B layout (transformers PR #45283 only validated on 0.8B). Tracked in [INTERNALS.md](../models/qwen3.6-27b/INTERNALS.md#qwen36-27b-gguf-on-vllm). Use llama.cpp for GGUF on this model.
|
||||
Multiple gates blocked. Qwen3.6-27B GGUF on vLLM hits a chain of "fixed but-not-quite" issues — multimodal config routing, ParallelLMHead skip, the `Qwen35TensorProcessor._reverse_reorder_v_heads` weight loader producing garbage output on the 27B layout (transformers PR #45283 only validated on 0.8B). Tracked in [INTERNALS.md](../models/qwen3.6-27b/INTERNALS.md#qwen36-27b-gguf-on-vllm). Use llama.cpp for GGUF on this model. **Note (v0.8.0):** `pull` evaluates *safetensors* repos only — GGUF→llama.cpp is **not** served via `pull` (it stays the curated/manual path; cross-engine generation is deliberately deferred). A GGUF/`.bin` repo aborts cleanly at the deriver stage as `unsupported-format` (the message is generic — it does not yet say "GGUF, use llama.cpp"; a clearer message is a tracked v0.8.1 follow-up), not a crash.
|
||||
|
||||
### Why AutoRound INT4 not GPTQ / AWQ?
|
||||
|
||||
@@ -134,7 +138,7 @@ If your numbers on the same compose look different from ours by >15%, the most l
|
||||
|
||||
For a first install, run `bash scripts/setup.sh` with no model argument in a normal terminal. It opens a hardware-aware model picker, marks Qwen / Gemma / Both as eligible or not for your detected GPUs, then continues into the existing download flow.
|
||||
|
||||
After setup, run `bash scripts/launch.sh`. The wizard asks which model (filtered to what you've downloaded), then which GPU(s) to use, auto-picks TP for homogeneous sets (PP for heterogeneous), filters variants by hardware fit, shows a per-card VRAM projection from `tools/kv-calc.py` for the suggested default, then boots and runs `verify-full.sh`. Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, partial flags like `bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1` (skips prompts), `--tp 4 --pp 2` to override parallelism, plus `setup.sh --help` / `launch.sh --help` for the full flag list.
|
||||
After setup, run `bash scripts/launch.sh`. The wizard asks which model (filtered to what you've downloaded), then which GPU(s) to use, auto-picks TP for homogeneous sets (PP for heterogeneous), filters variants by hardware fit, shows a per-card VRAM projection from `tools/kv-calc.py` for the suggested default, then boots and runs `verify-full.sh`. Power-user forms still work: `bash scripts/setup.sh qwen3.6-27b`, `bash scripts/launch.sh --variant vllm/dual`, partial flags like `bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1` (skips prompts), `--tp 4 --pp 2` to override parallelism, plus `setup.sh --help` / `launch.sh --help` for the full flag list. This wizard covers the **curated catalog**; for a model *not* in the catalog (any safetensors HF repo), use `scripts/pull.sh` instead — see [docs/PULL.md](PULL.md).
|
||||
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
|
||||
|
||||
|
||||
@@ -6,6 +6,16 @@ Plain-language definitions for terms used throughout the docs. Roughly grouped b
|
||||
|
||||
---
|
||||
|
||||
## Universal `pull` (v0.8.0)
|
||||
|
||||
| Term | What it means |
|
||||
|---|---|
|
||||
| **`pull`** | `scripts/pull.sh <hf-repo> --profile-like <key>` — evaluates *any* safetensors HF repo against this stack's KV math and, if it passes the gates, downloads + generates a minimal compose + boots it. The model-agnostic front door; the curated catalog still works unchanged. See [PULL.md](PULL.md). |
|
||||
| **dry-run** | `pull … --dry-run` — evaluate only: never downloads, never boots, just prints the verdict. |
|
||||
| **confidence tier** | How much the fit verdict is trusted: `exact` (a measured/curated profile) vs `estimated-lower-bound` (derived from the repo's own config — a floor, likely under-modeled). Always shown with the verdict. |
|
||||
| **boot-fit ≠ runtime-stability** | A "fits" verdict is a *boot-time* allocation check. It's necessary-not-sufficient: a config that boots clean can still degrade/OOM under sustained accumulated-context agent workloads (see [CLIFFS.md](CLIFFS.md)). Validate with soak-continuous before relying on it. |
|
||||
| **calibration backbone** | The curated catalog's role under v0.8.0 — the measured anchor set the KV math is calibrated against (vs. being the only supported models). |
|
||||
|
||||
## Throughput / latency
|
||||
|
||||
| Term | What it means |
|
||||
|
||||
+1
-1
@@ -53,7 +53,7 @@ Use `--force` only when you are intentionally testing an unsupported combo. Exam
|
||||
|
||||
On 20 GB cards (modded 3080) the cudagraph-profiling overhead is a meaningful slice of available VRAM. Drop `--gpu-memory-utilization` to **0.82** (vs shipped 0.95 for 24 GB). vLLM nightly's `gpu_worker.py` reports the equivalent effective KV size in the boot log; tune to keep activation headroom for the ~15K tool-prefill peak (verify-full check 8). Credit: [@troymroberts](https://github.com/troymroberts).
|
||||
|
||||
**4090s with attached display — env-override the compose defaults.** Some 4090 rigs land at ~23.5 GB usable VRAM with X server + driver overhead, vs the headless 3090s the composes are calibrated for. Boot may fail with `No available memory for the cache blocks` at default `max-model-len`. Cross-rig data: @laurimyllari's 4090 single-card on `long-text.yml` needed `MAX_MODEL_LEN=90000` (down from 180K default) to fit cleanly ([disc #62](../../../noonghunna/club-3090/discussions/62) / [issue #71](../../../noonghunna/club-3090/issues/71)). Pattern:
|
||||
**4090s with attached display — env-override the compose defaults.** Some 4090 rigs land at ~23.5 GB usable VRAM with X server + driver overhead, vs the headless 3090s the composes are calibrated for. Boot may fail with `No available memory for the cache blocks` at default `max-model-len`. Cross-rig data: @laurimyllari's 4090 single-card on `long-text.yml` needed `MAX_MODEL_LEN=90000` (down from 180K default) to fit cleanly ([disc #62](../../../noonghunna/club-3090/discussions/62) / [issue #71](../../../noonghunna/club-3090/issues/71)). **Newer driver shrinks the budget the same way even on a headless 3090:** @sethbrasile's controlled 9-run matrix on a headless 3090 with driver 595.71.05 / CUDA 13.2 capped `long-text.yml` at `MAX_MODEL_LEN=105000` — the newer driver's activation-profile reserve measured ~2.87 GiB vs ~1.5 GiB on the bare-metal reference rig, shrinking the KV pool by the difference ([issue #149](../../../noonghunna/club-3090/issues/149)). On a newer-driver 3090, start at `MAX_MODEL_LEN=105000` rather than the 180K default. Pattern:
|
||||
|
||||
```bash
|
||||
MAX_MODEL_LEN=90000 bash scripts/switch.sh vllm/long-text
|
||||
|
||||
@@ -7,6 +7,8 @@ This page explains what scales (and what doesn't) when going beyond TP=2,
|
||||
the constraints to know, and how to derive your own compose when `multi4.yml`
|
||||
isn't your topology.
|
||||
|
||||
> **Model not in the configs here / want any HF safetensors repo?** → [`docs/PULL.md`](PULL.md): `scripts/pull.sh` evaluates any model against the KV math (honest, no download) and boots it if it passes. The recipes on this page are the measured/derivation path; both work.
|
||||
|
||||
> **Validation note:** the maintainer rig is **2× RTX 3090 PCIe**, but
|
||||
> Whamp's 4× RTX 3090 PCIe rig validated the TP=4 fp8/MTP baseline in
|
||||
> [discussion #26](https://github.com/noonghunna/club-3090/discussions/26)
|
||||
|
||||
+6
-2
@@ -37,6 +37,8 @@ What you'll see — exactly one of:
|
||||
| `hard-block` | `2` | Honest stop with a precise reason (unsupported engine/arch, won't-fit, disk, needs `--trust-remote-code`). Nothing downloaded. |
|
||||
| `override-accepted` | `0` | You explicitly accepted a non-pass path (e.g. `--force-download`); proceeds with the caveat recorded. |
|
||||
|
||||
> **First-run heads-up:** many common models (anything `Qwen2ForCausalLM` — Qwen2.5 & a large family, plus other custom-code archs) hard-block at `[C0] needs-trust-remote-code-ack` on the *very first* try — **even with `--dry-run`**. That's the gate working, not a failure. After you've checked what code the repo would run, add **`--trust-remote-code`** to that same command to clear it. See [`--trust-remote-code` — a security decision](#--trust-remote-code--a-security-decision) below.
|
||||
|
||||
It is **honest about confidence and never silently passes.** A "fits"
|
||||
verdict is a *boot-time* check — read [Boot-fit ≠ runtime-stability](#boot-fit--runtime-stability--read-this)
|
||||
before relying on it for sustained agent workloads. Full detail below.
|
||||
@@ -116,8 +118,10 @@ scripts/pull.sh some-org/Some-Llama-7B --profile-like vllm/minimal --dry-run
|
||||
|---|---|
|
||||
| `0` | Download-eligible / clean verdict. |
|
||||
| `3` | Needs a flag — a `confirm→proceed` or advisory terminal that is not yet satisfied (re-run with the named flag). |
|
||||
| `2` | Honest hard-stop (a gate aborted, or `hard-block`). |
|
||||
| `64` | Usage error. |
|
||||
| `2` | Honest hard-stop — a gate aborted, or a `hard-block` terminal. |
|
||||
| `64` | Usage error — missing/unknown argument (distinct from `2`, so a typo is distinguishable from an honest gate-block). |
|
||||
|
||||
> *Note: the `64` usage-vs-`2` hard-stop split is a post-`v0.8.0` fix — present on `master`/the next release; the `v0.8.0` release tag still exits `2` for argument errors.*
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
You have **one RTX 3090 (24 GB VRAM)**. This page is the front door for picking a config and knowing what to expect. The model-specific deep dives (quants, Genesis patches, engine internals) live elsewhere — links at the bottom.
|
||||
|
||||
> **Model not in the configs below / want any HF safetensors repo?** → [`docs/PULL.md`](PULL.md): `scripts/pull.sh` evaluates any model against the KV math (honest, no download) and boots it if it passes. The curated configs on this page are the measured path; both work.
|
||||
|
||||
---
|
||||
|
||||
## ⚠️ Critical — read first if you're running an agentic coding client
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
# vLLM PR #41800 overlay — `truncate_prompt_tokens` kwarg on `get_max_tokens`
|
||||
|
||||
## What this fixes
|
||||
|
||||
Agentic clients (opencode, codex-cli, and similar IDE/agent runtimes) send `truncate_prompt_tokens` on chat-completion requests. Pre-[vLLM PR #41800](https://github.com/vllm-project/vllm/pull/41800), `vllm.entrypoints.utils.get_max_tokens()` doesn't accept that kwarg — and the kwarg propagates from the request handler down into the function call — so requests fail with:
|
||||
|
||||
```
|
||||
HTTP 400: {"error":{"message":"get_max_tokens() got an unexpected keyword argument 'truncate_prompt_tokens'",...}}
|
||||
```
|
||||
|
||||
The fix is upstream PR #41800 (merged 2026-05-06 at commit `d5b31c95`). It adds the kwarg to the function signature and a small body block that clamps `input_length` to `min(input_length, truncate_prompt_tokens or max_model_len)` before the existing length check.
|
||||
|
||||
## When this overlay is needed
|
||||
|
||||
This overlay is needed on engines pinned to vLLM SHAs that **predate `d5b31c95`**:
|
||||
|
||||
| Engine | Pinned SHA | Pre-fix? |
|
||||
|---|---|---|
|
||||
| `vllm-nightly-mtp` | `01d4d1ad` (2026-05-04) | ✅ needs overlay |
|
||||
| `vllm-nightly-dflash` | `e47c98ef` (~2026-05-05) | ✅ needs overlay (20 commits behind d5b31c95) |
|
||||
| `vllm-nightly-full` | `e47c98ef` | ✅ needs overlay |
|
||||
| `vllm-nightly-clean` | `bf610c2f` (2026-05-15) | ❌ already includes fix |
|
||||
|
||||
If a compose routes through `vllm-nightly-clean`, the overlay is unnecessary — the function signature already accepts the kwarg upstream.
|
||||
|
||||
## How the overlay works
|
||||
|
||||
`install.sh` is a Python anchor-based in-place patcher. It does two surgical edits to the in-container `/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/utils.py`:
|
||||
|
||||
1. **Signature**: adds `truncate_prompt_tokens: int | None = None,` to `get_max_tokens`'s signature, anchored to the existing `override_max_tokens: int | None = None,` line.
|
||||
2. **Body**: inserts a 6-line truncation-aware `input_length` adjustment block before the existing `if max_model_len < input_length:` check, anchored to that line.
|
||||
|
||||
Each insertion carries a sentinel comment (`# PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)`) so re-running the install on an already-patched file is a no-op. Post-patch the file is AST-validated before write.
|
||||
|
||||
Why anchor-based and not full-file replacement: the PR diff is +14 / -0 across a 200-line file — replacing the full file would shadow other upstream changes in `utils.py`. Anchor-based insertion is drift-resistant to unrelated upstream movement.
|
||||
|
||||
## Canonical source
|
||||
|
||||
This is a vendored, byte-identical mirror of the canonical overlay at
|
||||
`models/qwen3.6-27b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/`.
|
||||
`install.sh` is model-agnostic (it patches engine-level `vllm/entrypoints/utils.py`),
|
||||
so it is copied verbatim per the per-model-tree `../../patches/` convention.
|
||||
Keep `install.sh` identical to the qwen3.6-27b copy; track upstream/drop state in `docs/UPSTREAM.md` only.
|
||||
|
||||
## Composes that wire this overlay in (Gemma 4 31b tree)
|
||||
|
||||
* `models/gemma-4-31b/vllm/compose/dual/int8.yml`
|
||||
* `models/gemma-4-31b/vllm/compose/dual/awq.yml`
|
||||
* `models/gemma-4-31b/vllm/compose/dual/int8-tq3.yml`
|
||||
* `models/gemma-4-31b/vllm/compose/dual/dflash.yml`
|
||||
* `models/gemma-4-31b/vllm/compose/dual/dflash-int8.yml`
|
||||
|
||||
## How to add this overlay to another affected compose
|
||||
|
||||
In any compose that routes through `vllm-nightly-mtp` / `vllm-nightly-dflash` / `vllm-nightly-full`, add:
|
||||
|
||||
1. **Volume mount** in the `volumes:` block:
|
||||
|
||||
```yaml
|
||||
- ../../patches/vllm-pr41800-truncate-prompt-tokens/install.sh:/etc/club3090/install-pr41800.sh:ro
|
||||
```
|
||||
|
||||
2. **Install line** in the `entrypoint:` bash script, before `exec vllm serve`:
|
||||
|
||||
```bash
|
||||
bash /etc/club3090/install-pr41800.sh
|
||||
```
|
||||
|
||||
Run `bash install.sh` (the file in this directory) standalone to test against a transient vLLM container before wiring into a compose. See the smoke test in the next section.
|
||||
|
||||
## Smoke test
|
||||
|
||||
```bash
|
||||
docker run --rm --entrypoint /bin/bash \
|
||||
-v $(pwd)/install.sh:/install.sh:ro \
|
||||
vllm/vllm-openai:nightly-01d4d1ad375dc5854779c593eee093bcebb0cada \
|
||||
-c '
|
||||
python3 -c "from vllm.entrypoints.utils import get_max_tokens; import inspect; print(inspect.signature(get_max_tokens))"
|
||||
bash /install.sh
|
||||
python3 -c "from vllm.entrypoints.utils import get_max_tokens; import inspect; print(inspect.signature(get_max_tokens))"
|
||||
'
|
||||
```
|
||||
|
||||
Expected: signature lacks `truncate_prompt_tokens` BEFORE install, has it AFTER. Verified on `01d4d1ad` (2026-05-15).
|
||||
|
||||
## When to drop this overlay
|
||||
|
||||
When **both** are true:
|
||||
|
||||
1. PR #41800 has merged upstream (it has — 2026-05-06 at `d5b31c95`)
|
||||
2. The engine's pinned nightly SHA bumps past `d5b31c95`
|
||||
|
||||
For the Genesis-anchored engines, the bump happens with Sander's next Genesis release cycle (v7.73.x). For `vllm-nightly-dflash` and `vllm-nightly-full`, the bump happens when their respective overlays (PR #41703 DFlash, PR #42102 INT8 PTH KV) are re-validated against a newer nightly.
|
||||
|
||||
Track in `docs/UPSTREAM.md`.
|
||||
|
||||
## Source
|
||||
|
||||
- vLLM PR #41800: https://github.com/vllm-project/vllm/pull/41800
|
||||
- Merged commit: `d5b31c95`
|
||||
- Tracking issue: noonghunna/club-3090#139
|
||||
- Triggered by: noonghunna/club-3090#138 (SEVENID's opencode boot failure)
|
||||
- Patch summary: +7 lines in `vllm/entrypoints/utils.py` (the actual fix) + 5 call-site forward-compat additions in other files (we skip those — the signature fix alone unblocks all known TypeError reports)
|
||||
+141
@@ -0,0 +1,141 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install vLLM PR #41800 — `truncate_prompt_tokens` kwarg on get_max_tokens.
|
||||
#
|
||||
# WHY THIS OVERLAY EXISTS:
|
||||
# opencode (and other agentic clients like codex-cli) send `truncate_prompt_tokens`
|
||||
# on chat-completion requests. Pre-#41800, vLLM's `get_max_tokens()` doesn't
|
||||
# accept that kwarg — and somewhere upstream of the function the kwarg gets
|
||||
# unpacked into the call — so requests fail with:
|
||||
# HTTP 400: get_max_tokens() got an unexpected keyword argument 'truncate_prompt_tokens'
|
||||
#
|
||||
# PR: https://github.com/vllm-project/vllm/pull/41800
|
||||
# Merged: 2026-05-06 at commit d5b31c95
|
||||
# Affected pins on master:
|
||||
# - vllm-nightly-mtp (01d4d1ad, 2026-05-04) — pre-fix
|
||||
# - vllm-nightly-dflash (e47c98ef) — pre-fix
|
||||
# - vllm-nightly-full (e47c98ef) — pre-fix
|
||||
# (vllm-nightly-clean at bf610c2f is POST-fix; doesn't need the overlay)
|
||||
#
|
||||
# Tracking issue: #139 (noonghunna/club-3090)
|
||||
# Triggered by: #138 — SEVENID's opencode boot failure on dual-dflash-noviz.
|
||||
#
|
||||
# WHY A PYTHON ANCHOR-BASED PATCHER:
|
||||
# The PR is +7 lines in `vllm/entrypoints/utils.py` (the actual fix) plus a
|
||||
# handful of forward-compat call-site additions in 5 other files. The
|
||||
# function-signature change in utils.py is the ONLY thing required to fix
|
||||
# the TypeError — once `get_max_tokens` accepts the kwarg, requests stop
|
||||
# crashing. The call-site changes are nice-to-have semantic completeness
|
||||
# (actually applying the truncation), so we patch those too via anchors.
|
||||
#
|
||||
# Idempotent: each anchor checks for a sentinel marker before inserting.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
# Container's vLLM install path. Override via env if vLLM moves.
|
||||
SITE_PACKAGES="${CLUB3090_PR41800_SITE_PACKAGES:-/usr/local/lib/python3.12/dist-packages}"
|
||||
|
||||
UTILS_PY="$SITE_PACKAGES/vllm/entrypoints/utils.py"
|
||||
|
||||
if [ ! -f "$UTILS_PY" ]; then
|
||||
echo "[club3090/pr41800] ERROR: $UTILS_PY not found; aborting overlay install" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
python3 - <<'PY'
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
|
||||
site_packages = os.environ.get(
|
||||
"CLUB3090_PR41800_SITE_PACKAGES",
|
||||
"/usr/local/lib/python3.12/dist-packages",
|
||||
)
|
||||
utils_py = f"{site_packages}/vllm/entrypoints/utils.py"
|
||||
|
||||
SENTINEL = "# PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)"
|
||||
|
||||
# Function signature update: add `truncate_prompt_tokens: int | None = None,`
|
||||
# as a kwarg on `get_max_tokens`. Anchor on the existing line that closes
|
||||
# the signature (`override_max_tokens: int | None = None,` line right before `) -> int:`).
|
||||
SIGNATURE_ANCHOR_RE = re.compile(
|
||||
r'^(?P<indent>[ \t]+)override_max_tokens: int \| None = None,\n(?P<close>[ \t]*\) -> int:)',
|
||||
re.MULTILINE,
|
||||
)
|
||||
SIGNATURE_INSERT = ''' override_max_tokens: int | None = None,
|
||||
truncate_prompt_tokens: int | None = None, # PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)
|
||||
) -> int:'''
|
||||
|
||||
# Body update: insert truncation-aware input_length adjustment BEFORE the
|
||||
# `if max_model_len < input_length:` check. Anchor on that line.
|
||||
BODY_ANCHOR_RE = re.compile(
|
||||
r'^(?P<indent>[ \t]+)if max_model_len < input_length:',
|
||||
re.MULTILINE,
|
||||
)
|
||||
BODY_INSERT = ''' # PATCH: truncate_prompt_tokens kwarg (club3090/pr41800)
|
||||
if truncate_prompt_tokens is not None:
|
||||
limit = truncate_prompt_tokens
|
||||
input_length = min(
|
||||
input_length,
|
||||
max_model_len if limit == -1 else limit,
|
||||
)
|
||||
if max_model_len < input_length:'''
|
||||
|
||||
with open(utils_py, "r", encoding="utf-8") as f:
|
||||
src = f.read()
|
||||
|
||||
if SENTINEL in src:
|
||||
print(f"[club3090/pr41800] {utils_py}: sentinel present, patch already applied; no-op", file=sys.stderr)
|
||||
sys.exit(0)
|
||||
|
||||
# Upstream-fix detection: if the function signature already accepts the kwarg
|
||||
# (i.e. the engine pinned a post-#41800 nightly), the overlay is unnecessary
|
||||
# and should no-op gracefully so composes that mount it on a post-fix image
|
||||
# (e.g. via vllm-nightly-clean) still boot cleanly.
|
||||
UPSTREAM_RE = re.compile(
|
||||
r'def get_max_tokens\([^)]*truncate_prompt_tokens\b',
|
||||
re.DOTALL,
|
||||
)
|
||||
if UPSTREAM_RE.search(src):
|
||||
print(f"[club3090/pr41800] {utils_py}: upstream get_max_tokens() already accepts truncate_prompt_tokens; no-op", file=sys.stderr)
|
||||
sys.exit(0)
|
||||
|
||||
# Apply signature patch first (so the function accepts the kwarg)
|
||||
m = SIGNATURE_ANCHOR_RE.search(src)
|
||||
if not m:
|
||||
print(
|
||||
f"[club3090/pr41800] ERROR: signature anchor "
|
||||
f"'override_max_tokens: int | None = None, ... ) -> int:' not found in {utils_py}. "
|
||||
f"vLLM nightly may have changed entrypoints/utils.py — overlay needs re-anchoring.",
|
||||
file=sys.stderr,
|
||||
)
|
||||
sys.exit(1)
|
||||
|
||||
src = SIGNATURE_ANCHOR_RE.sub(SIGNATURE_INSERT, src, count=1)
|
||||
|
||||
# Apply body patch
|
||||
m = BODY_ANCHOR_RE.search(src)
|
||||
if not m:
|
||||
print(
|
||||
f"[club3090/pr41800] ERROR: body anchor 'if max_model_len < input_length:' not found in {utils_py} "
|
||||
f"after signature patch. vLLM nightly diverged unexpectedly — overlay needs re-anchoring.",
|
||||
file=sys.stderr,
|
||||
)
|
||||
sys.exit(1)
|
||||
|
||||
src = BODY_ANCHOR_RE.sub(BODY_INSERT, src, count=1)
|
||||
|
||||
with open(utils_py, "w", encoding="utf-8") as f:
|
||||
f.write(src)
|
||||
|
||||
# Quick validity check
|
||||
import ast
|
||||
try:
|
||||
ast.parse(src)
|
||||
except SyntaxError as e:
|
||||
print(f"[club3090/pr41800] ERROR: post-patch utils.py is not valid Python: {e}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
print(f"[club3090/pr41800] {utils_py}: signature + body patches applied (truncate_prompt_tokens kwarg)", file=sys.stderr)
|
||||
PY
|
||||
|
||||
echo "[club3090/pr41800] install complete" >&2
|
||||
@@ -71,6 +71,7 @@ arches:
|
||||
- gemma-vllm-gemma4-tool-parser-fixes
|
||||
- gemma-vllm-gemma4-dflash
|
||||
- gemma-vllm-gemma4-dflash-int8
|
||||
- gemma-vllm-pr41800-truncate-prompt-tokens
|
||||
valid_tp:
|
||||
tp_divisors: [1, 2, 4, 8, 16]
|
||||
marlin_alignment_required: false
|
||||
|
||||
@@ -279,6 +279,44 @@ patches:
|
||||
drop_when: "each affected engine profile pins a vLLM nightly after d5b31c95"
|
||||
status: verified
|
||||
|
||||
- id: gemma-vllm-pr41800-truncate-prompt-tokens
|
||||
model: [gemma-4-31b]
|
||||
files:
|
||||
- models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens
|
||||
load_bearing_when:
|
||||
- composes:
|
||||
- vllm/gemma-int8
|
||||
- vllm/gemma-int8-262k
|
||||
- vllm/gemma-awq
|
||||
- vllm/gemma-int8-tq3
|
||||
- vllm/gemma-dflash
|
||||
- vllm/gemma-dflash-int8
|
||||
reason: "Pre-d5b31c95 engine pins reject agent clients that send truncate_prompt_tokens."
|
||||
evidence: "docs/UPSTREAM.md#41800 row; models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/README.md"
|
||||
delivery: # DEPRECATED/READ-ONLY (test-only) — see header
|
||||
dockerfile_bake: false
|
||||
entrypoint_invoke: true
|
||||
genesis: false
|
||||
delivery_gaps: []
|
||||
delivery_mechanism: install_script
|
||||
delivery_spec:
|
||||
script: models/gemma-4-31b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/install.sh
|
||||
mounted_at: /etc/club3090/install-pr41800.sh
|
||||
invoke: "bash /etc/club3090/install-pr41800.sh"
|
||||
invoked_before: vllm-serve
|
||||
wired_at: [volumes, entrypoint]
|
||||
drift_guard:
|
||||
kind: behavioral
|
||||
check: "agent client sending truncate_prompt_tokens does not get HTTP 400 on the selected nightly (no-op if nightly post-d5b31c95)"
|
||||
on_fail: capability-degraded
|
||||
capability: truncate-prompt-tokens-kwarg
|
||||
foundational: false
|
||||
upstream:
|
||||
ref: vllm-project/vllm#41800
|
||||
status: merged
|
||||
drop_when: "each affected engine profile pins a vLLM nightly after d5b31c95"
|
||||
status: verified
|
||||
|
||||
- id: gemma-vllm-gemma4-dflash
|
||||
model: [gemma-4-31b]
|
||||
files:
|
||||
|
||||
@@ -1316,7 +1316,19 @@ _EXIT_USAGE = 64
|
||||
def main(argv: list[str]) -> int:
|
||||
import argparse
|
||||
|
||||
ap = argparse.ArgumentParser(
|
||||
class _UsageExit64Parser(argparse.ArgumentParser):
|
||||
"""argparse's default `error()` hard-exits 2 — which collides with
|
||||
`_EXIT_ABORT` (honest gate hard-stop), so a typo and a legitimate
|
||||
block are indistinguishable to callers/automation. Override to exit
|
||||
`_EXIT_USAGE` (64) on argument/usage errors, restoring the
|
||||
documented contract. `--help` is unaffected (it goes through
|
||||
`exit()`, not `error()`, and still returns 0)."""
|
||||
|
||||
def error(self, message):
|
||||
self.print_usage(sys.stderr)
|
||||
self.exit(_EXIT_USAGE, f"{self.prog}: error: {message}\n")
|
||||
|
||||
ap = _UsageExit64Parser(
|
||||
prog="pull.sh",
|
||||
description="v0.8.0 Pull-Gate — derive an HF repo, gate it through "
|
||||
"the locked 6-stratum taxonomy, and (Path A, curated+emittable) "
|
||||
|
||||
@@ -1335,4 +1335,18 @@ print("\nSUMMARY: all Pull-Gate P4 truth-table assertions passed "
|
||||
"CONTRACT-5 reject).")
|
||||
PY
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# CLI-contract: exit-code boundary (the pure truth-table above can't cover
|
||||
# argv parsing / process exit). #370 regression lock: argparse usage errors
|
||||
# MUST exit 64 (not 2 — argparse default), so a typo is distinguishable
|
||||
# from an honest gate hard-stop (2); --help stays 0; hard-stop stays 2.
|
||||
_clifail=0
|
||||
_ec(){ bash scripts/pull.sh "$@" >/dev/null 2>&1; echo $?; }
|
||||
[ "$(_ec)" = 64 ] || { echo "FAIL: no-args -> 64 (#370)" >&2; _clifail=1; }
|
||||
[ "$(_ec Qwen/Qwen2.5-0.5B-Instruct)" = 64 ] || { echo "FAIL: missing required --profile-like -> 64 (#370)" >&2; _clifail=1; }
|
||||
[ "$(_ec --nope x)" = 64 ] || { echo "FAIL: unknown flag -> 64 (#370)" >&2; _clifail=1; }
|
||||
[ "$(_ec --help)" = 0 ] || { echo "FAIL: --help -> 0 (#370 must not regress help)" >&2; _clifail=1; }
|
||||
[ "$(_ec definitely/nonexistent-xyz123 --profile-like vllm/minimal --dry-run)" = 2 ] || { echo "FAIL: honest hard-stop -> 2 (must stay 2, not 64) (#370)" >&2; _clifail=1; }
|
||||
[ "$_clifail" = 0 ] && echo "PASS: CLI exit-code contract (#370): usage=64, --help=0, hard-stop=2" || { echo "1+ CLI-contract assertion(s) failed." >&2; exit 1; }
|
||||
|
||||
echo "test-pull.sh OK"
|
||||
|
||||
Reference in New Issue
Block a user