b0eeb21ff53a2c0d89677be8d0a3af4eef8a2f4f
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ed1507122c |
llama.cpp single: thinking-off policy alignment + MTP profile family
Closes the dataextract 0/15 quality regression caused by mounting the vLLM-only froggeric Jinja template on the llama.cpp engine, which silently suppressed --reasoning off. Native template + --reasoning off restores it. Stack-wide Qwen3.6 thinking-off-by-default policy now shipped on the llama.cpp path (matches 22/24 vLLM composes). Adds three named single-card profiles: - llamacpp/default (262K vanilla, unchanged) — cliff-immune fallback - llamacpp/mtp (NEW: 131K + MTP n=2) — single-card workhorse, ~60 TPS code - llamacpp/mtp-vision (NEW: 49K + MTP + vision) — first stack profile combining MTP + vision on build 9235 (the older "strip mmproj when MTP" rule was obsolete; sweep-verified MTP + vision coexist) Retires single/concurrent.yml — single-card concurrency is anti-value (per-slot ctx ~48K, worse than every other profile). Concurrency belongs on dual. Fixes a latent user-facing bug: engines/llama-cpp-mainline.yml had supported_drafters: [draft-mtp] (the llama.cpp CLI flag value) while the compat-layer C7 gate compares against drafter spec_method (mtp). Net effect pre-fix: llamacpp/mtp would have been silently filtered out of launch.sh candidate lists. Fixed [draft-mtp] -> [mtp]. Lowers -ub default 2048 -> 1024 on the llama.cpp single composes. The per-pass activation peak halves; verify-stress goes 5/7 -> 7/7 including the 60K + 91K needle rungs previously treated as architectural Cliff 2 territory. Cliff 2 single-prompt at 50-60K was config-driven on llama.cpp, not architectural; CLIFFS.md note added (vLLM Cliff 2 narrative unchanged — different kernel-level failure). Migrates 6 stale call sites for the engine-id rename (llama-cpp-mainline -> llama-cpp-local) that earlier work lagged. Measured impact, Config A (llamacpp/mtp, 131K, MTP n=2, ub=1024): - bench: 51.28 narr / 59.72 code decode TPS (n=3, CV 1.9%/0.5%) - verify-stress: 7/7 PASS (incl. 60K + 91K needle recall) - quality 8-pack: 102/150 (68%) — beats every Qwen vLLM dual config in discussion #119 by 6-16 pp - aider-polyglot-30: 17/30 (56.7%) — matches Qwen vLLM bf16 dual exactly (17/30) on half the hardware - per-GPU code TPS (59.72) ~equal to vLLM dual configs (60-63), confirming the engine-side per-card rate is identical and vLLM dual's aggregate advantage is purely from the second card Config B (llamacpp/mtp-vision, 49K, MTP + vision): - multimodal vision probe passed - bench: 56.52 narr / 66.17 code decode TPS (n=5) - verify-stress: 7/7 PASS Tests green: test-launch-compat, test-profiles-compat, test-switch-registry-parity, test-pullgate-gates, test-patch-attribution. Leak-grep clean. YAML lint clean. Performance charts regenerated to include the two new entries. Doc alignment: SINGLE_CARD, CLIFFS, FAQ, EXAMPLES, README, llama-cpp README, COMPOSE_GENERATOR, PULL_GATE all reflect the new family. CHANGELOG.md + UPSTREAM.md left intact per append-only history rule. Followups (queued, not blocking): #400 MTP retune A/B (n=4 + p-min=0 on the fixed config), #403 soak-test container glob (unlocks Cliff 2b validation on llama.cpp), #405 benchlocal-cli port-offset for parallel rebench, #401 + #404 ik_llama + Tom's TurboQuant trackers. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
a891b3921f | docs: reorder docsindex (GSD first), add FAQ TOC + promote troubleshooting ladder, add tool-calling example | ||
|
|
46bb271737 |
docs(examples): correct "thinking on by default" — shipped composes set enable_thinking=false (#372)
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m17s
Release / release (push) Failing after 50s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
EXAMPLES.md asserted thinking is on by default (lines 22/43/228), but
every shipped Qwen3.6 compose sets
--default-chat-template-kwargs '{"enable_thinking": false}'
(bounded-thinking.yml is the only exception). Same docs-vs-shipped
class as the v0.8.0 docs-fidelity gaps. Corrected the 3 inaccurate
spots, added a canonical "thinking is OFF by default + how to enable
per-request + bounded-thinking exception" note under the max_tokens
table. Consistent with the disc #151 public answer and the
enable_thinking-default rationale; does not pre-judge the #150
froggeric re-eval. Doc-only.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
b956c85477 |
feat(bounded-thinking): Phase 3 grammar A/B complete; DeepSeek scratchpad is the new recommended grammar
After the full Phase 3 5-grammar A/B (HE+ 164 + LCB v6 50, n=214 problems
× 5 conditions = 1070 generations), bounded-thinking.yml is updated to
recommend the DeepSeek scratchpad grammar (PLAN/NOTE×0-15/VERDICT FSM at
tools/grammar-eval/deepseek-scratchpad.gbnf) as the default.
Phase 3 results (combined HE+ + LCB v6, all FSM-enforced):
- DeepSeek scratchpad: 87.4% Pass@1 (+1 net vs andthattoo, +4pp on LCB)
- andthattoo G/A/E: 86.9% (the originally-published technique)
- Holiday tagline: 86.4%
- PROMPT_TERSE (no FSM): 82.2% (Phase 2's n=30 win was subset-selection bias)
- FREE (no constraint): 78.0% (baseline)
Phase 1 reproducibility is exact: HE+ FSM Δ +4.3pp / LCB v6 Δ +24.0pp,
both match Phase 1's published numbers — validates the bench harness.
Compose ships unchanged engine-side (same vLLM image, same Genesis stack,
same TQ3 KV, same MTP n=3, same enable_in_reasoning flag). Only the
docstring's recommended-grammar pointer + the on-disk grammar files in
tools/grammar-eval/ change. Three grammars are now validated and available
client-side via extra_body={"structured_outputs": {"grammar": ...}}:
- DeepSeek scratchpad (default, best LCB)
- andthattoo G/A/E (originally-published, ~4× tighter think budget)
- Holiday tagline (extreme 24-token compression, wins LCB by 4pp too)
Combined-accuracy spread is within noise (0.5pp at n=214), so we ship one
compose rather than three siblings — choice is at the client, not at the
compose level.
Bug fix bundled: tools/grammar-eval/subset-bench.py --full --include-lcb
mode now correctly threads dataset kind through run_condition so LCB
problems use mod.run_tests_livecodebench instead of HE+ assertion-based
testing. Without this, the Phase 3 LCB shard crashed with KeyError:
'prompt' on the HE→LCB transition.
This wraps active research on bounded-thinking. Reopen if upstream FSM-
regress cluster behavior changes (Genesis pin bump, vLLM grammar engine
swap, model-family change), or if user demand surfaces for a sibling
compose pinning a non-default grammar.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
f8c9c365e0 |
docs: full sync to v7.69 + Cliff 2 60K closure recipes
Sweep all stale v7.66 / fc89395 substrate references to v7.69 (commit 2db18df) + local vllm#35975 inputs_embeds backport. Ship the Balanced MTP (long-text.yml, 180K + 0.93) and Max-context (long-text-no-mtp.yml, 200K + 0.95, no MTP) variants as the Cliff 2 closure recipes — both PASS the 60K single-prompt envelope (623s and 537s wall respectively). Updates: - CHANGELOGs (root + model) — new v7.69 PM entry above v7.66 - README + SINGLE_CARD + HARDWARE + EXAMPLES + FAQ + INTERNALS + VLLM engine doc — Cliff 2 status, substrate pins, mem-util defaults, variant table, sidecar list - vllm/README.md compose menu refreshed for the new ctx envelopes - model README patch surface table — added PN30 part3, PN32, P103, PN34 rows; collapsed P98 reference to PN34 env-gate - tools/charts/gen-perf.py + gen-vram.py — substrate label bumped to v7.69 + #35975, panel labels for the long-text variants updated, long-text-no-mtp 200K Max-context noted as bench-pending in chart - All performance + VRAM charts (svg + png) regenerated Cliff 2 60K closure: Genesis v7.69 (PN32 GDN chunked-prefill + P103 worker self-install + PN30 part3 + PN34 workspace_lock relax) plus local backport of vllm#35975 (~444 MiB freed on text-only paths). 3 sidecars dropped on long-text variants; 2 sidecars retained on master (patch_inputs_embeds_optional.py, patch_tolist_cudagraph.py). >60K single-prompt still hits the 24 GB hardware-physical wall on single-card. For those: dual-card TP=2 (verified at 237K) or llama.cpp single-card (262K, different engine). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
5aa97a25d9 |
v0.20 migration + Genesis v7.65 dev tip + cold-start cache + env-var alignment
This branch migrates the entire vLLM stack from `dev205+g07351e088` + Genesis
v7.64 to `0.20.1rc1.dev16+g7a1eb8ac2` + Genesis v7.65 dev tip (commit
`d89a089`). v7.65 is on Sandermage's `dev` branch — explicitly the cross-rig
testing surface he requested in discussion #19; he'll merge dev→main once we
both confirm stable. Pin gates restated when that lands.
What changes
------------
Pin migration:
- vLLM image: nightly-07351e08... → nightly-7a1eb8ac2... (dev205 → v0.20.1rc1.dev16)
- Genesis: 64dd18b (v7.64) → d89a089 (v7.65 dev tip)
Sidecar churn:
- DROPPED: patch_pn12_ffn_pool_anchor.py (PN12 native on v0.20)
- DROPPED: patch_pn12_compile_safe_custom_op.py (Genesis P38B in-source hook)
- DROPPED: patch_fa_max_seqlen_clamp.py (Genesis PN17 + P15B)
- ADDED: patch_workspace_lock_disable.py (relaxes vllm#39226 strict assertion;
P98 covers same surface but auto-skips on v0.20 due to drift-marker false
positive — pending Sandermage marker fix)
Env-var alignment to Sandermage's PROD set (start_27b_int4_TQ_k8v4.sh@dev):
- FIXED naming bugs that silently no-op'd patches:
- PN9_INDEPENDENT_DRAFTER_ATT → _ATTN (was silently OFF)
- PN22 → PN22_LOCAL_ARGMAX_TP (was silently OFF)
- PN26_BLOCK_KV → PN26_SPARSE_V_BLOCK_KV (fell back to default 4, not 8)
- PN26_NUM_WARPS → PN26_SPARSE_V_NUM_WARPS
- PN26_THRESHOLD → PN26_SPARSE_V_THRESHOLD (fell back to default 0.001, not 0.01)
- ADDED explicit-OFFs to match Sander's PROD verbatim:
- P78_TOLIST_CAPTURE_GUARD=0 (we use our own patch_tolist_cudagraph.py)
- P81_FP8_BLOCK_SCALED_M_LE_8=0 (FP8-specific, no-op on TQ3)
- P82=0, P82_THRESHOLD_SINGLE=0.3
- Cap divergence (justified): PROFILE_RUN_CAP_M=4128 + PREALLOC_TOKEN_BUDGET=4128
(Sander uses 4096 — vLLM `interface.py:639` forces our config's Mamba
block_size to 4128 due to TQ3 + TP=1 page-size math; lower values
AssertionError at boot)
- Carry-forward (intentional): P4 (hybrid TQ required), P65 (TQ spec-CG
downgrade — pending v0.20 verification that #40880 closure makes it
redundant)
Cold-start cache mounts (closes #22):
- All 10 composes now mount torch_compile_cache + Triton cache from
`models/qwen3.6-27b/vllm/cache/`. First boot warms (~6 min); warm boot
drops to ~3.2 min (47% faster). Per-stage savings on long-text:
- Dynamo bytecode transform: 18s → 5s (-73%)
- torch.compile: 57s → 9s (-85%)
- Initial profiling/warmup: 51s → 7s (-87%)
Mamba block_size cap fix:
- v0.20 enforces `long_prefill_token_threshold >= block_size`; on hybrid
Mamba+TQ3, vLLM forces block_size=4128. Bumped GENESIS_PROFILE_RUN_CAP_M
and PREALLOC_TOKEN_BUDGET 4096→4128 across all 5 main composes.
Default 48K compose:
- Required workspace_lock_disable sidecar after initial v0.20 boot hit
vllm#39226 strict assertion. Caught during validation, fixed.
Context restored vs dev205 backoffs (validated 33K + 50K stress on v0.20):
- long-text: 185K → 214K (+16%)
- long-vision: 140K → 198K (+41%)
- bounded-thinking: 185K → 214K (+16%)
Bench results (n=5, results/v0.20-migration/):
- long-text 214K narr 49.74 / code 67.39 (CV 2.6/2.7%)
- long-vision 198K narr 50.32 / code 66.12 (CV 2.3/4.1%)
- bounded-thinking 214K narr 49.77 / code 65.80 (CV 1.4/2.3%)
- tools-text 75K (fp8) narr 53.32 / code 69.66 (CV 2.3/1.4%)
- dual-turbo 262K (TP=2) narr 58.33 / code 76.01 per-stream
269 TPS aggregate at n=4 streams (3.63x speedup)
- default 48K narr 48.82 / code 65.98 (n=3)
Validation: verify-full 8/8 on every variant. verify-stress 33K AND 50K
tool-prefill PASS on every variant — the cliff that fired on EVERY dev205
config no longer reproduces.
Docs + charts:
- README + SINGLE_CARD + DUAL_CARD + CLIFFS + EXAMPLES + STRUCTURED_COT
+ FAQ + UPSTREAM + 3 engine docs + model README + INTERNALS + CHANGELOG
all updated with new pin, ctx, TPS numbers, and "v0.20 unblock" section
- performance.{png,svg} + variants regenerated with measured TPS
- vram-budget.{png,svg} + variants regenerated with measured VRAM
- UPSTREAM tracker: 5 issues moved ✅ closed (PR #12, #13, #14, #15, P104
superseded by PN17 + P15B)
Issues addressed:
- #16 (Cliff 1 mech B leaks past PN12 on inductor-compiled FFN) — partial:
v0.20's revised TQ FA paths close the synthetic stress; PN25 (Sander's
proper compile-path opaque-op fix) is on dev but explicitly opt-in pending
worker-fork registration fix. Workarounds documented (tools-text fp8 path
/ --enforce-eager) until Sander ships PN25 default-on.
- #20 (launch.sh port + container-name mismatch) — already closed by
|
||
|
|
cc4f0835e6 |
docs+composes: refresh long-text/long-vision/bounded-thinking headers + max_tokens guidance
The compose headers were stale (long-text said 218K + 0.985, long-vision
192K + 0.98, bounded-thinking 218K) and didn't reference v7.64 or the
new patch stack. Updated all three to reflect the shipped 185K + 0.975
(text) / 140K + 0.95 (vision) configs and link to docs/CLIFFS.md "Update
2026-05-01 PM" for the bisection rationale.
Each compose header now also documents the recommended client max_tokens
default:
- long-text + long-vision (FREE thinking): 8192 (16384 for hard
reasoning / competition problems). 4096 was the trap that bit our
LCB v6 baseline mid-think.
- bounded-thinking (FSM grammar caps think): 4096 is sufficient — the
grammar bounds think to ~150-300 structured tokens.
docs/EXAMPLES.md gets a new top-of-doc "max_tokens defaults" table so
copy-paste users land on the right number without reading the bench
forensics. Two existing examples bumped: math reasoning 400 → 2048 (easy
math but FREE thinking can run that), Quicksort code 800 → 4096 (code
gen with FREE thinking traps at 800).
Smoke-test "Capital of France" examples kept at max_tokens=200 — that's
the documented intentional headroom for thinking + short answer.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
2f8bade82c |
fix(docs): bump curl smoke-test max_tokens 30 → 200 (#14)
Qwen3.6 thinks before answering by default, so a "Capital of France?" smoke with max_tokens=30 returns truncated mid-`<think>` content. apnar hit this on a working stack (verify-full.sh all green) and wasted time debugging a non-bug. Bump all 7 user-facing curl examples to max_tokens=200 (covers a typical think block + the one-sentence answer with headroom). verify-full.sh / verify.sh / verify-stress.sh stay at max_tokens=30 because they already pass chat_template_kwargs.enable_thinking=false, which skips the think block entirely. EXAMPLES.md gets an inline note explaining the headroom + the alternative (disable thinking via chat_template_kwargs) for users who want a tighter smoke. |
||
|
|
48f93e550f |
docs: demote 48K/tools-text/minimal to fallback; lead with long-* + llama.cpp
User feedback: the small-ctx variants (48K default, tools-text 75K, minimal 32K) "offer very little context and not many people will find that as viable options." Now that Cliff 1 is closed on the long-* variants via the PN12 anchor sidecar, they're strictly more useful than the 48K/75K alternatives for the workloads most users come for. The single residual limitation is Cliff 2 on single-prompt >50K, addressed by llama.cpp. SINGLE_CARD.md: - TL;DR table reduced to 3 recommended options (long-vision · long-text · llamacpp/default). - Cliff 2 caveat promoted to a prominent ⚠️ callout right under the table — the one limitation users need to know. - Old per-variant sections folded; small-ctx variants moved to an "Other variants in the repo" section as fallback / diagnostic. scripts/launch.sh: - Wizard leads with the 3 primary options (long-vision · long-text · llamacpp/default). Diagnostic / niche options bundled at the end with a "[fallback]" prefix so they don't dominate the menu. models/qwen3.6-27b/README.md: - Single-card recommended-options bullet list now leads with the 3 primary variants and explicitly names Cliff 2 as the single shipped limitation. docs/EXAMPLES.md: - Cline section: stop pointing at tools-text; long-* now handle Cline's tool returns. Cliff 2 is the only remaining caveat to flag. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
26ac8118de |
Restructure docs around hardware axis: SINGLE_CARD.md + DUAL_CARD.md
User feedback: navigating to relevant docs was cumbersome. The natural first decision is "1 GPU or 2 GPUs?" and the existing docs mixed model- specific reference with deployment guidance. New navigation: - README.md adds a "Pick your path" pivot pointing at hardware-axis pages - docs/SINGLE_CARD.md — 1× 3090 deployment menu (workload → compose → TPS) with all single-card configs (vLLM + llama.cpp), VRAM budget, prefill cliffs explained operationally, what single-card can't do - docs/DUAL_CARD.md — 2× 3090 mirror (4 dual variants + TP=2 explainer + what dual unlocks vs single + Marlin pad fork dependency) Slimming: - models/qwen3.6-27b/README.md: dropped duplicated variant tables (now in GPU-count pages); kept model-specific content (quants, Genesis patch surface table, what's working / not, VRAM diagram) - models/qwen3.6-27b/USE_CASES.md: deleted. Per-workload content absorbed into the GPU-count pages (deduplicated). Troubleshooting list moved to docs/FAQ.md as a new "Troubleshooting" subsection. Image-token cost / vision specifics absorbed into SINGLE_CARD.md. Reference updates: 8 files updated (engines/VLLM.md, engines/README.md, COMPARISONS.md, ARCHITECTURE.md, EXAMPLES.md, FAQ.md, INTERNALS.md, top-level README) — all USE_CASES.md links re-pointed to SINGLE_CARD/ DUAL_CARD where appropriate. Net delta: -167 lines (was 197 in USE_CASES + duplicated tables in model README; now 342 lines split between SINGLE_CARD + DUAL_CARD with content deduplicated against each other). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
91b817fa72 |
Add docs/EXAMPLES.md — client snippets + IDE / Open WebUI connection
One-stop reference for "how do I actually call this from code?" Sections: - Curl sanity test (one-liner, jq-extracts the answer) - Python via openai SDK: chat / streaming / tool calls / vision / reasoning mode (with the llama.cpp parser-gap caveat called out) - Python via raw `requests` (no SDK) — for environments where installing openai isn't an option, including SSE streaming parse - TypeScript / Node — same flows - Connection settings for Open WebUI, Cline / Roo, Cursor, with the Cliff-1 warning for tool-using agents that send big returns - Security note re 0.0.0.0:8020 binding (LAN exposure caveat) with the 127.0.0.1 opt-in Linked from top-level README and models/qwen3.6-27b/README.md. The Cline/Cursor sections specifically call out the Cliff 1 risk (25K+ tool returns OOM on vLLM single-card 192K) and recommend either vllm/default (48K) or llamacpp/default (cliff-free at 21 TPS). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |