This branch migrates the entire vLLM stack from `dev205+g07351e088` + Genesis
v7.64 to `0.20.1rc1.dev16+g7a1eb8ac2` + Genesis v7.65 dev tip (commit
`d89a089`). v7.65 is on Sandermage's `dev` branch — explicitly the cross-rig
testing surface he requested in discussion #19; he'll merge dev→main once we
both confirm stable. Pin gates restated when that lands.
What changes
------------
Pin migration:
- vLLM image: nightly-07351e08... → nightly-7a1eb8ac2... (dev205 → v0.20.1rc1.dev16)
- Genesis: 64dd18b (v7.64) → d89a089 (v7.65 dev tip)
Sidecar churn:
- DROPPED: patch_pn12_ffn_pool_anchor.py (PN12 native on v0.20)
- DROPPED: patch_pn12_compile_safe_custom_op.py (Genesis P38B in-source hook)
- DROPPED: patch_fa_max_seqlen_clamp.py (Genesis PN17 + P15B)
- ADDED: patch_workspace_lock_disable.py (relaxes vllm#39226 strict assertion;
P98 covers same surface but auto-skips on v0.20 due to drift-marker false
positive — pending Sandermage marker fix)
Env-var alignment to Sandermage's PROD set (start_27b_int4_TQ_k8v4.sh@dev):
- FIXED naming bugs that silently no-op'd patches:
- PN9_INDEPENDENT_DRAFTER_ATT → _ATTN (was silently OFF)
- PN22 → PN22_LOCAL_ARGMAX_TP (was silently OFF)
- PN26_BLOCK_KV → PN26_SPARSE_V_BLOCK_KV (fell back to default 4, not 8)
- PN26_NUM_WARPS → PN26_SPARSE_V_NUM_WARPS
- PN26_THRESHOLD → PN26_SPARSE_V_THRESHOLD (fell back to default 0.001, not 0.01)
- ADDED explicit-OFFs to match Sander's PROD verbatim:
- P78_TOLIST_CAPTURE_GUARD=0 (we use our own patch_tolist_cudagraph.py)
- P81_FP8_BLOCK_SCALED_M_LE_8=0 (FP8-specific, no-op on TQ3)
- P82=0, P82_THRESHOLD_SINGLE=0.3
- Cap divergence (justified): PROFILE_RUN_CAP_M=4128 + PREALLOC_TOKEN_BUDGET=4128
(Sander uses 4096 — vLLM `interface.py:639` forces our config's Mamba
block_size to 4128 due to TQ3 + TP=1 page-size math; lower values
AssertionError at boot)
- Carry-forward (intentional): P4 (hybrid TQ required), P65 (TQ spec-CG
downgrade — pending v0.20 verification that #40880 closure makes it
redundant)
Cold-start cache mounts (closes #22):
- All 10 composes now mount torch_compile_cache + Triton cache from
`models/qwen3.6-27b/vllm/cache/`. First boot warms (~6 min); warm boot
drops to ~3.2 min (47% faster). Per-stage savings on long-text:
- Dynamo bytecode transform: 18s → 5s (-73%)
- torch.compile: 57s → 9s (-85%)
- Initial profiling/warmup: 51s → 7s (-87%)
Mamba block_size cap fix:
- v0.20 enforces `long_prefill_token_threshold >= block_size`; on hybrid
Mamba+TQ3, vLLM forces block_size=4128. Bumped GENESIS_PROFILE_RUN_CAP_M
and PREALLOC_TOKEN_BUDGET 4096→4128 across all 5 main composes.
Default 48K compose:
- Required workspace_lock_disable sidecar after initial v0.20 boot hit
vllm#39226 strict assertion. Caught during validation, fixed.
Context restored vs dev205 backoffs (validated 33K + 50K stress on v0.20):
- long-text: 185K → 214K (+16%)
- long-vision: 140K → 198K (+41%)
- bounded-thinking: 185K → 214K (+16%)
Bench results (n=5, results/v0.20-migration/):
- long-text 214K narr 49.74 / code 67.39 (CV 2.6/2.7%)
- long-vision 198K narr 50.32 / code 66.12 (CV 2.3/4.1%)
- bounded-thinking 214K narr 49.77 / code 65.80 (CV 1.4/2.3%)
- tools-text 75K (fp8) narr 53.32 / code 69.66 (CV 2.3/1.4%)
- dual-turbo 262K (TP=2) narr 58.33 / code 76.01 per-stream
269 TPS aggregate at n=4 streams (3.63x speedup)
- default 48K narr 48.82 / code 65.98 (n=3)
Validation: verify-full 8/8 on every variant. verify-stress 33K AND 50K
tool-prefill PASS on every variant — the cliff that fired on EVERY dev205
config no longer reproduces.
Docs + charts:
- README + SINGLE_CARD + DUAL_CARD + CLIFFS + EXAMPLES + STRUCTURED_COT
+ FAQ + UPSTREAM + 3 engine docs + model README + INTERNALS + CHANGELOG
all updated with new pin, ctx, TPS numbers, and "v0.20 unblock" section
- performance.{png,svg} + variants regenerated with measured TPS
- vram-budget.{png,svg} + variants regenerated with measured VRAM
- UPSTREAM tracker: 5 issues moved ✅ closed (PR #12, #13, #14, #15, P104
superseded by PN17 + P15B)
Issues addressed:
- #16 (Cliff 1 mech B leaks past PN12 on inductor-compiled FFN) — partial:
v0.20's revised TQ FA paths close the synthetic stress; PN25 (Sander's
proper compile-path opaque-op fix) is on dev but explicitly opt-in pending
worker-fork registration fix. Workarounds documented (tools-text fp8 path
/ --enforce-eager) until Sander ships PN25 default-on.
- #20 (launch.sh port + container-name mismatch) — already closed by
77ca576 (post-issue-filing).
- #22 (cold-start caching) — closed by cache mounts above.
Remaining caveats:
- Cliff 2 (DeltaNet GDN forward, single prompt ≥50-60K) unchanged —
architectural, applies to all single-card vLLM TQ3 paths. Mitigation:
dual-turbo TP=2 (state splits across cards) or llama.cpp 262K.
- Default 48K narr_TPS (48.82) slightly under chart's 55 reference —
bench variance + sample size n=3; not regression.
- Dual.yml / dual-dflash* not re-benched on v0.20; numbers carry forward
from dev205 (fp8 paths were not TPS-changed by the migration).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
46 lines
3.1 KiB
Plaintext
46 lines
3.1 KiB
Plaintext
|
|
========== NARRATIVE (prompt=65 chars, max_tokens=1000) ==========
|
|
=== warmups (3) ===
|
|
warm-1 wall= 46.30s ttft= 29048ms toks=1000 wall_TPS= 21.60 decode_TPS= 57.97
|
|
warm-2 wall= 17.87s ttft= 150ms toks=1000 wall_TPS= 55.97 decode_TPS= 56.44
|
|
warm-3 wall= 18.14s ttft= 153ms toks=1000 wall_TPS= 55.14 decode_TPS= 55.61
|
|
|
|
=== measured (5) ===
|
|
run-1 wall= 18.42s ttft= 153ms toks=1000 wall_TPS= 54.30 decode_TPS= 54.75
|
|
run-2 wall= 18.10s ttft= 148ms toks= 950 wall_TPS= 52.49 decode_TPS= 52.92
|
|
run-3 wall= 19.29s ttft= 150ms toks=1000 wall_TPS= 51.84 decode_TPS= 52.25
|
|
run-4 wall= 18.28s ttft= 148ms toks=1000 wall_TPS= 54.71 decode_TPS= 55.15
|
|
run-5 wall= 18.77s ttft= 149ms toks=1000 wall_TPS= 53.28 decode_TPS= 53.71
|
|
|
|
=== summary [narrative] (n=5) ===
|
|
wall_TPS mean= 53.32 std= 1.20 CV= 2.3% min=51.84 max=54.71
|
|
decode_TPS mean= 53.76 std= 1.22 CV= 2.3% min=52.25 max=55.15
|
|
TTFT mean= 150ms std= 2ms min=148ms max=153ms
|
|
|
|
========== CODE (prompt=78 chars, max_tokens=800) ==========
|
|
=== warmups (3) ===
|
|
warm-1 wall= 12.13s ttft= 156ms toks= 800 wall_TPS= 65.98 decode_TPS= 66.84
|
|
warm-2 wall= 7.05s ttft= 147ms toks= 481 wall_TPS= 68.26 decode_TPS= 69.71
|
|
warm-3 wall= 11.38s ttft= 155ms toks= 800 wall_TPS= 70.32 decode_TPS= 71.29
|
|
|
|
=== measured (5) ===
|
|
run-1 wall= 11.34s ttft= 154ms toks= 800 wall_TPS= 70.52 decode_TPS= 71.49
|
|
run-2 wall= 10.32s ttft= 150ms toks= 721 wall_TPS= 69.84 decode_TPS= 70.87
|
|
run-3 wall= 11.33s ttft= 157ms toks= 800 wall_TPS= 70.63 decode_TPS= 71.62
|
|
run-4 wall= 9.63s ttft= 150ms toks= 663 wall_TPS= 68.86 decode_TPS= 69.95
|
|
run-5 wall= 6.76s ttft= 152ms toks= 463 wall_TPS= 68.47 decode_TPS= 70.04
|
|
|
|
=== summary [code] (n=5) ===
|
|
wall_TPS mean= 69.66 std= 0.97 CV= 1.4% min=68.47 max=70.63
|
|
decode_TPS mean= 70.79 std= 0.78 CV= 1.1% min=69.95 max=71.62
|
|
TTFT mean= 153ms std= 3ms min=150ms max=157ms
|
|
|
|
=== GPU state ===
|
|
0, 95 %, 22220 MiB, 24576 MiB, 229.42 W, 62
|
|
1, 0 %, 4 MiB, 24576 MiB, 21.48 W, 40
|
|
|
|
=== Last 3 SpecDecoding metrics ===
|
|
(APIServer pid=1) INFO 05-01 16:46:34 [metrics.py:101] SpecDecoding metrics: Mean acceptance length: 3.53, Accepted throughput: 49.90 tokens/s, Drafted throughput: 59.10 tokens/s, Accepted: 499 tokens, Drafted: 591 tokens, Per-position acceptance rate: 0.970, 0.858, 0.706, Avg Draft acceptance rate: 84.4%
|
|
(APIServer pid=1) INFO 05-01 16:46:44 [metrics.py:101] SpecDecoding metrics: Mean acceptance length: 3.59, Accepted throughput: 51.09 tokens/s, Drafted throughput: 59.09 tokens/s, Accepted: 511 tokens, Drafted: 591 tokens, Per-position acceptance rate: 0.964, 0.883, 0.746, Avg Draft acceptance rate: 86.5%
|
|
(APIServer pid=1) INFO 05-01 16:46:54 [metrics.py:101] SpecDecoding metrics: Mean acceptance length: 3.54, Accepted throughput: 49.80 tokens/s, Drafted throughput: 58.80 tokens/s, Accepted: 498 tokens, Drafted: 588 tokens, Per-position acceptance rate: 0.954, 0.847, 0.740, Avg Draft acceptance rate: 84.7%
|