Root cause of club-3090 #548, pinned from the stack trace: A1's promotion accidentally committed the whole torch_compile cache (3,667 files, ~20 MB) compiled on THIS sm_86 rig, and the compose warm-start mount fed it to every fresh pull. On sm_120 the AOT graph (baked against Marlin-processed FP8 layers -> reads layer.workspace) loads onto Cutlass-processed layers -> AttributeError -> restart loop. vLLM's AOT cache key doesn't include arch/kernel selection, so the cross-arch hit is silent (upstream issue to follow; UPSTREAM.md row updated with the pinned mechanism). Fix = the house pattern every other model already has: cache contents gitignored (cache/.gitignore + README), directory kept for the mount, local files untouched (our warm-start intact). Fresh users pay one ~60-90s compile on first boot and warm-start locally thereafter -- against THEIR OWN silicon's kernel selection. Also retires the recurring dirty-tree noise from best_config files updating during our own runs. Full scripts gate 65/65. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EfF565T9eSLaqGzidyJ1Pm
87 KiB
Upstream tracker
Issues and PRs in upstream repos that affect this stack — what we depend on, what we've filed, what unblocks for us when each lands.
This file is the single source of truth for upstream status. When you file or notice an upstream issue / PR / commit relevant to club-3090, add a row here. When status changes (closed, merged, propagated), update it. Don't scatter the same link across multiple docs without coming back here first.
If you're adding a new compose that depends on an unmerged upstream patch (volume-mount of a fork, monkey-patch script), it MUST link to a row in this file so future readers know when the workaround can drop.
How rows work
Each row covers one upstream link with: title • status • our dependency / impact • workaround (if any).
Status vocabulary:
- 🟢 Landed — merged upstream + propagated to our pinned versions (pin-bump done)
- 🔵 Merged, awaiting propagation — merged upstream but our nightly / commit pin hasn't picked it up yet
- 🟡 Open / in review — PR open, no merge yet; we depend on it landing
- 🟠 Open / blocked or stalled — PR exists but progress stalled
- 🔴 Open, no PR yet — issue acknowledged but no fix in progress (us or upstream)
- ⚫ Workaround locally, no plan to merge — fixed in our patches, upstream not pursuing
- ✅ Resolved — closed and resolved (kept for historical context)
- ❌ Closed without fix — closed, won't fix, kept for context
Active follow-ups (next-week revisit queue) 🗓️
Items deferred for review next week (week of 2026-05-10). Audit at the start of that week — most should be either ready to action or have new upstream signal worth re-evaluating.
| Item | Why deferred | Trigger to revisit |
|---|---|---|
| 🟡 Sandermage/genesis-vllm-patches#22 — PN59 streaming-GDN doesn't engage on chunked-prefill (single-card 24 GB Cliff 2b stays open under v7.72.2). Filed 2026-05-05 with reproducer + 4 fix proposals. | Awaiting Sander's review. The cleanest of our 4 proposals is making has_no_chunk_metadata rejection optional (env-gated), letting single-seq chunked-prefill take the streaming path. |
Sander posts a candidate fix or comment on the issue; or we run a one-line A/B if he requests it. Until resolved, single-card 24 GB long-context users should run dual.yml / dual-turbo.yml / llamacpp/default. |
2db18df → f2147ad |
✅ Done 2026-05-05 — bumped to 7b9fd319 (v7.72.2) on branch v7.72.2-uplift. Drops patch_inputs_embeds_optional.py (PN35 native), patch_pn30_dst_shaped_temp_fix.py (PN30 v7.68), patch_pn25_genesis_register_fix.py (PN25), patch_tolist_cudagraph.py (P78), patch_workspace_lock_disable.py (PN34), patch_pr40798_workspace.py (research artifact). |
— |
| Superseded by v7.72.1 P68 auto-skip + v7.72.2 PN70 schema-subset filter — both ship default-aware behavior. Closed #57 along the way. | — | |
✅ Closed-superseded 2026-06-12 — NOT rebased. Upstream merged #45295 (mgoin) "Consolidate Marlin thread-tile padding across all dense Marlin paths" — one shared marlin_padded_nk mechanism covering WNA16/AWQ/GPTQ + FP4/FP8, which subsumes our per-case W4A16 sub-tile-n pad (issue #40354). Closed our PR in its favor (it also explicitly replaces the per-dtype #43806/#43910 — maintainers chose one general fix). |
✅ DONE 2026-07-02 — #45295 is native in v0.24.0, which vllm-stable now pins, so the vendored vllm-marlin-pad overlay is retired. De-registered in patches.yml (upstream.status: merged, load_bearing_when: [], delivery: none); no live compose mounts it (archive-only). See the #40361 row below. |
|
✅ vLLM stable pin: v0.22.0 → v0.24.0 — BUMPED 2026-06-30 (vllm-stable engine) |
Done. vllm-stable install.spec bumped v0.22.0 → vllm/vllm-openai:v0.24.0 (the first release that includes #45295), so the bump also retires the vllm-marlin-pad overlay (native now). 10 engine=="vllm-stable" compose :- defaults bumped; arch_patches.yml carries vllm-stable@v0.24.0 loads:true rows (v0.22.0 kept as history); pin-asserting fixtures updated (gemma asserts left on v0.22.0). Live-validated on 2× 3090 (caps 370/420 W): TP=2 device-selection intact (v0.24.0 dropped internal CUDA_VISIBLE_DEVICES), MTP accept 3.51, KV 622K/2.37×, verify-full 8/8, verify-stress NIAH→240K, soak-continuous PASS, bench 70/93 ≈ par, quality toolcall 11/15 · instructfollow 15/15. Full suite green (59/60; the 1 = worktree-fixture-absent test-submit-bench, green on master). |
Done — follow-ups: vllm-gemma-stable (#40391/#42006) consolidation onto v0.24.0 = separate PR (gated on the gemma KV + parser re-validation); the vllm-marlin-pad patches.yml entry is now de-registered (✅ 2026-07-02 — upstream.status: merged, load_bearing_when: [], delivery: none; archive-only, no live compose mounts it); streaming tool-calls validated clean on v0.24.0 → the qwen3coder deferred-commit sidecar (#72) is droppable (verify on its own). |
| 🟡 beellama — validate sm_89/sm_120; drop our unofficial image when Anbeeld's official lands | beellama is a registry engine (beellama-local profile + two composes). beellama/dflash (Qwen3.6-27B) was promoted to the single-card DEFAULT on 2026-05-30 (⚠️ caveats + DEFAULTS[(qwen3.6-27b, beellama, single)]; it is #1 in ENGINE_PREFERENCE[single], so the resolver now picks it over ik-llama). Basis: code-throughput leader (~100 TPS vs ik 69), slight 8-pack quality edge (107 vs 99 think-off), output-lossless DFlash, 160K validated ceiling. beellama/gemma-dflash (Gemma-4-31B) was likewise promoted to the single-card DEFAULT (⚠️ caveats + DEFAULTS[(gemma-4-31b, beellama, single)]) — it is the only viable fast single-card Gemma-4 path (vLLM FA-walled at head_dim=512, ik-llama walls ~24K), so Gemma single had no default before; re-point to the no-fork mainline path when llama.cpp#23398 (Gemma-4 MTP) merges. Both default to our unofficial multi-arch image ghcr.io/noonghunna/beellama-cpp:multiarch-b9459-07ac3ce — a club-3090 build of Anbeeld/beellama.cpp (MIT; .devops/cuda.Dockerfile with CUDA_DOCKER_ARCH="86;89;120" + -DGGML_CUDA_FA_ALL_QUANTS=ON) covering sm_86 / sm_89 / sm_120 = RTX 3090 / 4090 / 5090. sm_89 / sm_120 are COMPILED but UNVALIDATED — only sm_86 / 3090 is verified on our rig. CUDA_VISIBLE_DEVICES/WSL device-selection interaction, the compose pins GPU 0). Root-caused the 50-series gap while investigating: the CI builds on the Dockerfile-default CUDA 12.4, whose nvcc cannot target sm_120 (no cubin, max PTX compute_90) → every official tag (preview AND stable) lacks Blackwell. Filed Anbeeld#85 (CUDA_VERSION 12.8.1 + explicit arch list incl 120). Our multiarch snapshot is v0.3.0-feature-level and unmaintained (predates KVarN + the v0.3.1 fixes) — interim 50-series path only; self-build recipe (2 build-args, FA_ALL_QUANTS now hardcoded upstream via #48) verified + documented in engines/beellama-local.yml notes. |
(1) A 4090 / 5090 user confirms boots + verify (numbers-from-your-rig) → drop the "unvalidated" caveat. (2) |
🟢 beellama multi-GPU DFlash — dual composes RELEASED 🧪 experimental on v0.3.0 (Anbeeld's official image, PR #296); promote → ⚠️/✅ on a tagged stable release + prose recovery |
Evaluated dual-card (layer-split, 262K) Gemma-4 compose 2026-05-31 (broken on our pinned 07ac3ce: decode -1, accept 0.357, ~24/38 TPS, --device-draft ggml_abort). 2026-06-01: built the v0.3.0 dev branch (efe856397) + validated on 2× 3090 — multi-GPU DFlash is FIXED. No decode failed -1 (boot logs drafter=1 devices; enabling GPU cross ring); --spec-draft-device CUDA0 no longer ggml_aborts. Code DFlash up to ~157 TPS dual (≈4.2× the ~37 no-spec baseline), accept 0.33–0.64. Two residual issues, neither multi-GPU: (a) ~efe856397/e0663be/63abcd3) shows DFlash prose net-positive everywhere: qwen single +27% / dual +52%, gemma single +28–31%. The same efe856397 image we logged at ~0.07 AR now reads ~0.32 at the same tok/s → it was an AR over-read (adaptive-DM AR is noisy/prompt-dependent — Anbeeld's #288 caution) + a wrong no-spec baseline (real dual-Q8 no-spec = 23.4, not the ~37 we used). New adaptive-DM HEAD 63abcd3 is neutral. See learnings/qwen3.6-27b.md + gemma-4-31b.md 2026-06-03; (b) MTP unsupported — --spec-type mtp + radamanthys-assistant GGUF fails: unknown model architecture: gemma4_mtp (distinct from CLOSED #36 = garbage MTP output on Qwen, not arch-load). Built + published multi-arch image ghcr.io/noonghunna/beellama-cpp:multiarch-v0.3.0-efe856397 (same public package as stable b9459) + filed PR Anbeeld#48 (FA_ALL_QUANTS Dockerfile fix + CUDA→GHCR nightly workflow). 2026-06-01 RELEASE (PR #296): flipped the dual gemma compose ⏸️→🧪 experimental + shipped 3 new Q8_K_XL duals (beellama/qwen-mtp-dual :8064 embedded-MTP, beellama/qwen-dflash-dual :8065 @ 262K, beellama/gemma-q8-dflash-dual :8066 @ 192K) for community v0.3.0 testing (#288). BEELLAMA_IMAGE is now injected centrally from engines/beellama-local.yml install.spec, repointed at Anbeeld's OFFICIAL ghcr.io/anbeeld/beellama.cpp:server-cuda-v0.3.0-* commit tag (was our self-hosted snapshot); bump via scripts/beellama-pin-bump.sh. 2026-06-12 PIN BUMP: install.spec → server-cuda-preview-v0.3.2-317c65e27e1e (v0.3.2 PREVIEW — commit-suffixed for an immutable pin; adds experimental KVarN KV-compression). Maintainer chose the preview over the v0.3.1 stable for the newer build; validated single-card (q5ks-dflash, GPU0) — boots on the preview image, verify-full all-pass (Paris/tool_calls/streaming/thinking), prose coherent, DFlash spec-dec active (acceptance ~0.12–0.17 on short tasks, not collapsed). Composes STAY 🧪 (preview = rolling pre-release, not stable). Findings: #288, beellama #39. |
Anbeeld tags a STABLE release → promote 🧪→⚠️/✅, keep BEELLAMA_IMAGE on upstream's image (now done), retire our snapshot fallback. The first stable tag now exists: v0.3.1 (server-cuda-v0.3.1, non-prerelease — Qwen3 MTP post-norm + CUDA KV-quant fixes). Repoint install.spec v0.3.2-preview → v0.3.1 stable to UN-PARK (#455) once it passes the full gate (verify-stress + bench + soak + quality). (The "prose-recovery" half of this gate was DROPPED 2026-06-03 — DFlash prose is net-positive on tok/s; the regression was a measurement artifact, not real.) |
See the platform-specific tables below for the rows these reference.
Pinned images
What container image each compose pins, why each pin exists, and which pins
are candidates for retirement when their reason resolves. This section answers
"where the engine pins live" and drives the work in
NIGHTLY_BUMP_RUNBOOK.md.
Run bash scripts/maintenance/list-image-pins.sh for a live snapshot.
vLLM compose files do not hardcode full nightly image tags. They use:
image: ${VLLM_IMAGE:-vllm/vllm-openai:nightly-${VLLM_NIGHTLY_SHA}}
scripts/launch.sh, scripts/switch.sh, and estate boot resolve
VLLM_NIGHTLY_SHA from scripts/lib/profiles/engines/<engine-id>.yml → install.spec. VLLM_IMAGE is a full-image override for users who want to pin a
specific upstream nightly, or run a current image when a pinned nightly has been
purged (e.g. VLLM_IMAGE=vllm/vllm-openai:latest).
| Pin source | Composes using it | Reason for pin | Retirement candidate? |
|---|---|---|---|
scripts/lib/profiles/engines/vllm-nightly-mtp.yml → vllm/vllm-openai:nightly-01d4d1ad... (PURGED) |
🗑️ DEPRECATED 2026-06-05 (#254) — Genesis/TQ3 composes archived to compose/_archive/, zero registry users |
Genesis MTP nightly; image 404 on Docker Hub + Genesis path on hold upstream. Retained as the genesis_equipped test anchor (required_genesis:true). |
Do not bump; revive only if Genesis re-anchors on a post-#42521 nightly. |
scripts/lib/profiles/engines/vllm-nightly-dflash.yml → vllm/vllm-openai:nightly-e47c98ef... |
DFlash vLLM composes | DFlash overlay baseline. | Bump this YAML after DFlash overlay drift is revalidated. |
scripts/lib/profiles/engines/vllm-nightly-full.yml → vllm/vllm-openai:nightly-e47c98ef... (PURGED) |
🗑️ DEPRECATED 2026-06-05 (#254) — vllm/dual-int8 archived, zero registry users |
Experimental full-overlay nightly; image 404 on Docker Hub. | Do not bump; retained for revival. |
scripts/lib/profiles/engines/vllm-stable.yml → vllm/vllm-openai:v0.22.0 (overlay-free) |
Qwen3-Next vLLM: vllm/minimal, vllm/dual (27B), vllm/qwen-35b-a3b-dual (35B-A3B); generic dense; uncurated derived-emission bases |
Stock immutable image, no vendored overlays — the load-bearing distinction (CONTRACT-5) that lets it be a derived-emission base. Successor to the purged vllm-nightly-clean (#254). |
Bump install.spec only; never the compose literals (drift-guarded by test-compose-image-drift). |
scripts/lib/profiles/engines/vllm-gemma-stable.yml → vllm/vllm-openai:v0.22.0 (overlay-carrying; split off vllm-stable in #324) |
Gemma 4 31B vLLM: vllm/gemma-bf16-mtp (v0.22.0, #42006 tool-parser overlay only, 131K) + vllm/gemma-int8-mtp (v0.22.0 as of 2026-05-31, #40391 + #42006 overlays, 262K); vllm/gemma-mtp-tp1 DEPRECATED 2026-05-31 (fp8 dead on Ampere) |
Both gemma duals now on immutable v0.22.0, both carry the #42006 streaming-multi-tool fix (lean diff-apply); gemma-int8-mtp ALSO carries the #40391 per-head-KV overlay. #40391 DID need a real rebase onto v0.22.0 — the old v0.21.0-era full-module copies ImportError'd on v0.22.0 (cannot import get_kv_cache_spec_kind — v0.22.0 added that symbol; the stale copies lacked it). Rebased = v0.22.0-stock + the ~240-line #40391 delta (2 club-3090 resolutions: hand-merged attn_utils is_mla() branch onto v0.22.0's refactored _reshape_kv_cache; kept kv_cache_utils's replace import that the PR removed but v0.22.0 still uses). Re-delivered LEAN (boot-time diff-apply install_script, not 7 full-module mounts). Live-validated 2026-05-31: boots, int8_per_token_head KV inits, KV pool 447K @ 262K, bench 95.7/125.8, coherent generation. The 2 other overlays int8 used to carry (#41800, #41991) are in stock v0.22.0; #42006 is vendored on both duals (see its row). |
Drop the #40391 overlay entirely when PR #40391 merges + lands in a release. UPDATE 2026-07-02 (v0.24.0 consolidation): the 31b DEFAULT moved to overlay-free bf16 (vllm/gemma-31b-dual, stock v0.24.0, ~224K). On v0.24.0 int8-PTH allocates 262K but SILENTLY craters recall past ~32K — #40391 is NOT native there (both cyankiwi + w4a16 crater identically; the SAME cyankiwi weights recall clean on v0.22.0+#40391). The #40391 + #42006 overlays and vllm-gemma-stable v0.22.0 STAY — still used by the 26b-a4b-single (gemma-a4b copy) + the now-deprecated 31b composes (gemma-int8-mtp / gemma-bf16-mtp / qat-w4a16). The 262K int8-PTH 31b path returns overlay-free when #40391 merges. |
ghcr.io/ggml-org/llama.cpp:server-cuda |
2 (Qwen 3.6-27B llama-cpp) | Stable tag, no hash drift on upstream side. No patches mounted. | Not a retirement candidate — drift-free. Capture digest if reproducibility matters. |
Retirement workflow: see NIGHTLY_BUMP_RUNBOOK.md.
Retired pins
| Pin | Former use | Notes |
|---|---|---|
vllm/vllm-openai:nightly-01d4d1ad |
Qwen 3.6-27B Genesis baseline | Retired from default profile pins; retained as historical validation context. |
vLLM (vllm-project/vllm)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
(to file) — torch_aot_compile cache key ignores GPU arch / resolved quant kernel → cross-arch cache load crashes (AttributeError: 'QKVParallelLinear' object has no attribute 'workspace') |
🔴 To file upstream (mechanism pinned 2026-07-04; draft pending maintainer review) | Root cause pinned by the stack trace in club-3090 #548: the crash fires inside torch/_dynamo/aot_compile.py __call__ immediately after Directly load AOT compilation from path …/torch_aot_compile/61bda05f… — an AOT graph compiled on our sm_86 rig (Marlin-processed FP8 layers → graph bakes layer.workspace reads) loading onto sm_120 (Cutlass-processed layers → no workspace attr). vLLM's AOT cache key (61bda05f…) is IDENTICAL across the two arches despite different per-layer kernel processing — cross-arch cache hits should be impossible. The trigger on user rigs was us: A1's promotion accidentally git-shipped the whole compiled cache (3,667 files), and the compose warm-start mount fed it to every fresh pull (fixed — cache un-tracked, house .gitignore pattern applied). |
Two independent unblocks: (a) delete the mounted torch_compile cache contents → fresh AOT compile against the rig's own kernel selection (native Cutlass FP8 on sm_89+ — the preferred path, likely faster than #567's numbers); (b) VLLM_TEST_FORCE_FP8_MARLIN=1 forces layer processing to match the sm_86-baked graph (validated in #567, but costs native FP8 GEMMs). Upstream fix = include compute capability / resolved kernel config in the AOT cache key. |
#42609 — custom_all_reduce IPC handle fails with expandable_segments |
🟡 Open / local fix active | PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True crashes vLLM custom all-reduce at graph-buffer registration — custom_all_reduce.cuh:455 'invalid argument' (cudaIpcGetMemHandle on a cuMemMap VA range is not a valid IPC source). Upstream's title scopes the repro to DP>1 AND TP>1, but the same crash class fires on TP-only startup paths (#42583 Qwen3.5+LoRA TP=4, #45604 MiniMax TP=4). On our 2-card NVLink / patched-PCIe-P2P rigs custom AR is ON at TP=2, so the expandable_segments:True that dual composes inject for the PCIe (custom-AR-OFF) path crashes boot when P2P is enabled. Earlier signal: JusefPol's NVLink boot crash (club-3090 PR #31). |
scripts/detect_nvlink.sh strips only the expandable_segments token from PYTORCH_CUDA_ALLOC_CONF on the P2P / custom-AR path (other knobs like max_split_size_mb preserved); the PCIe path keeps it. Guard: scripts/tests/test-detect-nvlink-alloc-conf.sh. Drop when #43923 (the open in-engine fix targeting custom_allreduce + expandable_segments) merges and lands in our pin. NB: #40812 (merged v0.20.1, already in our v0.22.0 pin) only auto-disables expandable_segments around the cumem sleep-mode pool, not the custom_all_reduce path — it does not cover this crash. |
| #45163 — DiffusionGemma (vLLM's first discrete-diffusion LM) | 🟢 #45163 MERGED (in v0.24.0) — on stock v0.24.0 + 3 vendored Ampere/TP fixes | The DiffusionGemma 26B-A4B arch (DiffusionGemmaForBlockDiffusion, V2 model runner) is in no released/nightly vLLM and only in this unmerged draft — BUT vLLM publishes an official vllm/vllm-openai:gemma image (a stock build of the dgemma branch commit 74b5964f) that has it baked in. Needed to serve vllm/diffusiongemma-dual (the dLLM studio model). |
Engine vllm-diffusion-gemma pins vllm/vllm-openai:gemma BY DIGEST (sha256:9c719fc0…, purge-resistant) + bind-mounts 3 fix files that are NOT upstream (vLLM tests H100/B200 + TP=1): models/diffusiongemma-26b-a4b/vllm/patches/gemma-image-fixes/ — marlin.py+marlin_utils_fp8.py (sm_86 fp8 Marlin sub-tile-K pad; :gemma clean dies in warmup on the K=352/1056 wall) + diffusion_gemma.py (TP-vocab soft-embed + dtype fix). Status 🧪 experimental (no DEFAULTS row; launch --force) — soak-continuous PASS (0-growth, 0/25 silent-empty) + 8-pack 100/150, but held at experimental because the arch is in no release (unmerged + mutable :gemma tag) and the dLLM SSE/streaming-NIAH checks can't cleanly gate. Supersedes the former 123-file sideload-on-nightly (PR #358 → simplified here: same perf, 120 fewer files). Re-pin trigger: vLLM re-pushes :gemma → re-pin the digest + rebase the 3 fixes. Promote trigger (decision 2026-06-11): when #45163 merges into a pinnable release → re-pin to it, re-run the gate, flip 🧪 → ⚠️ caveats (documented caveats: long-ctx exact-recall degrades past ~30K, block-streaming SSE, dual-only) — see club-3090 task. Drop trigger: the K-pad lands upstream (our #40361 / an Ampere Marlin fix) AND the TP-vocab fix merges into :gemma → mount nothing. (2026-06-12: the K-pad half landed on vLLM main via #45295 — general marlin_padded_nk; our #40361 closed-superseded. Drop still gated on a stock image that includes #45295 + the TP-vocab merge.) UPDATE 2026-07-02: #45163 MERGED (shipped in v0.24.0) → engine vllm-diffusion-gemma bumped off the :gemma branch digest to stock vllm/vllm-openai:v0.24.0 (arch native). The 3 Ampere/TP fixes STAY vendored + re-validated on stock v0.24.0 (boot + coherent TP=2). #45295's marlin-pad (native in v0.24.0) is the dense/linear path — it does NOT cover DiffusionGemma's marlin_moe_wna16 sm_86 fp8 K-pad (native-no-mounts test still crashed there), so marlin.py+marlin_utils_fp8.py remain required; native diffusion_gemma.py is still TP=1 → the TP-vocab fix remains required. Compose also gained the --model flag fix (v0.24.0 positional-arg change). Stays 🧪 (still 3 vendored fixes + no soak on v0.24.0). |
| club-3090 #451 / #250 / #167 — Gemma 4 31B vLLM nightly prune + stable repoint | 🟢 Local cleanup in progress | Gemma 31B vLLM composes referenced purged Docker Hub nightlies (bf610c2f, e47c98ef), recreating the same fresh-user failure class as #250/#167. |
Active set pruned from 9 to 3 and repointed to immutable vllm/vllm-openai:v0.21.0 via vllm-gemma-stable; Qwen remains on its separate nightly/Genesis profiles. Live-check DONE 2026-05-31 on stock v0.22.0: vllm/gemma-mtp-tp1 DEPRECATED — fp8 KV is hardware-impossible for Gemma 4 on Ampere sm_86 (gemma4 attention asserts kv_cache_dtype ∈ {fp8, fp8_e4m3, nvfp4} → rejects fp8_e5m2; fp8/fp8_e4m3 need the fp8e4nv kernel sm_86 lacks; nvfp4 Blackwell-only). bf16 boots but is dominated by the beellama/gemma-dflash single-card default. Single-card Gemma → beellama; vllm/gemma-bf16-mtp (dual) bumped to v0.22.0, validated. |
#35936 — tool_choice="required" falls back to configured tool parser |
🟡 Open / local overlay active | Qwen3-Coder with --tool-call-parser qwen3_coder emits XML-style tool calls. On pinned nightly 1acd67a79, non-streaming tool_choice="required" validates JSON only, bypasses the configured parser, and returns tool_calls=[]. MLS-Bench hits this when thinking.enabled=false. |
Vendored overlay: models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md. Drop when #35936 or equivalent lands in our pinned image. |
#41800 — truncate_prompt_tokens kwarg on get_max_tokens() |
✅ Merged upstream 2026-05-06 at d5b31c95 / local overlay active on pre-fix engine pins |
opencode (and other agentic clients sending truncate_prompt_tokens) fail with HTTP 400 get_max_tokens() got an unexpected keyword argument on engines pinned to 01d4d1ad (Genesis MTP), e47c98ef (DFlash, full). All three SHAs predate d5b31c95. vllm-nightly-clean (bf610c2f, post-fix) doesn't need the overlay. Tracking issue: club-3090 #139. Triggered by club-3090 #138 (SEVENID's opencode failure). |
Vendored overlay: models/qwen3.6-27b/vllm/patches/vllm-pr41800-truncate-prompt-tokens/README.md. Wired into 18 affected composes (every compose routing through vllm-nightly-(mtp|dflash|full)). Install script has upstream-fix detection: no-ops cleanly when run against a post-d5b31c95 nightly. Drop trigger per engine: bump each affected engine's pin past d5b31c95. For vllm-nightly-mtp that requires Genesis v7.73.x; for vllm-nightly-dflash and vllm-nightly-full it requires re-validating PR #41703 and PR #42102 overlays on a newer base. |
| #42006 — Gemma 4 MTP streaming multi-tool calls | 🟡 Open upstream / VENDORED on BOTH gemma duals (v0.22.0, lean diff-apply) | Briefly dropped 2026-05-31 on a "minimal surface" call, then RE-INSTATED the same day once the bug was reproduced live on stock v0.22.0: a streamed response with 2+ tool calls drops the arguments of every call except the last (get_weather(Tokyo)+get_weather(London)+get_time(Paris) → get_weather()/get_weather()/get_time({"city":"Paris"}) — only the last kept args). Non-streaming returns all three correctly → confirmed streaming-parser bug. #41991 (formerly stacked) is now in stock v0.22.0 → dropped; only #42006 remains. |
Vendored as a single-file boot-time diff-apply: models/gemma-4-31b/vllm/patches/vllm-pr42006-v0.22.0/, carried by both gemma-bf16-mtp + gemma-int8-mtp (re-introduces the only overlay on bf16-mtp). Drop when #42006 merges + lands in the pin: gh api repos/vllm-project/vllm/pulls/42006 --jq '.state, .merged_at'. |
| #45413 — ParserEngine core (declarative streaming parser) | ✅ Merged 2026-06-15 / NOT in v0.23.0 (tagged before merge) | Declarative, config-driven streaming parser engine replacing hand-rolled per-model logic with a shared state machine. Handles reasoning/content/tool-call transitions, incremental JSON argument extraction, and detokenizer hold-back recovery via a single ParserEngineConfig. Potential to eliminate our vendored qwen3coder deferred-commit sidecar (#72) and the Gemma4 parser overlays — but likely does NOT cover the MTP spec-decode-streaming path (#39598). |
Wait for v0.24.0 (first release to include it). Adoption plan: boot new image on a test branch → run stream-toolcall-probe.py for Qwen3.6-27B (check #39598 residual) → run DiffusionGemma streaming tool replay (check #443 residual) → if clean, drop vendored overlays and update composes + patches.yml; if broken, keep overlays and file upstream issue. Flagged by @troymroberts in discussion #451. |
| #45588 — Gemma4 → ParserEngine migration | ✅ Merged 2026-06-15 / NOT in v0.23.0 (tagged before merge) | Migrates Gemma4 (incl. DiffusionGemma) reasoning + tool parsers to the new ParserEngine framework. Replaces the hand-coded gemma4_reasoning_parser.py + gemma4_tool_parser.py with a single declarative gemma4.py state machine. steamEngineer tested on pinned :gemma (PR #443 ablation): non-streaming tool_calls OK; streaming leaks raw <|tool_call>…<tool_call|> tokens into content under Hermes (block-canvas SSE chunking). Hybrid (engine reasoning + legacy tool mount) also broken. |
Keep legacy gemma4_*_parser.py overlays until upstream fixes block-canvas streaming tool extraction. See models/diffusiongemma-26b-a4b/vllm/patches/gemma-image-fixes/patch_docs/diffusionGemma_streaming_tool_regression.md for full ablation. Experiment preserved on fork branch experiment/dgemma-parser-engine-45588. Adopt trigger: lands in a release + streaming tool extraction clean on DiffusionGemma block-canvas. |
| #40361 — Marlin pad-sub-tile-n | 🟣 Closed 2026-06-12 — superseded by #45295 (mgoin, merged 2026-06-12: general marlin_padded_nk tile-pad across all dense Marlin paths incl. WNA16/AWQ/GPTQ — subsumes our per-case W4A16 sub-tile-n pad; explicitly the consolidated alternative to per-dtype #43806/#43910). |
✅ #45295 native in v0.24.0 — vllm-stable pins v0.24.0, so the tile-pad is stock. No live compose mounts the vendored patch (archive-only; the earlier dual/awq-bf16-int4/int8 / multi4/fp8/mtp / nex-n2-mini are all under _archive/). Patch de-registered 2026-07-02 (patches.yml: upstream.status: merged, load_bearing_when: [], delivery: none). |
Vendored mount: see models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md. |
#40807 — .tolist() cudagraph crash on continuation-prefill |
✅ Retired locally (2026-05-05 Genesis v7.72.2 bump) — Genesis ships P78 TOLIST_CAPTURE_GUARD as the equivalent fix. Currently disabled (=0) on tq3-mtp-genesis.yml after rebench-full leg 6 (2026-05-11) passed clean with it off — apparent root cause is now covered by Genesis PN34 (workspace-lock relax) + post-#41434 attention rework. Non-Genesis composes on 1acd67a79 pin run without any guard for this bug; unvalidated at long-context TurboQuant chunked-prefill (worth testing per cferra's vllm#41403 validation pass — see vllm#40798 row below). |
None active. Drop the Genesis env var permanently if a future v7.73.x rebench leaves it OFF without regression. | |
| #40798 + #42215 — share decode scratch workspace pre-CUDA-graph + decode-kernel warmup | 🟡 Open, validated cross-rig | Pair closes the AssertionError: Workspace is locked but allocation requires NMB crash that fires at turboquant_attn.py:_continuation_prefill for ≥48K-token chunked-prefill with TurboQuant KV. Independently validated on 2× 3090 sm_86 by cferra (vllm#41403 comment, 2026-05-12). |
Genesis PN34 WORKSPACE_LOCK_RELAX addresses the same symptom via a different mechanism (relax-lock vs reserve-before-capture). On non-Genesis composes (1acd67a79 pin) we currently have no guard — re-validate against this PR pair once they propagate to a nightly we pin to, then A/B PN34 vs upstream. |
| #40849 — MTP draft online-quant propagation | 🟡 Open / Genesis backport active | Closes Cliff 1 on FP8+MTP path (tools-text.yml). |
Genesis PN8 backport: GENESIS_ENABLE_PN8_MTP_DRAFT_ONLINE_QUANT=1. |
| #40914 — Sandermage K+1 verify routing | 🟡 Open, ❌ negative on our Qwen3.6-27B stack | Reframed 2026-05-11: the synthetic seq_lens K+1 route is not the P67-equivalent we need here. Local rebase on post-#41434 nightly made MTP acceptance look perfect (AL=4.0 / ~100%) but produced !-flood needle corruption plus tool/multi-turn timeouts. Dropping it improved verify-stress from 3/7 to 5/7, but TQ3/TQ4/k8v4 + MTP still fail long-context needles. |
Do not ship Genesis-free TQ+MTP on #40914 alone. Use dual/autoround-int4/tq3-nomtp.yml without Genesis, or dual/autoround-int4/tq3-mtp-genesis.yml with Genesis P67/P67b. |
#40334 — DFlash combine_hidden_states dtype mismatch |
🟡 Open | All dual-dflash*.yml need --dtype bfloat16 flag to work around. |
Composes set --dtype bfloat16. Drop when this lands. |
| #40382 — Gemma-4 + DFlash unservable on Ampere | 🟠 Open, no fix in progress | Blocks DFlash on Gemma-4 family. Not directly our problem (we serve Qwen3.6) but tracked because future model adds may hit it. | None — different attention backend selection. |
| #41559 — DFlash spec-decode incompatible with all KV cache quantization (seantechco, filed 2026-05-03) | ❌ OUR FIX PR #42102 CLOSED AS SLOP by @benchislett on 2026-05-15 (no comment, just closed-as-slop label). Issue #41559 still OPEN upstream. |
Local fix preserved: three-layer patch (4 files, +333/-35) on branch dflash-noncausal-kv-quant (commits cfb8f711 + 5cb61c60). (1) vllm/v1/core/kv_cache_utils.py partitions DFlash drafter specs into independent KV groups before unify, allocator extended to size isolated tensors by their own page_size; (2) vllm/model_executor/models/qwen3_dflash.py overrides drafter cache_dtype to "auto" when engine global is quantized; (3) vllm/v1/attention/backends/flash_attn.py FA metadata scheduler uses per-spec dtype when spec's kv_quant_mode is NONE. Validated locally on dual 3090 Ampere: Gemma 4 + z-lab DFlash drafter + INT8 PTH KV target boots HEALTHY at 65K, narrative 95.89 / code 168.09 TPS, AL 5.0-5.3 preserved, NIAH PASS at 32K, KV pool 149,345 tokens (4× lift). |
Vendor permanently. Patch lives at retired Gemma DFlash INT8 overlay (removed from the active tree in #451; recover from git history if revisited) and is baked into vllm-nightly-full + vllm-nightly-dflash EngineProfiles. Re-engagement with upstream NOT recommended (vLLM has hardened anti-AI-PR policy). Watch issue #41559 for any newer maintainer-blessed PR; drop our overlay then. |
| #40354 — Marlin TP=2 W4A16 < 64 | 🟢 Resolved upstream by #45295 (general tile-pad; the <64 WNA16 sub-tile-n case is now padded). Our #40361 closed-superseded. | ✅ Effective in v0.24.0 (vllm-stable pinned 2026-06-30); vendored patch de-registered 2026-07-02. |
See #40361 row. |
| #39931 — DeltaNet rollback support | 🔴 Open, architectural | Blocks all spec-decode (EAGLE / DFlash) on Qwen3-Next family across engines. The reason "speculative decoding doesn't work" on this stack. | Use MTP (no rollback needed) until this lands. |
| #40124 — related architectural | 🔴 Open | Pairs with #39931 for DeltaNet rollback. | Same as above. |
| #40880 — MTP × TQ × cudagraph cascade | ✅ Closed upstream issue, but not solved by direct upstream vLLM | Genesis P65 removed the CUDA-graph-specific failure mode; P67/P67b is the correctness path for K+1 multi-query TurboQuant attention. Round-4 testing showed --enforce-eager alone does not close TQ+MTP needles without P67-equivalent behavior. |
Use Genesis P67/P67b or disable MTP on TurboQuant. |
| #40831 — TQ × spec-decode corruption | ✅ Closed issue, open upstream gap | Our new matrix reproduces the same class across TQ3, TQ4, and k8v4 under MTP. TQ3 no-MTP passes 7/7, so the bug is the MTP x TurboQuant multi-query path, not precision. | Same as #40880. |
| #40798 — workspace-manager refactor | ❌ Negative result | Hypothesized fix for #40831 / #40880; backporting it (Probe 8) didn't resolve the bug. Kept for context — saved future time on the same dead end. | n/a |
#40875 — ngram spec-dec default prompt_lookup_min=2 corrupts tool-call output (Qwen3-class + structured output) |
🟡 Open / mitigated | Upstream-confirmed bug: the ngram (prompt-lookup) default of 2 mangles tool-call args; config-only fix is prompt_lookup_min=8. mgabor3141 (#178) independently reports a likely-distinct corruption mode under MTP-3 + --tool-call-parser qwen3_coder on chained agent calls — that mode, plus whether to flip the vLLM agent-facing parser default to qwen3_xml, is tracked separately (vLLM-only; llama.cpp/ik use the native template + --jinja, unaffected). |
prompt_lookup_min=8 set in composes that use ngram. |
| #41142 — Quentin-M streaming tool-call IndexError | 🟡 Open / Genesis backport active | Closes a streaming tool-call crash on Hermes / similar templates. | Genesis PN11 backport (auto-enabled where REC). |
#39056 — Qwen3-Next tool-calls inside <think> reasoning region lost |
🟡 Open / auto fixed on v0.22.0; required residual = #39598 | Tool-call XML inside <think>…</think> swallowed by the reasoning-parser before the tool-parser sees it. |
v0.22.0 status (verified 2026-06-13, club-3090#145): tool_choice=auto + streaming + thinking is clean on stock v0.22.0 — and the qwen3_coder→qwen3_xml swap is a no-op (A/B byte-identical, both parsers; club-3090 PR #400 closed on this). The earlier "qwen3_xml pivot, less <think>-sensitive" workaround is disproven. The only residual is tool_choice=required+thinking+streaming, which is MTP-gated → tracked at #39598. |
| #39598 — kotori-yan qwen3coder MTP streaming early-return | 🟡 Open / UN-MITIGATED on v0.22.0 (P64 retired) | MTP spec-decode drops streamed tool-calls (empty tool_calls[], finish_reason=stop) when MTP bundles the call in one delta. Was fixed by Genesis P64 (GENESIS_ENABLE_P64_QWEN3CODER_MTP_STREAMING=1); Genesis retired (#182 / #254) → P64 gone → the drop is back on stock v0.22.0 (shipped silently into the dual composes — invisible because benchlocal is non-streaming + verify-full's streaming check has no tools). Confirmed 2026-06-13 (scripts/stream-toolcall-probe.py): MTP n=3 dropped 13/20 tool_choice=required+thinking+streaming requests; no-MTP clean 0/20; parser-independent (qwen3_coder ≡ qwen3_xml, byte-identical drop set). auto+streaming is clean even with MTP. |
Mitigation: tool_choice=auto (clean w/ MTP) or disable MTP for required+streaming workloads. Durable fix: upstream — #45413 rewrites the parser and likely does NOT cover this spec-decode-streaming path; needs a P64-equivalent / MTP-streaming fix. |
qwen3coder tool-parser SSE-silence on prose <tool_call> (upstream #22975 closed-as-stale; reported on club-3090 as #72) |
⚫ Local workaround / upstream PR deferred until cross-rig validation lands | When the model's prose mentions the literal <tool_call> text (e.g. agent reasoning that describes the markup), extract_tool_calls_streaming flips is_tool_call_started=True permanently on either the special-token-id or the string match. Subsequent deltas return None; the serving layer skips them; SSE wire goes silent for 30-120s while tokens decode server-side and never reach the client. Verified bug still present in vLLM main as of 2026-05-07 (no deferred-commit guard in current source). Upstream issue #22975 reports a related symptom (<tool_call> markup remains as plain content) but was closed-as-stale 90+ days ago without a fix — different observed surface, likely shared root cause. |
models/qwen3.6-27b/vllm/patches/local/qwen3coder_tool_parser_deferred_commit.py runs after apply_all in the entrypoint of all 8 Genesis-equipped composes. Defers is_tool_call_started=True until <function= confirms within a 64-char slack window past the <tool_call> tag. Direct-cmd composes (dual.yml, dual-dflash*.yml, minimal.yml, multi4*.yml, carnice-bf16mtp.yml, qwopus-bf16mtp.yml) don't currently receive the sidecar — they have no entrypoint script. Plan: ship local sidecar → validate cross-rig → file upstream PR (with cross-rig evidence and the V2 deferred-commit logic) once the local fix has held up under multi-rig real-world traffic. |
| #40961 — Preserve max_seq_len in ubatch metadata during CUDA graph capture | 🟡 Open PR | Confirms the cap-leak pattern: cudagraph capture passes max_model_len as max_seq_len through ubatch metadata. PR is fixing a missing pass-through for SWA models (where seqlen=1 at capture broke kernel selection) — by establishing that max_model_len is what gets carried through capture metadata, it cements the source of Cliff 1's max-ctx-dependent FA2 workspace sizing. |
Stay at default 48K — see FA2 #1011 row + INTERNALS.md Cliff 1 mechanism. |
| #40069 — [Tracking] TurboQuant / HIGGS Attention follow-ups | 🟡 Open tracker | Umbrella tracking for TurboQuant + attention backend issues on our stack class. | Watch for cross-references when Cliff 1/2 work lands upstream. |
#25543 — [V0 Deprecation] Remove max_seq_len_to_capture |
✅ Merged 2025-09-24 | Important to know: the --max-seq-len-to-capture flag (commonly suggested as a Cliff 1 mitigation) does not exist in V1. Don't recommend it. |
n/a — flag removed. |
| #39226 — workspace-resize GPU memory leak fix | 🔵 Merged into v0.20.0; covered by sidecar | Strict WorkspaceManager.lock() semantics. After our 2026-05-01 v0.20 + Genesis v7.65 dev tip migration, the surfaces that locked at 0 MB on our config are largely covered by v0.20's revised TQ FA paths (#40092). For the residual cases, our local patch_workspace_lock_disable.py sidecar (mounted on every TQ3 compose) downgrades the strict assertion to a one-shot WARNING. P98 covers the same surface but auto-skips on v0.20 due to a drift-marker false-positive (filed as side-note, awaiting Sandermage marker fix). |
Drop the sidecar when Sandermage ships the marker fix that re-enables P98 on v0.20. |
| #40092 — TurboQuant FA3/FA4 prefill paths | 🔵 Merged into v0.20.0 | TQ + flash-attention 3/4 prefill support. Relevant if/when v0.20 unblocks for us — the FA varlen workspace allocator behavior may change under FA3/FA4 vs the FA2 path we currently hit. | Track. Re-evaluate the flash_attn_interface.py:300 cliff (Genesis #15) once v0.20 unblocks since FA3/FA4 may have different workspace semantics. FA3/FA4 not enabled on Ampere SM 8.6 anyway (Hopper+ only). |
| #40941 — TurboQuant share buffers | 🔵 Merged into v0.20.0 | Sandermage's bare_metal_27b_int4_TQ_k8v4.sh comments call out P98 as the workaround for "WorkspaceManager fix vs vllm#40941". Same WorkspaceManager class that vllm#39226 made strict. | Sandermage's P98 is the workaround — required for TQ k8v4 on hybrid. Worth a focused investigation: enable P98 on the v0.20-experimental compose to see if it also unblocks vllm#39226's path. |
#35975 — Skip inputs_embeds GPU buffer for text-only models ⭐ |
🟡 Open upstream / local backport active / Genesis PN35 lands same fix on dev f2147ad (2026-05-03) |
Frees ~444 MiB at boot on Qwen3.6-27B (both gpu_model_runner.py + llm_base_proposer.py call sites compound; PR claims ~64 MiB but our config has multiple residency points that benefit). Critical for Cliff 2 closure at 60K on TP=1 + 24GB — combined with mem-util 0.93, closes the late-stage 50 MiB activation peak. Diagnosed by ChatGPT/Codex as the missing margin; cross-rig validated 2026-05-02 PM. Upstream PR last updated 2026-03-13 (51d stale); Sandermage's PN35 is the practical replacement. |
patch_inputs_embeds_optional.py ships at compose-entrypoint time (mounted on long-text.yml and long-text-no-mtp.yml). Drops out when we bump GENESIS_PIN to dev tip — queued for next-week revisit. |
| #37429 — Hybrid Mamba/attention KV cache sizing | 🟡 Open | Could free more residency without trading mem-util. Larger/riskier than #35975 (architectural Mamba allocation change). Untested on this stack. | Not currently backported. Test on a separate branch when CI signals stabilize. |
| #37521 — Spec-decode warmup memory accounting | 🟡 Open | Profiling/KV sizing leaves less false headroom. Genesis PN33 already extends this beyond the original use_eagle() gate — so most of the surface is covered, but watch for upstream refinement. |
n/a — Genesis PN33 covers the path. |
| #36598 — Triton autotuner OOM on Qwen3.5/Qwen3-Next GDN layers (non-SM90 GPUs) | ✅ Closed 2026-03-12, fix shipped via #36599 | Original report of first-inference OOM during Triton autotuning on non-SM90 hardware. Closed because the warmup fix landed. Reading thread is useful context for understanding the GDN kernel autotuner pressure on our hardware class. | n/a — fix in our image. |
| #36599 — Warm up Triton autotuner for GDN layers during V1 profiling | ✅ Merged 2026-03-12 (in image SHA 7a1eb8ac) |
Adds _warmup_triton_kernels() at V1 profile phase. Warms with B=1, T=64 dummy tensors. Closes the boot-time first-inference autotuner OOM that #36598 reported. DOES NOT close Cliff 2b (multi-turn accumulated context) — the warmup uses T=64 but FLA kernels use do_not_specialize=["T"] so production T=4128 is the same autotune key, meaning runtime fragmentation isn't from missed autotune; it's from the per-shape Triton kernel binaries staying resident in CUDA context. Confirmed by Codex memo 2026-05-03. |
n/a — fix in image; doesn't help our remaining cliff. |
#36973 — _warmup_prefill_kernels leaks ~3.4 GiB despite empty_cache |
🟡 Open, RTX 5090-specific | jhsmith409's report — Triton autotuner cubin retention initially suspected but haosdent comment #18-19 traced the bulk to TMA overhead scaling with SM count (~22 MiB/SM × 170 SMs on 5090 = 3.7 GiB). Closed via #37700 (TMA-disable for SM12x). Doesn't apply to Ampere SM86 — no TMA hardware. Useful context though: thread comment #5 explicitly notes Triton autotuner keeps all variants loaded; empty_cache() only releases PyTorch's caching allocator, not CUDA-context cubins. |
n/a — RTX 3090 doesn't have TMA. |
| #37700 — Fix FLA Hopper/TMA misclassification on SM12x desktop Blackwell | 🟡 Open / closes #36973 for SM12x | Uses shared-memory threshold instead of major >= 9 checks for TMA path selection. SM12x desktop Blackwell only — RTX 5090, DGX Spark GB10. Doesn't apply to Ampere SM86 (no TMA hardware). |
n/a — different hardware family. |
| Cliff 2b — multi-turn accumulated-context OOM (we filed) | 🟡 Open, Sandermage genesis-vllm-patches#19 | DeltaNet chunk_gated_delta_rule_fwd holds ~500 MiB of simultaneous live tensors at T=4128. Under multi-turn agent traffic (hermes/openhands/etc.), accumulated KV + this kernel's working set + model + workspace exceeds 24 GiB on 1× 3090. Cliff fires at ~21-26K accumulated context. We tested mem-util tuning, MTP-off, max-num-batched-tokens reduction, TRITON_CACHE_AUTOTUNING, expandable_segments, empty_cache between turns — none close it. Validated 2026-05-03: 6 single-card vLLM variants FAIL v2 continuous soak; only TP=2 / dual.yml passes. Filed with Sandermage proposing streaming refactor of GDN forward intermediates. |
bash scripts/switch.sh vllm/dual (TP=2) for 2× rigs, llamacpp/default for 1× rigs. See club-3090#41 + docs/CLIFFS.md "Why TP=2 escapes" / "Why llama.cpp escapes" sections. |
| #41745 — Add Gemma4 MTP speculative decoding support (lucianommartins) | 🟢 Merged 2026-05-06, overlay dropped 2026-05-08 (commit 595be8f). Today's nightly tag 1acd67a795... (2026-05-08 06:10 UTC) contains the merge. dual.yml + single.yml bumped to post-merge nightly; overlay tree models/gemma-4-31b/vllm/patches/vllm-gemma4-mtp/ retained as fallback (drop in follow-up commit once Phase 2 cycle settles). |
First-party MTP for Google's Gemma 4 "assistant" drafter family. Validated on this stack 2026-05-05 (with overlay): 109/142 TPS soak PASS. Re-validated 2026-05-08 (overlay dropped, post-merge nightly): 105.91/141.11 TPS — within CV of the prior baseline → cleanup is parity-clean. | n/a — closed |
| Gemma 4 + per-token-head KV on Ampere (#40388, PR #40391) | 🟡 VENDORED + VALIDATED 2026-05-08 (commit f93d312 + bench 160e8fc). Local rebase of PR #40391 onto post-#41745 main resolved the conflict in vllm/v1/worker/gpu/attn_utils.py (combined main's hybrid attn/mamba dispatch with PR #40391's MLA-vs-standard-attention split for page_size_padded). Vendored as full 7-file overlay. Compose: dual/autoround-int4/int8.yml. Validated dual 3090 Ampere: 7/7 verify-stress at 98K AND at 262K, plus 137K NIAH recall PASS. Bench: 96/127 TPS at 98K, 95/126 at 262K (~10% TPS cost vs bf16 / 32K for 8.2× context lift). |
Earlier 2026-05-06 Codex investigation memo at perheadkv-overlay-comparison.md erroneously concluded "NOT split-able as an overlay" — that was based on PARTIAL overlays (worker-only or spec-only). A FULL PR #40391 overlay (all 8 files + post-#41745 rebase) works cleanly. Key insight: INT8 PTH (not FP8 PTH) is the Ampere-target dtype because Triton fp8e4nv kernel is not supported on sm_86 (only fp8e4b15/fp8e5); FP8 PTH crashes at _initialize_kv_caches on Ampere. INT8 PTH dispatches to standard torch.int8 ops which work on all consumer GPUs. PR #40391's page-size mismatch fix applies to ANY per-token-head KV format — the dtype choice is downstream. Cross-rig validators: cferra (sm_120 Blackwell, FP8 PTH), noonghunna (sm_86 Ampere, INT8 PTH). Phase 3 (PR #40391 + PR #41703 DFlash drafter combined) BOOT-BLOCKED 2026-05-08 — see also #41559 row below for the underlying upstream blocker. 17-file merged overlay parses + compiles, fails _init_minimal_kv_cache_for_profiling with NotImplementedError: page size of the layer is not divisible by the maximum page size at kv_cache_utils.py:1068. Diagnostic-print at unify_kv_cache_spec_page_size (2026-05-08) reproduced exactly the symptom seantechco described in #41559: drafter silently uses BF16 KV regardless of --kv-cache-dtype int8_per_token_head. Three page sizes seen: target Gemma 4 global INT8 PTH 33,280 (padded by PR #40391), target Gemma 4 local INT8 PTH 66,560, DFlash drafter 131,072 (= 16 × 8 × 1024 BF16 K+V at head_dim=256). 131072 / 66560 = 1.97, 131072 / 33280 = 3.94 → no integer ratios → unify rejects. Phase 3b validation (RedHatAI Gemma-aligned drafter, head_dim=256, num_kv_heads=16) failed identically — drafter weights/architecture irrelevant; the actual blocker is per #41559: DFlash mandates non-causal cross-attention and every KV-quant backend rejects KV-quant when causal=False. Why MTP gemma4_assistant works but DFlash doesn't: MTP doesn't require non-causal attention, AND gemma4_assistant shares Gemma 4 architecture (same gemma4.py:438 code path) so PR #40391's INT8 PTH padding propagates uniformly to drafter layers. Phase 3 stacks preserved as forensic artifacts at retired Gemma DFlash INT8 overlay (removed from the active tree in #451; recover from git history if revisited). |
Drop overlay when PR #40391 merges to vLLM main + propagates to a nightly tag. Track: gh api repos/vllm-project/vllm/pulls/40391 --jq '.state, .merged_at'. Local exploratory artifacts at models/gemma-4-31b/vllm/patches/{vllm-perheadkv-hybridpage-fix,vllm-pr40391-perheadkv,vllm-gemma4-fp8-ampere}/ (NOT committed; reference for future iterations). Phase 3 unblock paths: (a) patch qwen3_dflash.py:DFlashAttention.get_kv_cache_spec() to honor cache_config.cache_dtype and return kv_quant_mode=INT8_PER_TOKEN_HEAD with appropriate page_size_padded (single-file fix, most tractable upstream PR); (b) drafter-isolated KV groups extending DFlash's existing _get_dflash_isolated_group_ids to skip page-size unify for draft layers; (c) wait for PR #41703 to merge then re-attempt — fresher base may have unrelated KV-cache refactors that change the picture. As of #451 cleanup: long-context Gemma 4 on Ampere ships only via MTP (vllm/gemma-int8-mtp, 98K default with CTX=262144 MAX_NUM_SEQS=1 override). DFlash/AWQ vLLM composes were pruned from the active registry and remain recoverable from git history if revisited. |
| Gemma 4 + native TurboQuant KV on Ampere (bug #40094, tracking #41403; fix PRs #42637, #40108, #41123) | 🟡 Blocked on our pin, tracked upstream — NOT a permanent / architectural verdict | Exact mechanism (empirically captured 2026-05-16, full traceback): vLLM logs at config.py:101 — "Gemma4 model has heterogeneous head dimensions (head_dim=256, global_head_dim=512). Forcing TRITON_ATTN backend to prevent mixed-backend numerical divergence." → cuda.py:312 uses TRITON_ATTN → ValueError: Selected backend AttentionBackendEnum.TRITON_ATTN is not valid for this configuration. Reason: ['kv_cache_dtype not supported']. The force is unconditional and at config time, before backend selection — that is why FLASH_ATTN / FLASHINFER / XFORMERS / forced-TURBOQUANT / boundary-skip-disable are all moot (vLLM overrides every backend hint for Gemma-4's heterogeneous head_dim by design, as a numerical-correctness safeguard). Fails earlier than #41403 Gate-2 (boundary-skip — N/A here: Gemma-4 is hybrid so get_boundary_skip_layers returns [] already) and earlier than Gate-5 (page-size unify). Verified on Gemma-4-31B (text AutoRound-int4, Ampere sm_86, Genesis-free, pin bf610c2f). Gemma-4-26B-A4B fails even earlier at the WNA16/Marlin weight kernel (Genesis-free base lacks the Marlin-pad overlay — see #40361 row); turboquant verdict untested but predicted-same. No PR enables kv_cache_dtype inside TRITON_ATTN — that is the wrong frame; the upstream fixes operate at the backend-selection / TQ-spec layer so Gemma-4 is not force-pinned to TRITON_ATTN. Contrast: Qwen3-Next dense + MoE native turboquant works Genesis-free on the same pin (#39931 merged) — uniform-enough attention ⇒ no forced TRITON_ATTN ⇒ a turboquant-capable backend is used. Qwen3-Next-only capability today. |
Gemma 4 stays on Genesis / fp8_e5m2 KV — no change, no local action; no runtime kernel/backend/knob lever exists (closing probe exhausted them). Re-test trigger (any of, on merge into a nightly we pin): #42637 (open, freshest 2026-05-16 — most precisely targets this mechanism: adds TQSlidingWindowSpec + changes backend selection so TurboQuant is selected for heterogeneous-head attention instead of force-pinning TRITON_ATTN); #40108 (open, "ready/verified" — sliding-window + YOCO TurboQuant, different approach); #41123 (open — "allow TurboQuant on hybrid models"). Watch: gh pr view 42637 --repo vllm-project/vllm --json state,mergedAt (and 40108, 41123). PR #41497 (preserve sliding-window KV specs) closed without merge. Same unblock pattern as #39931 (merged → unblocked Qwen). |
| Gemma-4-12B unified — advertised 256K p-RoPE OOB past 131072 (#39914 p-RoPE large-prefill; arch PR #44429 merged 2026-06-03) | ✅ Resolved upstream 2026-06-04 by a config.json fix — overlay DROPPED. google/gemma-4-12B-it (+ the -assistant draft) now ship max_position_embeddings=262144, so the stock gemma4_unified image sizes the RoPE cache correctly and serves the full 256K with no patch. lucianommartins (gemma upstream) confirmed it was a config bug (comment); we'd shared the runtime-sizing fix + patch link earlier (comment). The unsloth/gemma-4-12b-it mirror still lagged at 131072 as of 2026-06-04 → bf16 source repointed unsloth→google. |
Root cause: the preview Gemma4 language model sizes the fixed RoPE cos/sin cache from config.max_position_embeddings (was 131072) even when serving with --max-model-len 262144; large prefill then indexed past the cache (< 131072 device assertions). The proper fix is the corrected config (262144), not engine code — so the former local vllm-gemma4-prope-longctx overlay (which sized the cache from runtime max_model_len) is now redundant and was removed. Overlay-free re-validation on dual 3090 TP=2 bf16 (2026-06-04): base + MTP both boot 262144 on the stock image, NIAH exact-recall PASS at 140K/170K/200K/230K/241K, MTP spec-decode active. |
Ship vllm/gemma-12b-dual-bf16-mtp (+ single-card vllm/gemma-12b-single-int8-mtp) at 262144 on the stock image, sourcing bf16 + assistant from google/ (gated; carries the 262144 config). No overlay. |
Genesis (Sandermage/genesis-vllm-patches)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| #22 — PN59 streaming-GDN never engages on chunked-prefill ⚠️ | 🟡 Open, filed 2026-05-05 by us | Genesis v7.72.2 advertises PN59 as the structural Cliff 2b fix on 24 GB single cards, but its eligibility check rejects calls with chunk_indices/chunk_offsets populated — which vLLM's mandatory --max-num-batched-tokens 4128 always sets. PN59 falls back to vanilla, OOMs at the same chunk_o.py:161 site. Single-card 24 GB long-context (long-text.yml / long-text-no-mtp.yml / long-vision.yml) regresses vs the prior workarounds. |
Use dual.yml / dual-turbo.yml (TP=2) or llamacpp/default (different engine, no Cliff 2b). Reproducer + 4 fix proposals in the issue body; awaiting Sander review. |
| #5 — P8 ImportError on vLLM v0.20.0 | ✅ Closed (now on v0.20 pin since 2026-05-01) | Originally about P8 ImportError on the v0.20.0 GA tag. We migrated master to v0.20.1rc1.dev16 + Genesis v7.65 dev tip — P8 path no longer fires on our configs. | n/a — pin already moved. |
| #6 — P65 PIECEWISE cost quantified | ✅ Closed | We characterized the +22 TPS narrative cost of P65 on Qwen3.6-27B + MTP. Sandermage acknowledged. Will recover when vllm#40914 lands. | Accept the cost on substrate-current; ampersandru's pre-P65 stack avoids it. |
| #7 — P67 Triton CompilationError on Qwen3.6-27B | ✅ Closed | Resolved in v7.64 — P67 generalized to non-power-of-2 GQA via BLOCK_QH = triton.next_power_of_2(HEADS_PER_KV) + lane_valid mask. Tool-call 0/5 → 7/7 on 2× A5000 validation. |
Now safe to enable on 27B configs with v7.64+. |
| #9 — P68/P69 8000-char threshold breaks IDE agents | ✅ Closed 2026-05-01 — fix shipped in v7.65+ (50K-char default); we're on v7.69 so the fix lives in our pin | P68 silently rewrites tool_choice: auto → required; P69 injects "must use a tool" hint. New 50K threshold clears typical IDE-agent contexts (Cline ~30K / Cursor ~25K / Copilot ~20-25K) while genuine long-history sessions still trigger the reminder. |
Composes still have P68/P69 env vars commented out — enabling them across composes is queued for next-week revisit (see "Active follow-ups" table above). Until then, manual override: GENESIS_ENABLE_P68_AUTO_FORCE_TOOL=1 GENESIS_ENABLE_P69_LONG_CTX_TOOL_REMINDER=1. |
| #11 — Cliff 1 mech A FA2 softmax_lse clamp request | ✅ Closed (PN17 in v7.64; default-on across all TQ3 composes since the v0.20 migration) | Sandermage's PN17 lands the clamp at flash_attn.py. Active on every TQ3 compose. |
n/a — default-on. |
| Local P104 FA max_seqlen_k runtime clamp | ✅ Dropped during v0.20 migration | Built 2026-04-30 as patch_fa_max_seqlen_clamp.py. Sandermage's PN17 + P15B together cover both layers (FA wrapper + TQ wrapper). Sidecar removed from compose mounts on 2026-05-01. |
n/a — Genesis-native. |
| PR #12 — P101 anchor drift fix | ✅ Closed; on v0.20 pin since 2026-05-01 | P101 anchor matches on 0.20.1rc1.dev16+g7a1eb8ac2. |
n/a — pin matches. |
| PR #13 — PN12 anchor drift fix | ✅ Closed; on v0.20 pin since 2026-05-01 | PN12 anchors match natively. Local patch_pn12_ffn_pool_anchor.py sidecar removed. |
n/a — pin matches. |
| #14 — P38 silently no-op'd on TurboQuant KV path (we filed 2026-05-01) | ✅ Closed via P38B in Genesis v7.65 dev tip | Sandermage shipped P38B — text-patches turboquant_attn.py source to inject a delegate hook at the start of _continuation_prefill body. Active via GENESIS_ENABLE_P38B_COMPILE_SAFE=1 on every TQ3 compose since 2026-05-01. |
n/a — Genesis-native. |
| #15 — FA varlen kernel workspace cliff at flash_attn_interface.py:300 (we filed 2026-05-01) | ✅ Closed via P15B in Genesis v7.65 dev tip | Sandermage shipped P15B — direct backport of our suggestion. Active via GENESIS_ENABLE_P15B_FA_VARLEN_CLAMP=1 on every TQ3 compose since 2026-05-01. Empirically the cliff also doesn't reproduce on v0.20 (vllm#40092 changed workspace allocator behavior) — covered from two directions. |
n/a — Genesis-native. |
| #16 — PN25 worker-spawn registration | 🟡 Sander shipped d92bcb3 (v7.65) + Library refactor (v7.66); both fail on TP=1 | v7.65 used @torch.library.custom_op (failed at infer_schema inside dynamo trace). v7.66 refactored to direct_register_custom_op + Library("genesis", "FRAGMENT") — fails at instantiate_user_defined_class_object inside dynamo trace. Same root cause: torch.library construction inside trace context disallowed on TP=1 spawn. Cross-rig data on Sander's discussion #19 reply. |
Local patch_pn25_genesis_register_fix.py v3 — text-patches activation.py to register at module-import time, BEFORE any trace. Survives both v7.65 and v7.66 mechanisms. PR-ready upstream. |
| #17 — DS conv state spec-decode crash | 🟡 Sander shipped a9977d8 (PN30) but .contiguous() is layout-incorrect |
Sander's PN30 materializes state[src, :, offset:].contiguous() (compact 10240×5) and raw-memcpys into state[dest] (strided 10240×6) → corrupts DS row strides → eventual TQ store CUDA assert several layers downstream. Diagnosis credit: ChatGPT/Codex CLI cross-check 2026-05-02. Sent corrected fix to Sander. |
Local patch_pn30_dst_shaped_temp_fix.py — patches collect_mamba_copy_meta to build dst-shaped temp instead of compact. Reuses Sander's _GENESIS_PN30_TEMP_TENSORS lifecycle. Validated on all 4 TQ3 composes; probes 4 + 5 pass cleanly. PR-ready upstream. |
| #15 — PN31 FA varlen persistent out | 🟡 Sander shipped 753344b (PN31, default OFF); doesn't fit on 24 GB | Per-shape persistent buffer growth + PN12+PN25 pool residence outpaces activation budget at DeltaNet chunk_fwd_o on 24 GB single GPU. Sander explicitly flagged he couldn't validate on 24 GB. |
Lower mem-util to 0.95 — gives enough activation headroom to close the 25K tool-RETURN path PN31 was meant to fix, without needing PN31. Cross-rig data on Sander #15 comment. |
| PN33 spec-decode warmup K-aware (Sander v7.66 fc89395, default ON) | 🟡 Partial close on TP=1 | Backport of vllm#37521 EXTENDED to MTP/ngram. Sander claimed it closes both ampersandru's mid-stream OOM AND our workspace_lock AssertionError. Cross-rig 2026-05-02: closes BOOT-time profile_run workspace_lock ✅, but runtime decode turboquant_attn.py:1350:_decode_attention AssertionError still fires ❌. |
Local patch_workspace_lock_disable.py sidecar still required for runtime decode. Drop when upstream covers the runtime path. |
| P98 marker false-positive on v0.20 (we filed 2026-05-01 in #9 thread) | 🟡 Side-noted to Sandermage; awaiting his call on fix | P98's drift detection auto-skips on v0.20 (UNIFORM_SINGLE_TOKEN_DECODE marker false-positive) but the strict workspace lock still fires rare paths P98 was supposed to revert. |
Local patch_workspace_lock_disable.py sidecar (mounted on every TQ3 compose) relaxes the strict assertion to a one-shot WARNING. Drop when Sandermage ships either a marker fix or P98 with explicit env-override. |
| PN30 v7.68 part3 drift-marker false-positive (we filed in noonghunna/club-3090#19 cross-rig retest) | ✅ Closed in v7.69 (commit 2db18df) | Part3's upstream_drift_markers=["[Genesis PN30"] (generic prefix) matched markers parts 1+2 wrote on the same file. Part3 skipped as upstream_merged → apply_all FAILS → vLLM aborts. v7.69 tightened to [Genesis PN30 v7.68 dst-shaped] (specific). |
n/a — fixed in v7.69. |
P103 setattr lost on exec vllm serve (we filed in noonghunna/club-3090#19) |
✅ Closed in v7.69 | v7.68 P103's setattr ran in entrypoint shell but was lost on exec vllm serve worker spawn (process image replaced). v7.69 ships chunk.py self-install hook appended to end-of-file — survives any startup mechanism. |
n/a — fixed in v7.69. |
| PN32 v1 chunked at wrong level (we filed in noonghunna/club-3090#19) | ✅ Closed in v7.69 (PN32 v2) | PN32 v1 chunked outer-level inputs but inner FLA call still got full-prompt cu_seqlens, allocating full h tensor regardless. v7.69 PN32 v2 patches _forward_core directly + threads last_recurrent_state between chunks. |
n/a — fixed in v7.69. |
| #18 — P103 cu_seqlens=[0,T] single-seq case is bypassed (we filed 2026-05-02 PM) | 🟡 Open / v7.70 proposal | P103's gate currently bypasses chunking for ANY non-None cu_seqlens, but cu_seqlens.shape[0] == 2 (single sequence boundary) is semantically dense B=1, not multi-seq varlen. Fix admits the chunked path on real serving. Diagnosis: ChatGPT/Codex CLI. Cross-rig observation: P103 chunked path never engages on real config because vLLM's outer chunked-prefill caps T at max_num_batched_tokens=4128 (well below _MAX_T=16384), so the gate-fix is semantically correct but doesn't independently close 60K Cliff 2 on TP=1+24GB. |
n/a yet — gate fix queued for v7.70. Real Cliff 2 closure on this config comes from vllm#35975 backport + mem-util 0.93 (see vLLM section above + docs/CLIFFS.md). |
FlashAttention 2 (Dao-AILab/flash-attention)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| #1011 — Variable memory allocation with varlen kernels | 🔴 Open since 2024, no fix | Cliff 1 root cause. softmax_lse is allocated as [num_seqs, num_heads, max_seqlen] — sized by max_seqlen parameter, NOT actual cu_seqlens. So a 25K-token chunked-prefill at max_model_len=86K allocates softmax_lse for 86K, not 25K. This is why Cliff 1 fires harder at higher max-ctx even when the actual prompt is the same. |
None. Stay at default 48K (or tools-text 75K with PN8 mitigation). FA2 redesign of softmax_lse format would be the upstream fix. |
flash-linear-attention (fla-org/flash-linear-attention)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| Cliff 2 — DeltaNet GDN forward OOM at 50–60K single-prompt | 🔴 Open, no upstream issue filed yet. Confirmed cleared on dual TP=2 (this rig, 2026-04-29 — see DUAL_CARD.md "237K single-prompt verified"). | The chunk_gated_delta_rule_fwd kernel allocates intermediate buffers proportional to seq_len. Fires on single-card regardless of mem-util. On dual TP=2 the activation memory splits across cards and the cliff doesn't fire — verified at 237K single-prompt prefill on dual.yml (~830 tok/s prefill, matches Sandermage's 262K @ 311s on 2× A5000). Sandermage explicitly punted on the single-card fix (genesis-vllm-patches issue #1: "can't fix this short of multi-GPU TP=2 or upstream fla.ops changes"). Likely the same architectural pattern as FA#1011 — recurrent state buffer pre-allocated by max_seq_len. |
Single-card: use tools-text.yml (75K cap) or llamacpp/default (262K, different engine). Dual: dual.yml clears at ≥237K. |
FlashQLA (QwenLM/FlashQLA)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| Ampere SM 8.6 / Ada SM 8.9 port | 🔴 No issue filed; tweet to @QwenLM drafted but not yet posted | FlashQLA is QwenLM's TileLang DeltaNet kernels — would fix Cliff 2 if it ran on Ampere. Currently SM90+ only. | None. Watch the repo for Ampere support; revisit when an issue is filed and a port is on the roadmap. |
Luce DFlash (Luce-Org/lucebox-hub) — separate llama.cpp fork (NOT our vLLM dual-dflash)
Heads-up — naming clarification:
- This section tracks
Luce-Org/lucebox-hub(a llama.cpp fork from Luce). As of 2026-05-04 this is no longer single-card-only — see "Dual-GPU split landed" below. - Our
dual/autoround-int4/dflash.yml/dual-dflash-noviz.yml(vLLM TP=2 dual-card) were the vLLM DFlash path but were deprecated 2026-05-31 — superseded bydual.yml(stock v0.22.0, 262K + vision + MTP + 2 streams) and stranded on a now-purged nightly (#297). DFlash on dual now ships on beellama (beellama/qwen-dflash-dual, v0.3.0 🧪). They consumed the same draft model (z-lab/Qwen3.6-27B-DFlash), but the engine + topology differ from Luce's fork. Don't confuse the two.
🆕 Dual-GPU split landed (2026-05-02 + 2026-05-04)
Two @weicj PRs shipped that change the lucebox-hub serving topology. Target weights on one GPU + DFlash draft (or PFlash drafter) on a separate GPU — heterogeneous spec-decode, not weight-sharded TP. Each model lives entirely on its own card; they communicate at spec-decode boundaries via peer copies.
- lucebox-hub PR #80 —
bench(dflash): add dual-GPU target/draft split harness(merged 2026-05-04). New flags--target-gpu/--draft-gpu(alsoDFLASH_TARGET_GPU/DFLASH_DRAFT_GPUenv). Validation on dual RTX 2080 Ti 22 GB: HumanEval 10-prompt at 51.86 tok/s, AL 7.09, 44.3% accept on Qwen3.5-27B Q4 target + z-lab DFlash draft. - lucebox-hub PR #78 —
bench(pflash): add dual-GPU PFlash phase-split harness(merged 2026-05-02). New flag--pflash-gpu+ persistentpflash_daemon. Validation on same hardware: single-GPU co-resident passes NIAH at 24,573 source tokens; dual-GPU phase split passes at 262,125 source tokens (10.7×). Compressed context reaches 13,229 tokens at 262K source.
Implication for our 2× 3090 stack: the single-card limitations we documented (65K max_ctx, draft VRAM competing with target activations) are addressed by dual-GPU split. Target Qwen3.5-27B Q4_K_M gets a full 24 GB on GPU 0; DFlash draft + PFlash drafter live on GPU 1. No NCCL/allreduce overhead per token since each model lives entirely on its own card — should be faster per-stream than SGLang TP=2 + DFlash for single-stream workloads. Bench tracked at task #229 (queued, not yet executed locally — PR #80 is hours old as of this entry). Qwen3.6-27B draft remains under training so the dual-GPU benefit applies primarily to the stable Qwen3.5-27B + DFlash pair today.
Re-benched 2026-04-30 PM on Qwen3.6-27B Q4_K_M + matched z-lab/Qwen3.6-27B-DFlash draft (under training). Open issues against single-card lucebox-hub follow:
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| z-lab/Qwen3.6-27B-DFlash — draft model still under training | 🟡 Snapshot 2026-04-26 | Narrative AL ~3.7, code AL ~7.0 on Luce-Org/lucebox-hub (single-card llama.cpp). When training finishes, expected to climb toward Qwen3.5 reference (8.31 HE, 7.04 Math). The same caveat applies to vLLM dual-dflash.yml — published 82/125 TPS in docs/DUAL_CARD.md was measured against this 2026-04-26 snapshot at peak code-prompt conditions; AL on real agent traffic will be lower until z-lab tags training-complete. |
Re-test when z-lab tags training-complete. The vLLM dual-dflash path remains shipping — see DUAL_CARD.md — but treat its numbers as a snapshot. For autonomous coding agents on dual-3090 today, dual.yml (FP8 + MTP) is the recommended robust path. |
Build fragility on dflash main HEAD |
🔴 Reproducible 2026-04-30 PM | cmake --build errors with ggml_turbo_wht and GGML_TYPE_TQ3_0 undefined. Required submodule commit b6ffab4a9 not auto-fetched. Cross-rig signal — fresh clone fails. |
After clone: cd dflash/deps/llama.cpp && git fetch origin && cd ../../.. && git submodule update --init. |
| Daemon-mode "empty prompt" regression | 🔴 Reproducible 2026-04-30 PM | After streaming requests, subsequent requests return "empty prompt" from the test_dflash daemon. Server keeps accepting requests but generates 0 tokens. Forces restart. |
Restart server between request flavors; avoid mixing streaming + non-streaming. |
enable_thinking chat_template_kwargs honored differently than vLLM |
🟡 Behavioural difference | Test sends enable_thinking=true and expects reasoning_content populated. Luce returns content directly. Not a missing feature, but breaks our verify-full.sh check 6. |
Don't treat the thinking-mode test as a Luce-correctness signal until the chat-template path is documented. |
| Greedy only | 🟡 Documented limitation | temperature / top_p accepted but ignored. Real downside for creative-writing workloads. |
Use vLLM long-text/long-vision when sampling matters. |
Prefill OOM in fattn-chunked.cu on 25K+ prompts at Q8_0 KV |
🟡 Open (configuration trade) | Chunked flash-attention CUDA OOMs on large prefill at default Q8_0. TQ3 KV (DFLASH27B_KV_TQ3=1) closes it at max_ctx=65K — verify-stress passes 791 chars / finish=stop. Higher max_ctx (131K) reopens it. |
Always set DFLASH27B_KV_TQ3=1 for stress-test-passing config. Cap max_ctx at ~65K. |
| PFlash — long-context prefill accelerator (sibling tech to DFlash, same Luce-Org/lucebox-hub repo) | 🟢 Public release 2026-04 + dual-GPU split shipped 2026-05-02 (PR #78) | Speculative prefill + block-sparse attention. Compresses 128K prompts to ~6.5K tokens (keep_ratio=0.05) before target prefill. Single-card claimed: TTFT 24.8s vs 257s vanilla llama.cpp at 128K (~10.4× speedup). Dual-GPU phase split (PR #78) extends the passing source-context ceiling from ~24K (single-card co-resident) to 262K (~10.7×) on dual 22 GB cards — NIAH key/answer retained at 262K. C++/CUDA only, lives inside the lucebox-hub server stack. PFlash sits in front of DFlash decode: PFlash accelerates prefill, DFlash accelerates generation. For 2× 3090 deployments: pin PFlash drafter to GPU 1 via --pflash-gpu, target on GPU 0. The single-card-coresident limit (was the binding blocker for our use) no longer applies. MIT license. Open exploration: bench PFlash + DFlash dual-GPU vs vLLM dual-dflash.yml (185K, 82/125 TPS on 2× 3090) on TTFT-bound workloads. Tracked at task #229. |
Re-evaluate as a club-3090 shipping option once we (a) reproduce the 262K passing source-ctx claim on 2× 3090 with verify-stress + soak-continuous + bench, OR (b) an upstream-vLLM port lands. The dual-GPU split removes the single-card co-residency blocker; remaining blockers are daemon-mode bugs (greedy-only, no vision, "empty prompt" regression) carried over from the single-card history. |
llama.cpp (ggml-org/llama.cpp)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| PR #21089 — TurboQuant KV mainline | 🟡 Open (CPU first, CUDA follow-on) | When CUDA path lands, turbo3 becomes a first-class option on llama.cpp. Naming will migrate from turbo3 → tbq3_0. |
Use Tom's fork for now: llama-cpp-turboquant. |
| Q3_K_XL TPS regression (28.5 TPS @ 262K → 21 TPS today) | 🔴 Suspected, no upstream issue filed | Measured 2026-04-23 vs 2026-04-28: same model, same hardware, 28.5 TPS dropped to 21 TPS between commits 9ab47e7d8 and 0d0764dfd. Bisect or file. |
None — we're on the slower commit. Tracked in club-3090 TODO (private). |
PR #22673 — MTP support (am17an, mtp-clean) |
🟡 Open, unmerged | First-party MTP for llama.cpp via an MTP head baked into the GGUF (RDson republished Qwen3.6-27B-MTP-Q4_K_M-GGUF with the head wired). Benched on 1× 3090 (2026-05-05): +34% narrative TPS at n-max=3 (22.83 → 30.69), ~57% accept. Code at n-max=5 hit 31.9 TPS. NOT a club-3090 recommendation yet. Reasons: (1) unmerged → forces every cross-rig user to compile am17an's fork or maintain a custom image; (2) q8_0 KV ceiling caps context at ~64-80K — current llamacpp/default ships 262K, trading that for +34% TPS isn't worth it for the cliff-immune audience; (3) MTP forces n_parallel=1 (kills llamacpp/concurrent.yml); (4) RDson GGUF doesn't bundle mmproj (vision regression). Audience for this is empty — vLLM dual-turbo already gives 170 TPS for users wanting max single-stream throughput. |
None recommended. Re-evaluate when PR merges + q4_0 KV variant tests recover 128K+ context + cross-rig data lands. Detailed bench + reasoning is documented out-of-tree in this stack's learnings/qwen3.6-35b-a3b.md ("llama.cpp MTP — PR #22673 path" subsection). |
transformers (huggingface/transformers)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| #45283 — Qwen3.5 GGUF support | ❌ Closed without fix 2026-04-28 (no associated PR; closing event has source: null, last comment was just cc @SunMarc — looks won't-fix or stale-bot) |
Was tracked as the missing piece (alongside vllm#38140 / vllm#37797) for Qwen3.5/3.6 GGUF on vLLM/SGLang. Won't be picked up via transformers — llama.cpp remains the only GGUF path for this model family. | llama.cpp path. Don't expect a vLLM/SGLang GGUF route for Qwen3-Next family. |
transformers ≥ 5.8.0 required for gemma4_assistant |
✅ Resolved 2026-05-08 (ships in pinned images) | transformers 5.8.0 (released 2026-05-05) was the first version with the native gemma4_assistant model class (Google's Gemma 4 MTP drafter). The historical break: vLLM nightly :nightly-01d4d1ad3 shipped transformers 5.7.0 → AutoConfig rejected the drafter checkpoint at validation time. |
No longer needed — the dep now ships in the pinned stable images. Verified 2026-06-20 live in the running container: vllm/vllm-openai:v0.22.0 carries transformers 5.9.0 (≥ 5.8.0); the single-card pin vllm/vllm-openai:v0.21.0 is likewise post-Gemma4-merge (compose-header verification 2026-05-08). The former pip install --upgrade transformers==5.8.0 entrypoint line was removed 2026-05-08 — see the Gemma dual compose models/gemma-4-31b/vllm/compose/dual/autoround-int4/bf16-mtp.yml ("entrypoint upgrade line dropped"). |
SGLang (sgl-project/sglang)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| Same Marlin pad-sub-tile-n bug as vllm#40361 | 🔴 Not filed; same kernel-line fix applies | Blocks Lorbus INT4 + EAGLE on SGLang. We haven't filed an SGLang PR. | None on SGLang. Use vLLM (with our patched fork) or wait for SGLang to pick up the upstream Marlin fix. |
| DeltaNet KV rollback (vllm#39931 cross-engine) | 🔴 Same architectural issue | Blocks EAGLE on Qwen3-Next family in SGLang too. | None — see vllm#39931. |
Community templates / model assets (Hugging Face)
External-but-load-bearing resources that aren't issue trackers (no PR / merge state to track). Watch list — re-check when upstream Qwen / Gemma official templates change, or when these resources update.
| Resource | Status | Why it matters | Drop trigger |
|---|---|---|---|
froggeric/Qwen-Fixed-Chat-Templates — community fork of the default Qwen 3.5 / 3.6 chat templates fixing seven documented bugs (empty <think></think> spam in past turns, </thinking> hallucination on Qwen 3.6, unclosed thinking before tool call, no-user-query crash in agentic loops, developer role rejection, |items filter for C++ engines, type-aware tojson). Surfaced by @troymroberts in discussion #121. Vendored snapshot at models/qwen3.6-27b/vllm/patches/froggeric-chat-template/chat_template.jinja; mounted default-on across all 22 vanilla Qwen 3.6 composes via --chat-template. Carnice and Qwopus composes intentionally excluded (ship their own bespoke templates). ADOPTED 2026-05-18 for issue #150: froggeric v19, HF c31fd393e531dbacd92b6deb99a2037cc949f950 (2026-05-16) — re-eval PASSED maintainer rig A/B. |
✅ Re-eval PASSED (#150, 2026-05-18) — v19 ADOPTED. Today-anchored A/B on vllm/dual (Qwen3.6-27B INT4, identical engine, template-only delta): hermesagent-20 +10pp (50% → 60%, +2 scenarios); 7 other packs flat (toolcall/instructfollow/structoutput/dataextract/reasonmath/bugfind/cli-40 unchanged); streaming tool-calls OK both arms (#145 clear); soak-continuous PASS 0/25 silent-empty both arms. TPS-neutral — symmetric restart+settle, 3 bench/arm: CODE 89.06 → 87.62 (−1.6%, within run-to-run noise), NARRATIVE flat. The first-pass −7% did not reproduce — it was an asymmetric-protocol artifact (candidate benched ~110s post-restart vs baseline post-settled-boot), not a template effect. Reproduces the prior froggeric finding on a clean today-anchored baseline. (Superseded prior: 🟡 first-pass A/B 2026-05-12 — +15pp hermesagent-20 45%→60%; the 45% base was stale, today's is 50%, control-run confound now resolved by the symmetric re-eval.) |
Replace with default model template if upstream Qwen pushes equivalent fixes. Watch for: froggeric updates the template (Qwen 4 support, additional bug fixes), or Qwen upstream lands their own version. |
Filing conventions
When you file or learn of a new upstream issue:
- Add a row to the appropriate section of this file. Include the link, status emoji, one-line "why it matters," and the local workaround (if any).
- Cross-link from any code, compose comment, or doc that depends on the workaround back to the row in this file (e.g.,
# See docs/UPSTREAM.md — vllm#40361). - Update the row when status changes — closed, merged, propagated, replaced. Don't delete; if a row is no longer load-bearing, mark it ✅ Resolved or ❌ Closed without fix and leave it as historical context.
- Bump the relevant pin when an upstream lands (Genesis commit, vLLM nightly, llama.cpp commit). Add a CHANGELOG entry citing the upstream PR.
When you file an issue against an upstream repo from this work, link back to club-3090 in the body so the upstream maintainer can see the affected user surface and re-test if needed.
Related reading
models/qwen3.6-27b/INTERNALS.md— model-specific deep dives (DFlash forensics, MTP head, AutoRound rationale)models/qwen3.6-27b/vllm/patches/README.md— local patches (tolist, Marlin pad fork, Genesis env-var matrix)AGENTS.md— repo-wide conventions, including the rule that this file is the upstream-tracking single source of truth