Files
club-3090/docs/UPSTREAM.md
T
noonghunnaandClaude Opus 4.7 84498d47aa
Release / release (push) Failing after 50s
feat(qwen): ship froggeric chat-template fixes as default-on
Vendored snapshot of froggeric/Qwen-Fixed-Chat-Templates qwen3.6
template, mounted into all 22 vanilla Qwen 3.6-27B composes via
--chat-template. Replaces the model's default Jinja template with
the community-patched one. Carnice and Qwopus composes intentionally
excluded — they ship bespoke Hermes-JSON templates that must not be
overwritten.

Upstream fixes seven documented bugs in the default Qwen 3.5 / 3.6
templates:
  - empty <think></think> blocks polluting past-turn context
  - </thinking> closing-tag hallucination on Qwen 3.6
  - unclosed <think> before tool_call (mangled output)
  - raise_exception crash when no user query in messages (kills
    agentic loops)
  - "developer" role rejection (blocks modern API clients)
  - |items Jinja filter unsupported in C++ runtimes (llama.cpp,
    LM Studio, MLX)
  - type-aware tojson serialization for tool arguments
Source: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

Surfaced by @troymroberts in club-3090 discussion #121.

First-pass A/B 2026-05-12 vs the matched-config Qwen INT8 PTH n=4
rebench baseline (2026-05-10):

  Pack             Baseline   Froggeric   Δ
  ---------------- --------   ---------   --
  toolcall-15       10/15      10/15       0
  instructfollow    13/15      13/15       0
  structoutput      13/15      13/15       0
  dataextract       15/15      15/15       0
  reasonmath         6/15       6/15       0
  bugfind           11/15      11/15       0
  hermesagent-20     9/20      12/20      +3 (+15pp)
  cli-40            17/40      17/40       0
  TOTAL             94/150     97/150     +3 (+2pp)

hermesagent-20 is the multi-turn agentic pack — exactly where the
empty-think + no-user-query + unclosed-think-before-tool-call fixes
compound. 7 other packs flat = no regression on single-turn flows.

Added a new "Community templates / model assets" section to
docs/UPSTREAM.md tracking the resource + drop trigger (replace when
upstream Qwen pushes equivalent fixes).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-12 21:21:37 +00:00

241 lines
54 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Upstream tracker
Issues and PRs in upstream repos that affect this stack — what we depend on, what we've filed, what unblocks for us when each lands.
This file is the **single source of truth** for upstream status. When you file or notice an upstream issue / PR / commit relevant to club-3090, add a row here. When status changes (closed, merged, propagated), update it. Don't scatter the same link across multiple docs without coming back here first.
If you're adding a new compose that depends on an unmerged upstream patch (volume-mount of a fork, monkey-patch script), it MUST link to a row in this file so future readers know when the workaround can drop.
---
## How rows work
Each row covers one upstream link with: **title • status • our dependency / impact • workaround (if any)**.
**Status vocabulary:**
- 🟢 **Landed** — merged upstream + propagated to our pinned versions (pin-bump done)
- 🔵 **Merged, awaiting propagation** — merged upstream but our nightly / commit pin hasn't picked it up yet
- 🟡 **Open / in review** — PR open, no merge yet; we depend on it landing
- 🟠 **Open / blocked or stalled** — PR exists but progress stalled
- 🔴 **Open, no PR yet** — issue acknowledged but no fix in progress (us or upstream)
- ⚫ **Workaround locally, no plan to merge** — fixed in our patches, upstream not pursuing
- ✅ **Resolved** — closed and resolved (kept for historical context)
- ❌ **Closed without fix** — closed, won't fix, kept for context
---
## Active follow-ups (next-week revisit queue) 🗓️
Items deferred for review next week (week of 2026-05-10). Audit at the start of that week — most should be either ready to action or have new upstream signal worth re-evaluating.
| Item | Why deferred | Trigger to revisit |
|---|---|---|
| **🟡 [Sandermage/genesis-vllm-patches#22](https://github.com/Sandermage/genesis-vllm-patches/issues/22)** — PN59 streaming-GDN doesn't engage on chunked-prefill (single-card 24 GB Cliff 2b stays open under v7.72.2). Filed 2026-05-05 with reproducer + 4 fix proposals. | Awaiting Sander's review. The cleanest of our 4 proposals is making `has_no_chunk_metadata` rejection optional (env-gated), letting single-seq chunked-prefill take the streaming path. | Sander posts a candidate fix or comment on the issue; or we run a one-line A/B if he requests it. **Until resolved, single-card 24 GB long-context users should run `dual.yml` / `dual-turbo.yml` / `llamacpp/default`.** |
| ~~**Genesis pin bump `2db18df` → `f2147ad`**~~ | ✅ **Done 2026-05-05** — bumped to `7b9fd319` (v7.72.2) on branch `v7.72.2-uplift`. Drops `patch_inputs_embeds_optional.py` (PN35 native), `patch_pn30_dst_shaped_temp_fix.py` (PN30 v7.68), `patch_pn25_genesis_register_fix.py` (PN25), `patch_tolist_cudagraph.py` (P78), `patch_workspace_lock_disable.py` (PN34), `patch_pr40798_workspace.py` (research artifact). | — |
| ~~**Enable P68/P69 across composes**~~ | Superseded by v7.72.1 P68 auto-skip + v7.72.2 PN70 schema-subset filter — both ship default-aware behavior. Closed [#57](https://github.com/noonghunna/club-3090/issues/57) along the way. | — |
| **Rebase + ping vllm#40361** (our Marlin pad-sub-tile-n PR) | 13 days stale on vLLM upstream as of 2026-05-03; not blocking anything locally (we vendor the patched files in-repo) but worth keeping on the maintainer queue. | Open a "still relevant" comment + rebase if behind main, link to the JusefPol NVLink merge as additional cross-rig evidence the patch is in real use. |
See the platform-specific tables below for the rows these reference.
---
## Pinned images
What container image each compose pins, why each pin exists, and which pins
are candidates for retirement when their reason resolves. This section answers
"why do we have N distinct vLLM nightlies cached" and drives the work in
[`NIGHTLY_BUMP_RUNBOOK.md`](./NIGHTLY_BUMP_RUNBOOK.md).
Run `bash scripts/maintenance/list-image-pins.sh` for a live snapshot.
| Pin | Composes using it | Reason for pin | Retirement candidate? |
|---|---|---|---|
| `vllm/vllm-openai:nightly-01d4d1ad` | 19 (all Qwen 3.6-27B) | Working baseline for Marlin-pad patch + Genesis v7.51 + Qwen 3.6-27B AutoRound. | When Marlin PR #40361 merges + propagates to a newer nightly, bump to that nightly + drop the Marlin patch mount. Tracked in vLLM section below. |
| `vllm/vllm-openai:nightly-1acd67a7` | 4 (Gemma 4 base / AWQ / INT8) | Post PR #41745 merge — Gemma 4 MTP "assistant" drafter. | When PR #42102 (DFlash + KV-quant unblock) propagates AND PR #40391 (per-head KV) absorbs, can consolidate with `nightly-e47c98ef` to a single Gemma pin. |
| `vllm/vllm-openai:nightly-e47c98ef` | 2 (Gemma 4 DFlash + DFlash-INT8) | Working baseline for our DFlash + DFlash-INT8 patch stack. Heaviest patch surface (~25 patched files spanning `v1/spec_decode/`, `v1/attention/`, `v1/worker/`). | When PR #42102 lands + DFlash patches absorb upstream, this pin retires. Until then, **do not bump blindly** — the patches are tightly coupled to this nightly's internals. |
| `ghcr.io/ggml-org/llama.cpp:server-cuda` | 2 (Qwen 3.6-27B llama-cpp) | Stable tag, no hash drift on upstream side. No patches mounted. | Not a retirement candidate — drift-free. Capture digest if reproducibility matters. |
**Retirement workflow:** see [`NIGHTLY_BUMP_RUNBOOK.md`](./NIGHTLY_BUMP_RUNBOOK.md).
### Retired pins
(none yet — first entry will land when the first consolidation happens)
---
## vLLM (`vllm-project/vllm`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [#35936](https://github.com/vllm-project/vllm/pull/35936) — `tool_choice="required"` falls back to configured tool parser | 🟡 Open / **local overlay active** | Qwen3-Coder with `--tool-call-parser qwen3_coder` emits XML-style tool calls. On pinned nightly `1acd67a79`, non-streaming `tool_choice="required"` validates JSON only, bypasses the configured parser, and returns `tool_calls=[]`. MLS-Bench hits this when `thinking.enabled=false`. | Vendored overlay: [`models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md). Drop when #35936 or equivalent lands in our pinned image. |
| [#40361](https://github.com/vllm-project/vllm/pull/40361) — Marlin pad-sub-tile-n | 🟡 Open, mergeable, **stale 13d** (last update 2026-04-20) | All 4 dual-card composes + `dual-nvlink.yml` + `dual-nvlink-turbo.yml` mount the patched files vendored in-repo at `models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/`. Drops out as a setup dependency when this merges + propagates. | Vendored mount: see [`models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md). Queued for rebase + ping next week (see "Active follow-ups" table above). |
| [#40807](https://github.com/vllm-project/vllm/issues/40807) — `.tolist()` cudagraph crash on continuation-prefill | ⚫ Local workaround | Single-card TQ3 + spec-decode + chunked-prefill blocked without it. We ship a file-edit patch. | `patch_tolist_cudagraph.py` runs in `setup.sh`. Drop when upstream fixes the sync. |
| [#40849](https://github.com/vllm-project/vllm/pull/40849) — MTP draft online-quant propagation | 🟡 Open / Genesis backport active | Closes Cliff 1 on FP8+MTP path (`tools-text.yml`). | Genesis PN8 backport: `GENESIS_ENABLE_PN8_MTP_DRAFT_ONLINE_QUANT=1`. |
| [#40914](https://github.com/vllm-project/vllm/pull/40914) — Sandermage K+1 verify routing | 🟡 Open, ❌ negative on our Qwen3.6-27B stack | **Reframed 2026-05-11:** the synthetic `seq_lens` K+1 route is not the P67-equivalent we need here. Local rebase on post-#41434 nightly made MTP acceptance look perfect (AL=4.0 / ~100%) but produced `!`-flood needle corruption plus tool/multi-turn timeouts. Dropping it improved verify-stress from 3/7 to 5/7, but TQ3/TQ4/k8v4 + MTP still fail long-context needles. | Do not ship Genesis-free TQ+MTP on #40914 alone. Use `dual/tq3-nomtp.yml` without Genesis, or `dual/tq3-mtp-genesis.yml` with Genesis P67/P67b. |
| [#40334](https://github.com/vllm-project/vllm/pull/40334) — DFlash `combine_hidden_states` dtype mismatch | 🟡 Open | All `dual-dflash*.yml` need `--dtype bfloat16` flag to work around. | Composes set `--dtype bfloat16`. Drop when this lands. |
| [#40382](https://github.com/vllm-project/vllm/issues/40382) — Gemma-4 + DFlash unservable on Ampere | 🟠 Open, no fix in progress | Blocks DFlash on Gemma-4 family. Not directly our problem (we serve Qwen3.6) but tracked because future model adds may hit it. | None — different attention backend selection. |
| **[#41559](https://github.com/vllm-project/vllm/issues/41559) — DFlash spec-decode incompatible with all KV cache quantization** (seantechco, filed 2026-05-03) | 🟢 **OUR FIX PR OPEN: [#42102](https://github.com/vllm-project/vllm/pull/42102)** (filed 2026-05-08) | **REFRAMED 2026-05-08 PM via Codex investigation**: original allowlist-gating framing was partially outdated on current main. Current state: FLASH_ATTN gates dynamically via `flash_attn_supports_fp8()` (FA3-only); FLEX_ATTENTION raises `NotImplementedError` on quantized KV at impl construction; TRITON_ATTN remains causal-only via `assert causal` at `triton_unified_attention.py:542`. The KV-quant write path itself (`triton_reshape_and_cache_flash_per_token_head_quant`) IS causal-mask-independent — but no current backend actually executes both quantized KV AND non-causal attention. **Sharper framing for the common case (BF16 DFlash drafter alongside quantized target KV)**: don't need any backend to "support quantized KV in non-causal mode" — just need the engine to stop forcing target+drafter to share a single page-size unify pass. Three-layer local fix at `/opt/ai/engines/vllm/primary` branch `dflash-noncausal-kv-quant` (commit `cfb8f711`, 4 files, +333/-35): (1) `vllm/v1/core/kv_cache_utils.py` partition DFlash drafter specs into independent KV groups before unify, allocator extended to size isolated tensors by their own page_size; (2) `vllm/model_executor/models/qwen3_dflash.py` override drafter cache_dtype to "auto" when engine global is quantized; (3) `vllm/v1/attention/backends/flash_attn.py` FA metadata scheduler uses per-spec dtype when spec's kv_quant_mode is NONE. **Validated end-to-end on dual 3090 Ampere**: Gemma 4 + z-lab DFlash drafter + INT8 PTH KV target boots HEALTHY at 65K, Paris smoke clean, narrative 95.89 / code 168.09 TPS (matches bf16 32K baseline within CV — long-context unlocked at zero perf cost), AL 5.0-5.3 long-ctx code preserved, NIAH PASS at 32K prompt, KV pool 149,345 tokens (4× lift over baseline). | Local commit `cfb8f711` ready for review + push to `noonghunna/vllm` fork + upstream PR submission. PR description draft at `/tmp/dflash-int8-pr-description.md` (covers non-duplication checks, AI-assistance disclosure, validation matrix). Forensic Phase 3a/3b stacks remain at `models/gemma-4-31b/vllm/patches/vllm-gemma4-dflash-int8/` as historical record (the wrong-fix path that helped diagnose). Container artifacts cleaned up; Qwen production restored. |
| [#40354](https://github.com/vllm-project/vllm/issues/40354) — Marlin TP=2 W4A16 < 64 | ✅ Same root-cause as #40361 | Our PR #40361 resolves this. | See #40361 row. |
| [#39931](https://github.com/vllm-project/vllm/issues/39931) — DeltaNet rollback support | 🔴 Open, architectural | Blocks **all** spec-decode (EAGLE / DFlash) on Qwen3-Next family across engines. The reason "speculative decoding doesn't work" on this stack. | Use MTP (no rollback needed) until this lands. |
| [#40124](https://github.com/vllm-project/vllm/issues/40124) — related architectural | 🔴 Open | Pairs with #39931 for DeltaNet rollback. | Same as above. |
| [#40880](https://github.com/vllm-project/vllm/issues/40880) — MTP × TQ × cudagraph cascade | ✅ Closed upstream issue, but not solved by direct upstream vLLM | Genesis P65 removed the CUDA-graph-specific failure mode; P67/P67b is the correctness path for K+1 multi-query TurboQuant attention. Round-4 testing showed `--enforce-eager` alone does not close TQ+MTP needles without P67-equivalent behavior. | Use Genesis P67/P67b or disable MTP on TurboQuant. |
| [#40831](https://github.com/vllm-project/vllm/issues/40831) — TQ × spec-decode corruption | ✅ Closed issue, open upstream gap | Our new matrix reproduces the same class across TQ3, TQ4, and k8v4 under MTP. TQ3 no-MTP passes 7/7, so the bug is the MTP x TurboQuant multi-query path, not precision. | Same as #40880. |
| [#40798](https://github.com/vllm-project/vllm/pull/40798) — workspace-manager refactor | ❌ Negative result | Hypothesized fix for #40831 / #40880; backporting it (Probe 8) didn't resolve the bug. Kept for context — saved future time on the same dead end. | n/a |
| [#40875](https://github.com/vllm-project/vllm/issues/40875) — ngram + MTP coexistence | ✅ Closed | Routed via `prompt_lookup_min=8` flag. | Set in compose where applicable. |
| [#41142](https://github.com/vllm-project/vllm/pull/41142) — Quentin-M streaming tool-call IndexError | 🟡 Open / Genesis backport active | Closes a streaming tool-call crash on Hermes / similar templates. | Genesis PN11 backport (auto-enabled where REC). |
| [#39598](https://github.com/vllm-project/vllm/pull/39598) — kotori-yan qwen3coder MTP streaming early-return | 🟡 Open / Genesis backport active | Empty `tool_calls[]` when MTP bundles last param + `</function>` in same delta. | Genesis P64 backport: `GENESIS_ENABLE_P64_QWEN3CODER_MTP_STREAMING=1` (default-on in our composes). |
| **qwen3coder tool-parser SSE-silence on prose `<tool_call>`** ([upstream #22975 closed-as-stale](https://github.com/vllm-project/vllm/issues/22975); reported on club-3090 as [#72](https://github.com/noonghunna/club-3090/issues/72)) | ⚫ Local workaround / **upstream PR deferred until cross-rig validation lands** | When the model's prose mentions the literal `<tool_call>` text (e.g. agent reasoning that describes the markup), `extract_tool_calls_streaming` flips `is_tool_call_started=True` permanently on either the special-token-id or the string match. Subsequent deltas return `None`; the serving layer skips them; SSE wire goes silent for 30-120s while tokens decode server-side and never reach the client. Verified bug still present in vLLM main as of 2026-05-07 (no deferred-commit guard in current source). Upstream issue #22975 reports a related symptom (`<tool_call>` markup remains as plain content) but was closed-as-stale 90+ days ago without a fix — different observed surface, likely shared root cause. | `models/qwen3.6-27b/vllm/patches/local/qwen3coder_tool_parser_deferred_commit.py` runs after `apply_all` in the entrypoint of all 8 Genesis-equipped composes. Defers `is_tool_call_started=True` until `<function=` confirms within a 64-char slack window past the `<tool_call>` tag. **Direct-cmd composes (`dual.yml`, `dual-dflash*.yml`, `dual-nvlink.yml`, `minimal.yml`, `multi4*.yml`, `carnice-bf16mtp.yml`, `qwopus-bf16mtp.yml`) don't currently receive the sidecar** — they have no entrypoint script. Plan: ship local sidecar → validate cross-rig → file upstream PR (with cross-rig evidence and the V2 deferred-commit logic) once the local fix has held up under multi-rig real-world traffic. |
| [#40961](https://github.com/vllm-project/vllm/pull/40961) — Preserve max_seq_len in ubatch metadata during CUDA graph capture | 🟡 Open PR | Confirms the cap-leak pattern: cudagraph capture passes `max_model_len` as `max_seq_len` through ubatch metadata. PR is *fixing a missing pass-through* for SWA models (where seqlen=1 at capture broke kernel selection) — by establishing that `max_model_len` is what gets carried through capture metadata, it cements the source of Cliff 1's max-ctx-dependent FA2 workspace sizing. | Stay at `default` 48K — see FA2 #1011 row + INTERNALS.md Cliff 1 mechanism. |
| [#40069](https://github.com/vllm-project/vllm/issues/40069) — [Tracking] TurboQuant / HIGGS Attention follow-ups | 🟡 Open tracker | Umbrella tracking for TurboQuant + attention backend issues on our stack class. | Watch for cross-references when Cliff 1/2 work lands upstream. |
| [#25543](https://github.com/vllm-project/vllm/pull/25543) — [V0 Deprecation] Remove `max_seq_len_to_capture` | ✅ Merged 2025-09-24 | Important to know: the `--max-seq-len-to-capture` flag (commonly suggested as a Cliff 1 mitigation) **does not exist in V1**. Don't recommend it. | n/a — flag removed. |
| [#39226](https://github.com/vllm-project/vllm/pull/39226) — workspace-resize GPU memory leak fix | 🔵 Merged into v0.20.0; covered by sidecar | Strict `WorkspaceManager.lock()` semantics. After our 2026-05-01 v0.20 + Genesis v7.65 dev tip migration, the surfaces that locked at 0 MB on our config are largely covered by v0.20's revised TQ FA paths ([#40092](https://github.com/vllm-project/vllm/pull/40092)). For the residual cases, our local `patch_workspace_lock_disable.py` sidecar (mounted on every TQ3 compose) downgrades the strict assertion to a one-shot WARNING. P98 covers the same surface but auto-skips on v0.20 due to a drift-marker false-positive (filed as side-note, awaiting Sandermage marker fix). | Drop the sidecar when Sandermage ships the marker fix that re-enables P98 on v0.20. |
| [#40092](https://github.com/vllm-project/vllm/pull/40092) — TurboQuant FA3/FA4 prefill paths | 🔵 Merged into v0.20.0 | TQ + flash-attention 3/4 prefill support. Relevant if/when v0.20 unblocks for us — the FA varlen workspace allocator behavior may change under FA3/FA4 vs the FA2 path we currently hit. | Track. Re-evaluate the `flash_attn_interface.py:300` cliff (Genesis #15) once v0.20 unblocks since FA3/FA4 may have different workspace semantics. FA3/FA4 not enabled on Ampere SM 8.6 anyway (Hopper+ only). |
| [#40941](https://github.com/vllm-project/vllm/pull/40941) — TurboQuant share buffers | 🔵 Merged into v0.20.0 | Sandermage's bare_metal_27b_int4_TQ_k8v4.sh comments call out P98 as the workaround for "WorkspaceManager fix vs vllm#40941". Same WorkspaceManager class that vllm#39226 made strict. | Sandermage's **P98** is the workaround — required for TQ k8v4 on hybrid. Worth a focused investigation: enable P98 on the v0.20-experimental compose to see if it also unblocks vllm#39226's path. |
| [#35975](https://github.com/vllm-project/vllm/pull/35975) — Skip `inputs_embeds` GPU buffer for text-only models ⭐ | 🟡 Open upstream / **local backport active** / **Genesis PN35 lands same fix on dev `f2147ad`** (2026-05-03) | Frees ~444 MiB at boot on Qwen3.6-27B (both `gpu_model_runner.py` + `llm_base_proposer.py` call sites compound; PR claims ~64 MiB but our config has multiple residency points that benefit). **Critical for Cliff 2 closure** at 60K on TP=1 + 24GB — combined with mem-util 0.93, closes the late-stage 50 MiB activation peak. Diagnosed by ChatGPT/Codex as the missing margin; cross-rig validated 2026-05-02 PM. Upstream PR last updated 2026-03-13 (51d stale); Sandermage's PN35 is the practical replacement. | `patch_inputs_embeds_optional.py` ships at compose-entrypoint time (mounted on `long-text.yml` and `long-text-no-mtp.yml`). **Drops out when we bump GENESIS_PIN to dev tip — queued for next-week revisit.** |
| [#37429](https://github.com/vllm-project/vllm/pull/37429) — Hybrid Mamba/attention KV cache sizing | 🟡 Open | Could free more residency without trading mem-util. Larger/riskier than #35975 (architectural Mamba allocation change). Untested on this stack. | Not currently backported. Test on a separate branch when CI signals stabilize. |
| [#37521](https://github.com/vllm-project/vllm/pull/37521) — Spec-decode warmup memory accounting | 🟡 Open | Profiling/KV sizing leaves less false headroom. Genesis PN33 already extends this beyond the original `use_eagle()` gate — so most of the surface is covered, but watch for upstream refinement. | n/a — Genesis PN33 covers the path. |
| [#36598](https://github.com/vllm-project/vllm/issues/36598) — Triton autotuner OOM on Qwen3.5/Qwen3-Next GDN layers (non-SM90 GPUs) | ✅ Closed 2026-03-12, fix shipped via #36599 | Original report of first-inference OOM during Triton autotuning on non-SM90 hardware. Closed because the warmup fix landed. Reading thread is useful context for understanding the GDN kernel autotuner pressure on our hardware class. | n/a — fix in our image. |
| [#36599](https://github.com/vllm-project/vllm/pull/36599) — Warm up Triton autotuner for GDN layers during V1 profiling | ✅ Merged 2026-03-12 (in image SHA `7a1eb8ac`) | Adds `_warmup_triton_kernels()` at V1 profile phase. Warms with B=1, T=64 dummy tensors. Closes the boot-time first-inference autotuner OOM that #36598 reported. **DOES NOT close Cliff 2b** (multi-turn accumulated context) — the warmup uses T=64 but FLA kernels use `do_not_specialize=["T"]` so production T=4128 is the same autotune key, meaning runtime fragmentation isn't from missed autotune; it's from the per-shape Triton kernel binaries staying resident in CUDA context. Confirmed by Codex memo 2026-05-03. | n/a — fix in image; doesn't help our remaining cliff. |
| [#36973](https://github.com/vllm-project/vllm/issues/36973) — `_warmup_prefill_kernels` leaks ~3.4 GiB despite empty_cache | 🟡 Open, RTX 5090-specific | jhsmith409's report — Triton autotuner cubin retention initially suspected but haosdent comment #18-19 traced the bulk to **TMA overhead** scaling with SM count (~22 MiB/SM × 170 SMs on 5090 = 3.7 GiB). Closed via #37700 (TMA-disable for SM12x). **Doesn't apply to Ampere SM86** — no TMA hardware. Useful context though: thread comment #5 explicitly notes Triton autotuner keeps all variants loaded; `empty_cache()` only releases PyTorch's caching allocator, not CUDA-context cubins. | n/a — RTX 3090 doesn't have TMA. |
| [#37700](https://github.com/vllm-project/vllm/pull/37700) — Fix FLA Hopper/TMA misclassification on SM12x desktop Blackwell | 🟡 Open / closes #36973 for SM12x | Uses shared-memory threshold instead of `major >= 9` checks for TMA path selection. SM12x desktop Blackwell only — RTX 5090, DGX Spark GB10. Doesn't apply to Ampere SM86 (no TMA hardware). | n/a — different hardware family. |
| **Cliff 2b — multi-turn accumulated-context OOM (we filed)** | 🟡 Open, [Sandermage genesis-vllm-patches#19](https://github.com/Sandermage/genesis-vllm-patches/issues/19) | DeltaNet `chunk_gated_delta_rule_fwd` holds ~500 MiB of simultaneous live tensors at T=4128. Under multi-turn agent traffic (hermes/openhands/etc.), accumulated KV + this kernel's working set + model + workspace exceeds 24 GiB on 1× 3090. Cliff fires at ~21-26K accumulated context. We tested mem-util tuning, MTP-off, max-num-batched-tokens reduction, TRITON_CACHE_AUTOTUNING, expandable_segments, empty_cache between turns — none close it. Validated 2026-05-03: 6 single-card vLLM variants FAIL v2 continuous soak; only TP=2 / dual.yml passes. Filed with Sandermage proposing streaming refactor of GDN forward intermediates. | **`bash scripts/switch.sh vllm/dual` (TP=2)** for 2× rigs, **`llamacpp/default`** for 1× rigs. See [club-3090#41](https://github.com/noonghunna/club-3090/issues/41) + [docs/CLIFFS.md](CLIFFS.md) "Why TP=2 escapes" / "Why llama.cpp escapes" sections. |
| [#41745](https://github.com/vllm-project/vllm/pull/41745) — Add Gemma4 MTP speculative decoding support (lucianommartins) | 🟢 **Merged 2026-05-06, overlay dropped 2026-05-08** (commit [`595be8f`](https://github.com/noonghunna/club-3090/commit/595be8f)). Today's nightly tag `1acd67a795...` (2026-05-08 06:10 UTC) contains the merge. `dual.yml` + `single.yml` bumped to post-merge nightly; overlay tree `models/gemma-4-31b/vllm/patches/vllm-gemma4-mtp/` retained as fallback (drop in follow-up commit once Phase 2 cycle settles). | First-party MTP for Google's Gemma 4 "assistant" drafter family. Validated on this stack 2026-05-05 (with overlay): 109/142 TPS soak PASS. Re-validated 2026-05-08 (overlay dropped, post-merge nightly): **105.91/141.11 TPS** — within CV of the prior baseline → cleanup is parity-clean. | n/a — closed |
| **Gemma 4 + per-token-head KV on Ampere** ([#40388](https://github.com/vllm-project/vllm/issues/40388), [PR #40391](https://github.com/vllm-project/vllm/pull/40391)) | 🟡 **VENDORED + VALIDATED 2026-05-08** ([commit `f93d312`](https://github.com/noonghunna/club-3090/commit/f93d312) + bench [`160e8fc`](https://github.com/noonghunna/club-3090/commit/160e8fc)). Local rebase of PR #40391 onto post-#41745 main resolved the conflict in `vllm/v1/worker/gpu/attn_utils.py` (combined main's hybrid attn/mamba dispatch with PR #40391's MLA-vs-standard-attention split for `page_size_padded`). Vendored as full 7-file overlay. Compose: `dual/int8.yml`. **Validated dual 3090 Ampere**: 7/7 verify-stress at 98K AND at 262K, plus 137K NIAH recall PASS. Bench: 96/127 TPS at 98K, 95/126 at 262K (~10% TPS cost vs bf16 / 32K for **8.2× context lift**). | Earlier 2026-05-06 Codex investigation memo at [`perheadkv-overlay-comparison.md`](../models/gemma-4-31b/vllm/patches/perheadkv-overlay-comparison.md) erroneously concluded "NOT split-able as an overlay" — that was based on PARTIAL overlays (worker-only or spec-only). A FULL PR #40391 overlay (all 8 files + post-#41745 rebase) works cleanly. Key insight: **INT8 PTH (not FP8 PTH) is the Ampere-target dtype** because Triton `fp8e4nv` kernel is not supported on sm_86 (only `fp8e4b15`/`fp8e5`); FP8 PTH crashes at `_initialize_kv_caches` on Ampere. INT8 PTH dispatches to standard `torch.int8` ops which work on all consumer GPUs. PR #40391's page-size mismatch fix applies to ANY per-token-head KV format — the dtype choice is downstream. Cross-rig validators: cferra (sm_120 Blackwell, FP8 PTH), noonghunna (sm_86 Ampere, INT8 PTH). **Phase 3 (PR #40391 + PR #41703 DFlash drafter combined) BOOT-BLOCKED 2026-05-08** — see also [#41559](https://github.com/vllm-project/vllm/issues/41559) row below for the underlying upstream blocker. 17-file merged overlay parses + compiles, fails `_init_minimal_kv_cache_for_profiling` with `NotImplementedError: page size of the layer is not divisible by the maximum page size` at `kv_cache_utils.py:1068`. **Diagnostic-print at `unify_kv_cache_spec_page_size` (2026-05-08)** reproduced exactly the symptom seantechco described in #41559: drafter silently uses BF16 KV regardless of `--kv-cache-dtype int8_per_token_head`. Three page sizes seen: target Gemma 4 global INT8 PTH 33,280 (padded by PR #40391), target Gemma 4 local INT8 PTH 66,560, DFlash drafter 131,072 (= 16 × 8 × 1024 BF16 K+V at head_dim=256). 131072 / 66560 = 1.97, 131072 / 33280 = 3.94 → no integer ratios → unify rejects. **Phase 3b validation (RedHatAI Gemma-aligned drafter, head_dim=256, num_kv_heads=16)** failed identically — drafter weights/architecture irrelevant; the actual blocker is per #41559: DFlash mandates non-causal cross-attention and every KV-quant backend rejects KV-quant when causal=False. **Why MTP `gemma4_assistant` works but DFlash doesn't**: MTP doesn't require non-causal attention, AND `gemma4_assistant` shares Gemma 4 architecture (same `gemma4.py:438` code path) so PR #40391's INT8 PTH padding propagates uniformly to drafter layers. Phase 3 stacks preserved as forensic artifacts at `models/gemma-4-31b/vllm/patches/vllm-gemma4-dflash-int8/`. | Drop overlay when PR #40391 merges to vLLM main + propagates to a nightly tag. Track: `gh api repos/vllm-project/vllm/pulls/40391 --jq '.state, .merged_at'`. Local exploratory artifacts at `models/gemma-4-31b/vllm/patches/{vllm-perheadkv-hybridpage-fix,vllm-pr40391-perheadkv,vllm-gemma4-fp8-ampere}/` (NOT committed; reference for future iterations). **Phase 3 unblock paths**: (a) patch `qwen3_dflash.py:DFlashAttention.get_kv_cache_spec()` to honor `cache_config.cache_dtype` and return `kv_quant_mode=INT8_PER_TOKEN_HEAD` with appropriate `page_size_padded` (single-file fix, most tractable upstream PR); (b) drafter-isolated KV groups extending DFlash's existing `_get_dflash_isolated_group_ids` to skip page-size unify for draft layers; (c) wait for PR #41703 to merge then re-attempt — fresher base may have unrelated KV-cache refactors that change the picture. **Until then**: long-context Gemma 4 on Ampere via MTP (`dual-int8.yml` 262K, or `dual-awq.yml` 118K) only — DFlash code-optimal long-ctx unreachable. |
---
## Genesis (`Sandermage/genesis-vllm-patches`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [#22](https://github.com/Sandermage/genesis-vllm-patches/issues/22) — PN59 streaming-GDN never engages on chunked-prefill ⚠️ | 🟡 **Open, filed 2026-05-05 by us** | Genesis v7.72.2 advertises PN59 as the structural Cliff 2b fix on 24 GB single cards, but its eligibility check rejects calls with `chunk_indices`/`chunk_offsets` populated — which vLLM's mandatory `--max-num-batched-tokens 4128` always sets. PN59 falls back to vanilla, OOMs at the same `chunk_o.py:161` site. Single-card 24 GB long-context (`long-text.yml` / `long-text-no-mtp.yml` / `long-vision.yml`) regresses vs the prior workarounds. | **Use `dual.yml` / `dual-turbo.yml` (TP=2)** or `llamacpp/default` (different engine, no Cliff 2b). Reproducer + 4 fix proposals in the issue body; awaiting Sander review. |
| [#5](https://github.com/Sandermage/genesis-vllm-patches/issues/5) — P8 ImportError on vLLM v0.20.0 | ✅ Closed (now on v0.20 pin since 2026-05-01) | Originally about P8 ImportError on the v0.20.0 GA tag. We migrated master to v0.20.1rc1.dev16 + Genesis v7.65 dev tip — P8 path no longer fires on our configs. | n/a — pin already moved. |
| [#6](https://github.com/Sandermage/genesis-vllm-patches/issues/6) — P65 PIECEWISE cost quantified | ✅ Closed | We characterized the +22 TPS narrative cost of P65 on Qwen3.6-27B + MTP. Sandermage acknowledged. Will recover when vllm#40914 lands. | Accept the cost on substrate-current; ampersandru's pre-P65 stack avoids it. |
| [#7](https://github.com/Sandermage/genesis-vllm-patches/issues/7) — P67 Triton CompilationError on Qwen3.6-27B | ✅ Closed | Resolved in v7.64 — P67 generalized to non-power-of-2 GQA via `BLOCK_QH = triton.next_power_of_2(HEADS_PER_KV)` + lane_valid mask. Tool-call 0/5 → 7/7 on 2× A5000 validation. | Now safe to enable on 27B configs with v7.64+. |
| [#9](https://github.com/Sandermage/genesis-vllm-patches/issues/9) — P68/P69 8000-char threshold breaks IDE agents | ✅ Closed 2026-05-01 — fix shipped in v7.65+ (50K-char default); we're on v7.69 so the fix lives in our pin | P68 silently rewrites `tool_choice: auto → required`; P69 injects "must use a tool" hint. New 50K threshold clears typical IDE-agent contexts (Cline ~30K / Cursor ~25K / Copilot ~20-25K) while genuine long-history sessions still trigger the reminder. | Composes still have P68/P69 env vars commented out — **enabling them across composes is queued for next-week revisit** (see "Active follow-ups" table above). Until then, manual override: `GENESIS_ENABLE_P68_AUTO_FORCE_TOOL=1 GENESIS_ENABLE_P69_LONG_CTX_TOOL_REMINDER=1`. |
| [#11](https://github.com/Sandermage/genesis-vllm-patches/issues/11) — Cliff 1 mech A FA2 softmax_lse clamp request | ✅ Closed (PN17 in v7.64; default-on across all TQ3 composes since the v0.20 migration) | Sandermage's PN17 lands the clamp at `flash_attn.py`. Active on every TQ3 compose. | n/a — default-on. |
| **Local P104 FA max_seqlen_k runtime clamp** | ✅ Dropped during v0.20 migration | Built 2026-04-30 as `patch_fa_max_seqlen_clamp.py`. Sandermage's PN17 + P15B together cover both layers (FA wrapper + TQ wrapper). Sidecar removed from compose mounts on 2026-05-01. | n/a — Genesis-native. |
| [PR #12](https://github.com/Sandermage/genesis-vllm-patches/pull/12) — P101 anchor drift fix | ✅ Closed; on v0.20 pin since 2026-05-01 | P101 anchor matches on `0.20.1rc1.dev16+g7a1eb8ac2`. | n/a — pin matches. |
| [PR #13](https://github.com/Sandermage/genesis-vllm-patches/pull/13) — PN12 anchor drift fix | ✅ Closed; on v0.20 pin since 2026-05-01 | PN12 anchors match natively. Local `patch_pn12_ffn_pool_anchor.py` sidecar removed. | n/a — pin matches. |
| [#14](https://github.com/Sandermage/genesis-vllm-patches/issues/14) — P38 silently no-op'd on TurboQuant KV path (we filed 2026-05-01) | ✅ Closed via P38B in Genesis v7.65 dev tip | Sandermage shipped **P38B** — text-patches `turboquant_attn.py` source to inject a delegate hook at the start of `_continuation_prefill` body. Active via `GENESIS_ENABLE_P38B_COMPILE_SAFE=1` on every TQ3 compose since 2026-05-01. | n/a — Genesis-native. |
| [#15](https://github.com/Sandermage/genesis-vllm-patches/issues/15) — FA varlen kernel workspace cliff at flash_attn_interface.py:300 (we filed 2026-05-01) | ✅ Closed via P15B in Genesis v7.65 dev tip | Sandermage shipped **P15B** — direct backport of our suggestion. Active via `GENESIS_ENABLE_P15B_FA_VARLEN_CLAMP=1` on every TQ3 compose since 2026-05-01. Empirically the cliff also doesn't reproduce on v0.20 (vllm#40092 changed workspace allocator behavior) — covered from two directions. | n/a — Genesis-native. |
| [#16](https://github.com/Sandermage/genesis-vllm-patches/issues/16) — PN25 worker-spawn registration | 🟡 Sander shipped d92bcb3 (v7.65) + Library refactor (v7.66); both fail on TP=1 | v7.65 used `@torch.library.custom_op` (failed at `infer_schema` inside dynamo trace). v7.66 refactored to `direct_register_custom_op` + `Library("genesis", "FRAGMENT")` — fails at `instantiate_user_defined_class_object` inside dynamo trace. Same root cause: torch.library construction inside trace context disallowed on TP=1 spawn. Cross-rig data on Sander's [discussion #19 reply](https://github.com/noonghunna/club-3090/discussions/19#discussioncomment-16785590). | Local `patch_pn25_genesis_register_fix.py` v3 — text-patches `activation.py` to register at module-import time, BEFORE any trace. Survives both v7.65 and v7.66 mechanisms. PR-ready upstream. |
| [#17](https://github.com/Sandermage/genesis-vllm-patches/issues/17) — DS conv state spec-decode crash | 🟡 Sander shipped a9977d8 (PN30) but `.contiguous()` is layout-incorrect | Sander's PN30 materializes `state[src, :, offset:].contiguous()` (compact 10240×5) and raw-memcpys into `state[dest]` (strided 10240×6) → corrupts DS row strides → eventual TQ store CUDA assert several layers downstream. Diagnosis credit: ChatGPT/Codex CLI cross-check 2026-05-02. Sent corrected fix to Sander. | Local `patch_pn30_dst_shaped_temp_fix.py` — patches `collect_mamba_copy_meta` to build dst-shaped temp instead of compact. Reuses Sander's `_GENESIS_PN30_TEMP_TENSORS` lifecycle. Validated on all 4 TQ3 composes; probes 4 + 5 pass cleanly. PR-ready upstream. |
| [#15](https://github.com/Sandermage/genesis-vllm-patches/issues/15) — PN31 FA varlen persistent out | 🟡 Sander shipped 753344b (PN31, default OFF); doesn't fit on 24 GB | Per-shape persistent buffer growth + PN12+PN25 pool residence outpaces activation budget at DeltaNet `chunk_fwd_o` on 24 GB single GPU. Sander explicitly flagged he couldn't validate on 24 GB. | Lower mem-util to 0.95 — gives enough activation headroom to close the 25K tool-RETURN path PN31 was meant to fix, without needing PN31. Cross-rig data on Sander #15 comment. |
| **PN33 spec-decode warmup K-aware** (Sander v7.66 fc89395, default ON) | 🟡 Partial close on TP=1 | Backport of vllm#37521 EXTENDED to MTP/ngram. Sander claimed it closes both ampersandru's mid-stream OOM AND our workspace_lock AssertionError. Cross-rig 2026-05-02: closes BOOT-time profile_run workspace_lock ✅, but runtime decode `turboquant_attn.py:1350:_decode_attention` AssertionError still fires ❌. | Local `patch_workspace_lock_disable.py` sidecar still required for runtime decode. Drop when upstream covers the runtime path. |
| **P98 marker false-positive on v0.20** (we filed 2026-05-01 in [#9 thread](https://github.com/Sandermage/genesis-vllm-patches/issues/9#issuecomment-4359541875)) | 🟡 Side-noted to Sandermage; awaiting his call on fix | P98's drift detection auto-skips on v0.20 (`UNIFORM_SINGLE_TOKEN_DECODE` marker false-positive) but the strict workspace lock still fires rare paths P98 was supposed to revert. | Local `patch_workspace_lock_disable.py` sidecar (mounted on every TQ3 compose) relaxes the strict assertion to a one-shot WARNING. Drop when Sandermage ships either a marker fix or P98 with explicit env-override. |
| **PN30 v7.68 part3 drift-marker false-positive** (we filed in noonghunna/club-3090#19 cross-rig retest) | ✅ Closed in v7.69 (commit 2db18df) | Part3's `upstream_drift_markers=["[Genesis PN30"]` (generic prefix) matched markers parts 1+2 wrote on the same file. Part3 skipped as `upstream_merged` → apply_all FAILS → vLLM aborts. v7.69 tightened to `[Genesis PN30 v7.68 dst-shaped]` (specific). | n/a — fixed in v7.69. |
| **P103 setattr lost on `exec vllm serve`** (we filed in noonghunna/club-3090#19) | ✅ Closed in v7.69 | v7.68 P103's `setattr` ran in entrypoint shell but was lost on `exec vllm serve` worker spawn (process image replaced). v7.69 ships chunk.py self-install hook appended to end-of-file — survives any startup mechanism. | n/a — fixed in v7.69. |
| **PN32 v1 chunked at wrong level** (we filed in noonghunna/club-3090#19) | ✅ Closed in v7.69 (PN32 v2) | PN32 v1 chunked outer-level inputs but inner FLA call still got full-prompt cu_seqlens, allocating full h tensor regardless. v7.69 PN32 v2 patches `_forward_core` directly + threads `last_recurrent_state` between chunks. | n/a — fixed in v7.69. |
| [#18](https://github.com/Sandermage/genesis-vllm-patches/issues/18) — P103 cu_seqlens=[0,T] single-seq case is bypassed (we filed 2026-05-02 PM) | 🟡 Open / v7.70 proposal | P103's gate currently bypasses chunking for ANY non-None cu_seqlens, but `cu_seqlens.shape[0] == 2` (single sequence boundary) is semantically dense B=1, not multi-seq varlen. Fix admits the chunked path on real serving. Diagnosis: ChatGPT/Codex CLI. Cross-rig observation: P103 chunked path never engages on real config because vLLM's outer chunked-prefill caps T at `max_num_batched_tokens=4128` (well below `_MAX_T=16384`), so the gate-fix is semantically correct but doesn't independently close 60K Cliff 2 on TP=1+24GB. | n/a yet — gate fix queued for v7.70. Real Cliff 2 closure on this config comes from [vllm#35975 backport](https://github.com/vllm-project/vllm/pull/35975) + mem-util 0.93 (see vLLM section above + [`docs/CLIFFS.md`](CLIFFS.md)). |
---
## FlashAttention 2 (`Dao-AILab/flash-attention`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [#1011](https://github.com/Dao-AILab/flash-attention/issues/1011) — Variable memory allocation with varlen kernels | 🔴 Open since 2024, no fix | **Cliff 1 root cause.** `softmax_lse` is allocated as `[num_seqs, num_heads, max_seqlen]` — sized by `max_seqlen` parameter, NOT actual `cu_seqlens`. So a 25K-token chunked-prefill at `max_model_len=86K` allocates softmax_lse for 86K, not 25K. This is why Cliff 1 fires harder at higher max-ctx even when the actual prompt is the same. | None. Stay at `default` 48K (or `tools-text` 75K with PN8 mitigation). FA2 redesign of softmax_lse format would be the upstream fix. |
## flash-linear-attention (`fla-org/flash-linear-attention`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| **Cliff 2 — DeltaNet GDN forward OOM at 50–60K single-prompt** | 🔴 Open, **no upstream issue filed yet**. **Confirmed cleared on dual TP=2** (this rig, 2026-04-29 — see DUAL_CARD.md "237K single-prompt verified"). | The `chunk_gated_delta_rule_fwd` kernel allocates intermediate buffers proportional to `seq_len`. Fires on single-card regardless of mem-util. On dual TP=2 the activation memory splits across cards and the cliff doesn't fire — verified at 237K single-prompt prefill on `dual.yml` (~830 tok/s prefill, matches Sandermage's 262K @ 311s on 2× A5000). Sandermage explicitly punted on the single-card fix (genesis-vllm-patches issue #1: *"can't fix this short of multi-GPU TP=2 or upstream fla.ops changes"*). Likely the same architectural pattern as FA#1011 — recurrent state buffer pre-allocated by max_seq_len. | Single-card: use `tools-text.yml` (75K cap) or `llamacpp/default` (262K, different engine). Dual: `dual.yml` clears at ≥237K. |
---
## FlashQLA (`QwenLM/FlashQLA`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| **Ampere SM 8.6 / Ada SM 8.9 port** | 🔴 No issue filed; tweet to @QwenLM drafted but not yet posted | FlashQLA is QwenLM's TileLang DeltaNet kernels — would fix Cliff 2 if it ran on Ampere. Currently SM90+ only. | None. Watch the repo for Ampere support; revisit when an issue is filed and a port is on the roadmap. |
---
## Luce DFlash (`Luce-Org/lucebox-hub`) — separate llama.cpp fork (NOT our vLLM dual-dflash)
**Heads-up — naming clarification:**
- This section tracks **`Luce-Org/lucebox-hub`** (a llama.cpp fork from Luce). **As of 2026-05-04 this is no longer single-card-only** — see "Dual-GPU split landed" below.
- Our **`dual/dflash.yml` / `dual-dflash-noviz.yml`** (vLLM TP=2 dual-card) IS shipping and is the recommended DFlash path on this stack today. Both consume the **same draft model** (`z-lab/Qwen3.6-27B-DFlash`), but the engine + topology differ. Don't confuse the two.
### 🆕 Dual-GPU split landed (2026-05-02 + 2026-05-04)
Two @weicj PRs shipped that change the lucebox-hub serving topology. **Target weights on one GPU + DFlash draft (or PFlash drafter) on a separate GPU** — heterogeneous spec-decode, not weight-sharded TP. Each model lives entirely on its own card; they communicate at spec-decode boundaries via peer copies.
- [**lucebox-hub PR #80** — `bench(dflash): add dual-GPU target/draft split harness`](https://github.com/Luce-Org/lucebox-hub/pull/80) (merged 2026-05-04). New flags `--target-gpu` / `--draft-gpu` (also `DFLASH_TARGET_GPU` / `DFLASH_DRAFT_GPU` env). Validation on dual RTX 2080 Ti 22 GB: HumanEval 10-prompt at **51.86 tok/s, AL 7.09, 44.3% accept** on Qwen3.5-27B Q4 target + z-lab DFlash draft.
- [**lucebox-hub PR #78** — `bench(pflash): add dual-GPU PFlash phase-split harness`](https://github.com/Luce-Org/lucebox-hub/pull/78) (merged 2026-05-02). New flag `--pflash-gpu` + persistent `pflash_daemon`. Validation on same hardware: **single-GPU co-resident passes NIAH at 24,573 source tokens; dual-GPU phase split passes at 262,125 source tokens (10.7×).** Compressed context reaches 13,229 tokens at 262K source.
**Implication for our 2× 3090 stack:** the single-card limitations we documented (65K max_ctx, draft VRAM competing with target activations) are addressed by dual-GPU split. Target Qwen3.5-27B Q4_K_M gets a full 24 GB on GPU 0; DFlash draft + PFlash drafter live on GPU 1. **No NCCL/allreduce overhead per token** since each model lives entirely on its own card — should be faster per-stream than SGLang TP=2 + DFlash for single-stream workloads. Bench tracked at task #229 (queued, not yet executed locally — PR #80 is hours old as of this entry). **Qwen3.6-27B draft remains under training** so the dual-GPU benefit applies primarily to the stable Qwen3.5-27B + DFlash pair today.
Re-benched 2026-04-30 PM on Qwen3.6-27B Q4_K_M + matched z-lab/Qwen3.6-27B-DFlash draft (under training). Open issues against single-card lucebox-hub follow:
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [z-lab/Qwen3.6-27B-DFlash](https://huggingface.co/z-lab/Qwen3.6-27B-DFlash) — draft model still under training | 🟡 Snapshot 2026-04-26 | Narrative AL ~3.7, code AL ~7.0 on `Luce-Org/lucebox-hub` (single-card llama.cpp). When training finishes, expected to climb toward Qwen3.5 reference (8.31 HE, 7.04 Math). The same caveat applies to vLLM `dual-dflash.yml` — published 82/125 TPS in [`docs/DUAL_CARD.md`](DUAL_CARD.md) was measured against this 2026-04-26 snapshot at peak code-prompt conditions; AL on real agent traffic will be lower until z-lab tags training-complete. | Re-test when z-lab tags training-complete. The vLLM dual-dflash path remains shipping — see [DUAL_CARD.md](DUAL_CARD.md) — but treat its numbers as a snapshot. For autonomous coding agents on dual-3090 today, `dual.yml` (FP8 + MTP) is the recommended robust path. |
| **Build fragility on `dflash` main HEAD** | 🔴 Reproducible 2026-04-30 PM | `cmake --build` errors with `ggml_turbo_wht` and `GGML_TYPE_TQ3_0` undefined. Required submodule commit `b6ffab4a9` not auto-fetched. Cross-rig signal — fresh clone fails. | After clone: `cd dflash/deps/llama.cpp && git fetch origin && cd ../../.. && git submodule update --init`. |
| **Daemon-mode "empty prompt" regression** | 🔴 Reproducible 2026-04-30 PM | After streaming requests, subsequent requests return `"empty prompt"` from the test_dflash daemon. Server keeps accepting requests but generates 0 tokens. Forces restart. | Restart server between request flavors; avoid mixing streaming + non-streaming. |
| **`enable_thinking` chat_template_kwargs honored differently than vLLM** | 🟡 Behavioural difference | Test sends `enable_thinking=true` and expects `reasoning_content` populated. Luce returns `content` directly. Not a missing feature, but breaks our `verify-full.sh` check 6. | Don't treat the thinking-mode test as a Luce-correctness signal until the chat-template path is documented. |
| **Greedy only** | 🟡 Documented limitation | `temperature` / `top_p` accepted but ignored. Real downside for creative-writing workloads. | Use vLLM long-text/long-vision when sampling matters. |
| **Prefill OOM in `fattn-chunked.cu` on 25K+ prompts at Q8_0 KV** | 🟡 Open (configuration trade) | Chunked flash-attention CUDA OOMs on large prefill at default Q8_0. **TQ3 KV (`DFLASH27B_KV_TQ3=1`) closes it** at max_ctx=65K — verify-stress passes 791 chars / finish=stop. Higher max_ctx (131K) reopens it. | Always set `DFLASH27B_KV_TQ3=1` for stress-test-passing config. Cap max_ctx at ~65K. |
| [**PFlash — long-context prefill accelerator**](https://www.lucebox.com/blog/pflash) (sibling tech to DFlash, same Luce-Org/lucebox-hub repo) | 🟢 **Public release 2026-04 + dual-GPU split shipped 2026-05-02 (PR #78)** | **Speculative prefill + block-sparse attention.** Compresses 128K prompts to ~6.5K tokens (`keep_ratio=0.05`) before target prefill. Single-card claimed: TTFT 24.8s vs 257s vanilla llama.cpp at 128K (~10.4× speedup). **Dual-GPU phase split (PR #78) extends the passing source-context ceiling from ~24K (single-card co-resident) to 262K (~10.7×) on dual 22 GB cards** — NIAH key/answer retained at 262K. C++/CUDA only, lives inside the lucebox-hub server stack. PFlash sits *in front of* DFlash decode: PFlash accelerates prefill, DFlash accelerates generation. **For 2× 3090 deployments**: pin PFlash drafter to GPU 1 via `--pflash-gpu`, target on GPU 0. The single-card-coresident limit (was the binding blocker for our use) no longer applies. MIT license. **Open exploration**: bench PFlash + DFlash dual-GPU vs vLLM `dual-dflash.yml` (185K, 82/125 TPS on 2× 3090) on TTFT-bound workloads. Tracked at task #229. | Re-evaluate as a club-3090 shipping option once we (a) reproduce the 262K passing source-ctx claim on 2× 3090 with verify-stress + soak-continuous + bench, OR (b) an upstream-vLLM port lands. The dual-GPU split removes the single-card co-residency blocker; remaining blockers are daemon-mode bugs (greedy-only, no vision, "empty prompt" regression) carried over from the single-card history. |
---
## llama.cpp (`ggml-org/llama.cpp`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [PR #21089](https://github.com/ggerganov/llama.cpp/pull/21089) — TurboQuant KV mainline | 🟡 Open (CPU first, CUDA follow-on) | When CUDA path lands, `turbo3` becomes a first-class option on llama.cpp. Naming will migrate from `turbo3` → `tbq3_0`. | Use Tom's fork for now: [llama-cpp-turboquant](https://github.com/tdraxl/llama-cpp-turboquant). |
| **Q3_K_XL TPS regression (28.5 TPS @ 262K → 21 TPS today)** | 🔴 Suspected, no upstream issue filed | Measured 2026-04-23 vs 2026-04-28: same model, same hardware, 28.5 TPS dropped to 21 TPS between commits `9ab47e7d8` and `0d0764dfd`. Bisect or file. | None — we're on the slower commit. Tracked in [club-3090 TODO](https://github.com/noonghunna/club-3090) (private). |
| [PR #22673](https://github.com/ggml-org/llama.cpp/pull/22673) — MTP support (am17an, `mtp-clean`) | 🟡 Open, unmerged | First-party MTP for llama.cpp via an MTP head baked into the GGUF (RDson republished `Qwen3.6-27B-MTP-Q4_K_M-GGUF` with the head wired). Benched on 1× 3090 (2026-05-05): **+34% narrative TPS at `n-max=3` (22.83 → 30.69)**, ~57% accept. Code at `n-max=5` hit 31.9 TPS. **NOT a club-3090 recommendation yet.** Reasons: (1) unmerged → forces every cross-rig user to compile am17an's fork or maintain a custom image; (2) q8_0 KV ceiling caps context at ~64-80K — current `llamacpp/default` ships 262K, trading that for +34% TPS isn't worth it for the cliff-immune audience; (3) MTP forces `n_parallel=1` (kills `llamacpp/concurrent.yml`); (4) RDson GGUF doesn't bundle mmproj (vision regression). Audience for this is empty — vLLM dual-turbo already gives 170 TPS for users wanting max single-stream throughput. | None recommended. Re-evaluate when PR merges + q4_0 KV variant tests recover 128K+ context + cross-rig data lands. Detailed bench + reasoning is documented out-of-tree in this stack's `learnings/qwen3.6-35b-a3b.md` ("llama.cpp MTP — PR #22673 path" subsection). |
---
## transformers (`huggingface/transformers`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| [#45283](https://github.com/huggingface/transformers/issues/45283) — Qwen3.5 GGUF support | ❌ Closed without fix 2026-04-28 (no associated PR; closing event has `source: null`, last comment was just `cc @SunMarc` — looks won't-fix or stale-bot) | Was tracked as the missing piece (alongside vllm#38140 / vllm#37797) for Qwen3.5/3.6 GGUF on vLLM/SGLang. Won't be picked up via transformers — **llama.cpp remains the only GGUF path** for this model family. | llama.cpp path. Don't expect a vLLM/SGLang GGUF route for Qwen3-Next family. |
| **transformers ≥ 5.8.0 required for `gemma4_assistant`** | 🔵 Released 2026-05-05 | First version with native `gemma4_assistant` model class (Google's Gemma 4 MTP drafter). vLLM nightly `:nightly-01d4d1ad3` ships transformers 5.7.0 → AutoConfig rejects the drafter checkpoint at validation time. | `dual.yml` entrypoint runs `pip install --upgrade transformers==5.8.0` before exec'ing vllm serve. Drop the line when vLLM nightly rebuilds against transformers ≥ 5.8.0. |
---
## SGLang (`sgl-project/sglang`)
| Issue / PR | Status | Why it matters | Workaround |
|---|---|---|---|
| **Same Marlin pad-sub-tile-n bug as vllm#40361** | 🔴 Not filed; same kernel-line fix applies | Blocks Lorbus INT4 + EAGLE on SGLang. We haven't filed an SGLang PR. | None on SGLang. Use vLLM (with our patched fork) or wait for SGLang to pick up the upstream Marlin fix. |
| **DeltaNet KV rollback (vllm#39931 cross-engine)** | 🔴 Same architectural issue | Blocks EAGLE on Qwen3-Next family in SGLang too. | None — see vllm#39931. |
---
## Community templates / model assets (Hugging Face)
External-but-load-bearing resources that aren't issue trackers (no PR / merge state to track). Watch list — re-check when upstream Qwen / Gemma official templates change, or when these resources update.
| Resource | Status | Why it matters | Drop trigger |
|---|---|---|---|
| **[froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)** — community fork of the default Qwen 3.5 / 3.6 chat templates fixing seven documented bugs (empty `<think></think>` spam in past turns, `</thinking>` hallucination on Qwen 3.6, unclosed thinking before tool call, no-user-query crash in agentic loops, `developer` role rejection, `\|items` filter for C++ engines, type-aware `tojson`). Surfaced by @troymroberts in [discussion #121](https://github.com/noonghunna/club-3090/discussions/121). Vendored snapshot at [`models/qwen3.6-27b/vllm/patches/froggeric-chat-template/chat_template.jinja`](../models/qwen3.6-27b/vllm/patches/froggeric-chat-template/chat_template.jinja); mounted default-on across all 22 vanilla Qwen 3.6 composes via `--chat-template`. Carnice and Qwopus composes intentionally excluded (ship their own bespoke templates). | 🟡 First-pass A/B 2026-05-12 — **+15pp on `hermesagent-20` (45% → 60%)** on Qwen 3.6 27B INT4 + INT8 PTH KV (dual 3090). 7 other packs flat. Control run (revert template, same commit) pending to isolate PR #35936 confound. | Replace with default model template if upstream Qwen pushes equivalent fixes. **Watch for**: froggeric updates the template (Qwen 4 support, additional bug fixes), or Qwen upstream lands their own version. |
---
## Filing conventions
When you file or learn of a new upstream issue:
1. **Add a row** to the appropriate section of this file. Include the link, status emoji, one-line "why it matters," and the local workaround (if any).
2. **Cross-link** from any code, compose comment, or doc that depends on the workaround back to the row in this file (e.g., `# See docs/UPSTREAM.md — vllm#40361`).
3. **Update the row** when status changes — closed, merged, propagated, replaced. Don't delete; if a row is no longer load-bearing, mark it ✅ Resolved or ❌ Closed without fix and leave it as historical context.
4. **Bump the relevant pin** when an upstream lands (Genesis commit, vLLM nightly, llama.cpp commit). Add a CHANGELOG entry citing the upstream PR.
When you file an issue against an upstream repo from this work, **link back to club-3090** in the body so the upstream maintainer can see the affected user surface and re-test if needed.
---
## Related reading
- [`models/qwen3.6-27b/INTERNALS.md`](../models/qwen3.6-27b/INTERNALS.md) — model-specific deep dives (DFlash forensics, MTP head, AutoRound rationale)
- [`models/qwen3.6-27b/vllm/patches/README.md`](../models/qwen3.6-27b/vllm/patches/README.md) — local patches (tolist, Marlin pad fork, Genesis env-var matrix)
- [`AGENTS.md`](../AGENTS.md) — repo-wide conventions, including the rule that this file is the upstream-tracking single source of truth