Closes / addresses 3 reported issues + adds requested feature:
#7 vid (PORT not honored, MODEL_DIR vs MODELS_DIR confusion):
- All 8 vLLM compose files now use "${PORT:-XXXX}:8000" so .env PORT
flows through. Defaults preserved per-variant (8020 single, 8010-8013
dual). llama.cpp composes already had this pattern.
- scripts/switch.sh: load .env early; per-variant default-port table;
new resolve_ready_url() picks PORT > variant default for the readiness
probe.
- scripts/launch.sh: same default-port table; final endpoint URL printed
to user reflects actual mapped port.
- .env.example: ⚠ box callout that variable names are CASE-SENSITIVE
(MODEL_DIR singular, NOT MODELS_DIR plural — silently ignored).
New PORT section documenting per-variant defaults.
#4 timxx (tools-text.yml fails "Free memory ... less than desired"):
- docs/FAQ.md: new entry "Container fails to start: Free memory..."
explaining the vLLM startup check, the two workarounds (free VRAM /
lower mem-util), and which configs hit it most often (0.97+ mem-util).
- Compose defaults unchanged (0.97 stays the right design target on
headless rigs); the FAQ documents the workaround for users with X11.
#1 fabriciomalta (per-config VRAM column):
- docs/SINGLE_CARD.md: TL;DR table now has VRAM column with mem-util.
- docs/DUAL_CARD.md: TL;DR table same + footnote explaining per-card
semantics and which dual configs would/wouldn't fit on 2× 20 GB cards
(relevant to fabriciomalta's 2× 3080-20GB use case).
#2 tenitram (empty responses) — fixed in master via aab8ff4
(P68/P69 disabled). Closed with reply pointing at the fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Single canonical reference for everything we know about Cliff 1
(FA2 softmax_lse cap-leak) and Cliff 2 (fla.ops GDN forward
intermediate buffer): TL;DR table, empirical bisection with stack
traces, root-cause walk-through, why earlier "FFN intermediate
buffer" framing was wrong, why mem-util doesn't help, why PN8
closes Cliff 1 on tools-text but not on TQ3 paths, why llama.cpp
dodges both structurally, alternative attention backends with
feasibility, who-can-fix-it landscape (Sandermage, Tri Dao, fla-org,
QwenLM, us at any difficulty), recommended path forward, and
re-test triggers.
Cross-linked from FAQ.md and README.md.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
After bisecting long-vision config space (192K/128K/96K/86K at 0.98
and 0.92 mem-util) and second-opinion synthesis from ChatGPT +
DeepSeek + vLLM source review, the actual root cause is:
softmax_lse in flash_attn_varlen_func is allocated as
[num_seqs, num_heads, max_seqlen] — sized by the max_seqlen
parameter, NOT the actual cu_seqlens.
vLLM passes attn_metadata.max_seq_len; during cudagraph capture
that's set to max_model_len. So a 25K-token tool prefill at
max-model-len=192K allocates softmax_lse for 192K, eating the
activation headroom. The 50-138 MiB OOMs we'd been observing
are downstream of this leak.
Empirical OOM site (verified in our docker logs): _vllm_fa2_C.varlen_fwd
in flash_attn_varlen_func. Upstream root cause: Dao-AILab/flash-
attention#1011 (open since 2024). vLLM cap-leak path: vllm#40961.
Earlier "FFN intermediate buffer" characterization was wrong.
Updates:
- UPSTREAM.md: new FA2 section (Dao-AILab/flash-attention#1011);
added vllm#40961 (cudagraph capture max_seq_len pattern), vllm#40069
(TurboQuant follow-ups tracker), and vllm#25543 (V0 deprecation
removed max_seq_len_to_capture, so commonly-suggested mitigation
doesn't apply on V1 nightly)
- FAQ.md: corrected Cliff 1 explanation
- SINGLE_CARD.md: corrected "Cliff 1 still fires" caveat
- CHANGELOG: documented bisection + revision
- memory/qwen36_27b_prefill_cliffs.md: revised Cliff 1 mechanism;
noted Cliff 2 likely shares the same architectural pattern
Practical implication: no new variant ships. Default 48K + 0.92 +
TQ3 + vision stays the prefill-safe ceiling — pushing higher requires
upstream fix at FA repo, not config tuning. tools-text.yml (75K + FP8
+ PN8 closes Cliff 1) remains the IDE-agent path.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
fast-chat (20K, fp8, vision) and default docker-compose.yml (48K, TQ3,
vision) had effectively the same TPS post-PN8. fast-chat's only
remaining differentiator was "smaller context = ~3s faster boot," and
20K is actively bad for IDE-agent users (Copilot tool-schema preamble
alone hits 20K). Net negative — removed.
Default compose was missed in the previous P68/P69 fix — it had the
same env vars enabled and the same silent-stop bug above 8000 chars.
Both now disabled with the same explanatory comment.
Updated:
- scripts/switch.sh, scripts/launch.sh — drop the variant
- docs/SINGLE_CARD.md, FAQ.md, engines/VLLM.md, model + vllm + patches
READMEs — references removed or pointed to default/tools-text
- All sibling compose YAML "see also" tables — fast-chat row removed,
tools-text row repurposed for IDE-agent guidance
- CHANGELOG entry; old historical entries kept as-is (append-only)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
P68 (auto force tool_choice=required) and P69 (inject "must use a
tool" reminder) silently fired on prompts > 8000 chars — every IDE
agent (Cline, Cursor, OpenCode, Copilot Gateway) blew past that
threshold instantly and got silent finish_reason=stop with no
content + no tool_calls on greetings or clarifying questions.
Bisection on club-3090#2 (HoodOG1 + tenitram):
state A (P64+P68+P69+PN8): broken
state B (P64+P69+PN8, P68 off): still broken — model loops on
"I cannot respond with plain text" then stops mid-reasoning
state D (P64+PN8, P68+P69 off): clean — greeting → plain-text
reply; tool request → clean read_file call
P64 (qwen3coder MTP streaming early-return fix) and PN8 (FP8+MTP
draft online-quant memory savings) stay enabled — real bugfixes,
no user-intent override.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
User feedback: navigating to relevant docs was cumbersome. The natural
first decision is "1 GPU or 2 GPUs?" and the existing docs mixed model-
specific reference with deployment guidance.
New navigation:
- README.md adds a "Pick your path" pivot pointing at hardware-axis pages
- docs/SINGLE_CARD.md — 1× 3090 deployment menu (workload → compose →
TPS) with all single-card configs (vLLM + llama.cpp), VRAM budget,
prefill cliffs explained operationally, what single-card can't do
- docs/DUAL_CARD.md — 2× 3090 mirror (4 dual variants + TP=2 explainer +
what dual unlocks vs single + Marlin pad fork dependency)
Slimming:
- models/qwen3.6-27b/README.md: dropped duplicated variant tables
(now in GPU-count pages); kept model-specific content (quants, Genesis
patch surface table, what's working / not, VRAM diagram)
- models/qwen3.6-27b/USE_CASES.md: deleted. Per-workload content
absorbed into the GPU-count pages (deduplicated). Troubleshooting
list moved to docs/FAQ.md as a new "Troubleshooting" subsection.
Image-token cost / vision specifics absorbed into SINGLE_CARD.md.
Reference updates: 8 files updated (engines/VLLM.md, engines/README.md,
COMPARISONS.md, ARCHITECTURE.md, EXAMPLES.md, FAQ.md, INTERNALS.md,
top-level README) — all USE_CASES.md links re-pointed to SINGLE_CARD/
DUAL_CARD where appropriate.
Net delta: -167 lines (was 197 in USE_CASES + duplicated tables in
model README; now 342 lines split between SINGLE_CARD + DUAL_CARD with
content deduplicated against each other).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Documents the two compatibility issues users will hit with VS Code's
Copilot LLM Gateway:
1. Tool-schema preamble is ~20K tokens — fast-chat.yml's 20K cap is
too small. Recommended pick is tools-text.yml (75K + fp8 + PN8,
Cliff 1 closed since Genesis v7.62.x).
2. Copilot probe-style requests with max_tokens=64 truncate tool-call
JSON mid-string. With tool_choice: required + minItems: 1 in their
structured-outputs schema, the model must emit a tool call that
takes real arguments — won't fit in 64 tokens. Manifests as
"empty response" client-side. Server-side correct.
Background + debug-log analysis from tenitram on club-3090 #2.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
scripts/setup.sh — GENESIS_PIN bumped from bf667c7 (v7.54) to 917519b
(v7.62.x release, 2026-04-29). New patches: PN8 (MTP draft online-quant
propagation, backport of vllm#40849), PN11 (Quentin-M streaming tool-call
IndexError fix vllm#41142), per-GPU profile auto-rec, k8v4 unlock on
hybrid GDN via P4+P98.
PN8 enabled on FP8 paths only:
- tools-text.yml: -900 MiB at boot, Cliff 1 25K tool prefill closes,
-7% code TPS. Net win — production-safe for tool-using agents.
- fast-chat.yml: -800 MiB at boot, no cliff to test at 20K, -4.7% code
TPS. Free VRAM is useful for tighter mem-util configs.
PN8 not enabled on TQ3 paths (default 48K, long-vision, long-text) or
dual configs:
- default 48K: PN8 is no-op on TQ3 + 0.92 (plenty of headroom already)
- long-vision: PN8 grows KV pool 230 MiB and lifts engine ceiling 192K
→ 198K, but does NOT close Cliff 1 — the 138 MiB allocate is an FFN
intermediate-buffer activation peak (intermediate_size × max-num-
batched-tokens), not a draft-model footprint
- long-text: engine ceiling at 206K is gated by attention-block-size
divisor, not KV; PN8 has nothing to give
- dual.yml: deliberately Genesis-less by design; not worth restructuring
Verify-full passes on default 48K + v7.62.x without PN8 (8/8). Verify-
stress on tools-text + PN8 passes all checks including the 25K tool
prefill that was the launch-tweet headline caveat.
Cross-rig data shared with Sandermage:
https://github.com/noonghunna/qwen36-27b-single-3090/issues/1#issuecomment-4343317153
Docs updated: cross-cutting CHANGELOG, per-model CHANGELOG, USE_CASES.md
(Cliff 1 closure note on tools-text), FAQ.md (Cliff 1 entry + new PN8
entry).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>