This branch migrates the entire vLLM stack from `dev205+g07351e088` + Genesis
v7.64 to `0.20.1rc1.dev16+g7a1eb8ac2` + Genesis v7.65 dev tip (commit
`d89a089`). v7.65 is on Sandermage's `dev` branch — explicitly the cross-rig
testing surface he requested in discussion #19; he'll merge dev→main once we
both confirm stable. Pin gates restated when that lands.
What changes
------------
Pin migration:
- vLLM image: nightly-07351e08... → nightly-7a1eb8ac2... (dev205 → v0.20.1rc1.dev16)
- Genesis: 64dd18b (v7.64) → d89a089 (v7.65 dev tip)
Sidecar churn:
- DROPPED: patch_pn12_ffn_pool_anchor.py (PN12 native on v0.20)
- DROPPED: patch_pn12_compile_safe_custom_op.py (Genesis P38B in-source hook)
- DROPPED: patch_fa_max_seqlen_clamp.py (Genesis PN17 + P15B)
- ADDED: patch_workspace_lock_disable.py (relaxes vllm#39226 strict assertion;
P98 covers same surface but auto-skips on v0.20 due to drift-marker false
positive — pending Sandermage marker fix)
Env-var alignment to Sandermage's PROD set (start_27b_int4_TQ_k8v4.sh@dev):
- FIXED naming bugs that silently no-op'd patches:
- PN9_INDEPENDENT_DRAFTER_ATT → _ATTN (was silently OFF)
- PN22 → PN22_LOCAL_ARGMAX_TP (was silently OFF)
- PN26_BLOCK_KV → PN26_SPARSE_V_BLOCK_KV (fell back to default 4, not 8)
- PN26_NUM_WARPS → PN26_SPARSE_V_NUM_WARPS
- PN26_THRESHOLD → PN26_SPARSE_V_THRESHOLD (fell back to default 0.001, not 0.01)
- ADDED explicit-OFFs to match Sander's PROD verbatim:
- P78_TOLIST_CAPTURE_GUARD=0 (we use our own patch_tolist_cudagraph.py)
- P81_FP8_BLOCK_SCALED_M_LE_8=0 (FP8-specific, no-op on TQ3)
- P82=0, P82_THRESHOLD_SINGLE=0.3
- Cap divergence (justified): PROFILE_RUN_CAP_M=4128 + PREALLOC_TOKEN_BUDGET=4128
(Sander uses 4096 — vLLM `interface.py:639` forces our config's Mamba
block_size to 4128 due to TQ3 + TP=1 page-size math; lower values
AssertionError at boot)
- Carry-forward (intentional): P4 (hybrid TQ required), P65 (TQ spec-CG
downgrade — pending v0.20 verification that #40880 closure makes it
redundant)
Cold-start cache mounts (closes #22):
- All 10 composes now mount torch_compile_cache + Triton cache from
`models/qwen3.6-27b/vllm/cache/`. First boot warms (~6 min); warm boot
drops to ~3.2 min (47% faster). Per-stage savings on long-text:
- Dynamo bytecode transform: 18s → 5s (-73%)
- torch.compile: 57s → 9s (-85%)
- Initial profiling/warmup: 51s → 7s (-87%)
Mamba block_size cap fix:
- v0.20 enforces `long_prefill_token_threshold >= block_size`; on hybrid
Mamba+TQ3, vLLM forces block_size=4128. Bumped GENESIS_PROFILE_RUN_CAP_M
and PREALLOC_TOKEN_BUDGET 4096→4128 across all 5 main composes.
Default 48K compose:
- Required workspace_lock_disable sidecar after initial v0.20 boot hit
vllm#39226 strict assertion. Caught during validation, fixed.
Context restored vs dev205 backoffs (validated 33K + 50K stress on v0.20):
- long-text: 185K → 214K (+16%)
- long-vision: 140K → 198K (+41%)
- bounded-thinking: 185K → 214K (+16%)
Bench results (n=5, results/v0.20-migration/):
- long-text 214K narr 49.74 / code 67.39 (CV 2.6/2.7%)
- long-vision 198K narr 50.32 / code 66.12 (CV 2.3/4.1%)
- bounded-thinking 214K narr 49.77 / code 65.80 (CV 1.4/2.3%)
- tools-text 75K (fp8) narr 53.32 / code 69.66 (CV 2.3/1.4%)
- dual-turbo 262K (TP=2) narr 58.33 / code 76.01 per-stream
269 TPS aggregate at n=4 streams (3.63x speedup)
- default 48K narr 48.82 / code 65.98 (n=3)
Validation: verify-full 8/8 on every variant. verify-stress 33K AND 50K
tool-prefill PASS on every variant — the cliff that fired on EVERY dev205
config no longer reproduces.
Docs + charts:
- README + SINGLE_CARD + DUAL_CARD + CLIFFS + EXAMPLES + STRUCTURED_COT
+ FAQ + UPSTREAM + 3 engine docs + model README + INTERNALS + CHANGELOG
all updated with new pin, ctx, TPS numbers, and "v0.20 unblock" section
- performance.{png,svg} + variants regenerated with measured TPS
- vram-budget.{png,svg} + variants regenerated with measured VRAM
- UPSTREAM tracker: 5 issues moved ✅ closed (PR #12, #13, #14, #15, P104
superseded by PN17 + P15B)
Issues addressed:
- #16 (Cliff 1 mech B leaks past PN12 on inductor-compiled FFN) — partial:
v0.20's revised TQ FA paths close the synthetic stress; PN25 (Sander's
proper compile-path opaque-op fix) is on dev but explicitly opt-in pending
worker-fork registration fix. Workarounds documented (tools-text fp8 path
/ --enforce-eager) until Sander ships PN25 default-on.
- #20 (launch.sh port + container-name mismatch) — already closed by
77ca576 (post-issue-filing).
- #22 (cold-start caching) — closed by cache mounts above.
Remaining caveats:
- Cliff 2 (DeltaNet GDN forward, single prompt ≥50-60K) unchanged —
architectural, applies to all single-card vLLM TQ3 paths. Mitigation:
dual-turbo TP=2 (state splits across cards) or llama.cpp 262K.
- Default 48K narr_TPS (48.82) slightly under chart's 55 reference —
bench variance + sample size n=3; not regression.
- Dual.yml / dual-dflash* not re-benched on v0.20; numbers carry forward
from dev205 (fp8 paths were not TPS-changed by the migration).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
16 KiB
FAQ
Common questions about club-3090. If your question isn't here, open a GitHub Discussion — most things end up in this doc eventually.
Hardware
Can I use a 4090 instead of a 3090?
Yes — 4090 (Ada, sm_89) is strictly better than 3090 (Ampere, sm_86) for everything we ship. Slightly different kernel paths but no patches needed. Caveats: vLLM Genesis patches are tested on Ampere; tools should still work but TPS scaling is untested. Open an issue with numbers if you bench it.
Can I use a 5090?
Should work for vLLM (Blackwell adds new kernels but back-compat). The Marlin pad-sub-tile-n fork we mount targets Ampere edge cases — on Blackwell you can probably drop the /opt/ai/vllm-src/ mount. Not validated yet. We'd love numbers from a 5090 rig — use the Numbers from your rig issue template.
Do I need NVLink?
No. Our dual-card configs use PCIe-only, no NVLink. Custom all-reduce is disabled in the composes. NVLink would help dual-card TPS but it's not required, and the user has explicitly declined NVLink bridges as a default — adding the dependency would exclude most consumer rigs.
Does this work on AMD / Intel / Apple Silicon?
vLLM: NVIDIA-only (CUDA). llama.cpp: yes — pick the right Docker image (ghcr.io/ggml-org/llama.cpp:server-rocm for AMD, :server for CPU-only, or build from source for Apple Silicon). Update the image: line in the compose. The flags (--ngl, -fa on, --cache-type-k q4_0) work identically across backends.
Does this work on Windows / WSL2?
WSL2: yes, both engines. Make sure GPU passthrough is set up (nvidia-smi works inside WSL). Native Windows: vLLM doesn't support it; llama.cpp does — but use a native llama.cpp build, not Docker.
Engine choice
Why ship both vLLM and llama.cpp?
Different trades. vLLM is faster (51-89 TPS depending on config) and has full feature support (vision · tools · MTP spec-decode · streaming · reasoning). As of 2026-04-30 PM, Cliff 1 (the 25K-token tool-prefill OOM) is closed on every shipped vLLM single-card variant via Genesis PN8 on the FP8 path and PN12 anchor sidecar on the TQ3 paths. The remaining caveat is Cliff 2 — single prompts above 50-60K still OOM in DeltaNet GDN forward, and that's a different memory class with no upstream fix yet. llama.cpp is slower (~21 TPS) but passes every stress test cleanly at full 262K context with vision and tools — including the big single prompts that vLLM single-card can't yet handle. See the launch frame: vLLM dual = max throughput, llama.cpp single = max robustness.
Why not Ollama?
Ollama wraps llama.cpp with a different model registry and slightly easier UX. It's fine for chat. Two reasons we don't ship it:
- Ollama doesn't expose all llama.cpp flags we need (
--cache-type-k q4_0,--mmproj,--spec-type ngram-mod, custom--parallel). - Ollama's model registry doesn't have the exact Unsloth GGUF quants we ship (UD-Q3_K_XL). You can run Ollama against an Unsloth GGUF manually, but at that point you've reimplemented our llama.cpp compose with a different wrapper.
Why not LM Studio?
LM Studio is GUI-driven and great for hobbyist use. We ship CLI/Docker because:
- Reproducibility — pinned image SHAs + Genesis commit make exact bench runs across machines possible
- Headless deployment — homelab racks, dev backends
- Tool-call extraction across both engines on this exact model is non-trivial; LM Studio's defaults haven't been validated
Use LM Studio if you prefer a GUI and don't need the engineering. Use this repo if you want a tested config that another club-3090 user can match exactly.
Why MTP and not EAGLE?
We tried EAGLE — it's blocked on Qwen3-Next (the family Qwen3.5/3.6 belong to) by DeltaNet hybrid attention's lack of KV rollback support in vLLM/SGLang. MTP works because it's a different protocol (multi-token prediction at draft-head level, not a separate draft model). See INTERNALS.md "Speculative decoding" for the full forensic chain. Re-test triggers: if vllm#39931 lands or DeltaNet rollback support arrives upstream, EAGLE becomes viable again.
Why not GGUF on vLLM for this model?
Multiple gates blocked. Qwen3.6-27B GGUF on vLLM hits a chain of "fixed but-not-quite" issues — multimodal config routing, ParallelLMHead skip, the Qwen35TensorProcessor._reverse_reorder_v_heads weight loader producing garbage output on the 27B layout (transformers PR #45283 only validated on 0.8B). Tracked in INTERNALS.md. Use llama.cpp for GGUF on this model.
Why AutoRound INT4 not GPTQ / AWQ?
AutoRound (Lorbus) gave us +9% TPS over AWQ on this model. GPTQ has a similar quality bar but the AWQ + DFlash path failed (pad-Marlin × aux-layer interaction). AutoRound + Genesis + MTP is the production-validated path. AWQ is documented as a fallback for users who can't use AutoRound.
Performance
Why is single-card TPS lower than I expected?
Look at the TPS chart — single-card vLLM is 51-55 TPS narrative / 67-70 code at 48K, which beats most consumer-3090 numbers we've seen reported. If you're seeing materially lower, the most common causes are:
- Power cap < 230 W (this rig benches at 230 W; 280 W gives ~+5%, 350 W ~+10%)
- Wrong compose for your prompt shape (use the
docker-compose.yml48K default for chat — don't picklong-vision.ymlif you don't need 198K) - Genesis tree drift —
git pull origin mainbetween bench runs can change AL by ±15%. We pin to commitbf667c7for this reason.
My TPS dropped after switching to 198K context. Why?
It shouldn't, much — we measured 50.93 TPS narr at 192K vs 50.53 at 32K (within variance) on long-vision.yml pre-fix; the new 198K + 0.98 config is in the same range. If it dropped a lot, you're probably actually decoding into a long ctx (not just having KV pool reserved). Loaded-context decode is 2-4× cold short-prompt decode on any LLM. The TPS chart number is short-prompt cold; loaded numbers are in BENCHMARKS.md.
What's a "prefill cliff"?
VRAM-related OOM during prompt processing on single-card vLLM. Two cliffs documented:
- Cliff 1 — historical: FFN intermediate buffer (
SiluAndMuloutput, 138 MiB atmax_num_batched_tokens=4128 × intermediate_size=17408 × 2 bytes) fresh-allocated per layer. Plus a related FA2 softmax_lse cap-leak (Dao-AILab/flash-attention#1011). Closed on every shipped vLLM single-card variant as of 2026-04-30 PM:tools-text.ymlvia Genesis PN8 (frees ~900 MiB on FP8 path);long-vision.ymlandlong-text.ymlvia the PN12 anchor sidecar (PR #13 to Sandermage's repo) plus a local P104 FA softmax_lse clamp. Full diagnostic: docs/CLIFFS.md. - Cliff 2 — DeltaNet GDN forward OOM at ~50-60K single-prompt regardless of mem-util. Different memory class, lives in
fla.opsupstream, no file-replacement patch yet. Tracked in UPSTREAM.md. Mitigation: dual-card TP=2 (verified at 237K) or llama.cpp single-card (262K, different engine).
For the full deep dive — empirical bisection, root-cause walk-through, who-can-fix-it landscape, and what we could do at any difficulty level — see docs/CLIFFS.md.
vllm#40914 keeps coming up — what is it?
Sandermage's K+1 verify routing PR for vLLM. When it lands, the spec-verify cost we're paying on Ampere SM 8.6 (~22 TPS narrative regression vs pre-bug substrate) closes. Our default on 0.20.1rc1.dev16+g7a1eb8ac2 + Genesis v7.65 dev tip will jump from ~50 narr to ~70 narr, matching what ampersandru measures on the older dev21 + v7.13 cascade-prone substrate. We track it in INTERNALS.md "Upstream tracker".
What's PN8?
A Genesis patch (GENESIS_ENABLE_PN8_MTP_DRAFT_ONLINE_QUANT=1) added in v7.62.x — backport of vllm#40849 that makes the MTP draft head inherit the target model's online-quant config. We measured ~800-900 MiB freed on the FP8+MTP single-card path (tools-text.yml), which closes Cliff 1 there. No-op on TQ3 paths. Enabled by default in tools-text.yml since 2026-04-29; opt-in elsewhere via the env var if you want to test.
Setup
bash scripts/setup.sh qwen3.6-27b is downloading 20+ GB. Where does it go?
<repo>/models-cache/ by default. Override with MODEL_DIR=/path/to/your/scratch bash scripts/setup.sh qwen3.6-27b. See .env.example for all env vars.
My GPU isn't card 0 — how do I change it?
CUDA_VISIBLE_DEVICES=2 bash scripts/launch.sh --variant vllm/default (substitute your card index). For dual-card, pass two: CUDA_VISIBLE_DEVICES=2,3. The compose files inherit env from your shell.
Container fails to start: "Free memory ... is less than desired GPU memory utilization"
Looks like:
ValueError: Free memory on device cuda:0 (22.76/24.0 GiB) on startup
is less than desired GPU memory utilization (0.97, 23.28 GiB).
vLLM's startup check reserves mem-util × total VRAM of currently-free VRAM before booting. If something else on the GPU is holding memory (X11 / Wayland compositor, leftover container, Python process, browser GPU acceleration), the check fails. Most common on tools-text.yml (0.97) and the long-* variants (0.98 / 0.985).
Two fixes:
- Free the VRAM (preferred).
nvidia-smishows what's holding it. Common: log out of GUI, stop a leftover container (docker rm -f $(docker ps -aq --filter "name=vllm-")), or kill orphanedpythonprocesses. - Lower mem-util in the compose. e.g. on
tools-text.yml: drop--gpu-memory-utilization 0.97to0.94and reduce--max-model-lenproportionally (75K → ~70K). Loses ~6K context but works on any rig.
The 0.97 / 0.98 / 0.985 defaults assume a headless rig with ≥23.3 GiB consistently free. If you're running a desktop session on the same card, 0.92–0.94 is the safer ceiling.
Can I run multiple variants at once on the same machine?
You'd need different ports per variant. Set PORT=9876 in .env (or pass inline: PORT=9876 bash scripts/switch.sh vllm/default) — every shipped compose now reads ${PORT} for the host-side port mapping. Watch VRAM — two configs simultaneously typically don't fit on 24 GB.
Will this work behind Open WebUI?
Yes. Add a connection in Open WebUI's Settings → Connections → OpenAI: base URL http://localhost:8020/v1, any non-empty API key, model qwen3.6-27b-autoround. See docs/EXAMPLES.md.
Will this work with VS Code GitHub Copilot LLM Gateway?
Yes, but you need a compose with ≥48K context — Copilot's LLM Gateway sends ~20K tokens of tool-schema preamble (50+ VS Code tools enumerated in a structured-outputs JSON schema) on every request, which alone consumes most of a small context budget. Use tools-text.yml (75K + fp8 + PN8 enabled — Cliff 1 closed):
bash scripts/switch.sh vllm/tools-text
There's a second wrinkle: Copilot's LLM Gateway sometimes sends very low max_tokens (e.g. 64) on probe-style requests. With tool_choice: required (which Copilot enforces via minItems: 1 on its structured-outputs schema), the model must emit a tool-call JSON that wraps a real argument like a file path — and 64 tokens isn't enough to fit {"name": "read_file", "parameters": {"filePath": "/long/abs/path"}}. The truncated JSON arrives at the gateway as "empty response." If you see this pattern, it's a client-side limit, not the server. Other OpenAI-compat clients (Cline / Continue.dev / Cursor) tend to send realistic max_tokens by default and don't hit this.
Server-side fix landed 2026-04-29: the Genesis P68/P69 long-context tool-adherence patches were silently overriding tool_choice: auto → required and injecting "must use a tool" reminders whenever prompt > 8000 chars. That made greetings + clarifying questions stall on every IDE-agent setup (Cline, Cursor, OpenCode, and Copilot Gateway combined). We disabled both in tools-text.yml. Behavior now: greeting → plain-text reply ("Hello! How can I help you today?"); tool request → clean read_file({"path": "..."}) call. P64 and PN8 stay enabled (real targeted bugfixes, no user-intent override).
Background + bisection: club-3090 #2.
Community / contribution
Can I add my benchmark numbers from a different rig?
Please do — open an issue using the Numbers from your rig template. We collect cross-rig data points in BENCHMARKS for community signal.
Found a bug — what should I include?
The bug report template asks for the data we always need: docker logs --tail 100, verify-full.sh output, nvidia-smi, your compose variant, and the repo commit. Skipping these means the first reply will just ask for them, costing you a round-trip.
How do I bump Genesis to a newer commit?
GENESIS_PIN=<new-commit-sha> bash scripts/setup.sh qwen3.6-27b and re-run bash scripts/verify-full.sh to confirm tools still work. Don't bump in production without re-running the verify suite — Genesis releases sometimes change spec-verify routing in ways that affect tool-call extraction.
Troubleshooting
Quick recognition guide for common failure modes:
- Container dies at boot with
GPTQ_MARLIN_MIN_THREAD_N (64) > out_features— dual-card vllm#40361 patch didn't apply. Confirm/opt/ai/vllm-src/exists with the patched marlin kernel files. - Container dies during DFlash boot — vllm#40334 dtype mismatch. Verify the compose has
--dtype bfloat16. - Tool calls return
<tool_call>as plain text — Genesis didn't apply. CheckGenesis Results: 27 appliedin logs (boot-time). - OOM during prefill at 60K+ tokens — single-card Cliff 2 (DeltaNet GDN forward). Switch to lower max-model-len, dual-card, or llama.cpp + q4_0 KV.
- OOM during prefill at 25K+ tool response — historically Cliff 1 on TQ3 paths. Closed since 2026-04-30 PM via PN12 anchor sidecar on
long-vision.yml/long-text.yml. If you're hitting it, check your compose has the sidecar wired in (patch_pn12_ffn_pool_anchor.pyin entrypoint). - "Empty response" through VS Code Copilot LLM Gateway — Copilot sends ~20K tokens of tool schemas + sometimes uses
max_tokens=64which truncates tool-call JSON. Switch totools-text.yml(75K) and check Copilot's max_tokens setting. See #2 for full debug-log analysis. - Per-stream TPS lower than expected — re-run
bench.shwith 3+ warmups + 5 measured runs first. Run-to-run variance is ~5%.
If none match, open an issue with docker logs <container> 2>&1 | tail -200 + nvidia-smi — see bug-report.yml template.
See also
- README — top-level overview + quick start
- docs/SINGLE_CARD.md — 1× 3090 deployment menu
- docs/DUAL_CARD.md — 2× 3090 deployment menu
- models/qwen3.6-27b/README.md — model-specific reference
- models/qwen3.6-27b/INTERNALS.md — engineering deep dive
- docs/EXAMPLES.md — Python / TS / curl client snippets
- docs/HARDWARE.md — Ampere notes, NVLink, power caps
- docs/GLOSSARY.md — TPS / KV / MTP / TP / etc. plain-language definitions