After 383b5cc shipped 175K + 0.97 (text) and 140K + 0.95 (vision),
audit showed the cliffs the backoff was protecting against fire on
every config we ship — they're independent of max-model-len. So the
context capacity was wasted protection.
Push text-only ceilings up:
long-text: 175K + 0.97 → 185K + 0.975
bounded-thinking: 175K + 0.97 → 185K + 0.975
Vision stays at 140K + 0.95: tried 185K + 0.98, 185K + 0.975, 160K
+ 0.97; all reopened Cliff 2 (DeltaNet GDN forward buffer) at the
130K-char stress class. Vision tower's ~1 GiB persistent + the new
patches' persistent allocations (P38 K_full/V_full ~750 MiB at 185K
+ compile-safe sidecar ~138 MiB) leave too little headroom for the
GDN intermediate buffer at 30K+ token prefills on this variant.
P37 disabled on vision (was on for parity with long-text but P37's
MoE intermediate cache pool is no-op on dense Qwen3.6-27B and the
env gate doesn't free memory anyway).
Verification at the new ceilings:
long-text 185K + 0.975: verify-full 8/8 (MTP AL 2.66),
130K-char tool-prefill stress PASS
long-vision 140K + 0.95: verify-full 8/8 (MTP AL 3.27),
130K-char tool-prefill stress PASS
bounded-thinking 185K + 0.975: not re-booted in this final state
(config identical to long-text +
one --structured-outputs flag,
no memory delta expected)
Docs updated: SINGLE_CARD.md picker table + activation-budget +
per-variant blurbs; engines/VLLM.md TL;DR + KV cache table; engines/
LLAMA_CPP.md "when to use vLLM"; STRUCTURED_COT.md "When to pick
this over long-text"; models/qwen3.6-27b/README.md per-variant lines;
docs/CLIFFS.md "Update 2026-05-01 PM" with full bisection sweep and
final decision.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
13 KiB
Single 3090 — what fits, how to run it
You have one RTX 3090 (24 GB VRAM). This page is the front door for picking a config and knowing what to expect. The model-specific deep dives (quants, Genesis patches, engine internals) live elsewhere — links at the bottom.
TL;DR — pick by workload
Four recommended options:
| What you're doing | Compose | Max ctx | Narr / Code TPS | VRAM (24 GB / card) |
|---|---|---|---|---|
| Long ctx + vision (chat, agents, image input) | long-vision.yml |
140K | ~50 / ~67 | ~22.4 GB (mem-util 0.95) |
| Long ctx, text-only (RAG, codebase, books) | long-text.yml |
185K | ~50 / ~66 | ~22.9 GB (mem-util 0.975) |
| Bounded thinking (coding agents, structured-CoT, cost-bounded thinking) — see STRUCTURED_COT.md | bounded-thinking.yml |
185K | ~52 / ~56 | ~22.9 GB (mem-util 0.975) |
| Bulletproof, no cliffs (production service, unpredictable inputs) | llamacpp/default |
262K | 21 / 21 | ~20 GB |
Run via bash scripts/launch.sh (interactive) or bash scripts/switch.sh <variant>.
⚠️ The one limitation to know
vLLM single-card variants will crash if you send a single prompt above ~50K tokens.
This is Cliff 2 — DeltaNet GDN forward OOMs at 50–60K single-shot regardless of how much VRAM you have left. Both
long-vision.yml(140K) andlong-text.yml(185K) are designed for steady-state accumulation across many turns — context that builds up across tool calls, replies, retrieved chunks. They are NOT designed for "paste an 80K-token document and ask one question."If your workload ever sends single big prompts: use
llamacpp/default(262K, no cliffs anywhere — different engine entirely) or move to dual-card (dual.ymlTP=2, verified at 237K).Cliff 1 mech B (FFN intermediate-buffer activation peak) was largely closed by PN12 + PN17 + P38 + our compile-safe sidecar. The downstream FA varlen workspace cliff at 50K-token tool prefills then surfaced — backing off to 130K + 0.95 (long-text) / 120K + 0.94 (long-vision) gives the necessary activation headroom. Re-push criteria in
docs/CLIFFS.md.
Measured TPS on single 3090
Bench protocol: 3 warm + 5 measured runs of the canonical narrative + code prompts on each config. Substrate: vLLM nightly dev205+g07351e088 + Genesis pinned to 917519b (v7.62.x), llama.cpp mainline 0d0764dfd, RTX 3090 sm_86 PCIe-only at 230 W. Per-config run-by-run + VRAM peaks: models/qwen3.6-27b/CHANGELOG.md.
VRAM budget on 24 GB
What this says about single-card constraints:
- Model weights consume ~14 GB (AutoRound INT4 / GGUF Q3_K_XL). Half the card.
- KV cache is the next biggest line; its size depends on
--kv-cache-dtype× ctx. fp8 ≈ 1 byte/token/(layer×head); TQ3 ≈ 0.4 bytes/token/(layer×head); fp16 ≈ 2 bytes/token/(layer×head). - Vision tower (mmproj) costs ~0.5–1.0 GB extra when on.
- Activations + cudagraph pools is what's left. At
--gpu-memory-utilization 0.92(default 48K) you have 2-3 GB of activation headroom — comfortable. 2026-05-01 PM — settled on 185K + 0.975 (long-text + bounded-thinking) and 140K + 0.95 (long-vision) after P37/P38 testing surfaced a downstream FA varlen workspace cliff (long-text) and a Cliff 2 GDN buffer regression on vision (the new persistent patches eat into the activation budget vision needs for 30K+ token prefills). Both ceilings pass realistic 130K-char (33K-token) tool-prefill stress. Synthetic 200K-char (50K-token) single-shot stress still cliffs — that's beyond what most agent workloads emit. Re-push criteria indocs/CLIFFS.md.
For the cross-card TP=2 picture, see DUAL_CARD.md.
Pick a config
Long ctx + vision — long-vision.yml ⭐
Workload: chat with images, vision-aware coding agents, multimodal RAG. Anything where the user might paste a screenshot.
140K + vision tower + TQ3 KV + Genesis MTP n=3 + PN17 + P104 + P37/P38 + compile-safe sidecar at mem-util 0.95. verify-full.sh 8/8 (MTP AL 2.49); verify-stress.sh 130K-char tool-prefill OK. 200K-char tool-prefill (50K tokens) still cliffs — vision tower's persistent overhead tightens the margin.
Long ctx, text-only — long-text.yml ⭐
Workload: RAG ingest, codebase analysis, book/document Q&A, long conversations without image input.
185K + no vision + TQ3 KV + same patch stack at mem-util 0.975. verify-full.sh 8/8 (MTP AL 2.66); verify-stress.sh 130K-char tool-prefill OK. Vision drop adds ~1 GB headroom over long-vision so this variant runs at higher ctx + mem-util safely.
Bulletproof / no cliffs — llamacpp/default ⭐
Workload: production service for unpredictable users. Inputs that might be 5K or might be 200K. Tool returns that might be 1K or might be 50K. Anywhere "predictable behavior" beats "peak TPS."
bash scripts/switch.sh llamacpp/default. Q3_K_XL (Unsloth dynamic) + q4_0 KV at 262K + vision (mmproj). Different attention library entirely (ggml-cuda, not FA2) → no Cliff 1 mechanism, no Cliff 2 mechanism. Trade is ~21 TPS (~2.5× slower than vLLM). Quant validated by Benjamin Marie's Kaitchup eval.
Other variants in the repo (not recommended for shipping)
These exist for troubleshooting, niche workloads, or historical comparison. Not promoted as primary because the long-* variants now cover their use cases:
docker-compose.yml— 48K + TQ3 + vision, mem-util 0.92. The "below both cliffs by definition" baseline (engine HTTP-400-rejects requests >48K, so Cliff 2 is unreachable). Useful when you want bulletproof error behavior on a specific small-ctx workload, or as a fast-boot diagnostic. Most users should picklong-visionorllamacpp/defaultinstead.tools-text.yml— 75K + FP8 KV + PN8. Was the only Cliff-1-safe single-card path before PN12 anchor fix landed. FP8 KV is closer in quality to FP16 than TQ3 is, so kept around for accuracy-sensitive comparisons. Most IDE-agent workloads now run fine onlong-text.yml.minimal.yml— 32K + FP8 + no Genesis + no spec-decode. Stripped-down stack for isolating "is this a Genesis bug?" questions. Half the throughput of any other variant.
Watch list — Luce DFlash (not yet a recommendation)
Re-tested 2026-04-30 PM against Luce-Org/lucebox-hub on Qwen3.6-27B Q4_K_M target + matched z-lab/Qwen3.6-27B-DFlash draft. Closer to parity than 2026-04-22 — but several gaps still keep it off the recommended list:
Measured TPS on this rig (RTX 3090, greedy, single-stream, n_gen=1000):
| Workload | Luce DFlash 3.6+3.6 (TQ3 KV, max_ctx=65K) | vLLM long-text 185K |
|---|---|---|
| Narrative essay | 37–47 TPS (mean ~40) | 50 TPS |
| Code (heap/LRU/AST) | 63–76 TPS (mean ~72) | 66 TPS |
| AL (code) | 5.9–7.1 | 3.4–3.8 (MTP) |
What works since 2026-04-22:
- ✅ Tool calls via
server_tools.py— parses Qwen<tool_call>format → returns OpenAItool_calls[]. The big server-UX gap from last bench is closed. - ✅ Streaming SSE with
reasoning_contentdeltas. - ✅ Daemon mode with cache-reuse for fast cold starts.
- ✅ Verify-stress 25K tool-prefill passes at TQ3 KV + max_ctx=65K.
What still keeps it off the recommended list:
- ❌ Greedy only —
temperature/top_pignored. Real downside for creative-writing workloads. - ❌ 3.6 draft under-trained (z-lab snapshot 2026-04-26). Narrative AL ~3.7 vs code ~7.0; narr loses ~20% TPS to vLLM until training completes.
- ❌ No vision tower.
- ❌
enable_thinkingchat_template_kwargs handled differently than vLLM — verify-full check 6 fails. - ❌ Prefill cliff at higher max_ctx — 25K tool prefill OOMs in
fattn-chunked.cuat Q8_0 KV + max_ctx=65K (TQ3 KV closes it). At max_ctx=131K + TQ3, the daemon subprocess crashes (broken-pipe to FastAPI) on 30K+ probes. - ❌ Build fragility — fresh
git cloneofdflashmain HEAD fails to compile (ggml_turbo_wht/GGML_TYPE_TQ3_0undefined) until yougit submodule update --initafter manualgit fetchindflash/deps/llama.cpp. - ❌ Daemon-mode "empty prompt" regression — after streaming requests, subsequent requests sometimes return 0 tokens; needs server restart.
Re-test trigger: z-lab tags the Qwen3.6-27B-DFlash draft as training-complete OR Luce-Org publishes a tagged release with the daemon-mode bug fixed. Track in docs/UPSTREAM.md.
What single-card can't do
| Want | Why not on 1× | What you'd need |
|---|---|---|
| 4 concurrent streams at 262K + vision | KV pool too small for 4 × full ctx | TP=2 (see DUAL_CARD.md) |
| Peak code TPS (>100 TPS on quicksort prompt) | DFlash N=5 needs head_size=256 + non-causal — vLLM head-dim split | TP=2 + DFlash |
| Single-prompt >60K tokens on vLLM | Cliff 2 (DeltaNet GDN forward), no fix yet | TP=2 OR llama.cpp 262K (different engine) |
Common pitfalls (single-card specifics)
Prefill cliffs
- Cliff 1 — FFN intermediate-buffer activation peak (138 MiB allocate at
intermediate_size × max-num-batched-tokens). Historically fired on long-ctx composes at >0.95 mem-util when prefill batch needed the buffer. Closed ontools-text.yml(FP8 KV path) since 2026-04-29 via Genesis PN8. Closed on TQ3 paths (long-vision.yml198K,long-text.yml218K) since 2026-04-30 PM via PN12 anchor sidecar — seedocs/CLIFFS.md. - Cliff 2 — DeltaNet GDN forward OOM at 50-60K single-prompt regardless of mem-util. In
fla.opsupstream, no file-replacement patch available. Watch vllm#40914 and FlashQLA for upstream fixes.
VRAM peak vs idle
nvidia-smi at boot ≠ peak. Boot shows weights + KV pool reservation. Peak adds activation buffers during prefill — typically +500-1500 MiB. If nvidia-smi shows 23.5/24 GB at idle, you have ~500 MiB for prefill activations — not enough for the 138 MiB-class buffer at long ctx. Drop mem-util by 0.03 if you need the headroom.
Tool-call extraction needs --enable-auto-tool-choice
vLLM ships this off by default. Our composes set --tool-call-parser qwen3_coder + --enable-auto-tool-choice. If you're rolling your own compose, both are required.
Quick start
# 1. Setup (downloads model, clones Genesis, ~20 min cold)
bash scripts/setup.sh qwen3.6-27b
# 2. Pick + boot via wizard (asks engine + workload)
bash scripts/launch.sh
# 3. Or skip the wizard:
bash scripts/launch.sh --variant vllm/tools-text # IDE agent path
bash scripts/launch.sh --variant llamacpp/default # easy mode
# 4. Sanity test
curl -sf http://localhost:8020/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.6-27b-autoround","messages":[{"role":"user","content":"Capital of France?"}],"max_tokens":200}'
# 5. Switch later without re-running setup
bash scripts/switch.sh vllm/long-vision # for example
bash scripts/switch.sh --list # show all variants
Models supported on single 3090
- Qwen3.6-27B — primary model. Quant choices, Genesis patches, engine internals all in the model directory.
- More models coming. As they're added, this section will list which single-card configs each one supports.
Deep dives
- Model README — quant choices (AutoRound INT4 / GGUF Q3_K_XL), Genesis patch surface, what's working / what's not.
- INTERNALS.md — engineering rationale (Genesis P65/P66/PN8, Marlin pad fork, MTP, the cascade bug, upstream tracker).
- VRAM allocation diagram — full per-config breakdown across single + dual.
- FAQ.md — common questions (4090 / 5090 support, why MTP not EAGLE, Copilot Gateway, what's a cliff, etc.).
- EXAMPLES.md — Python / TS / curl client snippets + IDE connection settings.
- HARDWARE.md — Ampere SM 8.6 specifics, NVLink (declined), power caps.
- DUAL_CARD.md — when you need what single-card can't deliver.

