Files
club-3090/docs/SINGLE_CARD.md
T
noonghunnaandClaude Opus 4.7 df91d641c4 push long-text/bounded-thinking back to 185K + 0.975; long-vision stays 140K + 0.95
After 383b5cc shipped 175K + 0.97 (text) and 140K + 0.95 (vision),
audit showed the cliffs the backoff was protecting against fire on
every config we ship — they're independent of max-model-len. So the
context capacity was wasted protection.

Push text-only ceilings up:
  long-text:        175K + 0.97  → 185K + 0.975
  bounded-thinking: 175K + 0.97  → 185K + 0.975

Vision stays at 140K + 0.95: tried 185K + 0.98, 185K + 0.975, 160K
+ 0.97; all reopened Cliff 2 (DeltaNet GDN forward buffer) at the
130K-char stress class. Vision tower's ~1 GiB persistent + the new
patches' persistent allocations (P38 K_full/V_full ~750 MiB at 185K
+ compile-safe sidecar ~138 MiB) leave too little headroom for the
GDN intermediate buffer at 30K+ token prefills on this variant.
P37 disabled on vision (was on for parity with long-text but P37's
MoE intermediate cache pool is no-op on dense Qwen3.6-27B and the
env gate doesn't free memory anyway).

Verification at the new ceilings:
  long-text 185K + 0.975:    verify-full 8/8 (MTP AL 2.66),
                              130K-char tool-prefill stress PASS
  long-vision 140K + 0.95:   verify-full 8/8 (MTP AL 3.27),
                              130K-char tool-prefill stress PASS
  bounded-thinking 185K + 0.975: not re-booted in this final state
                                  (config identical to long-text +
                                  one --structured-outputs flag,
                                  no memory delta expected)

Docs updated: SINGLE_CARD.md picker table + activation-budget +
per-variant blurbs; engines/VLLM.md TL;DR + KV cache table; engines/
LLAMA_CPP.md "when to use vLLM"; STRUCTURED_COT.md "When to pick
this over long-text"; models/qwen3.6-27b/README.md per-variant lines;
docs/CLIFFS.md "Update 2026-05-01 PM" with full bisection sweep and
final decision.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-01 11:15:01 +00:00

13 KiB
Raw Blame History

Single 3090 — what fits, how to run it

You have one RTX 3090 (24 GB VRAM). This page is the front door for picking a config and knowing what to expect. The model-specific deep dives (quants, Genesis patches, engine internals) live elsewhere — links at the bottom.


TL;DR — pick by workload

Four recommended options:

What you're doing Compose Max ctx Narr / Code TPS VRAM (24 GB / card)
Long ctx + vision (chat, agents, image input) long-vision.yml 140K ~50 / ~67 ~22.4 GB (mem-util 0.95)
Long ctx, text-only (RAG, codebase, books) long-text.yml 185K ~50 / ~66 ~22.9 GB (mem-util 0.975)
Bounded thinking (coding agents, structured-CoT, cost-bounded thinking) — see STRUCTURED_COT.md bounded-thinking.yml 185K ~52 / ~56 ~22.9 GB (mem-util 0.975)
Bulletproof, no cliffs (production service, unpredictable inputs) llamacpp/default 262K 21 / 21 ~20 GB

Run via bash scripts/launch.sh (interactive) or bash scripts/switch.sh <variant>.

⚠️ The one limitation to know

vLLM single-card variants will crash if you send a single prompt above ~50K tokens.

This is Cliff 2 — DeltaNet GDN forward OOMs at 50–60K single-shot regardless of how much VRAM you have left. Both long-vision.yml (140K) and long-text.yml (185K) are designed for steady-state accumulation across many turns — context that builds up across tool calls, replies, retrieved chunks. They are NOT designed for "paste an 80K-token document and ask one question."

If your workload ever sends single big prompts: use llamacpp/default (262K, no cliffs anywhere — different engine entirely) or move to dual-card (dual.yml TP=2, verified at 237K).

Cliff 1 mech B (FFN intermediate-buffer activation peak) was largely closed by PN12 + PN17 + P38 + our compile-safe sidecar. The downstream FA varlen workspace cliff at 50K-token tool prefills then surfaced — backing off to 130K + 0.95 (long-text) / 120K + 0.94 (long-vision) gives the necessary activation headroom. Re-push criteria in docs/CLIFFS.md.


Measured TPS on single 3090

Qwen3.6-27B TPS — single 3090 configs

Bench protocol: 3 warm + 5 measured runs of the canonical narrative + code prompts on each config. Substrate: vLLM nightly dev205+g07351e088 + Genesis pinned to 917519b (v7.62.x), llama.cpp mainline 0d0764dfd, RTX 3090 sm_86 PCIe-only at 230 W. Per-config run-by-run + VRAM peaks: models/qwen3.6-27b/CHANGELOG.md.


VRAM budget on 24 GB

Per-card VRAM allocation, single-card configs

What this says about single-card constraints:

  • Model weights consume ~14 GB (AutoRound INT4 / GGUF Q3_K_XL). Half the card.
  • KV cache is the next biggest line; its size depends on --kv-cache-dtype × ctx. fp8 ≈ 1 byte/token/(layer×head); TQ3 ≈ 0.4 bytes/token/(layer×head); fp16 ≈ 2 bytes/token/(layer×head).
  • Vision tower (mmproj) costs ~0.5–1.0 GB extra when on.
  • Activations + cudagraph pools is what's left. At --gpu-memory-utilization 0.92 (default 48K) you have 2-3 GB of activation headroom — comfortable. 2026-05-01 PM — settled on 185K + 0.975 (long-text + bounded-thinking) and 140K + 0.95 (long-vision) after P37/P38 testing surfaced a downstream FA varlen workspace cliff (long-text) and a Cliff 2 GDN buffer regression on vision (the new persistent patches eat into the activation budget vision needs for 30K+ token prefills). Both ceilings pass realistic 130K-char (33K-token) tool-prefill stress. Synthetic 200K-char (50K-token) single-shot stress still cliffs — that's beyond what most agent workloads emit. Re-push criteria in docs/CLIFFS.md.

For the cross-card TP=2 picture, see DUAL_CARD.md.


Pick a config

Long ctx + vision — long-vision.yml ⭐

Workload: chat with images, vision-aware coding agents, multimodal RAG. Anything where the user might paste a screenshot.

140K + vision tower + TQ3 KV + Genesis MTP n=3 + PN17 + P104 + P37/P38 + compile-safe sidecar at mem-util 0.95. verify-full.sh 8/8 (MTP AL 2.49); verify-stress.sh 130K-char tool-prefill OK. 200K-char tool-prefill (50K tokens) still cliffs — vision tower's persistent overhead tightens the margin.

Long ctx, text-only — long-text.yml ⭐

Workload: RAG ingest, codebase analysis, book/document Q&A, long conversations without image input.

185K + no vision + TQ3 KV + same patch stack at mem-util 0.975. verify-full.sh 8/8 (MTP AL 2.66); verify-stress.sh 130K-char tool-prefill OK. Vision drop adds ~1 GB headroom over long-vision so this variant runs at higher ctx + mem-util safely.

Bulletproof / no cliffs — llamacpp/default ⭐

Workload: production service for unpredictable users. Inputs that might be 5K or might be 200K. Tool returns that might be 1K or might be 50K. Anywhere "predictable behavior" beats "peak TPS."

bash scripts/switch.sh llamacpp/default. Q3_K_XL (Unsloth dynamic) + q4_0 KV at 262K + vision (mmproj). Different attention library entirely (ggml-cuda, not FA2) → no Cliff 1 mechanism, no Cliff 2 mechanism. Trade is ~21 TPS (~2.5× slower than vLLM). Quant validated by Benjamin Marie's Kaitchup eval.


These exist for troubleshooting, niche workloads, or historical comparison. Not promoted as primary because the long-* variants now cover their use cases:

  • docker-compose.yml — 48K + TQ3 + vision, mem-util 0.92. The "below both cliffs by definition" baseline (engine HTTP-400-rejects requests >48K, so Cliff 2 is unreachable). Useful when you want bulletproof error behavior on a specific small-ctx workload, or as a fast-boot diagnostic. Most users should pick long-vision or llamacpp/default instead.
  • tools-text.yml — 75K + FP8 KV + PN8. Was the only Cliff-1-safe single-card path before PN12 anchor fix landed. FP8 KV is closer in quality to FP16 than TQ3 is, so kept around for accuracy-sensitive comparisons. Most IDE-agent workloads now run fine on long-text.yml.
  • minimal.yml — 32K + FP8 + no Genesis + no spec-decode. Stripped-down stack for isolating "is this a Genesis bug?" questions. Half the throughput of any other variant.

Watch list — Luce DFlash (not yet a recommendation)

Re-tested 2026-04-30 PM against Luce-Org/lucebox-hub on Qwen3.6-27B Q4_K_M target + matched z-lab/Qwen3.6-27B-DFlash draft. Closer to parity than 2026-04-22 — but several gaps still keep it off the recommended list:

Measured TPS on this rig (RTX 3090, greedy, single-stream, n_gen=1000):

Workload Luce DFlash 3.6+3.6 (TQ3 KV, max_ctx=65K) vLLM long-text 185K
Narrative essay 37–47 TPS (mean ~40) 50 TPS
Code (heap/LRU/AST) 63–76 TPS (mean ~72) 66 TPS
AL (code) 5.9–7.1 3.4–3.8 (MTP)

What works since 2026-04-22:

  • ✅ Tool calls via server_tools.py — parses Qwen <tool_call> format → returns OpenAI tool_calls[]. The big server-UX gap from last bench is closed.
  • ✅ Streaming SSE with reasoning_content deltas.
  • ✅ Daemon mode with cache-reuse for fast cold starts.
  • ✅ Verify-stress 25K tool-prefill passes at TQ3 KV + max_ctx=65K.

What still keeps it off the recommended list:

  • ❌ Greedy only — temperature / top_p ignored. Real downside for creative-writing workloads.
  • ❌ 3.6 draft under-trained (z-lab snapshot 2026-04-26). Narrative AL ~3.7 vs code ~7.0; narr loses ~20% TPS to vLLM until training completes.
  • ❌ No vision tower.
  • ❌ enable_thinking chat_template_kwargs handled differently than vLLM — verify-full check 6 fails.
  • ❌ Prefill cliff at higher max_ctx — 25K tool prefill OOMs in fattn-chunked.cu at Q8_0 KV + max_ctx=65K (TQ3 KV closes it). At max_ctx=131K + TQ3, the daemon subprocess crashes (broken-pipe to FastAPI) on 30K+ probes.
  • ❌ Build fragility — fresh git clone of dflash main HEAD fails to compile (ggml_turbo_wht / GGML_TYPE_TQ3_0 undefined) until you git submodule update --init after manual git fetch in dflash/deps/llama.cpp.
  • ❌ Daemon-mode "empty prompt" regression — after streaming requests, subsequent requests sometimes return 0 tokens; needs server restart.

Re-test trigger: z-lab tags the Qwen3.6-27B-DFlash draft as training-complete OR Luce-Org publishes a tagged release with the daemon-mode bug fixed. Track in docs/UPSTREAM.md.


What single-card can't do

Want Why not on 1× What you'd need
4 concurrent streams at 262K + vision KV pool too small for 4 × full ctx TP=2 (see DUAL_CARD.md)
Peak code TPS (>100 TPS on quicksort prompt) DFlash N=5 needs head_size=256 + non-causal — vLLM head-dim split TP=2 + DFlash
Single-prompt >60K tokens on vLLM Cliff 2 (DeltaNet GDN forward), no fix yet TP=2 OR llama.cpp 262K (different engine)

Common pitfalls (single-card specifics)

Prefill cliffs

  • Cliff 1 — FFN intermediate-buffer activation peak (138 MiB allocate at intermediate_size × max-num-batched-tokens). Historically fired on long-ctx composes at >0.95 mem-util when prefill batch needed the buffer. Closed on tools-text.yml (FP8 KV path) since 2026-04-29 via Genesis PN8. Closed on TQ3 paths (long-vision.yml 198K, long-text.yml 218K) since 2026-04-30 PM via PN12 anchor sidecar — see docs/CLIFFS.md.
  • Cliff 2 — DeltaNet GDN forward OOM at 50-60K single-prompt regardless of mem-util. In fla.ops upstream, no file-replacement patch available. Watch vllm#40914 and FlashQLA for upstream fixes.

VRAM peak vs idle

nvidia-smi at boot ≠ peak. Boot shows weights + KV pool reservation. Peak adds activation buffers during prefill — typically +500-1500 MiB. If nvidia-smi shows 23.5/24 GB at idle, you have ~500 MiB for prefill activations — not enough for the 138 MiB-class buffer at long ctx. Drop mem-util by 0.03 if you need the headroom.

Tool-call extraction needs --enable-auto-tool-choice

vLLM ships this off by default. Our composes set --tool-call-parser qwen3_coder + --enable-auto-tool-choice. If you're rolling your own compose, both are required.


Quick start

# 1. Setup (downloads model, clones Genesis, ~20 min cold)
bash scripts/setup.sh qwen3.6-27b

# 2. Pick + boot via wizard (asks engine + workload)
bash scripts/launch.sh

# 3. Or skip the wizard:
bash scripts/launch.sh --variant vllm/tools-text   # IDE agent path
bash scripts/launch.sh --variant llamacpp/default  # easy mode

# 4. Sanity test
curl -sf http://localhost:8020/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-27b-autoround","messages":[{"role":"user","content":"Capital of France?"}],"max_tokens":200}'

# 5. Switch later without re-running setup
bash scripts/switch.sh vllm/long-vision    # for example
bash scripts/switch.sh --list              # show all variants

Models supported on single 3090

  • Qwen3.6-27B — primary model. Quant choices, Genesis patches, engine internals all in the model directory.
  • More models coming. As they're added, this section will list which single-card configs each one supports.

Deep dives

  • Model README — quant choices (AutoRound INT4 / GGUF Q3_K_XL), Genesis patch surface, what's working / what's not.
  • INTERNALS.md — engineering rationale (Genesis P65/P66/PN8, Marlin pad fork, MTP, the cascade bug, upstream tracker).
  • VRAM allocation diagram — full per-config breakdown across single + dual.
  • FAQ.md — common questions (4090 / 5090 support, why MTP not EAGLE, Copilot Gateway, what's a cliff, etc.).
  • EXAMPLES.md — Python / TS / curl client snippets + IDE connection settings.
  • HARDWARE.md — Ampere SM 8.6 specifics, NVLink (declined), power caps.
  • DUAL_CARD.md — when you need what single-card can't deliver.