Phase 3 of the v7.66 migration: documents the new state, regenerates performance/VRAM charts, posts cross-rig data to Sander on discussion #19 + issues #15/#16/#17. What changed ------------ **docs/SINGLE_CARD.md** - Updated TL;DR table: long-text 180K + 0.95, long-vision 145K + 0.95, bounded-thinking 180K + 0.95. - Removed Cliff 1 mech B "limitation to know" — now closed. - Added "What was Cliff 1 mech B (now closed) ✅" historical note. - Updated activation budget rationale to reflect PN12+PN25 pool residence. **docs/DUAL_CARD.md** - Bench protocol substrate: Genesis v7.65 → v7.66 dev tip. **docs/CLIFFS.md** - "vLLM pin compatibility status" rewritten for v7.66 + Cliff 1 mech B closure. - Replaced "v0.20 unblock" prose with "Cliff 1 mech B — what closed it" section explaining the two compounding fixes (PN25 v3 + PN30 dst-shaped). - Updated Genesis patches table with PN30 + PN33 + Sander v7.66 PN25. - Added "Local sidecars retained on master" table — 4 sidecars, why each one is still needed. - Updated "Validation across all 4 TQ3 variants" to 2026-05-02 numbers (180K / 145K / 180K / 262K — all 6/7 probes pass). **docs/UPSTREAM.md** - Genesis issue tracker updated with v7.66 cross-rig findings: - #16 PN25: both v7.65 and v7.66 mechanisms fail on TP=1 - #17 PN30: layout-correctness diagnosis + our corrected fix - #15 PN31: doesn't fit on 24 GB (lower mem-util sufficient) - PN33 partial (boot-time closes, runtime decode still fires) **docs/engines/VLLM.md, README.md, model README** - Genesis pin references bumped d89a089 → fc89395. **models/qwen3.6-27b/CHANGELOG.md** - New entry: "2026-05-02 — Genesis v7.66 + Cliff 1 mech B closed ⭐" with full validation matrix, sidecar inventory, and links to per-config result summaries. **tools/charts/gen-perf.py + gen-vram.py** - Substrate footnote: v7.65 dev (d89a089) → v7.66 dev (fc89395) - Compose ctx labels: long-text 214K → 180K, long-vision 198K → 145K, bounded-thinking 214K → 180K, mem-util 0.985 → 0.95 - Regenerated all 14 chart files (performance + vram, single + dual + combined). Cross-rig data posted to Sander ------------------------------- - [discussion #19 reply](https://github.com/noonghunna/club-3090/discussions/19#discussioncomment-16785590) — comprehensive update covering all 4 patches (PN33 partial, PN25 still TP=1-broken, PN30 layout diagnosis + corrected fix, PN31 still 24 GB-incompatible) - [genesis-vllm-patches#16 update](https://github.com/Sandermage/genesis-vllm-patches/issues/16#issuecomment-4362835267) — v7.66 mechanism still TP=1-broken - [genesis-vllm-patches#17 update](https://github.com/Sandermage/genesis-vllm-patches/issues/17#issuecomment-4362836196) — PN30 layout-correctness diagnosis + corrected fix offered - [genesis-vllm-patches#15 update](https://github.com/Sandermage/genesis-vllm-patches/issues/15#issuecomment-4362836881) — PN31 cross-rig confirmation Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> Co-Authored-By: Codex CLI (ChatGPT) <[email protected]>
Inference engines for Qwen3.6-27B — comparison + quick recipes
This repo's main path is vLLM because it has the deepest support for Qwen3-Next features (vision, MTP, TurboQuant, full OpenAI API parity). But the model also runs on llama.cpp and SGLang with different trade-offs. This page compares the three; per-engine pages have setup instructions.
🔁 Coming from the README's Quick start? It already shipped you the vLLM path. Skim this comparison to see what the alternatives look like, then pick a per-engine page if you want to try one.
At a glance
| Engine | Status on this stack | Per-stream TPS (1× 3090) | Max ctx (1× 3090) | Vision | Tool calls | Spec-decode | OpenAI API parity |
|---|---|---|---|---|---|---|---|
| vLLM ⭐ | Validated, production-grade (this repo) | 50-53 narr / 66-70 code | 48K default · 75K IDE-agent · 198K vision · 214K text-only | ✅ | ✅ | ✅ MTP n=3 | ✅ Full |
| llama.cpp | Works mainline + Luce DFlash fork for spec-decode | 35-60 (varies by quant + KV type) | 262K (Q4_K_M + q4_0 KV) | ✅ (via mmproj) | ⚠️ Limited (no auto-tool-choice in server) | ✅ DFlash N=5 in fork | ⚠️ Partial |
| SGLang | Blocked by same Marlin pad-sub-tile-n bug (vllm#40361 / sglang equivalent); EAGLE spec-decode separately blocked by GDN/DeltaNet rollback | n/a (untested at this state) | n/a | ✅ | ✅ | ⚠️ EAGLE blocked on hybrid | ✅ Full |
Pros / cons matrix
vLLM ⭐
Pros:
- Deepest Qwen3-Next feature support upstream
- TurboQuant 3-bit KV cache (lets us reach 198K + vision or 214K text-only on a single 3090; 262K dual-card)
- MTP speculative decoding works out of the box
- Genesis patch ecosystem (Sandermage's tree fixes many compatibility edges)
- Full OpenAI API parity (chat, vision, tools, streaming, reasoning, structured output)
- Active development — bugs we hit get triaged within days
Cons:
- Heavyweight — Docker image is ~9 GB
- Longer cold-start (~2 min for compile + cudagraph capture)
- Sensitive to upstream API drift across nightly versions (we pin to a specific nightly SHA —
7a1eb8ac=0.20.1rc1.dev16since 2026-05-01 — to avoid this) - Frontier features sometimes ship with bugs we have to patch around (the whole reason this repo exists)
When to pick: Production / serious local work / anything that needs the full feature set.
llama.cpp
Pros:
- Lightweight — single binary, ~50 MB
- Fastest cold-start (~30 sec)
- Lowest VRAM overhead (no inference framework taxes)
- GGUF support for many quant formats (Q4_K_M, Q5_K_S, IQ4_XS, etc.)
- Works on AMD + Intel + Apple Silicon (vLLM is NVIDIA-only)
- Active community, lots of distros / wrappers (Ollama, LM Studio, LocalAI, etc.)
Cons:
- Qwen3-Next family support is a moving target — needs the right binary build
- Server feature parity behind vLLM (no auto-tool-choice in upstream
server; need wrapper) - DFlash spec-decode requires a fork (Luce's llama-cpp-dflash)
- Concurrent serving is single-threaded by default (the server forks per request — sluggish under concurrent load)
- No TurboQuant equivalent → max usable context is much lower (~64K with Q4_K_M on 24 GB)
When to pick: Quick experiments, embedded use, non-NVIDIA hardware, when you want simplicity over feature completeness.
SGLang
Pros:
- Designed for high-throughput serving — RadixAttention prefix sharing, structured-output-aware scheduling
- Often beats vLLM by 10-30% on multi-tenant throughput when both work
- First-class OpenAI API
- Good support for batched structured output (constraint decoding)
Cons:
- Currently blocked on this stack by the same Marlin pad-sub-tile-n bug we hit on vLLM TP=2. Same kernel-line fix applies (would need a similar patch on SGLang's side or for them to pick up the upstream fix).
- EAGLE spec-decode (their MTP equivalent) is separately blocked by the DeltaNet/GDN hybrid layer not supporting KV rollback — this is a Qwen3-Next architectural issue, not SGLang-specific.
- Smaller community than vLLM; fewer eyes on Qwen3-Next bugs.
When to pick: Production multi-tenant serving on models that work cleanly on it (not yet Qwen3.6-27B-int4-AutoRound — track the unblock list below).
Watch list to unblock SGLang on this stack:
- Marlin pad-sub-tile-n landing (we filed PR #40361 on vLLM; the same fix applies to SGLang's Marlin call site)
- DeltaNet KV rollback support upstream (vllm#39931 / issue #40124 land would unblock EAGLE on Qwen3-Next family across engines)
How to choose
| Your priority | Pick | Why |
|---|---|---|
| Full feature set, MTP spec-decode, OpenAI API parity | vLLM + Lorbus AutoRound | This repo's path. 51-70 TPS depending on workload, all features, prefill-safe at 48K default. |
| Maximum context (262K) on one 3090 | llama.cpp + UD-Q3_K_XL or Q4_K_M + q4_0 KV | Smaller quants leave 8-10 GB headroom for KV at 262K. ~35-45 TPS sustained. |
| Best concurrent throughput on dual 3090 | vLLM TP=2 + Turbo (TQ3) | 4 streams at full 262K, ~200 TPS aggregate. See companion repo. |
| Non-NVIDIA hardware (AMD / Intel / Apple) | llama.cpp | Only engine with cross-platform support. |
| Lightest setup, fastest cold start | llama.cpp | Single binary, ~30s cold start. Good for embedded use, quick experiments. |
| High-throughput multi-tenant serving | SGLang (when unblocked — currently blocked on Qwen3.6) | RadixAttention prefix sharing wins at scale. Watch list in SGLANG.md. |
Quant choice (orthogonal to engine choice)
The model itself comes in several quant formats. Engine-quant compatibility:
| Quant | Disk size | Engine fit | Notes |
|---|---|---|---|
| AutoRound int4 (Lorbus) | ~18-19 GB | vLLM ✅ · llama.cpp ❌ · SGLang (when unblocked) | This repo's choice. W4A16, group_size=128, BF16 mtp.fc head. Required for vLLM's MTP spec-decode. |
| GPTQ int4 | ~16.5-17 GB | vLLM ✅ · llama.cpp ❌ · SGLang ✅ | Mature, broadly supported. Slightly smaller disk than AutoRound. |
| AWQ int4 | ~16-17 GB | vLLM ✅ · llama.cpp ❌ · SGLang ✅ | Strong baseline, compatible with Marlin kernels. |
| GGUF Q4_K_M | ~16.8 GB | llama.cpp ✅ · vLLM ⚠️ experimental · SGLang ❌ | The default GGUF mid-range quant. Strong quality, broad ecosystem (Ollama, LM Studio, etc). |
| GGUF UD-Q3_K_XL (Unsloth) | ~14.5 GB | llama.cpp ✅ | Smaller than 4-bit options. Quality cost is small on Qwen3.6 (quantization-friendly), buys substantial KV cache room. |
| GGUF Q3_K_M | ~13.6 GB | llama.cpp ✅ | More aggressive 3-bit; quality cost real but acceptable for many workloads. |
AutoRound vs GPTQ vs AWQ (within vLLM)
All three are 4-bit weight-only quantization for vLLM. Differences:
| Aspect | AutoRound | GPTQ | AWQ |
|---|---|---|---|
| Method | Signed gradient descent jointly optimizing rounding + scaling | Layer-wise Hessian-based error minimization | Activation-aware salience scaling, then RTN |
| Calibration set | Small (~128-512 samples) | Larger (~1024-2048) | Small-medium |
| Quantization time | Minutes to ~1-2 hours for 27B | Slower for same model | Fast |
| Accuracy at 4-bit | Typically slightly best on hard reasoning (MMLU/GPQA/Math style) | Strong baseline; 0.5-2% behind AutoRound on average | Comparable to GPTQ; depends on tuning |
| Ultra-low bits (3, 2) | Strongest at <4 bit | Degrades faster below 4 bit | Middle of the pack |
| Marlin kernel support | ✅ (via the kernel-line fix in our vllm#40361) | ✅ (mature) | ✅ |
| Ecosystem | Newer, growing fast (Intel-maintained) | Most mature, broadest tool support | Strong vLLM/SGLang support |
Why we picked AutoRound for this repo: Lorbus's AutoRound quant ships mtp.fc.weight as BF16 (preserved at higher precision), which lets vLLM's Qwen3_5MTP loader actually load the head and run multi-token prediction at high acceptance rates (~80% per-position-1, AL ~3.5). GPTQ-quantized variants of the MTP head silently fail to load → 0% draft acceptance. So AutoRound isn't just "slightly better quality" here — it's the only path to working MTP spec-decode in vLLM today.
If MTP isn't a priority for your workload, GPTQ or AWQ are equally valid.
Per-engine pages
- VLLM.md — current setup (what this repo ships). Brief recap + tuning levers.
- LLAMA_CPP.md — quick GGUF recipe, vision via mmproj, Luce DFlash fork pointer for spec-decode, gotchas around server feature parity.
- SGLANG.md — current blocked state, what would unblock, when to revisit. TBD recipe placeholder until either Marlin pad lands upstream or DeltaNet rollback lands.
See also
- docs/INTERNALS.md — why this repo picked vLLM specifically (the 9-probe forensics + upstream tracker)
- docs/SINGLE_CARD.md and docs/DUAL_CARD.md — workload-specific configs by hardware count
- LEARNINGS.md (parent stack) — why vLLM, why these patches