Consolidates and supersedes:
- noonghunna/qwen36-27b-single-3090
- noonghunna/qwen36-dual-3090
The two predecessor repos partitioned by card count (1× vs 2×). This
repo partitions by engine instead, which matches how users actually
decide ("vLLM or llama.cpp?" before "1 card or 2"). Card count becomes
a config variant within each engine.
Structure (model-agnostic from day 1):
docs/ cross-model engine + hardware docs
engines/ vLLM / llama.cpp / SGLang comparison + per-engine deep dives
HARDWARE.md Ampere SM 8.6+, NVLink, power, VRAM ceilings
GLOSSARY.md plain-language definitions
img/ illustrations (vram-budget.svg)
ARCHITECTURE.md how this stack thinks about LLM serving on 24 GB
models/<model-name>/ everything specific to a model
qwen3.6-27b/ today's only model
README.md / INTERNALS.md / USE_CASES.md / CHANGELOG.md
vllm/ vLLM-specific configs for this model
compose/ docker-compose files (single + dual variants)
patches/ tolist_cudagraph + Marlin pad notes
llama-cpp/ llama.cpp recipes for this model
recipes/ shell scripts (single-card default + 262K max-ctx)
sglang/ SGLang status (currently blocked)
scripts/ shared, model-aware
setup.sh bash setup.sh <model> → downloads + verifies
verify.sh / verify-full.sh smoke + functional tests
bench.sh canonical TPS bench
vLLM compose variants (all under models/qwen3.6-27b/vllm/compose/):
Single-card:
docker-compose.yml ⭐ DEFAULT — TQ3 + Genesis P65, 48K, 51/68 TPS
docker-compose.fast-chat.yml fp8 + 20K, 55/70 TPS — fastest at small ctx
docker-compose.tools-text.yml fp8 + 75K, 53/70 TPS — best for long single prompts
docker-compose.no-genesis-mtp.yml control variant
docker-compose.minimal.yml no spec-decode
Dual-card:
docker-compose.dual.yml ⭐ fp8 + 262K + MTP + vision, 71/89 TPS
docker-compose.dual-turbo.yml TQ3 + Genesis v7.14 — 4-stream concurrency
docker-compose.dual-dflash.yml DFlash N=5 + 185K + vision — 78/128 TPS
docker-compose.dual-dflash-noviz.yml DFlash + 200K text-only
llama.cpp recipes (under models/qwen3.6-27b/llama-cpp/recipes/):
single-card-default.sh Q4_K_M + 65K
single-card-max-ctx.sh Q4_K_M + q4_0 KV at full 262K — the standout recipe
Old repos remain readable for issue history + external links (Medium,
Reddit, Twitter, Sandermage's PR threads). New issues should be filed
here.
Credits in README. Apache 2.0.
5.6 KiB
SGLang — currently blocked, watch list
SGLang is a strong alternative to vLLM for high-throughput multi-tenant serving — RadixAttention prefix sharing, structured-output-aware scheduling, batch structured decoding. It often beats vLLM by 10-30% on multi-tenant aggregate throughput when both work.
Currently SGLang doesn't run cleanly on this stack (Qwen3.6-27B-int4-AutoRound + 3090). Below: what's blocked, why, and what would unblock it.
TL;DR
- ❌ Blocked by the same Marlin pad-sub-tile-n bug we hit on vLLM TP=2 (same kernel-line fix applies).
- ❌ EAGLE spec-decode (their MTP equivalent) blocked separately by DeltaNet/GDN hybrid layer not supporting KV rollback.
- ✅ Will likely unblock when (a) Marlin pad lands on SGLang, and (b) DeltaNet KV rollback support lands upstream (vllm#39931 / issue #40124 cross-engine).
- ⚠️ TBD recipe — we don't ship a working SGLang config yet. We'll add one when the blockers clear.
Pros (when it works)
| Pro | Detail |
|---|---|
| High-throughput multi-tenant serving | RadixAttention shares prefix KV across requests automatically. Multi-tenant aggregate often beats vLLM by 10-30%. |
| Structured-output-aware scheduling | Prioritizes batched constraint-decoding requests for better GPU utilization. |
| First-class OpenAI API | Same level of API parity as vLLM. |
| Active development | Smaller community than vLLM but engaged maintainers. |
Cons (right now, on this model class)
| Con | Detail |
|---|---|
| Marlin pad-sub-tile-n bug | Same INT4 kernel issue we filed vllm#40361 for. SGLang's Marlin call site has the equivalent constraint and would crash on Lorbus's quant + TP=2. We haven't filed an SGLang PR — would need someone to port the fix or wait for SGLang to pick up the upstream Marlin fix. |
| EAGLE spec-decode blocked on hybrid attention | Qwen3-Next's DeltaNet layers don't support KV rollback the way standard attention does. EAGLE (and any speculative-decode method that needs rollback) breaks. This is architectural; needs upstream flash-linear-attention to add rollback support, then SGLang to integrate. |
| Smaller community than vLLM | Fewer eyes on Qwen3-Next bugs. When something breaks here, we may be on our own. |
Watch list — what would unblock SGLang on this stack
1. Marlin pad-sub-tile-n landed on SGLang
Two paths:
- Upstream Marlin landing — if vllm-project's Marlin gets the fix and SGLang picks it up via shared upstream code, we get it for free.
- SGLang-side patch — same kernel line check + pad logic, applied in SGLang's Marlin call site. Could be filed as an SGLang PR (we haven't yet — vLLM was the priority).
Track: our vllm#40361. When it lands and propagates, SGLang can pick it up.
2. DeltaNet KV rollback support for EAGLE / spec-decode
EAGLE (SGLang's MTP equivalent) requires the model to support rolling back the KV cache when speculative tokens are rejected. Standard attention layers do this trivially — KV cache is just discarded for the rejected positions. DeltaNet layers maintain a recurrent state that doesn't roll back cleanly.
Track:
- vllm#39931 — Qwen3-Next hybrid attention rollback support
- vllm#40124 — related upstream issue
flash-linear-attentionlibrary — the underlying linear-attention impl needs rollback hooks added; this is the architectural change
When this lands, EAGLE on Qwen3-Next becomes possible across all engines (vLLM, SGLang, etc).
3. (Optional) FlashKDA Hopper kernels for Ampere
A separate research direction — Kimi Delta Attention (KDA) / FlashKDA brings prefill-speedup CUTLASS kernels for DeltaNet-family models, but they're targeting Hopper today. Not usable on Ampere. If/when an Ampere port appears or vLLM/SGLang adopt the flash-linear-attention backend, this would substantially reduce the GDN forward cost on this stack. Watch list, not unblocker.
Recipe — TBD
We don't have a working SGLang config yet for this model. Here's what one would look like (untested, just as a placeholder):
# IF the Marlin pad fix were already in SGLang:
docker run --gpus all --shm-size 16g \
-v /mnt/models/huggingface/qwen3.6-27b-autoround-int4:/model \
-p 8020:30000 \
lmsysorg/sglang:latest \
python -m sglang.launch_server \
--model-path /model \
--quantization awq_marlin \
--tp-size 1 \
--mem-fraction-static 0.92 \
--context-length 65536 \
--speculative-algorithm EAGLE \
--speculative-draft-model-path <eagle-draft-path> \
--speculative-num-steps 3
Don't run this — it'll crash on the Marlin pad bug. Listed here as the shape of the recipe we'd ship once unblocked.
When to revisit
We'll add a working recipe (and lift this page from "blocked" to "validated alternative") when both:
- A Marlin pad-equivalent fix lands on SGLang (upstream or via PR), AND
- DeltaNet KV rollback support lands upstream (or we accept running without spec-decode)
Until then, the vLLM path is the validated option for serious local use.
See also
- VLLM.md — current validated path
- LLAMA_CPP.md — alternative for lighter setups
- SGLang docs — official documentation
- Cross-engine architecture issue tracking: vllm#39931 (DeltaNet rollback)