Files
club-3090/docs/engines/SGLANG.md
T
noonghunna 3fa33332ce Initial commit — club-3090: model-agnostic LLM serving recipes for RTX 3090
Consolidates and supersedes:
  - noonghunna/qwen36-27b-single-3090
  - noonghunna/qwen36-dual-3090

The two predecessor repos partitioned by card count (1× vs 2×). This
repo partitions by engine instead, which matches how users actually
decide ("vLLM or llama.cpp?" before "1 card or 2"). Card count becomes
a config variant within each engine.

Structure (model-agnostic from day 1):

  docs/                       cross-model engine + hardware docs
    engines/                    vLLM / llama.cpp / SGLang comparison + per-engine deep dives
    HARDWARE.md                 Ampere SM 8.6+, NVLink, power, VRAM ceilings
    GLOSSARY.md                 plain-language definitions
    img/                        illustrations (vram-budget.svg)
    ARCHITECTURE.md             how this stack thinks about LLM serving on 24 GB

  models/<model-name>/        everything specific to a model
    qwen3.6-27b/                today's only model
      README.md / INTERNALS.md / USE_CASES.md / CHANGELOG.md
      vllm/                     vLLM-specific configs for this model
        compose/                  docker-compose files (single + dual variants)
        patches/                  tolist_cudagraph + Marlin pad notes
      llama-cpp/                llama.cpp recipes for this model
        recipes/                  shell scripts (single-card default + 262K max-ctx)
      sglang/                   SGLang status (currently blocked)

  scripts/                    shared, model-aware
    setup.sh                    bash setup.sh <model> → downloads + verifies
    verify.sh / verify-full.sh  smoke + functional tests
    bench.sh                    canonical TPS bench

vLLM compose variants (all under models/qwen3.6-27b/vllm/compose/):

  Single-card:
    docker-compose.yml             ⭐ DEFAULT — TQ3 + Genesis P65, 48K, 51/68 TPS
    docker-compose.fast-chat.yml   fp8 + 20K, 55/70 TPS — fastest at small ctx
    docker-compose.tools-text.yml  fp8 + 75K, 53/70 TPS — best for long single prompts
    docker-compose.no-genesis-mtp.yml control variant
    docker-compose.minimal.yml     no spec-decode

  Dual-card:
    docker-compose.dual.yml             ⭐ fp8 + 262K + MTP + vision, 71/89 TPS
    docker-compose.dual-turbo.yml       TQ3 + Genesis v7.14 — 4-stream concurrency
    docker-compose.dual-dflash.yml      DFlash N=5 + 185K + vision — 78/128 TPS
    docker-compose.dual-dflash-noviz.yml DFlash + 200K text-only

llama.cpp recipes (under models/qwen3.6-27b/llama-cpp/recipes/):

  single-card-default.sh    Q4_K_M + 65K
  single-card-max-ctx.sh    Q4_K_M + q4_0 KV at full 262K — the standout recipe

Old repos remain readable for issue history + external links (Medium,
Reddit, Twitter, Sandermage's PR threads). New issues should be filed
here.

Credits in README. Apache 2.0.
2026-04-28 10:24:14 +00:00

5.6 KiB

SGLang — currently blocked, watch list

SGLang is a strong alternative to vLLM for high-throughput multi-tenant serving — RadixAttention prefix sharing, structured-output-aware scheduling, batch structured decoding. It often beats vLLM by 10-30% on multi-tenant aggregate throughput when both work.

Currently SGLang doesn't run cleanly on this stack (Qwen3.6-27B-int4-AutoRound + 3090). Below: what's blocked, why, and what would unblock it.


TL;DR

  • ❌ Blocked by the same Marlin pad-sub-tile-n bug we hit on vLLM TP=2 (same kernel-line fix applies).
  • ❌ EAGLE spec-decode (their MTP equivalent) blocked separately by DeltaNet/GDN hybrid layer not supporting KV rollback.
  • ✅ Will likely unblock when (a) Marlin pad lands on SGLang, and (b) DeltaNet KV rollback support lands upstream (vllm#39931 / issue #40124 cross-engine).
  • ⚠️ TBD recipe — we don't ship a working SGLang config yet. We'll add one when the blockers clear.

Pros (when it works)

Pro Detail
High-throughput multi-tenant serving RadixAttention shares prefix KV across requests automatically. Multi-tenant aggregate often beats vLLM by 10-30%.
Structured-output-aware scheduling Prioritizes batched constraint-decoding requests for better GPU utilization.
First-class OpenAI API Same level of API parity as vLLM.
Active development Smaller community than vLLM but engaged maintainers.

Cons (right now, on this model class)

Con Detail
Marlin pad-sub-tile-n bug Same INT4 kernel issue we filed vllm#40361 for. SGLang's Marlin call site has the equivalent constraint and would crash on Lorbus's quant + TP=2. We haven't filed an SGLang PR — would need someone to port the fix or wait for SGLang to pick up the upstream Marlin fix.
EAGLE spec-decode blocked on hybrid attention Qwen3-Next's DeltaNet layers don't support KV rollback the way standard attention does. EAGLE (and any speculative-decode method that needs rollback) breaks. This is architectural; needs upstream flash-linear-attention to add rollback support, then SGLang to integrate.
Smaller community than vLLM Fewer eyes on Qwen3-Next bugs. When something breaks here, we may be on our own.

Watch list — what would unblock SGLang on this stack

1. Marlin pad-sub-tile-n landed on SGLang

Two paths:

  • Upstream Marlin landing — if vllm-project's Marlin gets the fix and SGLang picks it up via shared upstream code, we get it for free.
  • SGLang-side patch — same kernel line check + pad logic, applied in SGLang's Marlin call site. Could be filed as an SGLang PR (we haven't yet — vLLM was the priority).

Track: our vllm#40361. When it lands and propagates, SGLang can pick it up.

2. DeltaNet KV rollback support for EAGLE / spec-decode

EAGLE (SGLang's MTP equivalent) requires the model to support rolling back the KV cache when speculative tokens are rejected. Standard attention layers do this trivially — KV cache is just discarded for the rejected positions. DeltaNet layers maintain a recurrent state that doesn't roll back cleanly.

Track:

When this lands, EAGLE on Qwen3-Next becomes possible across all engines (vLLM, SGLang, etc).

3. (Optional) FlashKDA Hopper kernels for Ampere

A separate research direction — Kimi Delta Attention (KDA) / FlashKDA brings prefill-speedup CUTLASS kernels for DeltaNet-family models, but they're targeting Hopper today. Not usable on Ampere. If/when an Ampere port appears or vLLM/SGLang adopt the flash-linear-attention backend, this would substantially reduce the GDN forward cost on this stack. Watch list, not unblocker.


Recipe — TBD

We don't have a working SGLang config yet for this model. Here's what one would look like (untested, just as a placeholder):

# IF the Marlin pad fix were already in SGLang:
docker run --gpus all --shm-size 16g \
  -v /mnt/models/huggingface/qwen3.6-27b-autoround-int4:/model \
  -p 8020:30000 \
  lmsysorg/sglang:latest \
  python -m sglang.launch_server \
    --model-path /model \
    --quantization awq_marlin \
    --tp-size 1 \
    --mem-fraction-static 0.92 \
    --context-length 65536 \
    --speculative-algorithm EAGLE \
    --speculative-draft-model-path <eagle-draft-path> \
    --speculative-num-steps 3

Don't run this — it'll crash on the Marlin pad bug. Listed here as the shape of the recipe we'd ship once unblocked.


When to revisit

We'll add a working recipe (and lift this page from "blocked" to "validated alternative") when both:

  • A Marlin pad-equivalent fix lands on SGLang (upstream or via PR), AND
  • DeltaNet KV rollback support lands upstream (or we accept running without spec-decode)

Until then, the vLLM path is the validated option for serious local use.


See also