Files
club-3090/docs/INFERENCE_ENGINES.md
T
noonghunnaandClaude Opus 4.7 acd7ffb67c restructure: promote topology to a directory level (single/dual/multi4)
Compose files now live under `<model>/<engine>/compose/<topology>/<file>.yml`,
with topology as a folder rather than a filename prefix. Solves all 7
inconsistencies surfaced in the post-rename audit (single-card composes
without `single-` prefix, unsuffixed `docker-compose.yml` ambiguity,
fine-tunes encoding model name in filename, etc.) by making the directory
hierarchy enforce the convention.

Layout:
  models/<model>/<engine>/compose/<topology>/<feature>.yml

Where:
  - <model>:    qwen3.6-27b, gemma-4-31b
  - <engine>:   vllm, llama-cpp, sglang
  - <topology>: single, dual, multi3, multi4, multi8
  - <feature>:  docker-compose.yml (default) | turbo.yml | dflash.yml | etc.

Each topology subdir has a `docker-compose.yml` for the recommended
starter — bare `cd <topology> && docker compose up` works because
docker compose finds that filename automatically. Variants drop the
`docker-compose.` prefix since they're invoked via `-f` flag.

27 compose file moves total:
- 18 Qwen vLLM composes redistributed across single/dual/multi4
- 2 Qwen llama-cpp composes into single/
- 6 Gemma vLLM composes redistributed across single/dual
- 1 untracked qwopus-bf16mtp moved to dual/

Inside each compose: relative paths to `../patches/` and `../cache/`
bumped to `../../patches/` / `../../cache/`, and `../../../../models-cache`
to `../../../../../models-cache` (one extra `..` for the new depth).

Reference updates across 148 files (BENCHMARKS, all docs, CHANGELOGs,
sibling-table cross-references in compose headers, scripts, patch
READMEs, .github issue templates, tools/residency-instrument).

scripts/switch.sh VARIANTS map updated; tags themselves unchanged
(`vllm/dual` → `dual/docker-compose.yml`, `vllm/dual4` → `multi4/docker-compose.yml`,
`vllm/gemma-mtp` → `gemma-4-31b/.../dual/docker-compose.yml`, etc.).

AGENTS.md "Compose layout" section rewritten to describe the new
hierarchy, with concrete examples and the fine-tune exception
(`dual/carnice-bf16mtp.yml` carries the fine-tune name as a filename
prefix until the fine-tune graduates to its own model directory).

All switch.sh paths verified to resolve to actual files post-move.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-09 12:11:17 +00:00

26 KiB
Raw Blame History

Inference engine comparison — vLLM, llama.cpp, SGLang, ktransformers, ik_llama.cpp

Pragmatic comparison of the five engines this stack interacts with most. Versions captured 2026-05-07 — all five projects move fast; verify against upstream release notes before making a long-horizon decision. Linked sources at the bottom.

This doc isn't an evangelism piece. It's a picker — which engine to reach for when the workload demands a specific feature, and where each engine still has structural gaps. Coverage is biased toward what matters on club-3090's hardware class (consumer Ampere/Ada/Blackwell + 24-32 GB VRAM cards + community workstations), not enterprise H100/B200 deployments.

Note on ik_llama.cpp: it's a Iwan Kawrakow fork of llama.cpp. Most rows mirror mainline because it inherits the codebase. Differences are flagged where they matter — chiefly MTP merged on main (vs mainline's open PR), fused MoE kernels for DeepSeek-R1 / Kimi, and the IQ_K quant series (IQ4_KT, IQ3_K_R4) that mainline doesn't have.


TL;DR — pick by workload

Workload Engine Why
Production multi-tenant chat / API (24-48 GB VRAM, fits-VRAM models) vLLM Continuous batching + paged attention + extensive tool-call/structured-output support; Genesis patches close most Qwen3-Next gaps
Single-card max-quality on small-VRAM rigs (12-24 GB), bulletproof llama.cpp Engine-agnostic memory model; no Cliff 1/2 footguns; broadest model coverage; q4_0/q8_0 KV
Hyper-optimized prefix caching, structured outputs, multimodal at scale SGLang RadixAttention-class prefix cache + FSM-native structured outputs + best-in-class spec-decode V2
Big MoE (>VRAM) on consumer hardware (M2/Kimi/DeepSeek-class) ktransformers (or SGLang+kt-kernel) Router-aware hot-expert caching; 1.5-2× decode TPS over llama.cpp's layer-uniform --n-cpu-moe
Apple Silicon llama.cpp (Metal) or SGLang (MLX backend, v0.5.10+) Both work; MLX backend is newer but native
Qwen3.x + MTP on llama.cpp without PR-branch building ik_llama.cpp Qwen MTP merged on main (vs mainline's open PR #22673); GLM-4.x MTP also working
DeepSeek-R1 / Kimi-K2 / large MoE on consumer hardware (alternative to ktransformers) ik_llama.cpp Fused MoE kernels + smart expert reduction (-ser) + on-the-fly MLA tensors; ships as a llama.cpp fork rather than separate engine

Versions + release cadence (2026-05-07)

Engine Latest Release cadence Maturity License
vLLM v0.20.1 (May 4) + nightlies Major/month, nightlies/day Production Apache 2.0
llama.cpp b9050 (May 7) Tagged releases multi-daily Production MIT
SGLang v0.5.11 (May 5) ~Monthly minors, weekly patches Production Apache 2.0
ktransformers v0.6.2 (May 3) ~Monthly minors Research → graduating Apache 2.0
ik_llama.cpp rolling (active main, no tagged releases) PR-by-PR, smaller community Production fork MIT

All five are actively developed. ktransformers is positioned as research but production paths exist (kt-kernel + SGLang). ik_llama.cpp tracks mainline llama.cpp's API surface but adds quant types (IQ_K series), MTP-merged-on-main, and fused MoE kernels for big-MoE workloads — chiefly used as a drop-in llama-server replacement when those features matter.


Hardware support

Feature vLLM llama.cpp SGLang ktransformers ik_llama.cpp
NVIDIA CUDA (CC 7.0 Volta) ❌ (needs ≥7.5) ✅ ❌ (needs ≥8.0) ❌ (needs ≥8.0) ✅
NVIDIA CC 7.5 Turing ✅ ✅ ⚠️ Limited ❌ ✅
NVIDIA CC 8.0+ Ampere (3090) ✅ ✅ ✅ ✅ ✅
NVIDIA CC 8.6 Ampere consumer ✅ ✅ ✅ ✅ ✅
NVIDIA CC 8.9 Ada (4090) ✅ ✅ ✅ ✅ ✅
NVIDIA CC 9.0+ Hopper (H100) ✅ + FA3 ✅ ✅ + TRT-LLM NSA ✅ ✅
NVIDIA CC 12.0 Blackwell (5090) ✅ ✅ ✅ + 8× 5090 validated ✅ + kt-kernel validated ✅
AMD ROCm ✅ ✅ ✅ + DFLASH on ROCm ⚠️ Limited ✅
Apple Silicon (Metal) ❌ ✅ first-class ✅ MLX backend (v0.5.10) ❌ ✅ first-class
Intel GPU (XPU/SYCL) ⚠️ ✅ ⚠️ ❌ ✅
CPU-only inference ⚠️ Slow path ✅ ⚠️ ❌ (kernel library) ✅
Vulkan (cross-platform) ❌ ✅ ❌ ❌ ✅

Recommendation for our stack (RTX 3090 sm_86 + 2× consumer-board PCIe 4.0): all five work. Pick by feature, not hardware support.


Quantization formats

Weight quants

Format vLLM llama.cpp SGLang ktransformers ik_llama.cpp
GGUF (Q, K, IQ*)** ⚠️ Experimental (official docs flag "under-optimized"; single-file only — multi-file needs gguf-split merge; UD-* prefixes via PR #39471 merged 2026-04-10; tokenizer conversion unstable on large-vocab models like Qwen3.6) ✅ Native ⚠️ Recent ⚠️ ✅ Native + IQ_K series
AutoRound INT4 (Marlin) ✅ + our PR #40361 ❌ ✅ ⚠️ ❌
GPTQ INT4 ✅ ❌ ✅ ✅ ❌
AWQ INT4 ✅ ❌ ✅ ⚠️ ❌
FP8 (e4m3, e5m2) ✅ multi-backend (Marlin/TRTLLM/MXFP8) ⚠️ Limited ✅ FlashInfer MXFP8 ✅ Native (M2.x) ⚠️ Limited (inherits llama.cpp)
FP4 / MXFP4 ✅ (online MoE quant) ⚠️ ✅ MXFP4 kernels ✅ DeepSeek-V4 / kt-kernel ⚠️ Limited
NVFP4 (Blackwell) ✅ rescaled weight scales ❌ ✅ ⚠️ ❌
INT8 ✅ ✅ ✅ ✅ AVX2-VNNI RAWINT4 (consumer CPU) ✅
BF16 ✅ ✅ ✅ ✅ ✅
FP16 ✅ ✅ ✅ ✅ ✅

KV cache types

KV format vLLM llama.cpp SGLang ktransformers ik_llama.cpp
FP16 / BF16 ✅ ✅ ✅ ✅ ✅
FP8 e4m3 / e5m2 ✅ ⚠️ ✅ ✅ ⚠️ (inherits llama.cpp)
INT8 per-token-head ⚠️ (PR #40391 for hybrid pages) ❌ ✅ ⚠️ ✅
Q4_0 (4.5 bit) ❌ ✅ ❌ ❌ ✅
Q8_0 (8.5 bit) ❌ ✅ ❌ ❌ ✅
TurboQuant 3-bit (TQ3) ✅ via Genesis ✅ (PR #21089 WIP) ⚠️ Same kernel bug as vLLM PR #40361 ❌ 🟡 Issue #1509 — CPU complete + CUDA written, awaiting merge
2-bit KV ✅ (PR #38479) ❌ ❌ ❌ ❌
CPU KV offload ✅ pluggable policies (PR #37160) ⚠️ via mmap ✅ Decode Radix Cache (v0.5.11) N/A ⚠️ via mmap
Disk KV offload ✅ FlexKV (PR #34328) + LMCache ⚠️ ⚠️ ❌ ⚠️

Notes: llama.cpp's q4_0 KV is the only engine choice when you need 4-bit KV at 200K+ ctx on 24 GB VRAM. vLLM's TurboQuant 3-bit KV (TQ3) is comparable ratio but Qwen3-Next-only via Sandermage's Genesis patches.


Speculative decoding

Method vLLM llama.cpp SGLang ktransformers ik_llama.cpp
MTP (Multi-Token Prediction) ✅ broad model coverage (Qwen3.x, Gemma4, MiniMax-M2, DeepSeek) ✅ via PR #22673 (am17an) — community-built ✅ Spec V2 default (overlap scheduling) ⚠️ via SGLang front-end ✅ Merged on main (GLM-4.x + Qwen MTP, no PR-branch needed)
EAGLE / EAGLE-3 ✅ Eagle3 (Qwen3.5, Gemma4, MiniMax-M2) ❌ ✅ Day-0 for newest models ❌ ❌
DFlash (block-diffusion) ✅ (PR #41703 Codex-rebased; Luce z-lab) ⚠️ Luce fork (server-only) ✅ DFLASH cross-backend incl. ROCm (v0.5.11) ❌ ❌
N-gram prompt-lookup ✅ GPU impl + async scheduler (PR #29184) ✅ ✅ ❌ ✅ (inherits llama.cpp)
Draft model (separate small) ✅ ✅ ✅ ❌ ✅
PFlash (prompt-compression) ❌ ⚠️ Luce experimental ❌ ❌ ❌
Async / overlap scheduling ✅ Zero-bubble (PR #32951) ❌ ✅ Spec V2 with overlap (default v0.5.11) ⚠️ via SGLang ❌

Notes on Qwen3-Next family (DeltaNet hybrid attention): only MTP works today. EAGLE/DFlash/draft/ngram all blocked on KV rollback support — see vLLM #39931. MTP is the rolling default for our shipped Qwen3.6-27B composes.


MoE features

Feature vLLM llama.cpp SGLang ktransformers ik_llama.cpp
Tensor Parallel ✅ ✅ ✅ ✅ ⚠️ Limited (inherits llama.cpp)
Expert Parallel (EP) ✅ Elastic EP M2 (PR #35627) ❌ ✅ Independent MoE/attention tuning (v0.5.11) ✅ kt-kernel ❌
Layer-uniform expert offload to CPU ⚠️ generic --cpu-offload-gb (catastrophic for MoE) ✅ --n-cpu-moe N / -ot regex ⚠️ via kt-kernel N/A ✅ -ot regex + -ser smart expert reduction (key strength)
Router-aware hot-expert caching ❌ ❌ (feature request #20757 open) ✅ via kt-kernel integration ✅ Native (purpose-built) ❌ (same gap as mainline)
Disk-offload for cold experts ⚠️ via FlexKV / LMCache ⚠️ via mmap ⚠️ via kt-kernel ⚠️ ⚠️ via mmap
Shared-expert handling ✅ (PR #35153 Oracle Flow) ✅ ✅ ✅ ✅ + fused MoE kernels (DeepSeek-R1 / Kimi optimized) ⭐
MXFP4 MoE online quant ✅ ❌ ✅ MXFP8/MXFP4 ✅ DeepSeek-V4-Flash (kt-kernel MXFP4 op) ❌
Disaggregated prefill/decode ✅ ~5% overhead reduction ❌ ✅ Decode Radix Cache (v0.5.11) ⚠️ via SGLang ❌

Bottom line on big-MoE: ktransformers (or kt-kernel via SGLang) is the only engine with router-aware hot-expert caching. Sustainable benchmark on DeepSeek-V2 (2× 3090): 9 TPS @ 45% VRAM (kt) vs 7.5 TPS @ 95% VRAM (llama.cpp --n-cpu-moe) — +20% TPS at half the VRAM.


Distributed / multi-card

Feature vLLM llama.cpp SGLang ktransformers ik_llama.cpp
Tensor Parallel (TP) ✅ ✅ -sm tensor ✅ ✅ ⚠️ Limited (inherits llama.cpp)
Pipeline Parallel (PP) ✅ + cudagraphs (PR #35162) ⚠️ Limited ✅ ✅ ⚠️ Limited
Expert Parallel (EP) ✅ Elastic EP M2 ❌ ✅ + all-reduce fusion ✅ kt-kernel ❌
Data Parallel (DP) ✅ ❌ ✅ ⚠️ ❌
Context Parallel (long-ctx) ✅ DeepSeek context-parallel ❌ ✅ Enhanced (v0.5.11) ❌ ❌
Disaggregated PD ✅ ~5% reduction ❌ ✅ Decode Radix Cache ❌ ❌
NVLink awareness ✅ ⚠️ Generic NCCL ✅ ⚠️ ⚠️ Generic NCCL
Custom all-reduce ✅ (must disable on PCIe-only) N/A ✅ N/A N/A
Patched-P2P (consumer GPUs) ✅ (Sam McLeod's guide) ✅ ✅ ✅ ✅

Cross-rig validation on club-3090's 2× 3090 PCIe: NVLink lift = +15-19% on DFlash paths (disc #19); patched-P2P captures ~60-80% of NVLink's lift on DFlash, ~13% on plain dual (issue #91/#95).


Memory / KV cache features

Feature vLLM llama.cpp SGLang ktransformers ik_llama.cpp
Paged attention ✅ Original (Berkeley) ⚠️ via continuous-batch slots ✅ ✅ via SGLang ⚠️ via parallel slots (inherits)
Prefix caching (in-VRAM) ✅ default-on ✅ slot-based ✅ RadixAttention (best-in-class) ✅ via SGLang ✅ slot-based (inherits)
CPU-tier prefix cache (warm restart) ✅ via LMCache connector ⚠️ via mmap ✅ Decode Radix Cache (v0.5.11) ⚠️ ⚠️ via mmap
Disk-tier prefix cache (cold restart) ✅ LMCache + FlexKV ⚠️ ✅ Decode Radix Cache ⚠️ ⚠️
Sparse attention (long ctx) ⚠️ HiSparse research ⚠️ ✅ HiSparse (v0.5.10) ❌ ⚠️
Mamba/SSM hybrid handling ✅ Mamba state corruption fixes (PR #37728) ⚠️ Limited ✅ SSM-FA hybrid via NIXL ⚠️ ⚠️ Limited
TMA / async copies (Hopper+) ✅ ⚠️ ✅ ✅ ⚠️
Chunked prefill ✅ default ⚠️ via parallel slots ✅ ✅ + Layerwise Prefill ⚠️ via parallel slots
Continuous batching ✅ ✅ via parallel slots ✅ ✅ via SGLang ✅ via parallel slots

Multimodal

Modality vLLM llama.cpp SGLang ktransformers ik_llama.cpp
Vision (image) ✅ broad VLM coverage (Gemma4, Granite Vision, Hunyuan v3, ViT cudagraphs) ✅ via --mmproj ✅ optimized encoders ⚠️ Limited ✅ via --mmproj + on-the-fly MLA tensors for DeepSeek
Audio ✅ Nemotron / Qwen3-Omni ✅ Granite Speech (b9045) ⚠️ Limited ❌ ✅ inherits Granite Speech support
Video ⚠️ Frame-by-frame ⚠️ ✅ Diffusion (LTX-2, FLUX) ❌ ⚠️
Image generation ❌ ❌ ✅ Diffusion (FLUX, Qwen-Image fused kernels) ❌ ❌

Practical for our stack: Gemma 4 vision + Qwen3.6 vision both work via vLLM (see models/gemma-4-31b/ and models/qwen3.6-27b/long-vision.yml); llama.cpp --mmproj is the single-card fallback.


Structured output / tool calling

Feature vLLM llama.cpp SGLang ktransformers ik_llama.cpp
OpenAI tool-call API compat ✅ ✅ via --jinja ✅ ✅ via SGLang ✅ via --jinja
Custom tool parsers ✅ Per-model (Gemma4, Kimi-K2.5, GigaChat 3.1, qwen3_coder) ⚠️ Generic ✅ ⚠️ ⚠️ Generic
Grammar (GBNF / EBNF) ✅ via outlines/lm-format-enforcer/xgrammar ✅ Native GBNF ✅ FSM-based ⚠️ ✅ Native GBNF
JSON schema mode ✅ ✅ ✅ Best-in-class FSM ⚠️ ✅
Function calling enforcement ✅ tool-choice ⚠️ Soft ✅ Strict ⚠️ ⚠️ Soft
Reasoning-channel separation ✅ qwen3 reasoning parser ⚠️ Default ON via peg-native + <think> parsing → routes to reasoning_content field. Most clients (incl. opencode) ignore this and hang (issue #97). Workaround: --reasoning-format none flag (now default in our llamacpp/default compose). ✅ ⚠️ ⚠️ Inherits llama.cpp default
Streaming tool-call deltas ✅ + Anthropic API compat ✅ ✅ + Responses API streaming ⚠️ ✅

Notes: SGLang's RadixAttention + native FSM makes structured output the headline strength. Our bounded-thinking compose uses vLLM's xgrammar; could re-do on SGLang for a probable speedup but vLLM is the daily-driver here.


Model coverage (latest architectures, 2026 lens)

Model family vLLM llama.cpp SGLang ktransformers ik_llama.cpp
Qwen3.5 / Qwen3.6 (incl. 80B-A3B) ✅ ✅ ✅ Day-0 (Qwen3.6 v0.5.11) ✅ ✅
Qwen3-Next family (DeltaNet hybrid) ✅ via Genesis patches ✅ ✅ ⚠️ ✅
Gemma 4 / Gemma-4 31B ✅ + MTP (PR #41745) + DFlash (PR #41703) ✅ via mmproj ✅ Day-0 ⚠️ ✅ via mmproj
DeepSeek V3 / R1 / V4-Flash ✅ ✅ ✅ + TRT-LLM NSA (3-5× on Blackwell) ✅ Native (kt-kernel MXFP4 for V4-Flash) ✅
Kimi-K2.5 / K2.6 ✅ tool parser (PR #37438) ✅ ✅ Day-0 K2.6 (v0.5.11) ✅ ✅
MiniMax-M2 / M2.5 / M2.7 (incl. REAP) ⚠️ via custom code ✅ via REAP'd GGUF ✅ Day-0 M2.5 / M2.6 ✅ Day-0 + native FP8 ✅ via REAP'd GGUF
GLM-4.5 / GLM-5 / GLM-5.1 ✅ ✅ ✅ Day-0 GLM-5.1 + MoE ✅ GLM-5 (v0.6.x) ✅
Mistral Medium 3.5 ⚠️ ✅ ✅ Day-0 ⚠️ ✅
GPT-OSS-120B ✅ ✅ ✅ ✅ via kt-kernel ✅
Llama 3 / Llama 4 ✅ ✅ ✅ ⚠️ Experimental L4 ✅
EXAONE-4.5 / Phi-4-reasoning-vision ✅ (v0.20) ⚠️ ⚠️ ❌ ⚠️

Day-0 wins for new MoE releases: SGLang and ktransformers both ship same-day support for new MiniMax / Kimi / GLM releases. vLLM tends to lag 1-2 weeks on new MoE architectures (PR-driven). llama.cpp typically lags 2-4 weeks (community-driven).


API / serving surface

Feature vLLM llama.cpp SGLang ktransformers ik_llama.cpp
OpenAI-compatible HTTP ✅ ✅ (llama-server) ✅ ✅ via SGLang ✅ (llama-server compatible CLI)
Anthropic API compat ⚠️ ❌ ✅ Direct (v0.5.9) ⚠️ ❌
Native Python SDK ✅ ✅ ✅ ✅ ✅
gRPC ✅ (PR #36169) ❌ ⚠️ ❌ ❌
Responses API streaming ✅ ⚠️ ✅ ⚠️ ⚠️
Docker images (official) ✅ vllm/vllm-openai ✅ ghcr.io/ggml-org/llama.cpp ✅ lmsysorg/sglang ❌ No official Docker image ❌ No official Docker image
Pluggable backends ✅ ✅ ✅ kt-kernel pluggable ✅ kt-kernel as backend ✅ + runtime quant repacking

Note on ktransformers Docker absence: this is a real friction point for our compose-based stack. The pip install path works in a conda env but breaks the "every service is a compose dir" convention.


When to pick which engine

Decision tree for a new model deployment on club-3090's hardware class:

Q1: Does the model fit your VRAM at desired quant?
├── YES, fits comfortably → 
│      Q2: Need max throughput + multi-tenant?
│      ├── YES → vLLM (production daily-driver)
│      └── NO → 
│             Q3: Want bulletproof / no-cliffs?
│             ├── YES → llama.cpp (single-card)
│             └── NO → vLLM still
│
└── NO, doesn't fit → 
       Q4: Is it MoE?
       ├── NO (dense doesn't fit) → No good options on consumer rig.
       │      Lower quant, smaller model, or rent cloud.
       └── YES → 
              Q5: Big MoE (>2× VRAM, < 96 GB RAM)?
              ├── YES → ktransformers (or SGLang+kt-kernel)
              └── NO (just slightly over) → llama.cpp `--n-cpu-moe`
                  OR ik_llama.cpp (better DeepSeek/Kimi MoE kernels)

Specific picks for club-3090's shipped models

Model Daily driver Reason
Qwen3.6-27B (dense hybrid, fits VRAM) vLLM + Genesis patches Multi-tenant, full feature set, Cliff 1/2 closed on TP=2
Qwen3.6-27B (single-card no-cliffs path) llama.cpp Different memory model; no Cliff 2b under multi-turn
Qwen3.6-27B (single-card with MTP, no PR-branch building) ik_llama.cpp MTP merged on main — get the ~+34% TPS lift without rebuilding from PR #22673
Gemma 4 31B vLLM + MTP/DFlash overlays Best spec-decode story, vision support
Carnice / Qwopus / variants vLLM Same daily-driver path
(future) MiniMax-M2.7-REAP-172B ktransformers + SGLang Big-MoE > VRAM, router-aware caching is the unlock
(future) GPT-OSS-120B ktransformers or llama.cpp --n-cpu-moe Either works; ktransformers ~+20% TPS but harder to deploy
(future) DeepSeek-V4-Flash ktransformers (kt-kernel native MXFP4) Best support for V4-Flash architecture

Honest gaps / what each isn't great at

vLLM

  • Single-card cliffs on Qwen3-Next: Cliff 2 / Cliff 2b on long-ctx single-card (24 GB) — see docs/CLIFFS.md. Mitigated by Genesis but not fully closed.
  • Layer-uniform CPU offload only (no router-aware MoE caching).
  • Patch-heavy for Qwen3-Next family — needs Genesis-vllm-patches for production-grade behavior.

llama.cpp

  • Single-stream-only on dense models (parallel slots exist but TP/EP are limited).
  • No FP8 weight quants — stuck with GGUF Q* / K* / IQ*.
  • MTP only via community PR (#22673) — not yet merged after months.
  • Layer-uniform --n-cpu-moe — no router-aware caching upstream yet (feature request #20757).

SGLang

  • More fragile than vLLM on production — fewer cross-rig anchor data points in the wild.
  • Marlin pad-sub-tile-n bug mirrors vLLM's PR #40361 — blocks AutoRound INT4 + EAGLE on TP=2 sub-tile-N shards (same kernel-line fix applies).
  • Less mature on Apple Silicon despite MLX backend.

ktransformers

  • No Docker image — breaks compose-based deployment patterns.
  • Smaller community — issues/PRs move slower than vLLM/llama.cpp.
  • NVIDIA-only — no AMD/Apple/Vulkan paths.
  • Yaml-based layer-placement config per model — non-trivial first-time setup.
  • Targets 200 GB RAM + Sapphire Rapids CPU for headline numbers; consumer rigs land in degraded-mode (AVX2 backend, REAP'd quants).

ik_llama.cpp

  • Smaller community than mainline llama.cpp — bugs take longer to surface, fewer cross-rig data points.
  • No tagged releases — rolls on main; no version pinning story for production users.
  • No official Docker image — same friction as ktransformers (would need to containerize ourselves).
  • Diverging quant naming from mainline — IQ_K series flags differ; cross-engine GGUF compatibility caveats.
  • Spec-decode coverage narrow — MTP merged but no EAGLE3 / DFlash; lags mainline on those research paths.
  • Smaller maintainer surface — primarily Iwan Kawrakow + a handful of contributors. Not the right pick for production where you need >1 person to debug a kernel issue.

Cross-engine bug / feature parity tracker

Issue vLLM llama.cpp SGLang ktransformers ik_llama.cpp
Marlin pad-sub-tile-n (output-dim shards <64 on TP=2 W4A16) 🟡 PR #40361 open N/A 🟡 Same bug, same fix applies N/A N/A
DeltaNet rollback (blocks EAGLE/DFlash/draft on Qwen3-Next) 🔴 #39931 🔴 🔴 🔴 🔴
MTP for Gemma 4 🟢 PR #41745 merged ❌ ✅ Day-0 ⚠️ ❌
DFlash for Gemma 4 🟡 PR #41703 Codex-rebased ⚠️ Luce fork ✅ Day-0 ROCm + CUDA ❌ ⚠️ Luce fork
qwen3coder <tool_call>-in-prose silent SSE ⚫ Local sidecar (issue #72) N/A ⚫ Same root cause N/A N/A
Per-token-head KV on hybrid pages 🟡 PR #40391 N/A 🟡 N/A N/A
Workspace lock strictness (vLLM 0.20) 🟢 PR #39226 merged N/A N/A N/A N/A
TurboQuant CPU + CUDA 🟢 ✅ via Genesis 🟡 PR #21089 (CPU first; CUDA no PR) 🟡 ⚠️ Same kernel bug as PR #40361 ❌ 🟡 Issue #1509 — CPU complete + CUDA written, awaiting validation/merge
MTP merged-not-PR-branch on llama.cpp family N/A 🟡 PR #22673 open 3d N/A N/A 🟢 ✅ Merged on main (Qwen + GLM-4.x)

🟢 Landed / 🟡 Open / 🔴 Blocked / ⚫ Local workaround


Sources


Maintenance note

This file ages quickly — all five projects ship frequently (some multiple releases per month, ik_llama.cpp rolls on main). Refresh quarterly (or whenever a major release substantively changes the comparison shape). When updating:

  1. Re-fetch each project's latest release notes
  2. Update version + date row at top
  3. Update any 🟡 → 🟢 transitions in the bug-parity tracker
  4. Add new feature rows where a project ships something the others don't
  5. Don't preserve historical "this used to be ❌, now ✅" — that lives in each project's release notes, not here

Last refresh: 2026-05-07.