Capture the quant/KV findings so others don't re-hit them:
DTYPE_MATRIX.md — new "KV-quant × checkpoint compatibility — the two Ampere traps":
(1) fp8 KV is rejected for compressed-tensors checkpoints (AWQ/FP8/INT8 weights),
flag-independent; auto_round/GPTQ unaffected. (2) int8_per_token_head is NATIVE
in stock v0.22.0 for uniform-head-dim models — #40391 is the Gemma-4-only
(interleaved 256/512 head-dim → page-size unification) adapter, don't copy it.
+ an Ampere KV-dtype picker table; tag the INT8-PTH row native/overlay status.
QUANTIZATION.md — new §4a "Picking a quant by fidelity (KLD) — and where QAT fits":
Phaelon74 KLD ranking (INT8 0.009 < FP8 0.023 < AWQ-BF16-INT4 0.042 < AWQ-INT4
0.051 < AutoRound 0.063); KLD is weights-only (KV-quant adds separate error);
QAT only out-earns PTQ at <=4-bit (8-bit PTQ already near-lossless); dual=fidelity /
single=fit tiering. + int8_per_token_head row + the fp8-guard caveat in §5.
FIX the stale §4 FP8 line: FP8 *weights* DO run on Ampere via Marlin W8A16 (not
"emulated / KV-only") and are a top-fidelity option.
FAQ.md — new Q: "My AWQ/FP8 model errors on --kv-cache-dtype fp8" → use int8-PTH.
Gate 42/42.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
12 KiB
Quantization — a field guide for the club-3090 community
Quantization is how a 27B model that would need ~54 GB at FP16 fits in 24 GB of VRAM. This page explains the quant families you'll see in the wild, what actually differs between them, and which ones this stack ships — including the IQK imatrix quants that exist only in ik_llama.cpp.
The one idea to take away: at the same bits-per-weight, not all quants are equal. The two levers that separate good from bad are (1) non-linear levels that match the weight distribution and (2) calibration (an "importance matrix") that spends bits where they matter. The best quants use both.
See also: GLOSSARY.md · DTYPE_MATRIX.md (KV/compute dtypes) · engines/IK_LLAMA.md.
1. The vocabulary
- bpw (bits per weight): the headline number. FP16 = 16 bpw. A "4-bit" quant is ~4-4.5 bpw once you count the per-block scale/zero-point overhead. Lower bpw = smaller file = more context room, but more quality risk.
- Block / group: quants don't store one scale for the whole tensor — they chunk weights into blocks (e.g. 32 weights) and store a scale per block. Smaller blocks = finer = more accurate, but more overhead.
- imatrix (importance matrix): a calibration pass over real text that records which weights matter most for the model's outputs, so the quantizer protects those and compresses the rest harder. "i-quant" / "IQ" prefixes signal imatrix use.
- Weight quant vs KV-cache quant: two independent knobs. One shrinks the model; the other shrinks the context (see §5). You pick both.
2. The GGUF ladder (llama.cpp + ik_llama.cpp)
GGUF is the llama.cpp-family weight format. Roughly in order of quality-per-bit (worst → best at a given bpw):
| Family | Examples | Calibrated? | Where | Notes |
|---|---|---|---|---|
| Legacy | Q4_0, Q4_1, Q5_0, Q8_0 |
❌ | mainline | Simple round-to-nearest, one scale/block. Q8_0 is still a great near-lossless choice; the low-bit legacy ones are superseded. |
| K-quants | Q3_K_M, Q4_K_M, Q5_K_M, Q6_K |
❌ (data-free) | mainline | Mixed precision per tensor-type + 2-level block scales. The mainstream default. Q4_K_M is what our shipped llamacpp/mtp runs. Good, but data-free — no calibration. |
| i-quants | IQ2_XXS … IQ3_M, IQ4_XS |
✅ imatrix | mainline | Non-linear lattice codebooks + importance matrix. Clearly better quality-per-bit than k-quants, especially below 4 bpw. Slightly slower dequant than k-quants. |
| IQK quants ⭐ | IQ4_KS, IQ5_KS, IQ4_K, IQ2_K … |
✅ imatrix | ik_llama.cpp only | Refined grids + imatrix + kernels co-designed for those grids. Best quality-per-bit in the GGUF world and fast (the dequant path is hand-tuned). Fork-exclusive. |
The progression that matters: Q4_K_M (data-free) → IQ4_XS (imatrix, mainline) → IQ4_KS (imatrix + co-designed kernels, ik fork). Each step is better quality at similar bpw. Our shipped llamacpp/mtp is at the first rung (Q4_K_M); the ik_llama track is at the last (IQ4_KS).
3. What "imatrix" actually buys you
A data-free quant treats every weight as equally important and rounds uniformly. But in a trained model, a small fraction of weights carry most of the signal. An importance matrix is computed by running calibration text through the model and measuring how much each weight influences activations. The quantizer then:
- protects high-importance weights (more bits / closer grid points), and
- compresses low-importance weights harder.
Result: at 4 bpw, an imatrix quant loses noticeably less quality than a data-free one — and the gap widens as you go lower (at 2-3 bpw, imatrix is the difference between usable and broken). The cost is a one-time calibration step when building the quant; inference is the same speed.
Calibration corpus matters. An imatrix calibrated on chat+code+tool-calling preserves those skills; one calibrated on Wikipedia may quietly drop tool-call formatting. This is exactly the kind of thing that shows up in our 8-pack quality tests (see QUALITY_TEST.md).
4. The vLLM / safetensors side (not GGUF)
vLLM and SGLang don't use GGUF — they load safetensors with these quant schemes:
| Quant | Bits | Calibrated? | Notes |
|---|---|---|---|
| AutoRound ⭐ | INT4 | ✅ (sign-gradient) | Intel's method; what our shipped vllm/dual runs (qwen3.6-27b-autoround-int4). Strong 4-bit quality. |
| AWQ | INT4 | ✅ (activation-aware) | Protects salient channels by inspecting activations. Widely available. |
| GPTQ | INT3/4/8 | ✅ (second-order) | Older, well-supported; AWQ/AutoRound usually edge it at 4-bit. |
| FP8 (e4m3/e5m2) | 8 | ❌ | FP8 weights DO run on Ampere — via Marlin W8A16 (weights dequant for the matmul: real VRAM saving, no compute speedup; only Hopper+ multiply FP8 directly). Among the highest-fidelity practical quants (see §4a). An FP8 weight checkpoint is compressed-tensors → it can't also use fp8 KV — pair it with int8-PTH (DTYPE_MATRIX). Don't confuse FP8 weights with the fp8 KV cache (§5). |
| bitsandbytes | 4/8 | ❌ | Easy/on-the-fly; lower quality-per-bit than AWQ/AutoRound. |
These are conceptually the same idea as imatrix i-quants (calibrate, protect what matters) in a different ecosystem. There is no GGUF↔safetensors interchange — a quant is tied to its engine family.
4a. Picking a quant by fidelity (KLD) — and where QAT fits
Method labels don't rank fidelity — measure it. The cleanest signal is KL-divergence vs the BF16 model (lower = closer to full precision). Phaelon74's logit-capture sweep on Qwen3.6-27B (one consistent instrument across quants):
| Quant | Mean KLD | ~Size | Ampere path |
|---|---|---|---|
| INT8 (W8A16) | 0.009 | 34 GiB | Marlin W8A16 |
| FP8 | 0.023 | 29 GiB | Marlin W8A16 |
| AWQ-BF16-INT4 | 0.042 | 27 GiB | Marlin int4 |
| AWQ-INT4 | 0.051 | 20 GiB | Marlin int4 |
AutoRound INT4 (our vllm/dual) |
0.063 | 18 GiB | Marlin int4 |
| NVFP4 variants | 0.06–0.19 | — | Blackwell-only — Ampere-dead |
- Fidelity rises with bits; 8-bit PTQ (FP8/INT8) is already near-lossless — hence "accuracy → FP8" as a default. INT4 trades fidelity for size/speed.
- Our shipped AutoRound INT4 is near the bottom of the practical pack — not because AutoRound is bad, but because it's the smallest (18 GiB). A method label ("calibrated"/"QAT-adjacent") does not guarantee best KLD; the bigger/higher-bit quants win. Always measure for the specific model.
- KLD here is weights-only (measured with full/BF16 KV). KV-quant (fp8 / int8-PTH) adds a separate error this number doesn't include → deployed fidelity = weight-KLD + KV-quant drift. The purest accuracy config is best-weights + BF16 KV; a 1-byte KV (to reach max context) spends a little of that back.
Where QAT fits: quantization-aware training fine-tunes with quantization simulated, so weights adapt → near-BF16 at low bits. Its advantage is concentrated at ≤4-bit — at 8-bit, PTQ is already near-lossless, so QAT adds little. So: want 4-bit + max fidelity → a QAT-int4 (e.g. Google's Gemma QAT) beats AutoRound/AWQ int4; want max fidelity, size-flexible → just use 8-bit PTQ (FP8/INT8), no QAT needed. PTQ methods that approach QAT without retraining: AutoRound (learns rounding), AWQ (activation-aware), GPTQ (2nd-order), QuIP#/AQLM (codebook, 2-bit), SpinQuant/QuaRot (rotation); the GGUF analogue is imatrix (§3).
Tiering principle: dual-card = fidelity tier, single-card = fit/speed tier. A dual's distinct value is capacity — spend it on a higher-fidelity quant (FP8/INT8/AWQ) than the single's small INT4, not the same one.
5. KV-cache quantization (a separate knob)
Independent of the weight quant, you can quantize the KV cache — this is what sets your max context, not your model quality:
| KV type | Bits | Engine | Notes |
|---|---|---|---|
f16 |
16 | all | Lossless, biggest. Rarely needed. |
q8_0 |
8 | llama.cpp / ik | Near-lossless; good default when context is moderate. |
q4_0 |
4 | llama.cpp / ik | Halves KV vs q8_0 → enables 262K on one 3090 (ik IQ4_KS). Tiny quality cost. |
fp8_e5m2 |
8 | vLLM | Our vllm/dual default (AutoRound weights). |
int8_per_token_head |
8 | vLLM | ~1 byte/tok like fp8; native in stock v0.22.0 for standard models (Gemma-4 needs the #40391 overlay). The KV path for compressed-tensors weights (AWQ/FP8/INT8) at long context — those can't use fp8 KV. |
| TQ3 (TurboQuant) | 3 | vLLM (Genesis) | 3-bit KV — beats fp8 on long-context memory; powers our dual-turbo. See TQ3_MTP_GENESIS.md + CLIFFS.md. |
-khad (modifier) |
— | ik only | Hadamard transform on the K-cache → recovers accuracy lost to KV quantization, so you keep quality at q4_0/q8_0. |
⚠️ fp8 KV is rejected for compressed-tensors checkpoints (AWQ / FP8 / INT8 weights):
--kv-cache-dtype fp8_e5m2→ValueError: … not supported with fp8 checkpoints, regardless of the--quantizationflag. Useint8_per_token_headthere —auto_round/GPTQ weights are unaffected (they take fp8 KV fine). Full picker + the Gemma-4 #40391 caveat: DTYPE_MATRIX.
6. Engine × quant support
| Quant family | vLLM | mainline llama.cpp | ik_llama.cpp | SGLang |
|---|---|---|---|---|
K-quants (Q4_K_M…) |
❌ | ✅ | ✅ | ❌ |
i-quants (IQ4_XS…) |
❌ | ✅ | ✅ | ❌ |
IQK (IQ4_KS…) |
❌ | ❌ | ✅ only | ❌ |
| AutoRound / AWQ / GPTQ | ✅ | ❌ | ❌ | ✅ |
| FP8 weights | ✅ | ❌ | ❌ | ✅ |
7. Why doesn't every GGUF repo ship IQK?
If IQK is the best quality-per-bit, why are most community GGUFs still Q4_K_M?
- It's fork-locked. IQK quants run only on ik_llama.cpp. A
Q4_K_Mruns on mainline llama.cpp, Ollama, LM Studio, LocalAI, Jan — everything. Quant authors optimize for reach. - Kernel co-design. IQK's quality comes partly from kernels written for its grids. Porting that to mainline isn't a small patch, and upstreaming has been slow.
- Inertia + tooling.
Q4_K_Mis the well-trodden default; build pipelines, docs, and "recommended download" buttons all point at it.
So IQK is a deliberate "I'll run the fork to get the better quant" choice — which is exactly the niche the ik_llama track fills on this stack.
8. What this stack ships (and why)
| Path | Quant (weights) | KV | Rationale |
|---|---|---|---|
vllm/dual |
AutoRound INT4 | fp8_e5m2 | Production dual-card; deepest Qwen3-Next feature support |
vllm/dual-turbo |
AutoRound INT4 | TQ3 | Max throughput + long context (3-bit KV) |
llamacpp/mtp |
Q4_K_M | q4_0 | Conservative, mainline image, cliff-immune single-card |
ik-llama/iq4ks-mtp ⭐ |
IQ4_KS (imatrix) | q4_0 + -khad |
Advanced-quant track: best quality-per-bit + 262K single-card |
Rule of thumb for your own rig:
- Tightest VRAM / lowest bpw → reach for an imatrix quant (
IQ4_XSmainline, orIQ4_KSon ik_llama), not a data-freeQ4_K_M. - Want maximum quality-per-bit and willing to run the fork → ik_llama + IQK.
- Multi-tenant / vision / tools at scale → vLLM + AutoRound.
- "Just works everywhere, no fork" → mainline llama.cpp + Q4_K_M.
See also
- engines/IK_LLAMA.md — the engine that unlocks IQK
- INFERENCE_ENGINES.md — engine comparison
- DTYPE_MATRIX.md — compute/KV dtype matrix
- CLIFFS.md + TQ3_MTP_GENESIS.md — KV-cache quant deep-dives
- BENCHMARKS.md — measured quality + TPS per quant/engine