fp8-mtp.yml was only ever an A/B reference -- MTP is net-negative on this MoE
at vLLM TP=2 (n2 ~= -45% / n3 ~= -51% wall TPS; the built-in-MTP draft forward
pays inter-GPU sync the acceptance can't amortize). It served no production
purpose, so remove it and its wiring.
Removed:
- models/qwen3.6-35b-a3b/vllm/compose/dual/autoround-int4/fp8-mtp.yml
- compose_registry.py entry `vllm/qwen-a3b-preview-mtp` (port 8052)
- its profile_runtime.yml capture block + calibration anchor (-> 17 anchors)
- the kv-calc COMPOSE_ALIAS_TEXT token
Test/doc updates for the new counts:
- test-kv-generic-dense + test-submit-pull: calibration contract 18/18 -> 17/17
- test-compose-registry-disk: registry 50->49, disk 51->50 composes
- test-report-calib: container ref -> live vllm-qwen36-35b-a3b-dual
- BENCHMARKS.md + KV_MATH.md notes
The MTP-net-negative finding stays recorded in the BENCHMARKS production row
and learnings/qwen3.6-35b-a3b.md. Full tests/ suite green (sole red is the
pre-existing fixture-dependent test-submit-bench, identical on clean master).
Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Earlier passes harmonized the 4 GROWING-KV formulas to the
'× k_v_tensors=N × bpe' form, but the 2 SLIDING-KV formulas (Gemma 4
31B + Gemma 4 26B-A4B sliding-window portions) still used the old
'× 1 (K=V tied)' parenthetical.
All 6 formulas (4 growing + 2 sliding) now use consistent notation.
(Note: Grok's third-pass review claimed a leftover at Qwen 27B line
~211 — that's actually a Python comment in the config.json extraction
section, not a formula. Grok was hallucinating against a cached older
state. But the sliding-KV inconsistency was real and worth fixing.)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Per Grok third-pass nit: parenthetical (N, K=V tie note) was redundant
with the architecture summary above each section. Switched to compact
'k_v_tensors=N' form across all 4 formulas:
- Qwen 27B: × k_v_tensors=2
- Qwen 35B-A3B: × k_v_tensors=2
- Gemma 31B: × k_v_tensors=1
- Gemma 26B-A4B: × k_v_tensors=1
K=V tie context still present in each section's architecture summary;
no information lost.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
1. Quick reference table — new "Vs Qwen 27B" column anchored at fp8_e5m2
TP=2 (common production config). Makes the family hierarchy
self-evident in the table without needing to read the commentary:
- Qwen 3.6 27B: 1.00× (baseline)
- Qwen 3.6 35B-A3B: 0.31× (~3.2× lighter)
- Gemma 4 31B: 2.50× (~2.5× heavier)
- Gemma 4 26B-A4B: 0.16× (~6.4× lighter)
Footnote explains the anchor choice + that ratios shift slightly
under other formats but family hierarchy is stable.
2. DeltaNet general formula — added worked example for Qwen 3.6 35B-A3B
showing plug-and-chug of the formula with real config.json values
(30 layers, 16 k_heads × 128, 32 v_heads × 128, conv_kernel=4, fp32).
Result: ~3.5 MB at max_num_seqs=1, ~14 MB at seqs=4. Makes the formula
actionable for future DeltaNet onboardings.
Skipped Grok #1 (Qwen 27B notation parenthetical tightening) — Grok
explicitly says "already good" and didn't suggest a concrete form.
Aesthetic-perfection rabbit hole; not worth a further commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
1. Formula notation — flip from '× 2 (k_v_tensors — no K=V tie) × bpe'
to '× k_v_tensors (2, no K=V tie) × bpe'. Variable name leads the
factor; value is parenthetical. Applied across all 4 model formulas
for consistency with the general formula's reading order.
2. Quick reference table — new section after the per-card budget
composition. At-a-glance per-token growing-KV bytes for all 4 models
across bf16/fp8/TQ3-or-INT8 at TP=1/2. Includes 'what jumps out'
commentary calling out:
- Gemma 26B-A4B ~16x smaller per token vs 31B
- Qwen 35B-A3B ~3.2x smaller per token vs 27B
- Sliding-window KV is fixed (not per-token) for Gemma family
- TQ3 vs INT8 PTH family applicability
3. DeltaNet general formula — new subsection inside the general KV
formula block. Single-line formula covering K state + V state +
conv1d kernel state across num_gdn_layers × max_num_seqs streams,
using the linear_* config fields. Future DeltaNet model onboardings
can compute their state size from this formula directly without
re-deriving from scratch.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
1. Formula consistency — all per-token formulas now annotate the k_v_tensors
factor explicitly:
- Qwen 27B: '× 2 (k_v_tensors — no K=V tie)'
- Qwen 35B-A3B: '× 2 (k_v_tensors — no K=V tie)'
- Gemma 31B: '× 1 (k_v_tensors — K=V tied)'
- Gemma 26B-A4B: '× 1 (k_v_tensors — K=V tied)'
Ties the per-model math back to the general formula's named variable.
2. MoE column format — summary table now uses compact 'Yes (256×8)' /
'Yes (128×8)' notation (experts × active-per-token). Note added
explaining the convention.
3. Qwen 35B-A3B router workspace math — corrected for real 256 experts
(was using estimate-era 128). Now shows ~1 MB per router × 40 layers
≈ 40 MB total.
4. DeltaNet recurrent state — new dedicated subsection in both Qwen 3.6
27B and Qwen 3.6 35B-A3B sections. Concrete sizes (~7.5 MB for 27B,
~3.5 MB for 35B-A3B per stream at seqs=1) computed from
linear_num_*_heads + linear_*_head_dim + linear_conv_kernel_dim.
Distinguishes the recurrent state (small, persistent) from the
block-wise intermediate during forward (large, captured in activation
peak §). Grok suggested ~50-150 MB total but that overstates by ~10×;
I used computed values from the real config.json fields instead.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
- KV_MATH.md: add general formula, 4-model architecture summary table,
HF config.json extraction guidance, full math derivations for Qwen3.6
35B-A3B + Gemma4 26B-A4B MoE variants (calibration-pending), best
practices for KV calculator implementations, and sources-of-error
decomposition.
- ADDING_MODELS.md (new): end-to-end workflow for onboarding a new model
into the v0.7.0 profile catalog — architecture sourcing, ModelProfile
authoring, compose env-hook requirements, COMPOSE_REGISTRY entries,
calibration capture, validation chain, common pitfalls, worked example
for the Qwen3.6 35B-A3B MoE case.
Pairs with the v0.7.0 profile data model: KV_MATH stays math-reference,
ADDING_MODELS owns the workflow.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
A directional estimator for "will this compose fit on my hardware?" that
runs before boot. Useful for:
- Comparative analysis (TQ3 vs fp8 KV trade-off, TP=1 vs TP=2 budget)
- Max-ctx solver (binary search the largest max_ctx that fits given
vram/tp/kv_format/max_num_seqs/mem_util)
- Educational breakdown (weights / KV pool / activation peak / cudagraph
workspace per-card components)
Anchored to published literature where applicable:
- PerfMamba (arxiv 2511.22849) — O(γ·D·N·L) scaling for the GDN forward
intermediate (this is our Cliff 2 mechanism)
- TurboQuant (arxiv 2504.19874, ICLR 2026) — TQ3 byte savings
- PagedAttention (arxiv 2309.06180) — KV pool layout
Calibrated empirically against BENCHMARKS.md cross-rig data. Current
verdict accuracy: 9/11 shipped composes (82%) with ±1.5 GB error band.
The two miscalibrations are over-predictions on max_num_seqs > 1
configs; vLLM's actual KV pool allocation rate-limits internally in
ways my naive demand formula doesn't capture. Documented as a known
limitation in docs/KV_MATH.md.
Usage examples:
bash tools/kv-calc.py --compose dual-turbo --vram 20 --mem-util 0.82
bash tools/kv-calc.py --solve-max-ctx --tp 2 --kv-format fp8_e5m2 --vram 16
bash tools/kv-calc.py --calibration
Documents the math at docs/KV_MATH.md including:
- Per-component formulas (weights / KV pool / activation peak / overhead)
- Per-KV-format byte tables
- The TQ3→fp8 swap rule on 20 GB cards (validated by @efschu, #47)
- Known limitations + when to trust kv-calc.py vs vLLM's boot log
Closes#226 (the v1 calculator). Future work that's NOT in this commit:
multi-model support (DeepSeek-V3, GLM-4.5 — requires per-architecture
spec tables), driver-class overhead modeling (4090 vs 3090 deltas
beyond what we currently fold into the ±1.5 GB error band), Cliff 2b
fragmentation modeling.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>