Commit Graph
13 Commits
Author SHA1 Message Date
a49944162b chore(35b-a3b): drop the built-in-MTP A/B compose (vllm/qwen-a3b-preview-mtp) (#262)
fp8-mtp.yml was only ever an A/B reference -- MTP is net-negative on this MoE
at vLLM TP=2 (n2 ~= -45% / n3 ~= -51% wall TPS; the built-in-MTP draft forward
pays inter-GPU sync the acceptance can't amortize). It served no production
purpose, so remove it and its wiring.

Removed:
- models/qwen3.6-35b-a3b/vllm/compose/dual/autoround-int4/fp8-mtp.yml
- compose_registry.py entry `vllm/qwen-a3b-preview-mtp` (port 8052)
- its profile_runtime.yml capture block + calibration anchor (-> 17 anchors)
- the kv-calc COMPOSE_ALIAS_TEXT token

Test/doc updates for the new counts:
- test-kv-generic-dense + test-submit-pull: calibration contract 18/18 -> 17/17
- test-compose-registry-disk: registry 50->49, disk 51->50 composes
- test-report-calib: container ref -> live vllm-qwen36-35b-a3b-dual
- BENCHMARKS.md + KV_MATH.md notes

The MTP-net-negative finding stays recorded in the BENCHMARKS production row
and learnings/qwen3.6-35b-a3b.md. Full tests/ suite green (sole red is the
pre-existing fixture-dependent test-submit-bench, identical on clean master).

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-05-30 17:52:09 +05:00
noonghunnaandClaude Opus 4.7 eaa7a8c1ad docs+scripts: finish <quant>/ path migration across full repo sweep
Repo-wide follow-up to the compose quant-layer move (9821c94). The
mechanical move + registry/launch/test rewire covered the launchable
surface; a full-tree sweep found compose-path refs the move invalidated
in docs, two functional scripts, and one half-migrated mapping.

Functional fixes:
- bench-row-formatter.sh infer_compose_path(): 6 dual entries
  (int8-tq3, tq3-mtp-genesis, tq3-nomtp, tq3-mtp, int8, bf16) were left
  at bare dual/<file>.yml while the rest were migrated -> would emit
  dead compose paths into BENCHMARKS rows.
- residency-instrument/run-instrumented-soak.sh: case->COMPOSE_FILE
  paths (long-text, long-text-no-mtp, tools-text, dual default) now
  resolve under <quant>/.

Docs: README layout line + tree, engine/model READMEs, patch-README
quick-recipes, diagnostics, FAQ/CLIFFS/KV_MATH/DTYPE/MULTI/SINGLE/
STRUCTURED_COT/TQ3/UPSTREAM, issue template, sglang cross-refs
(-> vllm prod path). Per-model targets: qwen-vllm->autoround-int4,
llama-cpp->unsloth-q4km, gemma defaults (bf16-mtp/fp8-mtp),
carnice->own slug dir.

Intentionally left as historical/append-only records: CHANGELOG x2,
BENCHMARKS row-labels (live paths already correct in row bodies),
calibration source: provenance citations, switch.sh/parity history
comments. Separate follow-ups: gpu-mode.sh (#417 deprecated-repo
repoint), bench-row-formatter compose_display() docker-compose.yml
branch (PR-B). Flagged pre-existing-stale: dual/int8-tq3.yml in
pr40798/pr40914 READMEs (predate this refactor; ambiguous target).

Guard tests (registry-disk, mounts-resolve, switch-parity,
launch-compat) all PASS post-edit. Leak-clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-26 16:50:13 +00:00
noonghunnaandClaude Opus 4.7 9821c94efb refactor(compose): insert <quant> layer + make the registry the single source of truth
Restructure all 47 composes to
models/<model>/<engine>/compose/<topology>/<quant-slug>/<serving>.yml and centralize
defaults as registry pointers. Move + path-rewire only — no compose runtime-config
changes (plus the +1 ../ depth bump each moved file requires, and descriptive names
for the former docker-compose.yml defaults).

- <quant-slug> = 1:1 with a weights artifact == compose_registry weights_variant ==
  weights.py key (provider-quant form). Fixes the prior coarse `gguf` and the
  bf16/int8-files-mislabeled-as-autoround_int4 weights_variant.
- default = DEFAULTS[(model,engine,topology)] pointer in compose_registry.py; no
  default.yml / docker-compose.yml. KV-dtype variants (bf16/int8) stay as files
  under autoround-int4/.
- +1 ../ on every relative mount in moved composes (cache/patches/scripts/models-cache).
- Rewired: compose_registry (47 compose_path + weights_variant + DEFAULTS), launch.sh,
  gpu-mode.sh, bench-row-formatter, profile *.yml, generate_compose.py, weights.py
  (+ back-compat ALIASES), AGENTS.md, ADDING_MODELS.md, BENCHMARKS + doc path refs.
- New guard tests: test-compose-registry-disk.sh, test-compose-mounts-resolve.sh.
- Variant keys (vllm/dual, ik-llama/iq4ks-mtp, ...) unchanged — launch/switch UX preserved.

Validation: docker compose config 47/47 · parity/launch-compat/registry-disk/
mounts-resolve/profiles-compat/model-weights-registry PASS · kv-calc 22/22 · leak-clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-26 15:58:08 +00:00
5b3793b85e feat(kv-calc): opt-in --kv-breakdown architecture cache planning layer (#213)
Additive + reporting-only. Default text and --json output are byte-identical
unless --kv-breakdown is passed; predict()/raw_verdict()/--calibration remain
the calibrated fit authority. Adds --gpus N (informational), a conflict-guarded
--sequences alias, and opt-in compressed-KV / indexer-cache / draft-KV estimate
buckets scoped to single-node 1-8 GPU home/workstation planning.

Implemented via Codex collab; independently validated on-rig by Claude:
- default output IDENTICAL (HEAD vs change) across text / --json / --solve-max-ctx
- --calibration 22/22, both kv-calc test scripts pass, py_compile clean, diff --check clean
- 4 error cases exit 2 (sequences/max-num-seqs conflict, --gpus bounds,
  indexer/compressed missing head-dim)

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-24 08:57:05 +05:00
noonghunnaandClaude Opus 4.7 b1c68b4fe3 docs(kv-math): extend k_v_tensors=N notation to sliding-KV formulas
Earlier passes harmonized the 4 GROWING-KV formulas to the
'× k_v_tensors=N × bpe' form, but the 2 SLIDING-KV formulas (Gemma 4
31B + Gemma 4 26B-A4B sliding-window portions) still used the old
'× 1 (K=V tied)' parenthetical.

All 6 formulas (4 growing + 2 sliding) now use consistent notation.

(Note: Grok's third-pass review claimed a leftover at Qwen 27B line
~211 — that's actually a Python comment in the config.json extraction
section, not a formula. Grok was hallucinating against a cached older
state. But the sliding-KV inconsistency was real and worth fixing.)

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 11:01:55 +00:00
noonghunnaandClaude Opus 4.7 54d8bf0b3e docs(kv-math): tighten k_v_tensors notation across all 4 formulas
Per Grok third-pass nit: parenthetical (N, K=V tie note) was redundant
with the architecture summary above each section. Switched to compact
'k_v_tensors=N' form across all 4 formulas:

- Qwen 27B: × k_v_tensors=2
- Qwen 35B-A3B: × k_v_tensors=2
- Gemma 31B: × k_v_tensors=1
- Gemma 26B-A4B: × k_v_tensors=1

K=V tie context still present in each section's architecture summary;
no information lost.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:56:49 +00:00
noonghunnaandClaude Opus 4.7 67de3eca3e docs(kv-math): third-pass Grok polish
1. Quick reference table — new "Vs Qwen 27B" column anchored at fp8_e5m2
   TP=2 (common production config). Makes the family hierarchy
   self-evident in the table without needing to read the commentary:
   - Qwen 3.6 27B: 1.00× (baseline)
   - Qwen 3.6 35B-A3B: 0.31× (~3.2× lighter)
   - Gemma 4 31B: 2.50× (~2.5× heavier)
   - Gemma 4 26B-A4B: 0.16× (~6.4× lighter)
   Footnote explains the anchor choice + that ratios shift slightly
   under other formats but family hierarchy is stable.

2. DeltaNet general formula — added worked example for Qwen 3.6 35B-A3B
   showing plug-and-chug of the formula with real config.json values
   (30 layers, 16 k_heads × 128, 32 v_heads × 128, conv_kernel=4, fp32).
   Result: ~3.5 MB at max_num_seqs=1, ~14 MB at seqs=4. Makes the formula
   actionable for future DeltaNet onboardings.

Skipped Grok #1 (Qwen 27B notation parenthetical tightening) — Grok
explicitly says "already good" and didn't suggest a concrete form.
Aesthetic-perfection rabbit hole; not worth a further commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:55:56 +00:00
noonghunnaandClaude Opus 4.7 78f94cacf0 docs(kv-math): second-pass Grok polish
1. Formula notation — flip from '× 2 (k_v_tensors — no K=V tie) × bpe'
   to '× k_v_tensors (2, no K=V tie) × bpe'. Variable name leads the
   factor; value is parenthetical. Applied across all 4 model formulas
   for consistency with the general formula's reading order.

2. Quick reference table — new section after the per-card budget
   composition. At-a-glance per-token growing-KV bytes for all 4 models
   across bf16/fp8/TQ3-or-INT8 at TP=1/2. Includes 'what jumps out'
   commentary calling out:
   - Gemma 26B-A4B ~16x smaller per token vs 31B
   - Qwen 35B-A3B ~3.2x smaller per token vs 27B
   - Sliding-window KV is fixed (not per-token) for Gemma family
   - TQ3 vs INT8 PTH family applicability

3. DeltaNet general formula — new subsection inside the general KV
   formula block. Single-line formula covering K state + V state +
   conv1d kernel state across num_gdn_layers × max_num_seqs streams,
   using the linear_* config fields. Future DeltaNet model onboardings
   can compute their state size from this formula directly without
   re-deriving from scratch.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:53:38 +00:00
noonghunnaandClaude Opus 4.7 3114399983 docs(kv-math): address Grok review feedback
1. Formula consistency — all per-token formulas now annotate the k_v_tensors
   factor explicitly:
   - Qwen 27B: '× 2 (k_v_tensors — no K=V tie)'
   - Qwen 35B-A3B: '× 2 (k_v_tensors — no K=V tie)'
   - Gemma 31B: '× 1 (k_v_tensors — K=V tied)'
   - Gemma 26B-A4B: '× 1 (k_v_tensors — K=V tied)'
   Ties the per-model math back to the general formula's named variable.

2. MoE column format — summary table now uses compact 'Yes (256×8)' /
   'Yes (128×8)' notation (experts × active-per-token). Note added
   explaining the convention.

3. Qwen 35B-A3B router workspace math — corrected for real 256 experts
   (was using estimate-era 128). Now shows ~1 MB per router × 40 layers
   ≈ 40 MB total.

4. DeltaNet recurrent state — new dedicated subsection in both Qwen 3.6
   27B and Qwen 3.6 35B-A3B sections. Concrete sizes (~7.5 MB for 27B,
   ~3.5 MB for 35B-A3B per stream at seqs=1) computed from
   linear_num_*_heads + linear_*_head_dim + linear_conv_kernel_dim.
   Distinguishes the recurrent state (small, persistent) from the
   block-wise intermediate during forward (large, captured in activation
   peak §). Grok suggested ~50-150 MB total but that overstates by ~10×;
   I used computed values from the real config.json fields instead.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:50:22 +00:00
noonghunnaandClaude Opus 4.7 6ec6a6761f docs(kv-math): config-verify Qwen 35B-A3B + Gemma 26B-A4B MoE sections
Both MoE model sections rewritten from "math-ready estimates" to
"config-verified, calibration pending". Architectural facts now sourced
from on-disk config.json + layer_types arrays rather than guesses.

Qwen 3.6 35B-A3B corrections (real values from
/mnt/models/huggingface/qwen3.6-35b-a3b-autoround-int4/config.json):
- num_experts: 256 (was estimated 128 — 2x correction)
- Confirmed full_attention indices [3,7,11,15,19,23,27,31,35,39]
  via full_attention_interval=4 and layer_types array
- Built-in MTP (mtp_num_hidden_layers=1) confirmed
- attn_output_gate=True (gated attention) confirmed
- 5 quant variants on disk (autoround_int4 production, plus gptq_int4,
  gguf, dflash, dflash-gguf)

Gemma 4 26B-A4B major rewrite (real values from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json):
- 30 total layers (was estimated 40-48 — smaller than expected)
- 5 full_attention layers at indices [5,11,17,23,29]
- 25 sliding_attention layers
- *** ASYMMETRIC KV HEADS *** (the architectural surprise):
  - num_key_value_heads: 8 for sliding layers
  - num_global_key_value_heads: 2 for global layers (distinct field)
- Per-token growing KV at bf16 = 5 × 2 × 512 × 1 = 5,120 bytes
  vs Gemma 4 31B's 81,920 bytes — *** ~16x smaller per token ***
- 128 experts (was estimated 64), 8 active per token
- moe_intermediate_size: 704
- Multimodal: vision + audio token IDs
- requires_genesis: false (Gemma 4 has no DeltaNet quirks)
- Per-card budget projection: ~10-12 GB at 200K + fp8 KV →
  massive headroom on 24 GB; single-card serving may be viable at 262K

Top intro table + architecture summary table both updated with
verified values; status changed from "Math-ready" to "Config-verified".

Calibration remains pending: empirical activation-peak coefficients
need ≥4 measured BENCHMARKS rows per model before locking in.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 10:42:43 +00:00
noonghunnaandClaude Opus 4.7 1f8aaa2acc docs: expand KV_MATH + add ADDING_MODELS workflow
- KV_MATH.md: add general formula, 4-model architecture summary table,
  HF config.json extraction guidance, full math derivations for Qwen3.6
  35B-A3B + Gemma4 26B-A4B MoE variants (calibration-pending), best
  practices for KV calculator implementations, and sources-of-error
  decomposition.

- ADDING_MODELS.md (new): end-to-end workflow for onboarding a new model
  into the v0.7.0 profile catalog — architecture sourcing, ModelProfile
  authoring, compose env-hook requirements, COMPOSE_REGISTRY entries,
  calibration capture, validation chain, common pitfalls, worked example
  for the Qwen3.6 35B-A3B MoE case.

Pairs with the v0.7.0 profile data model: KV_MATH stays math-reference,
ADDING_MODELS owns the workflow.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-14 20:10:22 +00:00
noonghunnaandClaude Opus 4.7 0d48dac818 feat(tools): extend kv-calc.py to multi-model (Qwen 3.6 + Gemma 4 31B)
- Refactor MODEL_SPECS / COMPOSES / CALIBRATION into per-model dicts
- Add --model flag (defaults qwen3.6-27b for back-compat; infers from --compose)
- Add Gemma 4 31B architecture (10 full + 50 sliding layers, K==V tying,
  global_head_dim=512 asymmetry)
- KV pool capping models vLLM's PagedAttention allocator: predicted KV is
  min(requested, available_budget); verdict becomes TIGHT when capped
  (resolves the original limitation #1 — max_num_seqs>1 over-prediction)
- 8 Gemma composes registered; calibration accuracy 18/18 across both models
  (was 9/11 Qwen-only)
- Update docs/KV_MATH.md: per-model sections, K==V tying empirical finding,
  multi-model calibration table, limitation #1 marked resolved

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-13 23:15:13 +00:00
noonghunnaandClaude Opus 4.7 4e89c6aa40 feat(tools): kv-calc.py — predict per-card VRAM budget for Qwen3.6-27B (#226)
A directional estimator for "will this compose fit on my hardware?" that
runs before boot. Useful for:
- Comparative analysis (TQ3 vs fp8 KV trade-off, TP=1 vs TP=2 budget)
- Max-ctx solver (binary search the largest max_ctx that fits given
  vram/tp/kv_format/max_num_seqs/mem_util)
- Educational breakdown (weights / KV pool / activation peak / cudagraph
  workspace per-card components)

Anchored to published literature where applicable:
- PerfMamba (arxiv 2511.22849) — O(γ·D·N·L) scaling for the GDN forward
  intermediate (this is our Cliff 2 mechanism)
- TurboQuant (arxiv 2504.19874, ICLR 2026) — TQ3 byte savings
- PagedAttention (arxiv 2309.06180) — KV pool layout

Calibrated empirically against BENCHMARKS.md cross-rig data. Current
verdict accuracy: 9/11 shipped composes (82%) with ±1.5 GB error band.
The two miscalibrations are over-predictions on max_num_seqs > 1
configs; vLLM's actual KV pool allocation rate-limits internally in
ways my naive demand formula doesn't capture. Documented as a known
limitation in docs/KV_MATH.md.

Usage examples:
  bash tools/kv-calc.py --compose dual-turbo --vram 20 --mem-util 0.82
  bash tools/kv-calc.py --solve-max-ctx --tp 2 --kv-format fp8_e5m2 --vram 16
  bash tools/kv-calc.py --calibration

Documents the math at docs/KV_MATH.md including:
- Per-component formulas (weights / KV pool / activation peak / overhead)
- Per-KV-format byte tables
- The TQ3→fp8 swap rule on 20 GB cards (validated by @efschu, #47)
- Known limitations + when to trust kv-calc.py vs vLLM's boot log

Closes #226 (the v1 calculator). Future work that's NOT in this commit:
multi-model support (DeepSeek-V3, GLM-4.5 — requires per-architecture
spec tables), driver-class overhead modeling (4090 vs 3090 deltas
beyond what we currently fold into the ±1.5 GB error band), Cliff 2b
fragmentation modeling.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-04 12:37:04 +00:00