Commit Graph
21 Commits
Author SHA1 Message Date
noonghunna 99328b4cda feat(estate): add parallel boot mode 2026-05-15 22:18:09 +05:00
noonghunna 127f4f6d8f test(launch): align engine pin expectations 2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 3b2d940d26 fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean
Per the TQ3-only Genesis policy (Genesis is strictly required only for
turboquant_3bit_nc KV — Cliff 2 mitigations are recommended but not
required to boot), the Qwen 27B composes that don't use TQ3 KV no longer
need to be anchored to the Genesis-locked SHA. They can ride the latest
unconstrained nightly (vllm-nightly-clean → bf610c2f).

Composes moved from vllm-nightly-mtp to vllm-nightly-clean (8 entries):
- vllm/tools-text         single/tools-text.yml      fp8_e5m2
- vllm/minimal            single/minimal.yml         fp8_e5m2
- vllm/dual               dual/docker-compose.yml    fp8_e5m2
- vllm/dual-bf16          dual/bf16.yml              bf16
- vllm/dual-carnice-bf16mtp dual/carnice-bf16mtp.yml fp8_e5m2
- vllm/dual-qwopus-bf16mtp  dual/qwopus-bf16mtp.yml  fp8_e5m2
- vllm/dual-nvlink        dual/nvlink.yml (extends)  fp8_e5m2
- vllm/dual4              multi4/docker-compose.yml  fp8_e5m2

TQ3-using composes (vllm/default, vllm/long-text, vllm/long-text-no-mtp,
vllm/long-vision, vllm/bounded-thinking, vllm/dual-turbo,
vllm/dual-tq3-mtp, vllm/dual-tq3-mtp-genesis, vllm/dual-tq3-nomtp,
vllm/dual-nvlink-turbo) stay on vllm-nightly-mtp.

Model profile update:
- qwen3.6-27b.requires_genesis flipped true → false.
  Strictly bootable on any qwen3-next-hybrid-capable vLLM nightly.
  Genesis is required only for TQ3 KV format; that's enforced at the
  compose level via Engine-profile selection.

Engine profile update:
- vllm-nightly-clean.supported_model_families adds qwen3-next-hybrid.
- Notes corrected to reflect the TQ3-only policy and broader family
  coverage.

Test updates:
- to_compose_name strict match: updated to expect vllm-nightly-clean
  for fp8/tp=2 long-ctx Qwen.
- C6 test reframed: under the TQ3-only policy no model declares
  requires_genesis=true, so the Genesis enforcement happens at C15
  (engine feature) level, not C6. Test now asserts positive (Qwen +
  fp8 on non-Genesis engine is valid) AND negative (Qwen + TQ3 on
  non-Genesis engine fails C15).

All compat tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 f7f6f444b9 feat(moe): wire Gemma 4 26B-A4B + Qwen 3.6 35B-A3B composes through fits()
Adds the v0.7.3 critical-path plumbing for both MoE models:

* Compose files:
  - models/gemma-4-26b-a4b/vllm/compose/single/docker-compose.yml
    (vllm-nightly-clean, TP=1, port 8040, bf16 KV, no drafter)
  - models/qwen3.6-35b-a3b/vllm/compose/single/preview.yml
    (vllm-nightly-clean, TP=1, port 8050, fp8_e5m2 KV, preview-only)

* COMPOSE_REGISTRY entries 'vllm/gemma-a4b' and 'vllm/qwen-a3b-preview'.

* DrafterProfile: new gemma-26b-it-assistant for the 26B-A4B
  (google/gemma-4-26B-A4B-it-assistant — separate weights from 31B);
  qwen-mtp-builtin.model_compat extended with qwen3.6-35b-a3b.

* vllm-nightly-clean.supported_model_families adds qwen3-next-moe as a
  preview/smoke route; notes clarify Cliff 2 / TQ3 unavailable here.

* Qwen 3.6 35B-A3B model: requires_genesis flipped false → upstream is
  strictly bootable on post-#42521 nightlies; Genesis is RECOMMENDED for
  production (TQ3 KV + Cliff 2 mitigations), not strictly required to boot.

* ModelProfile gains kv_calc_supported (defaults true). Both new MoE
  models set false because tools/kv-calc.py only models qwen3.6-27b and
  gemma-4-31b today (MoE + asymmetric KV + DeltaNet+attention hybrid all
  require new MODEL_SPECS + activation coefficients — queued for Codex).
  fits() now skips C12 for kv_calc_supported=false instead of failing.

All compat tests pass; both new composes fit on 1x-3090 (Gemma) and
1x-3090/2x-3090/1x-5090 (Qwen preview).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 15eda8a823 feat(profiles): split engine-pin policy by Genesis dependency
Add vllm-nightly-clean engine profile (latest nightly, no Genesis) for
model families that don't need DeltaNet stabilization patches.

Rule:
- required_genesis: true  → SHA pinned to Genesis anchor, moves only on
  Sander's pin-bump cycle (vllm-nightly-mtp/dflash/full)
- required_genesis: false → free to ride latest nightly (vllm-nightly-clean)

Routes:
- qwen3-next-hybrid + qwen3-next-moe → vllm-nightly-mtp (Genesis-locked)
- gemma4-swa-moe                     → vllm-nightly-clean (unconstrained)
- gemma4-swa-dense                   → either
- dense                              → either

Also adds qwen3-next-moe + gemma4-swa-moe to vllm-nightly-mtp's family
list (factual, independent of SHA), corrects misleading notes that
referenced an unmerged pin bump, and bumps engine-count assertion in
test-profiles-compat.sh from 6 to 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 abf0e327f9 feat(profiles): add Gemma 4 26B-A4B ModelProfile + num_global_kv_heads field
Onboards Gemma 4 26B-A4B (gemma4-swa-moe family) as the fourth
ModelProfile. Real values sourced from
/mnt/models/huggingface/gemma-4-26b-a4b-autoround-int4-mixed/config.json
(file downloaded but checkpoint still streaming in background as of
this commit).

Schema extension: new optional field `num_global_kv_heads`. Gemma 4
26B-A4B has asymmetric KV head counts (8 sliding / 2 global) — distinct
from Gemma 4 31B which uses 16 for both. Field is None for symmetric
models; existing Gemma 4 31B + Qwen models work unchanged.

Confirmed architectural facts (from config.json):
- 30 total hidden layers (smaller than estimated 40-48)
- 5 full_attention layers at indices [5, 11, 17, 23, 29]
  (one global per 6-layer block; last layer always global)
- 25 sliding_attention layers (sliding_window=1024)
- num_kv_heads=8 (sliding), num_global_kv_heads=2 (global) — ASYMMETRIC
- head_dim=256 (sliding), global_head_dim=512 (global)
- attention_k_eq_v: TRUE (K=V tied)
- num_experts: 128, top_k_experts: 8 (active per token)
- moe_intermediate_size: 704
- Multimodal: vision_config + audio token IDs present
- requires_genesis: false (Gemma 4 doesn't need Genesis like Qwen)

KV math implication: per-token growing KV at bf16 is
5 × 2 × 512 × 1 (K=V tied) = 5,120 bytes
vs Gemma 4 31B's 10 × 16 × 512 × 1 = 81,920 bytes per token
=> Gemma 4 26B-A4B has ~16x smaller growing-KV per token than 31B,
   making long context dramatically cheaper. KV_MATH.md MoE section
   estimates need refresh.

Drafter compat note: references gemma-it-assistant (the existing 31B
drafter). TODO Codex: extend that DrafterProfile's model_compat to
include gemma-4-26b-a4b, OR split into a separate
gemma-26b-it-assistant DrafterProfile (google/gemma-4-26B-A4B-it-assistant
is downloading in background alongside the main model).

Tests pass: 41 tests, kv-calc 18/18 invariant held.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunnaandClaude Opus 4.7 937871492a feat(profiles): add Qwen 3.6 35B-A3B ModelProfile (MoE schema extensions)
Onboards Qwen 3.6 35B-A3B (qwen3-next-moe family) as the third
ModelProfile in the v0.7.0 profile catalog. Architectural values sourced
from the actual config.json at /mnt/models/huggingface/qwen3.6-35b-a3b-
autoround-int4/config.json, NOT estimates.

Schema extensions (additive, optional — existing dense models work
unchanged):
- num_experts, num_experts_per_tok
- moe_intermediate_size, shared_expert_intermediate_size
- active_params_b (documentation only)
- mtp_num_hidden_layers (built-in MTP indicator)
- attn_output_gate (gated attention bool)
- vision_capable (multimodal support flag)

Confirmed architectural facts (from config.json):
- 40 total hidden layers
- 10 full_attention layers at indices [3,7,11,15,19,23,27,31,35,39]
  (full_attention_interval=4)
- 30 linear_attention (Gated DeltaNet) layers
- num_kv_heads=2 caps valid TP at [1, 2]
- 256 experts, 8 active per token (corrects KV_MATH.md's 128 estimate)
- mtp_num_hidden_layers=1 — built-in MTP drafter
- vision_config present — multimodal

5 weight variants on disk:
- autoround_int4 (production, 20 GB)
- gptq_int4 (experimental, 22 GB)
- gguf (production, 90 GB)
- dflash + dflash_gguf (experimental)

What's NOT in this commit (still to do for v0.7.3):
- Build first vllm/dual compose for the new model
- Take down llama estate fixtures + boot the new compose
- Capture boot log + back-solve per_token_bytes against KV math
- 4+ calibration anchors → calibration YAML
- kv-calc activation coefficients for Qwen 3.6 35B-A3B
- BENCHMARKS rows
- learnings/qwen3.6-35b-a3b.md
- KV_MATH.md MoE section update (remove "calibration pending", fill in
  measured coefficients)

Tests:
- test-profiles-compat.sh: 41 tests pass (was 34 pre-v0.7.2 topology
  + 7 topology + bumped model count 2 → 3)
- kv-calc --calibration: 18/18 = 100% (held)
- All 7 other test suites pass

Branch reasoning: feature branch for Codex review before merge to
master. Live boot + calibration discipline benefits from Codex's
focused-session bench rigor.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 22:18:09 +05:00
noonghunna d116ba9ba3 feat(launch): add hardware topology advisor
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m22s
Release / release (push) Failing after 48s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-15 09:54:04 +00:00
noonghunna 2a148d702b feat(bench): surface prompt processing throughput
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m23s
Release / release (push) Failing after 45s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-15 09:03:10 +00:00
noonghunna c2adb3970d feat(scripts): add diagnose-profile triage
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 22m22s
Release / release (push) Failing after 46s
Build vLLM Club3090 Image / Detect self-hosted GPU runner (push) Canceled after 0s
Build vLLM Club3090 Image / GPU smoke and alias promotion (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
2026-05-14 22:07:00 +00:00
noonghunna 40f1ef78f8 feat(profiles): resolve vllm nightly pins 2026-05-14 21:38:47 +00:00
noonghunna 52e43470c6 fix(launch): persist estate source of truth 2026-05-14 20:18:12 +00:00
noonghunna c9b153f91f feat(launch): add estate planner orchestration 2026-05-14 17:04:38 +00:00
noonghunna a142b1ce10 feat(launch): validate single-model profiles 2026-05-14 16:38:50 +00:00
noonghunna 6581ccca9e feat(compat): add profile validator and estate self-test 2026-05-14 15:25:28 +00:00
noonghunna 98f0406d0f fix(launch): project TP greater than four
Release / release (push) Failing after 48s
2026-05-14 00:23:32 +00:00
noonghunna 5882bbef6f feat(launch): add hardware-aware launcher 2026-05-13 23:43:32 +00:00
noonghunna 12a33fbdd1 feat(scripts): add hardware-aware setup picker 2026-05-13 20:44:35 +00:00
noonghunna 22bf2e9398 fix(scripts): make submit-bench issue-first
Release / release (push) Failing after 53s
2026-05-13 18:40:43 +00:00
noonghunna ef770322f4 feat(scripts): add submit-bench flow
Release / release (push) Failing after 51s
2026-05-13 18:24:14 +00:00
noonghunna 26985527f7 Add hardware-aware compose preflight
Release / release (push) Failing after 50s
2026-05-13 18:03:37 +00:00