7 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 421114b9dd fix(vllm-pr35936): make overlay tolerate bf610c2f upstream drift (closes #144)
Build vLLM Club3090 Image / Build and push dated image (push) Failing after 1m28s
Release / release (push) Failing after 50s
Build vLLM Club3090 Image / Promote latest and nightly-stable (push) Canceled after 0s
Build vLLM Club3090 Image / Retain four weeks of dated nightlies (push) Canceled after 0s
The PR #35936 required-tool-fallback overlay was captured against vLLM
source HEAD a2696294 / image nightly-1acd67a7. The v0.7.3 engine-pin
policy split (3b2d940) routed 8 non-TQ3 Qwen 3.6-27B composes to
vllm-nightly-clean (bf610c2f), where three upstream API drifts broke the
overlay. All three are now defensively patched and validated end-to-end
on bf610c2f (boot + /v1/chat/completions + /v1/completions all green,
zero errors, system_fingerprint vllm-...gbf610c2f5):

1. engine/serving.py — `speech_to_text` relocated from
   entrypoints/openai/speech_to_text → entrypoints/speech_to_text and
   split into transcription/ + translation/ subpackages. try/except
   imports old path first (matches captured SHA), falls back to new.

2. chat_completion/serving.py — `RequestOutput.prompt_routed_experts`
   removed on the dense path in bf610c2f. getattr(..., None) so the
   overlay tolerates both layouts; dense path produces None as before.

3. engine/serving.py — bf610c2f's chat_completion / completion /
   responses serving call `self._with_kv_transfer_rejection_cleanup(...)`,
   a method absent from the captured base. Added a passthrough stub
   (club-3090 doesn't run disaggregated KV-transfer / remote-prefill).

Drift surface was bounded to these three points — verified by live boot
on bf610c2f, not just import-time. Engine-pin policy split (3b2d940)
stays intact; reverts d32e168/e7bca8e were undone.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-15 20:58:14 +00:00
noonghunna d7804107c9 Reapply "fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean"
This reverts commit d32e168a89.
2026-05-15 20:52:21 +00:00
noonghunna 0c4260f53d Revert "test(launch): re-align engine pin expectations after revert"
This reverts commit e7bca8e035.
2026-05-15 20:52:21 +00:00
noonghunna e7bca8e035 test(launch): re-align engine pin expectations after revert
Companion to the revert of 3b2d940. The test-launch-compat.sh
expectations were updated in 127f4f6 to expect CLEAN_SHA for vllm/dual;
since the routing went back to vllm-nightly-mtp (MTP_SHA = 01d4d1ad),
flip back. Removes the now-redundant CLEAN_SHA assertion in the
estate_cli compose_env block (both vllm/dual and vllm/dual-tq3-mtp now
resolve to MTP_SHA, so one assertion covers both).

All 8 test suites green.
2026-05-15 18:45:23 +00:00
noonghunna d32e168a89 Revert "fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean"
This reverts commit 3b2d940d26.
2026-05-15 18:40:12 +00:00
noonghunna 77802a3d8e docs(BENCHMARKS): @OVDEN13 dual.yml — PCIe Gen 4 x4+x8 asymmetric (#142)
First asymmetric PCIe-lane data point on the matrix. 75.80 narr / 99.04
code wall TPS, soak PASS (20×5 fresh, 0 errors, 100.7% retention,
0 MiB growth). −15% narr / −14% code vs @danbedford's same-compose
Gen 4 x16 PCIe-only A/B (#77) — gap consistent with the Gen 4 x4 card
being the bottleneck for cross-card NCCL allreduce. MTP per-position
acceptance also lower-floor (0.81/0.56/0.35) vs reference stable
(0.94/0.84/0.72), suggesting draft-forward sync stalls on the slow lane.
2026-05-15 17:30:44 +00:00
github-actions[bot] c3d1c5d67d chore(changelog): regenerate for v0.7.3 [skip ci] 2026-05-15 17:19:12 +00:00
4 changed files with 90 additions and 7 deletions
+1
View File
@@ -115,6 +115,7 @@ Primary serving model. Hybrid Qwen3-Next architecture (DeltaNet GDN + standard a
| `bounded-thinking.yml` | @lolren (2× 3090 PCIe + Ryzen 9 5950X, 250W cap, **MTP-disabled-suspected**) | TQ3 | 180K | 64.86 / 64.96 (CV **0.1%**) | — | ~22.3 GB | 2026-05-05 | **Anomaly:** lolren reports "no spec-decode" on this run despite `bounded-thinking.yml` shipping `--speculative-config mtp n=3` by default. Near-identical narr=code TPS + extreme CV stability (0.1%) suggests MTP was inactive — likely because his image was older `nightly-7a1eb8ac2` (pre-v7.72.2 + pre-PN35). Re-test on `nightly-01d4d1ad3` should restore MTP path → expect ~50/66 narr/code with normal CV. Tracked. [Disc #18](https://github.com/noonghunna/club-3090/discussions/18#discussioncomment-16820303). |
| `dual.yml` | @JDWarner (**Mixed RTX A5000 + RTX 3090**, both **Razer Core X eGPU enclosures over Thunderbolt 3**, Intel NUC11TNH i5-1135G7, **16 GB RAM**, headless, A5000=230W cap / 3090=290W cap, PCIe **x4 Gen 3** per card) | fp8 | 262K | **56.83 / 72.47** (soak p50 93.09) | — | ~23.6 GB/card | 2026-05-09 | **Soak: ✓ PASS** (5×5, 0 errors, 0 silent-empty, 100% TPS retention, 0 MiB growth). **First TB3 dual-eGPU + mixed-arch cross-rig data**. The setup that "shouldn't work": each card on a separate TB3 controller → ~3.94 GB/s effective per card vs ~32 GB/s on PCIe x16 Gen 4 (~8× cut), mixed Ampere SKUs (workstation A5000 + consumer 3090 with different mem bandwidth + clocks), 16 GB system RAM total. **Result: matches `dual.yml` PCIe x16 baseline within run-to-run noise** — confirms decode on Qwen3.6-27B is per-card-bandwidth bound, cross-card NCCL allreduce is small enough that even an 8× link cut doesn't dominate. Extends @aaronlockhartdev's #91/#95 finding (patched-P2P only +2%/+9% on `dual.yml`) in the opposite direction: even with 8× *less* cross-card bandwidth, decode holds. MTP AL 3.39-3.52, per-pos accept 0.93/0.83/0.70 (89% avg). verify-full + verify-stress all PASS. Genesis pin `7b9fd319` (v7.72.2). [Issue #107](https://github.com/noonghunna/club-3090/issues/107). |
| `dual/docker-compose.yml` (default) | @ygafarov (**3090 via USB4 eGPU dock + 5070 Ti via OCuLink** — heterogeneous Ampere + Blackwell consumer dual-eGPU, AMD Ryzen AI MAX+ 395 / Strix Halo miniPC, CachyOS, 123 GB RAM, 290 W cap both cards, PCIe **x4** per card — USB4 ≈ 3.94 GB/s, OCuLink ≈ 7.88 GB/s) | fp8 | 200K | **65.10 / 85.81** | — | 17.1 / 15.7 GB | 2026-05-12 | **First heterogeneous Ampere + Blackwell consumer dual-eGPU on the matrix.** TP=2 bound by the slower USB4 link in allreduce + sm_86 kernels (5070 Ti spends back-half of step waiting — 91% util but only 125 W out of 290 W cap). KV pool 200K @ 1.00× concurrency — VRAM cap from the 5070 Ti's 16 GiB (model takes 13.8 GiB/card → only ~2.2 GiB left for KV on the smaller card). verify-stress 7/7 incl. **91K needle recall** (Cliff 2 clean). Soak ⚠ borderline (360 MiB > 200 MiB threshold — same eGPU-bus accretion as ygafarov's own #113 single-card row above at 240 MiB; 100% TPS retention + 0 silent-empty + 0 errors so not a leak). MTP AL 3.50, per-pos accept 0.94/0.86/0.70. CV 4.5%/1.8%. **Slower than ygafarov's own single-3090 #113 row** (68.86/91.70 at 48K) — on this rig the single-card path is recommended; the 5070 Ti adds VRAM cap pain without TPS gain. Driver 595.71.05, vLLM `nightly-1acd67a79`, no Genesis (Blackwell consumer not on allowlist). [Issue #120](https://github.com/noonghunna/club-3090/issues/120). |
| `dual.yml` | @OVDEN13 (2× 3090 + Ryzen 7 5700X, **PCIe Gen 4 x4 + x8 asymmetric**, 300 W cap, no NVLink, Ubuntu 26.04 / driver 595.58.03 / CUDA 13.2, 92 GB RAM, bare metal) | fp8 | 262K | **75.80 / 99.04** | — | 22.3 GB/card | 2026-05-15 | **First asymmetric PCIe-lane data point on the matrix** (Gen 4 x4 to one card, x8 to the other — half the cross-card bandwidth of @danbedford's Gen 4 x16 baseline). **−15% narr / −14% code vs @danbedford's same-compose PCIe-only A/B** (89.24 / 114.57, [#77](https://github.com/noonghunna/club-3090/issues/77)) — gap consistent with the Gen 4 x4 card being the bottleneck for cross-card NCCL allreduce. MTP AL 2.71–3.51, per-pos accept ranges 0.81/0.56/0.35 → 0.95/0.85/0.72; the wide variance and lower floor vs @lolren's stable 0.94/0.84/0.72 ([disc #18](https://github.com/noonghunna/club-3090/discussions/18#discussioncomment-16820303)) suggest draft-forward sync stalls on the slow lane. Both GPUs pegged at 297 W power cap (80% / 87% util — mildly asymmetric, consistent with x4+x8 imbalance). CV 1.2% narr / 4.2% code. **Soak: ✓ PASS** (20×5 fresh mode, p50 101.41 TPS, p95 TTFT 1645 ms, 0 errors, 0 silent-empty, 100.7% retention, 0 MiB growth). Stack: master @ `b1c68b4` (v0.7.2-7-g, pre-v0.7.3 tag), vLLM `nightly-1acd67a795eb`, Genesis pin `7b9fd319`. [Issue #142](https://github.com/noonghunna/club-3090/issues/142). |
### Quad-card (4× RTX 3090, TP=4)
+52
View File
@@ -16,6 +16,58 @@ history; SemVer takes over from `v0.3.0` onward.
---
## v0.7.3 — 2026-05-15
### ✨ Features
- feat(report): surface kv-calc calibration verdict ([#143](https://github.com/noonghunna/club-3090/pull/143) by @noonghunna)
- feat(kv-calc): model v0.7.3 MoE architectures ([39e1873](https://github.com/noonghunna/club-3090/commit/39e18733aa8f14c83a1f76d5bed23156fda61568))
- feat(gemma-4-26b-a4b): AWQ + MTP n=4 — +12% narr / +49% code over no-MTP baseline ([6dc9a0d](https://github.com/noonghunna/club-3090/commit/6dc9a0dce1b4ecef787e50808a8a15439cabe209))
- feat(qwen-35b-a3b): preview-MTP compose + bench row — MTP measured SLOWER on MoE ([e1d44bd](https://github.com/noonghunna/club-3090/commit/e1d44bd732cef0757b8d3f0f872b5b0a1a4fe5cc))
- feat(vllm-pr41800): vendor truncate_prompt_tokens overlay across all pre-fix engines (closes #139) ([1d7aad1](https://github.com/noonghunna/club-3090/commit/1d7aad112c097709d954750ab0e725a785ad062a))
- feat(gemma-4-26b-a4b): AWQ path via vLLM PR #40886 overlay ([0053444](https://github.com/noonghunna/club-3090/commit/0053444e84f80ce9f3fb910435c8d627591f59ed))
- feat(estate): add parallel boot mode ([99328b4](https://github.com/noonghunna/club-3090/commit/99328b4cda7ec9efc723a001295525a6d1f86e2a))
- feat(moe): add dual-card composes for Gemma 26B-A4B + Qwen 35B-A3B preview ([2d1b1dc](https://github.com/noonghunna/club-3090/commit/2d1b1dc347c3897f8b902f7ab60c8bd9c97abbeb))
- feat(moe): wire Gemma 4 26B-A4B + Qwen 3.6 35B-A3B composes through fits() ([f7f6f44](https://github.com/noonghunna/club-3090/commit/f7f6f444b9e937396c3de70df7dd6a5b957be80b))
- feat(profiles): split engine-pin policy by Genesis dependency ([15eda8a](https://github.com/noonghunna/club-3090/commit/15eda8a823dd273ee9b238369cba365021b61a29))
- feat(profiles): add Gemma 4 26B-A4B ModelProfile + num_global_kv_heads field ([abf0e32](https://github.com/noonghunna/club-3090/commit/abf0e327f9b877f989fe16c88bda1a457c144e0a))
- feat(profiles): add Qwen 3.6 35B-A3B ModelProfile (MoE schema extensions) ([9378714](https://github.com/noonghunna/club-3090/commit/937871492ae08b58b041a1fded41ab793f0c8549))
### 🐛 Bug fixes
- fix(gpu-mode): mode_off tears down estate-managed instances ([9cd854d](https://github.com/noonghunna/club-3090/commit/9cd854dbcb9b8b215876e0f88443c269072b93b9))
- fix(qwen3.6-27b): route non-TQ3 composes to vllm-nightly-clean ([3b2d940](https://github.com/noonghunna/club-3090/commit/3b2d940d26ec2ffe6daf631e77986898c1d2849d))
- fix(gemma-4-31b): route default/bf16 composes to vllm-nightly-clean ([cf0451a](https://github.com/noonghunna/club-3090/commit/cf0451a7ad5c56bd7b63327aac7178df70cba1fb))
- fix(engines): vllm-nightly-mtp anchors to 01d4d1ad (Sander v7.72.2 PROD pin) ([87f0a0c](https://github.com/noonghunna/club-3090/commit/87f0a0c528473a225f06997bdcb9b79fefa13b04))
### 📝 Documentation
- docs(soak-test): clarify PASS verdict semantics — closes #140 ([9a039d8](https://github.com/noonghunna/club-3090/commit/9a039d8c922d3141102e8244d5468e8de46c6674))
- docs(UPSTREAM): add PR #41800 truncate_prompt_tokens row ([273c017](https://github.com/noonghunna/club-3090/commit/273c017646087fe508626be7ecea5e2106be54be))
- docs(README): add v0.7.3 MoE models to Supported Models table ([e49c939](https://github.com/noonghunna/club-3090/commit/e49c9397481b5fba226323b59c4fa37bdc2aeeab))
- docs(BENCHMARKS): Gemma 4 26B-A4B AWQ first row + AutoRound row demoted ([92b69bd](https://github.com/noonghunna/club-3090/commit/92b69bd220ce8360d2dfd5fdcf89d26ee2264791))
- docs(BENCHMARKS): add v0.7.3 MoE preview section ([bdfb939](https://github.com/noonghunna/club-3090/commit/bdfb939edd98c3872439a7a05dcde13cce7ccaa2))
- docs(HARDWARE): add note on PCIe Gen 3 + older CPU TP=2 headwind ([8cf38b0](https://github.com/noonghunna/club-3090/commit/8cf38b05906e3954405af3db09a822ac87a25ae5))
- docs(KERNEL_MATRIX): add KV Cache Impact subsection ([1a233cd](https://github.com/noonghunna/club-3090/commit/1a233cd9fd05a6ec51adfadae79813bc80cea4b0))
- docs: add KERNEL_MATRIX.md (attention backend + engine support matrix) ([97195fe](https://github.com/noonghunna/club-3090/commit/97195fe443770bc8adc883755fd3baf9933df874))
- docs(kv-math): extend k_v_tensors=N notation to sliding-KV formulas ([b1c68b4](https://github.com/noonghunna/club-3090/commit/b1c68b4fe3211af3442cd3bb7e9fd011e30e5554))
- docs(kv-math): tighten k_v_tensors notation across all 4 formulas ([54d8bf0](https://github.com/noonghunna/club-3090/commit/54d8bf0b3e2b322dbd67e6939ff9f9995c642437))
- docs(kv-math): third-pass Grok polish ([67de3ec](https://github.com/noonghunna/club-3090/commit/67de3eca3eb4b75b43fd0f83d2ed304976b3f1b5))
- docs(kv-math): second-pass Grok polish ([78f94ca](https://github.com/noonghunna/club-3090/commit/78f94cacf006064f68d945122907d30be8723eec))
- docs(kv-math): address Grok review feedback ([3114399](https://github.com/noonghunna/club-3090/commit/31143999837639b5a7d492d2f2b64e164787bcfe))
- docs(kv-math): config-verify Qwen 35B-A3B + Gemma 26B-A4B MoE sections ([6ec6a67](https://github.com/noonghunna/club-3090/commit/6ec6a6761f36cca44ea819434be7205bcb6fdbc6))
### 🧹 Maintenance
- test(launch): align engine pin expectations ([127f4f6](https://github.com/noonghunna/club-3090/commit/127f4f6d8fe104a954b7865a4d7550017a1c629b))
[Pin: `git checkout v0.7.3`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.7.2...v0.7.3)
## v0.7.2 — 2026-05-15
@@ -1358,9 +1358,13 @@ class OpenAIServingChat(OpenAIServing):
request_metadata.final_usage_info = usage
# club-3090 patch: `prompt_routed_experts` was removed from
# `RequestOutput` in vLLM bf610c2f. Use getattr so the overlay tolerates
# both pre-bf610c2f and post-bf610c2f layouts. Dense path produces None.
_prompt_routed_experts_raw = getattr(final_res, "prompt_routed_experts", None)
prompt_routed_experts = None
if final_res.prompt_routed_experts is not None:
prompt_routed_experts = final_res.prompt_routed_experts.tolist()
if _prompt_routed_experts_raw is not None:
prompt_routed_experts = _prompt_routed_experts_raw.tolist()
response = ChatCompletionResponse(
id=request_id,
@@ -38,11 +38,24 @@ from vllm.entrypoints.openai.engine.protocol import (
)
from vllm.entrypoints.openai.models.serving import OpenAIServingModels
from vllm.entrypoints.openai.responses.protocol import ResponsesRequest
from vllm.entrypoints.openai.speech_to_text.protocol import (
TranscriptionRequest,
TranscriptionResponse,
TranslationRequest,
)
# club-3090 patch: speech_to_text was relocated from
# `vllm.entrypoints.openai.speech_to_text` (pre-bf610c2f) to
# `vllm.entrypoints.speech_to_text` and split into transcription/ + translation/
# subpackages. Try old path first (matches captured SHA), fall back to new.
try:
from vllm.entrypoints.openai.speech_to_text.protocol import (
TranscriptionRequest,
TranscriptionResponse,
TranslationRequest,
)
except ImportError:
from vllm.entrypoints.speech_to_text.transcription.protocol import (
TranscriptionRequest,
TranscriptionResponse,
)
from vllm.entrypoints.speech_to_text.translation.protocol import (
TranslationRequest,
)
from vllm.entrypoints.serve.disagg.protocol import GenerateRequest, GenerateResponse
from vllm.entrypoints.serve.tokenize.protocol import (
DetokenizeRequest,
@@ -615,6 +628,19 @@ class OpenAIServing:
except ValueError:
return None
# club-3090 patch: bf610c2f's `chat_completion/serving.py`, `completion/serving.py`,
# and `responses/serving.py` call `self._with_kv_transfer_rejection_cleanup(...)`.
# The pre-bf610c2f base this overlay was captured against didn't define it.
# club-3090 doesn't run disaggregated KV-transfer / remote-prefill workflows,
# so this stub just delegates to the awaitable — the cleanup branch never fires.
async def _with_kv_transfer_rejection_cleanup(
self,
awaitable,
request,
raw_request,
):
return await awaitable
@staticmethod
def _parse_tool_calls_from_content(
request: ResponsesRequest | ChatCompletionRequest,