Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6617e1e090 | ||
|
|
1678ca0c8b | ||
|
|
14ffe45667 | ||
|
|
722f998ff3 |
@@ -16,6 +16,44 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.5.0 — 2026-05-12
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(qwen): ship froggeric chat-template fixes as default-on ([84498d4](https://github.com/noonghunna/club-3090/commit/84498d47aaf7a2fdb7c0203d53bb64414a64b6c1))
|
||||
- feat(vllm): add PR #35936 required-tool fallback overlay ([28b16b5](https://github.com/noonghunna/club-3090/commit/28b16b5dc9d602a8e1c4b8d5496aa82cbca7f95d))
|
||||
- feat(qwen-tq3): add CLUB3090_TQ_K1_SKIP_MTP layer-filter for PR #40914 K+1 dispatch ([6b2a7d5](https://github.com/noonghunna/club-3090/commit/6b2a7d553b6164bb4854e80fc0acdf8dcec18a87))
|
||||
|
||||
|
||||
### 🎯 New models + serving paths
|
||||
|
||||
- compose(tq3-mtp-genesis): pin to Genesis v7.72.2 known-good vLLM nightly ([570fa71](https://github.com/noonghunna/club-3090/commit/570fa71240a12ab642693c0570851611e933f8c4))
|
||||
|
||||
|
||||
### 📊 Benchmarks + cross-rig data
|
||||
|
||||
- bench(matrix): @ygafarov first heterogeneous Ampere + Blackwell eGPU dual ([1770931](https://github.com/noonghunna/club-3090/commit/1770931729a354bad319c8b58bdee143fe6ebce2))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs(dtype-matrix): more polish — RDNA naming, FP8 maturity caveats, AMD detection ([62b3b45](https://github.com/noonghunna/club-3090/commit/62b3b455a9ea5146cbf5576febd0c8b25a8c0fa1))
|
||||
- docs(dtype-matrix): polish nuances + add Intel and AMD vendor sections ([3d4548c](https://github.com/noonghunna/club-3090/commit/3d4548c50422da07c16fcba2a59d6f42f268355b))
|
||||
- docs(dtype-matrix): per-arch hardware accelerator matrix for compose optimization ([9c6d3cf](https://github.com/noonghunna/club-3090/commit/9c6d3cfba1de9d9079d6eafc9ff68c352cac7197))
|
||||
- docs(faq): add 'INT8 PTH doesn't scale at concurrency — is that a bug?' ([df53287](https://github.com/noonghunna/club-3090/commit/df53287b1c26ec83be8a30eec24baf2bddc993eb))
|
||||
- docs(tq3-mtp): writeup + charts for the Genesis-backed TQ3+MTP path ([c2b1c93](https://github.com/noonghunna/club-3090/commit/c2b1c93872f84fa9afa3bbe41360dc42be28c066))
|
||||
- docs(qwen-tq3): close round-4 — #40914 not shippable, route to nomtp + Genesis ([9fba037](https://github.com/noonghunna/club-3090/commit/9fba03788e30151f9bf8c85f279260696954d094))
|
||||
- docs(qwen-tq3): re-tombstone tq3-mtp.yml after round-3 MTP-skip validation ([063d3e9](https://github.com/noonghunna/club-3090/commit/063d3e943ce8da9cc69bd30c68b87426aca6202e))
|
||||
|
||||
|
||||
### 🧹 Maintenance
|
||||
|
||||
- refactor(qwen): rename int8-tq3 → tq3-* family + add no-MTP + Genesis variants ([6182922](https://github.com/noonghunna/club-3090/commit/6182922225dfeee1c28084d1ff917bfd25539520))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.5.0`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.4.0...v0.5.0)
|
||||
## v0.4.0 — 2026-05-11
|
||||
|
||||
|
||||
|
||||
@@ -67,6 +67,9 @@ git clone https://github.com/noonghunna/club-3090.git
|
||||
cd club-3090
|
||||
|
||||
# 2. Download + SHA-verify the model (~20 GB; clones Genesis patches too)
|
||||
# (asks you where to put model weights — pick in-repo default, ~/models, or
|
||||
# a custom path on a different drive. To skip the prompt:
|
||||
# `export MODEL_DIR=/mnt/your-drive/models` before running. See FAQ + .env.example.)
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
|
||||
# 3. Pick a config + boot it (interactive wizard — asks engine / cards / workload)
|
||||
|
||||
+19
-2
@@ -130,9 +130,26 @@ If your numbers on the same compose look different from ours by >15%, the most l
|
||||
|
||||
## Setup
|
||||
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go?
|
||||
### `bash scripts/setup.sh qwen3.6-27b` is downloading 20+ GB. Where does it go? / Can I put models on a different drive?
|
||||
|
||||
`<repo>/models-cache/` by default. Override with `MODEL_DIR=/path/to/your/scratch bash scripts/setup.sh qwen3.6-27b`. See [`.env.example`](../.env.example) for all env vars.
|
||||
Yes. The knob is `MODEL_DIR`, with **four ways** to set it (priority order):
|
||||
|
||||
1. **`MODEL_DIR` env var in your shell** — takes precedence over everything:
|
||||
```bash
|
||||
export MODEL_DIR=/mnt/your-second-drive/models
|
||||
bash scripts/setup.sh qwen3.6-27b
|
||||
```
|
||||
2. **`.env` file at repo root** — picked up automatically on every script run. See [`.env.example`](../.env.example).
|
||||
3. **Interactive prompt** — `bash scripts/setup.sh qwen3.6-27b` with nothing set offers three choices: in-repo default, `~/models`, or custom path. After you pick custom, it asks "Save `MODEL_DIR=/your/path` to `.env` so we skip this next time?" — say `Y` and it persists for every subsequent `launch.sh` / `switch.sh` / `bench.sh` call.
|
||||
4. **Silent fallback** — `<repo>/models-cache/`. Functional but pollutes the git tree; not recommended.
|
||||
|
||||
Every script that touches model paths reads from the same `MODEL_DIR`. The compose YAMLs' volume mount is `${MODEL_DIR:-...}:/root/.cache/huggingface` — once set, every container reads + writes there.
|
||||
|
||||
**HF env-var integration** — we don't directly respect `HF_HOME` / `HF_HUB_CACHE` because we mount a host directory INTO the container's `/root/.cache/huggingface`, not the host's HF cache. The internal layout inside `MODEL_DIR` matches HF's repo-cache convention (`<MODEL_DIR>/<repo-subdir>/`), so models downloaded by `setup.sh` are byte-compatible with anything that reads HF's local cache layout. Two clean workarounds if you already have an HF cache you want to reuse:
|
||||
- Set `MODEL_DIR=$HF_HOME/hub`
|
||||
- Or symlink between them
|
||||
|
||||
**On Windows / WSL2** — same mechanism. Docker Desktop handles path translation. Use Windows paths (`D:\models`) from PowerShell or WSL paths (`/mnt/d/models`) from WSL. If you flip between Linux and Windows on the same rig, point `MODEL_DIR` at a drive both OSes can see — the model files themselves are OS-agnostic.
|
||||
|
||||
### How do I keep my install up-to-date?
|
||||
|
||||
|
||||
+2
-1
@@ -69,7 +69,8 @@ Run `bash scripts/maintenance/list-image-pins.sh` for a live snapshot.
|
||||
|---|---|---|---|
|
||||
| [#35936](https://github.com/vllm-project/vllm/pull/35936) — `tool_choice="required"` falls back to configured tool parser | 🟡 Open / **local overlay active** | Qwen3-Coder with `--tool-call-parser qwen3_coder` emits XML-style tool calls. On pinned nightly `1acd67a79`, non-streaming `tool_choice="required"` validates JSON only, bypasses the configured parser, and returns `tool_calls=[]`. MLS-Bench hits this when `thinking.enabled=false`. | Vendored overlay: [`models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-pr35936-required-fallback/README.md). Drop when #35936 or equivalent lands in our pinned image. |
|
||||
| [#40361](https://github.com/vllm-project/vllm/pull/40361) — Marlin pad-sub-tile-n | 🟡 Open, mergeable, **stale 13d** (last update 2026-04-20) | All 4 dual-card composes + `dual-nvlink.yml` + `dual-nvlink-turbo.yml` mount the patched files vendored in-repo at `models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/`. Drops out as a setup dependency when this merges + propagates. | Vendored mount: see [`models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md`](../models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/README.md). Queued for rebase + ping next week (see "Active follow-ups" table above). |
|
||||
| [#40807](https://github.com/vllm-project/vllm/issues/40807) — `.tolist()` cudagraph crash on continuation-prefill | ⚫ Local workaround | Single-card TQ3 + spec-decode + chunked-prefill blocked without it. We ship a file-edit patch. | `patch_tolist_cudagraph.py` runs in `setup.sh`. Drop when upstream fixes the sync. |
|
||||
| [#40807](https://github.com/vllm-project/vllm/issues/40807) — `.tolist()` cudagraph crash on continuation-prefill | ✅ **Retired locally** (2026-05-05 Genesis v7.72.2 bump) — Genesis ships [P78 `TOLIST_CAPTURE_GUARD`](../models/qwen3.6-27b/vllm/compose/dual/tq3-mtp-genesis.yml) as the equivalent fix. Currently disabled (`=0`) on `tq3-mtp-genesis.yml` after rebench-full leg 6 (2026-05-11) passed clean with it off — apparent root cause is now covered by Genesis PN34 (workspace-lock relax) + post-#41434 attention rework. Non-Genesis composes on `1acd67a79` pin run without any guard for this bug; unvalidated at long-context TurboQuant chunked-prefill (worth testing per cferra's vllm#41403 validation pass — see vllm#40798 row below). | None active. Drop the Genesis env var permanently if a future v7.73.x rebench leaves it OFF without regression. |
|
||||
| [#40798](https://github.com/vllm-project/vllm/pull/40798) + [#42215](https://github.com/vllm-project/vllm/pull/42215) — share decode scratch workspace pre-CUDA-graph + decode-kernel warmup | 🟡 Open, validated cross-rig | Pair closes the `AssertionError: Workspace is locked but allocation requires NMB` crash that fires at `turboquant_attn.py:_continuation_prefill` for ≥48K-token chunked-prefill with TurboQuant KV. Independently validated on 2× 3090 sm_86 by cferra (vllm#41403 [comment](https://github.com/vllm-project/vllm/issues/41403#issuecomment-4435164709), 2026-05-12). | Genesis [PN34 `WORKSPACE_LOCK_RELAX`](../models/qwen3.6-27b/vllm/compose/dual/tq3-mtp-genesis.yml) addresses the same symptom via a different mechanism (relax-lock vs reserve-before-capture). On non-Genesis composes (`1acd67a79` pin) we currently have no guard — re-validate against this PR pair once they propagate to a nightly we pin to, then A/B PN34 vs upstream. |
|
||||
| [#40849](https://github.com/vllm-project/vllm/pull/40849) — MTP draft online-quant propagation | 🟡 Open / Genesis backport active | Closes Cliff 1 on FP8+MTP path (`tools-text.yml`). | Genesis PN8 backport: `GENESIS_ENABLE_PN8_MTP_DRAFT_ONLINE_QUANT=1`. |
|
||||
| [#40914](https://github.com/vllm-project/vllm/pull/40914) — Sandermage K+1 verify routing | 🟡 Open, ❌ negative on our Qwen3.6-27B stack | **Reframed 2026-05-11:** the synthetic `seq_lens` K+1 route is not the P67-equivalent we need here. Local rebase on post-#41434 nightly made MTP acceptance look perfect (AL=4.0 / ~100%) but produced `!`-flood needle corruption plus tool/multi-turn timeouts. Dropping it improved verify-stress from 3/7 to 5/7, but TQ3/TQ4/k8v4 + MTP still fail long-context needles. | Do not ship Genesis-free TQ+MTP on #40914 alone. Use `dual/tq3-nomtp.yml` without Genesis, or `dual/tq3-mtp-genesis.yml` with Genesis P67/P67b. |
|
||||
| [#40334](https://github.com/vllm-project/vllm/pull/40334) — DFlash `combine_hidden_states` dtype mismatch | 🟡 Open | All `dual-dflash*.yml` need `--dtype bfloat16` flag to work around. | Composes set `--dtype bfloat16`. Drop when this lands. |
|
||||
|
||||
@@ -44,8 +44,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -74,6 +80,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -73,8 +73,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
environment:
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
@@ -103,6 +109,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -62,8 +62,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -97,6 +103,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -86,8 +86,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -121,6 +127,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -74,8 +74,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -112,6 +118,8 @@ services:
|
||||
# hardware where Cliff 2 GDN activation spikes occur at runtime
|
||||
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
|
||||
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -44,8 +44,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -74,6 +80,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -94,8 +94,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -131,6 +137,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -90,8 +90,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -127,6 +133,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -87,8 +87,14 @@ services:
|
||||
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
|
||||
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -222,6 +228,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -80,8 +80,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -119,6 +125,8 @@ services:
|
||||
# hardware where Cliff 2 GDN activation spikes occur at runtime
|
||||
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
|
||||
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -49,8 +49,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
environment:
|
||||
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
||||
- VLLM_WORKER_MULTIPROC_METHOD=spawn
|
||||
@@ -69,6 +75,14 @@ services:
|
||||
- driver: nvidia
|
||||
count: all
|
||||
capabilities: [gpu]
|
||||
entrypoint:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
- --model
|
||||
- /root/.cache/huggingface/qwopus3.6-27b-int4-recipe-d-bf16mtp
|
||||
|
||||
@@ -194,7 +194,14 @@ services:
|
||||
# PN54 — GDN contiguous-call deduplication (Cliff 2b OOM mitigation).
|
||||
- GENESIS_ENABLE_PN54=${GENESIS_ENABLE_PN54:-0}
|
||||
# Explicit OFFs to match Sandermage's PROD env-var set verbatim:
|
||||
# P78 (P78_TOLIST_CAPTURE_GUARD) — superseded by our patch_tolist_cudagraph.py
|
||||
# P78 (P78_TOLIST_CAPTURE_GUARD) — vllm#40807 `.tolist()` cudagraph guard.
|
||||
# Was OFF here because we previously shipped our own patch_tolist_cudagraph.py
|
||||
# sidecar covering the same bug; that sidecar was retired 2026-05-05 (Genesis
|
||||
# v7.72.2 bump). We keep P78=0 because rebench-full leg 6 (2026-05-11) passed
|
||||
# clean with it off — 7/7 verify-stress, 0/100 silent-empty soak, 18/30 aider.
|
||||
# PN34 (workspace-lock relax) + post-#41434 attention rework appear to cover
|
||||
# the symptom on this Genesis-pinned nightly. Flip to 1 if a future bench
|
||||
# surfaces .tolist() cudagraph corruption.
|
||||
# P81 (FP8 block-scaled M<=8) — FP8-specific, no-op on our TQ3 path
|
||||
# P82 — biased on small-batch single-stream Lorbus INT4 + MTP K=3 (Sander PROD)
|
||||
- GENESIS_ENABLE_P78_TOLIST_CAPTURE_GUARD=0
|
||||
|
||||
@@ -91,8 +91,14 @@ services:
|
||||
# cudagraph path off the table entirely.
|
||||
- ../../patches/vllm-pr40798-rebased/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -122,6 +128,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -75,8 +75,14 @@ services:
|
||||
# used; nightly already has the kernel + caller for the rest.
|
||||
- ../../patches/vllm-pr40798-rebased/v1/worker/gpu_model_runner.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_model_runner.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -105,6 +111,8 @@ services:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -71,8 +71,14 @@ services:
|
||||
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
|
||||
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -204,6 +210,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -79,8 +79,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -114,6 +120,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -82,8 +82,14 @@ services:
|
||||
- ../../patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
|
||||
- ../../patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -120,6 +126,8 @@ services:
|
||||
# hardware where Cliff 2 GDN activation spikes occur at runtime
|
||||
# (~50-65K active context tokens). Costs ~20-30% TPS in exchange for
|
||||
# stability. See docs/HARDWARE.md "Note for WSL2 / Windows users".
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -148,8 +148,14 @@ services:
|
||||
# workspace_lock_disable — relaxes vllm#39226 strict assertion until
|
||||
# Sandermage's P98 marker fix lands. v0.20-only requirement.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -275,6 +281,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -119,8 +119,14 @@ services:
|
||||
# workspace_lock_disable — relaxes vllm#39226 strict assertion until
|
||||
# Sandermage's P98 marker fix lands. v0.20-only requirement.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -206,6 +212,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -136,8 +136,14 @@ services:
|
||||
# surface natively (PN12 native + PN25 + PN17 + P15B). See migration
|
||||
# commit history if you need to resurrect.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -300,6 +306,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -146,8 +146,14 @@ services:
|
||||
# surface natively (PN12 native + PN25 + PN17 + P15B). See migration
|
||||
# commit history if you need to resurrect.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -317,6 +323,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -114,8 +114,14 @@ services:
|
||||
# workspace_lock_disable — relaxes vllm#39226 strict assertion until
|
||||
# Sandermage's P98 marker fix lands. v0.20-only requirement.
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -241,6 +247,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -46,8 +46,14 @@ services:
|
||||
- ../../cache/torch_compile:/root/.cache/vllm/torch_compile_cache
|
||||
- ../../cache/triton:/root/.triton/cache
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -82,6 +88,8 @@ services:
|
||||
- |
|
||||
# VLLM_ENFORCE_EAGER=1 in .env disables CUDA graphs — use on
|
||||
# hardware where graph capture causes OOM or instability (e.g. WSL2).
|
||||
# Install PR #35936 overlay before vllm imports (drop when upstream lands).
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
|
||||
- --
|
||||
command:
|
||||
|
||||
@@ -51,8 +51,14 @@ services:
|
||||
- ../../patches/genesis/vllm/_genesis:/usr/local/lib/python3.12/dist-packages/vllm/_genesis:ro
|
||||
- ../../patches/local/qwen3coder_tool_parser_deferred_commit.py:/patches/qwen3coder_tool_parser_deferred_commit.py:ro
|
||||
# vLLM PR #35936 required-tool fallback (drop when upstream lands).
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py:ro
|
||||
# vLLM PR #35936 required-tool fallback — sidecar pattern (drop when upstream lands).
|
||||
# Bind-mounted at side paths so install.sh can copy into vLLM's site-packages
|
||||
# BEFORE Genesis runs — avoids the RO-mount conflict with Genesis P64/P68/P69
|
||||
# which write to chat_completion/serving.py at vllm-import time. See
|
||||
# patches/vllm-pr35936-required-fallback/install.sh + README.md.
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/chat_completion/serving.py:/etc/club3090/pr35936-chat-completion-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/vllm/entrypoints/openai/engine/serving.py:/etc/club3090/pr35936-engine-serving.py:ro
|
||||
- ../../patches/vllm-pr35936-required-fallback/install.sh:/etc/club3090/install-pr35936.sh:ro
|
||||
# froggeric/Qwen-Fixed-Chat-Templates qwen3.6 — fixes 7 default-template
|
||||
# bugs (empty <think></think> spam, </thinking> hallucination, unclosed
|
||||
# think before tool call, no-user-query crash, developer role, etc.).
|
||||
@@ -130,6 +136,9 @@ services:
|
||||
echo " bash scripts/setup.sh qwen3.6-27b" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Install PR #35936 overlay BEFORE Genesis runs so Genesis can write hooks
|
||||
# to chat_completion/serving.py without hitting RO-mount errors.
|
||||
bash /etc/club3090/install-pr35936.sh
|
||||
python3 -m vllm._genesis.patches.apply_all
|
||||
# Tool-parser deferred-commit fix for qwen3coder SSE-silence bug (issue #72).
|
||||
# Drops out when vllm-project/vllm lands the upstream fix.
|
||||
|
||||
@@ -49,10 +49,48 @@ This overlay therefore lands the PR's intent in **both** places in
|
||||
as before so XML-emitting paths keep working.
|
||||
|
||||
The chat completion file is vendored unchanged from the pinned nightly as
|
||||
a matching drop-in mount target. The PR's streaming-side hunks target a
|
||||
pre-parser-manager code path that has been refactored on `nightly-1acd67a79`;
|
||||
non-streaming clients (MLS-Bench, our curl repro, most agent harnesses)
|
||||
exercise only the engine-side fix.
|
||||
a future-ready slot. The PR's streaming-side hunks target a pre-parser-manager
|
||||
code path that has been refactored on `nightly-1acd67a79`; non-streaming
|
||||
clients (MLS-Bench, our curl repro, most agent harnesses) exercise only the
|
||||
engine-side fix.
|
||||
|
||||
## Installation: sidecar pattern (v0.5.1+)
|
||||
|
||||
The overlay is **not** mounted directly at vLLM's site-packages paths. Instead:
|
||||
|
||||
1. Both files are bind-mounted at side paths under `/etc/club3090/`:
|
||||
- `pr35936-chat-completion-serving.py`
|
||||
- `pr35936-engine-serving.py`
|
||||
2. `install.sh` is bind-mounted at `/etc/club3090/install-pr35936.sh`
|
||||
3. The compose entrypoint invokes `bash /etc/club3090/install-pr35936.sh`
|
||||
**before** any other patch step (`python3 -m vllm._genesis.patches.apply_all`
|
||||
in Genesis-loaded composes, or `exec vllm serve` in Genesis-less ones).
|
||||
4. `install.sh` copies our files into vLLM's site-packages with `cp`. The
|
||||
destination becomes a writable file in the container's RW layer.
|
||||
|
||||
### Why a sidecar instead of an RO bind-mount
|
||||
|
||||
v0.5.0 originally mounted both files directly at vLLM's site-packages paths
|
||||
with `:ro`. That broke 8 Genesis-loaded composes because Genesis P64
|
||||
(qwen3coder MTP streaming early-return), P68 (auto force tool_choice=required),
|
||||
and P69 (long-context tool-format reminder) all write hooks into
|
||||
`chat_completion/serving.py` at vllm-import time. The RO mount blocked those
|
||||
writes with `Errno 30: Read-only file system`, and Genesis explicitly warned
|
||||
"partial state risk; container should be torn down." Reported by @ygafarov in
|
||||
[#120](https://github.com/noonghunna/club-3090/issues/120#issuecomment-4443236686);
|
||||
sidecar pattern shipped in v0.5.1.
|
||||
|
||||
The sidecar resolves it cleanly:
|
||||
- Our patched files land in the container's RW layer (not bind-mounted RO)
|
||||
- Genesis can write its hooks on top freely
|
||||
- Host patches dir stays canonical (no Genesis hooks bleeding into our git tree)
|
||||
- Each container restart starts fresh; Genesis re-applies on a clean copy
|
||||
|
||||
If a future patch to our overlay also touches `chat_completion/serving.py`
|
||||
(e.g. when PR #35936's streaming hunks land for a nightly we pin to), the
|
||||
sidecar layout supports it without further changes — our patched content
|
||||
sits on disk in the install.sh source path, gets installed before Genesis
|
||||
runs, Genesis layers its hooks on top.
|
||||
|
||||
## Validation
|
||||
|
||||
@@ -68,6 +106,12 @@ End-to-end checked 2026-05-12:
|
||||
agent completes loop with non-zero steps (was "No action returned after
|
||||
3 attempts" pre-overlay).
|
||||
|
||||
Sidecar pattern validated 2026-05-13 on `single/long-text.yml` (Genesis-loaded):
|
||||
- `install.sh` runs before Genesis: "chat_completion/serving.py installed from /etc/club3090/..."
|
||||
- Genesis P64 succeeds: "P64 applied: 2 files modified, 0 idempotent" (was failing with `Read-only file system` pre-v0.5.1)
|
||||
- Genesis P68/P69 succeed: "Hook injected into create_chat_completion" (was failing pre-v0.5.1)
|
||||
- Zero `Read-only file system` errors anywhere in the boot log.
|
||||
|
||||
## Drop Trigger
|
||||
|
||||
Remove this overlay when vLLM PR #35936, or an equivalent fix, is merged
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install PR #35936 required-tool fallback overlay into vLLM's site-packages.
|
||||
#
|
||||
# WHY a sidecar instead of an RO bind-mount:
|
||||
# Genesis P64 (qwen3coder MTP streaming early-return), P68 (auto-force
|
||||
# tool_choice=required), and P69 (long-context tool-format reminder) all
|
||||
# write hooks INTO chat_completion/serving.py at vllm-import time. An RO
|
||||
# bind-mount blocks those writes with `Errno 30: Read-only file system`,
|
||||
# and Genesis explicitly warns "partial state risk; container should be
|
||||
# torn down."
|
||||
#
|
||||
# Sidecar copies our files into the container's RW layer BEFORE Genesis
|
||||
# runs, so Genesis can write its hooks freely on top. Same pattern we
|
||||
# used for `patch_tolist_cudagraph.py` before Genesis absorbed it.
|
||||
#
|
||||
# Today chat_completion/serving.py is byte-identical to the upstream
|
||||
# nightly (PR #35936's streaming hunks don't apply on our pin), so
|
||||
# Genesis writes its hooks onto a clean copy. Slot stays future-ready
|
||||
# for when streaming hunks land — our changes will sit underneath
|
||||
# Genesis writes.
|
||||
#
|
||||
# Idempotent: cp overwrites unconditionally (per-boot fresh copy is
|
||||
# correct behaviour — Genesis re-applies on each fresh container).
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
SRC_CC="${CLUB3090_PR35936_CHAT_COMPLETION_SRC:-/etc/club3090/pr35936-chat-completion-serving.py}"
|
||||
DST_CC="/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py"
|
||||
|
||||
SRC_ENGINE="${CLUB3090_PR35936_ENGINE_SRC:-/etc/club3090/pr35936-engine-serving.py}"
|
||||
DST_ENGINE="/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/engine/serving.py"
|
||||
|
||||
if [ -r "$SRC_CC" ]; then
|
||||
cp "$SRC_CC" "$DST_CC"
|
||||
echo "[club3090/pr35936] chat_completion/serving.py installed from $SRC_CC" >&2
|
||||
else
|
||||
echo "[club3090/pr35936] WARN: $SRC_CC not found; chat_completion/serving.py left untouched" >&2
|
||||
fi
|
||||
|
||||
if [ -r "$SRC_ENGINE" ]; then
|
||||
cp "$SRC_ENGINE" "$DST_ENGINE"
|
||||
echo "[club3090/pr35936] engine/serving.py installed from $SRC_ENGINE" >&2
|
||||
else
|
||||
echo "[club3090/pr35936] WARN: $SRC_ENGINE not found; engine/serving.py left untouched (PR #35936 fix INACTIVE)" >&2
|
||||
fi
|
||||
Reference in New Issue
Block a user