Compare commits
11
Commits
v2026.05.10
...
v0.3.1
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9db8b2603c | ||
|
|
88eb67aa18 | ||
|
|
7080f1f89b | ||
|
|
e08988e614 | ||
|
|
7d91ac75e0 | ||
|
|
534d29f1b1 | ||
|
|
29d17ed82d | ||
|
|
3909c2d6b8 | ||
|
|
cc3a717524 | ||
|
|
fbf343129c | ||
|
|
cf7f1959fd |
@@ -244,3 +244,24 @@ Cross-rig data on Google's official Gemma 4 MTP "assistant" drafter (released 20
|
||||
None close the **-13% narr / -11% code gap to 3dluvr's anchor**. Remaining gap likely rig-specific (3dluvr's EPYC 7J13 / different PCIe topology / 275W cap / etc) rather than tunable via flags. Cross-rig productionizable settings: stick with default `--dtype bfloat16`, default cudagraph, default scheduling. Custom override `cudagraph_capture_sizes [9]` worth it ONLY if you serve >95% n=8-MTP-single-stream code traffic (e.g. dedicated coding-agent endpoint) where the +2.5% code lift exceeds the -3% narrative loss. |
|
||||
| `dual.yml`-shape forced TP=1 | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | 32K | **159.67 / 215.10** (decode 160.71 / 217.30) | 27.5 GB | 2026-05-07 | **First single-5090 Gemma 4 MTP data point.** First non-OOM single-card Gemma 4 result on the matrix — the 32 GB Blackwell envelope clears the 24 GB Ampere boot OOM. CV 1.9%/1.8%, peak 426 W. **+46% narr / +51% code over @noonghunna's 2× 3090 TP=2 baseline (109/142)** — single-card 5090 beats dual-3090 on Gemma 4. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
|
||||
| `dual-dflash.yml`-shape forced TP=1 (mem-util 0.96, max-model-len 12000) | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | **12K** | **150.40 / 261.06** (decode 151.16 / 264.62) | 28.8 GB | 2026-05-07 | **First single-5090 Gemma 4 DFlash data point.** Trade vs MTP row above: ~6% narr loss, **+21% code lift** (215→261). 1st-warmup TTFT outlier (73 s) suggests cudagraph warmup taking longer on first request; subsequent warmups stable at <40 ms. CV 3.6%/2.8%, peak 440 W. **Required mem-util 0.96 + max-model-len 12K** to fit BF16 weights + DFlash N=5 drafter on 32 GB — DFlash drafter footprint pushes out ctx ceiling vs MTP's 32K. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Quality benches — Aider Polyglot 30
|
||||
|
||||
Pass rate on a curated 30-exercise subset of [aider-polyglot-benchmark](https://github.com/Aider-AI/polyglot-benchmark) (5 per language across cpp/go/java/javascript/python/rust, mix of easy/medium/hard). Tests **edit-format reliability** AND **algorithmic correctness** — does the model emit diffs aider can apply, AND do the resulting tests pass.
|
||||
|
||||
Run via [`benchlocal-cli`](https://github.com/noonghunna/benchlocal-cli) `aider-polyglot-30` pack. Different from the TPS rows above — this is a quality / agentic-coding signal, not a throughput measurement.
|
||||
|
||||
| Model | Compose | Rig | Pass / Total | % | Wall (real) | Wall (sum-dur) | Tokens (P+C) | Date | Notes |
|
||||
|---|---|---|---:|---:|---:|---:|---:|---|---|
|
||||
| **Qwen 3.6 27B** (AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | **20 / 30** | **66.7%** | 19.0 min | 34.0 min | 436K + 111K = 547K | 2026-05-10 | **`enable_thinking=false`** (server-side `--default-chat-template-kwargs '{"enable_thinking": false}'` + per-request `extra_body` belt). With thinking ON: 0/30 (1500s timeout exceeded before any exercise completed — Qwen burns the token budget on hidden CoT). Per-language: cpp 3/5 · go 4/5 · **java 4/5** · js 4/5 · python 2/5 · rust 3/5. `threads=2`. |
|
||||
| **Gemma 4 31B** (Intel AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | 17 / 30 | 56.7% | 19.2 min | 19.2 min | 380K + 72K = 452K | 2026-05-10 | Default thinking off (Gemma 4's chat template requires explicit `enable_thinking=true` to enable). Per-language: cpp 2/5 · **go 4/5** · java 1/5 · **js 4/5** · python 3/5 · rust 3/5. `threads=2`. |
|
||||
|
||||
**Notable**:
|
||||
- Qwen 3.6 27B beats Gemma 4 31B by **+10 pp** despite 4 GB fewer parameters. Java is the biggest swing (4/5 vs 1/5 — `affine-cipher` specifically tripped Gemma).
|
||||
- Gemma is meaningfully faster wall-clock (sum-of-exercise-durations 19 vs 34 min), suggesting Qwen produces longer per-turn answers but they convert to passes more reliably.
|
||||
- Qwen with thinking ON is unusable for this kind of bench on club-3090 hardware: hits the 1500s subprocess timeout cap (now bumped to 2700s) before completing any exercises — the hidden CoT eats the per-exercise token budget. **Set `enable_thinking=false` for any agentic / multi-turn workload.**
|
||||
|
||||
**To re-run cross-rig**: `bash scripts/quality-test.sh --pack aider-polyglot-30 --enable-sandboxed-packs` against your endpoint. See [docs/QUALITY_TEST.md](docs/QUALITY_TEST.md) for the harness setup.
|
||||
|
||||
@@ -2,8 +2,83 @@
|
||||
|
||||
Changes that span the entire stack — engine version pins, script behavior, repo structure. Per-model dated history lives in `models/<name>/CHANGELOG.md`.
|
||||
|
||||
## Versioning
|
||||
|
||||
club-3090 ships docker composes, system scripts, and patch bundles that downstream rigs run as-is — it's effectively software, not just rolling recipes. **SemVer** in `0.x` from `v0.3.0` onward; treat any minor bump as potentially breaking until `1.0`.
|
||||
|
||||
Past CalVer tags (preserved for history):
|
||||
|
||||
| CalVer tag | SemVer equivalent | Date |
|
||||
|---|---|---|
|
||||
| `v2026.05.09` | (≈ v0.1.0) | 2026-05-09 — first tagged release |
|
||||
| `v2026.05.10` | (≈ v0.2.0) | 2026-05-10 — stack reorg + Gemma 4 INT8 PTH unblock |
|
||||
|
||||
CHANGELOG entries below 2026-05-10 use date-prose headings (pre-SemVer convention). New entries from `v0.3.0` use `## v0.X.Y (YYYY-MM-DD) — title`.
|
||||
|
||||
## v0.3.1 (2026-05-10) — soak-helper captures `delta.reasoning` (closes cross-rig silent-empty turn-5 mystery)
|
||||
|
||||
One-line patch behind a meaningful finding: vLLM nightly (`0.20.2rc1.dev9+`) emits the qwen3 reasoning parser's output under `delta.reasoning` (legacy field name), not `delta.reasoning_content` that `soak-helper.py` was watching. Result: any thinking-on response whose `<think>` block didn't close within `max_tokens` showed up as a silent-empty turn (`ttft_ms == t_ms`, `decode_tps = 0.0`, empty `content` + empty `reasoning_content`) — the model was generating correctly, the harness just couldn't see the wire.
|
||||
|
||||
This closes the cross-rig "silent-empty turn 5" pattern that was parked behind the Cliff 2b investigation. It was a harness measurement bug, not a model/rig issue.
|
||||
|
||||
Validation on JDWarner's exact repro request (#107 turn 5, math problem + `enable_thinking=true` + `max_tokens=2000`):
|
||||
|
||||
| | Before | After |
|
||||
|---|---|---|
|
||||
| `ttft_ms` | 22709 (== `t_ms`) | **234** |
|
||||
| `decode_tps` | **0.0** | **88.985** |
|
||||
| `reasoning_content` chars | 0 | 3959 |
|
||||
| `completion_tokens` | 2000 | 2000 (unchanged) |
|
||||
|
||||
Validation soak (`qwen3.6-27b dual.yml`, fresh-mode, 20 sessions × 5 turns):
|
||||
|
||||
```
|
||||
verdict PASS
|
||||
silent_empty 0 / 100 (0.0%) ← was ~3-5/40 baseline
|
||||
p50_decode_tps 90.22
|
||||
p95_ttft_ms 1389
|
||||
errors 0
|
||||
max_growth_mib 0 / 200
|
||||
```
|
||||
|
||||
**Patch** ([88eb67a](https://github.com/noonghunna/club-3090/commit/88eb67a)): `soak-helper.py` accumulates `delta.reasoning_content || delta.reasoning` (covers both vLLM streaming-output shapes, current and legacy). Three-line change inside `cmd_run`'s SSE chunk loop.
|
||||
|
||||
## v0.3.0 (2026-05-10) — Qwen thinking-off default, MODEL_DIR UX overhaul, docs cleanup, aider-polyglot bench rows
|
||||
|
||||
Eight commits since `v2026.05.10`. Headline: Qwen 3.6 27B now defaults `enable_thinking=false` server-side via `--default-chat-template-kwargs` across all 21 vLLM Qwen composes (closes the recurring "Qwen looks much slower than Gemma" agentic-bench skew). MODEL_DIR docs no longer bake the dev-rig path. New `--include-commit` flag on `power-cap-sweep.sh`.
|
||||
|
||||
**Compose change — Qwen 3.6 27B thinking OFF by default** ([534d29f](https://github.com/noonghunna/club-3090/commit/534d29f), [29d17ed](https://github.com/noonghunna/club-3090/commit/29d17ed)):
|
||||
- All 17 vLLM Qwen 27B composes (single + dual + multi4 + nvlink variants + carnice + qwopus) gain `--default-chat-template-kwargs '{"enable_thinking": false}'` after `--reasoning-parser qwen3`. Server-side default — no per-request kwarg needed.
|
||||
- llama.cpp single composes: `DISABLE_THINKING` default flipped 0 → 1; `concurrent.yml` gains `--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}`.
|
||||
- `bounded-thinking.yml` intentionally keeps thinking ON (its whole purpose is constrained CoT).
|
||||
- Per-request override still works for users who want thinking on: `chat_template_kwargs: {enable_thinking: true}` in OpenAI request body.
|
||||
- Empirical impact: aider-polyglot-30 bench went from 0/30 (1500s timeout, hidden CoT eats budget) → 20/30 (66.7%) on the same hardware once thinking defaulted off.
|
||||
- Note: 29d17ed used the wrong flag name (`--chat-template-kwargs`); 534d29f corrected to `--default-chat-template-kwargs` after vLLM nightly source confirmed the actual CLI option (referenced in upstream PR #37739).
|
||||
|
||||
**MODEL_DIR UX overhaul** (closes club-3090#116 — RobH589's "downloads landing in root of drive" report):
|
||||
- `cf7f195` — fixed 4 stale refs missed in the 2026-05-10 reorg push (preflight.sh hf-download hint, README.md example, compose header Q5_K_XL→Q3_K_XL stale comment).
|
||||
- `fbf3431` — replaced the dev-rig path `/mnt/models/huggingface/...` with `$MODEL_DIR/...` placeholders across `models/qwen3.6-27b/llama-cpp/README.md`, `docs/engines/LLAMA_CPP.md`, the compose header, and the `preflight.sh` hint. Cross-rig users no longer see our path baked into instructions.
|
||||
- `cc3a717` — recipe scripts (`single-card-default.sh`, `single-card-max-ctx.sh`) default `MODEL_PATH` to `${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/...`.
|
||||
- `3909c2d` — `setup.sh` now **prompts interactively** when `MODEL_DIR` isn't set and stdin is a TTY: pick `<repo>/models-cache`, `$HOME/models`, or custom path; optionally persist to `.env` (gitignored). CI / scripted runs (no TTY) get the silent fallback unchanged.
|
||||
|
||||
**`power-cap-sweep.sh` — `--include-commit` flag** ([7d91ac7](https://github.com/noonghunna/club-3090/commit/7d91ac7), closes [#112](https://github.com/noonghunna/club-3090/issues/112)):
|
||||
- Off by default. When set, stamps the club-3090 git short SHA in the report header next to the date. Useful for cross-rig sweep correlation per @laurimyllari's request.
|
||||
- Suppresses the field entirely (no "n/a") when run from a non-clone — closes the curl-pipe-from-docs UX hole.
|
||||
|
||||
**BENCHMARKS.md — Aider Polyglot 30 quality bench** ([e08988e](https://github.com/noonghunna/club-3090/commit/e08988e)):
|
||||
- New "Quality benches — Aider Polyglot 30" section with cross-model rows: Qwen 3.6 27B `dual.yml` 20/30 (66.7%) vs Gemma 4 31B `dual.yml` 17/30 (56.7%).
|
||||
- Documents the thinking-on/thinking-off Qwen failure mode + per-language breakdown.
|
||||
- Run via [`benchlocal-cli`](https://github.com/noonghunna/benchlocal-cli)'s `aider-polyglot-30` pack.
|
||||
|
||||
**cliff.toml + CHANGELOG meta**: switched from CalVer rolling-stack framing to SemVer software framing. Past CalVer tags preserved for history.
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-10 — Stack reorg: services consolidation, gpu-mode under git, ComfyUI in services/, pin tracker 🧹
|
||||
|
||||
(This entry covers `v2026.05.10` — see SemVer mapping above.)
|
||||
|
||||
|
||||
The rig had grown organically across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`, `/home/wasif/`), three repos (single-3090, dual-3090, club-3090), and ad-hoc paths under `/opt/ai/vllm-src/`, `/opt/ai/vendor/`, `/home/wasif/lucebox-hub/`, `/home/wasif/llama.cpp*`. Disk-out on `/` (97% used) forced a clean-up; rather than just prune, we consolidated the layout while we were at it.
|
||||
|
||||
**What landed in this repo:**
|
||||
|
||||
+1
-1
@@ -34,7 +34,7 @@ body = """
|
||||
|
||||
## Pinning to this release
|
||||
|
||||
This is a snapshot of the rolling stack — not a versioned API. To pin to this exact state:
|
||||
**Versioning:** SemVer in `0.x` — treat any minor bump as potentially breaking until `1.0`. Past CalVer tags (`v2026.05.09`, `v2026.05.10`) are preserved for history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
```bash
|
||||
git checkout {{ version }}
|
||||
|
||||
+10
-10
@@ -86,7 +86,7 @@ This is exactly why our launch frame is **two routes, not one** ([README](../../
|
||||
|
||||
```bash
|
||||
# Use hf CLI (pip install 'huggingface-hub[hf_transfer]')
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
```
|
||||
|
||||
Confirm size matches the HuggingFace listing. If a `sha256` is published, verify it.
|
||||
@@ -108,7 +108,7 @@ For a sane mid-context default (65K, plenty for chat + light agent work):
|
||||
|
||||
```bash
|
||||
/opt/llama.cpp/build/bin/llama-server \
|
||||
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-c 65536 \
|
||||
--host 0.0.0.0 --port 8020 \
|
||||
-ngl 999 \
|
||||
@@ -131,7 +131,7 @@ Recipe (community-reported, validated by multiple users on r/LocalLLaMA):
|
||||
|
||||
```bash
|
||||
/opt/llama.cpp/build/bin/llama-server \
|
||||
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-ngl 99 \
|
||||
-c 262144 \
|
||||
-np 1 \
|
||||
@@ -154,12 +154,12 @@ Sustained throughput at 262K with this config is typically **35-45 tok/s** on a
|
||||
|
||||
Download the `mmproj` model:
|
||||
```bash
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
```
|
||||
|
||||
Add to launch:
|
||||
```bash
|
||||
--mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
--mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
```
|
||||
|
||||
### 5. Tool calls (limited)
|
||||
@@ -185,12 +185,12 @@ cmake -B build -DGGML_CUDA=ON
|
||||
cmake --build build --config Release -j
|
||||
|
||||
# Download draft model (~500 MB)
|
||||
hf download z-lab/Qwen3.6-27B-DFlash --local-dir /mnt/models/huggingface/z-lab/Qwen3.6-27B-DFlash/
|
||||
hf download z-lab/Qwen3.6-27B-DFlash --local-dir $MODEL_DIR/z-lab/Qwen3.6-27B-DFlash/
|
||||
|
||||
# Launch
|
||||
/opt/lucebox-hub/build/bin/llama-server \
|
||||
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
--draft /mnt/models/gguf/qwen3.6-27b-dflash/dflash-N5.gguf \
|
||||
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
--draft $MODEL_DIR/qwen3.6-27b-dflash-gguf/dflash-N5.gguf \
|
||||
--draft-max 5 \
|
||||
--draft-min 1 \
|
||||
-c 65536 \
|
||||
@@ -210,8 +210,8 @@ If you have two GPUs (e.g. 2× 3090), lucebox-hub now supports a heterogeneous-s
|
||||
```bash
|
||||
# Target on GPU 0, DFlash draft on GPU 1
|
||||
/opt/lucebox-hub/build/bin/llama-server \
|
||||
-m /mnt/models/gguf/qwen3.5-27b/Qwen3.5-27B-Q4_K_M.gguf \
|
||||
--draft /mnt/models/gguf/qwen3.5-27b-dflash/dflash-N5.gguf \
|
||||
-m $MODEL_DIR/qwen3.5-27b-gguf/Qwen3.5-27B-Q4_K_M.gguf \
|
||||
--draft $MODEL_DIR/qwen3.5-27b-dflash-gguf/dflash-N5.gguf \
|
||||
--target-gpu 0 --draft-gpu 1 \
|
||||
--draft-max 16 --draft-min 1 \
|
||||
-c 262144 \
|
||||
|
||||
@@ -30,7 +30,7 @@ Showcase: full **262K context** on one 3090 with vision + q4_0 KV.
|
||||
|
||||
```bash
|
||||
cd models/qwen3.6-27b/llama-cpp/compose
|
||||
MODEL_DIR=/mnt/models/gguf docker compose up -d
|
||||
MODEL_DIR=/your/models/dir docker compose up -d
|
||||
```
|
||||
|
||||
Memory budget: 14.5 GB (Q3_K_XL) + 4.5 GB KV @ 262K + 0.8 GB mmproj ≈ 20 GB / 24 GB.
|
||||
@@ -66,7 +66,7 @@ The Q3_K_XL number at 262K is **lower than community-reported 35-45 tok/s** ([Re
|
||||
|
||||
```bash
|
||||
# 1. Get a GGUF quant (recommended: Unsloth's Q4_K_M)
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
|
||||
# 2. Build llama.cpp with CUDA support
|
||||
git clone https://github.com/ggerganov/llama.cpp /opt/llama.cpp
|
||||
@@ -99,9 +99,9 @@ GGUFs of this model are at [unsloth/Qwen3.6-27B-GGUF](https://huggingface.co/uns
|
||||
## Vision (mmproj)
|
||||
|
||||
```bash
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
|
||||
# Add to launch: --mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
# Add to launch: --mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
```
|
||||
|
||||
Vision works via the mmproj model. Sample text+image queries are OpenAI-compat.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
|
||||
# Engine: llama.cpp (NOT vLLM)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: none
|
||||
@@ -66,6 +66,7 @@ services:
|
||||
--cont-batching
|
||||
--jinja
|
||||
--reasoning-format ${REASONING_FORMAT:-none}
|
||||
--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
|
||||
# Engine: llama.cpp (NOT vLLM — different engine, different memory model)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: none (vanilla llama.cpp; MTP via PR #22673 not adopted yet)
|
||||
@@ -43,15 +43,15 @@
|
||||
# 1. Get the GGUF + mmproj:
|
||||
# hf download unsloth/Qwen3.6-27B-GGUF \
|
||||
# --include "Qwen3.6-27B-UD-Q3_K_XL.gguf" "mmproj-F16.gguf" \
|
||||
# --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/unsloth-q3kxl
|
||||
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-q3kxl
|
||||
# (mmproj sometimes ships at the repo root rather than under unsloth-q3kxl;
|
||||
# the MMPROJ env var below points at the canonical sidecar location.)
|
||||
# 2. From this directory:
|
||||
# MODEL_DIR=/mnt/models/huggingface docker compose up -d
|
||||
# MODEL_DIR=/your/models/dir docker compose up -d
|
||||
# 3. curl http://localhost:8020/v1/models → should list the model.
|
||||
#
|
||||
# Override defaults via .env or shell:
|
||||
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /mnt/models/huggingface on this stack)
|
||||
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /path/to/your/models on your stack)
|
||||
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf)
|
||||
# MMPROJ_FILE path under /models (default: qwen3.6-27b-gguf/mmproj-F16.gguf)
|
||||
# CTX_SIZE total KV pool (default: 262144)
|
||||
@@ -60,7 +60,7 @@
|
||||
# Override to `auto` to get separate `reasoning_content` field
|
||||
# (Qwen3.6 thinking trace) — useful for clients that render
|
||||
# reasoning_content (most don't). Issue: club-3090#97.
|
||||
# DISABLE_THINKING set to 1 in .env to add `--chat-template-kwargs '{"enable_thinking":false}'`
|
||||
# DISABLE_THINKING default: 1 (thinking OFF). Set to 0 to opt INTO thinking
|
||||
# which forces empty <think></think> blocks in responses. Useful for
|
||||
# clients (e.g. opencode) that display <think> content as the response.
|
||||
# Tradeoff: applies to ALL clients on this server — Hermes/agents that
|
||||
@@ -92,7 +92,7 @@ services:
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
# DISABLE_THINKING=1 in compose/.env appends --chat-template-kwargs to disable
|
||||
# DISABLE_THINKING=1 (default) appends --chat-template-kwargs to disable
|
||||
# Qwen3 thinking server-side. Forces the chat template to insert empty
|
||||
# <think></think> blocks → output goes straight to the response. Useful for
|
||||
# clients (e.g. opencode) that display <think> content as the response.
|
||||
@@ -100,7 +100,7 @@ services:
|
||||
# use thinking lose reasoning capability. See docs/HARDWARE.md and disc club-3090#97.
|
||||
# Note: $$VAR is YAML-escape for $VAR (compose passes literal $ to bash).
|
||||
EXTRA_ARGS=()
|
||||
if [ "$${DISABLE_THINKING:-0}" = "1" ]; then
|
||||
if [ "$${DISABLE_THINKING:-1}" = "1" ]; then
|
||||
EXTRA_ARGS+=("--chat-template-kwargs" '{"enable_thinking":false}')
|
||||
echo "[entrypoint] DISABLE_THINKING=1 — chat template will produce empty <think></think>"
|
||||
fi
|
||||
|
||||
@@ -4,15 +4,15 @@
|
||||
#
|
||||
# Prereqs:
|
||||
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
|
||||
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/)
|
||||
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir ${MODEL_DIR}/qwen3.6-27b-gguf/)
|
||||
#
|
||||
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
|
||||
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
PORT="${PORT:-8020}"
|
||||
CTX="${CTX:-65536}"
|
||||
|
||||
|
||||
@@ -18,14 +18,14 @@
|
||||
#
|
||||
# Prereqs:
|
||||
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
|
||||
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
#
|
||||
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX, KV_TYPE
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
|
||||
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
PORT="${PORT:-8020}"
|
||||
CTX="${CTX:-262144}"
|
||||
KV_TYPE="${KV_TYPE:-q4_0}"
|
||||
|
||||
@@ -127,6 +127,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_xml
|
||||
|
||||
@@ -118,6 +118,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -141,6 +141,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -130,6 +130,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -154,6 +154,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -149,6 +149,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -257,6 +257,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -139,6 +139,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -237,6 +237,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -134,6 +134,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -138,6 +138,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -247,6 +247,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -357,6 +357,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -374,6 +374,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -274,6 +274,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -97,6 +97,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -157,6 +157,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
# sudo bash scripts/power-cap-sweep.sh --load-mode decode-concurrent --concurrency 8 --bench-runs 3
|
||||
# sudo bash scripts/power-cap-sweep.sh --load-mode prefill-heavy
|
||||
# sudo bash scripts/power-cap-sweep.sh --no-reset # leave at last cap (you reset manually)
|
||||
# sudo bash scripts/power-cap-sweep.sh --include-commit # stamp club-3090 git short SHA in report header
|
||||
#
|
||||
# Load modes:
|
||||
# decode-single:
|
||||
@@ -157,6 +158,10 @@ PREFILL_FILLER_REPEATS=""
|
||||
PREFILL_PROMPT_TOKENS=""
|
||||
DECODE_CONCURRENT_RUN_SECONDS=""
|
||||
CALIBRATION_NOTE=""
|
||||
INCLUDE_COMMIT=0 # --include-commit stamps the club-3090 git short SHA in
|
||||
# the report header. Off by default — `curl ... | bash`
|
||||
# users have no clone, and stamping "n/a" is confusing
|
||||
# (better to suppress the field entirely there).
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
@@ -171,6 +176,7 @@ while [ $# -gt 0 ]; do
|
||||
--load-target) LOAD_TARGET="$2"; shift 2 ;;
|
||||
--concurrency-stretch) CONCURRENCY_STRETCH="$2"; shift 2 ;;
|
||||
--target-cap-seconds) TARGET_CAP_SECONDS="$2"; shift 2 ;;
|
||||
--include-commit) INCLUDE_COMMIT=1; shift ;;
|
||||
--no-reset) RESET=0; shift ;;
|
||||
-h|--help)
|
||||
sed -n '1,/^set -euo/p' "$0" | grep '^#' | sed 's/^# \?//'
|
||||
@@ -1120,7 +1126,17 @@ RESULTS_FILE=/tmp/power-cap-summary.md
|
||||
echo "**Model:** \`${MODEL}\` **Engine:** \`${CONTAINER}\` **Endpoint:** ${URL}"
|
||||
echo "**Load mode:** \`${LOAD_MODE}\`$([ "$LOAD_MODE" = "decode-single" ] && echo " (${TARGET_CAP_SECONDS}s × 2 timed streams)")$([ "$LOAD_MODE" = "decode-concurrent" ] && echo " (concurrency=${CONCURRENCY}, ${DECODE_CONCURRENT_RUN_SECONDS}s/run × ${BENCH_RUNS} runs × 2 timed batches)")$([ "$LOAD_MODE" = "prefill-heavy" ] && echo " (target-prefill=${TARGET_PREFILL_SECONDS}s, filler_repeats=${PREFILL_FILLER_REPEATS})")$([ "$LOAD_MODE" != "decode-single" ] && echo " (bench-runs=${BENCH_RUNS})")"
|
||||
[ -n "$CALIBRATION_NOTE" ] && echo "**Calibration:** ${CALIBRATION_NOTE}"
|
||||
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z"
|
||||
# --include-commit: stamp club-3090 git short SHA next to the date if requested.
|
||||
# Suppress entirely (rather than show "n/a") when run from a non-clone or
|
||||
# when git isn't reachable — closes the curl-pipe-from-docs UX hole.
|
||||
COMMIT_FRAGMENT=""
|
||||
if [ "${INCLUDE_COMMIT:-0}" = "1" ]; then
|
||||
COMMIT_SHA=$(git -C "$REPO_ROOT" rev-parse --short HEAD 2>/dev/null || true)
|
||||
if [ -n "$COMMIT_SHA" ]; then
|
||||
COMMIT_FRAGMENT=" **club-3090 commit:** \`${COMMIT_SHA}\`"
|
||||
fi
|
||||
fi
|
||||
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z${COMMIT_FRAGMENT}"
|
||||
echo ""
|
||||
if [ "$COOLING" = "unspecified" ]; then
|
||||
echo "> ⚠️ Cooling class not specified at run time. Add **air / water / AIO** when posting"
|
||||
|
||||
@@ -325,8 +325,8 @@ preflight_compose_deps() {
|
||||
mmproj_in_container="${mmproj_in_container//\$\{MMPROJ_FILE:-/}"
|
||||
mmproj_in_container="${mmproj_in_container%\}}"
|
||||
|
||||
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
|
||||
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b/mmproj-F16.gguf"
|
||||
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
|
||||
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b-gguf/mmproj-F16.gguf"
|
||||
|
||||
if [[ -n "${GGUF_FILE:-}" ]]; then gguf_in_container="$GGUF_FILE"; fi
|
||||
if [[ -n "${MMPROJ_FILE:-}" ]]; then mmproj_in_container="$MMPROJ_FILE"; fi
|
||||
@@ -375,9 +375,10 @@ preflight_compose_deps() {
|
||||
if [[ $hint_gguf -eq 1 ]]; then
|
||||
echo "[preflight] hf download unsloth/Qwen3.6-27B-GGUF \\" >&2
|
||||
echo "[preflight] Qwen3.6-27B-UD-Q3_K_XL.gguf mmproj-F16.gguf \\" >&2
|
||||
echo "[preflight] --local-dir ${model_dir}/qwen3.6-27b/unsloth-q3kxl" >&2
|
||||
echo "[preflight] --local-dir \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl" >&2
|
||||
echo "[preflight] # (set MODEL_DIR first: export MODEL_DIR=\${MODEL_DIR:-/path/to/your/models})" >&2
|
||||
echo "[preflight] # mmproj lands at unsloth-q3kxl/ — move it up so the default --mmproj path resolves:" >&2
|
||||
echo "[preflight] # mv ${model_dir}/qwen3.6-27b/unsloth-q3kxl/mmproj-F16.gguf ${model_dir}/qwen3.6-27b/" >&2
|
||||
echo "[preflight] # mv \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl/mmproj-F16.gguf \${MODEL_DIR}/qwen3.6-27b-gguf/" >&2
|
||||
echo "[preflight] (~16 GB total. setup.sh today only fetches the vLLM AutoRound weights;" >&2
|
||||
echo "[preflight] GGUF must be fetched separately for any llamacpp/* variant.)" >&2
|
||||
fi
|
||||
|
||||
@@ -90,6 +90,73 @@ case "${MODEL_NAME}" in
|
||||
esac
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
# ---------- MODEL_DIR resolution ----------
|
||||
# Order of precedence:
|
||||
# 1. MODEL_DIR already exported in the calling shell → use as-is
|
||||
# 2. .env at repo root sets MODEL_DIR → source it
|
||||
# 3. Interactive prompt (only if stdin is a TTY) → ask user
|
||||
# 4. Silent fallback to <repo>/models-cache → in-repo default
|
||||
#
|
||||
# The prompt only fires for fresh users on a TTY who haven't set anything.
|
||||
# CI / scripted runs (no TTY) get the silent fallback, preserving prior behavior.
|
||||
|
||||
# Step 2: source repo-root .env if present (lets a saved choice persist)
|
||||
if [[ -z "${MODEL_DIR:-}" && -f "${ROOT_DIR}/.env" ]]; then
|
||||
# shellcheck source=/dev/null
|
||||
set -a; source "${ROOT_DIR}/.env"; set +a
|
||||
fi
|
||||
|
||||
# Step 3: prompt if still unset + interactive
|
||||
if [[ -z "${MODEL_DIR:-}" && -t 0 && -t 1 ]]; then
|
||||
echo ""
|
||||
echo "Where should I put model weights?"
|
||||
echo " Models are large (Qwen3.6-27B AutoRound: ~14 GB; Gemma 4 31B: ~21 GB)."
|
||||
echo " This dir lives outside the git tree — pick a location with sufficient free space."
|
||||
echo ""
|
||||
echo " 1) ${ROOT_DIR}/models-cache (in-repo, default — pollutes git tree)"
|
||||
echo " 2) ${HOME}/models (recommended for cross-rig — outside repo)"
|
||||
echo " 3) custom path"
|
||||
echo ""
|
||||
while true; do
|
||||
read -rp "Choice [1-3] (or set MODEL_DIR env var to skip): " pick
|
||||
case "${pick}" in
|
||||
1) MODEL_DIR="${ROOT_DIR}/models-cache"; break ;;
|
||||
2) MODEL_DIR="${HOME}/models"; break ;;
|
||||
3)
|
||||
read -rp " Enter absolute path: " custom
|
||||
if [[ "${custom}" =~ ^/ ]]; then
|
||||
MODEL_DIR="${custom}"; break
|
||||
else
|
||||
echo " ! must be an absolute path (start with /)" >&2
|
||||
fi
|
||||
;;
|
||||
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
|
||||
esac
|
||||
done
|
||||
echo ""
|
||||
|
||||
# Offer to persist the choice so future runs skip the prompt
|
||||
read -rp "Save MODEL_DIR=${MODEL_DIR} to .env so we skip this next time? [Y/n]: " save
|
||||
if [[ "${save:-y}" =~ ^[Yy]$ || -z "${save:-}" ]]; then
|
||||
if [[ -f "${ROOT_DIR}/.env" ]]; then
|
||||
# Update existing .env (replace MODEL_DIR= line if present, else append)
|
||||
if grep -qE "^MODEL_DIR=" "${ROOT_DIR}/.env"; then
|
||||
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=${MODEL_DIR}|" "${ROOT_DIR}/.env"
|
||||
else
|
||||
echo "MODEL_DIR=${MODEL_DIR}" >> "${ROOT_DIR}/.env"
|
||||
fi
|
||||
else
|
||||
echo "MODEL_DIR=${MODEL_DIR}" > "${ROOT_DIR}/.env"
|
||||
fi
|
||||
echo " → saved. (.env is gitignored.)"
|
||||
else
|
||||
echo " → not saved. Set MODEL_DIR=... when re-running, or you'll get this prompt again."
|
||||
fi
|
||||
echo ""
|
||||
fi
|
||||
|
||||
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
|
||||
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
|
||||
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
|
||||
|
||||
|
||||
@@ -557,15 +557,21 @@ def cmd_run(endpoint, req_path, timeout_s, metrics_path):
|
||||
choices = chunk.get("choices") or []
|
||||
if choices:
|
||||
delta = choices[0].get("delta") or {}
|
||||
if ttft is None and (delta.get("content") or delta.get("reasoning_content") or delta.get("tool_calls")):
|
||||
# vLLM emits reasoning under either `delta.reasoning_content`
|
||||
# (older qwen3 reasoner path) or `delta.reasoning` (current
|
||||
# nightly as of vllm-0.20.2rc1+; legacy field name). Watch
|
||||
# both so the soak harness doesn't go silent when the
|
||||
# underlying field name shifts under us.
|
||||
reasoning_delta = delta.get("reasoning_content") or delta.get("reasoning")
|
||||
if ttft is None and (delta.get("content") or reasoning_delta or delta.get("tool_calls")):
|
||||
ttft = time.time() - t0
|
||||
# Accumulate streamed parts. vLLM splits content/reasoning
|
||||
# across many small deltas; tool_calls stream as indexed
|
||||
# objects whose fields (name, arguments) arrive in pieces.
|
||||
if delta.get("content"):
|
||||
content_parts.append(delta["content"])
|
||||
if delta.get("reasoning_content"):
|
||||
reasoning_parts.append(delta["reasoning_content"])
|
||||
if reasoning_delta:
|
||||
reasoning_parts.append(reasoning_delta)
|
||||
for tc in (delta.get("tool_calls") or []):
|
||||
idx = tc.get("index", 0)
|
||||
slot = tool_calls_acc.setdefault(idx, {"id": "", "type": "function", "name": "", "args": ""})
|
||||
|
||||
Reference in New Issue
Block a user