Compare commits
14
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
255c743dff | ||
|
|
64b0474a62 | ||
|
|
83bf73d3ec | ||
|
|
9db8b2603c | ||
|
|
88eb67aa18 | ||
|
|
7080f1f89b | ||
|
|
e08988e614 | ||
|
|
7d91ac75e0 | ||
|
|
534d29f1b1 | ||
|
|
29d17ed82d | ||
|
|
3909c2d6b8 | ||
|
|
cc3a717524 | ||
|
|
fbf343129c | ||
|
|
cf7f1959fd |
@@ -16,17 +16,54 @@ jobs:
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0 # full history needed for git-cliff
|
||||
token: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
- name: Generate release notes
|
||||
id: cliff
|
||||
# ---- (1) Generate the GitHub Release body (latest only, no header) ----
|
||||
# `--strip header` drops the static SemVer preamble so the release page
|
||||
# shows just the per-version section, while CHANGELOG.md (below) keeps
|
||||
# the header at the top of the file.
|
||||
- name: Generate release notes (latest only)
|
||||
id: cliff-release
|
||||
uses: orhun/git-cliff-action@v4
|
||||
with:
|
||||
config: cliff.toml
|
||||
args: --latest --github-repo noonghunna/club-3090
|
||||
args: --latest --strip header --github-repo noonghunna/club-3090
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
OUTPUT: RELEASE_NOTES.md
|
||||
|
||||
# ---- (2) Regenerate full CHANGELOG.md (all tags, with header) ----
|
||||
- name: Regenerate CHANGELOG.md
|
||||
uses: orhun/git-cliff-action@v4
|
||||
with:
|
||||
config: cliff.toml
|
||||
args: --github-repo noonghunna/club-3090
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
OUTPUT: CHANGELOG.md
|
||||
|
||||
# ---- (3) Commit the regenerated CHANGELOG.md back to master ----
|
||||
# The tag's commit doesn't include this auto-regen — but the GitHub
|
||||
# Release page is correct (step 1), and master's CHANGELOG.md catches
|
||||
# up ~1 min after tag push. Skipped if nothing changed.
|
||||
- name: Commit CHANGELOG.md back to master
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
|
||||
git fetch origin master
|
||||
git checkout -B master origin/master
|
||||
# Re-run cliff against master HEAD so the regen reflects what master
|
||||
# actually contains (the tag may not yet be on master if pushed from
|
||||
# a feature branch; uncommon but handled).
|
||||
git add CHANGELOG.md
|
||||
if git diff --staged --quiet; then
|
||||
echo "CHANGELOG.md unchanged — skipping commit."
|
||||
else
|
||||
git commit -m "chore(changelog): regenerate for ${{ github.ref_name }} [skip ci]"
|
||||
git push origin master
|
||||
fi
|
||||
|
||||
# ---- (4) Publish GitHub Release ----
|
||||
- name: Create GitHub Release
|
||||
uses: softprops/action-gh-release@v2
|
||||
with:
|
||||
|
||||
@@ -244,3 +244,24 @@ Cross-rig data on Google's official Gemma 4 MTP "assistant" drafter (released 20
|
||||
None close the **-13% narr / -11% code gap to 3dluvr's anchor**. Remaining gap likely rig-specific (3dluvr's EPYC 7J13 / different PCIe topology / 275W cap / etc) rather than tunable via flags. Cross-rig productionizable settings: stick with default `--dtype bfloat16`, default cudagraph, default scheduling. Custom override `cudagraph_capture_sizes [9]` worth it ONLY if you serve >95% n=8-MTP-single-stream code traffic (e.g. dedicated coding-agent endpoint) where the +2.5% code lift exceeds the -3% narrative loss. |
|
||||
| `dual.yml`-shape forced TP=1 | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | 32K | **159.67 / 215.10** (decode 160.71 / 217.30) | 27.5 GB | 2026-05-07 | **First single-5090 Gemma 4 MTP data point.** First non-OOM single-card Gemma 4 result on the matrix — the 32 GB Blackwell envelope clears the 24 GB Ampere boot OOM. CV 1.9%/1.8%, peak 426 W. **+46% narr / +51% code over @noonghunna's 2× 3090 TP=2 baseline (109/142)** — single-card 5090 beats dual-3090 on Gemma 4. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
|
||||
| `dual-dflash.yml`-shape forced TP=1 (mem-util 0.96, max-model-len 12000) | @apnar (1× **RTX 5090** 32 GB, air-cooled, 600 W) | bf16 | **12K** | **150.40 / 261.06** (decode 151.16 / 264.62) | 28.8 GB | 2026-05-07 | **First single-5090 Gemma 4 DFlash data point.** Trade vs MTP row above: ~6% narr loss, **+21% code lift** (215→261). 1st-warmup TTFT outlier (73 s) suggests cudagraph warmup taking longer on first request; subsequent warmups stable at <40 ms. CV 3.6%/2.8%, peak 440 W. **Required mem-util 0.96 + max-model-len 12K** to fit BF16 weights + DFlash N=5 drafter on 32 GB — DFlash drafter footprint pushes out ctx ceiling vs MTP's 32K. [Disc #67](https://github.com/noonghunna/club-3090/discussions/67#discussioncomment-16832042). |
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Quality benches — Aider Polyglot 30
|
||||
|
||||
Pass rate on a curated 30-exercise subset of [aider-polyglot-benchmark](https://github.com/Aider-AI/polyglot-benchmark) (5 per language across cpp/go/java/javascript/python/rust, mix of easy/medium/hard). Tests **edit-format reliability** AND **algorithmic correctness** — does the model emit diffs aider can apply, AND do the resulting tests pass.
|
||||
|
||||
Run via [`benchlocal-cli`](https://github.com/noonghunna/benchlocal-cli) `aider-polyglot-30` pack. Different from the TPS rows above — this is a quality / agentic-coding signal, not a throughput measurement.
|
||||
|
||||
| Model | Compose | Rig | Pass / Total | % | Wall (real) | Wall (sum-dur) | Tokens (P+C) | Date | Notes |
|
||||
|---|---|---|---:|---:|---:|---:|---:|---|---|
|
||||
| **Qwen 3.6 27B** (AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | **20 / 30** | **66.7%** | 19.0 min | 34.0 min | 436K + 111K = 547K | 2026-05-10 | **`enable_thinking=false`** (server-side `--default-chat-template-kwargs '{"enable_thinking": false}'` + per-request `extra_body` belt). With thinking ON: 0/30 (1500s timeout exceeded before any exercise completed — Qwen burns the token budget on hidden CoT). Per-language: cpp 3/5 · go 4/5 · **java 4/5** · js 4/5 · python 2/5 · rust 3/5. `threads=2`. |
|
||||
| **Gemma 4 31B** (Intel AutoRound INT4) | `dual.yml` (TP=2) | @noonghunna (2× 3090 PCIe, 230W cap) | 17 / 30 | 56.7% | 19.2 min | 19.2 min | 380K + 72K = 452K | 2026-05-10 | Default thinking off (Gemma 4's chat template requires explicit `enable_thinking=true` to enable). Per-language: cpp 2/5 · **go 4/5** · java 1/5 · **js 4/5** · python 3/5 · rust 3/5. `threads=2`. |
|
||||
|
||||
**Notable**:
|
||||
- Qwen 3.6 27B beats Gemma 4 31B by **+10 pp** despite 4 GB fewer parameters. Java is the biggest swing (4/5 vs 1/5 — `affine-cipher` specifically tripped Gemma).
|
||||
- Gemma is meaningfully faster wall-clock (sum-of-exercise-durations 19 vs 34 min), suggesting Qwen produces longer per-turn answers but they convert to passes more reliably.
|
||||
- Qwen with thinking ON is unusable for this kind of bench on club-3090 hardware: hits the 1500s subprocess timeout cap (now bumped to 2700s) before completing any exercises — the hidden CoT eats the per-exercise token budget. **Set `enable_thinking=false` for any agentic / multi-turn workload.**
|
||||
|
||||
**To re-run cross-rig**: `bash scripts/quality-test.sh --pack aider-polyglot-30 --enable-sandboxed-packs` against your endpoint. See [docs/QUALITY_TEST.md](docs/QUALITY_TEST.md) for the harness setup.
|
||||
|
||||
+8747
-413
File diff suppressed because it is too large
Load Diff
+42
-20
@@ -8,15 +8,41 @@
|
||||
# → release published.
|
||||
|
||||
[changelog]
|
||||
# Header rendered once at top of every release body.
|
||||
# Header rendered once at top of CHANGELOG.md (preserved across full regens).
|
||||
# Use `--strip header` on `--latest` for GitHub Release bodies so they don't
|
||||
# duplicate this intro on every release page.
|
||||
header = """
|
||||
# Changelog
|
||||
|
||||
Auto-generated from commit messages by [git-cliff](https://git-cliff.org/).
|
||||
Update flow: write rich commit message bodies → tag → CI regenerates this file
|
||||
and the GitHub Release notes from the same source. Don't hand-edit below the
|
||||
header — your changes will be overwritten on the next tag.
|
||||
|
||||
**Versioning:** SemVer in `0.x` — treat any minor bump as potentially breaking
|
||||
until `1.0`. Past CalVer tags (`v2026.05.09`, `v2026.05.10`) are preserved for
|
||||
history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
| CalVer tag | SemVer equivalent | Date |
|
||||
|---|---|---|
|
||||
| `v2026.05.09` | (≈ v0.1.0) | 2026-05-09 — first tagged release |
|
||||
| `v2026.05.10` | (≈ v0.2.0) | 2026-05-10 — stack reorg + Gemma 4 INT8 PTH unblock |
|
||||
|
||||
---
|
||||
|
||||
"""
|
||||
|
||||
# Body template — rendered per release (we use --latest so only one).
|
||||
# Tera templating syntax. Each commit shows as a bullet with PR link if present.
|
||||
# Body template — rendered per release. For `--latest` (GitHub Release) only
|
||||
# the most recent block renders; for full regen of CHANGELOG.md, every tagged
|
||||
# block renders in reverse-chronological order.
|
||||
#
|
||||
# Commit message body (everything after subject + blank line) renders below
|
||||
# the subject bullet so rich narrative (tables, validation data, before/after
|
||||
# numbers) ends up in both CHANGELOG.md and the GitHub Release page from the
|
||||
# same source. Tera templating; see https://keats.github.io/tera/docs/.
|
||||
body = """
|
||||
{% if version %}\
|
||||
## What's in {{ version }}
|
||||
## {{ version }}{% if timestamp %} — {{ timestamp | date(format="%Y-%m-%d") }}{% endif %}
|
||||
|
||||
{% else %}\
|
||||
## Unreleased
|
||||
@@ -26,25 +52,21 @@ body = """
|
||||
### {{ group }}
|
||||
|
||||
{% for commit in commits %}\
|
||||
- {{ commit.message | split(pat="\\n") | first | trim }}{% if commit.github.pr_number %} ([#{{ commit.github.pr_number }}](https://github.com/noonghunna/club-3090/pull/{{ commit.github.pr_number }}) by @{{ commit.github.username }}){% else %} ([{{ commit.id | truncate(length=7, end="") }}](https://github.com/noonghunna/club-3090/commit/{{ commit.id }})){% endif %}
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
- **{{ commit.message | split(pat="\\n") | first | trim }}**{% if commit.github.pr_number %} ([#{{ commit.github.pr_number }}](https://github.com/noonghunna/club-3090/pull/{{ commit.github.pr_number }}) by @{{ commit.github.username }}){% else %} ([{{ commit.id | truncate(length=7, end="") }}](https://github.com/noonghunna/club-3090/commit/{{ commit.id }})){% endif %}
|
||||
{% set body_lines = commit.message | split(pat="\\n") %}\
|
||||
{% if body_lines | length > 1 %}\
|
||||
{% set body = body_lines | slice(start=1) | join(sep="\\n") | trim %}\
|
||||
{% if body %}
|
||||
|
||||
---
|
||||
{{ body | replace(from="\\n", to="\\n ") }}
|
||||
|
||||
## Pinning to this release
|
||||
|
||||
This is a snapshot of the rolling stack — not a versioned API. To pin to this exact state:
|
||||
|
||||
```bash
|
||||
git checkout {{ version }}
|
||||
```
|
||||
|
||||
When posting cross-rig benchmark numbers ([disc #86](https://github.com/noonghunna/club-3090/discussions/86)), please include this version tag (or commit SHA) so others can reproduce against the same script revision.
|
||||
|
||||
{% if previous.version %}\
|
||||
**Full diff:** [{{ previous.version }}...{{ version }}](https://github.com/noonghunna/club-3090/compare/{{ previous.version }}...{{ version }})
|
||||
{% endif %}\
|
||||
{% endif %}\
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
|
||||
{% if version %}[Pin: `git checkout {{ version }}`]{% if previous.version %} · [Full diff](https://github.com/noonghunna/club-3090/compare/{{ previous.version }}...{{ version }}){% endif %}
|
||||
{% endif %}
|
||||
"""
|
||||
|
||||
footer = ""
|
||||
|
||||
+10
-10
@@ -86,7 +86,7 @@ This is exactly why our launch frame is **two routes, not one** ([README](../../
|
||||
|
||||
```bash
|
||||
# Use hf CLI (pip install 'huggingface-hub[hf_transfer]')
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
```
|
||||
|
||||
Confirm size matches the HuggingFace listing. If a `sha256` is published, verify it.
|
||||
@@ -108,7 +108,7 @@ For a sane mid-context default (65K, plenty for chat + light agent work):
|
||||
|
||||
```bash
|
||||
/opt/llama.cpp/build/bin/llama-server \
|
||||
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-c 65536 \
|
||||
--host 0.0.0.0 --port 8020 \
|
||||
-ngl 999 \
|
||||
@@ -131,7 +131,7 @@ Recipe (community-reported, validated by multiple users on r/LocalLLaMA):
|
||||
|
||||
```bash
|
||||
/opt/llama.cpp/build/bin/llama-server \
|
||||
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
-ngl 99 \
|
||||
-c 262144 \
|
||||
-np 1 \
|
||||
@@ -154,12 +154,12 @@ Sustained throughput at 262K with this config is typically **35-45 tok/s** on a
|
||||
|
||||
Download the `mmproj` model:
|
||||
```bash
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
```
|
||||
|
||||
Add to launch:
|
||||
```bash
|
||||
--mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
--mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
```
|
||||
|
||||
### 5. Tool calls (limited)
|
||||
@@ -185,12 +185,12 @@ cmake -B build -DGGML_CUDA=ON
|
||||
cmake --build build --config Release -j
|
||||
|
||||
# Download draft model (~500 MB)
|
||||
hf download z-lab/Qwen3.6-27B-DFlash --local-dir /mnt/models/huggingface/z-lab/Qwen3.6-27B-DFlash/
|
||||
hf download z-lab/Qwen3.6-27B-DFlash --local-dir $MODEL_DIR/z-lab/Qwen3.6-27B-DFlash/
|
||||
|
||||
# Launch
|
||||
/opt/lucebox-hub/build/bin/llama-server \
|
||||
-m /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
--draft /mnt/models/gguf/qwen3.6-27b-dflash/dflash-N5.gguf \
|
||||
-m $MODEL_DIR/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf \
|
||||
--draft $MODEL_DIR/qwen3.6-27b-dflash-gguf/dflash-N5.gguf \
|
||||
--draft-max 5 \
|
||||
--draft-min 1 \
|
||||
-c 65536 \
|
||||
@@ -210,8 +210,8 @@ If you have two GPUs (e.g. 2× 3090), lucebox-hub now supports a heterogeneous-s
|
||||
```bash
|
||||
# Target on GPU 0, DFlash draft on GPU 1
|
||||
/opt/lucebox-hub/build/bin/llama-server \
|
||||
-m /mnt/models/gguf/qwen3.5-27b/Qwen3.5-27B-Q4_K_M.gguf \
|
||||
--draft /mnt/models/gguf/qwen3.5-27b-dflash/dflash-N5.gguf \
|
||||
-m $MODEL_DIR/qwen3.5-27b-gguf/Qwen3.5-27B-Q4_K_M.gguf \
|
||||
--draft $MODEL_DIR/qwen3.5-27b-dflash-gguf/dflash-N5.gguf \
|
||||
--target-gpu 0 --draft-gpu 1 \
|
||||
--draft-max 16 --draft-min 1 \
|
||||
-c 262144 \
|
||||
|
||||
@@ -30,7 +30,7 @@ Showcase: full **262K context** on one 3090 with vision + q4_0 KV.
|
||||
|
||||
```bash
|
||||
cd models/qwen3.6-27b/llama-cpp/compose
|
||||
MODEL_DIR=/mnt/models/gguf docker compose up -d
|
||||
MODEL_DIR=/your/models/dir docker compose up -d
|
||||
```
|
||||
|
||||
Memory budget: 14.5 GB (Q3_K_XL) + 4.5 GB KV @ 262K + 0.8 GB mmproj ≈ 20 GB / 24 GB.
|
||||
@@ -66,7 +66,7 @@ The Q3_K_XL number at 262K is **lower than community-reported 35-45 tok/s** ([Re
|
||||
|
||||
```bash
|
||||
# 1. Get a GGUF quant (recommended: Unsloth's Q4_K_M)
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
|
||||
# 2. Build llama.cpp with CUDA support
|
||||
git clone https://github.com/ggerganov/llama.cpp /opt/llama.cpp
|
||||
@@ -99,9 +99,9 @@ GGUFs of this model are at [unsloth/Qwen3.6-27B-GGUF](https://huggingface.co/uns
|
||||
## Vision (mmproj)
|
||||
|
||||
```bash
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/
|
||||
hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf --local-dir $MODEL_DIR/qwen3.6-27b-gguf/
|
||||
|
||||
# Add to launch: --mmproj /mnt/models/huggingface/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
# Add to launch: --mmproj $MODEL_DIR/qwen3.6-27b-gguf/mmproj-F16.gguf
|
||||
```
|
||||
|
||||
Vision works via the mmproj model. Sample text+image queries are OpenAI-compat.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
|
||||
# Engine: llama.cpp (NOT vLLM)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: none
|
||||
@@ -66,6 +66,7 @@ services:
|
||||
--cont-batching
|
||||
--jinja
|
||||
--reasoning-format ${REASONING_FORMAT:-none}
|
||||
--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth Q5_K_XL GGUF)
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
|
||||
# Engine: llama.cpp (NOT vLLM — different engine, different memory model)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: none (vanilla llama.cpp; MTP via PR #22673 not adopted yet)
|
||||
@@ -43,15 +43,15 @@
|
||||
# 1. Get the GGUF + mmproj:
|
||||
# hf download unsloth/Qwen3.6-27B-GGUF \
|
||||
# --include "Qwen3.6-27B-UD-Q3_K_XL.gguf" "mmproj-F16.gguf" \
|
||||
# --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/unsloth-q3kxl
|
||||
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-q3kxl
|
||||
# (mmproj sometimes ships at the repo root rather than under unsloth-q3kxl;
|
||||
# the MMPROJ env var below points at the canonical sidecar location.)
|
||||
# 2. From this directory:
|
||||
# MODEL_DIR=/mnt/models/huggingface docker compose up -d
|
||||
# MODEL_DIR=/your/models/dir docker compose up -d
|
||||
# 3. curl http://localhost:8020/v1/models → should list the model.
|
||||
#
|
||||
# Override defaults via .env or shell:
|
||||
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /mnt/models/huggingface on this stack)
|
||||
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /path/to/your/models on your stack)
|
||||
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf)
|
||||
# MMPROJ_FILE path under /models (default: qwen3.6-27b-gguf/mmproj-F16.gguf)
|
||||
# CTX_SIZE total KV pool (default: 262144)
|
||||
@@ -60,7 +60,7 @@
|
||||
# Override to `auto` to get separate `reasoning_content` field
|
||||
# (Qwen3.6 thinking trace) — useful for clients that render
|
||||
# reasoning_content (most don't). Issue: club-3090#97.
|
||||
# DISABLE_THINKING set to 1 in .env to add `--chat-template-kwargs '{"enable_thinking":false}'`
|
||||
# DISABLE_THINKING default: 1 (thinking OFF). Set to 0 to opt INTO thinking
|
||||
# which forces empty <think></think> blocks in responses. Useful for
|
||||
# clients (e.g. opencode) that display <think> content as the response.
|
||||
# Tradeoff: applies to ALL clients on this server — Hermes/agents that
|
||||
@@ -92,7 +92,7 @@ services:
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
# DISABLE_THINKING=1 in compose/.env appends --chat-template-kwargs to disable
|
||||
# DISABLE_THINKING=1 (default) appends --chat-template-kwargs to disable
|
||||
# Qwen3 thinking server-side. Forces the chat template to insert empty
|
||||
# <think></think> blocks → output goes straight to the response. Useful for
|
||||
# clients (e.g. opencode) that display <think> content as the response.
|
||||
@@ -100,7 +100,7 @@ services:
|
||||
# use thinking lose reasoning capability. See docs/HARDWARE.md and disc club-3090#97.
|
||||
# Note: $$VAR is YAML-escape for $VAR (compose passes literal $ to bash).
|
||||
EXTRA_ARGS=()
|
||||
if [ "$${DISABLE_THINKING:-0}" = "1" ]; then
|
||||
if [ "$${DISABLE_THINKING:-1}" = "1" ]; then
|
||||
EXTRA_ARGS+=("--chat-template-kwargs" '{"enable_thinking":false}')
|
||||
echo "[entrypoint] DISABLE_THINKING=1 — chat template will produce empty <think></think>"
|
||||
fi
|
||||
|
||||
@@ -4,15 +4,15 @@
|
||||
#
|
||||
# Prereqs:
|
||||
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
|
||||
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir /mnt/models/huggingface/qwen3.6-27b-gguf/)
|
||||
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
# (download via: hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir ${MODEL_DIR}/qwen3.6-27b-gguf/)
|
||||
#
|
||||
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
|
||||
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
PORT="${PORT:-8020}"
|
||||
CTX="${CTX:-65536}"
|
||||
|
||||
|
||||
@@ -18,14 +18,14 @@
|
||||
#
|
||||
# Prereqs:
|
||||
# - llama.cpp built with -DGGML_CUDA=ON at /opt/llama.cpp
|
||||
# - Q4_K_M GGUF at /mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
# - Q4_K_M GGUF at ${MODEL_DIR}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf
|
||||
#
|
||||
# Override defaults via env: LLAMA_DIR, MODEL_PATH, PORT, CTX, KV_TYPE
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
LLAMA_DIR="${LLAMA_DIR:-/opt/llama.cpp}"
|
||||
MODEL_PATH="${MODEL_PATH:-/mnt/models/huggingface/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
MODEL_PATH="${MODEL_PATH:-${MODEL_DIR:-$HOME/models}/qwen3.6-27b-gguf/Qwen3.6-27B-Q4_K_M.gguf}"
|
||||
PORT="${PORT:-8020}"
|
||||
CTX="${CTX:-262144}"
|
||||
KV_TYPE="${KV_TYPE:-q4_0}"
|
||||
|
||||
@@ -127,6 +127,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_xml
|
||||
|
||||
@@ -118,6 +118,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -141,6 +141,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -130,6 +130,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -154,6 +154,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -149,6 +149,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -257,6 +257,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -139,6 +139,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -237,6 +237,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -134,6 +134,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -138,6 +138,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -247,6 +247,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -357,6 +357,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -374,6 +374,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -274,6 +274,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -97,6 +97,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -157,6 +157,8 @@ services:
|
||||
- --trust-remote-code
|
||||
- --reasoning-parser
|
||||
- qwen3
|
||||
- --default-chat-template-kwargs
|
||||
- '{"enable_thinking": false}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
# sudo bash scripts/power-cap-sweep.sh --load-mode decode-concurrent --concurrency 8 --bench-runs 3
|
||||
# sudo bash scripts/power-cap-sweep.sh --load-mode prefill-heavy
|
||||
# sudo bash scripts/power-cap-sweep.sh --no-reset # leave at last cap (you reset manually)
|
||||
# sudo bash scripts/power-cap-sweep.sh --include-commit # stamp club-3090 git short SHA in report header
|
||||
#
|
||||
# Load modes:
|
||||
# decode-single:
|
||||
@@ -157,6 +158,10 @@ PREFILL_FILLER_REPEATS=""
|
||||
PREFILL_PROMPT_TOKENS=""
|
||||
DECODE_CONCURRENT_RUN_SECONDS=""
|
||||
CALIBRATION_NOTE=""
|
||||
INCLUDE_COMMIT=0 # --include-commit stamps the club-3090 git short SHA in
|
||||
# the report header. Off by default — `curl ... | bash`
|
||||
# users have no clone, and stamping "n/a" is confusing
|
||||
# (better to suppress the field entirely there).
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
@@ -171,6 +176,7 @@ while [ $# -gt 0 ]; do
|
||||
--load-target) LOAD_TARGET="$2"; shift 2 ;;
|
||||
--concurrency-stretch) CONCURRENCY_STRETCH="$2"; shift 2 ;;
|
||||
--target-cap-seconds) TARGET_CAP_SECONDS="$2"; shift 2 ;;
|
||||
--include-commit) INCLUDE_COMMIT=1; shift ;;
|
||||
--no-reset) RESET=0; shift ;;
|
||||
-h|--help)
|
||||
sed -n '1,/^set -euo/p' "$0" | grep '^#' | sed 's/^# \?//'
|
||||
@@ -1120,7 +1126,17 @@ RESULTS_FILE=/tmp/power-cap-summary.md
|
||||
echo "**Model:** \`${MODEL}\` **Engine:** \`${CONTAINER}\` **Endpoint:** ${URL}"
|
||||
echo "**Load mode:** \`${LOAD_MODE}\`$([ "$LOAD_MODE" = "decode-single" ] && echo " (${TARGET_CAP_SECONDS}s × 2 timed streams)")$([ "$LOAD_MODE" = "decode-concurrent" ] && echo " (concurrency=${CONCURRENCY}, ${DECODE_CONCURRENT_RUN_SECONDS}s/run × ${BENCH_RUNS} runs × 2 timed batches)")$([ "$LOAD_MODE" = "prefill-heavy" ] && echo " (target-prefill=${TARGET_PREFILL_SECONDS}s, filler_repeats=${PREFILL_FILLER_REPEATS})")$([ "$LOAD_MODE" != "decode-single" ] && echo " (bench-runs=${BENCH_RUNS})")"
|
||||
[ -n "$CALIBRATION_NOTE" ] && echo "**Calibration:** ${CALIBRATION_NOTE}"
|
||||
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z"
|
||||
# --include-commit: stamp club-3090 git short SHA next to the date if requested.
|
||||
# Suppress entirely (rather than show "n/a") when run from a non-clone or
|
||||
# when git isn't reachable — closes the curl-pipe-from-docs UX hole.
|
||||
COMMIT_FRAGMENT=""
|
||||
if [ "${INCLUDE_COMMIT:-0}" = "1" ]; then
|
||||
COMMIT_SHA=$(git -C "$REPO_ROOT" rev-parse --short HEAD 2>/dev/null || true)
|
||||
if [ -n "$COMMIT_SHA" ]; then
|
||||
COMMIT_FRAGMENT=" **club-3090 commit:** \`${COMMIT_SHA}\`"
|
||||
fi
|
||||
fi
|
||||
echo "**Date:** $(date -u +%Y-%m-%dT%H:%M:%S)Z${COMMIT_FRAGMENT}"
|
||||
echo ""
|
||||
if [ "$COOLING" = "unspecified" ]; then
|
||||
echo "> ⚠️ Cooling class not specified at run time. Add **air / water / AIO** when posting"
|
||||
|
||||
@@ -325,8 +325,8 @@ preflight_compose_deps() {
|
||||
mmproj_in_container="${mmproj_in_container//\$\{MMPROJ_FILE:-/}"
|
||||
mmproj_in_container="${mmproj_in_container%\}}"
|
||||
|
||||
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
|
||||
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b/mmproj-F16.gguf"
|
||||
[[ -z "$gguf_in_container" ]] && gguf_in_container="qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf"
|
||||
[[ -z "$mmproj_in_container" ]] && mmproj_in_container="qwen3.6-27b-gguf/mmproj-F16.gguf"
|
||||
|
||||
if [[ -n "${GGUF_FILE:-}" ]]; then gguf_in_container="$GGUF_FILE"; fi
|
||||
if [[ -n "${MMPROJ_FILE:-}" ]]; then mmproj_in_container="$MMPROJ_FILE"; fi
|
||||
@@ -375,9 +375,10 @@ preflight_compose_deps() {
|
||||
if [[ $hint_gguf -eq 1 ]]; then
|
||||
echo "[preflight] hf download unsloth/Qwen3.6-27B-GGUF \\" >&2
|
||||
echo "[preflight] Qwen3.6-27B-UD-Q3_K_XL.gguf mmproj-F16.gguf \\" >&2
|
||||
echo "[preflight] --local-dir ${model_dir}/qwen3.6-27b/unsloth-q3kxl" >&2
|
||||
echo "[preflight] --local-dir \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl" >&2
|
||||
echo "[preflight] # (set MODEL_DIR first: export MODEL_DIR=\${MODEL_DIR:-/path/to/your/models})" >&2
|
||||
echo "[preflight] # mmproj lands at unsloth-q3kxl/ — move it up so the default --mmproj path resolves:" >&2
|
||||
echo "[preflight] # mv ${model_dir}/qwen3.6-27b/unsloth-q3kxl/mmproj-F16.gguf ${model_dir}/qwen3.6-27b/" >&2
|
||||
echo "[preflight] # mv \${MODEL_DIR}/qwen3.6-27b-gguf/unsloth-q3kxl/mmproj-F16.gguf \${MODEL_DIR}/qwen3.6-27b-gguf/" >&2
|
||||
echo "[preflight] (~16 GB total. setup.sh today only fetches the vLLM AutoRound weights;" >&2
|
||||
echo "[preflight] GGUF must be fetched separately for any llamacpp/* variant.)" >&2
|
||||
fi
|
||||
|
||||
@@ -177,6 +177,19 @@ if [[ -n "$DETECTED_MODEL" && "$DETECTED_MODEL" != "$MODEL" ]]; then
|
||||
MODEL="$DETECTED_MODEL"
|
||||
fi
|
||||
|
||||
# hermesagent-20 runs its agent inside a Docker sandbox container. Localhost-style
|
||||
# URLs (localhost/127.x/[::1]) inside the container resolve to the container itself,
|
||||
# not the host's vLLM. Auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 so benchlocal-cli
|
||||
# (a) adds --add-host=host.docker.internal:host-gateway to the sandbox container, and
|
||||
# (b) rewrites the model endpoint URL to use host.docker.internal:<port> for the
|
||||
# hermes-agent's outbound API calls. Skip if already set (user override) or if URL
|
||||
# already uses host.docker.internal / a non-loopback host (real LAN IP, k8s service).
|
||||
if [[ -z "${BENCHLOCAL_HERMES_RESOLVE_LOCALHOST:-}" ]] \
|
||||
&& [[ "$URL" =~ ^https?://(localhost|127\.|\[::1\]) ]]; then
|
||||
export BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1
|
||||
echo "[quality-test] localhost URL detected — auto-set BENCHLOCAL_HERMES_RESOLVE_LOCALHOST=1 for hermes sandbox endpoint rewrite" >&2
|
||||
fi
|
||||
|
||||
# ---- run benchlocal-cli ------------------------------------------------------
|
||||
|
||||
RESULTS_DIR="${ROOT_DIR}/results/quality"
|
||||
|
||||
@@ -90,6 +90,73 @@ case "${MODEL_NAME}" in
|
||||
esac
|
||||
|
||||
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
# ---------- MODEL_DIR resolution ----------
|
||||
# Order of precedence:
|
||||
# 1. MODEL_DIR already exported in the calling shell → use as-is
|
||||
# 2. .env at repo root sets MODEL_DIR → source it
|
||||
# 3. Interactive prompt (only if stdin is a TTY) → ask user
|
||||
# 4. Silent fallback to <repo>/models-cache → in-repo default
|
||||
#
|
||||
# The prompt only fires for fresh users on a TTY who haven't set anything.
|
||||
# CI / scripted runs (no TTY) get the silent fallback, preserving prior behavior.
|
||||
|
||||
# Step 2: source repo-root .env if present (lets a saved choice persist)
|
||||
if [[ -z "${MODEL_DIR:-}" && -f "${ROOT_DIR}/.env" ]]; then
|
||||
# shellcheck source=/dev/null
|
||||
set -a; source "${ROOT_DIR}/.env"; set +a
|
||||
fi
|
||||
|
||||
# Step 3: prompt if still unset + interactive
|
||||
if [[ -z "${MODEL_DIR:-}" && -t 0 && -t 1 ]]; then
|
||||
echo ""
|
||||
echo "Where should I put model weights?"
|
||||
echo " Models are large (Qwen3.6-27B AutoRound: ~14 GB; Gemma 4 31B: ~21 GB)."
|
||||
echo " This dir lives outside the git tree — pick a location with sufficient free space."
|
||||
echo ""
|
||||
echo " 1) ${ROOT_DIR}/models-cache (in-repo, default — pollutes git tree)"
|
||||
echo " 2) ${HOME}/models (recommended for cross-rig — outside repo)"
|
||||
echo " 3) custom path"
|
||||
echo ""
|
||||
while true; do
|
||||
read -rp "Choice [1-3] (or set MODEL_DIR env var to skip): " pick
|
||||
case "${pick}" in
|
||||
1) MODEL_DIR="${ROOT_DIR}/models-cache"; break ;;
|
||||
2) MODEL_DIR="${HOME}/models"; break ;;
|
||||
3)
|
||||
read -rp " Enter absolute path: " custom
|
||||
if [[ "${custom}" =~ ^/ ]]; then
|
||||
MODEL_DIR="${custom}"; break
|
||||
else
|
||||
echo " ! must be an absolute path (start with /)" >&2
|
||||
fi
|
||||
;;
|
||||
*) echo " ! invalid — pick 1, 2, or 3" >&2 ;;
|
||||
esac
|
||||
done
|
||||
echo ""
|
||||
|
||||
# Offer to persist the choice so future runs skip the prompt
|
||||
read -rp "Save MODEL_DIR=${MODEL_DIR} to .env so we skip this next time? [Y/n]: " save
|
||||
if [[ "${save:-y}" =~ ^[Yy]$ || -z "${save:-}" ]]; then
|
||||
if [[ -f "${ROOT_DIR}/.env" ]]; then
|
||||
# Update existing .env (replace MODEL_DIR= line if present, else append)
|
||||
if grep -qE "^MODEL_DIR=" "${ROOT_DIR}/.env"; then
|
||||
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=${MODEL_DIR}|" "${ROOT_DIR}/.env"
|
||||
else
|
||||
echo "MODEL_DIR=${MODEL_DIR}" >> "${ROOT_DIR}/.env"
|
||||
fi
|
||||
else
|
||||
echo "MODEL_DIR=${MODEL_DIR}" > "${ROOT_DIR}/.env"
|
||||
fi
|
||||
echo " → saved. (.env is gitignored.)"
|
||||
else
|
||||
echo " → not saved. Set MODEL_DIR=... when re-running, or you'll get this prompt again."
|
||||
fi
|
||||
echo ""
|
||||
fi
|
||||
|
||||
# Step 4: silent fallback (preserves prior behavior for non-TTY contexts)
|
||||
MODEL_DIR="${MODEL_DIR:-${ROOT_DIR}/models-cache}"
|
||||
GENESIS_DIR="${ROOT_DIR}/models/${MODEL_NAME}/vllm/patches/genesis"
|
||||
|
||||
|
||||
@@ -557,15 +557,21 @@ def cmd_run(endpoint, req_path, timeout_s, metrics_path):
|
||||
choices = chunk.get("choices") or []
|
||||
if choices:
|
||||
delta = choices[0].get("delta") or {}
|
||||
if ttft is None and (delta.get("content") or delta.get("reasoning_content") or delta.get("tool_calls")):
|
||||
# vLLM emits reasoning under either `delta.reasoning_content`
|
||||
# (older qwen3 reasoner path) or `delta.reasoning` (current
|
||||
# nightly as of vllm-0.20.2rc1+; legacy field name). Watch
|
||||
# both so the soak harness doesn't go silent when the
|
||||
# underlying field name shifts under us.
|
||||
reasoning_delta = delta.get("reasoning_content") or delta.get("reasoning")
|
||||
if ttft is None and (delta.get("content") or reasoning_delta or delta.get("tool_calls")):
|
||||
ttft = time.time() - t0
|
||||
# Accumulate streamed parts. vLLM splits content/reasoning
|
||||
# across many small deltas; tool_calls stream as indexed
|
||||
# objects whose fields (name, arguments) arrive in pieces.
|
||||
if delta.get("content"):
|
||||
content_parts.append(delta["content"])
|
||||
if delta.get("reasoning_content"):
|
||||
reasoning_parts.append(delta["reasoning_content"])
|
||||
if reasoning_delta:
|
||||
reasoning_parts.append(reasoning_delta)
|
||||
for tc in (delta.get("tool_calls") or []):
|
||||
idx = tc.get("index", 0)
|
||||
slot = tool_calls_acc.setdefault(idx, {"id": "", "type": "function", "name": "", "args": ""})
|
||||
|
||||
Reference in New Issue
Block a user