Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7f7e41c6c4 | ||
|
|
37c739b4f9 | ||
|
|
ed1507122c | ||
|
|
b07b2f99e7 | ||
|
|
15bb3f307f |
@@ -80,7 +80,8 @@ Primary serving model. Hybrid Qwen3-Next architecture (DeltaNet GDN + standard a
|
||||
| Compose | Rig | Quant | Max ctx | Narr / Code TPS | PP tok/s | Peak VRAM | Date | Notes |
|
||||
| --- | --- | --- | ---: | ---: | ---: | ---: | --- | --- |
|
||||
| `llamacpp/default` | @noonghunna (1× 3090) | Unsloth Q5_K_XL | 262K | 21 / 21 | — | ~20 GB | 2026-04-21 | bulletproof — different engine, different memory allocator, no Cliff 1 / Cliff 2. Slow decode but cliff-immune. |
|
||||
| `llamacpp/concurrent` | @noonghunna (1× 3090) | Unsloth Q5_K_XL | 262K | TBD | — | TBD | — | concurrent-serving variant. |
|
||||
| `llamacpp/mtp` ⭐ | @noonghunna (1× 3090) | Unsloth Q4_K_M MTP (`unsloth/Qwen3.6-27B-MTP-GGUF`) | **131K** | **51.28 / 59.72** (decode, n=3, CV 1.9% / 0.5%) | **1063** | ~22.9 GB (~1.6 GB headroom) | 2026-05-19 | **Mainline llama.cpp build 9235 (d14ce3dab, PR #22673 merged)** + MTP `n=2` + q4_0 KV + `-ub 1024` + native template + `--reasoning off`. **verify-stress 7/7 PASS incl. 60K + 91K needle recall** — at this config the Cliff 2 single-prompt narrative is **config-driven, not architectural** (the prior "50–60K wall" was the `-ub 2048` activation-peak bound; `-ub 1024` halves it and walks past 91K cleanly). **Quality 8-pack 102/150 (68%)** — beats every Qwen vLLM-dual config in [#119](https://github.com/noonghunna/club-3090/discussions/119) by 6–16 pp. **Aider-polyglot-30 17/30 (56.7%)** — matches Qwen vLLM bf16 dual exactly on half the hardware (per-lang: cpp 3/5, go 4/5, java 2/5, js 4/5, python 3/5, rust 1/5). **Per-GPU code TPS (59.72) ≈ vLLM-dual configs (60–63)** — engine-side per-card rate is identical; vLLM dual's aggregate advantage is purely from the second card. Zero patches, no Genesis, no AutoRound. [PR #166](https://github.com/noonghunna/club-3090/pull/166). |
|
||||
| `llamacpp/mtp-vision` (NEW) | @noonghunna (1× 3090) | Unsloth Q4_K_M MTP + mmproj-F16 | **49K** | **56.52 / 66.17** (decode, n=5, CV 1.6% / 1.7%) | **1158** | ~20.5 GB (~3.5 GB headroom) | 2026-05-20 | **First stack profile combining MTP + vision** (sweep-verified on build 9235; the older "strip mmproj when MTP" rule was obsolete). Same Q4_K_M MTP GGUF + q4_0 KV + MTP `n=2` + `-ub 1024` + `--mmproj mmproj-F16.gguf` (vision projector loaded). **Multimodal probe ✅** (model answered "Red" on a synthetic 64×64 red PNG in 0.5s; mmproj pipeline functional end-to-end). **verify-stress 7/7 PASS** (`sapphire iguana 19` @ 9.8K, `crimson falcon 18` @ 29.3K, all 7 rungs clean at the 49K ctx). TPS being ~10% higher than the no-vision Config A above isn't strictly A/B-controlled across GPUs — Config A was on GPU 0 after ~2 hrs continuous load; this was a fresh GPU 1; plausibly KV-cache-locality (49K pool 2.6× smaller than 131K) + thermal-state. [PR #166](https://github.com/noonghunna/club-3090/pull/166). |
|
||||
| llama.cpp PR [#22673](https://github.com/ggml-org/llama.cpp/pull/22673) MTP, custom build (`Qwen3.6-27B-MTP-Q4_K_M-GGUF` + `--spec-type mtp --spec-draft-n-max 3`) | @efschu (**2× Tesla V100-SXM2-16GB**, Xeon Gold 6154, Debian 13, custom-built llama-server docker) | Q4_K_M MTP | 100K | **49.96 / 62.46** | — | 15.6 GB/card (15,596 MiB at 100K ctx) | 2026-05-06 | **First V100 (sm_70 Volta) cross-rig data on the matrix** — only non-3090/4090/5090 GPU class tested. vLLM blocked (V100=CC 7.0, vLLM needs ≥7.5); fell back to llama.cpp via am17an's PR #22673 with a custom-built docker. **All 7 stress checks PASS including 90K NIAH** (Cliff 2 territory). 2× cards via tensor split (`-sm tensor`). MTP n=3, accept rates not in log. ~80 W/card (V100 max 300 W). [Issue #80](https://github.com/noonghunna/club-3090/issues/80). |
|
||||
| llama.cpp PR [#22673](https://github.com/ggml-org/llama.cpp/pull/22673) MTP, host build (`havenoammo/Qwen3.6-27B-MTP-UD-GGUF` + `--spec-type mtp --spec-draft-n-max 3` + q4_0 KV) | @lamentofhighborne (1× RTX 3090, PCIe x8, 350W) | UD-Q4_K_XL + Q8_0 MTP head | **131K** | **47.12 / 60.42** | — | ~23.1 GiB | 2026-05-07 | **First 1× 3090 llama.cpp MTP data point** on Qwen3.6-27B. Decode 47.60 / 61.71 TPS, TTFT 212 / 194 ms. **`verify-full-mtp.sh` PASS 8/8** (locally-adapted), **`verify-stress-mtp.sh` PASS 7/7 including 91K needle at 131K ctx** — pushes the documented llama.cpp MTP ctx ceiling from ~64-80K (q8_0 KV) to 131K (q4_0 KV). MTP acceptance 78.7%; recurrent 65-layer bug from froggeric's earlier MTP GGUF did **NOT** reproduce on havenoammo's UD GGUF. Native host build (no Docker), surfaced engine-coupling shortcomings in our verify/soak harness — see [Issue #85](https://github.com/noonghunna/club-3090/issues/85). |
|
||||
| llama.cpp PR [#22673](https://github.com/ggml-org/llama.cpp/pull/22673) MTP, host build (`froggeric/Qwen3.6-27B-MTP-GGUF` + `--spec-type mtp --spec-draft-n-max 3` + q4_0 KV) | @lamentofhighborne (1× RTX 3090, PCIe x8, 350 W) | Q4_K_M MTP | **164K** | **47.49 / 55.09** | — | ~22.2 GiB | 2026-05-07 | **Second 1× 3090 llama.cpp MTP data point on same rig** — froggeric's Q4_K_M MTP GGUF vs havenoammo's UD-Q4_K_XL above. Decode 47.91 / 55.81 TPS, TTFT 96 / 98 ms. `verify-full-mtp.sh` PASS 8/8, `verify-stress-mtp.sh` PASS 7/7 incl. 91K needle at 164K ctx. Functional MTP acceptance **86.7%**; canonical acceptance 55.3% narr / 71.2% code. **Ctx-fit ladder**: 262K OOMed MTP, 229K served without MTP, 196K initialized MTP but daemon died at 90K stress; 164K was the stable stress-passing ceiling on this rig. **Beats havenoammo on narr (47.49 vs 47.12, +0.8%) and ctx ceiling (164K vs 131K) but trails on code (55.09 vs 60.42, −9%)**. Manual long-context needles also passed at **120K** (39.39 decode TPS, 81% MTP accept) and **150K** (35.44 decode TPS, 80% MTP accept). MTP+vision incompat (per froggeric's model card); separate no-MTP+vision path passed 65K and 150K. [Issue #94](https://github.com/noonghunna/club-3090/issues/94). |
|
||||
|
||||
@@ -16,6 +16,51 @@ history; SemVer takes over from `v0.3.0` onward.
|
||||
|
||||
---
|
||||
|
||||
## v0.8.2 — 2026-05-19
|
||||
|
||||
|
||||
### ✨ Features
|
||||
|
||||
- feat(pull): v0.8.2 STEP V5 — recommend UX + report-a-failed-pull doc + §9-reconciliation ([c5b5e9b](https://github.com/noonghunna/club-3090/commit/c5b5e9b27e24c04a880a0504b7497e266c2c8044))
|
||||
- feat(nvlink): auto-detect NVLink on N-GPU topologies; add detection to multi4 + gemma-4-26b dual ([8f8ec1c](https://github.com/noonghunna/club-3090/commit/8f8ec1cebe375cf687550ce1724d1545386f36e4))
|
||||
- feat(pull): v0.8.2 STEP V4 — optional whichllm hw-detect subprocess (CONTRACT-3, hw-detect-only) ([3917728](https://github.com/noonghunna/club-3090/commit/39177282b793549e6e4f2ac960358f0204c9cc5e))
|
||||
- feat(switch): v0.8.2 STEP V3 — switch.sh ↔ compose_registry parity (CONTRACT-2b-ii) ([e6503bc](https://github.com/noonghunna/club-3090/commit/e6503bc046e569a7d6adac46dc768418bd14f0a8))
|
||||
- feat(pull): v0.8.2 STEP V3 — arch-registry expansion + chat-template attribution/drift_guard ([999c93f](https://github.com/noonghunna/club-3090/commit/999c93fe8c3629b484d5901de7e58dda84ad39b4))
|
||||
- feat(pull): v0.8.2 STEP V2 — surface pointer + --submit-last/--submit (gh + gh-less, consented, F5 reuse) ([e1cdcb5](https://github.com/noonghunna/club-3090/commit/e1cdcb53c7c9ff32e5eb101df15944d0f99747d2))
|
||||
- feat(pull): v0.8.2 STEP V1 — capture-on-hard-block pt1-gate emitter + BaseCaptureBundle protocol lift ([20f1557](https://github.com/noonghunna/club-3090/commit/20f1557d2992fa2d0362e0e95ee9b9264199cc38))
|
||||
- feat(report): lspci PCIe/P2P diagnostics subsection (LnkSta/ACS/topology) (#148) ([af2e45a](https://github.com/noonghunna/club-3090/commit/af2e45ae96376a781b15f9cf8c6896e1b3ecf5ae))
|
||||
|
||||
|
||||
### 🐛 Bug fixes
|
||||
|
||||
- fix(pull): v0.8.2 STEP V5 — recommend must not label a fits-clean model "DOES NOT FIT" ([26949d7](https://github.com/noonghunna/club-3090/commit/26949d7fa845979f2234a961abb53b3fedef3062))
|
||||
- fix(pull): v0.8.2 STEP V3 — deliver CONTRACT-2's engine-supported broadening (TRC two-class) ([d78b9a9](https://github.com/noonghunna/club-3090/commit/d78b9a94963c581061c01b413827630cd65aafdd))
|
||||
- fix(pull): v0.8.2 STEP V2 — gh-less issue body must not carry the absolute capture path ([52451ca](https://github.com/noonghunna/club-3090/commit/52451ca0b0b190eb6c194917f9aa7a0dba8155e0))
|
||||
- fix(launch): force LC_NUMERIC=C so the VRAM-budget printf survives comma-decimal locales (#159) ([186dc93](https://github.com/noonghunna/club-3090/commit/186dc93fae60d7047a790033176f1c3c2e0bd54a))
|
||||
- fix(deriver): correct stale "GGUF not supported until v0.8.1" message — now misleading post-v0.8.1-ship ([344ab87](https://github.com/noonghunna/club-3090/commit/344ab87dd3723cf0fc30834141c2ccb17f25f507))
|
||||
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
- docs(architecture): bring current-state docs up to v0.8.2 (recommend / submit on-ramp / arch-registry / hwdetect) ([c5c8f46](https://github.com/noonghunna/club-3090/commit/c5c8f469b09ef17252cf2a4fd0a587ae9e9a2adf))
|
||||
- docs(generator): state plainly that generated-compose capacity is the reference profile's, NOT fit-adapted ([247b1dc](https://github.com/noonghunna/club-3090/commit/247b1dcfe859f0da439d1bc41d21cd3e9efb8970))
|
||||
- docs(pull): v0.8.2 STEP V6 — correct §9/headline to the true bundled release scope ([b791271](https://github.com/noonghunna/club-3090/commit/b79127176f34f67258e0dc50b1992032fa1c653e))
|
||||
- docs: fix duplicate MULTI_CARD.md entry in docs index ([966a8d1](https://github.com/noonghunna/club-3090/commit/966a8d142f0c66f9a5ccde3d9adca39b575e5482))
|
||||
- docs: reorder docsindex (GSD first), add FAQ TOC + promote troubleshooting ladder, add tool-calling example ([a891b39](https://github.com/noonghunna/club-3090/commit/a891b3921f8fd598237d4649b68dd552419583e8))
|
||||
- docs: add GETTING_STARTED.md, Gemma 4 model READMEs, restructure main README with quick start first ([6368bae](https://github.com/noonghunna/club-3090/commit/6368bae648684c6e2c9645dc96bcf5aa7f5d1b05))
|
||||
- docs: fix stale NVLINK_MODE comment, INTERNALS.md cliff status, and dead companion repo link ([28bd0e8](https://github.com/noonghunna/club-3090/commit/28bd0e89703b1c5052990dd071d4debb758867ce))
|
||||
- docs(container-runtimes): Proxmox passthrough — NVLink is the fragile path, not Proxmox (#161) ([3f066a0](https://github.com/noonghunna/club-3090/commit/3f066a044dceb918709fa32a931511980b2cb0fc))
|
||||
- docs(benchmarks): add @hlo-world dual-3090 PCIe x4 dual-dflash-noviz row (#158) ([135f2c4](https://github.com/noonghunna/club-3090/commit/135f2c48fd25b782a7f5c74b4ca828cb4682f812))
|
||||
- docs(upstream): froggeric v19 re-eval PASSED — ADOPTED (#150) ([ec1fd65](https://github.com/noonghunna/club-3090/commit/ec1fd652e8b02aa1f752d581f6e8c1fd5fdef0f3))
|
||||
|
||||
|
||||
### 🧹 Maintenance
|
||||
|
||||
- chore(chat-template): re-vendor latest froggeric Qwen3.6 template for re-eval (#150) ([8a9ea6c](https://github.com/noonghunna/club-3090/commit/8a9ea6ca45489ad8520d65a097203d4da1d78989))
|
||||
|
||||
|
||||
|
||||
[Pin: `git checkout v0.8.2`] · [Full diff](https://github.com/noonghunna/club-3090/compare/v0.8.1...v0.8.2)
|
||||
## v0.8.1 — 2026-05-17
|
||||
|
||||
|
||||
|
||||
@@ -29,7 +29,9 @@ bash scripts/launch.sh
|
||||
# Or skip the wizard:
|
||||
# bash scripts/launch.sh --variant vllm/default # single-card chat (recommended)
|
||||
# bash scripts/launch.sh --variant vllm/dual # dual-card 262K + vision
|
||||
# bash scripts/launch.sh --variant llamacpp/default # single-card 262K, no cliffs
|
||||
# bash scripts/launch.sh --variant llamacpp/default # single-card 262K vanilla, no cliffs
|
||||
# bash scripts/launch.sh --variant llamacpp/mtp # single-card 131K + MTP (fast, ~60 code TPS)
|
||||
# bash scripts/launch.sh --variant llamacpp/mtp-vision # single-card 49K + MTP + vision
|
||||
# Or partial flags (wizard fills the rest):
|
||||
# bash scripts/launch.sh --model qwen3.6-27b --gpus 0,1
|
||||
# bash scripts/launch.sh --tp 2 --pp 1 # override vLLM parallelism
|
||||
@@ -68,6 +70,7 @@ bash scripts/update.sh
|
||||
- **Model-agnostic**: today ships curated configs for Qwen3.6-27B and friends; structure scales as we add models
|
||||
- **Universal `pull`** (v0.8.0; extended in v0.8.2) — evaluate any safetensors HF repo, get an honest one-line fit verdict (`--recommend`), and when a pull hard-blocks, send the redacted diagnostic back in one consented step (`--submit-last`). Broader arch coverage each release. See [`docs/PULL.md`](docs/PULL.md)
|
||||
|
||||
**New to local AI itself?** → [`docs/LOCAL_AI_PRIMER.md`](docs/LOCAL_AI_PRIMER.md) — plain-English: how hardware / engines / model sizes / quants fit together.
|
||||
**New here?** → [`docs/GETTING_STARTED.md`](docs/GETTING_STARTED.md) — 5-minute clone-to-curl path.
|
||||
**Already running, want to compare engines?** → [docs/engines/](docs/engines/)
|
||||
**Picking an engine** (vLLM / llama.cpp / SGLang)? → [docs/INFERENCE_ENGINES.md](docs/INFERENCE_ENGINES.md)
|
||||
@@ -121,6 +124,7 @@ club-3090/
|
||||
├── CHANGELOG.md cross-cutting changes (engine pin bumps, script updates)
|
||||
├── LICENSE Apache-2.0
|
||||
├── docs/
|
||||
│ ├── LOCAL_AI_PRIMER.md plain-English on-ramp: hardware / engines / sizes / quants
|
||||
│ ├── ARCHITECTURE.md how this stack thinks about LLM serving on 24 GB
|
||||
│ ├── HARDWARE.md Ampere SM 8.6+, NVLink note, 24 GB ceilings
|
||||
│ ├── GLOSSARY.md plain-language definitions (TPS / KV / MTP / TP / etc.)
|
||||
|
||||
@@ -59,6 +59,8 @@ llama.cpp uses **different kernels and a different memory allocator** than vLLM.
|
||||
|
||||
Trade: llama.cpp gives up ~3× decode speed (21 TPS vs 67 on vLLM single when single works) for cliff-immunity at long context. Different engineering posture for different workload shapes.
|
||||
|
||||
> **2026-05-19 update — `llamacpp/mtp` config closes most of the speed gap and walks past 91K cleanly.** With MTP `n=2` + `-ub 1024` + native template + `--reasoning off` on a single 3090, the llama.cpp path now measures **~51 narr / ~60 code TPS** (vs ~21 vanilla), and `verify-stress.sh` 7/7 — *including* the 60K + 91K needle rungs we previously treated as Cliff-2-territory. So the "Cliff 2 single-prompt at 50–60K is architectural" framing was too strong: on llama.cpp at this config, the cliff was **config-driven, not architectural** (the per-pass activation peak at `-ub 2048` was the actual bound; halving it to 1024 + the larger KV pool closes it). The vLLM Cliff 2 narrative above is unchanged — those configs hit a different kernel-level failure path. See `llamacpp/mtp` in [docs/SINGLE_CARD.md](SINGLE_CARD.md).
|
||||
|
||||
---
|
||||
|
||||
## vLLM pin compatibility status (master ships on v0.20 + Genesis v7.69 — 2026-05-02 PM)
|
||||
@@ -522,7 +524,7 @@ This is why our launch frame is **two routes, not one**: vLLM dual-card for max
|
||||
| Cap `max-model-len` at 48K (TQ3) | Cliff 1 (under threshold) | `docker-compose.yml` (default) |
|
||||
| FP8 KV + PN8 + cap at 75K | Cliff 1 (PN8 absorbs leak) | `tools-text.yml` |
|
||||
| TP=2 (dual-card) | Cliff 2 (state splits across cards) | `dual.yml`, `dual-turbo.yml` |
|
||||
| llama.cpp engine swap | Both (different library entirely) | `llamacpp/default`, `llamacpp/concurrent` |
|
||||
| llama.cpp engine swap | Both (different library entirely) | `llamacpp/default`, `llamacpp/mtp`, `llamacpp/mtp-vision` |
|
||||
|
||||
### Workarounds that don't work or are unavailable
|
||||
|
||||
|
||||
@@ -87,7 +87,7 @@ $ scripts/generate-compose.sh --profile vllm/dual-tq3-mtp
|
||||
generation is permanently out of scope -> refuse
|
||||
|
||||
$ scripts/generate-compose.sh --profile llamacpp/default
|
||||
[generate-compose] REFUSE: engine llama-cpp-mainline type='llama.cpp' != vllm;
|
||||
[generate-compose] REFUSE: engine llama-cpp-local type='llama.cpp' != vllm;
|
||||
the #141 generator is non-Genesis vLLM only -> refuse (out of scope)
|
||||
```
|
||||
|
||||
|
||||
@@ -332,7 +332,7 @@ In the Cline settings panel:
|
||||
- **API Key:** `sk-local` (any non-empty string)
|
||||
- **Model ID:** `qwen3.6-27b-autoround`
|
||||
|
||||
Cline sends large tool returns (file reads, web fetches) up to ~25K tokens. As of 2026-05-02 PM (Genesis v7.69 dev tip + vllm#35975 backport), `vllm/long-text` (180K balanced + MTP K=3) handles these cleanly — 33K AND 50K tool-prefill stress PASS, and **60K single-prompt prefill PASS** (the Cliff 2 wall closed at 60K). For one-shot prompts beyond 60K, switch to `llamacpp/default` (262K, slower) or `dual-turbo.yml` (262K + 4 streams). See [docs/SINGLE_CARD.md](SINGLE_CARD.md), [docs/CLIFFS.md](CLIFFS.md), and the [VRAM diagram](../models/qwen3.6-27b/README.md#vram-allocation-across-configs).
|
||||
Cline sends large tool returns (file reads, web fetches) up to ~25K tokens. As of 2026-05-02 PM (Genesis v7.69 dev tip + vllm#35975 backport), `vllm/long-text` (180K balanced + MTP K=3) handles these cleanly — 33K AND 50K tool-prefill stress PASS, and **60K single-prompt prefill PASS** (the Cliff 2 wall closed at 60K). For one-shot prompts beyond 60K, switch to `llamacpp/default` (262K vanilla, slower), `llamacpp/mtp` (131K + MTP, ~60 code TPS, single-card, 7/7 verify-stress incl. 91K needle), or `dual-turbo.yml` (262K + 4 streams). See [docs/SINGLE_CARD.md](SINGLE_CARD.md), [docs/CLIFFS.md](CLIFFS.md), and the [VRAM diagram](../models/qwen3.6-27b/README.md#vram-allocation-across-configs).
|
||||
|
||||
### Cursor
|
||||
|
||||
|
||||
@@ -284,7 +284,9 @@ Symptoms users report: "performance degrades after 20 turns", "throughput drops
|
||||
bash scripts/switch.sh vllm/dual # 111+ TPS p50, 0 errors, 0 MiB growth across 5 sessions
|
||||
|
||||
# 1× 3090 — different engine, different kernels, different allocator
|
||||
bash scripts/switch.sh llamacpp/default # 21 TPS, 262K context, cliff-immune
|
||||
bash scripts/switch.sh llamacpp/default # 21 TPS, 262K context, cliff-immune, vision
|
||||
bash scripts/switch.sh llamacpp/mtp # ~60 code TPS, 131K, MTP, 7/7 verify-stress (incl. 91K needle)
|
||||
bash scripts/switch.sh llamacpp/mtp-vision # ~66 code TPS, 49K + vision (multimodal MTP)
|
||||
```
|
||||
|
||||
**Want to verify your rig hits the same class:**
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
The fastest path from `git clone` to serving your first response. No decisions, no menus — just commands.
|
||||
|
||||
> New to local AI and the terms below feel like jargon? Read [LOCAL_AI_PRIMER.md](LOCAL_AI_PRIMER.md) first — how hardware, engines, model sizes, and quants fit together in plain English.
|
||||
|
||||
```bash
|
||||
# 1. Clone
|
||||
git clone https://github.com/noonghunna/club-3090.git
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
Plain-language definitions for terms used throughout the docs. Roughly grouped by topic.
|
||||
|
||||
> Want the *narrative* version — how hardware / engines / sizes / quants fit together — instead of isolated definitions? See [LOCAL_AI_PRIMER.md](LOCAL_AI_PRIMER.md).
|
||||
|
||||
> Coming from a different background and don't see something here? Open an issue and we'll add it.
|
||||
|
||||
---
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
# A Layman’s Guide to Local AI: Hardware, Engines, and Quants
|
||||
|
||||
If you are new to running Large Language Models (LLMs) on your own hardware, the terminology can be overwhelming. You don't just "download a model and run it." You have to match your **Hardware** to an **Engine**, pick the right **Model Size**, and download the correct **Format/Quant**.
|
||||
|
||||
This page explains how these pieces fit together in plain English.
|
||||
|
||||
> **Scope:** This is general orientation for local AI as a whole. The club-3090 stack itself is built and tested specifically on **NVIDIA 24 GB consumer GPUs (the RTX 3090)**. Other vendors (Apple, AMD, Intel) are included here as context to help you understand the landscape — they are not a support commitment of this stack. Vendor-specific viability changes quickly; treat those sections as a map, not a guarantee.
|
||||
|
||||
---
|
||||
|
||||
### Step 1: Hardware (The Rule of VRAM)
|
||||
In local AI, processor speed is an afterthought. **VRAM (Video RAM) is the currency that matters.** If a model doesn't fit in your VRAM, it will either crash instantly or run so slowly on your system RAM that it’s unusable.
|
||||
|
||||
* **NVIDIA (CUDA):** The undisputed king of the local AI server. Because 99% of AI software is optimized for CUDA, "it just works." Used consumer cards with high VRAM (specifically the 24GB RTX 3090) are the gold standard for home labs.
|
||||
* **Apple Silicon (Unified Memory):** The "cheat code" for local AI. Macs share memory natively between the CPU and GPU. A Mac Studio with 128GB of RAM effectively has a 128GB graphics card, allowing users to run massive models quietly on a laptop/desktop using Apple's **MLX** framework.
|
||||
* **AMD (ROCm):** AMD has closed the gap massively. Modern ROCm works wonderfully as a CUDA alternative. High VRAM cards like the 7900 XTX (24GB) are now highly viable, first-class citizens.
|
||||
* **Intel (dGPU / Arc):** Intel's dedicated graphics cards (like the Arc A770 with 16GB VRAM) have entered the chat via the SYCL backend. While budget-friendly, the software ecosystem is still maturing.
|
||||
|
||||
*(For a comprehensive breakdown of hardware architectures, GPU generations, and feature support on this stack, see **[HARDWARE.md](./HARDWARE.md)**).*
|
||||
|
||||
---
|
||||
|
||||
### Step 2: Model Sizes and Architecture
|
||||
When you search for a model, you will see numbers like `8B`, `32B`, or `8x7B`. "B" stands for Billions of Parameters (the number of "synapses" in the AI's brain).
|
||||
|
||||
* **8B to 14B:** Fits on a basic 8GB-12GB GPU. Fast, great for basic chat and targeted coding, but prone to complex logic errors.
|
||||
* **32B to 35B:** The "home lab sweet spot" for a single 24GB RTX 3090 or a Mid-tier Mac. Extremely smart, highly coherent.
|
||||
* **70B+:** Nearing ChatGPT-4 levels of intelligence, but requires multiple GPUs or a very high-RAM Mac to run at reasonable speeds.
|
||||
|
||||
**Dense vs. MoE (Mixture of Experts)**
|
||||
You will also see models labeled as "Dense" or "MoE."
|
||||
* **Dense:** The model uses its *entire brain* for every single word it generates. (e.g., Llama-3 70B uses 70B parameters per word). Requires high VRAM *and* high processing power.
|
||||
* **MoE (e.g., Mixtral 8x7B, DeepSeek V3):** The model is split into "experts." It might have 50B total parameters, but only activates 12B of them to answer your specific question. MoE requires high VRAM to store the whole brain, but *very low compute power* to run, making them incredibly fast.
|
||||
|
||||
---
|
||||
|
||||
### Step 3: The Engines (The Software that runs the AI)
|
||||
To make the model talk, you need an "Inference Engine."
|
||||
|
||||
* **llama.cpp (Cross-platform):** *The rugged off-roader.* It runs on Windows, Linux, Macs, and regular CPUs. If you want to split a model between your GPU and standard RAM, you use this. (Apps like Ollama, LM Studio, and Msty are just user-friendly wrappers around `llama.cpp`).
|
||||
* **vLLM (NVIDIA/AMD Linux):** *The Formula 1 car.* Built for heavy-duty, dedicated GPU setups. It is strict—if you run out of memory, it crashes—but it is unimaginably fast and can serve dozens of users at exactly the same time.
|
||||
* **MLX (Apple Silicon):** Apple's native machine learning framework. If you are on an M-series Mac, using engines built on MLX (like `mlx-lm`) is the fastest, most battery-efficient way to run models.
|
||||
* **KTransformers:** A newer hybrid engine built specifically to run massive MoE models (like DeepSeek R1) by keeping the math-heavy parts on the GPU and offloading the rest to System RAM.
|
||||
|
||||
*(For a detailed technical comparison of server engine memory models and CLI surfaces, see **[INFERENCE_ENGINES.md](./INFERENCE_ENGINES.md)**).*
|
||||
|
||||
---
|
||||
|
||||
### Step 4: Quants, Cache, and Context (Shrinking the model)
|
||||
An uncompressed 70B model is roughly 140 GB. To make it fit your hardware, the community "quantizes" (compresses) it by rounding off the decimal points. **The Engine you chose in Step 3 dictates the Format you download.**
|
||||
|
||||
**1. Formats based on your Engine**
|
||||
* **GGUF (`llama.cpp`):** You **must** download `.gguf` files. `Q4_K_M` (4-bit) is the community sweet spot for size vs. smartness. `Q8_0` (8-bit) is nearly identical to the original but half the size.
|
||||
* **Safetensors (`vLLM` / `sglang`):** You must download a standard Hugging Face repo (a folder full of `.safetensors` files). Look for **AWQ** or **GPTQ** (4-bit VRAM savers) or **FP8** (the 8-bit standard for modern NVIDIA cards).
|
||||
> **RTX 3090 caveat:** FP8 *compute* (running FP8-weight models) requires Ada/Hopper (RTX 4090/5090, H100) — the **3090 (Ampere) cannot run FP8 weights** and will fall back or fail. On a 3090, use **AWQ/GPTQ** for 4-bit weights; FP8 is only useful here as a **KV-cache** storage format (`fp8_e5m2`, see Step 4.2), not as a weight format.
|
||||
* **MLX format (`mlx-lm`):** Macs use a custom natively optimized format (usually labeled as `-mlx` on Hugging Face).
|
||||
|
||||
**2. Tokens, Context, and KV Cache**
|
||||
Models don't read words; they read **Tokens** (roughly 3/4 of a word).
|
||||
The **Context Window** is the AI's short-term memory during your conversation. If a model has an 8,000-token context window and you paste a 10,000-token PDF into it, it will literally "forget" the beginning by the time it reaches the end.
|
||||
* **The VRAM Cost:** Your conversation history is stored in VRAM as the **KV Cache**. A 32,000-token context window can consume over 10GB of VRAM *on top* of the model weights. The KV Cache can also be quantized (e.g., `fp8_e5m2`) to squeeze massive conversations onto consumer GPUs.
|
||||
|
||||
*(For the deep-dive on exact byte-widths, capacity costs, and kernel constraints for all dtypes on this stack, see **[DTYPE_MATRIX.md](./DTYPE_MATRIX.md)** and **[KERNEL_MATRIX.md](./KERNEL_MATRIX.md)**).*
|
||||
|
||||
---
|
||||
|
||||
### Step 5: Prompt Templates (How to talk to it)
|
||||
If you boot up a model, say "Hello", and it responds with HTML code or an endless loop of `User: AI: User:`, your model isn't broken—you are using the wrong **Prompt Template**.
|
||||
|
||||
Base AI models are essentially super-powered autocorrects; they just predict the next word. To make them act like chat assistants, they are trained on specific hidden formatting tags (like `ChatML`, `Llama-3`, or `Alpaca`).
|
||||
Most modern frontends (like Open WebUI or LM Studio) detect this automatically, but if a model is speaking gibberish, checking the template is your first debugging step.
|
||||
|
||||
---
|
||||
|
||||
### Putting it together
|
||||
Figuring out *"Can I run this AWQ Safetensor in vLLM on my two 3090s with a 32,000 context window?"* requires manual math. You must calculate the weight size, the KV cache overhead, the sequence lengths, and the hardware limits.
|
||||
|
||||
*Note: The club-3090 stack does this math for you. Curated models ship heavily-tested, pre-calculated profiles; and the universal `pull` workflow evaluates **any** safetensors HF repo against this stack's KV math before downloading, so you get a fit verdict instead of a guess. See [PULL.md](./PULL.md) and [KV_MATH.md](./KV_MATH.md).*
|
||||
@@ -243,7 +243,7 @@ the §1 confidently-wrong outcome the design forbids.
|
||||
| `[D]`-emittable required | yes — reuses `[D]`'s scope-gate (`engine.type==vllm` ∧ `profile_runtime` entry exists ∧ `genesis_equipped==false`); a Genesis/TQ3 profile (`vllm/dual-turbo`) → stratum-2 `profile-not-emittable` (g0) | no |
|
||||
| `weights_variant` compat | Path-A only (`[C0]` checks the curated variant against arch constraints) | n/a — uses deriver-resolved `weight_format`/quant |
|
||||
| `drafter` | from the curated profile | `none` for non-curated (drafter profiles expose `model_compat`, not arch-compat) |
|
||||
| Non-vLLM `--profile-like` | refused stratum-2 `unsupported-runtime-engine` (g13: `llamacpp/default`, `engine=llama-cpp-mainline`, `mem_util=None`) | same — refused on both paths |
|
||||
| Non-vLLM `--profile-like` | refused stratum-2 `unsupported-runtime-engine` (g13: `llamacpp/default`, `engine=llama-cpp-local`, `mem_util=None`) | same — refused on both paths |
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -39,7 +39,9 @@ Five recommended options on Genesis v7.69 + vllm#35975 backport (2026-05-02 PM):
|
||||
| **Long ctx, text-only — Balanced MTP** ⭐ (RAG, codebase, IDE agents, default) | [`long-text.yml`](../models/qwen3.6-27b/vllm/compose/single/long-text.yml) | **180K** | 50 / 67 | ~22.3 GB (mem-util 0.93) |
|
||||
| **Long ctx, text-only — Max-context** (long single-shot RAG / codebase analysis) | [`long-text-no-mtp.yml`](../models/qwen3.6-27b/vllm/compose/single/long-text-no-mtp.yml) (NEW) | **200K** | TBD/TBD (slow decode, no MTP) | ~21.0 GB (mem-util 0.95) |
|
||||
| **Bounded thinking** (coding agents, structured-CoT — recommended grammar: DeepSeek scratchpad, 87.4% combined HE+/LCB v6) — see [STRUCTURED_COT.md](STRUCTURED_COT.md) | [`bounded-thinking.yml`](../models/qwen3.6-27b/vllm/compose/single/bounded-thinking.yml) | **180K** | 50 / 66 | ~21.7 GB (mem-util 0.95) |
|
||||
| **Bulletproof, no cliffs** (production service, unpredictable inputs) | [`llamacpp/default`](../models/qwen3.6-27b/llama-cpp/compose/docker-compose.yml) | **262K** | 21 / 21 | ~20 GB |
|
||||
| **Bulletproof, no cliffs** (production service, unpredictable inputs) | [`llamacpp/default`](../models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml) | **262K** | 21 / 21 | ~20 GB |
|
||||
| **llama.cpp + MTP, fast + long ctx** ⭐ (IDE agents, opencode, Hermes, long-multi-turn agentic) | [`llamacpp/mtp`](../models/qwen3.6-27b/llama-cpp/compose/single/mtp.yml) | **131K** | **51 / 60** | ~22.5 GB |
|
||||
| **llama.cpp + MTP + vision** (multimodal chat, screenshot-debugging, vision-aware review) | [`llamacpp/mtp-vision`](../models/qwen3.6-27b/llama-cpp/compose/single/mtp-vision.yml) | **49K** | **57 / 66** | ~20.5 GB |
|
||||
| **Small-context vLLM safe path** ([@stiggy2k16](https://github.com/noonghunna/club-3090/issues/43) data point) — IDE agents capped at <60K accumulated, when you need vLLM speed but llama.cpp is too slow | [`minimal.yml`](../models/qwen3.6-27b/vllm/compose/single/minimal.yml) at `--gpu-memory-utilization 0.95 --max-model-len 65536` | **64K** | ~32 / ~33 (no MTP) | ~22.4 GB |
|
||||
|
||||
Run via `bash scripts/launch.sh` (interactive) or `bash scripts/switch.sh <variant>`.
|
||||
@@ -115,6 +117,18 @@ For the cross-card TP=2 picture, see [`DUAL_CARD.md`](DUAL_CARD.md).
|
||||
|
||||
`bash scripts/switch.sh llamacpp/default`. Q3_K_XL (Unsloth dynamic) + q4_0 KV at 262K + vision (mmproj). Different attention library entirely (ggml-cuda, not FA2) → no Cliff 1 mechanism, no Cliff 2 mechanism. Trade is ~21 TPS (~2.5× slower than vLLM). Quant validated by [Benjamin Marie's Kaitchup eval](https://kaitchup.substack.com/p/summary-of-qwen36-gguf-evals-updating).
|
||||
|
||||
### llama.cpp + MTP, fast + long ctx — `llamacpp/mtp` ⭐
|
||||
|
||||
**Workload:** IDE agents (opencode, Cline), Hermes, multi-turn agentic. The single-card sweet spot for "fast enough + plenty of context + bulletproof."
|
||||
|
||||
`bash scripts/switch.sh llamacpp/mtp`. Q4_K_M MTP-enabled GGUF (`unsloth/Qwen3.6-27B-MTP-GGUF`) + q4_0 KV + MTP `n=2` + `-ub 1024`. Mainline llama.cpp build 9235 (PR #22673 MTP merged). **~51 narr / ~60 code TPS** — ~2.5× faster than `llamacpp/default`, only marginally slower than vLLM single-card paths *and* without the cliffs. **Quality 102/150 (68%) on the 8-pack matrix beats every Qwen vLLM dual config in [#119](https://github.com/noonghunna/club-3090/discussions/119)**; **aider-polyglot 17/30 exactly matches Qwen vLLM bf16 dual on half the hardware**. **Verify-stress 7/7** including 60K + 91K needle recall — the Cliff 2 narrative is **config-driven, not architectural**: at this config the model walks past 91K cleanly on a single 3090. No vision (mmproj cost trades against KV pool); for multimodal, use `llamacpp/mtp-vision` below.
|
||||
|
||||
### llama.cpp + MTP + vision — `llamacpp/mtp-vision`
|
||||
|
||||
**Workload:** multimodal chat, screenshot-debugging, vision-aware code review, UI-screenshot agents. The first stack profile combining MTP + vision on a single 3090 (the older "strip mmproj when MTP" rule was obsolete on build 9235 — sweep-verified 2026-05-19).
|
||||
|
||||
`bash scripts/switch.sh llamacpp/mtp-vision`. Same Q4_K_M MTP GGUF + q4_0 KV + MTP `n=2` + `-ub 1024`, **with `--mmproj mmproj-F16.gguf` mounted**. Ctx 49K (vision overhead trades against KV pool — 49K is the safe-headroom max with mmproj on 24 GB). **~57 narr / ~66 code TPS** (text path), vision overhead paid once at image-encode, not per decoded token. Verify-stress 7/7 at this config.
|
||||
|
||||
---
|
||||
|
||||
## Other variants in the repo (not recommended for shipping)
|
||||
|
||||
|
Before Width: | Height: | Size: 99 KiB After Width: | Height: | Size: 101 KiB |
|
Before Width: | Height: | Size: 95 KiB After Width: | Height: | Size: 87 KiB |
|
Before Width: | Height: | Size: 95 KiB After Width: | Height: | Size: 97 KiB |
|
Before Width: | Height: | Size: 90 KiB After Width: | Height: | Size: 82 KiB |
|
Before Width: | Height: | Size: 170 KiB After Width: | Height: | Size: 180 KiB |
|
Before Width: | Height: | Size: 151 KiB After Width: | Height: | Size: 141 KiB |
|
Before Width: | Height: | Size: 203 KiB After Width: | Height: | Size: 213 KiB |
|
Before Width: | Height: | Size: 171 KiB After Width: | Height: | Size: 159 KiB |
@@ -35,9 +35,13 @@ MODEL_DIR=/your/models/dir docker compose up -d
|
||||
|
||||
Memory budget: 14.5 GB (Q3_K_XL) + 4.5 GB KV @ 262K + 0.8 GB mmproj ≈ 20 GB / 24 GB.
|
||||
|
||||
### `single/concurrent.yml` — 4 parallel slots, vision
|
||||
### `single/mtp.yml` — MTP n=2, 131K ctx, no vision
|
||||
|
||||
Trade max context for parallelism. Same image, `--parallel 4` + smaller ctx pool.
|
||||
The single-card speed + context workhorse: ~51/60 TPS (narr/code), 131K ctx (sweep-verified safe-headroom max), 7/7 verify-stress boundary checks (incl. 60K + 91K needle recall), 102/150 (68%) on the 8-pack quality matrix. Best for IDE agents, opencode, Hermes, long-multi-turn agentic. Q4_K_M MTP GGUF (`unsloth/Qwen3.6-27B-MTP-GGUF` Q4_K_M).
|
||||
|
||||
### `single/mtp-vision.yml` — MTP n=2, 49K ctx, vision on
|
||||
|
||||
Multimodal speed profile — the first stack config combining MTP + vision (the older "strip mmproj when MTP" rule was obsolete on build 9235, sweep-verified 2026-05-19). 49K safe-headroom ceiling on 24 GB with mmproj F16 mounted.
|
||||
|
||||
### Tuning knobs
|
||||
|
||||
|
||||
@@ -1,79 +0,0 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
|
||||
# Engine: llama.cpp (NOT vLLM)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: none
|
||||
# KV: q4_0 (4-bit packed)
|
||||
# Vision: yes (mmproj F16)
|
||||
# Max ctx: 192K pool / 4 parallel slots
|
||||
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
|
||||
# Status: ✅ Production
|
||||
# Engine-profile: llama-cpp-mainline
|
||||
# Best for: Single-card multi-tenant llama.cpp — 4 concurrent agents
|
||||
# at smaller per-stream ctx; trade max-ctx for parallelism
|
||||
# ---------------------------------------------------------------------------
|
||||
# Qwen3.6-27B on llama.cpp — single 3090, max concurrency (4 slots), vision on.
|
||||
#
|
||||
# Trade max context for parallelism. Multi-tenant or batched-agent path:
|
||||
# four concurrent requests share a 192K KV pool (~48K each if split evenly,
|
||||
# but llama-server allocates dynamically per slot).
|
||||
#
|
||||
# Config differences from default (docker-compose.yml):
|
||||
# - --parallel 4 (was 1)
|
||||
# - --ctx-size 192000 (was 262144) — 4 slots fit in less total KV
|
||||
# - --cont-batching on (improves throughput when prompts arrive concurrently)
|
||||
# - everything else identical
|
||||
#
|
||||
# VRAM budget on 24 GB:
|
||||
# weights (Q3_K_XL): ~14.5 GB
|
||||
# KV at 192K (q4_0 K+V): ~3.3 GB (shared pool across 4 slots)
|
||||
# mmproj F16: ~0.8 GB
|
||||
# total: ~18.6 GB
|
||||
# headroom: ~5 GB
|
||||
#
|
||||
# When to pick this over docker-compose.yml:
|
||||
# ✅ Multi-tenant chat (e.g. small team, Open WebUI multi-user)
|
||||
# ✅ Agent farm — 4 agents working on tasks in parallel
|
||||
# ✅ A/B testing two prompts simultaneously
|
||||
# ❌ You need >65K per single conversation (default is better)
|
||||
# ❌ You need max single-stream TPS (more slots = some latency floor)
|
||||
#
|
||||
# Quick start: same as default, plus:
|
||||
# docker compose -f single/concurrent.yml up -d
|
||||
#
|
||||
# Env overrides identical to default — see docker-compose.yml header.
|
||||
|
||||
services:
|
||||
llama-cpp-qwen36-27b-concurrent:
|
||||
image: ghcr.io/ggml-org/llama.cpp:server-cuda
|
||||
container_name: "${ESTATE_CONTAINER:-llama-cpp-qwen36-27b-concurrent}"
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${ESTATE_PORT:-${PORT:-8020}}:8080"
|
||||
volumes:
|
||||
- "${MODEL_DIR:-../../../../../models-cache}:/models:ro"
|
||||
command: >-
|
||||
--host 0.0.0.0
|
||||
--port 8080
|
||||
-m /models/${GGUF_FILE:-qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf}
|
||||
--mmproj /models/${MMPROJ_FILE:-qwen3.6-27b-gguf/mmproj-F16.gguf}
|
||||
-c ${CTX_SIZE:-192000}
|
||||
-b ${BATCH_SIZE:-4096}
|
||||
-ub ${UBATCH_SIZE:-2048}
|
||||
-ngl 99
|
||||
-fa on
|
||||
--cache-type-k ${KV_TYPE:-q4_0}
|
||||
--cache-type-v ${KV_TYPE:-q4_0}
|
||||
-np ${PARALLEL:-4}
|
||||
--cont-batching
|
||||
--jinja
|
||||
--reasoning-format ${REASONING_FORMAT:-none}
|
||||
--chat-template-kwargs ${CHAT_TEMPLATE_KWARGS:-{"enable_thinking":false}}
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${ESTATE_GPUS:-${CUDA_VISIBLE_DEVICES:-0}}"]
|
||||
capabilities: [compute, utility]
|
||||
@@ -1,22 +1,34 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF)
|
||||
# Engine: llama.cpp (NOT vLLM — different engine, different memory model)
|
||||
# Model: Qwen3.6-27B (Unsloth Q3_K_XL GGUF, or MTP variant)
|
||||
# Engine: llama.cpp (local MTP-enabled build, build 9235)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: none (vanilla llama.cpp; MTP via PR #22673 not adopted yet)
|
||||
# Drafter: none (MTP opt-in via MTP_ENABLED=1, requires MTP-enabled GGUF)
|
||||
# KV: q4_0 (4-bit packed)
|
||||
# Vision: yes (mmproj F16)
|
||||
# Max ctx: 262K (full model native — no Cliff 1 / Cliff 2)
|
||||
# Vision: yes, also coexists with MTP (verified on build 9235)
|
||||
# Max ctx: 262K vanilla / 131K MTP no-vision / 49K MTP + vision (q4_0 KV, Q4_K_M, 24 GB safe-headroom)
|
||||
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
|
||||
# Status: ✅ Production
|
||||
# Engine-profile: llama-cpp-mainline
|
||||
# Engine-profile: llama-cpp-local
|
||||
# Best for: Bulletproof single-card path — slow decode (~21 TPS) but
|
||||
# cliff-immune; recommended fallback when vLLM hits OOM at long ctx
|
||||
# cliff-immune; MTP boosts to ~51/59 TPS (narr/code)
|
||||
# ---------------------------------------------------------------------------
|
||||
# Qwen3.6-27B on llama.cpp — single 3090, full 262K context, vision on.
|
||||
#
|
||||
# The "easy mode" path. No Genesis patches, no AutoRound, no patched vLLM
|
||||
# fork — just a stock llama.cpp Docker image and a single GGUF file.
|
||||
# fork — just a local llama.cpp Docker image and a single GGUF file.
|
||||
#
|
||||
# MTP spec-decode (--spec-type draft-mtp) supported. Set MTP_ENABLED=1 to
|
||||
# enable. Requires an MTP-enabled GGUF (e.g. unsloth/Qwen3.6-27B-MTP-GGUF).
|
||||
# When MTP is on:
|
||||
# - --mmproj is RETAINED (vision + MTP coexist on build 9235 — sweep-verified
|
||||
# 2026-05-19; the older "strip mmproj" rule is obsolete)
|
||||
# - --spec-draft-n-max ${MTP_DRAFT_N_MAX:-2} (2 is the sweet spot)
|
||||
# - VRAM-safe ctx ceilings (Q4_K_M + q4_0 KV on 24GB, ~1.5 GB headroom):
|
||||
# - no-vision: 131K (set CTX_SIZE=131072)
|
||||
# - vision: 49K (set CTX_SIZE=49152)
|
||||
# IMPORTANT: the default CTX_SIZE=262144 is sized for VANILLA (MTP off);
|
||||
# when MTP_ENABLED=1, override CTX_SIZE explicitly or you will OOM.
|
||||
#
|
||||
# Showcase: vLLM single-card caps at 192K opt-in (with prefill caveats);
|
||||
# this hits 262K (the model's architectural max) on a stock 3090 thanks to
|
||||
@@ -28,7 +40,7 @@
|
||||
# - mmproj: F16 (vision projector, ~0.8 GB)
|
||||
# - Context: 262144 (model max)
|
||||
# - KV: q4_0 K + q4_0 V
|
||||
# - Slots: 1 (single concurrent request; switch to single/concurrent.yml for 4)
|
||||
# - Slots: 1 (single concurrent request; for MTP-accelerated profiles see mtp.yml / mtp-vision.yml)
|
||||
# - FA: on
|
||||
#
|
||||
# VRAM budget on 24 GB:
|
||||
@@ -51,62 +63,84 @@
|
||||
# MODEL_DIR=/your/models/dir docker compose up -d
|
||||
# 3. curl http://localhost:8020/v1/models → should list the model.
|
||||
#
|
||||
# For MTP mode with an MTP-enabled GGUF:
|
||||
# MTP_ENABLED=1 GGUF_FILE=qwen3.6-27b-gguf/unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf docker compose up -d
|
||||
#
|
||||
# Override defaults via .env or shell:
|
||||
# MODEL_DIR host dir to mount as /models (default: ../../../../models-cache for repo, /path/to/your/models on your stack)
|
||||
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-q3kxl/Qwen3.6-27B-UD-Q3_K_XL.gguf)
|
||||
# MMPROJ_FILE path under /models (default: qwen3.6-27b-gguf/mmproj-F16.gguf)
|
||||
# CTX_SIZE total KV pool (default: 262144)
|
||||
# BATCH_SIZE llama.cpp -b (default: 4096)
|
||||
# UBATCH_SIZE llama.cpp -ub (default: 2048)
|
||||
# UBATCH_SIZE llama.cpp -ub (default: 1024 — cliff-survival.
|
||||
# Lowered from 2048 after empirically establishing that
|
||||
# 25K tool-prefill prompts OOM at -ub 2048 on single
|
||||
# card with 1.4-1.6 GB headroom. -ub 1024 halves the
|
||||
# per-pass activation peak. Override UBATCH_SIZE=2048
|
||||
# for raw prefill speed if no long-prompt workloads.)
|
||||
# KV_TYPE K and V quant type (default: q4_0)
|
||||
# REASONING_FORMAT reasoning channel routing (default: none — for opencode/IDE-agent compat)
|
||||
# Override to `auto` to get separate `reasoning_content` field
|
||||
# (Qwen3.6 thinking trace) — useful for clients that render
|
||||
# reasoning_content (most don't). Issue: club-3090#97.
|
||||
# DISABLE_THINKING default: 1 (thinking OFF). Set to 0 to opt INTO thinking
|
||||
# which forces empty <think></think> blocks in responses. Useful for
|
||||
# clients (e.g. opencode) that display <think> content as the response.
|
||||
# Tradeoff: applies to ALL clients on this server — Hermes/agents that
|
||||
# use thinking lose reasoning capability. Issue: club-3090#97.
|
||||
# MTP_ENABLED set to 1 to enable MTP spec-decode (default: 0)
|
||||
# MTP_DRAFT_N_MAX number of MTP draft tokens (default: 2)
|
||||
# Requires MTP-enabled GGUF (see unsloth/Qwen3.6-27B-MTP-GGUF)
|
||||
# Disables vision (--mmproj); reduces max ctx
|
||||
# REASONING thinking gate (llama.cpp --reasoning, default: off).
|
||||
# Stack-wide Qwen3.6 policy is thinking-off-by-default
|
||||
# (matches the 22/24 vLLM composes shipping
|
||||
# chat_template_kwargs.enable_thinking=false). Set
|
||||
# REASONING=on to opt into thinking traces.
|
||||
# REASONING_FORMAT reasoning channel routing (default: deepseek — hygiene
|
||||
# even with thinking off, in case the model emits stray
|
||||
# tags). With REASONING=off no <think> is emitted, so
|
||||
# `content` is clean for opencode/IDE-agent clients
|
||||
# (closes the issue #97 hang cleanly without needing
|
||||
# `--reasoning-format none`).
|
||||
# PORT host port (default: 8020)
|
||||
# CUDA_VISIBLE_DEVICES which GPU to use (default: 0)
|
||||
#
|
||||
# ─── opencode / IDE-agent compatibility (issue #97) ───
|
||||
# `--reasoning-format none` is the default because Qwen3.6's thinking mode
|
||||
# emits `<think>...</think>` blocks that llama.cpp's peg-native parser routes
|
||||
# to the OpenAI `reasoning_content` field by default. opencode (and most simple
|
||||
# OpenAI-compat clients) ignore `reasoning_content` and wait for `content`,
|
||||
# causing indefinite client hangs even though the server returns 200 cleanly.
|
||||
# `--reasoning-format none` collapses thinking into the content stream so all
|
||||
# clients work. Power users wanting reasoning_content separation: set
|
||||
# `REASONING_FORMAT=auto` in .env or shell.
|
||||
# ─── Thinking-off policy + opencode/IDE-agent compatibility (issue #97) ───
|
||||
# Qwen3.6 ships with thinking ON in its native chat template. The stack-wide
|
||||
# convention (already shipped on 22/24 vLLM composes) is thinking-off-by-
|
||||
# default. On llama.cpp the right lever is `--reasoning off` against the
|
||||
# *native* template — it correctly suppresses the <think> emission. We do NOT
|
||||
# mount the vLLM-only froggeric template here: froggeric has no llama.cpp
|
||||
# thinking-mode Jinja hook, so it silently suppresses `--reasoning off` and
|
||||
# forces thinking ON regardless. Native template + `--reasoning off` is the
|
||||
# correct llama.cpp interface and a superset of the old #97 `--reasoning-
|
||||
# format none` workaround: no <think> emitted → `content` is populated
|
||||
# directly → opencode + benchlocal-cli quality tests + Hermes all work.
|
||||
|
||||
services:
|
||||
llama-cpp-qwen36-27b:
|
||||
image: ghcr.io/ggml-org/llama.cpp:server-cuda
|
||||
image: llama-cpp:local
|
||||
container_name: "${ESTATE_CONTAINER:-llama-cpp-qwen36-27b}"
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${ESTATE_PORT:-${PORT:-8020}}:8080"
|
||||
volumes:
|
||||
- "${MODEL_DIR:-../../../../../models-cache}:/models:ro"
|
||||
# NOTE: NO chat-template mount. The native template embedded in the
|
||||
# GGUF is the correct template for llama.cpp — it exposes the Qwen3.6
|
||||
# thinking-mode Jinja gate that `--reasoning off` hooks into. The
|
||||
# froggeric template (vLLM-only, vendored under vllm/patches/) is
|
||||
# deliberately NOT mounted here: it has no llama.cpp thinking hook
|
||||
# and silently suppressed `--reasoning off`, forcing thinking ON.
|
||||
entrypoint:
|
||||
- bash
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
# DISABLE_THINKING=1 (default) appends --chat-template-kwargs to disable
|
||||
# Qwen3 thinking server-side. Forces the chat template to insert empty
|
||||
# <think></think> blocks → output goes straight to the response. Useful for
|
||||
# clients (e.g. opencode) that display <think> content as the response.
|
||||
# Tradeoff: applies to ALL clients on this server instance — Hermes/agents that
|
||||
# use thinking lose reasoning capability. See docs/HARDWARE.md and disc club-3090#97.
|
||||
# Note: $$VAR is YAML-escape for $VAR (compose passes literal $ to bash).
|
||||
EXTRA_ARGS=()
|
||||
if [ "$${DISABLE_THINKING:-1}" = "1" ]; then
|
||||
EXTRA_ARGS+=("--chat-template-kwargs" '{"enable_thinking":false}')
|
||||
echo "[entrypoint] DISABLE_THINKING=1 — chat template will produce empty <think></think>"
|
||||
|
||||
# MTP spec-decode toggle. We do NOT strip --mmproj anymore — MTP +
|
||||
# vision coexist on build 9235 (sweep-verified 2026-05-19). Earlier
|
||||
# composes stripped it defensively; that rule is obsolete.
|
||||
MTP_VAL="$${MTP_ENABLED:-0}"
|
||||
if [ "$$MTP_VAL" = "1" ]; then
|
||||
EXTRA_ARGS+=("--spec-type" "draft-mtp")
|
||||
EXTRA_ARGS+=("--spec-draft-n-max" "$${MTP_DRAFT_N_MAX:-2}")
|
||||
echo "[entrypoint] MTP_ENABLED=1 — added --spec-type draft-mtp --spec-draft-n-max $${MTP_DRAFT_N_MAX:-2}"
|
||||
fi
|
||||
|
||||
exec /app/llama-server "$$@" "$${EXTRA_ARGS[@]}"
|
||||
- --
|
||||
command:
|
||||
@@ -123,7 +157,7 @@ services:
|
||||
- -b
|
||||
- ${BATCH_SIZE:-4096}
|
||||
- -ub
|
||||
- ${UBATCH_SIZE:-2048}
|
||||
- ${UBATCH_SIZE:-1024}
|
||||
- -ngl
|
||||
- "99"
|
||||
- -fa
|
||||
@@ -134,9 +168,19 @@ services:
|
||||
- ${KV_TYPE:-q4_0}
|
||||
- -np
|
||||
- "1"
|
||||
# Native GGUF chat template (no --chat-template-file). Thinking-off
|
||||
# via --reasoning (modern llama.cpp lever) on the native template's
|
||||
# Qwen3.6 thinking_mode gate. --reasoning-format kept on `deepseek`
|
||||
# as hygiene — with REASONING=off the model emits no <think>, so
|
||||
# `content` is clean for all clients (incl. opencode, #97 obsolete).
|
||||
- --jinja
|
||||
- --reasoning
|
||||
- ${REASONING:-off}
|
||||
- --reasoning-format
|
||||
- ${REASONING_FORMAT:-none}
|
||||
- ${REASONING_FORMAT:-deepseek}
|
||||
environment:
|
||||
- MTP_ENABLED=${MTP_ENABLED:-0}
|
||||
- MTP_DRAFT_N_MAX=${MTP_DRAFT_N_MAX:-2}
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth MTP-enabled Q4_K_M GGUF) + mmproj-F16
|
||||
# Engine: llama.cpp (local MTP-enabled build, build 9235)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: MTP n=2 (--spec-type draft-mtp, sweet spot)
|
||||
# KV: q4_0 K + q4_0 V
|
||||
# Vision: YES (mmproj-F16 mounted — multimodal input enabled)
|
||||
# Max ctx: 49152 (safe-headroom max with vision on 24 GB; sweep-verified)
|
||||
# Genesis: N/A — llama.cpp engine
|
||||
# Status: ✅ Production (NEW — first stack profile combining MTP + vision)
|
||||
# Engine-profile: llama-cpp-local
|
||||
# Best for: Multimodal chat, screenshot-debugging, vision-aware code review.
|
||||
# ~51/60 TPS text + vision overhead paid once at image-encode.
|
||||
# ---------------------------------------------------------------------------
|
||||
# Qwen3.6-27B on llama.cpp — single 3090, MTP, 49K ctx, VISION ON.
|
||||
#
|
||||
# This is one of three named single-card profiles for this model:
|
||||
# - docker-compose.yml → vanilla 262K + vision, cliff-immune fallback
|
||||
# - mtp.yml → MTP n=2 + 131K + no vision (fast + ctx)
|
||||
# - mtp-vision.yml → MTP n=2 + 49K + vision (THIS file — multimodal + fast)
|
||||
#
|
||||
# Historical note: the older single compose stripped --mmproj when MTP was
|
||||
# enabled, treating MTP + vision as incompatible. **That rule was obsolete
|
||||
# on build 9235** — sweep-verified 2026-05-19 that MTP + vision coexist
|
||||
# cleanly, this profile is the first to ship the combination.
|
||||
#
|
||||
# Stack-wide thinking-off-by-default policy applies (same as mtp.yml). Native
|
||||
# template (NOT froggeric) + `--reasoning off`. Cliff-survival: -ub 1024.
|
||||
#
|
||||
# VRAM budget on 24 GB (Q4_K_M + vision):
|
||||
# weights (Q4_K_M): ~17.0 GB
|
||||
# KV at 49K (q4_0 K+V): ~1.9 GB
|
||||
# mmproj F16 (vision projector): ~0.8 GB
|
||||
# MTP draft head + overhead: ~0.5 GB
|
||||
# total: ~20.2 GB
|
||||
# headroom: ~2.0 GB for prompt + activation peaks +
|
||||
# image-encode tile buffer
|
||||
#
|
||||
# Why 49K ctx (vs mtp.yml's 131K): the mmproj F16 plus image-encode tile
|
||||
# buffers consume the headroom that the no-vision profile spends on KV.
|
||||
# 49K is the safe-headroom ceiling sweep-verified at this config; pushing
|
||||
# higher risks OOM on the first vision request.
|
||||
#
|
||||
# Quick start:
|
||||
# 1. Get the MTP-enabled GGUF + the vision projector:
|
||||
# hf download unsloth/Qwen3.6-27B-MTP-GGUF Qwen3.6-27B-Q4_K_M.gguf \
|
||||
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-mtp-q4km
|
||||
# hf download unsloth/Qwen3.6-27B-GGUF mmproj-F16.gguf \
|
||||
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf
|
||||
# 2. From this directory:
|
||||
# MODEL_DIR=/your/models/dir docker compose -f mtp-vision.yml up -d
|
||||
# 3. curl http://localhost:8020/v1/models → should list the model.
|
||||
#
|
||||
# Override defaults via .env or shell:
|
||||
# MODEL_DIR host dir to mount as /models
|
||||
# GGUF_FILE path under /models (default: …unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf)
|
||||
# MMPROJ_FILE path under /models (default: qwen3.6-27b-gguf/mmproj-F16.gguf)
|
||||
# CTX_SIZE KV pool size (default: 49152 — safe-headroom max with vision)
|
||||
# BATCH_SIZE llama.cpp -b (default: 4096)
|
||||
# UBATCH_SIZE llama.cpp -ub (default: 1024 — cliff-survival)
|
||||
# KV_TYPE K and V quant type (default: q4_0)
|
||||
# MTP_DRAFT_N_MAX MTP draft tokens (default: 2)
|
||||
# REASONING thinking gate (default: off)
|
||||
# REASONING_FORMAT reasoning routing (default: deepseek)
|
||||
# PORT host port (default: 8020)
|
||||
# CUDA_VISIBLE_DEVICES which GPU to use (default: 0)
|
||||
|
||||
services:
|
||||
llama-cpp-qwen36-27b-mtp-vision:
|
||||
image: llama-cpp:local
|
||||
container_name: "${ESTATE_CONTAINER:-llama-cpp-qwen36-27b}"
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${ESTATE_PORT:-${PORT:-8020}}:8080"
|
||||
volumes:
|
||||
- "${MODEL_DIR:-../../../../../models-cache}:/models:ro"
|
||||
command: >-
|
||||
--host 0.0.0.0
|
||||
--port 8080
|
||||
-m /models/${GGUF_FILE:-qwen3.6-27b-gguf/unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf}
|
||||
--mmproj /models/${MMPROJ_FILE:-qwen3.6-27b-gguf/mmproj-F16.gguf}
|
||||
-c ${CTX_SIZE:-49152}
|
||||
-b ${BATCH_SIZE:-4096}
|
||||
-ub ${UBATCH_SIZE:-1024}
|
||||
-ngl 99
|
||||
-fa on
|
||||
--cache-type-k ${KV_TYPE:-q4_0}
|
||||
--cache-type-v ${KV_TYPE:-q4_0}
|
||||
-np 1
|
||||
--spec-type draft-mtp
|
||||
--spec-draft-n-max ${MTP_DRAFT_N_MAX:-2}
|
||||
--jinja
|
||||
--reasoning ${REASONING:-off}
|
||||
--reasoning-format ${REASONING_FORMAT:-deepseek}
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${ESTATE_GPUS:-${CUDA_VISIBLE_DEVICES:-0}}"]
|
||||
capabilities: [compute, utility]
|
||||
@@ -0,0 +1,105 @@
|
||||
# ===========================================================================
|
||||
# Profile (at-a-glance):
|
||||
# Model: Qwen3.6-27B (Unsloth MTP-enabled Q4_K_M GGUF)
|
||||
# Engine: llama.cpp (local MTP-enabled build, build 9235)
|
||||
# Topology: Single 3090 (TP=1)
|
||||
# Drafter: MTP n=2 (--spec-type draft-mtp, sweet spot — see BENCHMARKS.md)
|
||||
# KV: q4_0 K + q4_0 V (densest mainline, Ampere-fast)
|
||||
# Vision: no (mmproj NOT mounted — for vision, see mtp-vision.yml)
|
||||
# Max ctx: 131072 (safe-headroom max on 24 GB; sweep-verified)
|
||||
# Genesis: N/A — llama.cpp engine; Genesis is vLLM/Qwen3-Next-specific
|
||||
# Status: ✅ Production
|
||||
# Engine-profile: llama-cpp-local
|
||||
# Best for: IDE agents, opencode, Hermes, long-multi-turn agentic — the
|
||||
# speed + ctx workhorse. ~51 narr / ~60 code TPS, 7/7 verify-stress
|
||||
# (incl. 60K + 91K needle), 102/150 quality (68%) on the 8-pack.
|
||||
# ---------------------------------------------------------------------------
|
||||
# Qwen3.6-27B on llama.cpp — single 3090, MTP, 131K ctx, no vision.
|
||||
#
|
||||
# This is one of three named single-card profiles for this model:
|
||||
# - docker-compose.yml → vanilla 262K + vision, cliff-immune fallback
|
||||
# - mtp.yml → MTP n=2 + 131K + no vision (THIS file — fast + ctx)
|
||||
# - mtp-vision.yml → MTP n=2 + 49K + vision (multimodal + fast)
|
||||
#
|
||||
# Why no vision here: mmproj F16 costs ~0.8 GB. Without it the MTP-safe ctx
|
||||
# ceiling jumps from ~49K to 131K (sweep-verified 2026-05-19 on build 9235).
|
||||
# If you don't need image input, this is the better MTP profile.
|
||||
#
|
||||
# Stack-wide thinking-off-by-default policy (matches the 22/24 vLLM composes
|
||||
# shipping chat_template_kwargs.enable_thinking=false). On llama.cpp the
|
||||
# right lever is `--reasoning off` against the *native* template (the one
|
||||
# embedded in the GGUF). The vLLM-only froggeric template is NOT mounted —
|
||||
# it has no llama.cpp thinking-mode Jinja hook and silently suppresses
|
||||
# `--reasoning off`. Native template + `--reasoning off` is the correct
|
||||
# llama.cpp interface; closes opencode hang (#97) cleanly without needing
|
||||
# the older `--reasoning-format none` workaround.
|
||||
#
|
||||
# Cliff-survival: -ub 1024 (lowered from the older 2048 default). The 25K
|
||||
# tool-prefill check in verify-stress fails at -ub 2048 on tight single-card
|
||||
# headroom; -ub 1024 halves the per-pass activation peak and the boundary
|
||||
# matrix goes 5/7 → 7/7. The "Cliff 2 single-prompt at 50–60K is
|
||||
# architectural" narrative is partially superseded by this config: at
|
||||
# -ub 1024 + 131K + MTP n=2 + thinking-off, verify-stress recalls needles
|
||||
# cleanly at 58K and 91K.
|
||||
#
|
||||
# VRAM budget on 24 GB (Q4_K_M):
|
||||
# weights (Q4_K_M): ~17.0 GB
|
||||
# KV at 131K (q4_0 K+V): ~5.0 GB
|
||||
# MTP draft head + overhead: ~0.5 GB
|
||||
# total: ~22.5 GB
|
||||
# headroom: ~1.6 GB for prompt + activation peaks
|
||||
#
|
||||
# Quick start:
|
||||
# 1. Get the MTP-enabled GGUF:
|
||||
# hf download unsloth/Qwen3.6-27B-MTP-GGUF Qwen3.6-27B-Q4_K_M.gguf \
|
||||
# --local-dir $MODEL_DIR/qwen3.6-27b-gguf/unsloth-mtp-q4km
|
||||
# 2. From this directory:
|
||||
# MODEL_DIR=/your/models/dir docker compose -f mtp.yml up -d
|
||||
# 3. curl http://localhost:8020/v1/models → should list the model.
|
||||
#
|
||||
# Override defaults via .env or shell:
|
||||
# MODEL_DIR host dir to mount as /models (default: ../../../../../models-cache)
|
||||
# GGUF_FILE path under /models (default: qwen3.6-27b-gguf/unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf)
|
||||
# CTX_SIZE KV pool size (default: 131072 — safe-headroom max)
|
||||
# BATCH_SIZE llama.cpp -b (default: 4096)
|
||||
# UBATCH_SIZE llama.cpp -ub (default: 1024 — cliff-survival)
|
||||
# KV_TYPE K and V quant type (default: q4_0)
|
||||
# MTP_DRAFT_N_MAX MTP draft tokens (default: 2 — sweet spot per BENCHMARKS.md)
|
||||
# REASONING thinking gate (default: off — stack-wide policy)
|
||||
# REASONING_FORMAT reasoning routing (default: deepseek — hygiene)
|
||||
# PORT host port (default: 8020)
|
||||
# CUDA_VISIBLE_DEVICES which GPU to use (default: 0)
|
||||
|
||||
services:
|
||||
llama-cpp-qwen36-27b-mtp:
|
||||
image: llama-cpp:local
|
||||
container_name: "${ESTATE_CONTAINER:-llama-cpp-qwen36-27b}"
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${ESTATE_PORT:-${PORT:-8020}}:8080"
|
||||
volumes:
|
||||
- "${MODEL_DIR:-../../../../../models-cache}:/models:ro"
|
||||
command: >-
|
||||
--host 0.0.0.0
|
||||
--port 8080
|
||||
-m /models/${GGUF_FILE:-qwen3.6-27b-gguf/unsloth-mtp-q4km/Qwen3.6-27B-Q4_K_M.gguf}
|
||||
-c ${CTX_SIZE:-131072}
|
||||
-b ${BATCH_SIZE:-4096}
|
||||
-ub ${UBATCH_SIZE:-1024}
|
||||
-ngl 99
|
||||
-fa on
|
||||
--cache-type-k ${KV_TYPE:-q4_0}
|
||||
--cache-type-v ${KV_TYPE:-q4_0}
|
||||
-np 1
|
||||
--spec-type draft-mtp
|
||||
--spec-draft-n-max ${MTP_DRAFT_N_MAX:-2}
|
||||
--jinja
|
||||
--reasoning ${REASONING:-off}
|
||||
--reasoning-format ${REASONING_FORMAT:-deepseek}
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${ESTATE_GPUS:-${CUDA_VISIBLE_DEVICES:-0}}"]
|
||||
capabilities: [compute, utility]
|
||||
@@ -228,7 +228,8 @@ declare -A LAUNCH_VARIANT_COMPOSE=(
|
||||
[vllm/gemma-mtp-tp1]="models/gemma-4-31b/vllm/compose/single/docker-compose.yml"
|
||||
[vllm/gemma-dflash]="models/gemma-4-31b/vllm/compose/dual/dflash.yml"
|
||||
[llamacpp/default]="models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml"
|
||||
[llamacpp/concurrent]="models/qwen3.6-27b/llama-cpp/compose/single/concurrent.yml"
|
||||
[llamacpp/mtp]="models/qwen3.6-27b/llama-cpp/compose/single/mtp.yml"
|
||||
[llamacpp/mtp-vision]="models/qwen3.6-27b/llama-cpp/compose/single/mtp-vision.yml"
|
||||
)
|
||||
declare -A LAUNCH_VARIANT_MODEL=(
|
||||
[vllm/default]="qwen3.6-27b" [vllm/long-vision]="qwen3.6-27b" [vllm/long-text]="qwen3.6-27b"
|
||||
@@ -238,7 +239,7 @@ declare -A LAUNCH_VARIANT_MODEL=(
|
||||
[vllm/dual-dflash-noviz]="qwen3.6-27b" [vllm/dual-nvlink]="qwen3.6-27b" [vllm/dual-nvlink-turbo]="qwen3.6-27b"
|
||||
[vllm/dual-nvlink-dflash]="qwen3.6-27b" [vllm/dual-nvlink-dflash-noviz]="qwen3.6-27b"
|
||||
[vllm/gemma-mtp]="gemma-4-31b" [vllm/gemma-mtp-tp1]="gemma-4-31b" [vllm/gemma-dflash]="gemma-4-31b"
|
||||
[llamacpp/default]="qwen3.6-27b" [llamacpp/concurrent]="qwen3.6-27b"
|
||||
[llamacpp/default]="qwen3.6-27b" [llamacpp/mtp]="qwen3.6-27b" [llamacpp/mtp-vision]="qwen3.6-27b"
|
||||
)
|
||||
declare -A LAUNCH_VARIANT_ENGINE=(
|
||||
[vllm/default]="vllm" [vllm/long-vision]="vllm" [vllm/long-text]="vllm" [vllm/long-text-no-mtp]="vllm"
|
||||
@@ -247,7 +248,7 @@ declare -A LAUNCH_VARIANT_ENGINE=(
|
||||
[vllm/dual-dflash-noviz]="vllm" [vllm/dual-nvlink]="vllm" [vllm/dual-nvlink-turbo]="vllm"
|
||||
[vllm/dual-nvlink-dflash]="vllm" [vllm/dual-nvlink-dflash-noviz]="vllm"
|
||||
[vllm/gemma-mtp]="vllm" [vllm/gemma-mtp-tp1]="vllm" [vllm/gemma-dflash]="vllm"
|
||||
[llamacpp/default]="llamacpp" [llamacpp/concurrent]="llamacpp"
|
||||
[llamacpp/default]="llamacpp" [llamacpp/mtp]="llamacpp" [llamacpp/mtp-vision]="llamacpp"
|
||||
)
|
||||
declare -A LAUNCH_VARIANT_KVCALC=(
|
||||
[vllm/default]="qwen3.6-27b:long-vision"
|
||||
@@ -271,7 +272,8 @@ declare -A LAUNCH_VARIANT_KVCALC=(
|
||||
[vllm/gemma-mtp-tp1]="gemma-4-31b:gemma-single"
|
||||
[vllm/gemma-dflash]="gemma-4-31b:gemma-dual-dflash"
|
||||
[llamacpp/default]="SKIP"
|
||||
[llamacpp/concurrent]="SKIP"
|
||||
[llamacpp/mtp]="SKIP"
|
||||
[llamacpp/mtp-vision]="SKIP"
|
||||
)
|
||||
LAUNCH_VARIANT_ORDER=(
|
||||
vllm/long-vision vllm/long-text vllm/long-text-no-mtp vllm/bounded-thinking
|
||||
@@ -279,7 +281,7 @@ LAUNCH_VARIANT_ORDER=(
|
||||
vllm/dual vllm/dual-turbo vllm/dual-dflash vllm/dual-dflash-noviz
|
||||
vllm/dual4 vllm/dual4-dflash
|
||||
vllm/gemma-mtp vllm/gemma-mtp-tp1 vllm/gemma-dflash
|
||||
llamacpp/default llamacpp/concurrent
|
||||
llamacpp/default llamacpp/mtp llamacpp/mtp-vision
|
||||
)
|
||||
|
||||
variant_hw_status() {
|
||||
@@ -1124,7 +1126,8 @@ declare -A LAUNCH_DEFAULT_PORT=(
|
||||
[vllm/gemma-mtp-tp1]=8031
|
||||
[vllm/gemma-dflash]=8032
|
||||
[llamacpp/default]=8020
|
||||
[llamacpp/concurrent]=8020
|
||||
[llamacpp/mtp]=8020
|
||||
[llamacpp/mtp-vision]=8020
|
||||
)
|
||||
declare -A LAUNCH_DEFAULT_CONTAINER=(
|
||||
[vllm/default]=vllm-qwen36-27b
|
||||
@@ -1148,7 +1151,8 @@ declare -A LAUNCH_DEFAULT_CONTAINER=(
|
||||
[vllm/gemma-mtp-tp1]=vllm-gemma-4-31b-mtp-tp1
|
||||
[vllm/gemma-dflash]=vllm-gemma-4-31b-dflash
|
||||
[llamacpp/default]=llama-cpp-qwen36-27b
|
||||
[llamacpp/concurrent]=llama-cpp-qwen36-27b-concurrent
|
||||
[llamacpp/mtp]=llama-cpp-qwen36-27b
|
||||
[llamacpp/mtp-vision]=llama-cpp-qwen36-27b
|
||||
)
|
||||
ENDPOINT_PORT="${PORT:-${LAUNCH_DEFAULT_PORT[$VARIANT]:-8020}}"
|
||||
ENDPOINT_URL="http://localhost:${ENDPOINT_PORT}"
|
||||
|
||||
@@ -227,16 +227,23 @@ COMPOSE_REGISTRY = {
|
||||
# Qwen 3.6 27B, llama.cpp single-card.
|
||||
"llamacpp/default": _entry(
|
||||
model="qwen3.6-27b", weights_variant="gguf", workload="long-ctx-single",
|
||||
engine="llama-cpp-mainline", drafter=None, kv_format="q4_0",
|
||||
engine="llama-cpp-local", drafter=None, kv_format="q4_0",
|
||||
tp=1, max_ctx=262144, max_num_seqs=1, mem_util=None,
|
||||
compose_path="models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml",
|
||||
default_port=8020,
|
||||
),
|
||||
"llamacpp/concurrent": _entry(
|
||||
model="qwen3.6-27b", weights_variant="gguf", workload="multi-stream-tenant",
|
||||
engine="llama-cpp-mainline", drafter=None, kv_format="q4_0",
|
||||
tp=1, max_ctx=192000, max_num_seqs=4, mem_util=None,
|
||||
compose_path="models/qwen3.6-27b/llama-cpp/compose/single/concurrent.yml",
|
||||
"llamacpp/mtp": _entry(
|
||||
model="qwen3.6-27b", weights_variant="gguf", workload="fast-chat",
|
||||
engine="llama-cpp-local", drafter="qwen-mtp-builtin", kv_format="q4_0",
|
||||
tp=1, max_ctx=131072, max_num_seqs=1, mem_util=None,
|
||||
compose_path="models/qwen3.6-27b/llama-cpp/compose/single/mtp.yml",
|
||||
default_port=8020,
|
||||
),
|
||||
"llamacpp/mtp-vision": _entry(
|
||||
model="qwen3.6-27b", weights_variant="gguf", workload="vision-coding",
|
||||
engine="llama-cpp-local", drafter="qwen-mtp-builtin", kv_format="q4_0",
|
||||
tp=1, max_ctx=49152, max_num_seqs=1, mem_util=None,
|
||||
compose_path="models/qwen3.6-27b/llama-cpp/compose/single/mtp-vision.yml",
|
||||
default_port=8020,
|
||||
),
|
||||
|
||||
|
||||
@@ -1,11 +1,11 @@
|
||||
schema_version: 1
|
||||
id: llama-cpp-mainline
|
||||
display_name: llama.cpp mainline CUDA server
|
||||
id: llama-cpp-local
|
||||
display_name: llama.cpp local CUDA server (MTP-enabled)
|
||||
type: llama.cpp
|
||||
stability: stable
|
||||
install:
|
||||
method: docker_image
|
||||
spec: ghcr.io/ggml-org/llama.cpp:server-cuda
|
||||
spec: llama-cpp:local
|
||||
min_sm: 6.0
|
||||
supported_model_families:
|
||||
- qwen3-next-hybrid
|
||||
@@ -15,11 +15,12 @@ supported_kv_formats:
|
||||
- q5_0
|
||||
- q8_0
|
||||
- k8v4
|
||||
supported_drafters: []
|
||||
supported_drafters:
|
||||
- mtp
|
||||
supported_weight_formats:
|
||||
- gguf
|
||||
required_overlays: []
|
||||
vendored_overlays: []
|
||||
required_genesis: false
|
||||
notes: "Cliff-immune single-card fallback for Qwen GGUF serving."
|
||||
notes: "Cliff-immune single-card fallback for Qwen GGUF serving. Built from ggerganov/llama.cpp @ d14ce3dab (MTP PR #22673)."
|
||||
|
||||
|
||||
@@ -191,7 +191,7 @@ def stratum2_profile_like(
|
||||
|
||||
Both paths: the named profile's `engine.type` must be `vllm` else
|
||||
structured refuse `unsupported-runtime-engine` (covers
|
||||
`llamacpp/default` engine=`llama-cpp-mainline`, mem_util=None).
|
||||
`llamacpp/default` engine=`llama-cpp-local`, mem_util=None).
|
||||
|
||||
Path A additionally:
|
||||
- the profile must be `[D]`-emittable: REUSE `[D]`'s scope-gate
|
||||
|
||||
@@ -41,7 +41,8 @@
|
||||
#
|
||||
# Single-card llama.cpp:
|
||||
# llamacpp/default Q3_K_XL + 262K + q4_0 KV + vision (max ctx, no cliffs)
|
||||
# llamacpp/concurrent Q3_K_XL + 192K pool + 4 parallel slots + vision
|
||||
# llamacpp/mtp Q4_K_M MTP + 131K + q4_0 KV (fast: ~60 TPS code; no vision)
|
||||
# llamacpp/mtp-vision Q4_K_M MTP + 49K + q4_0 KV + mmproj (fast + multimodal)
|
||||
#
|
||||
# Env overrides (rarely needed):
|
||||
# COMPOSE_BIN Default: "docker compose" (set to e.g. "podman compose" if needed)
|
||||
|
||||
@@ -41,14 +41,14 @@ assert_contains "$out" "vllm/minimal"
|
||||
assert_not_contains "$out" "vllm/long-text"
|
||||
|
||||
out="$(python3 "$HELPER" filter-candidates \
|
||||
--variants vllm/long-text,llamacpp/default,llamacpp/concurrent \
|
||||
--variants vllm/long-text,llamacpp/default,llamacpp/mtp \
|
||||
--model qwen3.6-27b \
|
||||
--gpu-spec "$GPU_3090" \
|
||||
--tp 1 \
|
||||
--pp 1 \
|
||||
--stable)"
|
||||
assert_contains "$out" "llamacpp/default"
|
||||
assert_contains "$out" "llamacpp/concurrent"
|
||||
assert_contains "$out" "llamacpp/mtp"
|
||||
assert_not_contains "$out" "vllm/long-text"
|
||||
|
||||
if out="$(python3 "$HELPER" validate-variant \
|
||||
|
||||
@@ -150,7 +150,7 @@ PY
|
||||
run_test "C4 engine KV support: llama.cpp rejects bf16 KV" <<'PY'
|
||||
from scripts.lib.profiles.compat import load_profiles, fits
|
||||
p = load_profiles()
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-mainline"], kv_format="bf16", weights_variant="gguf", tp=1, project_vram=False)
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-local"], kv_format="bf16", weights_variant="gguf", tp=1, project_vram=False)
|
||||
assert not r.valid
|
||||
assert any(reason.startswith("C4:") for reason in r.reasons), r.reasons
|
||||
PY
|
||||
@@ -210,7 +210,7 @@ PY
|
||||
run_test "C10 model family support rejected" <<'PY'
|
||||
from scripts.lib.profiles.compat import load_profiles, fits
|
||||
p = load_profiles()
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["gemma-4-31b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-mainline"], kv_format="q4_0", tp=1, project_vram=False)
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["gemma-4-31b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-local"], kv_format="q4_0", tp=1, project_vram=False)
|
||||
assert not r.valid
|
||||
assert any(reason.startswith("C10:") for reason in r.reasons), r.reasons
|
||||
PY
|
||||
@@ -244,7 +244,7 @@ PY
|
||||
run_test "C12 skipped for non-vLLM engines" <<'PY'
|
||||
from scripts.lib.profiles.compat import load_profiles, fits
|
||||
p = load_profiles()
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-mainline"], kv_format="q4_0", weights_variant="gguf", tp=1)
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-local"], kv_format="q4_0", weights_variant="gguf", tp=1)
|
||||
assert r.valid, r.reasons
|
||||
assert "C12" in r.diagnostics["constraints_skipped"]
|
||||
assert r.diagnostics["kv_calc_invoked"] is False
|
||||
@@ -261,7 +261,7 @@ PY
|
||||
run_test "C14 explicit unsupported weight variant rejected" <<'PY'
|
||||
from scripts.lib.profiles.compat import load_profiles, fits
|
||||
p = load_profiles()
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-mainline"], kv_format="q4_0", weights_variant="autoround_int4", tp=1, project_vram=False)
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-local"], kv_format="q4_0", weights_variant="autoround_int4", tp=1, project_vram=False)
|
||||
assert not r.valid
|
||||
assert any(reason.startswith("C14:") for reason in r.reasons), r.reasons
|
||||
PY
|
||||
@@ -277,7 +277,7 @@ PY
|
||||
run_test "weight fallthrough: Qwen llama.cpp resolves GGUF" <<'PY'
|
||||
from scripts.lib.profiles.compat import load_profiles, fits
|
||||
p = load_profiles()
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-mainline"], kv_format="q4_0", tp=1, project_vram=False)
|
||||
r = fits([p.hardware["rtx-3090"]], p.models["qwen3.6-27b"], p.workloads["long-ctx-single"], p.engines["llama-cpp-local"], kv_format="q4_0", tp=1, project_vram=False)
|
||||
assert r.valid, r.reasons
|
||||
assert r.weights_variant == "gguf"
|
||||
PY
|
||||
@@ -422,7 +422,7 @@ from scripts.lib.profiles.compat import load_profiles, InstanceSpec, validate_es
|
||||
p = load_profiles()
|
||||
instances = [
|
||||
InstanceSpec("llama-a", "llamacpp/default", (0,), 8020),
|
||||
InstanceSpec("llama-b", "llamacpp/concurrent", (1,), 8021),
|
||||
InstanceSpec("llama-b", "llamacpp/mtp", (1,), 8021),
|
||||
]
|
||||
r = validate_estate(instances, [p.hardware["rtx-3090"], p.hardware["rtx-3090"]], p, nvlink_active=False)
|
||||
assert r.valid, (r.cross_instance_failures, {k: v.reasons for k, v in r.per_instance.items()})
|
||||
|
||||
@@ -105,7 +105,7 @@ check(
|
||||
# ---------------------------------------------------------------------------
|
||||
# STRATUM 2
|
||||
# ---------------------------------------------------------------------------
|
||||
# non-vLLM --profile-like (llamacpp/default, engine=llama-cpp-mainline) ->
|
||||
# non-vLLM --profile-like (llamacpp/default, engine=llama-cpp-local) ->
|
||||
# unsupported-runtime-engine (both paths) before [C0]/[B].
|
||||
s2_llama = G.stratum2_profile_like("llamacpp/default", path="B")
|
||||
check(
|
||||
|
||||
@@ -39,7 +39,9 @@ configs_all = [
|
||||
("tools-text 75K\nfp8 IDE-agent", 53.32, 69.66, "single-vllm"),
|
||||
("bounded-thinking 180K\nstructured-CoT", 49.77, 65.80, "single-vllm"),
|
||||
("minimal\n(no spec-dec)", 32.41, 32.56, "single-vllm"),
|
||||
("llama.cpp Q3_K_XL\n262K + vision", 21.22, 20.79, "single-llama"),
|
||||
("llamacpp/mtp\nQ4_K_M MTP 131K", 51.28, 59.72, "single-llama"),
|
||||
("llamacpp/mtp-vision\nQ4_K_M MTP+vision 49K", 56.52, 66.17, "single-llama"),
|
||||
("llamacpp/default\nQ3_K_XL 262K + vision", 21.22, 20.79, "single-llama"),
|
||||
("llama.cpp Q4_K_M\n+ ngram-mod 32K",22.04, 26.11, "single-llama"),
|
||||
("Luce DFlash 3.6+3.6*\nTQ3, 65K, greedy", 40.00, 71.65, "single-luce-watch"),
|
||||
("dual.yml\n262K + vision", 69.05, 88.58, "dual-vllm"),
|
||||
@@ -104,7 +106,7 @@ def make_chart(configs, out_stem, title_subject, figsize):
|
||||
ax.set_xticks(x)
|
||||
ax.set_xticklabels(labels, fontsize=9)
|
||||
ax.set_ylabel("TPS (3 warm + 5 measured, canonical bench)", fontsize=10)
|
||||
ax.set_title(f"Qwen3.6-27B — measured TPS {title_subject} on noonghunna/club-3090 (2026-05-02)",
|
||||
ax.set_title(f"Qwen3.6-27B — measured TPS {title_subject} on noonghunna/club-3090 (updated 2026-05-20)",
|
||||
fontsize=12, pad=36)
|
||||
ax.set_ylim(0, max(max(narr), max(code)) * 1.30)
|
||||
ax.grid(axis="y", linestyle=":", alpha=0.4)
|
||||
@@ -120,7 +122,7 @@ def make_chart(configs, out_stem, title_subject, figsize):
|
||||
|
||||
substrate_parts = ["vLLM 0.20.1rc1.dev16+g7a1eb8ac2 + Genesis v7.69 dev (2db18df) + vllm#35975 backport"]
|
||||
if any(g == "single-llama" for g in groups):
|
||||
substrate_parts.append("llama.cpp mainline 0d0764dfd")
|
||||
substrate_parts.append("llama.cpp mainline d14ce3dab (build 9235, MTP)")
|
||||
if any(g == "single-luce-watch" for g in groups):
|
||||
substrate_parts.append("Luce DFlash dflash@f12a87c (greedy only)")
|
||||
substrate_parts.append("RTX 3090 sm_86, PCIe-only, 230W")
|
||||
|
||||