722f998ff31100a9bee2dfd5484e7f435b24b1cd
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fbf343129c |
docs: use \$MODEL_DIR placeholder, not the dev rig's /mnt/models/huggingface/
User docs were hardcoding the dev rig path (/mnt/models/huggingface/...) as
if it was canonical. It's not — cross-rig users have models at /data/models,
~/models, /mnt/nvme/llms, etc. Setting MODEL_DIR per their setup is the
intended UX (the compose already supports it via env-var default).
Updates:
- models/qwen3.6-27b/llama-cpp/README.md: download examples now use
\$MODEL_DIR/qwen3.6-27b-gguf/ instead of /mnt/models/huggingface/...
- models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml: header
comment uses \$MODEL_DIR/qwen3.6-27b-gguf/ for download examples + says
"MODEL_DIR=/your/models/dir docker compose up -d" instead of our path.
- docs/engines/LLAMA_CPP.md: same treatment + cleaned up Qwen3.5 + DFlash
draft path examples to also use \$MODEL_DIR.
- scripts/preflight.sh: hf download hint shows literal \${MODEL_DIR} so user
knows what to set, plus explicit "set MODEL_DIR first" line. Previously
echoed the resolved relative path (../../../../models-cache) which lands
outside the repo if pwd isn't the compose dir.
Caught by RobH589 in club-3090#116 — they hit the path-resolved-to-root-of-drive
case from the relative-path default. Closes the doc UX side; the compose's
env-var override mechanism was already correct.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
00366a58d7 |
reorg: services/ consolidation + gpu-mode under git + ComfyUI + pin tracker + path updates
Release / release (push) Failing after 53s
The dev rig had grown across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`,
`/home/wasif/`) and three repos (single-3090, dual-3090, club-3090). Disk-out
on `/` (97% used) forced a cleanup; rather than just prune, we consolidated
the layout across the whole stack while we were at it. This commit captures
what landed inside this repo.
Services consolidation:
- services/{ollama,openwebui,litellm,qdrant,searxng}/ migrated in from
/opt/ai/compose/<svc>/ (zero functional change — same docker-compose.yml).
- services/litellm/config.yaml rewritten: explicit routes for current
primaries (qwen3.6-27b-autoround → :8010, gemma-4-31b-autoround → :8030).
Removed `* → ollama/*` wildcard.
- services/comfyui/ migrated in (was /opt/ai/compose/comfyui/) — wired into
gpu-mode with full mutex against vLLM/SGLang.
scripts/gpu-mode.sh under git:
- Was a loose /opt/ai/gpu-mode.sh outside any repo. Symlinked at
/usr/local/bin/gpu-mode.
- Five Gemma 4 31B modes added: gemma, gemma-dflash, gemma-int8,
gemma-dflash-int8, gemma-awq.
- One ComfyUI mode (mutex with all LLM serving).
- prune / prune-all subcommands (safe image prune; aggressive variant adds
build cache --keep-storage 5GB + dangling networks).
- gpu-mode status now shows Docker disk + /var/lib/docker + /tmp sizes.
- compose_at() passes --env-file <repo>/.env so MODEL_DIR resolves
regardless of which compose dir gpu-mode cd's into. Fixes the recurring
"MODEL_DIR not set, defaulting to ../../../../../models-cache" warning.
- stderr no longer swallowed by compose_at() (real errors surface).
- Cross-model VRAM mutex: every Qwen mode stop_all_gemma + stop_comfyui
and vice-versa.
scripts/maintenance/ — new hygiene-tools subdir:
- list-image-pins.sh: engine-agnostic pin auditor. Scans every compose's
`image:` line, groups by `<repo>:<tag>`, flags pin-drift (multiple tags
per repo), ranks composes by patch surface.
Pin tracking:
- docs/UPSTREAM.md gains a "Pinned images" section: table of every pinned
image, why each pin exists, retirement candidate criteria.
- docs/NIGHTLY_BUMP_RUNBOOK.md (new): 7-step procedure for bumping pinned
engine images (scope → branch → patch survival → boot → verify-full +
verify-stress → bench delta → land → retire). Engine-specific notes
for vLLM nightly hashes, llama.cpp digest pinning, SGLang variants.
Path updates from the engine + model dir consolidation:
- /opt/ai/vllm-src/ → /opt/ai/engines/vllm/primary/
(in setup.sh, INTERNALS.md, several patch READMEs, docs/HARDWARE.md,
docs/FAQ.md, docs/DUAL_CARD.md, docs/UPSTREAM.md, models/qwen3.6-27b/
CHANGELOG.md)
- /mnt/models/gguf/qwen3.6-27b/ → /mnt/models/huggingface/qwen3.6-27b-gguf/
(in models/qwen3.6-27b/llama-cpp/{compose/single/*.yml, recipes/*.sh,
README.md}, docs/engines/LLAMA_CPP.md)
CHANGELOG.md narrative gap fill:
- 2026-05-10 entry for this reorg.
- 2026-05-09 entry for compose convention formalization (topology
promoted to dir level, profile schema, Status enum + Caveats, cliff
CI swap, Discord launch).
- 2026-05-08 entry for Gemma 4 INT8 PTH unblock + 262K validation.
- 2026-05-07 entry for power-cap-sweep campaign + HARDWARE.md cross-rig
charts + cross-rig benchmark rows.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
||
|
|
dec0f22dac |
docs(lucebox): record PRs #78 + #80 — dual-GPU PFlash + DFlash split shipped (May 2026)
Two @weicj PRs merged that change the lucebox-hub serving topology: - PR #78 (PFlash phase-split, merged 2026-05-02) — --pflash-gpu flag, persistent pflash_daemon. Validation: passing NIAH source ctx 24K → 262K (10.7× over single-card co-resident) on dual RTX 2080 Ti 22 GB. - PR #80 (DFlash target/draft split, merged 2026-05-04) — --target-gpu / --draft-gpu flags. Validation: 51.86 tok/s HE 10-prompt, AL 7.09, 44.3% accept on Qwen3.5-27B Q4 target + z-lab DFlash draft. This is heterogeneous spec-decode (each model on its own card), not weight-sharded TP. Removes the single-card co-residency limit that was the binding blocker for 2× 3090 users (target + draft + KV all competing for 24 GB → 65K max_ctx ceiling). Updated: - docs/UPSTREAM.md — Luce DFlash section gains a "🆕 Dual-GPU split landed" subsection with both PR links + @weicj's measured numbers. PFlash row status icon flipped from 🟡 to 🟢; "Re-evaluate" criteria reworked to focus on reproducing the 262K NIAH claim on 2× 3090. - docs/engines/LLAMA_CPP.md — added "🆕 Dual-GPU split" subsection under the existing DFlash recipe with the new flag-based recipe and carry-over caveat (Qwen3.6-27B draft still under training; the benefit applies primarily to Qwen3.5-27B + DFlash today). Memory updates (gitignored, not in this commit): - pflash_future_exploration.md — type=project, status flipped from "co-residency blocker" to "co-residency blocker addressed via dual-GPU; bench task #229 queued" - pflash_x_bounded_thinking_intersection.md — added 2026-05-04 update noting the new dual-GPU path and that the parked exploration is more concrete now Bench tracked at task #229 (queued, not executed yet — these PRs are hours old as of this commit). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
e9c658cbc6 |
fix(docs): replace dead luce-spec/llama-cpp-dflash links with Luce-Org/lucebox-hub
Reported in #39 by @clort81 — the `luce-spec/llama-cpp-dflash` repo returns 404. The DFlash work consolidated into Luce-Org/lucebox-hub (verified: github.com/Luce-Org/lucebox-hub returns 200, contains dflash/ + pflash/ subdirs and dflash/deps/llama.cpp submodule). Affected files: - docs/engines/README.md (2 link sites in comparison table) - docs/engines/LLAMA_CPP.md (4 sites: intro, "Pros" table, build clone command, "See also" links) - models/qwen3.6-27b/llama-cpp/README.md (2 link sites) Plus collateral updates: - Build clone path /opt/llama-cpp-dflash → /opt/lucebox-hub (matches the new repo name; was a 3-replace via path globbing) - HF model path luce-spec/dflash-qwen3.6-27b-N5 (401 gated) → z-lab/Qwen3.6-27B-DFlash (200 public, the actually-shipping draft) + local-dir adjusted to /mnt/models/huggingface/z-lab/... matching the canonical HF model path convention - `git clone --recurse-submodules` flag added since lucebox-hub uses submodules for its bundled llama.cpp fork (in dflash/deps/llama.cpp) Updates URL framing in user-facing prose to acknowledge that lucebox-hub is a separate harness containing a llama.cpp fork rather than just being a llama.cpp fork. The recipe section build commands should be re-verified against the lucebox-hub README before treating them as canonical — this commit only updates the URL/path; the multi- step build instructions in docs/engines/LLAMA_CPP.md may need a follow-up walkthrough. CHANGELOG references to luce-spec preserved as historical context (the links were valid at the time the CHANGELOG entries were written). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
5aa97a25d9 |
v0.20 migration + Genesis v7.65 dev tip + cold-start cache + env-var alignment
This branch migrates the entire vLLM stack from `dev205+g07351e088` + Genesis
v7.64 to `0.20.1rc1.dev16+g7a1eb8ac2` + Genesis v7.65 dev tip (commit
`d89a089`). v7.65 is on Sandermage's `dev` branch — explicitly the cross-rig
testing surface he requested in discussion #19; he'll merge dev→main once we
both confirm stable. Pin gates restated when that lands.
What changes
------------
Pin migration:
- vLLM image: nightly-07351e08... → nightly-7a1eb8ac2... (dev205 → v0.20.1rc1.dev16)
- Genesis: 64dd18b (v7.64) → d89a089 (v7.65 dev tip)
Sidecar churn:
- DROPPED: patch_pn12_ffn_pool_anchor.py (PN12 native on v0.20)
- DROPPED: patch_pn12_compile_safe_custom_op.py (Genesis P38B in-source hook)
- DROPPED: patch_fa_max_seqlen_clamp.py (Genesis PN17 + P15B)
- ADDED: patch_workspace_lock_disable.py (relaxes vllm#39226 strict assertion;
P98 covers same surface but auto-skips on v0.20 due to drift-marker false
positive — pending Sandermage marker fix)
Env-var alignment to Sandermage's PROD set (start_27b_int4_TQ_k8v4.sh@dev):
- FIXED naming bugs that silently no-op'd patches:
- PN9_INDEPENDENT_DRAFTER_ATT → _ATTN (was silently OFF)
- PN22 → PN22_LOCAL_ARGMAX_TP (was silently OFF)
- PN26_BLOCK_KV → PN26_SPARSE_V_BLOCK_KV (fell back to default 4, not 8)
- PN26_NUM_WARPS → PN26_SPARSE_V_NUM_WARPS
- PN26_THRESHOLD → PN26_SPARSE_V_THRESHOLD (fell back to default 0.001, not 0.01)
- ADDED explicit-OFFs to match Sander's PROD verbatim:
- P78_TOLIST_CAPTURE_GUARD=0 (we use our own patch_tolist_cudagraph.py)
- P81_FP8_BLOCK_SCALED_M_LE_8=0 (FP8-specific, no-op on TQ3)
- P82=0, P82_THRESHOLD_SINGLE=0.3
- Cap divergence (justified): PROFILE_RUN_CAP_M=4128 + PREALLOC_TOKEN_BUDGET=4128
(Sander uses 4096 — vLLM `interface.py:639` forces our config's Mamba
block_size to 4128 due to TQ3 + TP=1 page-size math; lower values
AssertionError at boot)
- Carry-forward (intentional): P4 (hybrid TQ required), P65 (TQ spec-CG
downgrade — pending v0.20 verification that #40880 closure makes it
redundant)
Cold-start cache mounts (closes #22):
- All 10 composes now mount torch_compile_cache + Triton cache from
`models/qwen3.6-27b/vllm/cache/`. First boot warms (~6 min); warm boot
drops to ~3.2 min (47% faster). Per-stage savings on long-text:
- Dynamo bytecode transform: 18s → 5s (-73%)
- torch.compile: 57s → 9s (-85%)
- Initial profiling/warmup: 51s → 7s (-87%)
Mamba block_size cap fix:
- v0.20 enforces `long_prefill_token_threshold >= block_size`; on hybrid
Mamba+TQ3, vLLM forces block_size=4128. Bumped GENESIS_PROFILE_RUN_CAP_M
and PREALLOC_TOKEN_BUDGET 4096→4128 across all 5 main composes.
Default 48K compose:
- Required workspace_lock_disable sidecar after initial v0.20 boot hit
vllm#39226 strict assertion. Caught during validation, fixed.
Context restored vs dev205 backoffs (validated 33K + 50K stress on v0.20):
- long-text: 185K → 214K (+16%)
- long-vision: 140K → 198K (+41%)
- bounded-thinking: 185K → 214K (+16%)
Bench results (n=5, results/v0.20-migration/):
- long-text 214K narr 49.74 / code 67.39 (CV 2.6/2.7%)
- long-vision 198K narr 50.32 / code 66.12 (CV 2.3/4.1%)
- bounded-thinking 214K narr 49.77 / code 65.80 (CV 1.4/2.3%)
- tools-text 75K (fp8) narr 53.32 / code 69.66 (CV 2.3/1.4%)
- dual-turbo 262K (TP=2) narr 58.33 / code 76.01 per-stream
269 TPS aggregate at n=4 streams (3.63x speedup)
- default 48K narr 48.82 / code 65.98 (n=3)
Validation: verify-full 8/8 on every variant. verify-stress 33K AND 50K
tool-prefill PASS on every variant — the cliff that fired on EVERY dev205
config no longer reproduces.
Docs + charts:
- README + SINGLE_CARD + DUAL_CARD + CLIFFS + EXAMPLES + STRUCTURED_COT
+ FAQ + UPSTREAM + 3 engine docs + model README + INTERNALS + CHANGELOG
all updated with new pin, ctx, TPS numbers, and "v0.20 unblock" section
- performance.{png,svg} + variants regenerated with measured TPS
- vram-budget.{png,svg} + variants regenerated with measured VRAM
- UPSTREAM tracker: 5 issues moved ✅ closed (PR #12, #13, #14, #15, P104
superseded by PN17 + P15B)
Issues addressed:
- #16 (Cliff 1 mech B leaks past PN12 on inductor-compiled FFN) — partial:
v0.20's revised TQ FA paths close the synthetic stress; PN25 (Sander's
proper compile-path opaque-op fix) is on dev but explicitly opt-in pending
worker-fork registration fix. Workarounds documented (tools-text fp8 path
/ --enforce-eager) until Sander ships PN25 default-on.
- #20 (launch.sh port + container-name mismatch) — already closed by
|
||
|
|
df91d641c4 |
push long-text/bounded-thinking back to 185K + 0.975; long-vision stays 140K + 0.95
After
|
||
|
|
383b5cc381 |
long-text/long-vision/bounded-thinking: middle-ground recovery 130K → 175K / 120K → 140K
After
|
||
|
|
d803278ebc |
docs + bounded-thinking: roll new context defaults across user-facing surfaces
Following
|
||
|
|
427d2f8aa9 |
docs+scripts+charts: propagate new ceilings (long-vision 198K, long-text 218K)
Sweep across all user-facing docs reflecting the post-PN12-anchor-fix ceilings established in |
||
|
|
17aff4ce05 |
LLAMA_CPP.md: add structural explanation of why prefill cliffs don't fire
User asked the obvious question: vLLM at 192K hits Cliff 1 on 25K tool prefills, but llama.cpp at 262K processes the same message cleanly — why? Three structural reasons documented: 1. ggml-cuda attention has no max_seqlen parameter; FA2 does 2. Static KV slab + dynamic workspace vs paged + varlen pre-alloc 3. Cudagraph capture is decode-only; no path for cap-leak Plus Cliff 2 doesn't fire because llama.cpp's Qwen3-Next GDN implementation uses online state updates instead of materializing the chunk_gated_delta_rule O(seq_len * chunk_size) intermediate. Reframes the 3-4× TPS gap as the necessary trade for batched worst-case-workspace optimization vs dynamic-shape per-call serving. This is the architectural defense of the two-routes launch frame. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> |
||
|
|
3fa33332ce |
Initial commit — club-3090: model-agnostic LLM serving recipes for RTX 3090
Consolidates and supersedes:
- noonghunna/qwen36-27b-single-3090
- noonghunna/qwen36-dual-3090
The two predecessor repos partitioned by card count (1× vs 2×). This
repo partitions by engine instead, which matches how users actually
decide ("vLLM or llama.cpp?" before "1 card or 2"). Card count becomes
a config variant within each engine.
Structure (model-agnostic from day 1):
docs/ cross-model engine + hardware docs
engines/ vLLM / llama.cpp / SGLang comparison + per-engine deep dives
HARDWARE.md Ampere SM 8.6+, NVLink, power, VRAM ceilings
GLOSSARY.md plain-language definitions
img/ illustrations (vram-budget.svg)
ARCHITECTURE.md how this stack thinks about LLM serving on 24 GB
models/<model-name>/ everything specific to a model
qwen3.6-27b/ today's only model
README.md / INTERNALS.md / USE_CASES.md / CHANGELOG.md
vllm/ vLLM-specific configs for this model
compose/ docker-compose files (single + dual variants)
patches/ tolist_cudagraph + Marlin pad notes
llama-cpp/ llama.cpp recipes for this model
recipes/ shell scripts (single-card default + 262K max-ctx)
sglang/ SGLang status (currently blocked)
scripts/ shared, model-aware
setup.sh bash setup.sh <model> → downloads + verifies
verify.sh / verify-full.sh smoke + functional tests
bench.sh canonical TPS bench
vLLM compose variants (all under models/qwen3.6-27b/vllm/compose/):
Single-card:
docker-compose.yml ⭐ DEFAULT — TQ3 + Genesis P65, 48K, 51/68 TPS
docker-compose.fast-chat.yml fp8 + 20K, 55/70 TPS — fastest at small ctx
docker-compose.tools-text.yml fp8 + 75K, 53/70 TPS — best for long single prompts
docker-compose.no-genesis-mtp.yml control variant
docker-compose.minimal.yml no spec-decode
Dual-card:
docker-compose.dual.yml ⭐ fp8 + 262K + MTP + vision, 71/89 TPS
docker-compose.dual-turbo.yml TQ3 + Genesis v7.14 — 4-stream concurrency
docker-compose.dual-dflash.yml DFlash N=5 + 185K + vision — 78/128 TPS
docker-compose.dual-dflash-noviz.yml DFlash + 200K text-only
llama.cpp recipes (under models/qwen3.6-27b/llama-cpp/recipes/):
single-card-default.sh Q4_K_M + 65K
single-card-max-ctx.sh Q4_K_M + q4_0 KV at full 262K — the standout recipe
Old repos remain readable for issue history + external links (Medium,
Reddit, Twitter, Sandermage's PR threads). New issues should be filed
here.
Credits in README. Apache 2.0.
|