Commit Graph
11 Commits
Author SHA1 Message Date
noonghunnaandClaude Opus 4.7 fbf343129c docs: use \$MODEL_DIR placeholder, not the dev rig's /mnt/models/huggingface/
User docs were hardcoding the dev rig path (/mnt/models/huggingface/...) as
if it was canonical. It's not — cross-rig users have models at /data/models,
~/models, /mnt/nvme/llms, etc. Setting MODEL_DIR per their setup is the
intended UX (the compose already supports it via env-var default).

Updates:
- models/qwen3.6-27b/llama-cpp/README.md: download examples now use
  \$MODEL_DIR/qwen3.6-27b-gguf/ instead of /mnt/models/huggingface/...
- models/qwen3.6-27b/llama-cpp/compose/single/docker-compose.yml: header
  comment uses \$MODEL_DIR/qwen3.6-27b-gguf/ for download examples + says
  "MODEL_DIR=/your/models/dir docker compose up -d" instead of our path.
- docs/engines/LLAMA_CPP.md: same treatment + cleaned up Qwen3.5 + DFlash
  draft path examples to also use \$MODEL_DIR.
- scripts/preflight.sh: hf download hint shows literal \${MODEL_DIR} so user
  knows what to set, plus explicit "set MODEL_DIR first" line. Previously
  echoed the resolved relative path (../../../../models-cache) which lands
  outside the repo if pwd isn't the compose dir.

Caught by RobH589 in club-3090#116 — they hit the path-resolved-to-root-of-drive
case from the relative-path default. Closes the doc UX side; the compose's
env-var override mechanism was already correct.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 17:13:29 +00:00
noonghunnaandClaude Opus 4.7 00366a58d7 reorg: services/ consolidation + gpu-mode under git + ComfyUI + pin tracker + path updates
Release / release (push) Failing after 53s
The dev rig had grown across three home dirs (`/opt/ai/compose/`, `/opt/ai/github/`,
`/home/wasif/`) and three repos (single-3090, dual-3090, club-3090). Disk-out
on `/` (97% used) forced a cleanup; rather than just prune, we consolidated
the layout across the whole stack while we were at it. This commit captures
what landed inside this repo.

Services consolidation:
- services/{ollama,openwebui,litellm,qdrant,searxng}/ migrated in from
  /opt/ai/compose/<svc>/ (zero functional change — same docker-compose.yml).
- services/litellm/config.yaml rewritten: explicit routes for current
  primaries (qwen3.6-27b-autoround → :8010, gemma-4-31b-autoround → :8030).
  Removed `* → ollama/*` wildcard.
- services/comfyui/ migrated in (was /opt/ai/compose/comfyui/) — wired into
  gpu-mode with full mutex against vLLM/SGLang.

scripts/gpu-mode.sh under git:
- Was a loose /opt/ai/gpu-mode.sh outside any repo. Symlinked at
  /usr/local/bin/gpu-mode.
- Five Gemma 4 31B modes added: gemma, gemma-dflash, gemma-int8,
  gemma-dflash-int8, gemma-awq.
- One ComfyUI mode (mutex with all LLM serving).
- prune / prune-all subcommands (safe image prune; aggressive variant adds
  build cache --keep-storage 5GB + dangling networks).
- gpu-mode status now shows Docker disk + /var/lib/docker + /tmp sizes.
- compose_at() passes --env-file <repo>/.env so MODEL_DIR resolves
  regardless of which compose dir gpu-mode cd's into. Fixes the recurring
  "MODEL_DIR not set, defaulting to ../../../../../models-cache" warning.
- stderr no longer swallowed by compose_at() (real errors surface).
- Cross-model VRAM mutex: every Qwen mode stop_all_gemma + stop_comfyui
  and vice-versa.

scripts/maintenance/ — new hygiene-tools subdir:
- list-image-pins.sh: engine-agnostic pin auditor. Scans every compose's
  `image:` line, groups by `<repo>:<tag>`, flags pin-drift (multiple tags
  per repo), ranks composes by patch surface.

Pin tracking:
- docs/UPSTREAM.md gains a "Pinned images" section: table of every pinned
  image, why each pin exists, retirement candidate criteria.
- docs/NIGHTLY_BUMP_RUNBOOK.md (new): 7-step procedure for bumping pinned
  engine images (scope → branch → patch survival → boot → verify-full +
  verify-stress → bench delta → land → retire). Engine-specific notes
  for vLLM nightly hashes, llama.cpp digest pinning, SGLang variants.

Path updates from the engine + model dir consolidation:
- /opt/ai/vllm-src/ → /opt/ai/engines/vllm/primary/
  (in setup.sh, INTERNALS.md, several patch READMEs, docs/HARDWARE.md,
   docs/FAQ.md, docs/DUAL_CARD.md, docs/UPSTREAM.md, models/qwen3.6-27b/
   CHANGELOG.md)
- /mnt/models/gguf/qwen3.6-27b/ → /mnt/models/huggingface/qwen3.6-27b-gguf/
  (in models/qwen3.6-27b/llama-cpp/{compose/single/*.yml, recipes/*.sh,
   README.md}, docs/engines/LLAMA_CPP.md)

CHANGELOG.md narrative gap fill:
- 2026-05-10 entry for this reorg.
- 2026-05-09 entry for compose convention formalization (topology
  promoted to dir level, profile schema, Status enum + Caveats, cliff
  CI swap, Discord launch).
- 2026-05-08 entry for Gemma 4 INT8 PTH unblock + 262K validation.
- 2026-05-07 entry for power-cap-sweep campaign + HARDWARE.md cross-rig
  charts + cross-rig benchmark rows.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-10 16:57:03 +00:00
noonghunnaandClaude Opus 4.7 dec0f22dac docs(lucebox): record PRs #78 + #80 — dual-GPU PFlash + DFlash split shipped (May 2026)
Two @weicj PRs merged that change the lucebox-hub serving topology:

- PR #78 (PFlash phase-split, merged 2026-05-02) — --pflash-gpu flag,
  persistent pflash_daemon. Validation: passing NIAH source ctx 24K →
  262K (10.7× over single-card co-resident) on dual RTX 2080 Ti 22 GB.
- PR #80 (DFlash target/draft split, merged 2026-05-04) — --target-gpu /
  --draft-gpu flags. Validation: 51.86 tok/s HE 10-prompt, AL 7.09,
  44.3% accept on Qwen3.5-27B Q4 target + z-lab DFlash draft.

This is heterogeneous spec-decode (each model on its own card), not
weight-sharded TP. Removes the single-card co-residency limit that was
the binding blocker for 2× 3090 users (target + draft + KV all
competing for 24 GB → 65K max_ctx ceiling).

Updated:
- docs/UPSTREAM.md — Luce DFlash section gains a "🆕 Dual-GPU split
  landed" subsection with both PR links + @weicj's measured numbers.
  PFlash row status icon flipped from 🟡 to 🟢; "Re-evaluate" criteria
  reworked to focus on reproducing the 262K NIAH claim on 2× 3090.

- docs/engines/LLAMA_CPP.md — added "🆕 Dual-GPU split" subsection
  under the existing DFlash recipe with the new flag-based recipe and
  carry-over caveat (Qwen3.6-27B draft still under training; the
  benefit applies primarily to Qwen3.5-27B + DFlash today).

Memory updates (gitignored, not in this commit):
- pflash_future_exploration.md — type=project, status flipped from
  "co-residency blocker" to "co-residency blocker addressed via
  dual-GPU; bench task #229 queued"
- pflash_x_bounded_thinking_intersection.md — added 2026-05-04 update
  noting the new dual-GPU path and that the parked exploration is
  more concrete now

Bench tracked at task #229 (queued, not executed yet — these PRs are
hours old as of this commit).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-04 14:15:10 +00:00
noonghunnaandClaude Opus 4.7 e9c658cbc6 fix(docs): replace dead luce-spec/llama-cpp-dflash links with Luce-Org/lucebox-hub
Reported in #39 by @clort81 — the `luce-spec/llama-cpp-dflash` repo
returns 404. The DFlash work consolidated into Luce-Org/lucebox-hub
(verified: github.com/Luce-Org/lucebox-hub returns 200, contains
dflash/ + pflash/ subdirs and dflash/deps/llama.cpp submodule).

Affected files:
- docs/engines/README.md (2 link sites in comparison table)
- docs/engines/LLAMA_CPP.md (4 sites: intro, "Pros" table, build clone
  command, "See also" links)
- models/qwen3.6-27b/llama-cpp/README.md (2 link sites)

Plus collateral updates:
- Build clone path /opt/llama-cpp-dflash → /opt/lucebox-hub (matches
  the new repo name; was a 3-replace via path globbing)
- HF model path luce-spec/dflash-qwen3.6-27b-N5 (401 gated) →
  z-lab/Qwen3.6-27B-DFlash (200 public, the actually-shipping draft)
  + local-dir adjusted to /mnt/models/huggingface/z-lab/... matching
  the canonical HF model path convention
- `git clone --recurse-submodules` flag added since lucebox-hub uses
  submodules for its bundled llama.cpp fork (in dflash/deps/llama.cpp)

Updates URL framing in user-facing prose to acknowledge that
lucebox-hub is a separate harness containing a llama.cpp fork rather
than just being a llama.cpp fork. The recipe section build commands
should be re-verified against the lucebox-hub README before treating
them as canonical — this commit only updates the URL/path; the multi-
step build instructions in docs/engines/LLAMA_CPP.md may need a
follow-up walkthrough.

CHANGELOG references to luce-spec preserved as historical context (the
links were valid at the time the CHANGELOG entries were written).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-03 09:44:59 +00:00
noonghunnaandClaude Opus 4.7 5aa97a25d9 v0.20 migration + Genesis v7.65 dev tip + cold-start cache + env-var alignment
This branch migrates the entire vLLM stack from `dev205+g07351e088` + Genesis
v7.64 to `0.20.1rc1.dev16+g7a1eb8ac2` + Genesis v7.65 dev tip (commit
`d89a089`). v7.65 is on Sandermage's `dev` branch — explicitly the cross-rig
testing surface he requested in discussion #19; he'll merge dev→main once we
both confirm stable. Pin gates restated when that lands.

What changes
------------

Pin migration:
- vLLM image: nightly-07351e08... → nightly-7a1eb8ac2... (dev205 → v0.20.1rc1.dev16)
- Genesis: 64dd18b (v7.64) → d89a089 (v7.65 dev tip)

Sidecar churn:
- DROPPED: patch_pn12_ffn_pool_anchor.py (PN12 native on v0.20)
- DROPPED: patch_pn12_compile_safe_custom_op.py (Genesis P38B in-source hook)
- DROPPED: patch_fa_max_seqlen_clamp.py (Genesis PN17 + P15B)
- ADDED: patch_workspace_lock_disable.py (relaxes vllm#39226 strict assertion;
  P98 covers same surface but auto-skips on v0.20 due to drift-marker false
  positive — pending Sandermage marker fix)

Env-var alignment to Sandermage's PROD set (start_27b_int4_TQ_k8v4.sh@dev):
- FIXED naming bugs that silently no-op'd patches:
  - PN9_INDEPENDENT_DRAFTER_ATT → _ATTN (was silently OFF)
  - PN22 → PN22_LOCAL_ARGMAX_TP (was silently OFF)
  - PN26_BLOCK_KV → PN26_SPARSE_V_BLOCK_KV (fell back to default 4, not 8)
  - PN26_NUM_WARPS → PN26_SPARSE_V_NUM_WARPS
  - PN26_THRESHOLD → PN26_SPARSE_V_THRESHOLD (fell back to default 0.001, not 0.01)
- ADDED explicit-OFFs to match Sander's PROD verbatim:
  - P78_TOLIST_CAPTURE_GUARD=0 (we use our own patch_tolist_cudagraph.py)
  - P81_FP8_BLOCK_SCALED_M_LE_8=0 (FP8-specific, no-op on TQ3)
  - P82=0, P82_THRESHOLD_SINGLE=0.3
- Cap divergence (justified): PROFILE_RUN_CAP_M=4128 + PREALLOC_TOKEN_BUDGET=4128
  (Sander uses 4096 — vLLM `interface.py:639` forces our config's Mamba
  block_size to 4128 due to TQ3 + TP=1 page-size math; lower values
  AssertionError at boot)
- Carry-forward (intentional): P4 (hybrid TQ required), P65 (TQ spec-CG
  downgrade — pending v0.20 verification that #40880 closure makes it
  redundant)

Cold-start cache mounts (closes #22):
- All 10 composes now mount torch_compile_cache + Triton cache from
  `models/qwen3.6-27b/vllm/cache/`. First boot warms (~6 min); warm boot
  drops to ~3.2 min (47% faster). Per-stage savings on long-text:
  - Dynamo bytecode transform: 18s → 5s (-73%)
  - torch.compile: 57s → 9s (-85%)
  - Initial profiling/warmup: 51s → 7s (-87%)

Mamba block_size cap fix:
- v0.20 enforces `long_prefill_token_threshold >= block_size`; on hybrid
  Mamba+TQ3, vLLM forces block_size=4128. Bumped GENESIS_PROFILE_RUN_CAP_M
  and PREALLOC_TOKEN_BUDGET 4096→4128 across all 5 main composes.

Default 48K compose:
- Required workspace_lock_disable sidecar after initial v0.20 boot hit
  vllm#39226 strict assertion. Caught during validation, fixed.

Context restored vs dev205 backoffs (validated 33K + 50K stress on v0.20):
- long-text:        185K → 214K (+16%)
- long-vision:      140K → 198K (+41%)
- bounded-thinking: 185K → 214K (+16%)

Bench results (n=5, results/v0.20-migration/):
- long-text 214K        narr 49.74 / code 67.39 (CV 2.6/2.7%)
- long-vision 198K      narr 50.32 / code 66.12 (CV 2.3/4.1%)
- bounded-thinking 214K narr 49.77 / code 65.80 (CV 1.4/2.3%)
- tools-text 75K (fp8)  narr 53.32 / code 69.66 (CV 2.3/1.4%)
- dual-turbo 262K (TP=2) narr 58.33 / code 76.01 per-stream
                         269 TPS aggregate at n=4 streams (3.63x speedup)
- default 48K           narr 48.82 / code 65.98 (n=3)

Validation: verify-full 8/8 on every variant. verify-stress 33K AND 50K
tool-prefill PASS on every variant — the cliff that fired on EVERY dev205
config no longer reproduces.

Docs + charts:
- README + SINGLE_CARD + DUAL_CARD + CLIFFS + EXAMPLES + STRUCTURED_COT
  + FAQ + UPSTREAM + 3 engine docs + model README + INTERNALS + CHANGELOG
  all updated with new pin, ctx, TPS numbers, and "v0.20 unblock" section
- performance.{png,svg} + variants regenerated with measured TPS
- vram-budget.{png,svg} + variants regenerated with measured VRAM
- UPSTREAM tracker: 5 issues moved ✅ closed (PR #12, #13, #14, #15, P104
  superseded by PN17 + P15B)

Issues addressed:
- #16 (Cliff 1 mech B leaks past PN12 on inductor-compiled FFN) — partial:
  v0.20's revised TQ FA paths close the synthetic stress; PN25 (Sander's
  proper compile-path opaque-op fix) is on dev but explicitly opt-in pending
  worker-fork registration fix. Workarounds documented (tools-text fp8 path
  / --enforce-eager) until Sander ships PN25 default-on.
- #20 (launch.sh port + container-name mismatch) — already closed by
  77ca576 (post-issue-filing).
- #22 (cold-start caching) — closed by cache mounts above.

Remaining caveats:
- Cliff 2 (DeltaNet GDN forward, single prompt ≥50-60K) unchanged —
  architectural, applies to all single-card vLLM TQ3 paths. Mitigation:
  dual-turbo TP=2 (state splits across cards) or llama.cpp 262K.
- Default 48K narr_TPS (48.82) slightly under chart's 55 reference —
  bench variance + sample size n=3; not regression.
- Dual.yml / dual-dflash* not re-benched on v0.20; numbers carry forward
  from dev205 (fp8 paths were not TPS-changed by the migration).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-01 18:33:01 +00:00
noonghunnaandClaude Opus 4.7 df91d641c4 push long-text/bounded-thinking back to 185K + 0.975; long-vision stays 140K + 0.95
After 383b5cc shipped 175K + 0.97 (text) and 140K + 0.95 (vision),
audit showed the cliffs the backoff was protecting against fire on
every config we ship — they're independent of max-model-len. So the
context capacity was wasted protection.

Push text-only ceilings up:
  long-text:        175K + 0.97  → 185K + 0.975
  bounded-thinking: 175K + 0.97  → 185K + 0.975

Vision stays at 140K + 0.95: tried 185K + 0.98, 185K + 0.975, 160K
+ 0.97; all reopened Cliff 2 (DeltaNet GDN forward buffer) at the
130K-char stress class. Vision tower's ~1 GiB persistent + the new
patches' persistent allocations (P38 K_full/V_full ~750 MiB at 185K
+ compile-safe sidecar ~138 MiB) leave too little headroom for the
GDN intermediate buffer at 30K+ token prefills on this variant.
P37 disabled on vision (was on for parity with long-text but P37's
MoE intermediate cache pool is no-op on dense Qwen3.6-27B and the
env gate doesn't free memory anyway).

Verification at the new ceilings:
  long-text 185K + 0.975:    verify-full 8/8 (MTP AL 2.66),
                              130K-char tool-prefill stress PASS
  long-vision 140K + 0.95:   verify-full 8/8 (MTP AL 3.27),
                              130K-char tool-prefill stress PASS
  bounded-thinking 185K + 0.975: not re-booted in this final state
                                  (config identical to long-text +
                                  one --structured-outputs flag,
                                  no memory delta expected)

Docs updated: SINGLE_CARD.md picker table + activation-budget +
per-variant blurbs; engines/VLLM.md TL;DR + KV cache table; engines/
LLAMA_CPP.md "when to use vLLM"; STRUCTURED_COT.md "When to pick
this over long-text"; models/qwen3.6-27b/README.md per-variant lines;
docs/CLIFFS.md "Update 2026-05-01 PM" with full bisection sweep and
final decision.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-01 11:15:01 +00:00
noonghunnaandClaude Opus 4.7 383b5cc381 long-text/long-vision/bounded-thinking: middle-ground recovery 130K → 175K / 120K → 140K
After d803278 (130K + 0.95 / 120K + 0.94) shipped, audit surfaced that the
backoff was driven by a synthetic 200K-char (50K-token) single-shot stress
that's heavier than typical agent workloads (ampersandru's repro was ~30K
real tokens; VolandBerlioz's was similar). Realistic agent workloads stay
in the 130K-char (33K-token) class which both 130K + 0.95 and 175K + 0.97
pass.

Recovery: middle-ground configs that keep ~360-720 MiB activation headroom
over the original 0.985 / 0.98 mem-util but recover meaningful context.

  long-text:        130K + 0.95 → 175K + 0.97   verify-full 8/8 (AL 2.87),
                                                 130K-char stress PASS
  bounded-thinking: 130K + 0.95 → 175K + 0.97   parity with long-text
                                                 (verified earlier in #134)
  long-vision:      120K + 0.94 → 140K + 0.95   verify-full 8/8 (AL 2.49),
                                                 130K-char stress PASS
                                                 (intermediate 160K + 0.96
                                                 booted but failed 130K
                                                 stress on vision tower
                                                 overhead; 150K + 0.95
                                                 wouldn't boot — engine
                                                 ceiling at 0.95 vision
                                                 is 140352)

200K-char (50K-token) single-shot synthetic stress still cliffs on all
three — that's the FA varlen workspace allocation we can't reach. The bar
that matters for real users (verify-full + 130K-char stress) is met.

Docs updated: SINGLE_CARD.md picker table + activation-budget rationale +
per-variant blurbs; engines/VLLM.md TL;DR + KV cache table; engines/
LLAMA_CPP.md "when to use vLLM instead"; STRUCTURED_COT.md "When to pick
this over long-text"; models/qwen3.6-27b/README.md per-variant lines.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-01 09:38:03 +00:00
noonghunnaandClaude Opus 4.7 d803278ebc docs + bounded-thinking: roll new context defaults across user-facing surfaces
Following 1a931b4 (long-text 130K + 0.95, long-vision 120K + 0.94), this
brings the rest of the user-visible surface in line:

bounded-thinking.yml gets the same backoff (was 218K + 0.985 → 130K + 0.95)
plus full patch parity with long-text (P37, PN17, compile-safe sidecar
mount + apply step, P104 already present).

User-facing docs updated:
- engines/VLLM.md TL;DR + KV cache table commentary.
- engines/LLAMA_CPP.md "when to use vLLM instead" (was citing 218K
  text-only; now 130K).
- STRUCTURED_COT.md "When to pick this over the standard long-text"
  (was 218K; now 130K).
- SINGLE_CARD.md picker table, the prominent ⚠️ box, the activation-
  budget rationale, and the long-vision / long-text per-variant blurbs.
- models/qwen3.6-27b/README.md long-text/long-vision/bounded-thinking
  one-liners.

Historical references (CLIFFS.md "Update 2026-04-30 PM" section, etc.)
left intact as record of what shipped at each pin.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-01 02:48:51 +00:00
noonghunnaandClaude Opus 4.7 427d2f8aa9 docs+scripts+charts: propagate new ceilings (long-vision 198K, long-text 218K)
Sweep across all user-facing docs reflecting the post-PN12-anchor-fix
ceilings established in 287de1c → f3e5b52:

Docs touched:
- models/qwen3.6-27b/README.md — VRAM allocation paragraph + What's not
  working list + Genesis patches table (PN12/PN13/P101/P103 added)
- models/qwen3.6-27b/INTERNALS.md — forward-looking note pointing at
  CLIFFS.md for current state
- models/qwen3.6-27b/CHANGELOG.md — new 2026-04-30 PM entry
- docs/SINGLE_CARD.md — TL;DR table, VRAM budget bullet, frontier-context
  section, cliff status footer
- docs/FAQ.md — vLLM-vs-llama.cpp framing, ctx-drop question, Cliff 1/2
  explanations, troubleshooting list
- docs/engines/README.md — engine comparison table
- docs/engines/VLLM.md — feature bullets, TQ3 table, ctx tier description
- docs/engines/LLAMA_CPP.md — "why no cliffs" framing, single-card switch
  decision

Charts regenerated:
- tools/charts/gen-perf.py — labels updated (long-vision 198K, long-text 218K)
- tools/charts/gen-vram.py — added 218K text-only row, relabeled 198K row
  with mem-util note. SVG/PNG outputs regenerated via Docker matplotlib.

Scripts:
- scripts/launch.sh — wizard option labels
- scripts/switch.sh — header documentation

Cliff 1 status across all variants: closed.
Cliff 2 status: still applies single-prompt >50–60K on single-card.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-30 12:59:57 +00:00
noonghunnaandClaude Opus 4.7 17aff4ce05 LLAMA_CPP.md: add structural explanation of why prefill cliffs don't fire
User asked the obvious question: vLLM at 192K hits Cliff 1 on 25K
tool prefills, but llama.cpp at 262K processes the same message
cleanly — why?

Three structural reasons documented:
1. ggml-cuda attention has no max_seqlen parameter; FA2 does
2. Static KV slab + dynamic workspace vs paged + varlen pre-alloc
3. Cudagraph capture is decode-only; no path for cap-leak

Plus Cliff 2 doesn't fire because llama.cpp's Qwen3-Next GDN
implementation uses online state updates instead of materializing
the chunk_gated_delta_rule O(seq_len * chunk_size) intermediate.

Reframes the 3-4× TPS gap as the necessary trade for batched
worst-case-workspace optimization vs dynamic-shape per-call serving.
This is the architectural defense of the two-routes launch frame.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-29 22:22:43 +00:00
noonghunna 3fa33332ce Initial commit — club-3090: model-agnostic LLM serving recipes for RTX 3090
Consolidates and supersedes:
  - noonghunna/qwen36-27b-single-3090
  - noonghunna/qwen36-dual-3090

The two predecessor repos partitioned by card count (1× vs 2×). This
repo partitions by engine instead, which matches how users actually
decide ("vLLM or llama.cpp?" before "1 card or 2"). Card count becomes
a config variant within each engine.

Structure (model-agnostic from day 1):

  docs/                       cross-model engine + hardware docs
    engines/                    vLLM / llama.cpp / SGLang comparison + per-engine deep dives
    HARDWARE.md                 Ampere SM 8.6+, NVLink, power, VRAM ceilings
    GLOSSARY.md                 plain-language definitions
    img/                        illustrations (vram-budget.svg)
    ARCHITECTURE.md             how this stack thinks about LLM serving on 24 GB

  models/<model-name>/        everything specific to a model
    qwen3.6-27b/                today's only model
      README.md / INTERNALS.md / USE_CASES.md / CHANGELOG.md
      vllm/                     vLLM-specific configs for this model
        compose/                  docker-compose files (single + dual variants)
        patches/                  tolist_cudagraph + Marlin pad notes
      llama-cpp/                llama.cpp recipes for this model
        recipes/                  shell scripts (single-card default + 262K max-ctx)
      sglang/                   SGLang status (currently blocked)

  scripts/                    shared, model-aware
    setup.sh                    bash setup.sh <model> → downloads + verifies
    verify.sh / verify-full.sh  smoke + functional tests
    bench.sh                    canonical TPS bench

vLLM compose variants (all under models/qwen3.6-27b/vllm/compose/):

  Single-card:
    docker-compose.yml             ⭐ DEFAULT — TQ3 + Genesis P65, 48K, 51/68 TPS
    docker-compose.fast-chat.yml   fp8 + 20K, 55/70 TPS — fastest at small ctx
    docker-compose.tools-text.yml  fp8 + 75K, 53/70 TPS — best for long single prompts
    docker-compose.no-genesis-mtp.yml control variant
    docker-compose.minimal.yml     no spec-decode

  Dual-card:
    docker-compose.dual.yml             ⭐ fp8 + 262K + MTP + vision, 71/89 TPS
    docker-compose.dual-turbo.yml       TQ3 + Genesis v7.14 — 4-stream concurrency
    docker-compose.dual-dflash.yml      DFlash N=5 + 185K + vision — 78/128 TPS
    docker-compose.dual-dflash-noviz.yml DFlash + 200K text-only

llama.cpp recipes (under models/qwen3.6-27b/llama-cpp/recipes/):

  single-card-default.sh    Q4_K_M + 65K
  single-card-max-ctx.sh    Q4_K_M + q4_0 KV at full 262K — the standout recipe

Old repos remain readable for issue history + external links (Medium,
Reddit, Twitter, Sandermage's PR threads). New issues should be filed
here.

Credits in README. Apache 2.0.
2026-04-28 10:24:14 +00:00