Files
club-3090/docs/MULTI_CARD.md
T
noonghunnaandClaude Opus 4.7 acd7ffb67c restructure: promote topology to a directory level (single/dual/multi4)
Compose files now live under `<model>/<engine>/compose/<topology>/<file>.yml`,
with topology as a folder rather than a filename prefix. Solves all 7
inconsistencies surfaced in the post-rename audit (single-card composes
without `single-` prefix, unsuffixed `docker-compose.yml` ambiguity,
fine-tunes encoding model name in filename, etc.) by making the directory
hierarchy enforce the convention.

Layout:
  models/<model>/<engine>/compose/<topology>/<feature>.yml

Where:
  - <model>:    qwen3.6-27b, gemma-4-31b
  - <engine>:   vllm, llama-cpp, sglang
  - <topology>: single, dual, multi3, multi4, multi8
  - <feature>:  docker-compose.yml (default) | turbo.yml | dflash.yml | etc.

Each topology subdir has a `docker-compose.yml` for the recommended
starter — bare `cd <topology> && docker compose up` works because
docker compose finds that filename automatically. Variants drop the
`docker-compose.` prefix since they're invoked via `-f` flag.

27 compose file moves total:
- 18 Qwen vLLM composes redistributed across single/dual/multi4
- 2 Qwen llama-cpp composes into single/
- 6 Gemma vLLM composes redistributed across single/dual
- 1 untracked qwopus-bf16mtp moved to dual/

Inside each compose: relative paths to `../patches/` and `../cache/`
bumped to `../../patches/` / `../../cache/`, and `../../../../models-cache`
to `../../../../../models-cache` (one extra `..` for the new depth).

Reference updates across 148 files (BENCHMARKS, all docs, CHANGELOGs,
sibling-table cross-references in compose headers, scripts, patch
READMEs, .github issue templates, tools/residency-instrument).

scripts/switch.sh VARIANTS map updated; tags themselves unchanged
(`vllm/dual` → `dual/docker-compose.yml`, `vllm/dual4` → `multi4/docker-compose.yml`,
`vllm/gemma-mtp` → `gemma-4-31b/.../dual/docker-compose.yml`, etc.).

AGENTS.md "Compose layout" section rewritten to describe the new
hierarchy, with concrete examples and the fine-tune exception
(`dual/carnice-bf16mtp.yml` carries the fine-tune name as a filename
prefix until the fine-tune graduates to its own model directory).

All switch.sh paths verified to resolve to actual files post-move.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-09 12:11:17 +00:00

14 KiB
Raw Blame History

Multi-card (3+ GPUs) — derivation, constraints, scaling recipe

You have 3 or more GPUs and want to know if club-3090 applies. Short answer: yes. We ship one community-validated 4×3090 baseline and keep other 3+ GPU configs as derivation recipes until someone measures them. This page explains what scales (and what doesn't) when going beyond TP=2, the constraints to know, and how to derive your own compose when multi4.yml isn't your topology.

Validation note: the maintainer rig is 2× RTX 3090 PCIe, but Whamp's 4× RTX 3090 PCIe rig validated the TP=4 fp8/MTP baseline in discussion #26 on 2026-05-03. The TP=8+ sections remain derived expectations. If you have 4× / 8× hardware and run additional configs, please share results via the Numbers from your rig issue template — bash scripts/report.sh --full > my-rig.md captures everything we'd want (verify + stress + soak-continuous + bench, ~35 min).


TL;DR — what scales, what doesn't

Aspect TP=1 TP=2 (measured) TP=4 (measured) TP=8 (derived)
Per-card weight share 100% (~14 GB) 50% (~7 GB) 25% (~3.5 GB) 12.5% (~1.75 GB)
KV pool capacity smallest 2× ~4× ~8×
Per-card peak VRAM (262K target) 23.5+ GB tight 23.6 GB tight 23.5 GB fp8 / 22.0 GB DFlash ~10-12 GB
Cliff 2 single-prompt fires at ~60K doesn't fire (verified at 237K) passes 91K needle shouldn't fire
Per-stream TPS (PCIe-only) baseline ~same as TP=1 63/76 fp8, 64/104 DFlash lower still
Concurrent throughput (multi-stream) 1× ~1.7-3.6× KV pre-check 6.77× fp8, 2.27× DFlash @ 262K derived ~3-12×
Marlin pad-sub-tile-n patch not needed required required required

Two key takeaways:

  1. More cards = much more headroom, especially for long-context single-prompt workloads. On TP=4 the 24 GB-per-card pressure that drives Cliff 2 disappears entirely — weights and KV pool both split.
  2. Per-stream TPS doesn't scale without NVLink. PCIe NCCL all-reduce overhead grows with TP count; per-stream decode at TP=4 may be lower than TP=2. Aggregate concurrent throughput still scales, but you don't get faster single-stream answers from more PCIe cards.

Valid TP values for Qwen3.6-27B

vLLM's tensor parallelism splits attention heads across cards. The TP value must divide both the attention head count AND the KV head count cleanly. Qwen3.6-27B has:

  • 80 attention heads (factors: 1, 2, 4, 5, 8, 10, 16, 20, 40, 80)
  • 5 KV heads (factors: 1, 5)

The intersection — TP values that work — is 1, 2, 4, 5, 8, 10. So:

GPUs Valid TP Notes
1 TP=1 Standard single-card. See SINGLE_CARD.md.
2 TP=2 Standard dual. See DUAL_CARD.md.
3 TP=2 only TP=3 would split 5 KV heads as 5/3 = 1.67 per card — vLLM errors at boot. Use TP=2 with 1 idle card (set CUDA_VISIBLE_DEVICES=0,1), or run 2 single-card stacks on different ports.
4 TP=4 Each card gets 20 attention heads + 1.25 KV heads — vLLM splits with replication for fractional KV (handled internally). Production-viable if your rig has the slots + power + cooling.
5 TP=5 Theoretically valid (1 KV head per card, 16 attention heads per card). Unusual rig count; not common.
6 or 7 TP=4 or TP=5 + spare cards TP=6/7 don't divide head count. Use TP=4 (idle 2-3 cards) or TP=5 (idle 1-2 cards).
8 TP=8 Datacenter-class. Each card gets 10 attention heads, splits KV heads via vLLM's internal handling.
10 TP=10 Server-class. Production-viable on data-center boards.

Critical: TP=3, TP=6, TP=7, TP=9 do NOT work. vLLM errors at boot ("number of attention heads must be divisible by tensor parallel size"). If you have an awkward GPU count, use the next-lower valid TP and leave the extras idle, or run separate stacks on different ports.

Picking which cards to use on awkward counts

On a rig with 3 cards (or more, where you only want to use 2 for vLLM), which two you select matters for both throughput and reliability.

# Inspect topology first — connectivity classes affect TP allreduce
nvidia-smi topo -m

The matrix shows pairwise links between GPUs. Best to worst for TP allreduce:

Class Meaning Implication
NV# NVLink-bonded Fastest. We don't have it on consumer 3090s by default.
PIX Same PCIe switch (one bridge hop) Optimal on PCIe-only stacks.
PXB Multiple PCIe bridges, no host bridge Acceptable; slightly higher latency.
PHB Crosses PCIe Host Bridge (the CPU) Common on consumer boards; works but ~10-15% allreduce overhead vs PIX.
SYS Crosses NUMA / SMP interconnect Avoid for TP if there's a same-NUMA pair available.

Pick a same-switch pair (PIX) if your rig has one — typically the two slots wired into the same PCIe expander on workstation boards. On consumer ATX, all GPUs usually traverse the host bridge (PHB) so it doesn't matter much; on Threadripper / EPYC / dual-CPU server boards, NUMA topology can make a measurable difference.

To run TP=2 on cards 1+2 (e.g., card 0 is reserved for ComfyUI / display):

# In your override compose file
services:
  vllm-qwen36-27b-dual:
    environment:
      - CUDA_VISIBLE_DEVICES=1,2

Cross-rig data: @lexhoefsloot runs TP=2 on host GPUs 1+2 of a 3× 3090 rig (same PCIe switch, PHB to GPU 0) — that's the right pattern for a rig where GPU 0 is doing other work or differs in topology.


Shipped TP=4 baselines — vllm/dual4 and vllm/dual4-dflash

For 4× RTX 3090 PCIe, start with the measured fp8/MTP compose:

bash scripts/switch.sh vllm/dual4

multi4/docker-compose.yml keeps the dual.yml fp8/MTP feature set and changes TP/streams from 2 → 4. Validation on Whamp's 4× 3090 PCIe rig:

  • boots at max_model_len=262144, max_num_seqs=4
  • vLLM reports GPU KV cache size 483,200 tokens and 6.77× maximum concurrency for 262K-token requests
  • verify-full.sh passes
  • verify-stress.sh passes 7/7; probe 7 recalls 58,569-token and 91,070-token needles
  • bench.sh: 63.01 narr / 76.25 code wall TPS, peak 23,494 MiB/card

Use the DFlash variant when code throughput matters more than stream count and you can download the gated z-lab/Qwen3.6-27B-DFlash draft:

WITH_DFLASH_DRAFT=1 bash scripts/setup.sh qwen3.6-27b
bash scripts/switch.sh vllm/dual4-dflash

multi4/dflash.yml keeps full 262K context but uses FP16 KV and admits two full-context streams:

  • boots at max_model_len=262144, max_num_seqs=2
  • vLLM reports GPU KV cache size 207,264 tokens and 2.27× maximum concurrency for 262K-token requests
  • verify-full.sh passes
  • verify-stress.sh passes 7/7; probe 7 recalls 58,570-token and 91,070-token needles
  • bench.sh: 64.00 narr / 104.40 code wall TPS, peak 21,960 MiB/card
  • DFlash AL during code bench: 4.43 / 4.37 / 4.35 last observed samples

Single-stream TPS is lower than the 2-card DFlash variants on PCIe-only allreduce, so use TP=4 DFlash for full-262K code-heavy work and two admitted streams — not as a replacement for the fastest 2-card short-prompt DFlash path.

Recipe — derive your own config from dual.yml

dual.yml is the tested 2-card baseline and multi4.yml is the measured 4-card baseline. To scale to another TP=N, copy one of those and change three lines:

  command:
    - --tensor-parallel-size
-   - "2"
+   - "4"      # or 8, etc. — must be a valid TP value from the table above
    - --max-num-seqs
-   - "2"
+   - "4"      # bump proportional to TP — more cards = more concurrent streams
    - --max-num-batched-tokens
-   - "8192"
+   - "16384"  # optionally bump proportional to TP for longer prefill chunks

Everything else stays the same:

  • --gpu-memory-utilization 0.92 — same per-card budget
  • --kv-cache-dtype fp8_e5m2 — same KV class
  • --max-model-len 262144 — same target context (more cards = more total KV pool, but per-request max stays at 262K unless you raise it)
  • MTP n=3 spec-decode — same
  • The Marlin pad-sub-tile-n patch mount stays — at higher TP, more out-features get sub-tile-split, so the patch is more likely to be needed, not less

Container name + port: pick something distinct so it doesn't collide with your other variants. multi4.yml uses vllm-qwen36-27b-multi4 and port 8015; reserve a different name/port for further experiments:

container_name: vllm-qwen36-27b-octa
ports:
  - "${PORT:-8016}:8000"

What we measured on TP=4 (4× 3090 PCIe)

Measured 2026-05-03 on Whamp's 4× RTX 3090 PCIe rig:

  • fp8/MTP boot time: 355s cold after model/image cache populated.
  • fp8/MTP pre-check: max_model_len=262144, max_num_seqs=4, GPU KV cache size 483,200 tokens, max concurrency 6.77× at 262K.
  • fp8/MTP VRAM: 21,714 MiB idle after boot; 23,494 MiB/card peak during canonical bench.
  • fp8/MTP TPS: 63.01 narrative / 76.25 code wall TPS.
  • MTP AL: last three code-bench metrics showed mean acceptance length 3.42 / 3.53 / 3.62.
  • DFlash boot time: 375s cold after model/image cache populated.
  • DFlash pre-check: max_model_len=262144, max_num_seqs=2, GPU KV cache size 207,264 tokens, max concurrency 2.27× at 262K.
  • DFlash VRAM: 21,940 MiB idle after boot; 21,960 MiB/card peak during canonical bench.
  • DFlash TPS: 64.00 narrative / 104.40 code wall TPS.
  • DFlash AL: last three code-bench metrics showed mean acceptance length 4.43 / 4.37 / 4.35.
  • Cliff 2: canonical verify-stress.sh probe 7 passes at both large rungs on both TP=4 variants: ~58.6K tokens and 91K tokens recalled correctly.
  • Trade-off: PCIe allreduce makes single-stream decode slower than TP=2, but TP=4 provides more full-context concurrency and the first published 4×3090 Cliff 2 boundary data.

What to expect on TP=8 (8× 3090 / A6000)

Server-class setup. Most users at this scale are on rack hardware (DGX, 4U server chassis, dedicated cooling). The per-card pressure essentially disappears:

  • Per-card peak VRAM: ~10-12 GB. You have headroom to do almost anything — bump max-num-seqs to 8+, push max-model-len higher, experiment with TQ3 + Genesis stack from dual-turbo.yml.
  • Per-stream decode TPS: without NVLink fabric, likely lower than TP=4. Server-class cards (A6000, A100) often have NVLink — that changes the per-stream calculus dramatically.
  • Aggregate throughput: scales near-linearly with N if you have multi-stream load.

If you're on a server-class rig with NVLink: per-stream TPS could approach 1.6-1.8× single-card vs the ~1.0× we see on PCIe TP=2. That makes TP=8 with NVLink a meaningfully different regime than what we measure.


Cross-rig data we'd love

If you have 4× / 8× hardware and run any config, please share via Numbers from your rig. The single command that captures everything:

bash scripts/report.sh --full > my-rig.md

That's verify-full + verify-stress 7/7 + SOAK_MODE=continuous + canonical bench in one ~35-min pass. Use --bench instead if you want bench numbers without the soak (faster but doesn't probe Cliff 2b).

Specifically interested in:

  • More TP=4 on 4× 3090 PCIe — does your motherboard / power cap / PCIe topology match or beat Whamp's 63 / 76 TPS baseline? What's concurrent throughput?
  • TP=4 with NVLink topology (e.g. NVLink across pairs) — how does per-stream TPS compare to PCIe-only TP=2?
  • TP=8 on 8× A6000 / A100 — first server-class data point we'd collect.
  • TP=4 on mixed cards (e.g. 2× 3090 + 2× 4090, or 4× modded 3080 20GB) — does vLLM's per-card weight balance handle asymmetric VRAM ceilings cleanly? Asymmetric setups need --gpu-memory-utilization tuned to the smallest card's free VRAM.

Why we ship only one pre-baked 4-card config

We now ship multi4.yml because a community rig validated that exact 4× RTX 3090 PCIe topology with verify-full.sh, verify-stress.sh, and bench.sh. We still avoid a broad matrix of untested 4+ GPU composes:

  1. Hardware combinations explode. 4× 3090 vs 4× A5000 vs 4× A6000 vs 2×3090 + 2×4090 vs 4× modded 3080 — each has different VRAM, topology, power profile, and allreduce characteristics.
  2. Variant count needs discipline. A single measured fp8/MTP TP=4 baseline is useful; a directory full of derived-but-unvalidated variants would create false confidence.
  3. Users at this scale are typically experienced. If you have a workstation chassis or rack with 4-8 GPUs, you've already done the hardware homework. What you need from us is the methodology, the constraints, and one validated starting point.

If a community member contributes another tested compose for a specific topology or workload (with verify-stress.sh passing + bench.sh numbers), we'll ship it with credit and a header noting which rig validated it.


See also