Counter-frames duart's #161 ("Proxmox needs HugePages/pinning, 3x"): our reference rig runs under Proxmox PCIe-only, no NVLink, untuned, at full dual baselines — out-of-box Proxmox is not a tax. The fragile element is NVLink across passed-through GPUs collapsing to a slow fallback on wrong IOMMU/ACS/NUMA; PCIe-only has no such path. Cross-ref #137 (NVLink-not-engaging under passthrough). Narrow, accurate framing — NOT a mandatory-tuning guide; report.sh --full is the diagnostic. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
11 KiB
Container runtimes — non-Docker / non-bare-metal notes
The shipped composes assume Docker (or Docker-compatible runtime) on bare-metal Linux. Most reporters run that exact stack and the verify-* scripts target it. This page captures what we know about non-default runtime/host combinations — not as recipes that are tested end-to-end, but as breadcrumbs for users hitting environmental issues that aren't club-3090 bugs.
Tested environment (the "known good" baseline): bare-metal Ubuntu 22.04 / 24.04, kernel 6.8.x, Docker Engine 27+ with NVIDIA Container Toolkit, no
default-runtime: nvidiaindaemon.json(opt-in via--gpus '"device=N,M"'per-command). Variants that materially diverge from this can surface separate bugs that aren't pin-fixable.
Soft-warn: the docker preflight in scripts/setup.sh is not load-bearing
scripts/setup.sh only fetches Genesis + models — no docker invocations until you actually run docker compose up later. As of 2f8ed19, the docker preflight in setup.sh is a soft warning rather than a hard fail, so users on non-Docker container runtimes can run setup without working around the gate.
scripts/launch.sh and scripts/switch.sh keep the hard check because they actually invoke docker commands directly.
Podman / Podman Compose
Already supported via env override — scripts/switch.sh reads COMPOSE_BIN and defaults to docker compose:
COMPOSE_BIN="podman compose" bash scripts/switch.sh vllm/dual
Caveats not validated end-to-end:
docker ps/docker inspectcalls insideswitch.sh(lines 119, 128-133) assume the binary is nameddockerand emit Docker-shaped JSON. Aliasingdocker → podmanmostly works but won't cover every case (e.g.,docker compose project labelsmay differ).- Per-command flag conventions differ:
--gpus '"device=0,1"'is Docker-specific syntax; podman uses--device nvidia.com/gpu=allor similar. Adjust your compose / override accordingly.
If you've shipped a Podman pipeline that works end-to-end through setup.sh → launch.sh → verify-full.sh, please file a docs PR with the diff so the next person inherits a working recipe.
microk8s
@apnar reported success running club-3090 under microk8s (disc #48, disc #51 for 5090 single-card data). Not pre-baked — the user translates the compose file to a k8s manifest manually. The setup.sh soft-warn (2f8ed19) is what unblocks setup; verify-full.sh works because it talks HTTP to the engine.
The path forward for microk8s:
- Models cache mounted into the k8s pod's volume (analogous to
~/.cache/huggingfacemount in compose) - GPUs assigned via the NVIDIA k8s device plugin, not
--gpus - vLLM arguments translated 1:1 from the compose
command:block into the pod spec'sargs: - Genesis / patch mounts (RO bind-mounts of the
models/qwen3.6-27b/vllm/patches/*files in the compose) → k8s ConfigMap or initContainer that places the same files
If you have a working microk8s pipeline and are willing to PR a manifest example, the docs should grow a docs/microk8s/ subfolder with one or two sample YAMLs. Open invitation.
Proxmox VE — known footgun on kernel 6.17.x (workaround: native venv)
Proxmox VE 8.x / kernel 6.17.x users: the Docker image
vllm/vllm-openai:nightly-7a1eb8ac2…(and likely current nightlies) crashes withRuntimeError: this event loop is already runningatvllm.entrypoints.cli.serve.cmd → uvloop.run(run_server(args))regardless of Genesis pin, TP, runtime config, or--init. Same image boots clean on bare-metal Ubuntu 6.8.x. Native venvpip install vllm==0.20.1on the same Proxmox host launches cleanly (@lexhoefsloot venv bisect — same kernel, samedefault-runtime: nvidiaruntime, native venv works end-to-end including 4-stream concurrent at 170-195 tok/s aggregate over 200K context). The bug is bounded to the Docker image × Proxmox container runtime interaction — not the kernel and not vLLM itself. Workaround: drop the Docker image, use a native Python venv.
The data trail (cite for upstream filing)
@lexhoefsloot's club-3090 #49 ran a thorough elimination sequence on the bug. What's been ruled out:
| Suspect | Status | Probe |
|---|---|---|
| Genesis patches | ❌ ruled out | Bare docker run (no Genesis) crashes |
torch.compile × uvloop |
❌ ruled out | --enforce-eager crashes |
| Multiproc spawn (TP > 1) | ❌ ruled out | TP=1 also crashes |
| Module import / lib-stack | ❌ ruled out | import vllm works |
| CLI dispatch / argparse | ❌ ruled out | vllm serve --help works |
| vLLM upstream regression | ❌ ruled out | Same image boots clean on bare-metal Ubuntu 2× 3090 PCIe (cross-rig, Probe B reproduction) |
default-runtime: nvidia in daemon.json |
❌ ruled out | Removed + per-command --runtime=nvidia still crashes |
| PID 1 / signal forwarding | ❌ ruled out | --init (tini PID 1) still crashes |
| Kernel 6.17.x | ❌ ruled out | Native venv on same kernel works end-to-end |
| vLLM as a Python package | ❌ ruled out | pip install vllm==0.20.1 venv works on same host |
| Proxmox VE at large | ❌ ruled out | Same Proxmox host runs the venv cleanly |
| Docker image × Proxmox container runtime interaction | ⚠️ remaining candidate | Sole surviving suspect after venv-vs-image cross-test |
What this means for you
If you're on bare-metal anything (Ubuntu, Debian, RHEL family) on a kernel 6.8-6.16-ish, none of this affects you — that's the tested baseline.
If you're on Proxmox VE LXC/VM and hit the same uvloop trace at boot:
- Confirm it's the same crash: container exits within ~5 seconds of
docker run(or compose up), the trace mentionsuvloop.loop.Loop.run_forever→RuntimeError: this event loop is already running. - If yes, the direct workaround is to drop the Docker image and use a native Python venv on the host:
This is what @lexhoefsloot's bisect proved works. Genesis can be applied via the same
python3 -m venv /opt/vllm-env source /opt/vllm-env/bin/activate pip install vllm==0.20.1 # Run vllm serve directly — same CLI args as the compose `command:` block, just without dockerapply_allscript against the venv install (Sander's setup supports it). Trade-off: you lose the consistency of a pinned Docker image, but you gain a working serving stack on Proxmox. - Cheapest in-Docker workaround attempt (still worth trying if you need Docker for orchestration reasons): boot with
--privileged. If that boots clean, the issue is namespace-policy-related and the workaround for your daily driver is--privileged(not ideal, but unblocks). - For deeper investigation / upstream filing: the bug is at the Docker image × Proxmox container runtime layer. Best filing target is NVIDIA Container Toolkit's issue tracker (since Proxmox uses it for GPU passthrough) or the Proxmox forum. Cite the venv-works datapoint to narrow scope.
If you find an in-Docker workaround that gets dual.yml booting on your Proxmox rig, please file a docs PR back here so the next Proxmox user inherits the answer faster.
Re-check triggers
This footgun page should be revisited when:
- Someone independently reproduces the asyncio crash on a non-Proxmox bare-metal Linux 6.17.x rig — would distinguish "kernel 6.17.x bug" from "Proxmox-specific" and re-route the diagnosis
- vLLM nightly bumps past
nightly-7a1eb8ac2…and a different SHA gets tested on the same Proxmox setup — would tell us whether it's nightly-7a1eb8ac2 specific or persistent across the nightly stream - Proxmox ships PVE kernel 6.18+ — kernel revisions historically resolve namespace × cgroup × user-space-runtime interactions
Proxmox passthrough performance — NVLink is the fragile path, not Proxmox itself
Out-of-box Proxmox VM GPU-passthrough is not an inherent performance tax. The reference rig runs under Proxmox (2× RTX 3090, PCIe-only, no NVLink, GPUs passed to the VM) with no HugePages / CPU-pinning / governor tuning and sustains the documented dual baselines (dual.yml ~69/89, dual-turbo ~81/108 tok/s). Pathologically low numbers (~20 tok/s on 2× 3090) indicate a specific misconfig, not "Proxmox".
The fragile element is NVLink across passed-through GPUs. Two NVLinked GPUs in a VM with wrong IOMMU/ACS/NUMA placement → NCCL silently drops the NVLink peer link and collapses to a slow cross-bridge/cross-NUMA fallback (a 2–3× throughput floor). PCIe-only rigs (the stack default — NCCL_P2P_DISABLE=1 / custom-all-reduce off) have no such path to misnegotiate, so they're robust out-of-box. Recurring class: #137 (NVLink-not-engaging under container/VM passthrough); #161 (Proxmox 2×3090+NVLink — NUMA-alignment + vCPU pinning restored a collapsed NVLink path, 20→60 tok/s, but still below the PCIe-only baseline → verify NVLink is actually engaged before chasing tuning folklore).
Diagnostic: bash scripts/report.sh --full includes a PCIe/NVLink topology + lspci + NVLink-detection section that shows whether NCCL is on the NVLink peer link or a fallback. On multi-socket hosts, check VM↔GPU NUMA alignment first; HugePages/governor are secondary.
What this page is NOT
- Not a recipe for Proxmox / microk8s / podman setups. Those are user-driven; we don't have CI runners to validate them.
- Not a list of all possible non-Docker setups. Just the ones that have surfaced via reporters with concrete data trails.
- Not a substitute for docs/HARDWARE.md which covers GPU/driver requirements regardless of runtime.
See also
- HARDWARE.md — GPU / driver / power / NVLink requirements
- CLIFFS.md — known failure modes (the engine-side bugs, not environmental ones)
- MULTI_CARD.md — TP topology and the
nvidia-smi topo -mranking - club-3090 #49 — full Proxmox VE asyncio investigation trail (parked)
- club-3090 disc #48 — original microk8s docker-soft-warn request