Files
club-3090/scripts/generate-compose.sh
T
noonghunnaandClaude Opus 4.8 45686808cb Prune Gemma vLLM variants (9->3) and pin survivors to v0.21.0
Every gemma-4-31b vLLM compose pinned a purged Docker Hub nightly
(e47c98ef / bf610c2f), un-bootable for fresh users (#250, #167 class).
Prune the 9-variant set to 3 and repoint survivors onto the immutable
stable release tag vllm/vllm-openai:v0.21.0 (never :latest).

Survivors: vllm/gemma-mtp (bf16 dual, default), vllm/gemma-int8 (int8-PTH
dual), vllm/gemma-mtp-tp1 (fp8 single). Dropped: gemma-dflash,
gemma-dflash-int8, gemma-int8-tq3, gemma-bf16, gemma-awq; the 262K path
folds into gemma-int8 via a CTX env override. Gemma is decoupled from the
shared Qwen nightly profiles via a new vllm-gemma-stable engine profile;
Qwen nightly profiles are untouched.

Diagnose decoupling: the pruned gemma dflash patch dirs doubled as the
diagnose tool's cross-model disk-source proxies for two Qwen overlays
(vllm-pr41703-dflash, vllm-pr42102-dflash-kv-quant). Those capabilities
are image-baked in the pinned nightly, not mountable files, so mark them
image_baked and have diagnose skip the disk-source check for image-baked
overlays. Runtime-neutral (no Qwen compose mounts them).

switch.sh launcher: handle the new VLLM_IMAGE engine-pin export (the twin
loop in launch.sh had it; switch.sh did not -> "unexpected engine pin
export"), and guard the VLLM_NIGHTLY_SHA echo for the image-only stable
profile (was unbound under set -u).

Validated on 2x RTX 3090 (Ampere, v0.21.0): full scripts/tests/*.sh suite
35/0; live boots of gemma-mtp (bf16) and gemma-int8 serve coherent output,
qwen vllm/dual regression-clean. gemma-mtp-tp1 (fp8_e4m3) is correctly
SM-gated (required_sm=9.0) -- confirmed Ampere-incompatible (Triton
"fp8e4nv not supported on sm_86"), a documented 32 GB+/Hopper variant;
beellama remains the single-card gemma path on Ampere.

Refs #451, #250, #167.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-05-29 07:59:42 +05:00

23 lines
1.0 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# generate-compose.sh — v0.8.0 #141 compose generator (PR #147).
#
# Emits a minimal-reproduction docker-compose for an in-scope (non-Genesis,
# vLLM-only) profile. Mission (locked decision #2): reproduce + flag, NEVER
# repair. The engine image is NEVER rewritten; a failed-drift-guard patch is
# NEVER wired; --trust-remote-code (a governed slot, locked §88) is NEVER
# blind-passed for an in-scope profile.
#
# Usage:
# scripts/generate-compose.sh --profile vllm/minimal [--out FILE]
# scripts/generate-compose.sh --profile vllm/gemma-int8 --accept-degraded
# scripts/generate-compose.sh --model gemma-4-31b --engine vllm-nightly-full
# # convenience tuple: prints candidate --profile values, exits non-zero
#
# All decision logic lives in scripts/lib/generate_compose.py (this is a
# thin argv pass-through, matching the diagnose-profile.sh pattern).
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
exec python3 "${ROOT_DIR}/scripts/lib/generate_compose.py" "$@"