Every gemma-4-31b vLLM compose pinned a purged Docker Hub nightly (e47c98ef / bf610c2f), un-bootable for fresh users (#250, #167 class). Prune the 9-variant set to 3 and repoint survivors onto the immutable stable release tag vllm/vllm-openai:v0.21.0 (never :latest). Survivors: vllm/gemma-mtp (bf16 dual, default), vllm/gemma-int8 (int8-PTH dual), vllm/gemma-mtp-tp1 (fp8 single). Dropped: gemma-dflash, gemma-dflash-int8, gemma-int8-tq3, gemma-bf16, gemma-awq; the 262K path folds into gemma-int8 via a CTX env override. Gemma is decoupled from the shared Qwen nightly profiles via a new vllm-gemma-stable engine profile; Qwen nightly profiles are untouched. Diagnose decoupling: the pruned gemma dflash patch dirs doubled as the diagnose tool's cross-model disk-source proxies for two Qwen overlays (vllm-pr41703-dflash, vllm-pr42102-dflash-kv-quant). Those capabilities are image-baked in the pinned nightly, not mountable files, so mark them image_baked and have diagnose skip the disk-source check for image-baked overlays. Runtime-neutral (no Qwen compose mounts them). switch.sh launcher: handle the new VLLM_IMAGE engine-pin export (the twin loop in launch.sh had it; switch.sh did not -> "unexpected engine pin export"), and guard the VLLM_NIGHTLY_SHA echo for the image-only stable profile (was unbound under set -u). Validated on 2x RTX 3090 (Ampere, v0.21.0): full scripts/tests/*.sh suite 35/0; live boots of gemma-mtp (bf16) and gemma-int8 serve coherent output, qwen vllm/dual regression-clean. gemma-mtp-tp1 (fp8_e4m3) is correctly SM-gated (required_sm=9.0) -- confirmed Ampere-incompatible (Triton "fp8e4nv not supported on sm_86"), a documented 32 GB+/Hopper variant; beellama remains the single-card gemma path on Ampere. Refs #451, #250, #167. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
23 lines
1.0 KiB
Bash
Executable File
23 lines
1.0 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
#
|
|
# generate-compose.sh — v0.8.0 #141 compose generator (PR #147).
|
|
#
|
|
# Emits a minimal-reproduction docker-compose for an in-scope (non-Genesis,
|
|
# vLLM-only) profile. Mission (locked decision #2): reproduce + flag, NEVER
|
|
# repair. The engine image is NEVER rewritten; a failed-drift-guard patch is
|
|
# NEVER wired; --trust-remote-code (a governed slot, locked §88) is NEVER
|
|
# blind-passed for an in-scope profile.
|
|
#
|
|
# Usage:
|
|
# scripts/generate-compose.sh --profile vllm/minimal [--out FILE]
|
|
# scripts/generate-compose.sh --profile vllm/gemma-int8 --accept-degraded
|
|
# scripts/generate-compose.sh --model gemma-4-31b --engine vllm-nightly-full
|
|
# # convenience tuple: prints candidate --profile values, exits non-zero
|
|
#
|
|
# All decision logic lives in scripts/lib/generate_compose.py (this is a
|
|
# thin argv pass-through, matching the diagnose-profile.sh pattern).
|
|
set -euo pipefail
|
|
|
|
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
|
exec python3 "${ROOT_DIR}/scripts/lib/generate_compose.py" "$@"
|