ik_llama.cpp serving path for Qwen3.6-27B (ubergarm MTP-IQ4_KS GGUF), validated on a single RTX 3090 (task #402): - compose/single/iq4ks-mtp.yml — text, 262K ctx (full native n_ctx_train), q4_0 KV, MTP n=2, froggeric v19. ~59 narr / 65 code TPS; verify-stress 7/7 to 90K; soak-continuous PASS. MTP n=2 vs n=3 A/B: n=2 wins (narrative +3.7%, code a wash within noise). - compose/single/iq4ks-mtp-vision.yml — vision, 160K ctx + full-res (4M-px / 2048²) images at -b 1024; bench + verify-stress + soak all PASS. 160K chosen for headroom (172K served images but OOM-crashed under heavy prefill). - patches/froggeric-chat-template/ — v19 template + PROVENANCE. - scripts/rebench-runtime.sh — bench+verify-stress+soak "runtime tier" suite (rebench-full minus the quality legs; for config-only changes). Engine quirk documented in-header: this ik_llama build forces n_ubatch = n_batch, so only -b moves the compute buffer (-ub is a no-op here — distinct from mainline llama.cpp). Eval track; not yet wired into COMPOSE_REGISTRY/switch.sh. Co-authored-by: noonghunna <[email protected]> Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
56 lines
2.4 KiB
Bash
Executable File
56 lines
2.4 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
#
|
|
# rebench-runtime.sh — the "runtime" tier of rebench: bench + verify-stress
|
|
# + soak, skipping the two generation-quality legs (quality-full + aider).
|
|
#
|
|
# WHEN TO USE THIS instead of rebench-full.sh:
|
|
# When you changed a serving-config knob that affects throughput / VRAM /
|
|
# stability but NOT per-token generation quality, so re-scoring the 8-pack
|
|
# + aider would be ~1.5 hr of wasted cycles that can only reproduce the
|
|
# prior scores. Examples:
|
|
# - context-size ceiling bump (KV *type* unchanged)
|
|
# - --ubatch-size / --batch-size tuning
|
|
# - --gpu-memory-utilization / --mem-fraction-static
|
|
# - power-limit sweep
|
|
# - MTP n / draft-p-min retune (TPS + accept-rate, not answer text)
|
|
# If the change DOES touch generation (quant swap, KV-cache *type* change,
|
|
# chat-template change, model swap) → run full rebench-full.sh instead.
|
|
#
|
|
# Three legs (≈30-40 min on single-card, vs ~1.75-2 hr for the full matrix):
|
|
# 1. bench.sh — TPS narrative + code
|
|
# 2. verify-stress.sh — long-context needle ladder + prefill-OOM boundary
|
|
# 3. soak-test.sh — accumulating-context endurance (Cliff 2b)
|
|
#
|
|
# Usage (identical surface to rebench-full.sh — all flags pass through):
|
|
# bash scripts/rebench-runtime.sh # auto-detect running compose
|
|
# bash scripts/rebench-runtime.sh --tag ik-262k # explicit tag
|
|
# bash scripts/rebench-runtime.sh --skip soak # ALSO skip soak (merged)
|
|
#
|
|
# llama.cpp / ik_llama (non-vllm container names) — soak-test.sh's container
|
|
# auto-detect only matches vllm-*; pass the endpoint + container explicitly
|
|
# (see #403):
|
|
# CONTAINER=ik-llama-qwen36-27b URL=http://localhost:8020 MODEL=ik-iq4ks-mtp \
|
|
# bash scripts/rebench-runtime.sh --engine llama-cpp
|
|
#
|
|
# This is a thin preset over rebench-full.sh: it injects
|
|
# `--skip quality-full,aider-polyglot` and merges any --skip you pass.
|
|
|
|
set -euo pipefail
|
|
|
|
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
|
|
|
BASE_SKIP="quality-full,aider-polyglot"
|
|
USER_SKIP=""
|
|
PASS_ARGS=()
|
|
while [[ $# -gt 0 ]]; do
|
|
case "$1" in
|
|
--skip) USER_SKIP="${2:-}"; shift 2 ;;
|
|
*) PASS_ARGS+=("$1"); shift ;;
|
|
esac
|
|
done
|
|
|
|
SKIP="$BASE_SKIP${USER_SKIP:+,$USER_SKIP}"
|
|
|
|
echo "[rebench-runtime] runtime tier (bench + verify-stress + soak); --skip=$SKIP"
|
|
exec bash "$ROOT_DIR/scripts/rebench-full.sh" --skip "$SKIP" "${PASS_ARGS[@]}"
|