Files
club-3090/scripts/rebench-runtime.sh
T
f1bb069eb5 feat(models): add ik_llama Qwen3.6-27B IQ4_KS composes — text 262K + vision 160K (#180)
ik_llama.cpp serving path for Qwen3.6-27B (ubergarm MTP-IQ4_KS GGUF), validated
on a single RTX 3090 (task #402):

- compose/single/iq4ks-mtp.yml — text, 262K ctx (full native n_ctx_train),
  q4_0 KV, MTP n=2, froggeric v19. ~59 narr / 65 code TPS; verify-stress 7/7
  to 90K; soak-continuous PASS. MTP n=2 vs n=3 A/B: n=2 wins (narrative +3.7%,
  code a wash within noise).
- compose/single/iq4ks-mtp-vision.yml — vision, 160K ctx + full-res (4M-px /
  2048²) images at -b 1024; bench + verify-stress + soak all PASS. 160K chosen
  for headroom (172K served images but OOM-crashed under heavy prefill).
- patches/froggeric-chat-template/ — v19 template + PROVENANCE.
- scripts/rebench-runtime.sh — bench+verify-stress+soak "runtime tier" suite
  (rebench-full minus the quality legs; for config-only changes).

Engine quirk documented in-header: this ik_llama build forces n_ubatch =
n_batch, so only -b moves the compute buffer (-ub is a no-op here — distinct
from mainline llama.cpp). Eval track; not yet wired into COMPOSE_REGISTRY/switch.sh.

Co-authored-by: noonghunna <[email protected]>
Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
2026-05-22 04:45:44 +05:00

56 lines
2.4 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# rebench-runtime.sh — the "runtime" tier of rebench: bench + verify-stress
# + soak, skipping the two generation-quality legs (quality-full + aider).
#
# WHEN TO USE THIS instead of rebench-full.sh:
# When you changed a serving-config knob that affects throughput / VRAM /
# stability but NOT per-token generation quality, so re-scoring the 8-pack
# + aider would be ~1.5 hr of wasted cycles that can only reproduce the
# prior scores. Examples:
# - context-size ceiling bump (KV *type* unchanged)
# - --ubatch-size / --batch-size tuning
# - --gpu-memory-utilization / --mem-fraction-static
# - power-limit sweep
# - MTP n / draft-p-min retune (TPS + accept-rate, not answer text)
# If the change DOES touch generation (quant swap, KV-cache *type* change,
# chat-template change, model swap) → run full rebench-full.sh instead.
#
# Three legs (≈30-40 min on single-card, vs ~1.75-2 hr for the full matrix):
# 1. bench.sh — TPS narrative + code
# 2. verify-stress.sh — long-context needle ladder + prefill-OOM boundary
# 3. soak-test.sh — accumulating-context endurance (Cliff 2b)
#
# Usage (identical surface to rebench-full.sh — all flags pass through):
# bash scripts/rebench-runtime.sh # auto-detect running compose
# bash scripts/rebench-runtime.sh --tag ik-262k # explicit tag
# bash scripts/rebench-runtime.sh --skip soak # ALSO skip soak (merged)
#
# llama.cpp / ik_llama (non-vllm container names) — soak-test.sh's container
# auto-detect only matches vllm-*; pass the endpoint + container explicitly
# (see #403):
# CONTAINER=ik-llama-qwen36-27b URL=http://localhost:8020 MODEL=ik-iq4ks-mtp \
# bash scripts/rebench-runtime.sh --engine llama-cpp
#
# This is a thin preset over rebench-full.sh: it injects
# `--skip quality-full,aider-polyglot` and merges any --skip you pass.
set -euo pipefail
ROOT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
BASE_SKIP="quality-full,aider-polyglot"
USER_SKIP=""
PASS_ARGS=()
while [[ $# -gt 0 ]]; do
case "$1" in
--skip) USER_SKIP="${2:-}"; shift 2 ;;
*) PASS_ARGS+=("$1"); shift ;;
esac
done
SKIP="$BASE_SKIP${USER_SKIP:+,$USER_SKIP}"
echo "[rebench-runtime] runtime tier (bench + verify-stress + soak); --skip=$SKIP"
exec bash "$ROOT_DIR/scripts/rebench-full.sh" --skip "$SKIP" "${PASS_ARGS[@]}"