Files
club-3090/models
noonghunna 755e5199ff bench(head-to-head): matched-config rebench + Qwen INT8 PTH KV compose
Adds qwen3.6-27b/vllm/compose/dual/int8.yml (new) — same vLLM nightly
1acd67a7 as gemma-int8, same int8_per_token_head KV class, MTP n=4
parameterized via SPEC_N_MAX env var. First validation of INT8 PTH KV
on Qwen3-Next DeltaNet hybrid attention on our stack.

Updates gemma-int8.yml to parameterize num_speculative_tokens via SPEC_N_MAX.

BENCHMARKS.md head-to-head section: matched-config addendum showing
Gemma's per-stream advantage shrinks to ~10% decode / ~45% TTFT (vs
+60% in the morning's mismatched run), and Qwen has +30% more KV pool /
concurrency. Original mismatched table preserved.

Side-effect: Qwen int8.yml is a candidate shipping path — +27/+33% TPS
over canonical dual.yml at -2.3 GiB/card VRAM.
2026-05-10 23:09:37 +00:00
..