Adds qwen3.6-27b/vllm/compose/dual/int8.yml (new) — same vLLM nightly
1acd67a7 as gemma-int8, same int8_per_token_head KV class, MTP n=4
parameterized via SPEC_N_MAX env var. First validation of INT8 PTH KV
on Qwen3-Next DeltaNet hybrid attention on our stack.
Updates gemma-int8.yml to parameterize num_speculative_tokens via SPEC_N_MAX.
BENCHMARKS.md head-to-head section: matched-config addendum showing
Gemma's per-stream advantage shrinks to ~10% decode / ~45% TTFT (vs
+60% in the morning's mismatched run), and Qwen has +30% more KV pool /
concurrency. Original mismatched table preserved.
Side-effect: Qwen int8.yml is a candidate shipping path — +27/+33% TPS
over canonical dual.yml at -2.3 GiB/card VRAM.