Files
club-3090/results/lucebox-pflash-niah-q8k-20260504-150600/gpu_monitor.csv
noonghunna ebca0c8921 docs(benchmarks): PFlash long-context bench — 131K source ceiling on 1× 3090 (#230)
Closes task #230. Measured PFlash NIAH compression at 16K-260K source
contexts on 1× 24 GB / 3090 single-card.

Result: PFlash works flawlessly up to 131K source. Compresses 131,068
tokens to 6,524 (5%) in 10.8s with NIAH key + answer both retained.
Vanilla llama.cpp pp131072 takes ~257s per Luce's published numbers,
so PFlash alone is ~24× faster at this context. End-to-end TTFT
(PFlash + target prefill on 6.5K) would be ~12-13s vs ~257s = ~20×.

Above 131K, drafter ephemeral forward-pass tensors (K_curr/V_curr/Q_last
at full sequence length) exceed 24 GB. K-cache quantization
(--pflash-k-type q8_0) doesn't help — the failing allocs are
forward-pass not cache, confirmed by separate bench at 200K/260K with
identical OOM at the same layer numbers.

@weicj's PR #78 claim of 24K → 262K dual-GPU phase split is neither
refuted nor reproduced. Their setup was 2× 22 GB Ti with target also
loaded co-resident on one card; the "24K" was target+drafter
combined. Our 131K is drafter-alone on 24 GB. Reproducing 262K
specifically would require investigation of their drafter config
(chunk_size, lookahead, BSA window) — drafter activation footprint
at 200K+ is the binding constraint regardless of GPU count.

Practical recommendation for 24 GB / 3090 single-card users: PFlash
is shippable for source contexts ≤ 131K. The ~24× TTFT speedup is
genuine and quality holds. Above 131K, fall back to vanilla llama.cpp
prefill or wait for upstream drafter optimizations.

Adds:
- BENCHMARKS.md "PFlash long-context compression on 1× 3090" subsection
  with full per-context table + drafter ceiling explanation
- results/lucebox-pflash-niah-20260504-150321/ (BF16 K cache run)
- results/lucebox-pflash-niah-q8k-20260504-150600/ (q8_0 K cache run)

This closes our active investigation of the Luce surface — three
benches done (DFlash same-card 73.97 mean, K8V4 same-card 74.68 mean,
PFlash compression ceiling 131K). Recommendation surface narrows to:
PFlash at ≤131K is the one piece of Luce that beats vLLM dual.yml on
TTFT for that workload class.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 15:08:12 +00:00

26 lines
1.6 KiB
CSV

ts,phase,index,temp_c,fan_pct,power_w,power_limit_w,mem_used_mib,mem_total_mib,util_pct
1777907191.039,pflash_load,0,43,0,27.92,230.00,1,24576,0
1777907191.039,pflash_load,1,49,0,22.60,230.00,1,24576,0
1777907192.089,niah_ctx131068,0,47,0,83.13,230.00,18690,24576,17
1777907192.089,niah_ctx131068,1,50,0,25.62,230.00,1,24576,0
1777907193.135,niah_ctx131068,0,51,0,214.50,230.00,18944,24576,99
1777907193.135,niah_ctx131068,1,50,0,26.32,230.00,1,24576,0
1777907194.195,niah_ctx131068,0,53,0,227.52,230.00,19138,24576,100
1777907194.195,niah_ctx131068,1,50,0,23.25,230.00,1,24576,0
1777907195.247,niah_ctx131068,0,53,0,229.01,230.00,19138,24576,99
1777907195.247,niah_ctx131068,1,49,0,22.55,230.00,1,24576,0
1777907196.291,niah_ctx131068,0,54,0,227.81,230.00,18944,24576,100
1777907196.291,niah_ctx131068,1,49,0,22.45,230.00,1,24576,0
1777907197.337,niah_ctx131068,0,53,56,229.25,230.00,19138,24576,99
1777907197.337,niah_ctx131068,1,49,0,22.45,230.00,1,24576,0
1777907198.384,niah_ctx131068,0,54,52,227.66,230.00,19138,24576,100
1777907198.384,niah_ctx131068,1,49,0,22.55,230.00,1,24576,0
1777907199.430,niah_ctx131068,0,54,54,229.18,230.00,18944,24576,99
1777907199.430,niah_ctx131068,1,49,0,22.46,230.00,1,24576,0
1777907200.474,niah_ctx131068,0,53,56,211.57,230.00,21024,24576,94
1777907200.474,niah_ctx131068,1,49,0,22.56,230.00,1,24576,0
1777907201.614,niah_ctx131068,0,53,58,186.76,230.00,21024,24576,93
1777907201.614,niah_ctx131068,1,49,0,22.60,230.00,1,24576,0
1777907202.805,niah_ctx131068,0,53,58,186.52,230.00,2048,24576,84
1777907202.805,niah_ctx131068,1,49,0,22.61,230.00,1,24576,0