Files
club-3090/results/lucebox-pflash-niah-20260504-150321/gpu_monitor.csv
noonghunna ebca0c8921 docs(benchmarks): PFlash long-context bench — 131K source ceiling on 1× 3090 (#230)
Closes task #230. Measured PFlash NIAH compression at 16K-260K source
contexts on 1× 24 GB / 3090 single-card.

Result: PFlash works flawlessly up to 131K source. Compresses 131,068
tokens to 6,524 (5%) in 10.8s with NIAH key + answer both retained.
Vanilla llama.cpp pp131072 takes ~257s per Luce's published numbers,
so PFlash alone is ~24× faster at this context. End-to-end TTFT
(PFlash + target prefill on 6.5K) would be ~12-13s vs ~257s = ~20×.

Above 131K, drafter ephemeral forward-pass tensors (K_curr/V_curr/Q_last
at full sequence length) exceed 24 GB. K-cache quantization
(--pflash-k-type q8_0) doesn't help — the failing allocs are
forward-pass not cache, confirmed by separate bench at 200K/260K with
identical OOM at the same layer numbers.

@weicj's PR #78 claim of 24K → 262K dual-GPU phase split is neither
refuted nor reproduced. Their setup was 2× 22 GB Ti with target also
loaded co-resident on one card; the "24K" was target+drafter
combined. Our 131K is drafter-alone on 24 GB. Reproducing 262K
specifically would require investigation of their drafter config
(chunk_size, lookahead, BSA window) — drafter activation footprint
at 200K+ is the binding constraint regardless of GPU count.

Practical recommendation for 24 GB / 3090 single-card users: PFlash
is shippable for source contexts ≤ 131K. The ~24× TTFT speedup is
genuine and quality holds. Above 131K, fall back to vanilla llama.cpp
prefill or wait for upstream drafter optimizations.

Adds:
- BENCHMARKS.md "PFlash long-context compression on 1× 3090" subsection
  with full per-context table + drafter ceiling explanation
- results/lucebox-pflash-niah-20260504-150321/ (BF16 K cache run)
- results/lucebox-pflash-niah-q8k-20260504-150600/ (q8_0 K cache run)

This closes our active investigation of the Luce surface — three
benches done (DFlash same-card 73.97 mean, K8V4 same-card 74.68 mean,
PFlash compression ceiling 131K). Recommendation surface narrows to:
PFlash at ≤131K is the one piece of Luce that beats vLLM dual.yml on
TTFT for that workload class.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 15:08:12 +00:00

2.5 KiB

1tsphaseindextemp_cfan_pctpower_wpower_limit_wmem_used_mibmem_total_mibutil_pct
21777907036.255pflash_load051028.92230.001245760
31777907036.255pflash_load152023.08230.001245760
41777907037.309niah_ctx16372054061.36230.004348245761
51777907037.309niah_ctx16372152025.83230.001245760
61777907038.355niah_ctx3276405857210.91230.00697224576100
71777907038.355niah_ctx32764152027.22230.001245760
81777907039.401niah_ctx3276405954223.68230.0069842457699
91777907039.401niah_ctx32764152023.30230.001245760
101777907040.447niah_ctx6552405856211.54230.001100424576100
111777907040.447niah_ctx65524152023.00230.001245760
121777907041.492niah_ctx6552405959224.04230.00109542457699
131777907041.492niah_ctx65524152022.92230.001245760
141777907042.536niah_ctx6552406062227.43230.00109542457699
151777907042.536niah_ctx65524152022.91230.001245760
161777907043.583niah_ctx6552406063226.14230.001148224576100
171777907043.583niah_ctx65524152022.81230.001245760
181777907044.715niah_ctx13106805964190.93230.001870224576100
191777907044.715niah_ctx131068152022.77230.001245760
201777907045.761niah_ctx13106806065221.91230.001894424576100
211777907045.761niah_ctx131068152022.94230.001245760
221777907046.805niah_ctx13106806065227.38230.00191382457699
231777907046.805niah_ctx131068152022.93230.001245760
241777907047.850niah_ctx13106806165227.67230.001894424576100
251777907047.850niah_ctx131068152022.81230.001245760
261777907048.894niah_ctx13106806166227.86230.00191382457699
271777907048.894niah_ctx131068152022.87230.001245760
281777907049.945niah_ctx13106806166227.93230.001913824576100
291777907049.945niah_ctx131068152022.92230.001245760
301777907051.001niah_ctx13106806166227.80230.001894424576100
311777907051.001niah_ctx131068152023.00230.001245760
321777907052.057niah_ctx13106806166228.89230.00191382457699
331777907052.057niah_ctx131068151022.97230.001245760
341777907053.102niah_ctx13106806067219.70230.00210242457694
351777907053.102niah_ctx131068151022.80230.001245760
361777907054.223niah_ctx13106806067191.47230.00210242457693
371777907054.223niah_ctx131068151022.93230.001245760
381777907055.359niah_ctx13106805967190.35230.0020462457692
391777907055.359niah_ctx131068151022.99230.001245760
401777907056.402cleanup05967178.23230.00408024576100
411777907056.402cleanup151022.87230.001245760