arXiv
2026

Characterizing High Bandwidth Flash for LLM Serving

Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami
Published
kv-cache
long-context

Abstract

This study examines how high-bandwidth flash (HBF) can supplement accelerator memory for large language model serving, particularly agentic workloads that repeatedly reuse growing KV caches. It evaluates the trade-offs between added capacity, access overhead, energy use, and flash endurance. The authors combine a three-tier HBM–HBF–host memory hierarchy with buffered cache-aware scheduling and assess the design through trace-driven simulations. In the tested workloads, the fastest HBF configurations shorten completion times by 36.1–87.0% compared with HBM-only systems. Energy reductions reach 55.8%, though some lighter workloads consume more energy with HBF. Buffered scheduling increases the estimated flash lifetime from 4.77 to 14.82 years. The findings highlight that memory placement and scheduling must be designed together to capture HBF’s performance benefits while maintaining practical durability.

‍

​
Read paper

Related Publications

Scaling Test-Time Compute for DLLMs via Parallel Search

COLM
2026
diffusion-llm
search
parallel-decoding
​
View Job

Characterizing High Bandwidth Flash for LLM Serving

arXiv
2026
kv-cache
long-context
​
View Job

AsyncOPD: How Stale Can On-Policy Distillation Be?

NeurIPS
2026
reasoning
distillation
asynchronous-execution
​
View Job