Characterizing High Bandwidth Flash for LLM Serving
Abstract
This study examines how high-bandwidth flash (HBF) can supplement accelerator memory for large language model serving, particularly agentic workloads that repeatedly reuse growing KV caches. It evaluates the trade-offs between added capacity, access overhead, energy use, and flash endurance. The authors combine a three-tier HBM–HBF–host memory hierarchy with buffered cache-aware scheduling and assess the design through trace-driven simulations. In the tested workloads, the fastest HBF configurations shorten completion times by 36.1–87.0% compared with HBM-only systems. Energy reductions reach 55.8%, though some lighter workloads consume more energy with HBF. Buffered scheduling increases the estimated flash lifetime from 4.77 to 14.82 years. The findings highlight that memory placement and scheduling must be designed together to capture HBF’s performance benefits while maintaining practical durability.