高带宽闪存提升大模型服务性能
Characterizing High Bandwidth Flash for LLM Serving
这篇论文展示了如何用高带宽闪存解决大模型内存瓶颈,能显著提升性能并延长硬件寿命。
研究评估高带宽闪存(HBF)在大语言模型服务中的应用效果。通过HBM-HBF分层存储系统和缓存感知调度,HBF增强系统将完成时间减少36.1-87.0%。模型显示可节省55.8%能源,缓存感知调度将HBF写入寿命从4.77年延长至14.82年。
Characterizing High Bandwidth Flash for LLM Serving
Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.0% relative to HBM-only systems. Modeled energy savings reach 55.8%, although HBF increases energy consumption on some light workloads. Buffered cache-aware scheduling extends estimated HBF write lifetime from 4.77 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.