Why has memory become the biggest cost in AI chips?
Modern AI models, especially large language models, require enormous amounts of memory and incredibly fast access to it. During training and inference, these models process billions or even trillions of parameters, which all need to be stored and constantly accessed by the processing cores. Unlike traditional CPUs, which prioritize latency for a few complex operations, AI chips like GPUs need to move vast quantities of data in parallel. This demand for high-bandwidth memory (like HBM - High Bandwidth Memory) is specialized and expensive to produce, making it a dominant factor in the overall cost of AI hardware. The sheer volume and speed required drive up both the price and physical size of these memory components.