// memory technology

All signals tagged with this topic

High-bandwidth flash could reshape GPU memory architecture

Researchers are developing storage-class memory that marries SSD capacities (multiple terabytes) with HBM speeds, potentially easing the current GPU bottleneck where limited VRAM forces constant data shuffling to system RAM. The constraint is economics and thermal overhead—whether the cost and power demands of high-bandwidth flash justify replacing traditional memory hierarchies, especially when chip designers can already optimize for larger models through other means. The technology matters for AI training at scale and real-time inference, but it succeeds only if it outpaces the incremental improvements chip makers are already shipping through better software and conventional memory stacking.

Intel's Optane Memory Could Have Solved AI's RAM Bottleneck

Intel discontinued Optane in 2022—years before the generative AI boom made its extreme write endurance and ultra-low latency valuable for KV cache acceleration. The timing was a costly product strategy failure: Optane was engineered for a shrinking problem (high-frequency trading, database writes), while the actual killer app (batching LLM inference requests) emerged too late for the investment case to survive. This created an opening for competitors like Nvidia (with NVLink-attached memory) and custom silicon makers to capture the AI memory acceleration market, locking in architectural choices that will persist for years.