Source: SiliconANGLE
The shift from training to inference-heavy AI workloads is creating a storage crisis at a specific, previously overlooked layer: the key-value caches that LLMs need to keep in memory during token generation. VAST's pivot here reflects real infrastructure pain—companies building AI systems are hitting memory limits faster than compute limits, and traditional cloud storage can't handle the random-access patterns required. Specialist vendors are now hunting the exabyte-scale cache market that didn't exist two years ago. Whoever controls the cache layer owns a critical chokepoint in AI deployment, much as GPU makers owned compute bottlenecks.