Source: The Register: Biting the hand that feeds
Traditional Ethernet-based datacenter networks designed for balanced compute/storage/network ratios are buckling under AI cluster demands, which require massive all-to-all bandwidth for model training and distributed inference. Vendors like Nvidia, Intel, and major cloud providers are deploying custom switching fabrics, optical interconnects, and new protocols (like Nvidia's InfiniBand dominance) to handle the skewed traffic patterns of tensor operations—a shift that fragments the ecosystem and locks customers into proprietary stacks. Infrastructure operators now choose between expensive specialized hardware or accepting training bottlenecks, a constraint absent from general-purpose networking for the past two decades.