Back to full agenda
Demo ARENA

The Hidden Infrastructure Tax: Why Hyperscaler Defaults Are Killing Your AI Training Performance

Session
810
Day & Date
Wednesday, Sep. 30, 2026
Time
Location
Expo Theater

About this session

At scale, AI training performance depends on more than GPU count or hourly price. Networking, storage, and cluster topology determine how efficiently accelerators communicate and whether expensive capacity translates into training progress. This session examines how infrastructure choices designed for general-purpose cloud workloads can bottleneck large AI clusters. Using benchmark data and a side-by-side comparison of default and purpose-built architectures, we’ll show where performance is lost and how fabric design, storage, and workload placement affect collective communication and end-to-end throughput. Attendees will leave with a practical framework for evaluating AI infrastructure: which architectural questions to ask, which metrics to measure, and how to determine whether a cluster is ready for large-scale training.

Demo: Live cluster topology comparison: starting from a reference architecture built on standard hyperscaler defaults (shared VPC, EFA-style networking, general-purpose storage), I'll walk through each layer — network fabric, storage tier, scheduler configuration — and show exactly where the performance degradation occurs and why. Using real benchmark data from collective communication tests at scale, the demo illustrates the gap between published specs and production throughput, and shows what a purpose-built AI infrastructure design looks like at the same node count. The goal is a side-by-side that any architect in the room can map directly to their own environment.

Share this session

Get in the room. San Francisco, September 29.

Fully Connected 2026 is where the engineers, leaders, and operators running AI in production come together for three days of depth, access, and real conversation.