Buses & I/O — Overview
A CPU, RAM, storage, and peripherals are separate physical chips — a bus is the shared electrical
A CPU, RAM, storage, and peripherals are separate physical chips — a bus is the shared electrical
Everything so far in this section covers bandwidth inside a single GPU — SM to L2, L2 to HBM. The moment a workload needs data on another device, whether that's the host CPU or a second GPU, a completely different and usually much slower link is in the critical path. Which link is available, and at what bandwidth, is not a software choice — it's a property of the physical topology of the machine, and designing a multi-GPU strategy without first knowing that topology is a common source of disappointing scaling.
The physical storage medium (platters, NAND cells) is only half the story — data still has to travel
Every wire diagram of a computer hides the same question: how do the CPU, RAM, and peripherals