AI networking overview
Training large AI models spreads work across many GPUs across many servers. The network between them — not any single GPU — often sets how fast a cluster can train, which is why AI fabrics lean on the highest-bandwidth optical transceivers available.
Why the network is critical
- Collective operations — GPUs constantly exchange gradients and parameters in
all-to-all patterns, producing heavy east-west traffic.
- Tail latency matters — a step waits for the slowest link, so consistent low
latency and lossless delivery are essential.
- Scale — thousands of GPUs mean dense, high-radix switches and enormous optics counts.
Two fabric styles
- InfiniBand — long established in HPC, with RDMA and lossless flow control (see
- Ethernet with RoCE — RDMA over Converged Ethernet, increasingly used for AI at
scale on lossless Ethernet.
Optics
AI links are typically 400G and 800G, on QSFP-DD and OSFP modules. The specific interconnect choices and optics are covered in AI interconnects & optics.
The same form factors and management memory apply — CodingBox reads these high-rate modules like any other.