CodingBox Dokumentasi

GPU fabric topologies

Tim kami sedang mengerjakan terjemahannya. Artikel ini untuk sementara belum tersedia dalam bahasa Indonesia dan ditampilkan dalam bahasa Inggris.

The network of an AI cluster is built in layers, and each layer uses different interconnects — and different optics. Understanding the layers explains why a cluster has so many transceivers and where a bad one hurts most.

Scale-up vs scale-out

Scale-upScale-out
ConnectsGPUs within a server or a rackservers with each other
Technologyproprietary GPU-to-GPU links (e.g. NVLink/NVSwitch)InfiniBand or Ethernet/RoCE
Mediacopper backplanes and cables, increasingly optics between rackspluggable optics: 400G/800G OSFP, QSFP-DD
Bandwidth per GPUhighesthigh, shared with the fabric

Pluggable transceivers live mostly in the scale-out layer — every GPU server has several 400G/800G ports towards the fabric, plus separate front-end/storage networking.

Fat-tree and non-blocking

The scale-out fabric is usually a fat-tree (Clos): leaf switches connect servers, spine switches connect leaves, with enough leaf–spine uplinks that any server can talk to any other at full rate (non-blocking). Most of a cluster's optics are these leaf–spine and server–leaf links — thousands of identical DR4/DR8 modules.

Rail-optimised design

Training traffic is dominated by collective operations between GPUs of the same rank across servers. A rail-optimised topology exploits this: GPU 0 of every server connects to leaf switch 0, GPU 1 to leaf 1 and so on, so GPUs of the same rank are one hop apart and collectives never cross the spine. The consequence for optics: a fault on a single rail affects one GPU position in every server — a pattern worth recognising.

Separate networks

A cluster typically runs several physically separate fabrics:

  • Compute fabric — GPU-to-GPU (IB or RoCE), the largest and most optic-intensive.
  • Storage fabric — to parallel file systems and object storage, often Ethernet (storage networking).
  • Front-end / management — user access, orchestration, telemetry.

Each has its own optics inventory and its own vendor-validation rules.

A collective operation completes at the speed of its slowest participant. One transceiver with a degrading lane — higher FEC corrections, occasional retransmits — stretches the tail latency of every step and can cost a measurable fraction of the whole cluster's throughput. That is why AI operators trend per-lane DDM across the fleet and replace optics proactively (Transceivers in AI clusters).

In CodingBox

Bench-qualify optics before they go into a rail: identity per CMIS, per-lane DDM baseline on the DDM screen, and a record in the code database so a module's history is available when a rail starts showing errors.

The adapters at the GPU end of these fabrics — BlueField, ConnectX-8 SuperNIC, Pensando — and their twin-port OSFP optics: SmartNICs and DPUs.


Jika Anda menemukan ketidakakuratan atau kesalahan dalam artikel ini, pilih bagian teks yang bersangkutan lalu tekan Ctrl+Enter untuk .