CodingBox Dokumentasi

Gambaran keseluruhan rangkaian AI

Pasukan kami sedang menyiapkan terjemahan. Artikel ini buat sementara waktu belum tersedia dalam bahasa Melayu dan dipaparkan dalam bahasa Inggeris.

Training large AI models spreads work across many GPUs across many servers. The network between them — not any single GPU — often sets how fast a cluster can train, which is why AI fabrics lean on the highest-bandwidth optical transceivers available.

Why the network is critical

  • Collective operations — GPUs constantly exchange gradients and parameters in all-to-all patterns, producing heavy east-west traffic.
  • Tail latency matters — a step waits for the slowest link, so consistent low latency and lossless delivery are essential.
  • Scale — thousands of GPUs mean dense, high-radix switches and enormous optics counts.

Two fabric styles

  • InfiniBand — long established in HPC, with RDMA and lossless flow control (see InfiniBand).
  • Ethernet with RoCE — RDMA over Converged Ethernet, increasingly used for AI at scale on lossless Ethernet.

Optics

AI links are typically 400G and 800G, on QSFP-DD and OSFP modules. The specific interconnect choices and optics are covered in AI interconnects & optics.

The same form factors and management memory apply — CodingBox reads these high-rate modules like any other.

Further reading


Jika anda menemui ketidaktepatan atau kesilapan dalam artikel ini, pilih bahagian teks yang berkenaan dan tekan Ctrl+Enter untuk .