CodingBox Tài liệu

AI networking overview

Đội ngũ của chúng tôi đang thực hiện bản dịch. Bài viết này tạm thời chưa có bản tiếng Việt và đang được hiển thị bằng tiếng Anh.

Training large AI models spreads work across many GPUs across many servers. The network between them — not any single GPU — often sets how fast a cluster can train, which is why AI fabrics lean on the highest-bandwidth optical transceivers available.

Why the network is critical

  • Collective operations — GPUs constantly exchange gradients and parameters in all-to-all patterns, producing heavy east-west traffic.
  • Tail latency matters — a step waits for the slowest link, so consistent low latency and lossless delivery are essential.
  • Scale — thousands of GPUs mean dense, high-radix switches and enormous optics counts.

Two fabric styles

  • InfiniBand — long established in HPC, with RDMA and lossless flow control (see InfiniBand).
  • Ethernet with RoCE — RDMA over Converged Ethernet, increasingly used for AI at scale on lossless Ethernet.

Optics

AI links are typically 400G and 800G, on QSFP-DD and OSFP modules. The specific interconnect choices and optics are covered in AI interconnects & optics.

The same form factors and management memory apply — CodingBox reads these high-rate modules like any other.

Further reading


Nếu bạn phát hiện thông tin chưa chính xác hoặc lỗi trong bài viết này, hãy chọn đoạn liên quan và nhấn Ctrl+Enter để .