CodingBox مستندات

مرور کلی شبکه‌های هوش مصنوعی

تیم ما در حال ترجمه است. این مطلب موقتاً به فارسی در دسترس نیست و به زبان انگلیسی نمایش داده می‌شود.

Training large AI models spreads work across many GPUs across many servers. The network between them — not any single GPU — often sets how fast a cluster can train, which is why AI fabrics lean on the highest-bandwidth optical transceivers available.

Why the network is critical

  • Collective operationsGPUs constantly exchange gradients and parameters in all-to-all patterns, producing heavy east-west traffic.
  • Tail latency mattersa step waits for the slowest link, so consistent low latency and lossless delivery are essential.
  • Scalethousands of GPUs mean dense, high-radix switches and enormous optics counts.

Two fabric styles

  • InfiniBand — long established in HPC, with RDMA and lossless flow control (see InfiniBand).
  • Ethernet with RoCERDMA over Converged Ethernet, increasingly used for AI at scale on lossless Ethernet.

Optics

AI links are typically 400G and 800G, on QSFP-DD and OSFP modules. The specific interconnect choices and optics are covered in AI interconnects & optics.

The same form factors and management memory apply — CodingBox reads these high-rate modules like any other.

Further reading


اگر در این مطلب نادرستی یا خطایی دیدید، بخش مربوط را انتخاب کنید و با فشردن Ctrl+Enter .