Перейти к содержимому

Building a GPU Fabric, Ep. 1: Why a GPU cluster needs a different network

MillenniumMesh

0:00 / 0:00

Building a GPU Fabric, Ep. 1: Why a GPU cluster needs a different network

30 просмотров · 5 дней назад
MillenniumMesh
1 подписчик
30 просмотров · 5 дней назад
A network engineer's lab notes: building the backend network for GPU servers on real hardware — Juniper QFX5K spines and leafs, two servers with 8× H100 and dual-port ConnectX-7 200G NICs, RoCEv2 traffic. In this episode: what RDMA and RoCEv2 are, why a GPU fabric has to be lossless, the three networks of a GPU cluster, rail-optimized design, and how to find which NIC sits next to which GPU. Chapters 0:00 Intro 0:16 The test stand 1:01 RoCEv2 and why packet loss hurts 2:03 Three networks and our limits 2:52 Rails 3:11 Pure L3 fabric, built with Apstra 3:48 Planning the rails 4:19 Server topology with nvidia-smi 5:16 Wrap-up and next episode Reference architecture (Juniper JVD): https://www. juniper. net/documentation/us/en/software/jvd/jvd-ai-dc-apstra-amd/solution_architecture.html Next episode: first bandwidth tests with perftest — RoCE over a lossy and a lossless network. #RoCE #RDMA #AINetworking #DataCenter