What do you benchmark?
NCCL all-reduce and all-to-all across the full cluster against theoretical bandwidth, storage throughput for checkpoints, and failure injection — results delivered with the repository.
Model labs, inference providers and GPU clouds need lossless, non-blocking back-end fabrics and a team that benchmarks before handover. Rail-optimized 400/800G RoCEv2 designs, tuned on the twin, delivered as code, with the numbers attached.
NCCL all-reduce and all-to-all across the full cluster against theoretical bandwidth, storage throughput for checkpoints, and failure injection — results delivered with the repository.
We build Ethernet. For most clusters under a few thousand GPUs a well-tuned RoCEv2 fabric reaches the same job performance with simpler operations and multi-vendor choice.
Yes. Most problems are PFC/ECN/DCQCN settings, buffer profiles or rail mistakes; we measure, tune on the twin and re-benchmark.
Sizing one to two weeks, design and twin two to four, build dependent on hardware arrival, benchmark and handover one week.
| Rail-optimized topology | Each GPU rail on its own leaf; non-blocking spine; 400G uplinks today, 800G ready. |
|---|---|
| Lossless Ethernet | PFC, ECN, DCQCN and buffer profiles tuned per platform and validated on the twin. |
| Storage and management fabrics | Separate GPUDirect storage, in-band and out-of-band networks, all as code. |
| Platforms | Arista 7060X/7800R, NVIDIA Spectrum-X, Juniper QFX, Cisco Nexus 9000. |
Rail-optimized 400G RoCEv2 fabric for 64 nodes / 512 GPUs: topology, oversubscription, QoS and congestion-control parameters, optics and cabling plan, benchmark methodology.