Workload-to-topology sizing
GPU count, collective patterns and model sizes translated into rail-optimized leaf-spine with the right oversubscription — usually none.
Rail-optimized 400G and 800G Ethernet fabrics with RoCEv2, designed around your training and inference workloads — PFC, ECN and QoS tuned on the twin, NCCL benchmarks run at handover, and the fabric delivered as code.
GPU count, collective patterns and model sizes translated into rail-optimized leaf-spine with the right oversubscription — usually none.
RoCEv2 with PFC, ECN and DCQCN parameters, buffer profiles and QoS classes tuned per platform and proven on the twin.
Arista 7060X/7800R, NVIDIA Spectrum-X, Juniper QFX or Cisco Nexus 9000 — chosen for the workload, with 800G and 1.6T-ready optics.
Separate fabrics for storage (GPUDirect), in-band management and out-of-band, each as code.
NCCL all-reduce and all-to-all benchmarks across the full cluster, failure injection, and a performance baseline you keep.
Telemetry for ECN marks, PFC pauses and buffer occupancy in the live portal; optional AIOps operations.
Workload interviews and cluster plan. One to two weeks.
Topology, QoS, optics, cabling, as code; twin validation.
Rack-and-stack, cabling certification, bring-up.
NCCL results against targets, documentation, repository.
We build Ethernet fabrics. For most clusters under a few thousand GPUs, a well-tuned RoCEv2 fabric reaches the same job performance with simpler operations and multi-vendor choice.
Yes. Most under-performing training clusters trace back to QoS and congestion-control settings we can measure and fix without new hardware.
Tell us about the project; a senior engineer responds the same business day.