
DriveNets and AMD unveil blueprint for large-scale AI infrastructure
The companies published a validated reference architecture designed to help customers build open AI clusters using AMD Instinct GPUs and DriveNets' networking platform.
Israeli networking company DriveNets and semiconductor giant AMD have unveiled a joint reference architecture for building large-scale artificial intelligence infrastructure, marking a deeper collaboration as AI customers increasingly seek alternatives to vertically integrated computing platforms.
The companies announced the publication of an end-to-end reference architecture designed around AMD's Instinct MI350 series GPUs and DriveNets' AI Fabric networking technology. The blueprint is intended to provide customers with a validated design for deploying production-scale AI clusters spanning training and inference workloads.
The announcement follows AMD's participation as a strategic investor in DriveNets' recent $410 million Series D funding round, underscoring the growing relationship between the two companies as they promote an open, multi-vendor approach to AI infrastructure.
Unlike proprietary AI stacks built around a single vendor, the joint architecture is designed to integrate computing, networking, storage and orchestration within an Ethernet-based platform that customers can scale as AI workloads grow.
Alongside the reference architecture, the companies also published a deployment guide describing how to design, deploy and optimize large GPU clusters. The documentation covers compute and networking design, as well as software optimization across AMD's ROCm ecosystem, RCCL collective communications libraries, network interfaces, servers and orchestration software.
The companies are positioning the platform not only as a technical blueprint but also as evidence that open AI infrastructure can compete on performance.
According to benchmark results included in the validation process, DriveNets' AI Fabric delivered approximately 5% higher throughput and 10% to 15% lower time to first token than publicly available industry results. The companies said the platform also met production targets for multi-node deployments, including inter-token latency below 20 milliseconds and output rates of at least 50 tokens per second per user.
The testing also focused on resilience, an increasingly important requirement as AI clusters grow in size. According to the companies, collective communication performance remained stable even during concurrent RDMA traffic and transient network disruptions, with no measurable degradation or recovery delays. They also said RCCL collective communication performance on AMD hardware was comparable to leading publicly available NCCL benchmarks on equivalent configurations.














