Kiwi Technology and Biren Technology GPU Direct RDMA Solution: Measured Throughput Soars, Latency Halved
2026-07-20 11:26
Favorite

en.Wedoany.com Reported - Kiwi Technology and Biren Technology showcase China's domestic GPU direct-connect solution at the World Artificial Intelligence Conference

At the 2026 World Artificial Intelligence Conference (WAIC 2026), held in Shanghai from July 17 to 20, 2026, AI network interconnect technology company Kiwi Technology and domestic GPU manufacturer Biren Technology jointly demonstrated the IBGDA communication solution for "domestic GPU direct connection to domestic RDMA network cards." Measured data shows that in small-packet, high-frequency scenarios, the solution achieved a significant leap in All-to-All communication throughput, with latency nearly halved. This achievement marks a key breakthrough in system-level collaboration for domestic AI computing clusters.

Kiwi Technology is a tech company specializing in full-stack AI network interconnect products and solutions, covering Scale-Out inter-node networking, Scale-Up supernode XPU interconnects, and next-generation optical interconnects. Biren Technology is a domestic high-performance general-purpose GPU chip design company dedicated to providing an autonomous and controllable computing foundation for AI computing. For a long time, efficient cluster communication for domestic GPUs has relied on NVIDIA network cards, creating a major bottleneck for the autonomy of domestic computing clusters. The IBGDA solution jointly demonstrated by the two companies achieves deep collaboration between domestic GPUs and domestic RDMA network cards through Kiwi Technology's self-developed single-channel 400G RDMA ASIC engine (based on the Kiwi SNIC 800G platform architecture).

In typical training and inference scenarios for MoE (Mixture of Experts) models, expert parallel communication involves two core stages: Dispatch and Combine, each requiring massive small-granularity data exchanges. Traditional cross-node communication methods require the GPU to relay instructions through the CPU, introducing additional latency and CPU overhead. By bypassing the CPU processing stage, IBGDA technology significantly reduces instruction relay wait times, fully unleashing the throughput advantages of GPUs in parallel computing. Measured data shows that IBGDA demonstrates generational advantages over traditional solutions in small-packet, high-frequency, strongly synchronized scenarios: single small-message bandwidth increases significantly, throughput for both raw small-packet sending and combined sending achieves a qualitative leap, and the "small-packet penalty" where small-message latency far exceeds large-message latency in traditional solutions is completely eliminated. Built on an open Ethernet ecosystem, this solution inherits the advantages of flexible deployment and cost control of Ethernet while providing a high-performance and autonomous network option for domestic computing clusters.

At this conference, Kiwi Technology, together with the Shanghai Computing Network Association, Shanghai Artificial Intelligence Laboratory, the East China Branch of the China Academy of Information and Communications Technology, as well as industry chain companies including Biren Technology, Muxi, Tianshu Zhixin, Enflame Technology, SenseTime, and Sugon Information, jointly launched the "Domestic Supernode Ecosystem Co-building Initiative." Additionally, Kiwi Technology, in collaboration with the Hong Kong University of Science and Technology (Guangzhou), XiWang Technology, Biren Technology, Turing Quantum, and Singular Photonics, jointly released the "Co-Packaged Optics (CPO) Technology White Paper," providing guidance for the standardization and industrialization of domestic optical interconnect technology.

This measured verification of the domestic GPU direct-connect solution proves for the first time with data that the "domestic chip + domestic network" technical path already possesses practical capabilities for system-level collaborative optimization. This breakthrough helps drive the leapfrog development of domestic AI computing power from "single-chip breakthroughs" to "system-level collaborative optimization."

 

This bulletin is compiled and reposted from information of global Internet and strategic partners, aiming to provide communication for readers. If there is any infringement or other issues, please inform us in time. We will make modifications or deletions accordingly. Unauthorized reproduction of this article is strictly prohibited. Email: news@wedoany.com