ZTE Introduces AI Supernode for System-Level Computing Power Optimization

Open Source Talent Scout
Futuristic data center with glowing server racks, symbolizing AI supernode technology and system-level computing.

ZTE has unveiled a new AI supernode technology designed to move beyond single-chip performance limitations and establish a system-level collaborative computing foundation. This initiative addresses the challenges posed by trillion-parameter large models, where traditional "chip stacking" approaches are becoming unsustainable due to bottlenecks in power consumption, interconnect bandwidth, and memory capacity of individual GPUs.

Close-up of interconnected GPU chips on a circuit board, illustrating system-level computing.

Close-up of interconnected GPU chips on a circuit board, illustrating system-level computing.

The company's approach aims to reconstruct the intelligent computing foundation by integrating dozens to hundreds of multi-vendor GPUs into a unified computing unit. This strategy focuses on system-level optimization of computing power rather than solely on individual chip development. The "ZTE Supernode White Paper" outlines this solution, which seeks to redefine AI computing power infrastructure.

System-Level Collaboration as Core Logic

The industry largely agrees that integrating multiple GPUs into a "supercomputer" through high-speed, lossless interconnection is key to overcoming single-chip performance limits. ZTE's supernode design aligns with this principle, shifting focus from single-chip performance competition to system-level computing power collaboration. This strategy avoids the high barriers and long development cycles associated with GPU chip research and development, instead addressing the inefficiency of multi-chip collaboration in current computing power models.

Abstract network of glowing nodes and lines, representing system-level collaboration in computing.

Abstract network of glowing nodes and lines, representing system-level collaboration in computing.

ZTE's supernode is an integrated system that combines multiple chips, hardware, high-speed interconnection, and supporting software. Its design is based on four prerequisites for system-level computing power collaboration:

  • Chip Capability Balance: Matching GPU computing power, video memory, and interconnect bandwidth to prevent resource waste.

  • Interconnection Architecture Effectiveness: Achieving interconnect bandwidth between any GPUs within the supernode that is approximately eight times greater than inter-machine interconnection, balancing communication efficiency, scalability, and adaptability.

  • Memory Access Convenience: Ensuring all GPUs support unified memory addressing, compatible with memory and message semantics for ease of programming and data access.

  • Architectural Expansion: Ensuring that expanded clusters remain within a high-bandwidth domain to support on-demand computing power configuration.

Hardware Architecture Innovation

Traditional GPU clusters often rely on cable tray architectures, which introduce signal loss, reduce computing power density, and increase maintenance and networking costs. ZTE's supernode hardware architecture introduces the Orthogonal Electrical eXchange (OEX) orthogonal backplane-free interconnection and switching architecture. This design, recognized as an ODCC "Annual Major Technological Breakthrough" in 2025, reconstructs the GPU physical interconnection system to support high-density, high-reliability GPU collaboration.

Detailed view of OEX orthogonal backplane-free interconnection architecture in a server rack.

Detailed view of OEX orthogonal backplane-free interconnection architecture in a server rack.

The OEX architecture enables vertical cross-physical direct connection between computing and switching trays, eliminating traditional high-speed cables. This cable-free system uses orthogonal connectors and a single-stage switching topology. According to the white paper, this design shortens the SerDes link length by over 30% in 112G high-speed signal scenarios, removing 6.5dB of insertion loss from cables and improving the end-to-end link insertion loss margin to over 3dB. This reduces the bit error rate and supports TB-level interconnect bandwidth.

The cable-free design also allows standard cabinets to integrate 64/128 or more GPUs, increasing computing power density per unit space. It reduces downtime risks from cable issues, shortening system fault repair times from hours to minutes, which is critical for 24/7 AI model training. Additionally, integrating parameter plane leaf switching within the switchboard removes the need for leaf-level switches, optical modules, and fibers, simplifying the system and reducing hardware costs.

Streaks of light representing high-speed data flow through fiber optic and copper cables.

Streaks of light representing high-speed data flow through fiber optic and copper cables.

High-Speed Interconnection Technology

Efficient GPU interconnection is crucial for system-level computing power. ZTE has developed high-speed interconnection technology across five dimensions: chips, physical layer, protocol layer, computing offload, and scalability. This aims to provide TB-level communication channels for AI computing demands.

ZTE's self-developed large-capacity switching chip is central to this. It offers TB-level bandwidth and hundreds of nanoseconds latency, supports large-scale all-to-all interconnection for dozens to hundreds of GPUs, and is compatible with mainstream interconnection protocols such as RDMA, CLink, OISA, Ethlink, SUE, and UEC.

The company opted for the Ethernet physical layer over the traditional PCIe bus. PCIe 5.0 x16 offers about 128GB/s bidirectional bandwidth, while Ethernet SerDes rates have reached 112Gbps, with 224Gbps products becoming available. This supports flexible multi-channel binding for TB/s-level port bandwidth, meeting AI training requirements.

At the protocol layer, ZTE supports open protocols like UALink and ESUN and participates in the CLink protocol formulation led by the Ministry of Industry and Information Technology. The company also integrates in-network computing into its switching chip, offloading high-load communication operations from GPUs. This reduces the complexity of All-Reduce operations in dense model training and decreases dispatch and reduction latency in MoE (Mixture of Experts) model training by 20-60%, while reducing trunk traffic by over 30%.

ZTE has also designed Scale-Up scalability across interconnection protocol, topology, physical form, and medium. It includes GPU ID identification bits for future clusters of tens of thousands of GPUs, uses linear non-convergent expansion topology, and adopts modular cabinet designs for "plug-and-play" expansion. The strategy of using "copper where possible, fiber for long distances" balances transmission efficiency and cost.

Server rack with visible liquid cooling plates and tubes, illustrating advanced power management.

Server rack with visible liquid cooling plates and tubes, illustrating advanced power management.

Power Management Innovation

High-density GPU integration increases power consumption and heat generation. NVIDIA data cited in the white paper indicates that GPU supernode cabinet power consumption is projected to rise from 50kW for H100 in 2022 to 120-150kW for GB300 NVL72 in 2025, potentially reaching 600kW or more.

ZTE's supernode incorporates a "forward-looking, all-dimensional adaptive" power management system. For heat dissipation, it uses single-phase cold plate liquid cooling, which is widely adopted and supports cabinets up to hundreds of kilowatts. For future single-chip power consumption exceeding 2000W, plans include silicon-based microchannel cold plates and two-phase cold plate liquid cooling. Immersion liquid cooling technology is also supported for future megawatt-level cabinets.

For power supply, ZTE uses a High-Voltage DC (HVDC) architecture, with mainstream evolution towards ±400V DC and 800V DC. This reduces current by 8-16 times for the same power, cutting copper consumption by 40-50% and freeing up space. It also improves overall end-to-end power supply efficiency by 3-5%, which can lead to significant operational cost savings in data centers where electricity costs are substantial. This architecture supports cabinet power levels from 100-150kW up to 1MW+ and reduces intermediate energy conversion layers.

Cluster Expansion and Software Stack

While a single supernode has limits, ZTE's Nebula Matrix cluster supernode allows for scalable expansion from hundreds to tens of thousands of cards. This system uses an "electrical switching + optical interconnection" route, with high-performance electrical switches for intra-cabinet GPU interconnection and optical fiber for inter-cabinet connections. This approach leverages the maturity of electrical switching and avoids the complexities of all-optical switching.

The existing Nebula X32 single supernode can expand into Nebula Matrix X256/800 cluster supernodes. Future, higher-density Nebula X128 single supernodes can further expand to ultra-large-scale clusters of X8192/16384. ZTE also offers a Scale-Up and Scale-Out network convergence design, building a unified supernode interconnection network that handles both high-bandwidth, low-latency communication (e.g., tensor parallelism) and lower-performance communication (e.g., data parallelism). Model calculations suggest this converged architecture can significantly reduce total cost of ownership (TCO) compared to independent networking.

ZTE's supernode also features a comprehensive software stack, described as its "operating system," for unified scheduling, management, optimization, and monitoring of hardware resources. This includes virtualized resource pools, intelligent orchestration, communication optimization, topology awareness, and unified scheduling for heterogeneous computing. The system also provides full-stack observability, intelligent operations and maintenance, high-reliability redundancy mechanisms, and "computing power-electricity" collaborative green scheduling.

Abstract digital brain with overlaid UI, symbolizing a comprehensive software stack for AI.

Abstract digital brain with overlaid UI, symbolizing a comprehensive software stack for AI.

A computing power simulation platform provides "digital twin" inference capabilities for supernode configuration. For instance, it indicates that for the Qwen3-235B model, a 256-card supernode can improve training performance by 15% compared to an 8-card server at a 2K card scale.

Multi-Vendor GPU Compatibility

ZTE's supernode emphasizes multi-vendor GPU compatibility, aiming to break ecosystem lock-in. This is achieved through systematic design across hardware, chips, protocols, ecosystem, and clusters.

At the hardware layer, the OEX orthogonal architecture of ZTE's Nebula single supernode uses a modular design where the GPU adaptation module (UBB) can be replaced for different manufacturers' GPUs without altering the overall architecture.

At the chip layer, ZTE's self-developed large-capacity switching chip is compatible with mainstream GPU interconnection protocols, supporting "one design, multi-card compatibility."

At the protocol layer, ZTE participates in the CLink protocol formulation and offers its open-standard OLink protocol.

At the ecosystem layer, ZTE provides open mechanical and electrical interface specifications for the OEX orthogonal architecture. The company also established the "Supernode Hardware System Based on Orthogonal Architecture" project within the ODCC Network Working Group in June 2025 to promote industry standardization.

At the cluster layer, multi-vendor GPU compatibility extends to the Nebula Matrix cluster supernode. Its converged networking architecture supports high-bandwidth, low-latency collaboration across cabinets and brands, even with mixed GPU brands within the same supernode or cluster.

Diverse GPU chips from multiple vendors integrated into a unified, glowing supernode.

Diverse GPU chips from multiple vendors integrated into a unified, glowing supernode.

ZTE's supernode technology focuses on being a TCO-optimal system-level integrator of computing power. By prioritizing system-level collaboration over chip R&D competition, the company aims to address performance bottlenecks of single GPU chips. The white paper indicates that MoE model dispatch latency can be reduced by 20-50%, and reduction latency by over 40-60%. This approach transforms computing power construction from "chip stacking" to "collaborative release" and from "single hardware performance competition" to "full-stack system optimization."

The company's open and compatible design aims to provide flexible GPU choices and promote openness in the intelligent computing industry. ZTE states that future computing power competition will shift from "FLOPS (floating-point operations per second)" to "Tokens per Watt," with its supernode designed to optimize computing power efficiency, expansion capability, and ecosystem compatibility.

ToolMesh
ToolMesh Weekly

Stay Ahead of the AI Curve

Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.

No spam, unsubscribe at any time.