ZTE server technologies were a key part of the company's TCO-optimal AI factory showcase at MWC Shanghai 2026. As large models move into scaled inference deployment, the cost per token is becoming an important measure of AI commercial value, making infrastructure efficiency a central concern for operators, enterprises, and cloud service providers.
The showcase focused on how computing, networks, storage, energy, software algorithms, and scheduling platforms can work together to improve token generation efficiency. By using multidimensional co-design and system-level optimization, ZTE presented a practical path for building AI factories that can support the token economy with better performance and cost control.
Cost Per Token Becomes a Core Metric
In the inference era, AI infrastructure is no longer evaluated only by peak computing power. The ability to generate tokens efficiently, reliably, and at a sustainable cost is becoming equally important. This is especially true as AI agents, multimodal applications, and real-time services create large and continuous inference demand.
ZTE proposed that meaningful improvements in token generation efficiency require architectural innovation and system-level synergy. Individual component upgrades are useful, but they are not enough when AI workloads need coordination across chips, servers, clusters, AIDC facilities, and scheduling software.

This shift changes the business logic of AI infrastructure. When inference becomes a daily production workload, every improvement in energy use, scheduling efficiency, cache performance, and hardware utilization can influence service economics. The AI factory model is designed to connect these factors into one measurable operating system.
OEX Architecture and SuperPOD Design
A central part of the showcase was the OEX architecture based SuperPOD. OEX, or Orthogonal Electrical eXchange, is designed to reduce computing power bottlenecks and maximize energy efficiency through a midplane-free and zero-cable structure.
This design supports physical decoupling and flexible replacement of core components such as GPUs, CPUs, and switch chips. It also supports mainstream high-speed interconnect protocols, including CLink and SUE, creating a framework for multi-chip synergy, open compatibility, and on-demand optimization.
Compared with traditional architectures, shorter communication paths and lower signal loss can improve interconnection efficiency, reduce latency, and enhance system reliability. These attributes are important for large-scale AI training and inference environments that require stable, high-throughput computing.
High-Density Computing for Large AI Workloads
The SuperPOD single rack achieves high-density integration of 128 GPUs and supports scale-up to 16,000 GPUs for extra-large clusters. This provides infrastructure for AI training and inference requirements ranging from thousand-card to ten-thousand-card scale.
Such scale is especially relevant for long-context and high-concurrency agent scenarios. As agentic AI becomes more widely used, infrastructure must support large volumes of simultaneous requests while maintaining predictable service quality and cost efficiency.
High-density design also has practical value for data center planning. By concentrating more computing capability in each rack while maintaining efficient interconnection, operators can improve space utilization and simplify the path toward larger clusters. This supports both current inference requirements and future expansion.
Full Series AI Server and Inference Pool
ZTE also highlighted its Full Series AI Server portfolio, which supports high-density deployment with 8 or 16 GPUs per server and 64 or 128 GPUs per rack. This enables adaptation to different deployment scenarios, from focused enterprise workloads to larger AI infrastructure environments.
The source information also describes an AI-native KV cache implemented through DPU hardware acceleration. By enabling direct GPU access to storage, the design supports zero-copy data transfer, microsecond-level latency, and PB-scale scalability. Combined with intelligent prefetching and dynamic eviction, the cache reaches a hit rate above 70 percent, improving inference efficiency.
Energy Efficiency and Open Evolution
AI factory development must balance performance, cost, and sustainability. ZTE's AIDC solution integrates 800V HVDC power supply, full-stack liquid cooling, and intelligent computing-electricity synergy to support efficient and lower-carbon operations.
The OEX-based SuperPOD also adopts a pre-integration model. Through pre-adaptation and pre-integration, the product adaptation and tuning cycle can be reduced from more than one year to within six months. This helps accelerate ecosystem convergence and commercial rollout while preserving flexibility for future hardware and software evolution.
Conclusion
The MWC Shanghai 2026 showcase demonstrated how ZTE server infrastructure can support more efficient AI factories. By combining high-density AI servers, OEX-based SuperPOD architecture, AIDC energy optimization, and software-driven scheduling, the company is addressing both performance and total cost of ownership.
As AI business models increasingly depend on token production and service delivery, infrastructure design will need to focus on efficiency per watt, efficiency per rack, and efficiency per token. ZTE's approach offers a blueprint for building computing systems that can evolve with the AI economy while keeping deployment and operating costs under control.
For enterprises and operators, the value of this approach lies in predictability. A TCO-oriented AI factory can help transform computing investment into sustained AI productivity, allowing organizations to scale intelligent services without treating infrastructure cost as an uncontrolled variable.
Media Contact
Company Name: ZTE
Contact Person: Leo
Email: Send Email
City: Shenzhen
Country: China
Website: https://www.zte.com.cn/global/index.html