Huawei Pushes Homegrown AI Stack with Optics Breakthrough and 960 SuperPoD
The Shenzhen firm is integrating near-packaged optics into its Ascend compute clusters, targeting the interconnect bottleneck that has constrained China's push for AI sovereignty.

An Interconnect Play in the AI Arms Race
At Huawei Connect 2026 in Shanghai this week, the Shenzhen-based technology group rolled out two pieces of infrastructure that clarify its AI strategy: the Ascend 960 SuperPoD compute cluster and an upgraded UnifiedBus interconnect built around near-packaged optics. The announcement, made on Thursday, underscores Huawei's determination to supply China's hyperscalers and research labs with an alternative to Nvidia's ecosystem, even as US semiconductor restrictions continue to limit access to cutting-edge chips and packaging technologies.
Near-packaged optics - or NPOs - sit at the heart of the new design. The approach places optical transceivers immediately adjacent to compute dies, shortening signal paths and reducing both latency and power draw. For training runs that shuffle terabytes of gradient updates across hundreds or thousands of accelerators, every nanosecond and every watt counts. Huawei's decision to integrate NPOs into UnifiedBus signals that it views the interconnect, not just raw compute throughput, as the lever that will determine who wins the race to train frontier models at scale.
What the 960 SuperPoD Brings
The Ascend 960 SuperPoD packages Huawei's latest AI accelerators into a rack-scale cluster optimised for distributed training. According to the company, the system delivers higher aggregate bandwidth than its predecessor, the 910 series, and supports larger model configurations without hitting the communication ceiling that often stalls multi-node runs. Huawei did not disclose peak FLOPS figures or memory capacity per node during the Shanghai event, but the emphasis on interconnect upgrades suggests that bandwidth, rather than arithmetic density, was the primary design constraint the engineering team sought to address.
The SuperPoD architecture mirrors the design philosophy seen in other hyperscale training clusters: tightly coupled nodes, high-radix switches, and a fabric that prioritises all-to-all collectives. What sets Huawei's offering apart is the integration layer. By moving optics closer to the package, the firm reduces the number of electrical-to-optical conversions and shortens the physical reach between chips, both of which translate into lower latency and tighter synchronisation across the cluster.
The NPO Advantage and Its Limits
Near-packaged optics is not a Huawei invention - hyperscalers in the United States and Europe have been experimenting with co-packaged and near-packaged designs for several years - but bringing the technology to volume production in China represents a meaningful step. Traditional pluggable optics sit on the edge of a server board, connected via copper traces that introduce signal degradation and thermal challenges. NPOs eliminate much of that distance, allowing designers to push data rates higher without breaching power budgets or reliability thresholds.
Still, the approach carries trade-offs. Placing optics so close to the die complicates thermal management, as both the compute engine and the transceiver generate significant heat in a confined area. Yield rates can suffer if the assembly process is not tightly controlled, and serviceability becomes more difficult when a failed optical component requires replacing an entire module rather than a pluggable transceiver. Huawei will need to demonstrate that its manufacturing partners can deliver the necessary precision at scale, particularly given that some of the advanced packaging equipment typically used for such tasks falls under export control regimes administered by the United States, Japan, and the Netherlands.
The Nvidia Shadow
Huawei's push into NPOs and high-bandwidth interconnects takes place in the shadow of Nvidia's dominance. The US chipmaker's H100 and H200 accelerators, paired with NVLink and InfiniBand fabrics, remain the reference architecture for most large language model training runs outside China. Chinese AI labs have access to downgraded variants - such as the A800 and H800 - but those chips carry reduced interconnect speeds, a deliberate hobbling intended to slow the development of advanced military and surveillance applications.
Export controls imposed in 2022 and tightened in subsequent rounds have made it harder for Chinese firms to procure not only the chips themselves but also the networking silicon and optical components that bind clusters together. This has created both urgency and opportunity for Huawei. If it can deliver a credible domestic alternative, the company stands to capture a substantial share of China's AI infrastructure spending, which analysts estimate will exceed USD twenty billion this year.
At Opentechwire, we have tracked how Beijing has responded to semiconductor restrictions by funnelling capital into indigenous chip design, fabrication, and packaging. Huawei's Ascend roadmap is a direct beneficiary of that policy pivot. The firm's HiSilicon subsidiary designs the accelerators, while domestic foundries - operating at process nodes that, while trailing TSMC and Samsung, are sufficient for many AI workloads - handle manufacturing. The addition of NPO capability suggests that Huawei is also making headway in advanced packaging, an area where China has historically lagged.
Ecosystem Gaps Remain
Hardware is only part of the equation. Nvidia's moat is as much software as silicon: CUDA, cuDNN, TensorRT, and a vast library of optimised kernels that researchers and engineers have spent years refining. Huawei's CANN - Compute Architecture for Neural Networks - aims to replicate that stack, but adoption has been slow outside state-backed institutions and companies with strong ties to the Chinese government. Developers accustomed to PyTorch and TensorFlow workflows face a steep learning curve when porting models to Ascend, and the documentation and community support lag far behind what Nvidia offers.
The Shanghai conference included sessions on CANN updates and new tools designed to ease migration from CUDA-based codebases. Whether those efforts will be enough to attract the broader developer community remains an open question. Inertia is powerful in infrastructure choices, and switching costs are high when training pipelines, monitoring tools, and operational runbooks are all built around a single vendor's ecosystem.
What Comes Next
Huawei has not disclosed a detailed roadmap beyond the 960 SuperPoD, but the trajectory is clear: tighter integration, higher bandwidth, and a relentless focus on closing the performance gap with US offerings. The company is also likely to lean harder on its networking heritage. Huawei built its reputation on telecom switches and routers, and that expertise translates naturally into the high-radix, low-latency fabrics that AI clusters demand.
For China's AI ambitions, the stakes are existential. The country's leading labs - Baidu, Alibaba, Tencent, ByteDance - are racing to train models that can compete with GPT-class systems developed in the United States. Doing so without access to the latest Nvidia hardware requires either accepting a performance penalty or betting on domestic alternatives like Ascend. Huawei's latest announcements make that bet incrementally more credible, though the company still has ground to cover in both silicon performance and software maturity.
The near-packaged optics story is particularly telling. It shows Huawei targeting not the headline metric - FLOPS per chip - but the plumbing that often determines real-world training throughput. In distributed AI, bottlenecks tend to emerge in the interconnect, the memory subsystem, or the I/O path long before raw compute becomes the limiting factor. By addressing those chokepoints, Huawei is signalling a more sophisticated understanding of what it takes to build systems that scale, not just chips that benchmark well in isolation.
Whether that sophistication will be enough to displace Nvidia, even within China, depends on factors beyond Huawei's control: the pace of US policy tightening, the willingness of Chinese hyperscalers to absorb switching costs, and the ability of domestic foundries and packaging houses to deliver the yields and volumes that a true Nvidia alternative would require. For now, the Ascend 960 SuperPoD and its NPO-equipped interconnect represent another incremental step in a long campaign to build an indigenous AI stack - one that assumes the door to Santa Clara will remain closed for the foreseeable future.


