OTWopentechwire
Tech Intelligence, Openly Wired
AI

DeepSeek Builds Software Bridge for Huawei's AI Chips

Hangzhou start-up open-sources six modules to help domestic processors compete as export controls tighten on US silicon

WZ
Wei Zhang
China Tech Correspondent · Hangzhou
Oct 2, 2026
5 min read
DeepSeek Builds Software Bridge for Huawei's AI Chips
Credit: Reuters

A Software Layer for Domestic Silicon

On 30 September, DeepSeek introduced six open-source software modules engineered specifically for Huawei Technologies' Ascend AI processors. The release mirrors the company's earlier toolkit for Nvidia hardware but targets a fundamentally different goal: establishing what DeepSeek calls an "independent and controllable" software ecosystem around Chinese-manufactured graphics processing units.

The Hangzhou-based AI laboratory framed the move as infrastructure work. Training and deploying large language models demands not only powerful chips but also mature software stacks that handle memory allocation, kernel optimisation, and distributed computation. Nvidia's CUDA platform has held that position for over a decade. DeepSeek's modules attempt to provide parallel functionality for Ascend silicon, lowering the friction for developers who want to migrate workloads away from US-origin hardware.

At Opentechwire, we've tracked similar ecosystem plays across the region - Samsung's push for its own AI accelerators, Japan's investments in post-K supercomputing software, and India's nascent semiconductor incentive schemes. What distinguishes the DeepSeek release is the combination of timing and ambition: it arrives as Washington continues to tighten export controls on advanced AI chips, and it comes from a company that has already demonstrated the ability to train competitive models under constrained hardware budgets.

Export Controls and the Inference Gap

US restrictions on AI chip exports to China have escalated in stages since 2022. The rules initially targeted cutting-edge GPUs with high interconnect bandwidth, then expanded to cover a wider performance envelope. Nvidia responded by designing export-compliant variants with reduced specifications, but those products themselves faced further curbs in 2023. The practical effect has been a widening gap between the compute available to Chinese AI labs and that accessible to their counterparts in North America and Europe.

Huawei's Ascend line represents the most advanced domestic alternative. The 910B chip, introduced in 2023, uses a seven-nanometre process and delivers performance that independent benchmarks place somewhere between Nvidia's A100 and H100 generations - sufficient for many inference tasks and smaller-scale training runs, though still behind the frontier for the largest multimodal models. Manufacturing remains constrained by limits on extreme ultraviolet lithography equipment, which China cannot yet produce domestically and cannot import under current Dutch and US export policies.

DeepSeek's software release does not change the hardware reality, but it addresses a different bottleneck. Even when Chinese organisations acquire Ascend chips, porting existing codebases from CUDA to Huawei's native framework has been labour-intensive. The six modules - covering automatic differentiation, distributed training, memory management, operator libraries, profiling, and model deployment - are intended to smooth that transition. DeepSeek has made the code available under a permissive licence, inviting contributions from the broader developer community.

Building a Parallel Ecosystem

The strategic logic resembles the early years of the Linux kernel or the Android Open Source Project: seed a public-good software layer, encourage adoption through openness, and use network effects to entrench an alternative standard. If enough research labs, universities, and enterprises begin developing on Ascend-compatible tooling, the platform gains momentum independent of any single vendor's roadmap.

DeepSeek's own models provide a proof point. The company's V3 architecture, disclosed earlier in 2026, was trained on a mixture of Nvidia GPUs and Ascend 910B chips. Internal documentation indicated that software optimisation - better scheduling, reduced memory fragmentation, and custom kernels - allowed the team to extract significantly more throughput from each chip than off-the-shelf frameworks delivered. By open-sourcing those optimisations, DeepSeek lowers the barrier for others to replicate the approach.

There are limits to what software can achieve. Certain workloads, particularly those requiring high-precision floating-point operations or massive inter-chip bandwidth, remain bound by silicon capabilities. Nvidia's H200 and the forthcoming Blackwell generation incorporate architectural features - such as transformer engine acceleration and fourth-generation NVLink - that cannot be emulated through clever code. For research groups pushing the frontier of model scale, those advantages matter.

But the bulk of production AI workloads do not sit at the frontier. Inference for conversational agents, recommendation systems, and vision models runs on older-generation hardware in many deployments. Fine-tuning smaller domain-specific models, a common enterprise pattern, requires less compute than pre-training from scratch. DeepSeek's toolkit targets this middle tier, where performance parity is achievable and cost or supply-chain considerations can tip the balance toward domestic chips.

Implications for the Regional Supply Chain

The release also carries implications beyond China's borders. Several Southeast Asian nations are weighing their own semiconductor strategies, balancing integration with US-led supply chains against the risk of future export controls. If Huawei's Ascend platform matures into a viable alternative, it offers a hedge - albeit one that comes with its own geopolitical dependencies.

Singapore's AI research institutes, for instance, have begun evaluating Ascend hardware in pilot projects, though Nvidia remains the dominant choice for production clusters. South Korea's hyperscalers continue to rely on US and domestic GPUs, but Samsung's AI chip division is watching the DeepSeek experiment closely. India's national AI mission, which aims to deploy large-scale compute infrastructure by 2027, has not publicly committed to any single vendor; open-source tooling that supports multiple backends increases optionality.

The software layer is where ecosystems are won or lost. CUDA's entrenchment stems not from technical superiority alone but from a decade of libraries, tutorials, university courses, and pre-trained models that assume its presence. DeepSeek's contribution is a bid to accelerate a parallel path. Whether it succeeds depends on adoption velocity, ongoing investment in the Ascend roadmap, and the trajectory of US export policy.

What Comes Next

DeepSeek has not disclosed how many engineers contributed to the six modules or what resources the company intends to commit to maintenance and documentation. Open-source projects can languish without sustained stewardship, and the AI infrastructure landscape is littered with abandoned frameworks. The company's track record with its earlier Nvidia-focused tools - widely used in Chinese academic labs - suggests it understands the importance of developer experience.

Huawei, for its part, has expanded its Ascend software team and launched a developer programme with cash incentives for contributors. The company has also begun offering cloud instances powered by Ascend chips at prices below comparable Nvidia-based offerings, a tactic to seed adoption. If DeepSeek's modules prove stable and performant, they could accelerate that flywheel.

The broader question is how much decoupling the AI hardware market can sustain. Export controls have created parallel ecosystems before - in telecommunications equipment, satellite navigation, and operating systems - but AI development is unusually iterative and collaborative. Research groups routinely share model weights, benchmark results, and training techniques. A fragmented tooling landscape imposes friction on that exchange.

For now, DeepSeek's release is a data point in a longer contest. It demonstrates that Chinese AI labs are investing not only in models but in the infrastructure that makes those models buildable on domestic silicon. The software works; the question is whether the ecosystem around it can reach critical mass before the next generation of hardware resets the terms of competition.

Read next
AI

Cerebras CEO to Address AI Scaling Limits at TechCrunch Disrupt

Sofia M. Reyes · 4 min
AI

China's Generative AI Users Surpass 700 Million as Adoption Accelerates

Wei Zhang · 5 min
AI

OpenAI Bets on Autonomous Agents That Work While You Sleep

Arjun S. Mehta · 6 min
Spot something wrong? Email corrections@opentechwire.com. We log every correction publicly.