Huawei Unveils Next-Gen Ascend AI Accelerators, Promises Major Performance Leap for China's AI Industry

Ascend 960DT arrives three quarters ahead of schedule, offering up to 4 petaFLOPS FP4 and 288 GB memory, but still lags behind Nvidia's latest Rubin chips

By LineZotpaper
Published
Read Time2 min
Huawei has unveiled its next-generation Ascend 960DT AI accelerator at its annual Connect conference, promising performance far beyond any Nvidia chip currently available in China. The chip, scheduled for Q1 2027 release, delivers up to 4 petaFLOPS of FP4 performance and 288 GB of custom HiZQ memory—double the capacity of its predecessor—though it still trails Nvidia's Rubin and AMD's MI455X in raw compute density.

At Huawei's Connect conference, the company revealed the Ascend 960DT, a new AI accelerator designed to meet the demands of Chinese model developers who are increasingly pressured to adopt domestic alternatives to Western hardware due to US export restrictions. The 960DT arrives three quarters ahead of schedule, according to Huawei, and offers up to 288 GB of the company's custom HiZQ memory—a homegrown high-bandwidth alternative—and up to 4 petaFLOPS of FP4 performance (half that at FP8).

Compared to Nvidia's B300 family, launched last year, the 960DT offers similar memory and bandwidth but only about half the FP8 and a third the FP4 compute. Nvidia's upcoming Rubin chip boasts between 35 and 50 petaFLOPS FP4, 288 GB of HBM4 memory, and 22 TB/s bandwidth, placing it in a different league. However, Rubin and AMD's MI455X are not available for sale in China, leaving the Ascend as a key option for domestic AI training and inference.

Huawei aims to compensate for lower per-chip compute through scale. Using near packaged optics (NPO), the company can scale its compute domain to as many as 4,096 chips, delivering up to 16 exaFLOPS of FP4 compute—a tactic similar to Google's TPU pods. Huawei claims NPO reduces power consumption by 550 kilowatts and halves failure rates compared to pluggable optics, achieving 99.8% uptime. This could address past issues where DeepSeek reportedly struggled with Huawei's earlier NPUs and had to revert to Nvidia GPUs.

Alongside the 960DT, Huawei is developing a compute-optimized variant called the 960PR, expected Q3 2027, featuring up to 8 petaFLOPS FP4 but with lower memory (192 GB). This variant targets prefill and recommender workloads, similar to Nvidia's now-cancelled Rubin CPX accelerators. While the Ascend series represents a significant step for Huawei, analysts note it still has a long way to catch up with Western competitors.

§

Analysis

Why This Matters

  • For Chinese AI developers, the Ascend 960DT offers a domestically produced alternative to Nvidia GPUs at a time when US sanctions restrict access to cutting-edge American hardware.
  • Huawei's scaling approach using near packaged optics could make large-scale AI training more viable in China, potentially reducing dependence on Western infrastructure.
  • The accelerated release schedule (three quarters early) suggests Huawei is under intense pressure to deliver competitive AI hardware quickly.

Background

Huawei has been developing its Ascend series of AI accelerators for several years, positioning them as a homegrown alternative to Nvidia's dominant GPUs. US export controls have progressively tightened, preventing Nvidia from selling its most advanced chips (such as the A100, H100, and now B200/Rubin) to China. This has created a market opportunity for domestic chipmakers like Huawei, though past iterations struggled with performance and reliability. DeepSeek, a prominent Chinese AI lab, reportedly had issues with earlier Huawei NPUs and switched back to Nvidia GPUs. Huawei's networking expertise—it is a leading telecom equipment supplier—gives it an edge in designing scalable interconnects for AI clusters.

Key Perspectives

Huawei: Positions the Ascend 960DT as a leap forward, doubling performance and memory of its predecessor. Emphasises scaling via NPO to overcome per-chip compute deficits. Chinese AI developers: Need reliable, performant domestic alternatives to train large models. The 960DT's improved specs and early availability could alleviate supply chain risks, but past issues with Huawei hardware may breed caution. Western competitors (Nvidia, AMD): Their latest chips (Rubin, MI455X) offer far higher compute density, but are barred from China. They face no direct competitive threat in the Chinese market, but Huawei's scaling approach could eventually challenge their dominance elsewhere if sanctions ease.

What to Watch

  • Adoption by major Chinese AI labs such as DeepSeek and Alibaba; whether they publicly endorse or report issues with the 960DT.
  • Huawei's ability to deliver the 960DT on its accelerated Q1 2027 timeline without further delays.
  • Any adjustments to US export controls that could allow Nvidia to sell Rubin or similar chips to China, potentially undermining Huawei's market opportunity.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.