At Huawei's Connect conference, the company revealed the Ascend 960DT, a new AI accelerator designed to meet the demands of Chinese model developers who are increasingly pressured to adopt domestic alternatives to Western hardware due to US export restrictions. The 960DT arrives three quarters ahead of schedule, according to Huawei, and offers up to 288 GB of the company's custom HiZQ memory—a homegrown high-bandwidth alternative—and up to 4 petaFLOPS of FP4 performance (half that at FP8).
Compared to Nvidia's B300 family, launched last year, the 960DT offers similar memory and bandwidth but only about half the FP8 and a third the FP4 compute. Nvidia's upcoming Rubin chip boasts between 35 and 50 petaFLOPS FP4, 288 GB of HBM4 memory, and 22 TB/s bandwidth, placing it in a different league. However, Rubin and AMD's MI455X are not available for sale in China, leaving the Ascend as a key option for domestic AI training and inference.
Huawei aims to compensate for lower per-chip compute through scale. Using near packaged optics (NPO), the company can scale its compute domain to as many as 4,096 chips, delivering up to 16 exaFLOPS of FP4 compute—a tactic similar to Google's TPU pods. Huawei claims NPO reduces power consumption by 550 kilowatts and halves failure rates compared to pluggable optics, achieving 99.8% uptime. This could address past issues where DeepSeek reportedly struggled with Huawei's earlier NPUs and had to revert to Nvidia GPUs.
Alongside the 960DT, Huawei is developing a compute-optimized variant called the 960PR, expected Q3 2027, featuring up to 8 petaFLOPS FP4 but with lower memory (192 GB). This variant targets prefill and recommender workloads, similar to Nvidia's now-cancelled Rubin CPX accelerators. While the Ascend series represents a significant step for Huawei, analysts note it still has a long way to catch up with Western competitors.