Huawei Atlas 960E SuperPoD vs NVIDIA Vera Rubin NVL72
The short answer
Huawei’s Atlas 960E SuperPoD (unveiled September 17, 2026) is a 4,096-NPU, 16 EFLOPS FP4 pod built around the first mass-produced near-packaged optics; it ships in 2027 on Ascend 960 silicon. NVIDIA’s Vera Rubin NVL72 is a 72-GPU, ~3.6 EFLOPS NVFP4 rack shipping now, with roughly 12x the per-chip compute and the mature software stack. Huawei competes on scale-up domain size and optical efficiency; NVIDIA on chip performance, availability and ecosystem.
| Huawei Atlas 960E SuperPoD | NVIDIA Vera Rubin NVL72 | |
|---|---|---|
| Announced / ships | September 17, 2026 / 2027 (Ascend 960 roadmap) | GTC 2026 / full production January 2026, shipping H2 2026 |
| Accelerators per scale-up domain | 4,096 Ascend NPUs | 72 Rubin GPUs + 36 Vera CPUs |
| FP4 compute (domain) | 16 EFLOPS (8 EFLOPS FP8) | ~3.6 EFLOPS NVFP4 inference |
| FP4 per accelerator | ~4 PFLOPS (Ascend 960DT); 8 PFLOPS on 960PR inference variant | 50 PFLOPS |
| Memory per accelerator | 288 GB HBM at 9.6 TB/s | 288 GB HBM4 at 22 TB/s |
| Domain memory | ~1.2 PB HBM (4,096 × 288 GB) | 20.7 TB HBM4 + 54 TB LPDDR5X (CPU) |
| Interconnect | UnifiedBus, unified memory addressing; Hi-ONE NPO optics, 7.2 Tbit/s per engine | NVLink 6, ~260 TB/s scale-up; ~28.8 TB/s scale-out |
| Cooling / form | Fully liquid-cooled, orthogonal, multi-cabinet | Liquid-only single rack, ~190–230 kW |
| Cluster ceiling | 512,000 NPUs (two-tier Clos); 1M with multi-rail | Scale-out via Spectrum-X / Quantum-X to AI-factory scale |
| Software | CANN (open-sourced), MindSpore, PyTorch via adapters | CUDA, TensorRT-LLM, Dynamo, SGLang, vLLM |
| Independent benchmarks | None yet | SemiAnalysis AgentX (September 15, 2026): 2.1x–7.2x GB300 tokens/MW |
What Huawei announced
At Huawei Connect 2026 in Shanghai, rotating chairman David Wang launched the Atlas 960E SuperPoD as the industry’s first SuperPoD built on near-packaged optics (NPO). The key numbers from Huawei’s release:
- 4,096 NPUs under one UnifiedBus domain with unified memory addressing, so any NPU can reach memory across physical servers.
- 8 EFLOPS FP8 / 16 EFLOPS FP4.
- 5,500 Hi-ONE optical engines (7.2 Tbit/s each, the only NPO with a built-in light source, per Huawei) replace the ~48,000 800G pluggable modules a conventional build would need, cutting >550 kW of power and doubling fault-free operating time to 99.8% availability.
- SuperCluster scaling to 512,000 NPUs over UnifiedBus or RoCE with a two-tier, four-plane Clos fabric, and to one million with a multi-rail topology.
Alongside it Huawei showed the TaiShan 950 SuperPoD (general-purpose, 4,096 nodes, 256 TB unified memory pool, aimed at agent sandboxes and vector search), OceanStor M900 (a petabyte-scale KV-cache tier for inference), and its Peerium architecture for making large processor groups behave like a single computer. Huawei’s chip roadmap puts the Ascend 960PR in Q3 2027 and 960DT in Q4 2027 (Da Vinci v4 architecture, 288 GB HBM at 9.6 TB/s, 2 PFLOPS FP8 / 4 PFLOPS FP4; the PR variant boosts FP4 inference to 8 PFLOPS), with 970 and 980 later in the decade. Reports at the show said the 960 timeline was pulled forward about nine months.
Where Rubin stands
NVIDIA’s Rubin GPU has 336 billion transistors, 50 PFLOPS of FP4 and 288 GB of HBM4 at 22 TB/s (2.75x Blackwell Ultra’s bandwidth). The Vera Rubin NVL72 rack pairs 72 of them with 36 Vera CPUs: 20.7 TB HBM4 at up to 1,580 TB/s, ~260 TB/s NVLink, 54 TB of LPDDR5X, liquid-only, roughly 190–230 kW per rack. It entered full production in January 2026 and is shipping to hyperscalers and neoclouds in H2 2026; India’s Yotta alone has announced an 80,000-GPU order.
The first verified performance numbers came from SemiAnalysis on September 15, 2026: on agentic inference of DeepSeek V4 Pro, Rubin NVL72 delivered 2.1x GB300’s tokens per megawatt at 100 tokens/second and 7.2x at 150 TPS, above NVIDIA’s own 3x claim — with the caveats that it was pre-release hardware, one model, one scenario and early TensorRT-LLM.
How to read the comparison
Aggregate vs per-chip. Atlas 960E’s 16 EFLOPS is about 4.4x a Rubin rack, from ~57x the accelerators. Per chip, Rubin is ~12x an Ascend 960 on FP4 and 2.3x on memory bandwidth. Huawei’s whole strategy is to make that gap irrelevant at system level: more chips, one memory space, cheaper optics, fewer failures.
Optics is the real innovation. Moving to NPO with an integrated light source attacks the two things that hurt very large scale-up domains — pluggable-module power and module failure rate. If Huawei’s 99.8% availability figure holds in production, that is a meaningful operational advantage at 4,096-accelerator scale, where NVIDIA still uses copper inside the rack and pluggable optics between racks.
Timing and software. Rubin is a 2026 product with a decade of CUDA behind it; Atlas 960E is a 2027 product whose CANN stack only recently moved to community-driven open source. For a buyer who can purchase NVIDIA, that settles it. For buyers inside China’s export perimeter — or governments pursuing sovereign compute — Atlas 960E is the roadmap that matters, and Huawei is explicitly positioning it as “a new option for the world.”
Watch for: independent Atlas 960E throughput numbers, actual Ascend 960 yields on SMIC’s non-EUV process, and whether the 4,096-NPU 960E replaces or complements the larger 15,488-card Atlas 960 configuration Huawei previewed in 2025.