Instinct MI455X and the Helios Rack: Can AMD beat Nvidia in AI?

AMD unveiled Zen 6-based server processors at the Advancing AI 2026 event, dubbing them the best generation it has ever released. The company is focusing on complete solutions including its own Helios server racks, though, which incorporate a new generation of GPUs: the Instinct MI455X, the first GPU based on an architecture derived from the RDNA gaming GPU family and a product that is intended to take AMD to a new level in the AI market.

The Helios rack-based infrastructure includes Epyc 9006 “Venice” processors with Zen 6 cores, Instinct MI455X GPUs, and products that AMD acquired some time ago through its acquisition of Pensando and has since incorporated into its portfolio: Salina DPUs and Vulcano 800Gb/s network adapters. This makes AMD’s AI server offering nearly as comprehensive as Nvidia’s.

Instinct MI455X: The first „UDNA“

The Instinct MI400X GPU generation succeeds the MI350X generation launched last year and represents a significant milestone. It is already based on the CDNA 5 architecture, which is something of a revolution in itself. The first four generations of accelerators based on the CDNA 1 through CDNA 4 architectures still used a GPU design that had evolved from the GCN architecture family, originally introduced in late 2011 and continued by the Polaris and Vega GPUs of 2016 and 2017. CDNA 5 is a significant departure and instead adopts as the foundation of its compute units the design concept used by the RDNA 1 through RDNA 4 architectures—the architecture family to which Radeon gaming GPUs transitioned in 2019. AMD has had two separate architecture lines since then.

AMD Instinct MI455X
AMD Instinct MI455X

However, CDNA 5 marks the beginning of a reunification and convergence of the compute and AI workloads GPU architectures and gaming and graphics architectures, which is something that was announced some time ago (and referred to as UDNA, although this term may ultimately not be used officially to designate the architectures). The established ROCm software stack continues to operate on top of the new hardware architecture and should hopefully bridge the differences for a large proportion of developers and users. CDNA 5 may already incorporate some features of RDNA 5-based gaming GPUs, which are expected to arrive next year.

The First 2nm GPU (for AI)

The Instinct MI450X is the first GPU to use TSMC’s 2nm process—the XCD compute chiplets are manufactured using this node and sit on a base die produced on a 3nm process. MI455X keeps using chiplet-based design using 3D interconnect technology, which Nvidia’s competing Rubin still does not employ (although it adopted disaggregated I/O chiplets connected through 2.5D packaging). Rubin also retains 3nm manufacturing node, giving the Instinct MI455X—or rather the entire MI400X series—a technological advantage in this respect. AMD states a total transistor count of 320 billion, while Nvidia lists 336 billion for Rubin.

AMD Instinct MI455X
AMD Instinct MI455X

The Instinct MI455X accelerator is expected to consist of a package containing two base dies and a total of 16 compute chiplets. Together, these expose a 12,288-bit HBM4 memory interface, allowing the accelerator to carry 12 HBM4 memory packages. This will be another advantage over Nvidia’s Rubin GPU, which has only an 8,192-bit memory interface: AMD’s memory interface is 50% wider and can accommodate 12 HBM4 packages, while Nvidia’s can accommodate only eight.

Up to 40 PFLOPS of Performance

One Instinct MI455X contains 256 WorkGroup Processors (these consist of two CUs in RDNA architectures, so this could correspond to the equivalent of 512 of the older CU blocks, although these may no longer be distinguished separately in newer architectures). AMD states a peak clock speed of 2400 MHz.

At this clock speed, the GPU delivers a theoretical compute performance of 40.3 PFLOPS using 4-bit OCP MXFP4 values and 20.1 PFLOPS using MXFP6 or MXFP8 values, 10.1 PFLOPS in FP16/BFloat16 calculations, or 10.1 petaOPS in INT8. These are figures for performance achieved using matrix cores for AI workloads and including the effect of Structured Sparsity. Without Sparsity, the raw matrix unit throughput is half as high—for example, 5 PFLOPS/POPS in FP16/BFloat16 and INT8.

Oficiální výkonnostní čísla pro AMD Instinct MI455X
Official performance figures for the AMD Instinct MI455X

In general-purpose compute workloads running on the “shader” units (although Instinct accelerators do not support graphics workloads, making the shader designation largely meaningless), performance is 315 TFLOPS and is the same for both FP32 and FP16 calculations. “Packed” operations, referred to as Rapid Math in earlier GPU generations, apparently are not supported natively.

Nvidia lists  Rubin GPU’s theoretical performance as up to 50 PFLOPS including the effect of Sparsity (and apparently 25 PFLOPS without it), so Rubin could be more powerful in terms of raw TFLOPS figures. However, theoretical TFLOPS/PFLOPS figures generally cannot be compared directly across different architectures, particularly when they come from entirely different vendors. The relationship between this theoretical peak performance and real-world performance typically varies considerably between architectures.

432 GB of Memory

The Instinct MI455X is equipped with a total of 432 GB of HBM4 memory, giving it an advantage over Nvidia’s solution thanks to its wider memory interface—Rubin has only 288 GB VRAM in total. MI455X uses HBM4 packages with a capacity of 36 GB each. The larger memory capacity means that larger AI models can be run and trained on the AMD solution. Memory bandwidth is 23.3 TB/s, and ECC protection is supported (the GPU also supports various RAS capabilities and SR-IOV). The accelerator has 192 MB of L2 cache directly on die—or rather, within its chiplets.

While Nvidia GPUs use NVLink to connect to other GPUs and CPUs, the Instinct MI455X uses the UALink standard, which should be capable of interoperating with accelerators and CPUs from other companies (at least in theory). For example, AMD announced a partnership with Cerebras that will enable servers containing Epyc processors, Instinct GPUs, and Cerebras’s Wafer Scale Engine. One Instinct MI455X has a transfer capacity of 600 GB/s (4.8 Tb/s) or 3.6 TB/s (28.8 Tb/s) via UALink over Ethernet. These figures represent the combined bandwidth in both directions of full-duplex communication; bandwidth in one direction is half as high.

Rack AMD Helios s procesory Epyc 9006 a GPU Instinct MI455X
AMD Helios rack with Epyc 9006 processors and Instinct MI455X GPUs

Are Helios Racks the Most Powerful AI Solution?

According to AMD, this GPU—or rather the Helios racks based on it—is the most powerful AI hardware solution in the world. Helios has more than 31 TB of total HBM4 memory capacity, total performance of 2.9 EFLOPS (exaFLOPS) in FP4/MXFP4 calculations, and 260 TB/s of connectivity bandwidth (terabytes per second, not terabits, which are more commonly used for network interfaces). It is expected to contain more than 4,600 CPU cores and more than 18,000 GPU compute units—if we are not mistaken, one Helios rack should contain 72 MI455X GPUs. It has 18 compute servers (“trays”), each carrying two Epyc 9006 processors and four GPUs, as well as six trays dedicated to network connectivity. The system is liquid-cooled.

Rack AMD Helios s procesory Epyc 9006 a GPU Instinct MI455X
AMD Helios rack with Epyc 9006 processors and Instinct MI455X GPUs

According to AMD, Helios with MI455X accelerators offers the highest total GPU compute performance, the greatest HBM4 memory bandwidth and capacity, and the highest interconnect and network connectivity bandwidth, even compared with Nvidia’s latest NVL72 rack solution based on Rubin. Whether this translates into real-world performance would naturally require independent testing, although results will certainly also vary depending on the workload.

Rack AMD Helios s procesory Epyc 9006 a GPU Instinct MI455X
AMD Helios rack with Epyc 9006 processors and Instinct MI455X GPUs

Reports have even emerged that AMD is charging significantly more for a single Helios rack than Nvidia charges for an NVL72 rack with Rubin GPUs, yet still expects it to sell on its merits. However, we do not know whether this information has been verified. It is also true that AMD’s rack contains 50% more HBM4 memory, so the higher price of the complete system makes sense.

Rack AMD Helios s procesory Epyc 9006 a GPU Instinct MI455X
AMD Helios rack with Epyc 9006 processors and Instinct MI455X GPUs

AMD also claims that companies can generate up to 30% more AI tokens per dollar from a single Helios rack. It is also said to be 10–15% more powerful than Nvidia’s NVL72 with Rubin GPUs when both systems are configured with the same power limit, which would indicate better energy efficiency.

Rack AMD Helios s procesory Epyc 9006 a GPU Instinct MI455X
AMD Helios rack with Epyc 9006 processors and Instinct MI455X GPUs

Customer Shipments Starting in September

According to AMD, Helios racks will be deployed by many leading names of the AI sector: Anthropic, OpenAI, Microsoft, Meta, and Oracle. AMD says they are already in full production, with customer shipments scheduled to begin toward the end of this quarter—meaning in September—with volumes said to be ramping during Q4 2026.

Sources: AMD (1, 2, 3)

English translation and edit by Jozef Dudáš</p

Contents

Leave a Reply

Your email address will not be published. Required fields are marked *