AMD unveiled the Instinct MI455X, a CDNA 5 AI accelerator with 320 billion transistors and 432GB of HBM4 memory, designed to power its Helios rack-scale systems and directly compete with Nvidia's Rubin GPUs in AI training and inference.
AMD used its Advancing AI event to detail the Instinct MI455X, a new data center accelerator built on the CDNA 5 architecture. According to Tom's Hardware, the chip contains 320 billion transistors packaged using TSMC's advanced CoWoS-L technology, with four Accelerator Complex Dies (XCDs) stacked on each Fabric and Cache Die (FCD) via hybrid bonding. AMD is using TSMC's most advanced 2nm gate-all-around process for the XCDs, while the FCDs and I/O dies rely on TSMC N3P — a chiplet strategy meant to apply the densest process node only where it delivers the most benefit.
The MI455X keeps the same 256 Work Group Processor count as its predecessor, the MI355X, but AMD says per-WGP throughput has increased substantially to drive large generational performance gains, particularly in lower-precision formats like OCP MXFP8 and MXFP4 used heavily in AI inference. AMD also overhauled the memory hierarchy, replacing the large Infinity Cache with a smaller but higher-bandwidth shared L2 cache totaling 192MB, and doubling the local data store within each WGP. Most notably, the chip moves to HBM4 memory, offering 432GB of capacity and, per Tom's Hardware, 23.3 TB/s of bandwidth per GPU — figures AMD says exceed what Nvidia has shown for its Rubin GPU so far, which Tom's Hardware reports at 288GB and up to 22 TB/s in its initial configuration.
At rack scale, AMD's Helios architecture links 72 MI455X GPUs into a single coherent domain, matching the scale of Nvidia's NVL72 design used in the Blackwell and Rubin generations. Tom's Hardware notes this gives AMD roughly 50% more aggregate HBM capacity across a rack than Nvidia's Vera Rubin setup. Wccftech reports the MI455X delivers up to 40 PFLOPs of FP4 and 20 PFLOPs of FP8 compute, compared to figures it cites for Rubin of 50 PFLOPs FP4 and 17.5 PFLOPs FP8. AMD itself has cautioned that real-world performance will likely fall short of these theoretical peak numbers, with software optimization remaining an ongoing challenge.
The MI455X is part of AMD's broader Instinct MI400 series, which also includes the MI430X for HPC and sovereign AI workloads. AMD says the entire lineup runs on its open ROCm software stack and has already translated into customer commitments, including deals with Microsoft and Anthropic reported around the announcement.
What We Know
| Spec | Detail |
|---|---|
| Architecture | CDNA 5 |
| Transistors | 320 billion |
| Process Node | TSMC 2nm (XCDs) + TSMC N3P (FCDs/IO dies) |
| Memory | 432GB HBM4 |
| Memory Bandwidth | 23.3 TB/s per GPU (per Tom's Hardware; Wccftech cites 22.3 TB/s) |
| WGPs | 256 Work Group Processors |
| Cache | 192MB total L2 (96MB per Fabric and Cache Die) |
| Rack-Scale System | Helios architecture joining 72 MI455X GPUs into one coherent domain |
Frequently Asked Questions
What architecture does the AMD Instinct MI455X use?
The MI455X is built on AMD's CDNA 5 architecture, using TSMC's 2nm process for its Accelerator Complex Dies and TSMC N3P for its Fabric and Cache Dies and I/O dies.
How much memory does the MI455X have?
The MI455X features 432GB of HBM4 memory, with Tom's Hardware citing 23.3 TB/s of bandwidth per GPU.
How does the MI455X compare to Nvidia's Rubin GPU?
According to the sources, the MI455X offers higher HBM4 memory capacity than Rubin's initial 288GB configuration and, per Wccftech, delivers up to 40 PFLOPs of FP4 compute versus 50 PFLOPs cited for Rubin, though FP8 figures favor AMD's chip.
What is AMD's Helios architecture?
Helios is AMD's rack-scale architecture that joins 72 MI455X GPUs into a single coherent accelerator domain, designed to compete with Nvidia's NVL72 rack-scale system.
