AMD's swim upstream
AMD launched Helios at Advancing AI 2026. Having been a silicon vendor for decades, AMD is making its first attempt at owning rack-scale design. Helios is also AMD’s answer to NVIDIA’s Vera Rubin NVL72. The first rack-scale design from NVIDIA, based on Blackwell, was deployed in 1H 2025, while the first AMD Helios is expected to be deployed in 2H 2026. NVIDIA has been through this cycle thrice already; that is a lot of ground for AMD to cover.
The specifications of the latest generation designs from both NVIDIA and AMD are comparable. NVIDIA chose the Vera Rubin superchip (CPU-GPU package) for its NVL72, while AMD chose separate EPYC Venice (CPU) and Instinct MI455X (GPU) for its Helios.
| Element | AMD Helios | NVIDIA Vera Rubin NVL72 |
|---|---|---|
| GPU | 72 × MI455X | 72 × Rubin |
| CPU | 18 × EPYC Venice | 36 × Vera |
| GPU:CPU ratio | 4:1 | 2:1 |
| HBM4 capacity | 31 TB | 20.7 TB |
| HBM4 bandwidth per GPU | 23.3 TB/s | 22 TB/s |
| Scale-up bandwidth | 260 TB/s, UALoE | 260 TB/s, NVLink 6 |
| Scale-out bandwidth | 43 TB/s, Ultra Ethernet | 28.8 TB/s, InfiniBand or Ethernet |
AMD claims that Helios is the most powerful rack-scale AI infrastructure. Though the leaderboard can only be confirmed once an independent benchmark has been performed with representative workloads.
Why AMD wanted to go from silicon to system?
1. Time-To-Market
NVIDIA has set the cadence of new architecture release every two years, and a memory upgrade in-between.
- New architecture
- Architecture update
Hover the chart, or focus it and use the arrow keys, to read values.
Data for “NVIDIA data-center GPU launches, 2017–2026”
| Announcement date | Architecture | GPU | Release type |
|---|---|---|---|
| 2017-05-10 | Volta | V100 | New architecture |
| 2020-05-14 | Ampere | A100 | New architecture |
| 2022-03-22 | Hopper | H100 | New architecture |
| 2023-11-13 | Hopper | H200 | Architecture update |
| 2024-03-18 | Blackwell | B200 | New architecture |
| 2025-03-18 | Blackwell Ultra | B300 | Architecture update |
| 2026-01-05 | Rubin | Rubin | New architecture |
Source: NVIDIA Newsroom launch announcements
AMD has released three architectures and updated one in the last 4 years. AMD is keeping up with the silicon release cadence. While AMD has had the APUs for consumer market since 2011, their first CPU-GPU package for AI servers hit the market closely after NVIDIA Superchip.
- New architecture
- Architecture update
Hover the chart, or focus it and use the arrow keys, to read values.
Data for “AMD datacenter GPU launches, 2020–2026”
| Announcement date | Architecture | GPU | Release type |
|---|---|---|---|
| 2020-11-16 | CDNA | MI100 | New architecture |
| 2021-11-08 | CDNA 2 | MI250X | New architecture |
| 2023-01-04 | CDNA 3 | MI300A | New architecture |
| 2024-10-10 | CDNA 3 | MI325X | Architecture update |
| 2025-06-12 | CDNA 4 | MI355X | New architecture |
| 2026-07-23 | CDNA 5 | MI455X | New architecture |
Source: AMD Newsroom launch announcements and AMD product specifications
Once the silicon is made available to the partners, it takes more than a year to build and deploy the platforms to datacenters. Each OEM, neocloud and hyperscaler has their own way of approaching AI infrastructure.
One thing to call out explicitly is that the below chart is not a fair representation of actual time-to-market for two reasons. First, the partners get early access to silicon before they were publicly announced. The public announcement does not equal Day-0. Second, some preferred partners get access to silicon sooner than the others. The actual date when each partner had their samples to start working on their designs is highly confidential. A partner or cloud provider having a longer time-to-deployment in the below charts does not necessarily mean they started at the same time as the one who went first to the market.
- Hopper
- Blackwell
- Vera Rubin
Hover the chart, or focus it and use the arrow keys, to read values.
Data for “NVIDIA generation rollout: Hopper, Blackwell and Vera Rubin”
| Group | Date | Event |
|---|---|---|
| Hopper | 2022-03-22 | GTC 2022: Hopper architecture and the H100 announced — 80 billion transistors, TSMC 4N, first GPU with HBM3 |
| Blackwell | 2024-03-18 | GTC 2024: Blackwell platform announced — B200 at 208 billion transistors, and the GB200 Grace Blackwell Superchip |
| Vera Rubin | 2026-01-05 | CES 2026: Rubin platform launched — six new chips, including the Vera CPU and Rubin GPU, and the Vera Rubin NVL72 |
| Generation | Date | Days after first product preview | Provider | System | Milestone | Category |
|---|---|---|---|---|---|---|
| Hopper | 2022-09-20 | 182 | NVIDIA | H100 Tensor Core GPU | NVIDIA announced H100 was in full production | Full production |
| Hopper | 2023-03-21 | 364 | CoreWeave | HGX H100 instances | General availability announced | AI cloud GA |
| Hopper | 2023-05-10 | 414 | Lambda | H100 PCIe instances | On-demand general availability announced | AI cloud GA |
| Hopper | 2023-07-26 | 491 | AWS | EC2 P5 instances | General availability announced | Hyperscaler GA |
| Hopper | 2023-08-07 | 503 | Microsoft Azure | ND H100 v5 VMs | General availability announced | Hyperscaler GA |
| Hopper | 2023-08-29 | 525 | Google Cloud | A3 VMs | September GA announced; no exact GA day given | Hyperscaler GA |
| Hopper | 2023-09-19 | 546 | Oracle Cloud | BM.GPU.H100.8 instances | General availability announced | Hyperscaler GA |
| Blackwell | 2024-11-19 | 246 | Microsoft Azure | ND GB200 v6 · GB200 NVL72 | Limited private preview for select partners | Private preview |
| Blackwell | 2024-11-20 | 247 | NVIDIA | Blackwell platform | NVIDIA reported Blackwell was in full production | Full production |
| Blackwell | 2025-02-04 | 323 | CoreWeave | GB200 NVL72 instances | General availability announced | Provider GA |
| Blackwell | 2025-03-18 | 365 | Microsoft Azure | ND GB200 v6 · GB200 NVL72 | General availability announced | Provider GA |
| Vera Rubin | 2026-06-01 | 147 | CoreWeave | Vera Rubin NVL72 | Industry-first cloud-provider bring-up and rack-level validation announced | Completed · Provider validation |
| Vera Rubin | 2026-08-21 | 228 | Microsoft | Production Vera Rubin systems | First production systems arrived at Microsoft datacenters; bring-up not confirmed | Completed · Production-system delivery |
Source: NVIDIA, CoreWeave, Lambda, AWS, Microsoft Azure, Google Cloud and Oracle announcements through August 24, 2026
AMD’s is no exception either. In the case of Helios, AMD statement that Meta’s validation is promising, shows that partner enablement pre-dates public announcement. Another reason Meta was the first to validate because Helios is based on Meta’s Open Rack Wide contribution to OCP.
- CDNA 3
- CDNA 4
- CDNA 5
Hover the chart, or focus it and use the arrow keys, to read values.
Data for “AMD generation rollout: CDNA 3, CDNA 4 and CDNA 5”
| Group | Date | Event |
|---|---|---|
| CDNA 3 | 2023-01-04 | CES 2023 keynote: MI300 previewed as an integrated data-center CPU and GPU — CDNA 3 with Zen 4 cores, 128GB of HBM3, 146 billion transistors |
| CDNA 4 | 2025-06-12 | Advancing AI 2025: MI350 Series launched, with MI400 and Helios previewed |
| CDNA 5 | 2026-07-23 | Advancing AI 2026: MI400 Series launched with HBM4, led by Helios rackscale solutions |
| Generation | Date or window | Days after first product preview | GPU or system | Provider or organization | Milestone | Category |
|---|---|---|---|---|---|---|
| CDNA 3 | 2023-12-06 | 336 | MI300A and MI300X | AMD | AMD announced MI300X and MI300A availability | Product availability |
| CDNA 3 | 2024-04-23 | 475 | MI300X | TensorWave | TensorWave said its MI300X cloud infrastructure was available at scale | AI cloud availability claimed |
| CDNA 3 | 2024-05-21 | 503 | MI300X | Microsoft Azure | Azure ND MI300X v5 generally available | Cloud GA |
| CDNA 3 | 2024-11-18 | 684 | MI300A | LLNL, HPE and AMD | El Capitan ranked first on TOP500 at 1.742 exaflops | Verified exascale milestone |
| CDNA 4 | 2025-06-12 | 0 | MI350X and MI355X | AMD | AMD launched the MI350 Series and said partner systems were rolling out | Product launch |
| CDNA 4 | 2025-09-09 | 89 | MI355X | Vultr | Vultr announced bare-metal and Cloud GPU availability | Cloud GA |
| CDNA 4 | 2025-10-14 | 124 | MI355X | Oracle Cloud Infrastructure | BM.GPU.MI355X.8 bare-metal instances, eight MI355X per node, generally available | Cloud GA |
| CDNA 4 | 2026-02-19 | 252 | MI350X | DigitalOcean | GPU Droplets generally available in the ATL1 region; MI355X said to follow the next quarter | Cloud GA |
| CDNA 5 | July 23, 2026 | 0 | MI455X and Helios | AMD | AMD launched MI400 and said Helios rack-scale solutions were in production | Completed · AMD |
| CDNA 5 | July 23, 2026 | 0 | Helios | Meta | AMD said Meta had begun testing and validating workloads on Helios racks | Completed · Meta |
Source: AMD, Lawrence Livermore, Microsoft Azure, TOP500, TensorWave, Vultr, MLCommons and Advancing AI announcements through August 24, 2026
As GPU architecture gets more powerful each generation, they demand higher power, cooling, and networking. Customers and partners are figuring out the power and coolant distribution, scale-up and scale-out domains in their own silos. NVIDIA started providing the rack-scale design arguably to shorten this development lifecycle.
NVIDIA has been gradually defining the higher order systems. Over the past decade, their strategic moves had been both lateral and vertical. Their footprint is on many strata of AI infrastructure now.
- 2016InterconnectNVLink 1.0A proprietary GPU-to-GPU link, defining the scale-up domain inside a node.
- 2016ServerDGX-1A turnkey integrated server: eight GPUs, sold assembled rather than as parts.
- 2017BaseboardHGX-1A GPU baseboard published with Microsoft as a reference architecture, so hyperscalers could build the server around it.
- 2020Cluster fabricMellanox acquisitionInfiniBand switching brought the scale-out domain in-house — reach across cabinets rather than a larger cabinet.
- 2023PackageGrace Hopper SuperchipCPU and GPU on one module in mass production, taking the level below the board.
- 2025RackGB200 NVL72Generally available: 18 compute trays around 9 NVLink switch trays, sold as one machine.
What NVIDIA sold, and at what level of the stack. Glyphs are not to scale.
NVIDIA product and newsroom announcements, 2016–2025
There is no official, public data to back up the claim that NVL72 or Helios improve time-to-market. However, there are public statements from partners, like this one from a Microsoft executive, that seem to corroborate the argument that reference architectures from silicon vendors help accelerate the deployment.
2. Rack is the new compute unit
For frontier AI, server is increasingly not the compute unit, mechanically, thermally, electrically, and digitally.
The industry has been disaggregating a compute unit to optimize for performance. A compute server at the turn of 2010 had CPU, GPU, memory, storage, network, cooling, power and battery, connected with a bunch of cables. The components that constituted a compute unit are no longer contained within a server. A few components have moved closer and a few higher. Now, AI infrastructure has the CPU-GPU-HBM as one integrated module in a compute tray with some storage, and the network, cooling, power and battery as separate trays of their own.
| Component | In a 2010 server | In an AI rack | |
|---|---|---|---|
| CPU | One or two sockets on the motherboard | Part of the CPU-GPU-HBM-LPDDR module | |
| GPU | An add-in card in a PCIe slot | Part of the CPU-GPU-HBM-LPDDR module | |
| Memory | DIMMs in banks beside the sockets | HBM and LPDDR on the same package as CPU-GPU | |
| Storage | Drive bays behind the front bezel | Stays local, on the compute tray | |
| Network | A NIC in a slot, cabled to a top-of-rack switch | Switch trays of its own | |
| Power | A PSU pair inside the chassis | A power shelf feeding the whole rack | |
| Cooling | A fan wall behind the drives | An in-rack CDU and a liquid loop | |
| Battery | A backup unit on the RAID controller | A backup shelf holding the rack up | |
Generic 2U server layout, and a generic AI rack; tray counts are illustrative
AI racks are drawing north of 100kW power and Google introduced 400VDC architecture to support 1MW racks last year. Air-only cooling is no longer sufficient. Direct Liquid Cooling is moving to either an in-rack or in-row CDU. Similarly, power delivery is moving to either an in-rack shelf or an in-row sidecar.
| Cabinet | Count | What it holds |
|---|---|---|
| Power sidecar | 1 | Rectifiers and the busbar |
| Battery sidecar | 1 | Ride-through to manage micro power drops |
| Compute rack | 6 | Eight compute trays and two network switch trays each |
| Cooling sidecar | 1 | The CDU and the liquid loop that runs the length of the row |
Evloving disaggregation of compute
A generic in-row sidecar layout; cabinet and tray counts are illustrative
Frontier models are pushing trillions of parameters and training requires a cluster of thousands of GPUs. Training a 405B parameter model took Meta more than 16,000 NVIDIA H100 GPUs. A model from a frontier lab cannot be hosted for inference on one server without quantization or distillation. For example, Moonshot recommends hosting its Kimi K3 on a supernode with 64 or more accelerators for inference.
Both scale-up and scale-out are critical for frontier AI training and inference. Silicon vendors cannot leave the scale-up and scale-out domain to system integrators as an afterthought. The UALoE/NVLink and InfiniBand/Ultra Ethernet are integral parts of the fabric.
Owning the rack-scale design becomes imperative to harness the full potential of the silicon.
3. Money, up for the grabs
Jensen Huang coined “5-layer cake”. Silicon companies are investing in and enabling every layer.
Many hyperscalers have their own custom silicon: Microsoft’s Maia, AWS’s Trainium and Inferentia, Google’s TPU, and Meta’s MTIA. As hyperscalers are moving down the stack, silicon vendors are countering by going up the stack. While the hyperscalers have in-house team to design custom silicon, the traditional silicon vendors are building racks and investing in neoclouds. The circular financing model of silicon vendors investing in their customers warrants a separate deep-dive.
Both AMD and NVIDIA are releasing their own open models. CUDA and ROCm are offered to enable the model layer. The software alone needs a separate discussion. Both companies have strategically invested in model and application layer companies. Their role in top layers is so far mostly enablement and driving demand, but they are monetizing the infrastructure layer.
| Layer | What it is | Drawn as |
|---|---|---|
| 5. Applications | Where the economic value is created | drug, legal, robotics, automobile |
| 4. Models | Trained on language, biology, chemistry, physics, finance, medicine and the physical world | model |
| 3. Infrastructure | Land, power delivery, cooling, construction and the rack-scale hardware that makes tens of thousands of processors one machine | land, power, cooling, construction, rack |
| 2. Chips | Processors that turn energy into computation at scale | cpu, gpu, memory |
| 1. Energy | Power generated in real time, the binding constraint on how much intelligence the system can produce | grid, solar, wind, nuclear |
Layers after Jensen Huang’s “AI five-layer cake”, NVIDIA blog
How far up the stack AMD sells
Everything that follows is easier to read against a single question: at which level of integration does the vendor hand the product over to somebody else?
- L1Parts manufacturingIndividual components and raw materialsComponents
- L2Piece-parts subassemblyManufactured parts combined into subassemblies
- L3Metals and plastics integrationStructural parts integrated into a chassis
- L4Kit assemblyChassis, PSU, cables and backplanes supplied as a kit
- L5Enclosure assembly and I/O testingComplete enclosure, cabling and I/O validation
- L6Motherboard integrationMotherboard installed in the chassis and powered on
- L7Add-on card integrationGPU, networking and expansion cards installed and tested
- L8Storage integrationStorage devices installed and validated
- L9CPU and memory integrationProcessors and DIMMs installed and validated
- L10Full server assembly and testingBootable server with system testing and software integrationServer
- L11Rack-level assemblyNodes, switches, cabling and PDUs integrated as a working rackHelios
- L12Cluster-level assemblyCross-rack cabling, cluster software, validation and optimisation
L7MI440X UBBAdd-on card integration
L7Pensando DPUAdd-on card integration
L9EPYC 9006CPU and memory integration
L10Helios compute trayFull server assembly and testing
L11Helios rackRack-level assembly
Manufacturing levels
Manufacturing levels after Glenn K. Lockwood and AMAX Engineering.
Compete and Collaborate
The role of silicon vendors now overlaps with that of ODMs and OEMs. First, NVIDIA offers the branded DGX Vera Rubin NVL72 to arguably capture some of the enterprise market while working alongside OEMs, hyperscalers and neoclouds on NVL72. Helios is a reference design which partners need to deliver and hence falls one step short of being an off-the-shelf appliance. Whether AMD will eventually offer its own first-party racks is an open question. Second, NVIDIA and AMD increase distribution through their respective specialized neocloud partners, CoreWeave and TensorWave. CoreWeave is also the first to bring H100 and GB200 to the public cloud, and validate Vera Rubin racks.
AMD and NVIDIA claim that their rack-scale designs support pre-training, post-training, and inference workloads. But each type of workload demands special consideration. Neither NVIDIA nor AMD has multiple rack-scale SKUs to cater to different use cases. That leaves room for OEMs and hyperscalers to design their customized systems for a variety of reasons, primarily, optimization and standardization.
AI infrastructure for pre-training workloads faces sustained near-peak power and cooling demands, and needs robust scale-out connectivity. Pre-training workloads are usually distributed across servers, racks, and clusters. There is a heavy penalty to pay for component failures when forced to restore from the last checkpoint. AI infrastructure for inference workloads needs higher HBM density for model weights and KV cache, some workloads benefit from a higher CPU-to-GPU ratio and most frontier-model inference requires low-latency scale-up bandwidth. Most vertically integrated hyperscale datacenter would require modified adaptation of silicon vendors’ rack-scale design because building blocks of hyperscalers are being meticulously standardized for manageability, modularity, reusability, and interoperability.
Long but scenic road ahead
AMD is a credible candidate for diversification and single-vendor risk mitigation, as part of the informal “NVIDIA Plus One” strategy adopted by many in the industry. The innovation ecosystem benefits and the costs decline with healthy competition. AMD has so far announced OpenAI, Meta, Microsoft, Anthropic, HUMAIN, Vultr, CirraScale, TensorWave, River, OneQode, Aligned, amp PBC, and Oracle as customers, early adopters, or partners. The specifications and announcements indicate that Helios can be an alternative to NVIDIA’s Vera Rubin NVL72. AMD has become a strong challenger to NVIDIA not just at the GPU level, but at rack scale. It may be a while before AMD catches up with NVIDIA on software and developer ecosystem.