infracalculus

AMD's swim upstream

AMD launched Helios at Advancing AI 2026. Having been a silicon vendor for decades, AMD is making its first attempt at owning rack-scale design. Helios is also AMD’s answer to NVIDIA’s Vera Rubin NVL72. The first rack-scale design from NVIDIA, based on Blackwell, was deployed in 1H 2025, while the first AMD Helios is expected to be deployed in 2H 2026. NVIDIA has been through this cycle thrice already; that is a lot of ground for AMD to cover.

The specifications of the latest generation designs from both NVIDIA and AMD are comparable. NVIDIA chose the Vera Rubin superchip (CPU-GPU package) for its NVL72, while AMD chose separate EPYC Venice (CPU) and Instinct MI455X (GPU) for its Helios.

AMD Helios and NVIDIA Vera Rubin NVL72, published rack specifications
ElementAMD HeliosNVIDIA Vera Rubin NVL72
GPU72 × MI455X72 × Rubin
CPU18 × EPYC Venice36 × Vera
GPU:CPU ratio4:12:1
HBM4 capacity31 TB20.7 TB
HBM4 bandwidth per GPU23.3 TB/s22 TB/s
Scale-up bandwidth260 TB/s, UALoE260 TB/s, NVLink 6
Scale-out bandwidth43 TB/s, Ultra Ethernet28.8 TB/s, InfiniBand or Ethernet

AMD claims that Helios is the most powerful rack-scale AI infrastructure. Though the leaderboard can only be confirmed once an independent benchmark has been performed with representative workloads.

Why AMD wanted to go from silicon to system?

1. Time-To-Market

NVIDIA has set the cadence of new architecture release every two years, and a memory upgrade in-between.

  • New architecture
  • Architecture update

Hover the chart, or focus it and use the arrow keys, to read values.

NVIDIA data-center GPU launches, 2017–2026Launch means NVIDIA's announcement date, not the later date when partner systems became generally available. Updates within an existing architecture are shown in blue.
Data for “NVIDIA data-center GPU launches, 2017–2026”
NVIDIA data-center GPU launches, 2017–2026
Announcement dateArchitectureGPURelease type
2017-05-10VoltaV100New architecture
2020-05-14AmpereA100New architecture
2022-03-22HopperH100New architecture
2023-11-13HopperH200Architecture update
2024-03-18BlackwellB200New architecture
2025-03-18Blackwell UltraB300Architecture update
2026-01-05RubinRubinNew architecture

Source: NVIDIA Newsroom launch announcements

AMD has released three architectures and updated one in the last 4 years. AMD is keeping up with the silicon release cadence. While AMD has had the APUs for consumer market since 2011, their first CPU-GPU package for AI servers hit the market closely after NVIDIA Superchip.

  • New architecture
  • Architecture update

Hover the chart, or focus it and use the arrow keys, to read values.

AMD datacenter GPU launches, 2020–2026Launch means AMD's announcement or published launch date, not the later date when partner systems became generally available. Updates within an existing architecture are shown in blue.
Data for “AMD datacenter GPU launches, 2020–2026”
AMD datacenter GPU launches, 2020–2026
Announcement dateArchitectureGPURelease type
2020-11-16CDNAMI100New architecture
2021-11-08CDNA 2MI250XNew architecture
2023-01-04CDNA 3MI300ANew architecture
2024-10-10CDNA 3MI325XArchitecture update
2025-06-12CDNA 4MI355XNew architecture
2026-07-23CDNA 5MI455XNew architecture

Source: AMD Newsroom launch announcements and AMD product specifications

Once the silicon is made available to the partners, it takes more than a year to build and deploy the platforms to datacenters. Each OEM, neocloud and hyperscaler has their own way of approaching AI infrastructure.

One thing to call out explicitly is that the below chart is not a fair representation of actual time-to-market for two reasons. First, the partners get early access to silicon before they were publicly announced. The public announcement does not equal Day-0. Second, some preferred partners get access to silicon sooner than the others. The actual date when each partner had their samples to start working on their designs is highly confidential. A partner or cloud provider having a longer time-to-deployment in the below charts does not necessarily mean they started at the same time as the one who went first to the market.

  • Hopper
  • Blackwell
  • Vera Rubin

Hover the chart, or focus it and use the arrow keys, to read values.

NVIDIA generation rollout: Hopper, Blackwell and Vera RubinDay 0 is the vendor's own headline announcement introducing the generation, in every series and for both vendors — the press release, not a cadence slide naming a codename years ahead. All six anchors are announcements made ahead of general availability, so the bars measure the same kind of interval and the two charts share one 0–780 day scale: a bar of a given length is the same number of days in either figure. Each anchor is dated and linked under “Bar zero for each group” in the data below. Hopper on March 22, 2022, announced with the H100; Blackwell on March 18, 2024, announced with B200 and the GB200 Superchip; Rubin on January 5, 2026, launched with six new chips and the Vera Rubin NVL72. Rows are grouped and labeled by generation, with color providing a second visual cue. Milestones are not equivalent to each other: they range from full production and private preview to cloud availability, rack validation and production-system delivery. The generation-opening bar is a manufacturing milestone — NVIDIA stating the part was in full production — which is not the same claim as the availability announcement that opens AMD's CDNA 3.
Data for “NVIDIA generation rollout: Hopper, Blackwell and Vera Rubin”
Bar zero for each group
GroupDateEvent
Hopper2022-03-22GTC 2022: Hopper architecture and the H100 announced — 80 billion transistors, TSMC 4N, first GPU with HBM3
Blackwell2024-03-18GTC 2024: Blackwell platform announced — B200 at 208 billion transistors, and the GB200 Grace Blackwell Superchip
Vera Rubin2026-01-05CES 2026: Rubin platform launched — six new chips, including the Vera CPU and Rubin GPU, and the Vera Rubin NVL72
NVIDIA generation rollout: Hopper, Blackwell and Vera Rubin
GenerationDateDays after first product previewProviderSystemMilestoneCategory
Hopper2022-09-20182NVIDIAH100 Tensor Core GPUNVIDIA announced H100 was in full productionFull production
Hopper2023-03-21364CoreWeaveHGX H100 instancesGeneral availability announcedAI cloud GA
Hopper2023-05-10414LambdaH100 PCIe instancesOn-demand general availability announcedAI cloud GA
Hopper2023-07-26491AWSEC2 P5 instancesGeneral availability announcedHyperscaler GA
Hopper2023-08-07503Microsoft AzureND H100 v5 VMsGeneral availability announcedHyperscaler GA
Hopper2023-08-29525Google CloudA3 VMsSeptember GA announced; no exact GA day givenHyperscaler GA
Hopper2023-09-19546Oracle CloudBM.GPU.H100.8 instancesGeneral availability announcedHyperscaler GA
Blackwell2024-11-19246Microsoft AzureND GB200 v6 · GB200 NVL72Limited private preview for select partnersPrivate preview
Blackwell2024-11-20247NVIDIABlackwell platformNVIDIA reported Blackwell was in full productionFull production
Blackwell2025-02-04323CoreWeaveGB200 NVL72 instancesGeneral availability announcedProvider GA
Blackwell2025-03-18365Microsoft AzureND GB200 v6 · GB200 NVL72General availability announcedProvider GA
Vera Rubin2026-06-01147CoreWeaveVera Rubin NVL72Industry-first cloud-provider bring-up and rack-level validation announcedCompleted · Provider validation
Vera Rubin2026-08-21228MicrosoftProduction Vera Rubin systemsFirst production systems arrived at Microsoft datacenters; bring-up not confirmedCompleted · Production-system delivery

Source: NVIDIA, CoreWeave, Lambda, AWS, Microsoft Azure, Google Cloud and Oracle announcements through August 24, 2026

AMD’s is no exception either. In the case of Helios, AMD statement that Meta’s validation is promising, shows that partner enablement pre-dates public announcement. Another reason Meta was the first to validate because Helios is based on Meta’s Open Rack Wide contribution to OCP.

  • CDNA 3
  • CDNA 4
  • CDNA 5

Hover the chart, or focus it and use the arrow keys, to read values.

AMD generation rollout: CDNA 3, CDNA 4 and CDNA 5Day 0 is the vendor's own headline announcement introducing the generation, in every series and for both vendors — the press release, not a cadence slide naming a codename years ahead. All six anchors are announcements made ahead of general availability, so the bars measure the same kind of interval and the two charts share one 0–780 day scale: a bar of a given length is the same number of days in either figure. Each anchor is dated and linked under “Bar zero for each group” in the data below. CDNA 3 on January 4, 2023, where MI300 was previewed as an integrated data-center CPU and GPU; CDNA 4 on June 12, 2025, where the MI350 Series launched; CDNA 5 on July 23, 2026, where the MI400 Series launched led by Helios. AMD's two most recent anchors are launches while CDNA 3's is a preview, so CDNA 3's bars start from an earlier point in its cycle than the other two. Rows are grouped and labeled by generation, with color as a second cue. Milestones are not equivalent: they range from product launch to cloud GA, a benchmark result and a deployment target. Three carry qualifications — the El Capitan entry confirms installation, not delivery of MI300A compute blades; TensorWave's availability entry is a provider claim rather than independently documented GA; and the MLPerf figure is aggregate multinode throughput, not a per-GPU result. AMD described Helios rack-scale solutions, not every MI455X and Venice component, as in production, and OpenAI's Q4 2026 entry is a target plotted at the start of the quarter, not confirmed public-cloud GA.
Data for “AMD generation rollout: CDNA 3, CDNA 4 and CDNA 5”
Bar zero for each group
GroupDateEvent
CDNA 32023-01-04CES 2023 keynote: MI300 previewed as an integrated data-center CPU and GPU — CDNA 3 with Zen 4 cores, 128GB of HBM3, 146 billion transistors
CDNA 42025-06-12Advancing AI 2025: MI350 Series launched, with MI400 and Helios previewed
CDNA 52026-07-23Advancing AI 2026: MI400 Series launched with HBM4, led by Helios rackscale solutions
AMD generation rollout: CDNA 3, CDNA 4 and CDNA 5
GenerationDate or windowDays after first product previewGPU or systemProvider or organizationMilestoneCategory
CDNA 32023-12-06336MI300A and MI300XAMDAMD announced MI300X and MI300A availabilityProduct availability
CDNA 32024-04-23475MI300XTensorWaveTensorWave said its MI300X cloud infrastructure was available at scaleAI cloud availability claimed
CDNA 32024-05-21503MI300XMicrosoft AzureAzure ND MI300X v5 generally availableCloud GA
CDNA 32024-11-18684MI300ALLNL, HPE and AMDEl Capitan ranked first on TOP500 at 1.742 exaflopsVerified exascale milestone
CDNA 42025-06-120MI350X and MI355XAMDAMD launched the MI350 Series and said partner systems were rolling outProduct launch
CDNA 42025-09-0989MI355XVultrVultr announced bare-metal and Cloud GPU availabilityCloud GA
CDNA 42025-10-14124MI355XOracle Cloud InfrastructureBM.GPU.MI355X.8 bare-metal instances, eight MI355X per node, generally availableCloud GA
CDNA 42026-02-19252MI350XDigitalOceanGPU Droplets generally available in the ATL1 region; MI355X said to follow the next quarterCloud GA
CDNA 5July 23, 20260MI455X and HeliosAMDAMD launched MI400 and said Helios rack-scale solutions were in productionCompleted · AMD
CDNA 5July 23, 20260HeliosMetaAMD said Meta had begun testing and validating workloads on Helios racksCompleted · Meta

Source: AMD, Lawrence Livermore, Microsoft Azure, TOP500, TensorWave, Vultr, MLCommons and Advancing AI announcements through August 24, 2026

As GPU architecture gets more powerful each generation, they demand higher power, cooling, and networking. Customers and partners are figuring out the power and coolant distribution, scale-up and scale-out domains in their own silos. NVIDIA started providing the rack-scale design arguably to shorten this development lifecycle.

NVIDIA has been gradually defining the higher order systems. Over the past decade, their strategic moves had been both lateral and vertical. Their footprint is on many strata of AI infrastructure now.

  1. 2016InterconnectNVLink 1.0A proprietary GPU-to-GPU link, defining the scale-up domain inside a node.
  2. 2016ServerDGX-1A turnkey integrated server: eight GPUs, sold assembled rather than as parts.
  3. 2017BaseboardHGX-1A GPU baseboard published with Microsoft as a reference architecture, so hyperscalers could build the server around it.
  4. 2020Cluster fabricMellanox acquisitionInfiniBand switching brought the scale-out domain in-house — reach across cabinets rather than a larger cabinet.
  5. 2023PackageGrace Hopper SuperchipCPU and GPU on one module in mass production, taking the level below the board.
  6. 2025RackGB200 NVL72Generally available: 18 compute trays around 9 NVLink switch trays, sold as one machine.

What NVIDIA sold, and at what level of the stack. Glyphs are not to scale.

NVIDIA product and newsroom announcements, 2016–2025

There is no official, public data to back up the claim that NVL72 or Helios improve time-to-market. However, there are public statements from partners, like this one from a Microsoft executive, that seem to corroborate the argument that reference architectures from silicon vendors help accelerate the deployment.

2. Rack is the new compute unit

For frontier AI, server is increasingly not the compute unit, mechanically, thermally, electrically, and digitally.

The industry has been disaggregating a compute unit to optimize for performance. A compute server at the turn of 2010 had CPU, GPU, memory, storage, network, cooling, power and battery, connected with a bunch of cables. The components that constituted a compute unit are no longer contained within a server. A few components have moved closer and a few higher. Now, AI infrastructure has the CPU-GPU-HBM as one integrated module in a compute tray with some storage, and the network, cooling, power and battery as separate trays of their own.

Where each component of a 2010 compute server lives in an AI rack
ComponentIn a 2010 serverIn an AI rack
CPUOne or two sockets on the motherboardPart of the CPU-GPU-HBM-LPDDR module
GPUAn add-in card in a PCIe slotPart of the CPU-GPU-HBM-LPDDR module
MemoryDIMMs in banks beside the socketsHBM and LPDDR on the same package as CPU-GPU
StorageDrive bays behind the front bezelStays local, on the compute tray
NetworkA NIC in a slot, cabled to a top-of-rack switchSwitch trays of its own
PowerA PSU pair inside the chassisA power shelf feeding the whole rack
CoolingA fan wall behind the drivesAn in-rack CDU and a liquid loop
BatteryA backup unit on the RAID controllerA backup shelf holding the rack up

Generic 2U server layout, and a generic AI rack; tray counts are illustrative

AI racks are drawing north of 100kW power and Google introduced 400VDC architecture to support 1MW racks last year. Air-only cooling is no longer sufficient. Direct Liquid Cooling is moving to either an in-rack or in-row CDU. Similarly, power delivery is moving to either an in-rack shelf or an in-row sidecar.

What stands in one AI row: six compute racks and three sidecars
CabinetCountWhat it holds
Power sidecar1Rectifiers and the busbar
Battery sidecar1Ride-through to manage micro power drops
Compute rack6Eight compute trays and two network switch trays each
Cooling sidecar1The CDU and the liquid loop that runs the length of the row

Evloving disaggregation of compute

A generic in-row sidecar layout; cabinet and tray counts are illustrative

Frontier models are pushing trillions of parameters and training requires a cluster of thousands of GPUs. Training a 405B parameter model took Meta more than 16,000 NVIDIA H100 GPUs. A model from a frontier lab cannot be hosted for inference on one server without quantization or distillation. For example, Moonshot recommends hosting its Kimi K3 on a supernode with 64 or more accelerators for inference.

Both scale-up and scale-out are critical for frontier AI training and inference. Silicon vendors cannot leave the scale-up and scale-out domain to system integrators as an afterthought. The UALoE/NVLink and InfiniBand/Ultra Ethernet are integral parts of the fabric.

Owning the rack-scale design becomes imperative to harness the full potential of the silicon.

3. Money, up for the grabs

Jensen Huang coined “5-layer cake”. Silicon companies are investing in and enabling every layer.

Many hyperscalers have their own custom silicon: Microsoft’s Maia, AWS’s Trainium and Inferentia, Google’s TPU, and Meta’s MTIA. As hyperscalers are moving down the stack, silicon vendors are countering by going up the stack. While the hyperscalers have in-house team to design custom silicon, the traditional silicon vendors are building racks and investing in neoclouds. The circular financing model of silicon vendors investing in their customers warrants a separate deep-dive.

Both AMD and NVIDIA are releasing their own open models. CUDA and ROCm are offered to enable the model layer. The software alone needs a separate discussion. Both companies have strategically invested in model and application layer companies. Their role in top layers is so far mostly enablement and driving demand, but they are monetizing the infrastructure layer.

The five layers of the AI stack, from energy at the bottom to applications at the top
LayerWhat it isDrawn as
5. ApplicationsWhere the economic value is createddrug, legal, robotics, automobile
4. ModelsTrained on language, biology, chemistry, physics, finance, medicine and the physical worldmodel
3. InfrastructureLand, power delivery, cooling, construction and the rack-scale hardware that makes tens of thousands of processors one machineland, power, cooling, construction, rack
2. ChipsProcessors that turn energy into computation at scalecpu, gpu, memory
1. EnergyPower generated in real time, the binding constraint on how much intelligence the system can producegrid, solar, wind, nuclear

Layers after Jensen Huang’s “AI five-layer cake”, NVIDIA blog

How far up the stack AMD sells

Everything that follows is easier to read against a single question: at which level of integration does the vendor hand the product over to somebody else?

  1. L1Parts manufacturingIndividual components and raw materialsComponents
  2. L2Piece-parts subassemblyManufactured parts combined into subassemblies
  3. L3Metals and plastics integrationStructural parts integrated into a chassis
  4. L4Kit assemblyChassis, PSU, cables and backplanes supplied as a kit
  5. L5Enclosure assembly and I/O testingComplete enclosure, cabling and I/O validation
  6. L6Motherboard integrationMotherboard installed in the chassis and powered on
  7. L7Add-on card integrationGPU, networking and expansion cards installed and tested
  8. L8Storage integrationStorage devices installed and validated
  9. L9CPU and memory integrationProcessors and DIMMs installed and validated
  10. L10Full server assembly and testingBootable server with system testing and software integrationServer
  11. L11Rack-level assemblyNodes, switches, cabling and PDUs integrated as a working rackHelios
  12. L12Cluster-level assemblyCross-rack cabling, cluster software, validation and optimisation
  • Eight lidded AMD Instinct MI440X modules mounted in two rows on a single universal baseboard, with power connectors along its lower edge.L7MI440X UBBAdd-on card integration
  • An AMD Pensando network card: a green PCIe board carrying a finned heatsink over its two cage connectors and the DPU package at its centre.L7Pensando DPUAdd-on card integration
  • A lidded AMD EPYC processor seated in its socket on a server motherboard, its gold contact edge and the socket retention hardware visible alongside.L9EPYC 9006CPU and memory integration
  • An open Helios compute tray seen from above: four AMD Instinct accelerator modules with copper cold plates along the rear, a finned AMD processor heatsink toward the front, and cabling running between them.L10Helios compute trayFull server assembly and testing
  • An AMD Helios rack with its side panel removed, showing five bays of stacked compute and switch trays with looms of yellow and blue cabling down the left side.L11Helios rackRack-level assembly

Manufacturing levels

Manufacturing levels after Glenn K. Lockwood and AMAX Engineering.

Compete and Collaborate

The role of silicon vendors now overlaps with that of ODMs and OEMs. First, NVIDIA offers the branded DGX Vera Rubin NVL72 to arguably capture some of the enterprise market while working alongside OEMs, hyperscalers and neoclouds on NVL72. Helios is a reference design which partners need to deliver and hence falls one step short of being an off-the-shelf appliance. Whether AMD will eventually offer its own first-party racks is an open question. Second, NVIDIA and AMD increase distribution through their respective specialized neocloud partners, CoreWeave and TensorWave. CoreWeave is also the first to bring H100 and GB200 to the public cloud, and validate Vera Rubin racks.

AMD and NVIDIA claim that their rack-scale designs support pre-training, post-training, and inference workloads. But each type of workload demands special consideration. Neither NVIDIA nor AMD has multiple rack-scale SKUs to cater to different use cases. That leaves room for OEMs and hyperscalers to design their customized systems for a variety of reasons, primarily, optimization and standardization.

AI infrastructure for pre-training workloads faces sustained near-peak power and cooling demands, and needs robust scale-out connectivity. Pre-training workloads are usually distributed across servers, racks, and clusters. There is a heavy penalty to pay for component failures when forced to restore from the last checkpoint. AI infrastructure for inference workloads needs higher HBM density for model weights and KV cache, some workloads benefit from a higher CPU-to-GPU ratio and most frontier-model inference requires low-latency scale-up bandwidth. Most vertically integrated hyperscale datacenter would require modified adaptation of silicon vendors’ rack-scale design because building blocks of hyperscalers are being meticulously standardized for manageability, modularity, reusability, and interoperability.

Long but scenic road ahead

AMD is a credible candidate for diversification and single-vendor risk mitigation, as part of the informal “NVIDIA Plus One” strategy adopted by many in the industry. The innovation ecosystem benefits and the costs decline with healthy competition. AMD has so far announced OpenAI, Meta, Microsoft, Anthropic, HUMAIN, Vultr, CirraScale, TensorWave, River, OneQode, Aligned, amp PBC, and Oracle as customers, early adopters, or partners. The specifications and announcements indicate that Helios can be an alternative to NVIDIA’s Vera Rubin NVL72. AMD has become a strong challenger to NVIDIA not just at the GPU level, but at rack scale. It may be a while before AMD catches up with NVIDIA on software and developer ecosystem.