Today in AI
Thursday, July 23, 2026
Today in AI: Nvidia Vera Rubin vs. AMD Helios — Inference Cost War and Grid Capacity Crisis — July 23, 2026
4
links
-
1blogs.nvidia.comNvidia's Vera Rubin NVL72 rack-scale system, now in full production across Taiwan server makers, delivers 10x lower inference cost per token and 10x higher throughput per megawatt versus Blackwell, with Google Cloud A5X instances already live. Developers targeting inference workloads on GCP now have a direct cost-per-token benchmark; builders on competing platforms need to recalibrate economics.
-
2ir.amd.comAMD unveiled Helios rack-scale platform at Advancing AI 2026 with $5.25M pricing, delivering 30% more inference tokens per dollar than competition, with immediate deployments expected from late 2026 at Microsoft Azure and via Cerebras Cloud. This is a credible second-source inference option for hyperscalers, directly challenging Nvidia's inference margin and giving enterprises negotiation leverage.
-
3www.cnbc.comAMD committed up to $5B in compute power to Anthropic, with the first gigawatt of capacity deploying in H1 2027, cementing AMD as a primary infrastructure vendor for frontier model inference beyond Nvidia. This capacity commitment signals Anthropic's confidence in AMD hardware maturity and represents a structural shift in infrastructure sourcing for large labs.
-
4www.tomshardware.comBloombergNEF raised its US data center power projection 83% to 194 GW by 2035 (one-fifth of total US electricity), with grid access now the primary constraint on new facility deployment. Infrastructure practitioners must factor power availability and grid interconnection timelines into capacity planning; companies without secured power access face multi-year delays.