Tech

Intel and NVIDIA pivot to CPU-centric architecture for agentic AI inference

Intel reports server Xeon demand outstripping supply while NVIDIA launches the Vera CPU, signalling a structural shift from GPU-heavy to CPU-balanced AI deployments.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
Major chipmakers and cloud providers are reshaping data centre infrastructure as agentic workloads drive demand for central processing units

Intel and NVIDIA are fundamentally altering their artificial intelligence infrastructure strategies, shifting focus from graphics processing units to central processing units to support agentic AI workloads. Intel reports that the CPU-to-GPU ratio in data centres is moving from a traditional 1:8 to 1:1, or up to 4:1 in agentic deployments, with server Xeon demand currently outstripping supply. This transition reflects a broader industry realignment where CPUs are increasingly tasked with orchestration and tool execution, while GPUs remain specialised for token generation.

NVIDIA has launched the Vera CPU, described as purpose-built for agentic AI, with first deliveries scheduled for 2026 to clients including Anthropic, OpenAI, SpaceX, and Oracle Cloud Infrastructure. The Vera Rubin NVL72 rack architecture utilises a 1:2 CPU-to-GPU ratio, optimising CPUs for orchestration and tool execution while GPUs handle token generation. This marks a significant departure from previous architectures, illustrating how the industry is redefining the division of labour between processor types.

Arm has also entered the fray with the launch of the AGI CPU, marking its return to production silicon after 35 years. This move underscores the growing importance of CPU-centric designs in the agentic era. Concurrently, Red Hat has released an open-source framework to evaluate CPU inference performance using vLLM, providing tools to standardise testing and deployment in this evolving landscape.

The shift is driven by the complex nature of agentic AI, which requires CPUs to handle multistep reasoning, tool calls, and orchestration across small, specialised models. Intel’s CEO Lip-Bu Tan noted that for reinforcement learning, orchestration, and agents, the CPU is a much better fit. Arm’s CEO Rene Haas estimates that the orchestration demands of agentic workloads could drive a fourfold increase in CPU cores required per gigawatt of capacity, rising from 30 million to 120 million CPU cores per GW.

Morgan Stanley estimates the agentic CPU shift represents $32.5–$60B in incremental CPU market growth by 2030. While GPUs remain essential for high-concurrency production serving of large models, the perimeter of their dominance is contracting. The industry is moving towards a hybrid model where CPUs manage the iterative, logic-heavy aspects of AI, creating new opportunities for chipmakers and infrastructure providers alike.

Continue reading

More from Tech

Read next: The Walrus warns of collapsing digital memory as AI erodes search reliability
Read next: Open-source tool claims 97 per cent token savings for AI agents
Read next: Valvoline Unveils August 2026 Promotional Offers for Service and Retail Buyers