NVIDIA Vera CPU enters the AI data-center race, signaling vertical integration against AMD/Intel
cnbc.com

NVIDIA Vera CPU enters the AI data-center race, signaling vertical integration against AMD/Intel

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRNVIDIA unveiled Vera, its first in-house server CPU designed for agentic AI, with deliveries to OpenAI, Anthropic, and SpaceX.

NVIDIA is now shipping its own data-center CPU, the Vera, marking a direct challenge to AMD and Intel in the server market. Designed from the ground up for agentic AI workloads, Vera prioritizes single-core speed and memory bandwidth over traditional core counts, aiming to keep expensive GPUs fully utilized in AI factories.

What happened

NVIDIA released new details about its Vera CPU, the first server processor the company has designed from the core rather than using an off-the-shelf Arm design. Vera chips were delivered to clients including OpenAI, Anthropic, and SpaceX in June. The chip is available in several configurations: as a standalone CPU, in a liquid-cooled rack of 256 Vera chips, in a two-chip server, and paired with NVIDIA GPUs in a system called Vera Rubin.

Vera consumes 250 to 450 watts and supports up to 1.5 terabytes of memory per chip. NVIDIA claims the chip delivers 50% better performance for AI agents than x86 CPUs, thanks to its Olympus core focused on single-core speed and low latency. The company said the whole server CPU market could eventually be worth $200 billion, while Wolfe Research estimated an average selling price of about $5,000 per Vera chip and projected 1.3 million shipments this year.

Why AI builders should care

Agentic AI workloads have shifted attention back to the CPU, which must feed data to GPUs and handle tasks like code compilation and tool execution. NVIDIA argues that traditional x86 CPUs, designed for high core counts, create bottlenecks in agent loops. Vera's focus on per-core speed and memory bandwidth is meant to let agents return to GPUs faster, keeping those expensive assets highly utilized.

For teams building AI agents or deploying inference at scale, Vera could change how data-center hardware is procured. If NVIDIA succeeds in selling full systems like Vera Rubin, cloud providers may offer integrated CPU-GPU instances optimized for agentic AI. This vertical integration could also lead to tighter software optimizations, similar to what Apple achieved with its own chips.

Practical implications

Vera is already in the hands of leading AI labs. OpenAI plans to deploy Vera chips in large quantities starting this quarter, and Oracle is listed as a partner. The chip is built on the NVIDIA MGX architecture, which integrates up to 256 Vera CPUs to run over 22,500 concurrent environments.

For builders, the key takeaway is that Vera is not a general-purpose server CPU. It is designed specifically for AI workloads, particularly agentic AI, reinforcement learning, and data processing. If you are deploying agents that require fast CPU response times, Vera-based systems could reduce latency and improve GPU utilization. However, pricing has not been publicly disclosed, and broad cloud availability remains uncertain.

Caveats

NVIDIA's performance claims are based on its own benchmarks, and WIRED noted that the tests appear to have used slightly older generations of competitor CPUs. Real-world performance against current AMD EPYC and Intel Xeon chips may differ. Adoption is still in the "early innings," as NVIDIA's product marketer noted, and the company has not listed major cloud providers beyond Oracle as partners. AMD and Intel hold deep relationships with hyperscalers, and AMD is reportedly gaining share in server CPUs. Vera's success will depend on whether cloud providers choose to adopt NVIDIA's full system stack over established x86 alternatives.

FAQs

Vera is NVIDIA's first in-house server CPU, designed from the core for agentic AI workloads. Unlike traditional x86 CPUs from AMD and Intel that focus on core count, Vera prioritizes single-core speed, memory bandwidth, and low latency. It uses a custom Arm-based Olympus microarchitecture and is built to keep GPUs highly utilized by reducing CPU bottlenecks in AI agent loops.

Sources

Latest Tech News