HomeTechNvidia Vera Rubin: Inside the agentic AI factory that rewrites the CPU...

Nvidia Vera Rubin: Inside the agentic AI factory that rewrites the CPU playbook

On the surface, this week’s Vera Rubin launch is another major platform moment for Nvidia Corp., as the company maintains a steady drumbeat of artificial intelligence infrastructure innovation.

Nvidia is positioning Vera Rubin as a full-stack system designed to improve performance per watt and reduce token costs, with production ramping across a broad global partner base, including cloud and AI infrastructure providers. Nvidia says the platform spans seven co-designed chips, integrates new networking and is already being deployed by partners such as CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure.

However, the bigger story isn’t that Nvidia launched another AI system, since that’s nothing new for the king of AI. Rather, it’s that the company is making an aggressive case that the AI era, especially the rise of agentic AI, requires rethinking the central processing unit, the network and the system architecture as a single, interdependent design problem rather than a pile of best-of-breed parts. That is why Vera (pictured) matters.

From cloud economics to agentic bottlenecks

For years, the CPU roadmap in the data center was largely shaped by cloud economics. Hyperscalers wanted more cores, higher throughput and lower costs, which rewarded chiplet-heavy designs optimized for scale-out efficiency. Nvidia’s analyst briefing, and another for reporters last week, framed that era as one in which core counts grew roughly nine time and single-thread performance gains doubled, leaving the market with processors well-suited for classic cloud workloads but less suited to the latency-sensitive, branch-heavy behavior of AI agents. As Nvidia’s Hannah Coutand explained in the briefing, “AI is asking for a new CPU,” because “agentic AI is putting CPU back on the critical path.”

That framing is important. In an agentic system, the graphics processing unit performs the reasoning, while the CPU is constantly in the loop, handling tool calls, code execution, queries, orchestration and data handling between reasoning steps. These workloads create continuous loops of reasoning, acting, observing and evaluating, which put simultaneous pressure on single-thread performance, memory bandwidth, latency and scaling efficiency. In other words, throwing more generic cores at the problem is not enough if the real bottleneck is the sequential work between GPU inference passes.

Inside Vera: A CPU designed for the AI loop

This is the central rationale behind Vera’s new CPU architecture. Rather than extending the conventional server CPU playbook, Nvidia built a custom Arm-based processor around its new Olympus core. Vera features 88 custom cores, 176 hardware threads, up to 1.2 terabits per second of LPDDR5X memory bandwidth, 164 megabytes of unified L3 cache, and up to 1.8 TB/s of coherent CPU-GPU bandwidth via NVLink-C2C.

Nvidia says the design target is what it calls the “max single-threaded CPU at scale,” meaning a processor that preserves strong per-core responsiveness while still scaling across highly concurrent agent workloads. The analyst briefing made the design intent unusually clear. Coutand stated that Nvidia “didn’t set out to go win CPUs,” but instead recognized that “the CPU was becoming a bottleneck in the AI factory” and designed a better processor to improve “AI factory economics.” That is classic Nvidia strategy: Start with the system bottleneck, then build the silicon needed to remove it.

Nvidia’s Ian Finder provided the technical details behind that claim. He described Olympus as a ground-up custom core featuring a 10-wide decode engine, aggressive reordering logic and a graph prefetcher tuned for pointer-chasing patterns common in compilers, graph structures and agent runtimes. His point was that CPU performance in this new era is less about chasing clock speeds and more about increasing the amount of useful work each cycle can do. That emphasis on instructions per cycle, branch handling and memory behavior is exactly what you would expect if the target is agent orchestration rather than old-school enterprise middleware.

Fabric and memory: Making data movement part of compute

Just as important, Vera is a reminder that in AI infrastructure, the network is no longer a peripheral technology but part of the compute architecture.

Inside the CPU, Nvidia’s second-generation Scalable Coherency Fabric serves as the data-movement backbone, linking cores, caches, LPDDR5X controllers, I/O and NVLink-C2C interfaces with multiterabyte-per-second bandwidth. Nvidia contrasts this monolithic fabric with chiplet-based designs that incur a “chiplet tax” in the form of higher latency and lower effective bandwidth as traffic crosses die boundaries. For agentic workloads, which are sensitive to loaded latency and cross-core data sharing, those differences translate directly into GPU utilization and end-to-end responsiveness.

The memory subsystem follows the same philosophy. By pairing LPDDR5X with an enterprise-ready module form factor, Vera aims to deliver high bandwidth per core and better bandwidth-per-watt than conventional DDR-based servers. In an AI factory with thousands of deployed servers, shaving tens of watts from the CPU-plus-memory envelope while increasing bandwidth frees more of the power budget for GPUs and high-speed networking.

Beyond the rack: The role of Spectrum-X

Outside the CPU, the platform extends to the rack and the cluster. Within AI infrastructure there is a significant distinction between scale-up and scale-out networking. NVLink connects GPUs within a rack, enabling them to act as a unified accelerator with all-to-all bandwidth and in-network compute. Spectrum-X Ethernet provides the scale-out fabric that ties those racks together across the AI factory.

Spectrum-X is more strategically important than many realize. In traditional enterprise infrastructure, Ethernet can be treated as a largely modular layer. In AI factories, the network directly affects token throughput, latency, utilization and ultimately economics. If mixture-of-experts models and agentic systems create much heavier east-west traffic and more distributed coordination, generic Ethernet becomes a tax on the entire system.

Nvidia’s answer is a purpose-built Ethernet stack: 102.4T Spectrum-6 switches, 1.6T ConnectX-9 SuperNICs, adaptive routing, congestion control, telemetry and open software, all tuned for RDMA and AI traffic patterns. The goal is to make Ethernet behave more like an AI-specific fabric while preserving operational familiarity, so that scale-out networking enhances AI factory performance rather than undermining it.

Extreme co-design as a competitive weapon

This brings us to Nvidia’s “extreme co-design.” The company says that Vera Rubin NVL72, the Vera CPU rack, BlueField-4 infrastructure processors, Spectrum-6 switching, and the rest of the platform were engineered as a single system rather than assembled from separate off-the-shelf products.

With agentic AI, infrastructure services such as networking, storage, telemetry, security and context handling are now part of the inference pipeline itself. That means CPUs, GPUs, DPUs and switches need to be tuned together to keep expensive accelerators fed and productive without burning host CPU cycles on infrastructure work.

Most semiconductor vendors can compete credibly in one layer of the stack; a few can reach two. NVIDIA now has meaningful assets across GPUs, CPUs, scale-up networking, scale-out networking, DPUs, interconnect software and system design. That breadth lets it optimize for delivered AI output — tokens per watt, cost per token and usable throughput — not just component specs.

Nvidia’s next share gain story: CPUs

The most interesting industry implication of Vera is that CPUs may become Nvidia’s next share-gain story. Nvidia is not trying to displace x86 across every general-purpose data center workload. It doesn’t need to. The company’s own sizing suggests a large incremental CPU opportunity tied specifically to AI-driven workloads and infrastructure patterns. If the CPU’s role in the AI factory is increasingly to orchestrate agents, feed GPUs, manage memory movement and support low-latency tool execution, then Nvidia can leverage its GPU dominance to pull its own CPU into the design.

This is the same playbook the company has used elsewhere: win the control point, then expand adjacencies. Because Nvidia already owns the strategic budget line in AI infrastructure through GPUs, it is uniquely positioned to define what the surrounding CPU, network and data processing unit should look like. Vera’s tight integration with NVLink-C2C, BlueField and Spectrum-X means buyers considering Rubin-class systems are not evaluating the CPU in isolation. They are evaluating a full AI factory architecture.

In the near term, that will matter most for AI clouds, hyperscalers and large model builders with agentic or reinforcement-learning-heavy workloads. Over time, though, the definition of a “good” data center CPU may shift more broadly. If Nvidia is right, the next important CPU category will not be the cheapest cloud workhorse or the highest-core-count generalist. It will be the processor that best removes friction from the AI loop.

Final thoughts

The fundamental tenet of my research has always been that rapid market share shifts occur when markets transition, and the CPU industry hasn’t seen a significant transition in a long time. But this is the AI era, and it’s seemingly redefining all industries.

Nvidia is making the case that the future CPU is no longer a standalone component decision. It is a systems decision, tightly coupled to GPUs, memory, networking and infrastructure processors — and Nvidia intends to own as much of that system as possible. This week, AMD is holding its own “Advancing AI” summit, and we should get a good look at how it plans to address the challenges Nvidia laid out above.

Zeus Kerravala is a principal analyst at ZK Research, a division of Kerravala Consulting. He wrote this article for SiliconANGLE.

Photo: Nvidia

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

https://siliconangle.com/aws-marketplace/

About SiliconANGLE Media

SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

 

Must Read

spot_img