NOVARIFT
NVIDIA Enters the Agentic Era
April 28, 2026·Technology·9 MIN READ

NVIDIA Enters the Agentic Era

NVIDIA spent two conferences this year proving it isn't a chip company anymore. Between GTC San Jose in March and GTC Taipei in June, Jensen Huang laid out a full stack, from silicon to the software running on your desktop, built around one bet: that AI agents, not chatbots, are what the next decade of computing runs on.

NVIDIA held two of the most consequential product events in its history this year, and neither one was really about a new GPU. At GTC San Jose in March and GTC Taipei in June, Jensen Huang laid out a case that the company has quietly become something bigger than a chipmaker: it's now trying to own every layer of the stack that AI agents run on, from the silicon up through the software that decides what those agents are allowed to touch.

The framing Huang used at both events was deliberate. Generative AI proved AI could be useful. Reasoning models made it capable. Agents, he argued, are what make it work autonomously, continuously, at scale. That's a bigger claim than a hardware refresh, and the announcements that followed backed it up with real products, real dollar figures, and at least one open-source project NVIDIA didn't build but is now racing to control.

What Actually Got Announced, and Where

GTC San Jose ran March 16 to 19 at the SAP Center, drawing more than 30,000 attendees for Huang's roughly two-and-a-half-hour keynote. That's where NVIDIA introduced the Vera Rubin platform, the Groq 3 LPU, NemoClaw, and RTX Spark, alongside a headline projection that purchase orders across Blackwell and Vera Rubin systems would reach $1 trillion through 2027.

Advertisement

GTC Taipei followed on June 1 at the Taipei Music Center, where Huang confirmed Vera Rubin had entered full production and was ramping toward broader shipments by fall 2026. He called it "the most ambitious endeavor in the history of our company," noting that all 40,000 of NVIDIA's engineers had touched some part of the project. The two events worked as a pair: San Jose delivered the vision and the specs, Taipei delivered proof that the hardware was real and shipping.

The Vera Rubin Platform, Piece by Piece

Vera Rubin isn't a single chip. It's a full rack-scale system built around seven co-designed components, and the specifics matter more than the marketing name.

The Rubin R100 GPU is the centerpiece: built on TSMC's 3nm process with 336 billion transistors, up from Blackwell's 208 billion, and carrying 288 GB of HBM4 memory at up to 22 TB/s of bandwidth. That's nearly triple Blackwell's 8 TB/s, and it matters because memory bandwidth, not raw compute, is what determines how fast a model can actually serve long, context-heavy requests. FP4 inference throughput lands at 50 petaflops per GPU, roughly 2.5 to 5 times Blackwell's output depending on configuration.

Pair two R100s with a Vera CPU, NVIDIA's second custom ARM chip carrying 88 Olympus cores, and you get a Vera Rubin Superchip. Scale that to a full NVL72 rack and you're looking at 20.7 TB of pooled HBM4 memory, 1.6 PB/s of aggregate bandwidth, and NVLink 6 interconnect running at roughly double the previous generation's speed. NVIDIA's own figures put the cost reduction at up to 10 times cheaper per inference token compared to Blackwell, and that number comes almost entirely from the memory bandwidth jump rather than raw FLOPs.

The Bottleneck Nobody Was Watching Two Years Ago

For most of the last three years, AI progress got measured by training metrics: parameter counts, training time, how much compute it took to build a bigger model. Those numbers mattered when the hard part was making capable models exist at all.

Advertisement

Agentic AI breaks that measurement stick. A single request to an autonomous agent doesn't trigger one inference pass, it triggers dozens or hundreds, chained together as the agent plans a task, calls tools, checks its own work, and revises course. That kind of workload is memory-bound in a way single-shot chatbot queries never were, which is exactly why NVIDIA built the Vera CPU with a coherent memory fabric linking it directly to the GPU over NVLink-C2C at 1.8 TB/s. Data handling agent state and context no longer has to cross the slower PCIe boundary between orchestration and compute.

NVIDIA Dynamo, the software layer coordinating all of this, claims up to 7x throughput per GPU on existing Blackwell hardware by splitting the prefill and decode stages of inference and routing each to whichever hardware handles it best. That detail matters for anyone running current-generation chips: the efficiency gains aren't locked behind a hardware upgrade, they run through existing deployments via TensorRT-LLM, vLLM, and SGLang.

OpenClaw's Wild Year, and Why NVIDIA Wanted In

The most unpredictable part of this whole story didn't come from NVIDIA at all. It came from an Austrian developer named Peter Steinberger, who released an autonomous AI agent under the name Warelay in November 2025. After a trademark dispute forced a rename to Moltbot and then, three days later, to OpenClaw (Steinberger later said Moltbot "never quite rolled off the tongue"), the project exploded. By January 2026 it had crossed 100,000 GitHub stars. By March it had passed 250,000, overtaking React to become the most-starred software project on GitHub in roughly 60 days, a faster climb than Linux managed in its early years.

Steinberger joined OpenAI in February 2026, and OpenClaw's governance shifted toward foundation status while the code stayed open. That popularity created a real problem for any company that wanted to actually deploy it: an autonomous agent with file system access, network permissions, and no fixed job description is exactly the kind of thing compliance teams veto on sight. There was no audit trail, no policy layer, nothing to show a regulator if something went wrong.

NemoClaw is NVIDIA's answer, announced at GTC San Jose and built with Steinberger's involvement rather than around him. It installs on top of OpenClaw with a single command, adding NVIDIA's OpenShell runtime to sandbox agents at the process level and enforce network and data-access policies. It isn't exclusive to NVIDIA's own models either: NemoClaw runs coding agents from OpenAI and Anthropic as easily as it runs NVIDIA's own Nemotron models, which can run locally on hardware ranging from RTX laptops to DGX Station. Adobe, Atlassian, Salesforce, and ServiceNow were named as launch partners. Huang's comparison was blunt: "Mac and Windows are the operating systems for the personal computer. OpenClaw is the operating system for personal AI."

The Other Six Chips: Groq, RTX Spark, and Filling in the Gaps

Vera Rubin's seventh component, the Groq 3 LPX, came out of a $20 billion deal, NVIDIA's largest acquisition ever, closed in December 2025 to bring most of Groq's team and technology in-house. Where the Rubin GPU relies on HBM4 for bandwidth, the Groq chip uses on-chip SRAM instead, delivering extremely low latency for the decode-heavy stretches of agentic workloads where a GPU's raw throughput matters less than how fast it can respond token by token. A rack of 256 Groq LPUs targets roughly 300 tokens per second on agentic tasks, a different job than the R100 is built for, and NVIDIA is positioning the two as complementary rather than competing.

RTX Spark pushes the same philosophy down to personal hardware. Built with MediaTek and running Windows, it packs 6,144 CUDA cores into a chip aimed at slim laptops and desktops, delivering roughly 1 petaflop of AI performance for agents that run locally instead of round-tripping to the cloud for every task. Adobe has already rearchitected Photoshop and Premiere around it.

The Trillion-Dollar Number and How Wall Street Actually Reacted

Huang's headline figure, $1 trillion in cumulative orders across Blackwell and Vera Rubin through 2027, is double what NVIDIA had projected for the same window just a year earlier, when the company was guiding toward roughly $500 billion. Quarterly revenue has been tracking around $78 billion, up 77% year over year.

What's notable is how little the market moved on any of it. Analysts described the reaction as muted, not because the numbers were bad, but because expectations had already climbed high enough that even a doubled projection didn't clear the bar investors had priced in. That's arguably the more interesting data point: when a company can double its own revenue guidance and barely move its stock, it says something about how large AI infrastructure spending has already become in the market's mental model.

Who's Actually Getting the Chips

Demand for Vera Rubin is running well ahead of what NVIDIA can currently produce. TSMC's 3nm capacity, shared with Apple and AMD, caps initial Rubin output at an estimated 200,000 to 300,000 units in 2026, and HBM4 supply from SK Hynix and Samsung is an additional constraint, with Huang publicly pushing memory makers to increase production as recently as June.

The eight confirmed cloud partners receiving initial shipments are AWS, Azure, Google Cloud, Oracle, CoreWeave, Lambda, Nebius, and Nscale. AWS alone has committed to deploying over a million GPUs alongside Groq LPUs, while Microsoft says it already has hundreds of thousands of liquid-cooled Grace Blackwell systems running across Azure. Historically, 60 to 70 percent of first-year supply on a new NVIDIA architecture goes to hyperscalers before anyone smaller gets meaningful allocation, which means most businesses outside that tier will be renting Rubin-era compute through a cloud provider rather than buying it directly for a while yet.

What This Actually Means If You're Not NVIDIA

Strip away the keynote theater and the honest takeaway is this: agentic AI isn't a future bet NVIDIA is hedging on, it's the assumption baked into every layer of what the company shipped this year, from the memory architecture up through the security software wrapped around a viral open-source project it didn't even build. The infrastructure underneath that bet is expensive, supply-constrained, and concentrated among a small number of cloud providers who get first access.

Advertisement

For anyone building on top of this rather than inside it, that concentration is worth watching closely. NVIDIA's parallel push into open model development through its Nemotron Coalition, which now includes groups like Mira Murati's Thinking Machines Lab, suggests the company understands that owning the hardware layer means little if the software layer building on top of it belongs entirely to someone else. Huang has said for years that NVIDIA isn't a GPU company, it's a data center company. This year is the clearest evidence yet that he means it literally: the goal isn't to sell you a chip, it's to be the platform every agent, every workload, and increasingly every developer runs on top of by default.

Share
novarift.org/blog/nvidia-enters-the-agentic-era

Leave a Comment

Comments (0)

No comments yet. Be the first to share your thoughts.

Advertisement
Back to all articles

Related