Tokenomics: Welcome to the future of AI economics

Nvidia Financial Analyst Q&A Highlights

At the recent NVIDIA Shareholder/Analyst Call on March 17, 2026, CEO Jensen Huang and CFO Colette Kress didn’t just discuss hardware updates—they offered a profound look into the fundamental roadmap of global computing.

During this deeply technical and highly strategic Q&A, Huang declared that we have officially entered AI’s third inflection point: the era of autonomous “agentic systems”. To support this massive evolutionary leap, NVIDIA revealed a staggering $1 trillion-plus in firm demand and visibility for their Blackwell and Rubin architectures stretching through 2027.

But perhaps the most transformative takeaway wasn’t the silicon itself; it was a complete shift in how the industry must value compute power. Introducing the concept of “Tokenomics,” Huang explained how data centers have evolved from hosting simple software tools into massive manufacturing plants dedicated to producing AI tokens.

Main Messages from Jensen

  • The Age of Agentic Systems: Jensen Huang stated that AI has reached its third inflection point, moving from generative AI to reasoning, and now to agentic systems. These systems operate autonomously to execute tasks, transforming computers from simple tools into manufacturing equipment that produces “tokens”.
Diagram illustrating a new computing platform titled 'Agents', featuring various components such as multi-modal prompts, memory, tools, and sub-agents. Central to the diagram is an 'Agent' character.

Renaissance from SaaS to Agent-as-a-Service

Diagram illustrating the evolution of enterprise IT from SaaS to Agent-as-a-Service, highlighting components like enterprise IT expense, software, and AI providers. Features illustrations of workers and data centers.

The $1 Trillion Visibility: NVIDIA has firm demand, purchase orders, and visibility for $1 trillion plus in its Blackwell and Rubin architectures through the end of 2027. This projection explicitly excludes newer offerings like stand-alone CPUs, Groq, and Rubin Ultra. Later he explained that including all that could be a upside of 50% to $1,5 trillion, not counting what else can be sold in the next 21 months.

A presentation slide featuring a bar graph illustrating strong growth in AI inference through NVIDIA's full-stack expansion, alongside a pie chart showing market distribution between hyperscalers and cloud AI natives.

Market Segmentation (60/40 Split): Historically, 60% of NVIDIA’s business is driven by hyperscalers, while the remaining 40% comes from enterprise, regional clouds, and industrial on-premise implementations. Successfully addressing the 40% segment is impossible without NVIDIA’s end-to-end full software stack.

Infographic announcing NVIDIA NemoClaw Reference OpenClaw, detailing the NVIDIA Agent Toolkit for building specialized agents, including elements like multi-modal prompts, cuDF, cuVS, and various tools.
  • OpenClaw Strategy: Huang emphasized that OpenClaw acts as the fundamental operating system for the new personal AI computer. He stated that every software company in the world will need to develop an OpenClaw strategy.

Ecosystem ROI and Market Expansion

  • Hyperscaler Revenue Upside: Huang noted that private AI companies are experiencing unprecedented growth, increasing revenues by billions weekly. He anticipates the $2 trillion IT software industry will grow to an $8 trillion industry by integrating open models, OpenAI, and Anthropic, essentially turning IT companies into resellers of tokens.
  • Evolution of Customer Mix: Huang expects digital AI (hyperscalers) and physical AI (industrial/enterprise) segments to grow at similar rates until a major physical AI inflection occurs. Eventually, the physical AI segment could outgrow the digital side because the real-world economy is significantly larger, but both would be significantly larger.
A presentation slide showing performance metrics for NVIDIA's GB300 NVL72 and H200 NVL8, highlighting tokens per watt and cost comparisons against competitors.

Token Pricing Trends: The fundamental cost of tokens will continuously decrease with new architectures, while the “smartness” per token increases. Different segments of the market will demand different tiers of token performance and pricing, ranging from free search tiers to premium agentic code generation.

Graph illustrating throughput in TPS/MW for different models, showcasing performance efficiency improvements from Rubin NVL72 and Blackwell NVL72 to Rubin + LPX, with annotations indicating performance multipliers.

The Dawn of Tokenomics

Jensen Huang heavily emphasized that the world needs to understand the new economics of AI, which he calls “Tokenomics”. In this new era, computers are no longer just tools; they are manufacturing equipment, and the product they manufacture is the “token”.

  • The Engineer’s Token Budget: In the future, companies won’t just give an engineer a laptop; they will give them a laptop and a “token budget”. Huang argued that it is illogical to pay an engineer $2,000 a day and not have them spend $1,000 a day on tokens to manage a fleet of autonomous agents doing their work.
  • Price of Computer vs. Cost of Token: Huang clarified a major misconception: the upfront price of an AI computer is only marginally related to the cost of the token it produces. By purchasing the most expensive, technologically advanced computer, a company can simultaneously produce tokens at the lowest possible cost because the machine generates them at such incredible rates.
  • The Ultimate Metric (Tokens/Second/Watt): When evaluating an AI system, the only metric that truly matters is tokens per second divided by watts. Because data centers are strictly limited by their power capacity (e.g., a 1-gigawatt data center cannot magically use 2 gigawatts), evaluating throughput without normalizing for power consumption is deceptive.
  • Market Segmentation: Just like cars or smartphones, AI tokens are becoming a new commodity with distinct tiers.
    • Free Tier: Typically utilized for basic search functions.
    • Premium/Extreme Tiers: Utilized for highly complex agentic tasks, like code generation, where high-paid engineers will happily pay premiums for massive throughput on large models.
  • The IT Industry Pivot: The $2 trillion IT industry historically relied on licensing software. Moving forward, 100% of these companies will transform their business models to integrate models like OpenAI, Anthropic, and open models, essentially becoming resellers that rent and generate tokens.

Groq Integration

The Role of Groq: Groq represents an extreme, low-latency architecture that NVIDIA will fuse with Vera Rubin. It will be utilized for the final, bandwidth-intensive stage of autoregressive models, targeting roughly 25% of the most premium inference workloads. Integrating Groq will increase a customer’s GPU compute spend by approximately 25%.

A presentation slide showing a graph comparing inference performance and efficiency of different models (Rubin NVL72, Blackwell NVL72, and Hopper) measured in TPS/MW. The graph highlights significant performance improvements and company results, with the speaker discussing the data.
A presentation slide showcasing a bar graph illustrating NVIDIA's Vera Rubin and LPX revenue potential, highlighting revenue opportunities in billions across different tiers: Free, Medium, High, Premium, and Ultra, with a graphical indication of unlocking a $300 billion revenue opportunity.

The Physics of Networking: Copper vs. Co-Packaged Optics (CPO)

Copper vs. Co-Packaged Optics: NVIDIA intends to scale with copper connections for as long as possible due to its reliability and ease of manufacturing. However, the physical limit of copper is around one meter. Future architectures like NVLink 144 (Rubin Ultra) will offer both copper and Co-Packaged Optics (CPO) options, while the larger NVLink 1152 will rely entirely on CPO.

  • The “Breathing Air” Philosophy: Huang emphasized that the industry should scale with copper interconnects for as long as physically possible. Copper is cheap, highly reliable, easy to manufacture, and humanity has a long history of mastering it. He likened using copper to breathing air—you should use it until you absolutely run out of it.
  • The Physical Limit of Copper: The fundamental physics of copper dictate that its maximum effective length for these extreme bandwidths is roughly one meter (plus or minus).
  • The Transition Timeline: * NVLink 144 (Vera Rubin Ultra): For the next generation, NVIDIA’s backplane will still support copper, offering customers two options: pure copper or copper combined with Co-Packaged Optics (CPO).
    • NVLink 1152 (Future Architectures): Two years out, scaling to an NVLink 1152 architecture will exceed copper’s physical reach. At this scale, the interconnects will be 100% CPO.
  • Copper Demand Will Still Grow: Even as massive scale-up fabrics transition to optics, NVIDIA’s total consumption of copper connectors will continue to rise. Copper will still be heavily utilized for Ethernet scale-up, storage arrays, and standard rack connections across their diverse MGX server lineup.

Capital Allocation and Supply Chain

  • Cash Flow Priorities: CFO Colette Kress outlined that strong free cash flows will primarily fund growth and support the supply chain, which may include prepayments. After addressing existing commitments in the first half of the year, NVIDIA plans to return capital to shareholders. The company anticipates allocating approximately 50% of its free cash flow toward stock repurchases and dividends.
  • Supply Constraints: Huang acknowledged that the world is somewhat supply-constrained across various components. However, he emphasized that NVIDIA’s massive scale allows them to manage harmony across multiple suppliers to ensure they meet their $1 trillion-plus demand forecast.

Competitive Moat and AI Modeling

  • Sustaining Margins: Huang defended NVIDIA’s pricing power by explaining that token generation fundamentally changes compute economics. Customers are willing to pay a premium for systems that deliver the highest “tokens per second per watt,” resulting in the most efficient token production.
  • Organizational Structure: NVIDIA’s ability to maintain a rapid annual release cadence relies on owning the entire technology stack, including chips, networking, storage, and software compilers. This full-stack ownership allows them to align all components to a single schedule without relying on third-party timelines.
  • State Space Models and Training Dynamics: NVIDIA’s architecture is versatile enough to support all AI models, including hybrid state space models like Nemotron 3, which handle extremely long context windows efficiently.
  • Regarding compute demand, Huang noted that post-training (skill acquisition) is vastly more compute-intensive than base pretraining. Ultimately, his goal is for 99% of future compute to be dedicated to inference, as that is where economic value is generated.

Grace Blackwell vs. Vera Rubin

To support this massive demand for token generation, NVIDIA has radically evolved its data center architecture. While Grace Blackwell was the king of standard inference, Vera Rubin is designed for the vastly more complex world of agentic systems.

Grace Blackwell (The Inference King):

  • Primary Focus: The Grace Blackwell architecture, specifically the NVL72 rack, was explicitly built to process and run large language models (LLMs) and dominate standard AI inference.
  • System Scale: It essentially operated as a single-rack solution focused on maximizing standard token throughput.

Vera Rubin (The Agentic Factory):

  • Workload Expansion: While LLMs still make up 90% of the workload, agentic systems (like Claude Code or Codex) require much more. Vera Rubin is designed to handle massive KV cache memory, structured and unstructured data (cuDF, cuVS), and CPU-reliant tool use like web browsing.
  • The MGX Architecture (5 Racks): To prevent data centers from becoming a “Frankenstein” mix of mismatched cooling and power systems, Vera Rubin expands from a single rack to a harmonized 5-rack MGX architecture.
  • Unified Infrastructure: Every component in the Vera Rubin cluster—whether it is computing, east-west storage, or networking—is 100% liquid-cooled and utilizes the exact same power delivery and cooling systems.
  • Fusing with Groq for Extreme Latency: To push the boundaries of token generation even further, NVIDIA is combining Vera Rubin’s high-throughput GPUs with Groq’s low-latency SRAM chips. Groq will process the final, hyper-bandwidth-intensive stage of auto-regressive models. While adding Groq increases a customer’s GPU compute spend by 25%, it allows them to maximize throughput and command much higher pricing in the “Best and Extreme” token tiers.

Ultimately, Huang noted that because NVIDIA delivers such massive generational leaps in value (tokens per second per watt), it is instantaneously smarter for customers to buy Vera Rubin at a higher price than to continue buying Grace Blackwell at a lower price.

The Blurring Lines: AI Training vs. Inference

Historically, the market viewed training and inference as two distinct, isolated phases. Huang argued that this view is completely outdated, predicting that the line between learning (training) and applying wisdom (inference) will soon disappear.

  • Pretraining (AI High School): This initial phase provides the model with basic vocabulary, grammar, and foundational reasoning. While today’s pretraining relies mostly on internet data, future pretraining will be driven largely by synthetic data.
  • Post-Training (Skill Acquisition): Once the foundation is set, post-training teaches the model specific skills like math, coding, and tool use through reinforcement learning and verifiable feedback. Huang noted that the computing intensity for this phase is staggering—potentially 1 million to 1 billion times greater than pretraining.
  • Continuous Learning: Moving forward, models will continuously learn and be fine-tuned to generalize and memorize information on a per-person basis.
  • The 99% Inference Goal: Huang stated his ultimate hope is that 99% of the world’s computing power will eventually be dedicated to inference.

“Nobody pays you for learning. Nobody pays for training, you pay for training. I want the world to be able to use these tokens for valuable outcome…” — Jensen Huang

Because inference translates generated tokens directly into economic value for industries like healthcare, manufacturing, and finance, NVIDIA went “all in” on their inference stack to capitalize on this massive future market.

Thank you,

CoreValue Research

Never miss the next one

Free articles and market notes by email, the moment they publish.

No spam. Unsubscribe in one click.

CoreValue Research Team's avatar

CoreValue Research Team

Related Articles

Free

Broadcom: The Custom Sillicon New Age

Broadcom FY 2Q26 and FY 3Q26 For those who came to this post maybe worried that we would say anything bad about Nvidia market share...

July 20, 2026 · CoreValue Research Team
Free

TSMC accelerated revenue growth August 2026!

Growth in USD increased from 31,8% in July to 44,4% in August, on track to deliver +40% in 2026 TSMC reported August revenue today in...

September 10, 2026 · CoreValue Research Team
Free

TSMC July-26 Revenues

Hold your breath! “Only” +32% yoy USD growth! TSMC released July sales this morning at NT$ 467.6 bi, a record month, +44.7% year on year...

August 10, 2026 · CoreValue Research Team

Discover more from CoreValue Research

Subscribe now to keep reading and get access to the full archive.

Continue reading