2026-09-01AITao
AI Demand Is Running Ahead of Compute Supply: Heavy Users, Data Centers, and Multi-Model Systems
Gavin Baker and David George use heavy-user demand, data-center paybacks, agent workflows, multi-model architecture, and NVIDIA's ecosystem to explain why AI demand may expand faster than compute supply and what must hold for that shortage thesis to work.
Contents9 sections
- A bubble can still form while supply stays tight
- Demand has reached only a small group of heavy users
- Training and inference bind revenue to capital spending
- Agents move token demand toward continuous work
- Short supply could create a compute divide first
- Multi-model systems compete for the intelligence layer
- NVIDIA's advantage extends into financing and supply coordination
- Orbital compute is a long-range stress test
- What must hold for the shortage thesis to work
Original video: Why AI Demand Is Outrunning Compute Supply
a16z · 2026-08-31 · 1 hour 14 minutes 26 seconds
Speakers: David George, general partner at Andreessen Horowitz · Gavin Baker, managing partner and CIO of Atreides Management
Caption note: The video provides only YouTube's English auto-generated captions. Names, products, and companies were repaired against the video metadata, repeated context, and primary sources. Payback periods, user counts, supply shares, and future schedules remain attributed to the speakers.
Additional primary sources: a16z: the speakers' earlier discussion of the AI bubble and infrastructure · xAI: Introducing Grok Bot
AI infrastructure debates often compress the problem into a binary choice: bubble or shortage.
David George and Gavin Baker offer a more complicated view. Every major technology cycle can accumulate overvaluation and excess construction while near-term physical supply still fails to meet demand. Both can happen, just at different points in the cycle.
They focus on the speed difference between two curves. One is the rate at which AI spreads from a small group of heavy users to a much wider population. The other is the rate at which power, chips, data centers, financing, and permits turn into usable compute.
If demand keeps moving faster, AI is more likely to encounter scarcity first, with a conventional glut arriving later.
This is a conditional investment thesis. The interview supplies estimates, examples, and an analytical framework, not a complete market dataset that independently proves the conclusion.
A bubble can still form while supply stays tight
Baker begins by accepting the historical pattern of bubbles. Automobiles, railroads, radio, personal computers, and the internet all attracted capital ahead of proven demand. Rising valuations encouraged construction, and construction eventually went too far.
The current AI cycle has one unusual feature. In the speakers' estimates, some infrastructure projects can pay back quickly enough to keep attracting capital.
Baker illustrates the point with Nebius. He assumes roughly $50 billion of construction cost per gigawatt, with customer prepayments covering 50% to 60% and remaining capacity available for the spot market. Under those assumptions, he calculates a payback period of about 9 to 10 months.
That is not a uniform audited return published by Nebius. Changes in cost definitions, prepayments, utilization, pricing, or financing would change the result. The illustration instead explains why capital continues chasing compute. If a project appears able to repay itself in about a year, suppliers have little reason to slow down voluntarily.
The funding source also changes the cycle. A debt-funded build needs cash quickly and exposes a timing error early. A large technology company can fund construction from operating cash flow, then place mature assets into financing structures backed by institutions such as Blackstone, KKR, and Apollo. That structure can absorb more short-term pressure.
Bubble risk therefore remains. Physical bottlenecks and fast paybacks can delay when excess capacity becomes visible, while giving frontier labs, clouds, open models, applications, and NVIDIA room to grow at the same time.
Demand has reached only a small group of heavy users
George estimates that no more than roughly 30,000,000 heavy AI users currently produce substantial paid value. Baker thinks the figure may be materially lower, perhaps below 10,000,000.
They compare that population with about 1,500,000,000 knowledge workers. No user database in the source package validates the exact ratio, but the orders of magnitude frame the question. Today's revenue and compute use may come from a small fraction of the potential population.
Usage inside a16z portfolio companies also follows a strong power law. George says the heaviest engineers can consume 10 to 100 times as many tokens as the median engineer. A traditional company using AI reasonably well may spend about 1% of human compensation on tokens. An AI-native company may reach the high single digits, with some above 10%.
These figures are portfolio observations from the interview, not industry-wide benchmarks. They reveal the central demand uncertainty. Depth of use varies enormously, and diffusion has a long way to go.
If more employees approach the working style of today's heavy users, efficiency gains from better models and chips may be consumed by more tasks and longer runtimes. If diffusion slows, supply pressure eases and data-center returns move closer to ordinary infrastructure economics.
Training and inference bind revenue to capital spending
A frontier lab's revenue also depends on how it allocates compute.
Baker uses a hypothetical 10-gigawatt company to show the relationship. If 8 gigawatts serve inference and each gigawatt produces $60 billion of annual revenue, the run rate reaches $480 billion. If a research breakthrough causes the company to reverse the allocation and leave only 2 gigawatts for inference, revenue could fall to $120 billion.
Those numbers illustrate a mechanism rather than forecast any real company's finances. The mechanism matters because a model lab can redirect compute from paid inference into training, exchanging current revenue for the possibility of a stronger next generation.
Traditional internet companies usually account for serving revenue and long-horizon research as separate activities. Frontier labs move the same chips among training, post-training, internal research, and customer inference. Revenue, capability progress, and capital spending therefore compete for one scarce resource.
Baker expects these companies to generate substantial operating cash flow while continuing to buy GPUs and other accelerators, leaving little reason to optimize for free cash flow soon. Public markets will have to judge which checkpoint a lab releases, where it prices along the capability curve, and how much capacity it reserves for the future.
Agents move token demand toward continuous work
The interview's most concrete demand example comes from Atreides itself. Baker says the firm's internal token consumption rose 100 times from March through August.
The growth extends beyond chat. He used Claude Code and Codex to create podcast, Substack, and X summarizers, then asked Grok Bot to combine what other agents had learned that day and recommend next actions. Tools that once waited for a manual prompt are becoming systems that can keep observing, organizing, and advancing work.
xAI's official introduction describes Grok Bot as a persistent agent with its own cloud computer that can complete multi-step work across applications. That source confirms the operating pattern. Atreides' 100-fold increase remains one firm's self-report and cannot establish a market-wide growth rate.
The demand model changes with the product. A chat system consumes tokens when a user asks a question. A persistent agent also reads new information, maintains state, checks results, proposes actions, and waits for approval. One employee can keep several workflows running, so token use is no longer limited by typing speed.
When agent value moves from better answers to continuously completed work, compute demand grows with task count, runtime, and verification frequency.
Short supply could create a compute divide first
George expects usable capacity to remain tight through 2028 even after planned projects, with political and permitting resistance adding delays. Baker is consequently more concerned about severe undersupply.
This is a forecast, and the source package does not include a project-by-project capacity ledger. The potential consequence is concrete. If frontier models continue creating more value than they cost, scarcity could raise frontier-token prices or direct the highest quotas toward large companies and wealthy users able to pay more.
Advertising can eventually subsidize low-cost consumer products, but an advertising system takes time to build. An interim compute divide could emerge in which advanced intelligence is useful but ordinary users and smaller companies cannot obtain stable, affordable supply.
Open-weight models can lower software margins, increase choice, and give companies more control over data. They still carry a physical cost. Similarly sized models require chips and power to produce tokens, while cost per completed task varies substantially with the model, quantization, caching, and inference method.
Open models can improve the competitive structure without automatically removing the supply constraint. Lower margins may also stimulate consumption and increase total compute demand.
Multi-model systems compete for the intelligence layer
The speakers expect large enterprises to adopt multi-model systems.
A company might start with an open-weight base model, apply reinforcement learning or supervised fine-tuning on proprietary data, then connect one or two frontier models. The strongest model handles planning, review, or difficult work. Lower-cost models handle high-volume execution, with a router assigning each task.
Companies want control over capability, cost, and data. Giving all core context to one outside model creates pricing, privacy, and supply risk. Building everything internally requires continuous base-model upgrades, quality evaluation, routing, and infrastructure operations.
The interview names Fireworks Nexus, Cursor, Harvey, Microsoft, Databricks, and Snowflake as different entry points. The vendor list will change, while the strategic target remains relatively stable: the intelligence abstraction layer between an organization and its models.
That layer must read enterprise data, choose models, control cost, preserve permissions, upgrade continuously, and make the workflow feel coherent. It is easy to describe and much harder to operate reliably than ordinary middleware.
NVIDIA's advantage extends into financing and supply coordination
Baker describes NVIDIA's strategy as vertically integrated and horizontally open. It provides accelerators, CPUs, networking chips, switches, and complete systems, while allowing major customers to use their own silicon in parts of the stack.
That position changes the unit of competition. A startup with a better accelerator solves one component. Land, power, racks, memory, networking, software, delivery capacity, and financing still have to arrive together.
Baker offers two aggressive estimates. Every 1% of accelerator share may represent $100 billion of value, and NVIDIA may have locked up 70% to 80% of relevant supply. The interview supplies no independent data validating either figure, so both should be read as his assessment of ecosystem leverage.
Financing amplifies that leverage. A chip company can invest in a customer, provide a residual value guarantee, or accept warrants tied to token pricing in exchange for orders. When established financial institutions offer cheaper capital to a particular system, the customer is buying both hardware and a more financeable asset.
In an extreme shortage, nearly every product can sell out, so shipment volume reveals little about preference. Baker suggests watching deal terms instead. Which supplier must subsidize a customer? Which system earns a residual value guarantee? Who can link warrants to usage prices? Contracts expose bargaining power and customer confidence.
NVIDIA's moat now reaches beyond a single chip into system completeness, supply coordination, and cost of capital.
Orbital compute is a long-range stress test
The interview spends substantial time on orbital data centers. The speakers picture aircraft-sized compute units with solar wings and radiators in sun-synchronous orbit, not giant buildings floating in space.
Baker's illustrative ledger again starts at $50 billion per gigawatt. About $35 billion represents computing equipment, while roughly $15 billion covers terrestrial power, cooling, land, and construction. He argues that if a fully reusable Starship drives launch cost below $1 billion, the economics of orbital inference could reverse.
He also says training clusters would remain on Earth. Distance between GPUs and speed-of-light latency directly affect large-scale training, leaving orbital capacity as a possible source of inference or swing capacity.
The conversation also relays a proposed fourth-quarter 2027 launch for a Rubin-generation rack. That schedule, the cost split, and the engineering status are not independently confirmed in this source package and should be treated as a highly uncertain long-range scenario.
Orbital compute currently works mainly as a stress test. It gains economic room only if terrestrial power, copper, cooling, labor, and permits keep becoming more expensive while launch costs keep falling. If either curve fails to move as assumed, terrestrial data centers remain the dominant source of compute.
What must hold for the shortage thesis to work
The interview can be reduced to a chain of conditions.
First, AI use must spread beyond fewer than 10,000,000 to 30,000,000 heavy users toward a much larger population.
Second, gains in model and chip efficiency must not fully offset demand from new tasks, persistent agents, and longer inference.
Third, power, wafers, memory, networking, land, construction, financing, and permits must continue limiting build speed.
Fourth, users must keep receiving more value from frontier models than they pay, allowing suppliers to maintain fast paybacks and reinvestment.
If all four conditions hold, frontier labs, open models, clouds, applications, and chip companies can benefit together during the growth phase, while compute may become more expensive before it becomes broadly cheap. If diffusion slows, compute per task falls sharply, or financing costs rise, conventional overbuilding can return.
This framework is more actionable than a general argument about whether AI is a bubble. The next stage depends on three curves: heavy-user diffusion, compute efficiency per task, and the rate of new usable supply. The gap among them determines whether the market encounters scarcity, higher prices, or excess capacity.
- Published from
- atlasnote-editorial
- Published
- 2026-09-01
- Tags
- AIcomputeinfrastructureAgentsnvidiainterview