The binding constraint on AI startups through 2025 was not talent or demand but compute — the supply of top-tier accelerators, the price of renting them, and the multi-year commitments required to guarantee either. Nvidia's datacenter revenue told the macro story — tens of billions of dollars per quarter through 2025 — but the startup-level story is about contract structure: who holds capacity when the market tightens, at what price, and with what rights. This is an analysis of the documented market; it is not investment advice.
What does compute actually cost a startup?
The documented price points of 2024-2025: a single top-tier accelerator rents on-demand from major clouds at roughly $3 to $10 per hour depending on class and configuration; training a frontier-adjacent model runs from millions of dollars for mid-scale efforts to hundreds of millions at the frontier; and inference at scale — the recurring cost that kills margins — scales linearly with usage unless efficiency work compounds. For a typical funded AI application startup, compute and inference costs land in the range of 20 to 40 percent of revenue, per the margin discussions that spread through investor writing in 2024-2025 — versus the 10-15 percent cloud cost of an equivalent SaaS company a decade earlier. That gap is the AI margin question in one line.
How does allocation actually work when supply is short?
The market's structure in shortage: hyperscalers and the big labs take capacity through multi-year contracts with prepayments — Microsoft's and Meta's disclosed capex ran to tens of billions per quarter through 2025 — and startups compete for the remainder through three channels: reserved instances on the public clouds, where the best price requires one-to-three-year commitments; specialized GPU clouds (CoreWeave, Lambda, Crusoe and peers), which raised debt against GPU-backed contracts to buy fleets and rent them out; and the programs of the chipmakers themselves — Nvidia's Inception and startup investments — which allocate scarce capacity to companies the ecosystem wants to exist. The documented pattern when demand exceeded supply in 2024-2025: allocation became a relationship business, with capacity going to customers with proven revenue, investor connections to the provider, or strategic value — market pricing by other means.
What did export controls change?
U.S. export controls on top-tier accelerators to China, tightened in stages through 2023-2025, did three documented things. They created a two-tier market — controlled chips and the performance-reduced variants sold into permitted markets. They redirected supply geopolitically, with the controls a recurring point of negotiation in U.S.-China trade talks through 2025, including reported proposals to loosen restrictions as part of broader agreements. And they motivated China's domestic accelerator push — Huawei's Ascend line being the most documented — whose success at mid-tier performance is part of the DeepSeek-era evidence that capability does not strictly require the top of the export-controlled stack. For startups, the controls matter mostly through price: any loosening softens rental prices market-wide; any tightening re-tightens them.
How should a startup manage compute as a supply chain?
The practices that survived contact with the 2024-2025 market, per engineering leadership writing and company case studies. Model-portability from day one: an inference stack locked to one provider is a negotiating position of zero. Efficiency before scale: quantization, batching, caching, and routing to smaller models routinely cut inference cost by half or more — the documented lever that separates profitable AI applications from demo-stage ones. Commitment calibration: reserve capacity for the baseline load, burst on-demand for peaks, and re-baseline quarterly. And the hedge the market keeps re-teaching: capable open-weights models runnable on owned mid-tier hardware are the exit from any pricing squeeze — the strategy whose economics the January 2025 DeepSeek release made impossible to dismiss.
What is the state of supply heading into 2026?
On the documented record: massive new capacity is being built — the Stargate venture and hyperscaler capex at a combined pace of hundreds of billions annually — which argues for softening prices as 2026-2027 capacity lands. Against that: demand has absorbed every previous supply increase, the frontier training runs of 2026 will consume multiples of 2025's, and power — not chips — is emerging as the next binding constraint, with grid interconnection queues now the long-lead item in datacenter plans. The honest forecast on the record is cyclical: relief windows followed by crunches, with the calendar unknowable and the hedged architecture the only durable answer.
Compute became the startup cost line that strategy is built around. The companies that treat it as a supply chain — sourced, hedged, engineered down — keep their margins; the ones that treat it as a utility bill learn the difference at their Series B.
For more context, read Datacenter Energy: The Wall AI Startups Are Heading Toward.
For more context, read open weights vs closed models.
For more context, read Open Source in the AI Era: How Projects Now Get Funded.

