GPU Price Trends for AI Projects
The most consequential number in AI compute right now is a date: October 5, 2026, when CME Group will list the first exchange-traded futures on Nvidia H100 and B200 rental prices, built on indexes published by Silicon Data, a GPU-market intelligence firm backed by trading house DRW. The announcement, made with Silicon Data on August 11, 2026, gives AI compute a public, tradable reference price for the first time. That matters more than any single per-GPU-hour quote because it converts what has been opaque, relationship-driven negotiation into a benchmark that enterprises, hyperscalers, and investors can actually hedge against.
Three months ago, when this site published its first GPU spot price and capacity outlook, the headline was a $1.35-per-hour H100 SXM5 spot rate and a market that had eased from an all-out shortage into something more complicated. The picture has shifted since. H100 rental prices have risen nearly 40% in six months, the bottleneck has migrated from memory to power, and the capital structure of the entire buildout is now in question after reports that Nvidia could backstop $250 billion of OpenAI’s data-center financing.
Key Takeaways
- Nvidia’s H100 one-year rental contracts jumped to $2.35 per GPU-hour in March 2026, up from a $1.70 low in October 2025, according to SemiAnalysis.
- CME Group and Silicon Data launch the first H100 and B200 rental-index futures on October 5, 2026, giving AI compute a public reference price.
- The binding constraint has moved from chips to memory-and-packaging stack, then to power: Gartner expects power shortages to restrict 40% of AI data centers by 2027.
- CoreWeave’s contracted backlog hit $129 billion in Q2 2026, growing $29.6 billion in six weeks, while pricing rose 25%.
- Self-hosting breaks even only at high, steady use; the neo-cloud versus hyperscaler price gap on the same card now runs 40% to 85%.
Spot Pricing: The H100 Reversal and Two-Tier Market
The cleanest signal that the shortage phase is not over comes from SemiAnalysis, which tracked Nvidia’s H100 one-year rental contract pricing surging almost 40% to $2.35 per GPU-hour in March 2026 from a low of $1.70 per GPU-hour in October 2025, as reported by Seeking Alpha. That is a reversal of the normalization this site described in May, when GridStackHub.ai was citing H100 SXM5 spot at $1.35 per hour and A100 80GB at $0.35 per hour.
Carmen Li, former Bloomberg executive who left in 2024 to found Silicon Data specifically to track GPU prices, put the dynamic bluntly in an April 2026 interview with Business Insider: “The price is going pretty nuts,” she said of H100 rental prices. Her explanation cuts against the usual pattern where prices spike at a chip’s launch and then ease as supply scales. Instead, B200 pricing has stayed raised and even moved higher, a sign that capacity constraints are still biting across chips, memory, power, and data-center space.
The market has bifurcated into two tiers with very different economics. ClusterBid’s mid-2026 analysis of 84 enterprise deployments found H100 spot pricing had collapsed to $1.20 to $1.80 per GPU-hour as data-center buildouts hit the market, while B200 on-demand pricing sat at $7 to $9 per GPU-hour because of persistent scarcity, with reserved B200 contracts at $4.50 to $6.00 per hour. The lead-time spread tells the same story: H100 lead times have shortened to 4 to 8 weeks, H200 is available in 8 to 12 weeks, but B200 SXM parts still run 36 to 52 weeks from purchase order to rack-level acceptance.

The wide pricing band is the single biggest lever on unit cost. According to the eCorpIT capacity playbook, on-demand rates for the newest cards roughly doubled over the past year, while mainstream H100, H200, and A100 pricing held a tighter band. The neo-cloud versus hyperscaler gap on the same card now runs 40% to 85%.
| GPU class | Typical on-demand rate (per GPU-hour) | Lead time | Source |
|---|---|---|---|
| H100 (neo-cloud) | $1.49 to $2.50 | 4 to 8 weeks | eCorpIT capacity playbook |
| H100 (hyperscaler) | $6.88 to $12.29 | Quota-bound | eCorpIT capacity playbook |
| H200 | $2.30 to $13.78 (median near $4.11) | 8 to 12 weeks | eCorpIT capacity playbook |
| B200 | $4.99 to $18.00 | 36 to 52 weeks | eCorpIT capacity playbook |
| H100 (1-year contract) | $2.35 | Reserved | SemiAnalysis via Seeking Alpha |
Li’s most striking observation is about depreciation, the fear that has hung over every AI cloud valuation. In the second year, a refurbished H100 can still be sold for 85 cents on the dollar, and in year three for 84 cents, she told Business Insider. “My car depreciates a lot faster than that,” she said. That residual-value floor is why Wall Street has begun treating Nvidia GPUs as near-perpetual cash-flow machines, a framing that 24/7 Wall St. flagged as fragile in August.
The CME Moment: Compute Becomes a Tradable Commodity
The October 5 launch is a structural event that changes how the whole market prices risk. CME Group and Silicon Data will list two contracts, the Silicon Data H100 Rental Index Futures and the Silicon Data B200 Rental Index Futures, each representing a month’s worth of rent for the respective chip. The contracts will be listed under NYMEX rules, pending regulatory review.
The rationale, laid out in the joint press release, is that compute has become the “currency of the AI age.” CME’s Global Head of Energy and Environmental Products Pete Keavey drew the oil analogy directly: just as crude evolved from spot trading into a global derivatives market, compute futures turn GPU rental capacity into a standardized, tradable commodity that lets AI developers and hyperscalers lock in costs.
Silicon Data CEO Carmen Li framed the transparency gap the product closes: “For years, two companies buying the exact same GPU capacity could pay wildly different prices with no way to know who got the better deal. They will now have a benchmark to check that against.” Her firm ingests hundreds of thousands of pricing points globally and normalizes them into indexes covering both spot rentals and longer-term value.
The strategic implication is that GPU pricing stops being relationship-driven negotiation and becomes a market with a forward curve. For an engineering manager planning a training run six months out, that means the ability to hedge the cost of compute the way an airline hedges jet fuel. For hyperscalers, it means a public price that will discipline both procurement and the pricing they pass through to customers.
Capacity Buildout: CoreWeave, Crusoe, and the Financing Question
The supply side is expanding at a pace that would have looked extreme a year ago, but the capital structure behind it is now the subject of real scrutiny. CoreWeave’s Q2 2026 earnings, reported by CNBC on August 11, showed revenue of $2.575 billion, up 112% year over year, with contracted customer demand pushing above $129 billion. The company raised its 2026 revenue guidance to $12.4 billion to $13.2 billion and lifted its year-end active power target to more than 1.85 gigawatts. Pricing rose 25% year over year, and the backlog grew $29.6 billion in just six weeks.
That backlog growth is the clearest evidence that demand is not slowing, but it also carries a warning. The same report noted roughly 50% default odds being priced into some of that contracted demand, reflecting skepticism about whether every customer who signed a multi-year commitment will actually take delivery. This is the tension at the heart of the AI compute market in August 2026: order books are enormous, but probability-weighted reality is murkier.
The financing question exploded into public view in late July when the Wall Street Journal reported that Nvidia was in talks to guarantee up to $250 billion in financing for OpenAI’s data-center buildout. The report, summarized by 24/7 Wall St., sent Nvidia down 5% and AMD down 8% in a single session, reigniting “circular financing” concerns: a vendor funding a customer who then buys that vendor’s GPUs. Nvidia’s Q1 FY2027 data-center revenue hit $75.25 billion, up 92% year over year, and CEO Jensen Huang called the buildout “the largest infrastructure expansion in human history.” The market’s worry is how much of that revenue is truly independent demand versus seller-funded demand.
Crusoe represents the energy-strategy side of the same buildout. The company’s Abilene, Texas campus is a multi-gigawatt AI data-center hub, and Crusoe has leaned into modular, prepackaged data centers with $200 million of funding earmarked for smaller deployments, according to Forbes. The company is also testing a sodium-cooled nuclear microreactor to power AI data centers entirely on-site, a bet that the binding constraint is increasingly power, not silicon.
The AMD angle has become material as well. In July 2026, Anthropic agreed to purchase up to 2 gigawatts of GPU capacity from AMD, adopting MI450 series in Helios rack configurations, with AMD investing up to $5 billion in Anthropic in return, according to SiliconANGLE. Microsoft is also deploying AMD Helios racks in Azure, and Meta announced plans in February to purchase up to 6 gigawatts of AMD graphics cards. The multi-supplier strategy is a direct response to Nvidia supply-chain bottlenecks.
The Moving Bottleneck: HBM3e, CoWoS, and Power
The most important supply story of 2026 is that the bottleneck keeps moving, and buyers who anchor on last quarter’s constraint will misread the next one. GlobalData’s August 2026 analysis, reported by InvestorIdeas, frames the shift explicitly: record earnings across chipmakers are masking a supply constraint where memory, advanced-node capacity, and power, not compute demand, are setting the pace of AI deployment.
The memory numbers are staggering. SK Hynix’s net income jumped 397.5% year over year, Samsung’s 486.7%, and Micron’s stunning 1,398.3% on 345.7% revenue growth, all riding an HBM shortage that has sold out DRAM through 2026. SK Hynix holds roughly 58% of the high-bandwidth memory market in Q1 2026, a position that gives it outsized influence over how fast premium GPUs can ship.
CoWoS packaging is the second chokepoint. TSMC’s net income grew 77.4% against a 33.8% capex rise as it races to add advanced-node capacity, and its CoWoS capacity is expanding nearly 50%, according to Seeking Alpha supply-chain analysis. But the throughput of packaging, not raw silicon, determines how quickly wafers become deployable accelerators.
Power has now overtaken both as the binding constraint. Gartner’s prediction, made in November 2024 and now aging into a base case, is that power shortages will restrict 40% of AI data centers by 2027, with electricity for incremental AI servers reaching 500 terawatt-hours per year, 2.6 times the 2023 level. Constellation Energy’s 1,247.5% net income growth in the GlobalData analysis is the market’s verdict: power, not chips, is where scarcity now clears.
The practical consequence is that announced capacity and available capacity can diverge for months. A buyer can have GPUs allocated and still be unable to rack them because megawatts, cooling, and grid interconnect are not there. The hyperscalers are responding by signing multi-year power deals ahead of everyone else, which pushes mid-size buyers to the back of the queue.
Self-Hosting Versus API Pricing: Where Break-Even Actually Sits
The question every serious AI buyer eventually asks is whether to keep paying per token or serve the model themselves. The arithmetic has not changed in principle since May, but the numbers have moved. A 1,024-GPU cluster of H100 SXM parts at peak pricing costs roughly $8.4 million annually in compute alone, and over $15 million once networking, storage, facilities, and personnel are added, according to ClusterBid’s budgeting analysis.
The break-even logic is use-dependent, not price-dependent. ClusterBid’s survey of 84 deployments found that GPU compute represents 48% to 62% of total AI infrastructure cost, supporting infrastructure adds 22% to 30%, and operations contribute 12% to 18%. A B200 that trains a model 2.3 times faster than an H100 can reduce the total cost per training run despite a higher hourly rate, once compute time, storage, and operations are included.
The interruption risk is a hidden tax on spot-only strategies. Teams that use spot exclusively for training face 15% to 25% job interruption rates on H100 and 30% to 40% on B200 due to scarcity, per ClusterBid. The recommended blend is to reserve 70% to 80% of baseline capacity on 12-month contracts and use spot or on-demand for the remaining 20% to 30%, which captures roughly 75% of the fully-reserved discount without locking in maximum capacity that may go unused.
There is also a human-capital dimension that generic cost comparisons ignore. Self-hosting requires people who can tune models, schedule workloads, watch use, and keep clusters healthy. APIs outsource most of that. A startup with a small platform team may rationally overpay per token because it cannot afford to underwrite infrastructure complexity, while a larger company with a strong infrastructure bench may rationally do the opposite.
The token-API side of the comparison has its own moving parts. Anthropic’s second-quarter revenue reportedly jumped to more than $11.5 billion, up 14-fold year over year, according to CNBC, evidence that the per-token model remains the dominant route for most AI consumption even as the compute market matures.
What to Watch Through the Rest of 2026
The next few quarters are unlikely to produce a clean collapse in GPU pricing. A more probable path is uneven softening, where mature hardware keeps getting cheaper in spot markets while the newest high-performance tiers stay sticky because memory, packaging, and power constraints hit those segments first. That widens the gap between practical inference economics and frontier training economics.
Three signals matter most. First, the CME compute futures launch on October 5 will reveal the forward curve for GPU rental prices for the first time, and the shape of that curve will tell buyers whether the market expects scarcity to persist or ease. Second, watch whether the $129 billion CoreWeave backlog converts to actual revenue or whether default odds materialize, since that conversion rate is the cleanest read on whether AI demand is real or financed. Third, track HBM and CoWoS cadence against power delivery, because the bottleneck’s location determines which suppliers hold pricing power in any given quarter.
The geopolitical layer has become a permanent cost variable. In January 2026, the Trump administration invoked Section 232 to impose a 25% tariff on select advanced AI chips, and the Bureau of Industry and Security shifted H200 and MI325X export reviews for China from presumptive denial to case-by-case approval capped near 50% of prior US sales. TSMC, Samsung, and SK Hynix lost their blanket Validated End-User exemptions and now need annual US licenses to ship equipment into China fabs. These are structural changes that will outlast any single quarter’s pricing print.
For builders, the actionable lesson from May still holds, but with a sharper edge. The shortage phase has eased in the mature tier, but the market still prices real constraints, and those constraints have migrated up the stack toward memory, packaging, and power. What matters now is not whether capacity exists somewhere, but whether the capacity you need can be procured under terms that match your workload and timeline. The CME futures market, when it opens in October, will finally make that question answerable with a number instead of a negotiation.
For more on how compute access shapes bargaining power across the AI sector, see this site’s earlier GPU spot price and capacity outlook, and for serving-side optimization that changes the cost-per-token math, the local AI inference engine guide. The broader capital story, including the $725 billion hyperscaler capex cycle behind this entire market, is covered in the AI data center buildout series.
Related Reading
More in-depth coverage from this blog on closely related topics:
Sources and References
Sources cited while researching and writing this article:
- Nvidia’s H100 GPU rental prices surge nearly 40% in 6 months: SemiAnalysis
- This CEO left Bloomberg to track GPUs. She explains why prices are ‘going nuts.’
- The 2026 AI compute crunch: a GPU capacity playbook
- CME Group and Silicon Data to Launch Compute Futures on October 5 to Unlock New Way to Hedge AI Risks
- AI Stocks Crash After NVIDIA Plans to Finance $250 Billion OpenAI Buildout Are Reported
- Anthropic to buy up to 2 gigawatts of GPU capacity from AMD
- AI Infrastructure Boom Shifts From Chip Race to Supply-Chain Race
- CNBC
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
