Rows of GPU server racks in a data center running AI training and inference workloads

GPU Pricing Forecast for AI Use

October 7, 2026 · 12 min read · By Rafael
  • Nebius raised on-demand GPU rates 17-21% on October 1, 2026: H100 to $4.50/hr, H200 to $5.40, B200 to $8.50, B300 to $9.50.
  • CoreWeave (CRWV) has pushed through roughly 37.5% in cumulative price increases since July, and is signing short contracts near $40 million per megawatt.
  • AWS raised EC2 Capacity Block reservation prices about 20% on July 1, 2026, with P6-B300 at $14.04 per accelerator-hour.
  • Runpod’s fleet data shows B200 list prices up 31% between January and August 2026, with high-bandwidth memory effectively sold out into 2027.
  • Custom ASIC shipments are growing 44.6% in 2026 versus 16.1% for merchant GPUs, per TrendForce.

The Price Reversal: H100 Is No Longer Cheap

The clearest evidence comes from providers themselves. Nebius disclosed its October 1 card through customer communications reshared on Reddit and X. H100 hourly rates rose about 17% to $4.50, H200 rose 20% to $5.40, B200 nearly 19% to $8.50, and B300 about 21% to $9.50 from $7.85. CPU pricing moved too, with AMD (AMD) EPYC Genoa climbing 25% to $0.015 per vCPU-hour and Genoa memory up roughly 41% to $0.0045 per GiB-hour. NBIS shares jumped 9.15% premarket on September 17 when the hike surfaced, and the stock is up about 110% over six months, per GuruFocus.

Nebius’s Q2 numbers explain the confidence. Revenue rose 454% year over year to $582.3 million, with AI Cloud revenue up 514% to $574.9 million, and adjusted EBITDA swung to $236.2 million profit from a $21 million loss. The company reaffirmed 2026 revenue guidance of $3 billion to $3.4 billion and expects annualized recurring revenue of $7 billion to $9 billion by year-end.

On-demand GPU rates are rising again in late 2026, reversing the spring normalization that briefly made H100 spot capacity feel abundant.

Rows of GPU server racks in a data center running AI training and inference workloads
Rows of GPU server racks in a data center running AI training and inference workloads.

The reversal is not limited to one neocloud. SemiAnalysis tracked Nvidia (NVDA) H100 one-year rental contracts climbing nearly 40% to $2.35 per GPU-hour in March 2026 from a $1.70 low in October 2025, as covered in this site’s August GPU price trends analysis. Runpod’s Fall 2026 State of AI Compute report, drawing on over 1.2 million users across 186 countries, found flagship GPU list prices rising through 2026, including a 31% B200 increase between January and August. The report’s framing is blunt: “Hardware is scarce, flagship GPU prices are rising, and AI bills are becoming material enough that teams can no longer treat the largest frontier model as the default answer to every problem,” according to the press release.

Hyperscalers Are Raising Prices Too

The neocloud hikes might be dismissed as a small-provider phenomenon, but AWS’s July 1 change confirms the trend at the top of the market. Amazon Web Services (AMZN) raised EC2 Capacity Block reservation prices for machine-learning GPU instances by about 20%, citing “supply and demand dynamics.” New per-accelerator hourly rates span the fleet: P6-B300 at $14.04, P6-B200 at $12.355, P5 in US regions at $5.191, P5e at $5.97, and P4de at $2.214, per Investing.com.

Capacity Blocks are a reservation product, so the hike is structurally significant. Customers pay a premium over spot specifically to guarantee availability for time-bound training runs, so a 20% step-up signals that even committed capacity is scarce. AWS revenue grew 28% year over year to $37.6 billion in Q1 2026, its fastest cloud growth in more than three years, and Amazon has committed roughly $200 billion in 2026 capex. Reuters reported in March that Amazon is set to receive 1 million Nvidia GPUs by end-2027, itself a measure of how supply-constrained the high-end market remains.

Hyperscaler data center GPU pricing comparison chart
Hyperscaler data center GPU pricing comparison.

The spread within a single card class is the real story: the same H100 a neocloud rents for $4.50 an hour costs AWS customers $5.19 to $6.87 on a reserved P5 instance, and the gap between neocloud and hyperscaler on identical silicon has run from 40% to 85% depending on card and term.

Provider / product Rate (per GPU-hour) Effective date Source
Nebius H100 (on-demand) $4.50 (from $3.85) Oct 1, 2026 GuruFocus
Nebius H200 (on-demand) $5.40 Oct 1, 2026 GuruFocus
Nebius B200 (on-demand) $8.50 Oct 1, 2026 GuruFocus
Nebius B300 (on-demand) $9.50 (from $7.85) Oct 1, 2026 GuruFocus
AWS P6-B300 (Capacity Block) $14.04 Jul 1, 2026 Investing.com
AWS P5 (US regions) $5.191 Jul 1, 2026 Investing.com
Nvidia H100 (1-year contract) $2.35 Mar 2026 SemiAnalysis via Seeking Alpha

Capacity Buildout: CoreWeave, Crusoe, and the Financing Question

The supply side is expanding at a pace that would have looked extreme a year ago, but the capital structure behind it is under close examination. CoreWeave scaled data center capacity from 70 MW to 1.5 GW by mid-2026, backed by a contracted revenue backlog that reached $129 billion in Q2 2026, according to 24/7 Wall St. That backlog grew $29.6 billion in six weeks, and the company raised its year-end active power target to more than 1.85 gigawatts.

At its Fully Connected 26 user conference, management disclosed another 10% price increase after a 25% hike in July, putting cumulative pricing up roughly 37.5% since July. Truist analyst Arvind Ramnani reiterated Buy and a $165 target, arguing pricing is rising faster than GPU costs. Management said 70% of Q2 2026 deals included prepayments, and highlighted an Nvidia A100 contract running through 2029, suggesting chips first launched in 2020 may generate revenue for nine years.

There is a counter-signal. JPMorgan’s Samik Chatterjee argues CoreWeave is deliberately signing shorter contracts, near $40 million per megawatt, because shorter terms reset pricing upward as the market tightens. That is a margin strategy, but it also carries renewal risk if demand softens. A backlog that resets every few quarters is not the same asset as a seven-year committed contract.

Crusoe sits on the energy side of the same buildout. It closed a $3.9 billion late-stage round at a $30.9 billion valuation in September, led by Atreides Management, Mubadala Capital, and Valor Equity. It opened a second Tulsa, Oklahoma manufacturing plant on September 14, adding 400,000 square feet for electrical switchgear behind its AI infrastructure, and it is developing a Google (GOOGL) data center campus in Armstrong County, Texas, built at the source of West Texas clean energy. Crusoe’s bet is that power, not silicon, is the real bottleneck.

Microsoft (MSFT) is the other side of the demand ledger. Bloomberg reported in September that Microsoft plans to build out to about 38 gigawatts of data center capacity by 2032, more than triple its current footprint, per U.S. News. That scale of commitment, coming after Microsoft reportedly lost clients to a capacity crisis, shows why demand keeps bidding prices up even as capacity comes online.

The Memory Bottleneck: HBM Sold Out Into 2027

The pricing dislocation is what matters for AI buyers. A memory module that costs $300 to $400 under a long-term supplier agreement trades at $2,100 on the spot market, a five-to-seven-fold premium. That is why Nvidia’s gross margin guidance is drifting toward 74% and why the company more than doubled its supply commitments to $279 billion, as detailed in this site’s Nvidia data center analysis. Memory is now the line item that sets GPU pricing.

The US HBM market was valued at $1.09 billion in 2025 and is forecast to grow from $1.43 billion in 2026 to $4.91 billion by 2031, a 27.98% CAGR, according to Research and Markets. HBM3E still accounted for 71.32% of the US market in 2025, but the transition to HBM4 is underway: Samsung began mass production of commercial HBM4 in February 2026 at up to 3.3 TB/s per stack, and Micron (MU) entered high-volume HBM4 production in March 2026 for Vera Rubin. Samsung began shipping HBM4E samples in May 2026 at up to 3.6 TB/s and 48 GB capacity.

Advanced packaging remains a second-order constraint. Frontier accelerators depend on a narrow group of qualified integration processes, and SK hynix’s Indiana facility is not expected to begin mass production until the second half of 2028. Micron’s Idaho, New York, and Virginia investments are multi-year programs rather than near-term additions. Even when wafer output rises, CoWoS-style packaging throughput determines how fast wafers become deployable accelerators.

Custom Silicon Is Eating the Inference Tier

The pricing pressure on merchant GPUs has accelerated a structural shift already underway. TrendForce data shows custom AI chip shipments from cloud providers growing 44.6% in 2026, against 16.1% for merchant GPUs, a nearly three-to-one gap that marks the first year custom silicon has meaningfully outpaced general-purpose processors, according to TechTimes. The economics are concrete: Midjourney reported cutting monthly compute costs from about $2.1 million to $700,000 after moving inference from Nvidia GPUs to Google’s seventh-generation TPUs, a 65% reduction.

Broadcom (AVGO) is the clearest listed beneficiary, with custom AI accelerator revenue surging 221% to $16.7 billion in its third fiscal quarter of 2026. OpenAI’s in-house “Jalapeño” chip reportedly beat Nvidia Blackwell systems on key inference-efficiency tests, and Anthropic confirmed an in-house chip team in August targeting roughly 50% cuts in per-token inference costs. The displacement is concentrated in inference, where workloads are predictable enough to justify custom design, while training remains more firmly with Nvidia because CUDA’s flexibility keeps diverse enterprise workloads on merchant GPUs.

Wafer-scale and commodity-memory approaches attack the bottleneck from a different angle. Cerebras sidesteps HBM, CoWoS, and 3nm constraints with 5nm wafer-scale chips, backed by a backlog that reportedly includes OpenAI. Positron AI raised $875 million at a $5 billion valuation to build inference chips using commodity memory rather than HBM, a direct bet that the memory bottleneck can be engineered around. Neither has displaced Nvidia at meaningful scale yet, but both target the specific cost line, memory, now driving merchant GPU prices higher.

Self-Hosting Versus API Pricing: Where Break-Even Sits Now

The price reversal changes the self-hosting arithmetic many teams ran in spring. When H100 spot sat at $1.35 an hour, the case for owning hardware was marginal outside heavy use. At $4.50 an hour on-demand, or $2.35 on a one-year contract, the break-even point against per-token APIs has shifted. The calculation still comes down to use and throughput, but the spread between the two paths has widened in favor of whoever can keep hardware busy.

Spot-only strategies carry real interruption risk. This site’s prior coverage cited ClusterBid data showing 15% to 25% job interruption rates on H100 and 30% to 40% on B200 under scarcity, with the recommended approach being to reserve 70% to 80% of baseline capacity on 12-month contracts and use spot for the remaining 20% to 30%. That guidance holds, but the reservation price has now risen, raising the floor on the entire self-hosting decision.

Runpod’s data shows how builders respond without renting more hardware. Quantization has become the default efficiency lever: 40.7% of pods running models over 70B params serve a quantized build, versus 32.5% for 8B-70B models. Qwen now runs on 74.2% of text endpoints, with Qwen3.x growing from 24.3% to 50.4% since November 2025, and lower-precision 4-bit and 8-bit builds gaining share as 16-bit deployments fell from 95.9% to 85.8%. Four in five pods that reference a frontier model API also run local weights or a local serving stack. The shift is from “pick the biggest model” to “pick the right model, precision, and amount of compute for each job,” as Runpod’s Charlotte Daniels put it.

What to Watch Into 2027

Three indicators will determine whether current price firmness persists or breaks. First, watch whether CME Group’s H100 and B200 rental-index futures, originally slated for October 5, actually clear the CFTC. The Information reported in late September that the launch had stalled at the regulator, meaning the forward curve that would make compute a hedgeable commodity is not yet live. If it clears, the curve’s shape will reveal whether the market expects scarcity to continue into 2027.

Second, watch CoreWeave’s backlog conversion. A $129 billion backlog growing $29.6 billion in six weeks is a demand signal, but it carries roughly 50% default risk priced into some contracts, and shorter terms mean renewal risk arrives sooner. If prepayments keep flowing and pricing sticks, neoclouds have genuine pricing power. If use softens, the same short contracts that boost margins now become a liability.

Third, watch memory and power in tandem. HBM is sold out into 2027, but Samsung’s HBM4E ramp and Micron’s high-volume HBM4 production will determine whether the constraint eases or tightens through the Vera Rubin cycle. Power is the quieter risk: several providers are waiting on grid interconnects for completed facilities, and Google was ordered to halt work on two data centers in Finland this week. The bottleneck has moved from silicon to memory to power, and buyers who anchor on last quarter’s constraint will misjudge the next one.

For builders, the actionable takeaway is unchanged from May but sharpened by the price reversal. The shortage has not eased; it has migrated up the stack. Mature H100 capacity is no longer cheap, premium Blackwell capacity is quota-bound, and memory in every card is the scarcest asset in the stack. Teams that secured long-horizon compute earlier now hold a structural cost advantage, and teams that waited are paying the difference in both price and lead time. In late 2026, the GPU-hour is not a commodity, it is a negotiation.

My forecast is specific: CoreWeave’s contracted backlog will exceed $150 billion by the end of 2026, because the backlog grew $29.6 billion in six weeks during Q2 and 37.5% cumulative price increases since July have not slowed customer signings. The mechanism is the one driving the whole market: memory is sold out into 2027, which keeps premium capacity scarce, which lets providers with secured supply keep raising prices while demand remains inelastic.

Prediction: CoreWeave (CRWV) will report contracted backlog above $150 billion by its Q4 2026 earnings, because its backlog grew $29.6 billion in six weeks during Q2 2026 and the 37.5% cumulative price increase since July has not slowed customer demand for scarce premium capacity.

Related reading: GPU Price Trends for AI Projects · Hyperscaler Capex and AI Infrastructure · Nvidia Data Center Growth in 2026 · Hyperscaler Investment Trends for 2026

Disclosure: This article is for informational and analytical purposes only. It is not investment advice, a price target, or a recommendation to buy, sell, or hold any security.

More in-depth coverage from this blog on closely related topics:

Sources and References

Sources cited while researching and writing this article:

Rafael

Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...