Nvidia GPU Innovations and AI Inference Costs
Key Takeaways:
- Independent MLPerf Inference v6.1 results, published September 16, 2026, show Vera Rubin NVL72 at up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL.
- Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s, 2.8x Blackwell’s bandwidth, the architectural change most responsible for long-context inference economics.
- Memory inflation is the counterforce: TrendForce projects DRAM and NAND at 68% of cloud hardware capex by 2027, and Nvidia reportedly told customers that Vera Rubin and Grace Blackwell servers would cost more than 15% more on early-2027 shipments.
- d-Matrix will integrate its Raptor inference XPUs into Nvidia’s MGX racks through NVLink Fusion, with initial rack availability expected in Q4 2027.
Nvidia’s own numbers say the Vera Rubin NVL72 rack delivers up to 30x more agentic AI inference throughput per megawatt than the GB300 NVL72 it replaces, while cutting cost per million tokens by up to 35x, figures published in its developer blog and repeated across GTC and Hot Chips messaging in 2026. The independent benchmarks that followed tell a more measured story.
Vera Rubin and the 30x Throughput Claim
The mechanism behind the claim is workload-specific. Nvidia measured Vera Rubin NVL72 on the SemiAnalysis AgentX benchmark, which replays recorded agentic coding sessions with real context growth, tool calls, and sub-agent spawning preserved. The company ran DeepSeek V4 Pro on the system and reported up to 30x higher throughput per megawatt than GB300 NVL72, according to Pulse 2.0’s report on the disclosure.

Agentic workloads explain the comparison. Nvidia cited OpenRouter data showing agentic AI workloads consume roughly 15 times more tokens than a simple chat request. An agent that queries databases, searches documents, invokes tools, and spawns sub-agents accumulates context at every step, pushing sessions into hundreds of thousands of input tokens. A rack optimized for short chat turns and one optimized for long agent loops diverge sharply on throughput per megawatt.
The cost claim follows the same logic: up to 35x lower cost per million tokens versus GB300 NVL72. Fudzilla’s summary of the developer blog notes the figures do not yet include Vera CPU performance for tool calling and focus on the Vera Rubin silicon rather than the full seven-chip platform. The 35x number describes a partial system, measured by the vendor, on a vendor-selected benchmark.
Inside the Rubin Architecture: HBM4 and NVFP4
The architectural change doing most of the work is memory bandwidth. Each Rubin GPU integrates up to 288 GB of HBM4 delivering 22 TB/s of aggregate bandwidth, a 2.8x improvement over Blackwell’s 8 TB/s, according to TechTimes’ breakdown of Nvidia’s July 17 technical blog. At rack level, a single Vera Rubin NVL72 combines 72 GPUs and 36 CPUs into 20.7 TB of HBM4 memory and 1.6 PB/s of cumulative bandwidth.

Decode is memory-bound. Generating each output token requires loading model weights from memory, so throughput tracks bandwidth more closely than compute. Doubling bandwidth does more for long-context inference than adding arithmetic units, which is why the HBM4 jump is the number to watch rather than transistor count.
The second lever is precision. NVFP4 quantization reduces model weights, attention, and KV cache to 4-bit precision, cutting memory footprint and raising throughput. Nvidia says this preserves output quality, though quantization always trades some accuracy for speed, and the tradeoff grows worse at aggressive bit widths. The third lever is the interconnect: NVLink 6 runs at 3.6 TB/s bidirectional per GPU, double Blackwell’s 1.8 TB/s, letting all 72 GPUs in a rack share memory coherently. Nvidia says its sixth-generation NVLink and switch silicon provide 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet.
The Software Layer Doing the Heavy Lifting
Hardware alone does not produce these numbers. The most important technique is disaggregated serving, which separates prefill (processing the input prompt) from decode (generating output tokens) onto separate GPU pools that scale independently. Prefill is compute-bound and parallelizes well; decode is memory-bound and serial. Running both on the same GPU forces them to compete for the same resources.
Nvidia’s Dynamo framework extends this to multimodal inputs with Encode-Prefill-Decode disaggregation. According to Blockchain.News’ report on the Dynamo EPD technique, separating the vision encoder from the language model stages delivers up to 7x faster end-to-end response times and 5x faster time-to-first-token in media-heavy scenarios, and supports 70% more traffic at the same latency threshold for image-heavy use cases. The gains shrink when output length dominates total latency, so long-form text generation sees much less benefit.
For mixture-of-experts models, large-scale expert parallelism distributes individual expert networks across the NVL72 scale-up domain, and distributed KV caching extends available memory across that domain. KV-cache offloading moves less-active context into host memory or storage without forcing recomputation, while KV-aware routing directs new requests toward GPUs that already hold relevant cached context. The snippet below shows how a team would route a long-context agent request using these primitives in a vLLM-style server.
from dataclasses import dataclass
@dataclass
class RouteDecision:
prefill_pool: str
decode_pool: str
cache_hit: bool
reason: str
def route_agent_request(request, cache_index, prefill_pool, decode_pool):
"""Route a multi-turn agent request using KV-aware cache lookup.
cache_index maps a prefix hash to the GPU pool holding that KV cache.
A cache hit means the decode pool can skip re-processing prior turns.
"""
prefix_hash = request.prefix_hash()
holder = cache_index.get(prefix_hash)
if holder is not None:
# Reuse cached context: decode can start immediately.
return RouteDecision(
prefill_pool=holder,
decode_pool=holder,
cache_hit=True,
reason="prefix cache resident on pool",
)
# Cache miss: prefill must process the full context first.
return RouteDecision(
prefill_pool=prefill_pool,
decode_pool=decode_pool,
cache_hit=False,
reason="no cached prefix, full prefill required",
)
# Note: production use needs a bounded cache with eviction (LRU or TTL),
# a fallback when the target pool is saturated, and correctness checks
# that the cached prefix still matches the incoming token sequence.
What Independent Benchmarks Actually Show
MLPerf Inference v6.1, published September 16, 2026, is the first peer-reviewed data on Vera Rubin NVL72. Nvidia submitted preview results showing up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios, and up to 2.5x higher throughput on DeepSeek-R1, according to Unite.AI’s summary of the submission. Those are auditable numbers, roughly an order of magnitude below the 30x figure Nvidia markets, because MLPerf measures standard prompt-response inference while the 30x claim measures long agentic sessions.
The MLCommons release itself is the more useful signal for buyers. This round set a participation record with 30 submitting organizations, and the MLCommons press release reports that the best per-accelerator server result for the DeepSeek-R1 test improved 5.7x compared with v5.1 one year earlier, while the VLM test improved 2.99x from v6.0 just six months earlier. Much of that came from software, not new silicon: Nvidia reported that GB300 NVL72 performance on Qwen3-VL improved up to 1.6x in v6.1 over v6.0 through lower KV cache precision, additional kernel fusion, and disaggregated serving.
| Platform | Claimed vs GB300 NVL72 | Workload / source | Measured by |
|---|---|---|---|
| Vera Rubin NVL72 | Up to 30x throughput per MW; up to 35x lower cost per million tokens | AgentX agentic coding, DeepSeek V4 Pro | Nvidia, via Pulse 2.0 |
| Vera Rubin NVL72 | Up to 3.7x higher throughput | MLPerf v6.1 Qwen3-VL, Closed division preview | MLCommons, via Unite.AI |
| Vera Rubin NVL72 | Up to 2.5x higher throughput | MLPerf v6.1 DeepSeek-R1, TensorRT-LLM | MLCommons, via Unite.AI |
| GB300 NVL72 | Up to 15x throughput per MW vs Hopper H200 | AgentX, DeepSeek V4 Pro 1.6T | Nvidia, via Fudzilla |
D-Matrix, NVLink Fusion, and Custom Silicon
Nvidia’s answer to custom inference silicon is to host it. On September 10, 2026, d-Matrix announced a multi-year roadmap collaboration to integrate its next-generation Raptor inference XPUs into Nvidia’s MGX rack-scale architecture through NVLink Fusion, with initial rack-integrated availability expected in Q4 2027, according to Unite.AI’s coverage of the announcement.
The arrangement is heterogeneous disaggregation. For AI coding, the compute-intensive prefill stage runs on Vera Rubin GPUs while the latency-sensitive decode stage runs on d-Matrix XPUs, with the racks designed to operate side by side. Nvidia reports NVLink Fusion delivers 3x lower XPU-to-XPU latency than off-the-shelf Ethernet and 10x higher packet rates, and lists AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung, and others among its Fusion partners.
This is a strategic concession dressed as an ecosystem play. By letting third-party accelerators plug into its rack architecture and software stack, Nvidia keeps the interconnect, the rack, and the orchestration layer on its own silicon even as customers route decode workloads to someone else’s chip. A d-Matrix rack that depends on Nvidia’s NVLink switches and MGX supply chain is not an independent path away from Nvidia, it is a different slot in the same system.

The Memory Cost Counterforce
Lower cost per token at the rack level does not mean lower cost per token for the buyer, because memory inflation is pushing system prices the other way. TrendForce’s August 25, 2026 research projects that DRAM and NAND Flash combined will account for 68% of major cloud providers’ hardware capital expenditure by 2027, up from 47% in 2026, according to TechTimes’ analysis of the TrendForce data.
The same reporting notes that Bloomberg reported on August 22, 2026 that Nvidia notified its largest customers that AI servers containing Vera Rubin and Grace Blackwell chips would cost more than 15% more on units shipped in early 2027, with memory costs cited as the primary driver. A single Vera Rubin GPU carries 288 GB of HBM4, and an NVL72 rack carries roughly 20 TB of HBM, placing HBM component cost in the range of $316,000 to $360,000 per rack before compute, networking, or power infrastructure. Nvidia has not publicly confirmed the specific increase figures.
Nvidia’s per-token efficiency gains and memory-driven system price increases pull in opposite directions. A buyer comparing Vera Rubin to GB300 on Nvidia’s throughput-per-megawatt metric sees a dramatic improvement. A buyer comparing the total cost of a deployed rack against the prior generation sees a smaller net gain, because the hardware that delivers the efficiency costs more to acquire. As covered in this site’s October analysis of GPU pricing trends, neocloud on-demand rates rose 17% to 21% on October 1, 2026, with B300 reaching $9.50 per GPU-hour, so the efficiency gains are landing in a market where capacity prices are already climbing.
Deployment Scale: AWS, Crusoe, and Groq 3 LPX
The efficiency claims only change industry economics if the hardware ships at scale. Nvidia’s dedicated inference accelerator, Groq 3 LPX, entered full production on August 24, 2026, purpose-built as an extension to the Vera Rubin platform. According to SiliconANGLE’s report from Hot Chips 2026, a rack-scale deployment can harness up to 256 LP30 accelerators, and Artificial Analysis benchmarked Groq 3 LPX at 3,400 tokens per second running the open-source Gemma 4 31B agentic model with a 100,000-token context window. Nebius signed on as the first customer to commit to the chip.
That 3,400 tokens-per-second figure comes from a third-party benchmarking firm, making it more credible than a vendor self-report, but it measures a single model on a single context configuration. It does not establish that the chip delivers similar speedups across the model and workload mix a production fleet actually serves.
What to Watch Into 2027
Three signals will determine whether Nvidia’s inference economics translate into real cost declines for buyers. First, watch whether AgentX-style agentic benchmarks get replicated independently at production scale. The 30x and 35x figures are vendor measurements on a vendor-selected workload; MLPerf’s 2.5x to 3.7x results are the audited number until someone runs AgentX on their own fleet.
Second, watch the memory line. Nvidia guided Q3 fiscal 2027 gross margin down to 74.0% from 75.0% because of rising memory costs, and its supply commitments more than doubled to $279 billion, a figure that secures HBM capacity but becomes a fixed obligation if demand softens. If HBM4 pricing keeps climbing, per-token cost improvements at the rack level may not reach the buyer’s invoice.
Third, watch whether heterogeneous disaggregation becomes the default architecture. If d-Matrix Raptor racks and Groq 3 LPX accelerators end up carrying a meaningful share of decode traffic in 2027, Nvidia’s position shifts from selling the whole inference stack to selling the interconnect and orchestration layer around someone else’s silicon, still a strong position, but a different business than the one the 35x cost claim implies. Nvidia’s Vera Rubin NVL72 will exceed 3.7x throughput versus GB300 NVL72 on an independently audited MLPerf Inference benchmark by the v7.0 results release, because the v6.1 figures are preview submissions and the software optimizations Nvidia reported post-submission have not yet been verified by MLCommons.
Related Reading
More in-depth coverage from this blog on closely related topics:
- GPU Pricing Forecast for AI Use
- Prime Big Deal Days 2026: Dates and Discounts
- What Is CyberLeek? An Explanation
- Future of Chip Supply Chains
- Cleo Mathematician Background
Sources and References
Sources cited while researching and writing this article:
- Vera Rubin makes Blackwell look pricey – Fudzilla.com
- NVIDIA Vera Rubin Cuts Post-Training Token Costs With Seven-Chip Codesign
- NVIDIA Dynamo EPD Boosts Multimodal AI Model Speed by 7x
- NVIDIA Vera Rubin NVL72 Posts First MLPerf Inference Preview Results
- MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Results
- D-Matrix Connects Raptor XPUs to NVIDIA AI Factories via NVLink Fusion
- Memory Will Cost Cloud Giants More Than GPUs by 2027: NVIDIA Uses That to Justify 15% Server Hike
- Crusoe Signs Multi-Year Deal to Power Perplexity Training and Inference
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
