RTX 5090 vs M3 Ultra Performance
Key Takeaways:
- An RTX 5090 street price of roughly $4,900 to $5,199 in the US in September 2026 reflects a buying-window issue, not a permanent price. The same card launched at $1,999.
- The 5090’s 32GB of GDDR7 limits the model size you can run whole, while the M3 Ultra (up to 512GB unified) and Strix Halo (up to 128GB) accommodate much larger models in memory.
- Token generation is limited by bandwidth, so the 5090 leads in raw decode speed on models that fit entirely; the other two perform better on larger models that do not fit.
- When a model must spill layers over PCIe, a single 5090’s throughput advantage decreases or reverses.
- Make purchasing decisions based on effective cost per token over three years, not just on the spec sheet.
Memory and Price, Not Clock Speed
An RTX 5090 with 32GB of GDDR7 now costs roughly two and a half times its $1,999 launch price in the US. TechSpot’s September 2026 pricing feature recorded the card around $4,900 in the US, with standard board-partner models at $5,090, and TechPowerUp’s summary of that tracker put the US figure at 145 percent above MSRP and the cross-region average at 136 percent. Treat that as a dated range rather than a fixed number. GPU prices rose 15 percent in a single month across the ten countries tracked, so any point price becomes outdated quickly.

The driver is memory supply, not scalpers. Data centers are outbidding consumer hardware for the same DRAM wafers. TrendForce reported that AI could consume nearly 20 percent of global DRAM supply by “equivalent wafer usage” in 2026, led by HBM and GDDR7 demand. The manufacturing multipliers explain why: 1GB of HBM consumes four times the wafer capacity of standard DRAM, and GDDR7 consumes 1.7 times. Wafers are being diverted toward HBM for AI accelerators, which reduces GDDR7 supply. That is a supply-allocation issue, not a shortage of HBM itself.
This matters for local inference because memory capacity, not core count, determines which models you can run at all. A 70B model at Q4_K_M needs about 42GB of weights plus a KV cache, which is why the 5090’s 32GB is the main limitation in this comparison. You can buy a faster card. You cannot buy more than 32GB of GDDR7 in a single consumer Nvidia GPU.
Three Architectures at a Glance
The three machines approach the same problem from different angles. The RTX 5090 is a discrete GPU with 32GB of GDDR7 on a 512-bit bus and 1,792 GB/s of bandwidth, the widest memory pipe of the three, at a 575W board power rating. The M3 Ultra is an Apple system-on-a-chip that Apple says ships with a 32-core CPU, up to an 80-core GPU, and up to 512GB of unified memory at over 800 GB/s, built by joining two M3 Max dies with UltraFusion across 10,000 connections for 184 billion transistors total. Strix Halo is AMD’s option in the middle: the Ryzen AI Max+ 395 pairs 16 Zen 5 cores with 40 RDNA 3.5 compute units and up to 128GB of unified LPDDR5X, of which up to 96GB can be reassigned to the GPU as VRAM.

The spec table conceals the trade-offs. The 5090’s bandwidth is more than twice the M3 Ultra’s, but its capacity is one-sixteenth of the Mac’s maximum. Strix Halo’s bandwidth is the narrowest of the three at a claimed 256 GB/s, but it fits models the 5090 cannot load. Each figure is accurate; none alone determines the best choice.
Throughput by Model Size
Token generation speed depends on memory bandwidth divided by how much of the model is stored in fast memory. The 5090 leads on sizes it can hold entirely. Measured on llama.cpp, a 3B coding model ran at 341.9 tokens per second on the 5090 versus 192.5 on a 60-core M3 Ultra running MLX, and a 20B model hit 318.4 versus 152.7, according to the llm-speed benchmark set. On a 32B model the gap narrows to 67.4 versus 34.5 tokens per second, but the 5090 still roughly doubles the Mac.
That advantage disappears when the model exceeds 32GB. A 70B model at Q4_K_M needs about 42GB of weights, so a single 5090 cannot hold it, and the card must offload layers to system RAM over PCIe. Benchmarks of that offloaded configuration vary widely depending on how much spills: Vucense measured 68 to 75 tokens per second with 24 of 80 layers on the GPU and the rest in system RAM, but a configuration with more spill performs much lower. The variance depends on a configuration detail, not the silicon.
The M3 Ultra and Strix Halo avoid that drop because the whole model stays in unified memory. Vucense’s testing put Strix Halo at 28 to 35 tokens per second on a 70B Q4 model with no offload, running on roughly 102 GB/s of effective memory bandwidth, well below AMD’s 256 GB/s peak claim. It is the slowest of the three on every size, but it never pays the PCIe penalty, which makes a 70B class model usable on a $1,800-class box.
| Platform | Memory | Bandwidth | 70B Q4_K_M behavior | Reported 70B Q4 decode |
|---|---|---|---|---|
| RTX 5090 | 32GB GDDR7 | 1,792 GB/s | Layers offload over PCIe | 68-75 tok/s with partial offload |
| Apple M3 Ultra | Up to 512GB unified | Over 800 GB/s | Fully resident, no offload | About 45-52 tok/s |
| AMD Strix Halo | Up to 128GB unified | About 102 GB/s effective | Fully resident, no offload | 28-35 tok/s |
Sources: bandwidth and capacity from llm-speed, Apple, and AMD; 70B decode figures from Vucense’s 2026 hardware tests. These are individual runs on different runtimes and quantizations, not matched controlled trials, so read them as directional rather than precise.
Effective Cost Per Token
The price increase changes the calculation that matters more than any single benchmark: dollars per token delivered per year. A 5090 build needs an 850W-plus ATX 3.0 power supply and cooling for a 575W card, which pushes a complete system past $5,000 at September 2026 street pricing. A Mac Studio with M3 Ultra costs more up front but runs the same generation at a fraction of the power, and Strix Halo systems cost about $1,800 to $2,000 for a 128GB mini PC.
Power consumption widens the difference. A 575W GPU under continuous inference uses roughly 13.2 kilowatt-hours per day, or about $770 a year at a US residential rate near 16 cents per kilowatt-hour. Apple’s M-series platform draws roughly 25 to 35W under local inference load per PromptQuorum’s 2026 power measurement guide, which amounts to a few hundred dollars a year. Over three years that difference approaches the cost of a second GPU, and it applies to a machine that may be idle between prompts.
Capacity per dollar is where the Mac and Strix Halo pull ahead on large models. LocalIA’s cost analysis put the 5090 at roughly EUR 110 per gigabyte of VRAM against about EUR 24 per gigabyte for a 256GB M3 Ultra, though the Mac delivers roughly a third of the throughput per byte. For a solo researcher running large models, capacity outweighs that trade-off. For a team serving many short requests, throughput per dollar is more important, and the 5090’s 1,792 GB/s pipe is difficult to beat.
The Gotchas: Offload, Prefill, Drivers
The 5090’s limitation is layer offload. Once a model exceeds 32GB, throughput stops being a hardware constant and depends on how many layers you keep resident. Vucense also measured 8 to 12ms of added PCIe latency per token when offloading, which causes interactive chat to feel sluggish even when the overall token rate looks good. If your model fits in 32GB, this never occurs. If it does not, benchmark the exact offload split before you buy.
The M3 Ultra’s limitation is prompt prefill. Unified memory is excellent at streaming generation but slower at processing a long prompt before the first token, and the delay grows with context length. Long-document retrieval or agent loops with large system prompts will experience that delay on every turn. Strix Halo’s limitation is driver maturity: AMD’s ROCm stack has improved but still trails CUDA in tooling depth, so a stack tied to a specific CUDA library will require a port. If your team already runs vLLM on Nvidia, moving to ROCm is a compatibility project, not just a card swap.
One configuration detail deserves a warning. Apple’s M3 Ultra starts at 96GB and Apple’s own UK spec listed configurations up to 256GB as of September 2026, while the launch announcement described up to 512GB, so confirm the exact memory tier you can actually order before buying. The measured machines in the benchmark tables above used 96GB, not 512GB. Historical availability and the machine in a benchmark are different facts.
Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.
# Estimate whether a model fits before buying hardware.
# Weight memory scales with parameter count x bytes per parameter.
# 70B at Q4_K_M is roughly 42GB of weights; add KV cache on top.
# Rule of thumb: 0.56GB per billion parameters at 4-bit quantization,
# plus 4-8GB of KV cache for a 70B model at moderate context.
# Check how much of a model actually fits on a 32GB card before offloading.
llama-server --model ./models/llama-3.3-70b.Q4_K_M.gguf \
--n-gpu-layers 24 \
--ctx-size 8192
# Note: production use should measure the real offload split and context
# length for your workload, not assume the benchmark numbers apply.
# Throughput varies widely with how many layers spill to system RAM.
Which One to Buy
Buy the 5090 if your models fit in 32GB and you serve multiple users. Its bandwidth and CUDA ecosystem make it the best throughput-per-dollar choice for models at the 32B-and-below tier, where it roughly doubles the M3 Ultra in measured decode speed. It also fits into an existing vLLM deployment with no porting work, which our comparison of local inference tools shows matters as much as raw speed. The case against is capacity and price: 32GB limits you, and the current street price means you are paying an AI-memory premium for a gaming-class card.
Buy the M3 Ultra if you need to run models the 5090 cannot hold and you want to save power. It is the only one of the three that reaches 512GB, and its 800-plus GB/s pipe means a large model stays genuinely usable rather than merely loadable. It is also the quietest and most power-efficient option, which matters for a machine running continuously. The cost is that you give up roughly half the decode speed on small models and you cannot add a second GPU later, a trade covered in more depth in our analysis of Apple Silicon AI acceleration.
Buy Strix Halo if budget and model size matter more than speed. A 128GB mini PC at $1,800 to $2,000 runs a 70B model entirely in memory, uses standard upgradable PC parts, and consumes much less power compared to the 5090. It is the slowest of the three on every measurement, and ROCm lags CUDA on tooling, but it is the cheapest way to run a large model without offload. For solo researchers and privacy-bound workloads that fit in 128GB, that trade is often the right choice.
The current price surge reflects memory allocation pressures rather than a permanent shift. TechSpot’s own tracker describes a market responding to memory supply, and the September 2026 figures sit far above the launch MSRP precisely because supply is limited. That said, the memory demand behind it is not a short-term spike: TrendForce’s projections put AI’s share of DRAM wafer capacity near 20 percent for 2026, and wafer capacity grows only 10 to 15 percent a year. A buyer on a fixed budget should compare the whole system cost, including power and cooling, against the model they actually need to run, and pick the platform whose maximum capacity they can live with rather than the one that wins the spec table.
One thing the earlier 70B-focused hardware analysis on this site got right and worth repeating: memory capacity comes first, raw compute second. This comparison adds the price dimension that has changed most since then. The 5090’s street price roughly doubled between mid-2026 and September 2026, while the unified-memory platforms held steadier because their memory is soldered and bought in volume ahead of the spot market. That price difference clearly shows where the value lies for large-model workloads right now.
Related Reading
More in-depth coverage from this blog on closely related topics:
- Best Hardware for Large AI Models
- Choosing the Best Local AI Inference Tools
- Best GPU for Local Large Language Models
- Apple Silicon AI Performance and Acceleration
- AI Inference Cost and Model Size Impact
Sources and References
Sources cited while researching and writing this article:
- GPU Prices Jumped 15% in One Month, RTX 5090 Now 136% Above MSRP | TechPowerUp
- [News] AI Reportedly to Consume 20% of Global DRAM Wafer Capacity in 2026, HBM and GDDR7 Lead Demand
- Apple reveals M3 Ultra, taking Apple silicon to a new extreme – Apple
- AMD Ryzen™ AI MAX+ 395 Processor: Breakthrough AI Performance in Thin and Light
- RTX 5090 vs M3 Ultra for local LLMs: speed vs memory
- Local LLM Hardware 2026: What Actually Runs 70B Models
- Local LLM Power Consumption 2026: RTX 4090 450W = $39/mo
- RTX 5090 vs Mac Studio M3 Ultra for local LLMs , LocalIA
Thomas A. Anderson
Mass-produced in late 2022, upgraded frequently. Has opinions about Kubernetes that he formed in roughly 0.3 seconds. Occasionally flops, but don't we all? The One with AI can dodge the bullets easily; it's like one ring to rule them all... sort of...
