Luxor's Compute NewsletterWEEKLY
Weekly coverage of the AI compute market: the news, the numbers, and the infrastructure economics that matter.
August 11, 2026 · 4 min read
Get every issue in your inbox
No spam. Unsubscribe anytime.
Compute Price Pulse
1-yr contract · median $/GPU-hr
| H100 | $2.56 ▼ 3.1% since last week |
Available on Hashrate Index. Powered by Ornn.
The Lead
NVIDIA Vera Rubin: The rack is the new GPU
NVIDIA’s next data center platform, Vera Rubin, has stopped selling you a chip and started selling you a rack. The NVL72 is designed and sold as one machine: 72 Rubin GPUs and 36 Vera CPUs wired to act as a single accelerator, with mandatory liquid cooling and 190–230 kW per rack, roughly five times a Hopper rack. Put simply, the rack is the new GPU.
The twist is inference itself, which now splits across two racks. Rubin GPUs handle the prompt, and racks of Groq’s inference chips handle the token-by-token output. NVIDIA bought a rival (Groq) so they could put its silicon inside their own platform because no single chip wins every job anymore. For anyone buying compute, that means planning around a facility of specialized racks, not a shelf of cards.
The SPOTLIGHT
NVIDIA opens its frontier self-driving model, focused on commercial use
NVIDIA released Alpamayo 2 Super — a top-ranked reasoning model for robotaxis and autonomous trucks — under a permissive license anyone can use commercially. It’s the same playbook that built NVIDIA’s dominance: move into a new domain, then open the model so the whole industry builds on it, which only deepens the reasons to keep buying NVIDIA. And the framing is squarely commercial — automakers, truckmakers, and fleet operators — with no mention of the everyday driver.
LUXOR'S TAKE
Autonomous vehicles need more inference compute, but the edge isn’t cheap power (worth ~3–4 margin points). And with Hopper sold out, it isn’t cheap hardware either — used H100 servers now trade near new. The durable edge is facility capital, above all the grid interconnect, and the ability to energize in months, not years. That’s a ~$50K/GPU greenfield build against ~$42K miner-converted: a ~$9,000 facility gap that decides who survives when rates soften.
AI Hardware availability
B300 and H200 in stock — start the revenue clock sooner
When the neocloud math works, sourcing is the next question. Luxor has B300 and H200 GPU hardware available now.
Around the Industry
This week’s throughline: compute is getting physical. The race is to own it, build it fast, and feed it power.
Morgan Stanley: cheaper AI models won't shrink compute demand
The bank applies the Jevons paradox to AI. When inference gets cheaper, people run far more of it, so total compute keeps climbing instead of falling. Across all three of its market scenarios, NVIDIA and on-site power providers come out ahead, which lands the same point twice: the real ceiling isn't chips, it's power. (Blockspace)
The scramble to build AI compute you can actually own
For site operators
As AI gives compute a physical address, more buyers want to know exactly where their model runs and who can legally reach it, and owning the building beats renting a slice of someone else's. The moat is speed: one operator turned a shuttered industrial building into a live data center in about six months, the exact conversion play miners are positioned to run. (Fast Company)
Inference is splitting into two markets: cheap tokens and premium ones
The chipmaker argues agentic AI has created a distinct tier of premium inference: fast, responsive output that buyers now pay up for. The proof is in the pricing, with providers like MiniMax charging roughly double for high-speed tokens and OpenAI and Anthropic both shipping paid fast modes. Serving that tier profitably means splitting the work between GPUs for the prompt and specialized chips for token-by-token decode, the same disaggregation NVIDIA baked into Vera Rubin. (SambaNova)
From the HRI Vault
Inside the used GPU market: Pricing, players, and the depreciation debate
Residual value is getting productized this week. Here’s the market underneath it: used A100 and H100 pricing, the ITAD channels, and the depreciation fight in full. Read now →
AI Hardware availability
Offtake is becoming a financed product.
Luxor’s compute offtake listings show real-time availability — fill capacity or secure it.