The two deals, side by side
- $45bn with Nscale. Bloomberg, CNBC and TechCrunch reported on 26 August 2026 a six-year commitment for compute capacity at Nscale's Monarch Compute Campus in West Virginia — a 2,250-acre site Nscale acquired in March 2026. About 460 megawatts, expected to start operating in late 2027, drawing on Nvidia's Vera Rubin chips.
- $35bn with Lambda. Bloomberg and the Wall Street Journal reported over 31 August and 1 September that Anthropic had signed on Monday 31 August with Lambda, an Nvidia-backed "neocloud". The capacity sits at Hut 8's Beacon Point campus in Nueces County, Texas — a one-gigawatt facility of which roughly 350 megawatts is the Anthropic portion.
- Nothing here is a press release. Neither Anthropic nor its counterparties have published terms. Every figure below is attributed to reporting, and we have flagged where the reporting is silent rather than filling the gap.
- There is no reconciled total. Other Anthropic compute agreements were reported during 2026, but coverage does not agree on which arrangements belong in a running tally, on their terms, or on a single aggregate figure. No reconciled total is public, so we are not going to publish one.
- The shape beats the size. Two commitments inside a single week, both sited where power and land are available rather than where the customers are, both routed through specialist GPU clouds rather than the hyperscalers that dominated the previous cycle.
| Counterparty | Amount | Term | Power | Site & location | Chips | Status |
|---|---|---|---|---|---|---|
| Nscale | $45bn | Six years | About 460MW | Monarch Compute Campus, West Virginia — 2,250 acres, acquired March 2026 | Nvidia Vera Rubin | Reported 26 Aug 2026; expected to start operating late 2027 |
| Lambda | $35bn | Not stated in reporting | About 350MW of a 1GW facility | Hut 8's Beacon Point campus, Nueces County, Texas | Nvidia-supplied; specific parts not stated in reporting | Signed Monday 31 Aug 2026; reported 31 Aug–1 Sep |
Eighty billion dollars inside a week is a number designed to be quoted, and it will be. But a headline total tells you almost nothing you can act on. What is informative sits one layer down: who owns what, who is exposed to whom, and where the concrete is being poured. Read that layer and both deals stop looking like a spending spree and start looking like a description of the constraints the industry now operates under.
Four seats at the table, and Nvidia is in three of them
Take the Lambda arrangement apart. As reported, four parties are involved and each has a distinct role.
Hut 8 develops the site. It was previously a bitcoin miner that moved into AI data-centre infrastructure — a transition that has become almost routine, because a company that spent years securing cheap power, negotiating with utilities and running sheds full of hot silicon turns out to have exactly the skills an AI campus requires. Nvidia holds the lease on the data centre itself. Lambda provides the compute to Anthropic. And Nvidia has invested in Lambda and supplies it with chips.
Count the seats. Nvidia is the chip supplier, the investor in the operator, and the landlord of the building. Three of the four positions in a single transaction. Anthropic, the customer, holds the fourth.
Why circularity is a fair question, not an accusation
Nothing about that structure is improper, and it is worth being precise rather than insinuating. Vendor financing is old, legal and often sensible — it is how aircraft and industrial plant have been sold for decades — and where the binding constraint is capital rather than demand, a supplier who helps fund the build-out genuinely accelerates deployment. Nvidia taking a lease also solves a real problem: a site developer needs a creditworthy tenant before financing closes.
The legitimate concern is narrower and it is about information, not ethics. When a chip vendor finances, leases to and supplies its own customers, some of the demand signal that reaches the market is partly generated by the party whose sales it validates. An outside observer looking at bookings cannot easily separate demand that would have existed anyway from demand the vendor helped bring into being. That does not make the demand fake. It makes it harder to measure — and forecasts, capacity plans and pricing all sit downstream of that measurement.
Do not read a multi-billion-dollar compute commitment as an independent verification of demand. When the supplier is also an investor in the operator and the landlord of the building, three of the four data points in that transaction come from the same balance sheet. Treat it as one signal, weighted accordingly, and look for corroboration in numbers that are not vendor-adjacent: utilisation rates, spot pricing, grid interconnection queues.
What a six-year, late-2027 commitment says about inference prices
The Nscale deal is the more legible of the two, because reporting gives both a term and a start date. Six years, roughly 460 megawatts, first operating in late 2027. That combination tells you something useful about the direction of unit costs, even without knowing the price per megawatt-hour.
A buyer signs a six-year, fixed-site commitment when it wants certainty of supply more than optionality. A seller accepts one when it needs guaranteed offtake to finance the build. Both are betting that capacity stays scarce enough to be worth locking in — which is not the same as betting that prices per token stay high. The resolution is volume: cost per unit of compute falls with each hardware generation while total spend rises because usage rises faster. That is the pattern we traced in our breakdown of how inference cost economics actually work, and nothing in these two deals contradicts it.
The chip detail matters here too. Vera Rubin is the generation that lands into this capacity, and generation-over-generation improvements in performance per watt are precisely what makes a six-year commitment survivable — if the silicon in 2029 delivers substantially more work per megawatt than the silicon in 2027, a fixed power envelope buys increasing amounts of compute over the life of the contract. Our builder's guide to the Vera Rubin generation covers what that means at the level of an individual workload.
The practical read for anyone budgeting: capacity being committed today starts serving traffic in late 2027 and beyond. Price movements you feel in 2026 are set by the capacity that came online in 2024 and 2025, not by anything announced last week. Plan your 2027 costs against the hardware roadmap, not against this month's headlines.
The two-speed GPU market
While the multi-billion-dollar deals dominate coverage, the market a small team actually transacts in has been moving in two directions at once — and the divergence is the single most useful thing on this page.
| Measure | Earlier | Later | Direction |
|---|---|---|---|
| H100, one-year reserved contract | About $1.70/GPU-hour (October 2025 low) | About $2.35/GPU-hour (March 2026) | Up roughly 40% (Oct 2025 to Mar 2026) |
| H100, on-demand (2026) | Roughly $2/GPU-hour at the low end | Roughly $12/GPU-hour at the top tier | Median near $3.37 across tracked providers |
| Used H100 card, secondary market | About $40,000 (late 2023) | As little as $6,000 (mid-2026) | Down roughly 85% |
Rent something and the price went up by about 40% in five months. Buy the same silicon outright and the asset lost the great majority of its value in under three years. Those facts are not in conflict — they are the same fact seen from two ends. What is scarce is not the chip. It is the chip installed in a building, connected to a grid, cooled, networked and operated. The card is the cheap part.
If you are a team of fewer than fifty people, buying GPUs in 2026 is very hard to justify. You would be taking depreciation risk on an asset that fell from about $40,000 to as little as $6,000 in under three years, in a market where reserved rental rates rose about 40% between October 2025 and March 2026 — meaning you carry the downside of ownership while the upside accrues to whoever operates the building. Buy only where a hard data-residency or security requirement genuinely rules out rented capacity, and price that requirement honestly before you commit.
Neoclouds are now a real line on a procurement sheet
The second consequence is more immediately useful. Nscale and Lambda are neoclouds: providers that sell GPU capacity and very little else — no managed database estate, no identity service, none of the enterprise scaffolding a hyperscaler carries. Until recently they were a niche option a serious buyer would raise an eyebrow at. A pair of commitments at this scale changes that perception, and it should change your shortlist.
For a team in Bengaluru, Pune, Manchester or London, the sensible pattern is a split. Latency-sensitive serving stays near your data in an AWS Mumbai or London region, because moving user-facing inference away from your database is a false economy that you pay for in every request. Training runs, fine-tuning jobs, evaluation sweeps and overnight batch inference are portable, tolerant of a few hundred milliseconds, and are exactly where the price gap against a hyperscaler is widest. Those go to whoever is cheapest this quarter.
Before you sign anything with a neocloud, make the workload portable: containerise the training job, keep weights and datasets in object storage you control, and confirm you can move to a different provider inside a week. Portability is what converts a volatile rental market from a risk into a negotiating position — and it is the difference between benefiting from price competition and being a hostage to it. Then negotiate reserved rates against a real utilisation forecast, not an aspirational one.
Working on inference economics? Compare notes with other verified Builders.
AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.
Become a Verified Builder →Power is the constraint, not silicon
Ask why the capacity is going to West Virginia and Nueces County rather than to Northern Virginia, Dublin or Mumbai, and you get the real story. Both sites were chosen because power and land are available there. Neither is an established technology metro. Both offer the one input that cannot be air-freighted.
The context makes the siting decision obvious. Global data-centre electricity demand is estimated to exceed 1,000 TWh in 2026, roughly double the 2023 baseline, and grid-connection waits reach up to seven years in some markets. Seven years is longer than the six-year term of the Nscale commitment. When the queue to plug in outlasts the contract you are trying to sign, land near a substation with spare capacity becomes the scarce asset, and everything else in the stack arranges itself around it.
This is not an American peculiarity, and it is the part builders in both our markets should be tracking. India's data-centre power demand is on a steep trajectory — we have covered the projected 26GW of grid demand from Indian data centres — which is why the siting of new Indian capacity is decided by state power policy as much as by proximity to Bengaluru or Hyderabad. In the UK, the same logic sits underneath the AI growth zones attached to the £500m sovereign AI fund: those zones exist because planning consent and grid connection, not capital or talent, are the binding constraints on British compute.
So if you are trying to forecast where inference capacity will be cheap and plentiful in 2028, watch interconnection queues and state or regional power policy. They are duller than chip launches and considerably more predictive.
What to do with all this on Monday morning
Four things follow, none of them requiring a balance sheet with a "bn" on it.
Rent, do not buy. The used-card collapse and the reserved-rate rise together say that ownership is the losing side of this trade for anyone without a data centre. Put the capital into people and product instead.
Widen the shortlist, keep the exit. Neoclouds now belong on your procurement sheet alongside AWS, Azure and GCP for portable workloads. Take the price advantage, and keep the migration path short enough that you could actually use it.
Read vendor-adjacent signals sceptically. When the chip supplier is also the investor and the landlord, a large booking is one data point wearing three hats. Corroborate against utilisation and spot pricing before you build a plan on it.
Watch the grid, not just the roadmap. Capacity announced now serves traffic from late 2027. Between here and there, the variable most likely to move your costs is how quickly power gets connected in Maharashtra, Tamil Nadu, Yorkshire and Wales — not which accelerator wins a benchmark this quarter.
Eighty billion dollars is a number that tells you the industry believes compute will stay scarce. The more useful conclusion is subtler: the scarcity has migrated from the chip to the building, from the building to the substation, and from the substation to the planning queue. Everything else in this stack — including your inference bill — is downstream of that.