The break-even is a variable, and it moved

Here is the whole argument in one sentence: the volume at which self-hosting beats buying inference is not a constant you look up, it is the output of a small model whose inputs became genuinely volatile through 2026, and hardware cost inflation moved that output against owning.

This matters because of how the decision is usually made. A team runs the comparison once, in a spreadsheet, during a planning cycle. They reach a conclusion. The conclusion becomes a slide, then an assumption, then a fact that nobody revisits. Eighteen months later the infrastructure reflects a calculation performed against prices that no longer exist. That failure mode is not a maths failure. It is a process failure, and the fix is a process: a model you own, a short list of inputs you re-quote on a schedule, and an explicit set of events that force a re-run.

If you have not yet made the decision at all, start with the static comparison — our guide to self-hosting versus an API for LLM inference lays out the cost components and the hard constraints that can override the arithmetic entirely, and the choice between reserved, spot and on-demand commitment covers which instrument to buy once you have decided to own capacity. This article assumes you have done that work and asks the next question: how do you keep the answer true?

Four things to take from what follows. First, the aggregate memory price figure that appeared in trade coverage through 2026 understates the exposure of small buyers, because the largest buyers are contractually shielded. Second, the cost line that actually moves your crossover is not the headline hardware price but your effective utilisation, which sits in the denominator. Third, the granularity of your commitment — a fraction of a rented GPU-hour versus a whole reserved node — changes the crossover by several multiples on identical inputs. Fourth, every published break-even token volume you have read, including the ones quoted in this article, is a single-source anchor rather than a finding; the only number that governs your decision is the one your own model produces from your own quotes.

Pro tip

Before you read any further, go and find the document where your team recorded its last self-host decision, and write the date of the hardware quote it used at the top of the page. If that date is more than two quarters old, the decision is not wrong yet, but it is unverified. Put a re-run on the calendar before you finish reading — the trigger list later in this guide gives you the conditions.

What happened to memory prices, and why the aggregate figure hides your exposure

Memory sits on the owning side of the ledger. When you buy or rent an accelerator, a substantial share of what you are paying for is memory: the high-bandwidth memory on the package, and the conventional DRAM in the host system that feeds it. That is the mechanical link between a commodity contract price and your break-even. It moves the capital cost of hardware you would own, and therefore your per-token floor. It does not, by itself, move the price of an API token, and this article makes no claim that any laboratory or provider changed published prices in response to memory costs. No such attribution is sourced.

The scale of the movement is worth stating plainly, with dates attached. TrendForce reported conventional DRAM contract prices climbing roughly 90 to 95 per cent quarter-on-quarter in Q1 2026, and forecast a further 58 to 63 per cent quarter-on-quarter increase in Q2 2026, with NAND rising by up to 75 per cent. Then the rate moderated: as of 9 July 2026, TrendForce expected server DRAM contract prices to rise 13 to 18 per cent quarter-on-quarter in 3Q26. Moderating is not the same as reversing, and TrendForce's own words on the following year are worth reading twice: "a server DRAM shortage is already anticipated for 2027." The supply picture behind that is structural rather than cyclical. RDIMM bit supply was initially projected to grow only 15 to 20 per cent year on year, lagging server CPU shipment growth, and HBM production consumes roughly three times the wafer capacity of standard DRAM per gigabyte, so every gigabyte of HBM the industry builds crowds out several gigabytes of conventional supply.

Vendors and analysts told a consistent story. SK Hynix said its HBM, DRAM and NAND capacity was "essentially sold out" for 2026. Samsung raised 32GB DDR5 module pricing to 239 US dollars from 149 US dollars, an increase of about 60 per cent, in September 2025 — worth dating precisely, because it was the leading edge of the move rather than part of it. Gartner forecast DRAM prices to rise 47 per cent across 2026 on significant undersupply; that is a forecast attributed to Gartner, not a measured outcome, and this article does not offer a forecast of its own. Our news coverage of memory being sold out and what it does to the inference cost floor sets out the supply-side detail in full.

Reference point (as reported) Figure Source and date
32GB DDR5 module list price $239, up from $149 (+60%) Samsung, September 2025
Conventional DRAM contract price, Q1 2026 ≈ +90–95% quarter-on-quarter TrendForce, reported for Q1 2026
Conventional DRAM contract price, Q2 2026 ≈ +58–63% QoQ forecast; NAND up to +75% TrendForce forecast for Q2 2026
Server DRAM contract price, 3Q26 ≈ +13–18% QoQ expected — rate moderating TrendForce, 9 July 2026
Full-year 2026 DRAM price direction +47% across 2026 (forecast) Gartner forecast
2026 supplier capacity HBM, DRAM and NAND "essentially sold out" SK Hynix statement
RDIMM bit supply growth ≈ +15–20% year on year (initial projection) TrendForce, initial projection

Figures as reported at the dates shown. Contract prices move; re-read the primary source at trendforce.com before using any of these numbers in a business case.

The two-tier market is the part that concerns you

Now the detail that changes how you should read every one of those percentages. TrendForce observed that "several U.S.-based CSPs have entered into multi-year long-term agreements (LTAs), which restrict suppliers from raising prices for these clients." From Q3 2026, TrendForce expected increases to shift toward customers without long-term agreements and toward supply sold outside those agreements.

Read that as an economics statement and the implication is uncomfortable. If the largest buyers hold price protection, then an industry-average increase is a blended figure across a protected tier and an unprotected tier. The unprotected tier absorbs more than the average, by arithmetic necessity. A startup in Chennai buying two nodes, or a scale-up in Manchester renewing a modest reservation, sits squarely in the unprotected tier. For those teams the aggregate figure is not a conservative estimate of their exposure — it is an optimistic one.

There is a second-order effect worth watching too. Neoclouds and specialist GPU providers sit between the memory market and you, and their own procurement terms determine how much of a supply-side move reaches your invoice. Aggregator estimates put the GPU rental market at roughly 52 billion US dollars in 2026, up from around 35 billion in 2025 — indicative rather than audited, and the capital raised against forward demand — visible in developments such as Nscale's IPO and Anyscale's reported backlog — is what funds the buildout you will eventually rent from. When you re-quote hardware, ask your provider directly whether their pricing is underwritten by a long-term supply agreement, and whether your contract permits a pass-through if theirs does not.

Watch out

Do not model your exposure from a headline industry percentage. Model it from your own supplier's quote, refreshed, with an explicit written answer to two questions: is this price protected by a long-term supply agreement upstream, and does my contract allow a cost pass-through if it is not? A team that assumed the industry average and discovered a pass-through clause mid-term has not mis-modelled the market; it has mis-modelled its own contract.

The break-even model, written out properly

A break-even model is only as honest as its cost lines, and the reason so many self-host business cases mislead is not bad arithmetic but missing rows. It helps to separate the two shapes that costs take, because they behave completely differently as volume changes.

The two shapes of cost

Variable cost per token is what each additional million tokens costs you once the machinery exists. On the API side it is simply the published price. On the self-host side it is the fully-loaded monthly cost of a GPU divided by the tokens that GPU actually serves in a month — which is where utilisation enters, in the denominator, and why it dominates everything else.

Fixed monthly cost is what you pay whether you serve one token or a billion. On the API side this is close to zero, which is the tier's structural advantage. On the self-host side it is engineer time, facility charges, monitoring, and the standby capacity you keep so a single node failure is not an incident.

The crossover is then the volume at which the API's higher variable cost has consumed the self-host option's fixed cost disadvantage. Written out: crossover volume equals fixed monthly self-host cost divided by the difference between the API price per million tokens and your self-host floor per million tokens. If that difference is zero or negative — if your floor is above the API price — there is no crossover at any volume, and the honest output of the model is the phrase "no crossover" rather than a very large number.

The lines people forget

Published analyses converge on hidden costs adding roughly 20 to 40 per cent on top of raw hardware amortisation, covering electricity, setup time, model management overhead and the absence of a service-level agreement. That range is a reasonable starting multiplier, but treat it as a placeholder to be replaced with your own itemised figures rather than a constant of nature. The table below is the row list; the point of it is the third column.

Cost line Which side it sits on Who typically forgets it
Published price per million tokens Buy Nobody — this is the one line everyone has
Hardware capital cost or hourly rental Own Nobody, but it is quoted stale far more often than it is quoted wrong
Effective utilisation (served ÷ peak-sized capacity) Own Almost everyone — assumed at 60–70%, observed far lower
Electricity and cooling Own Teams renting by the hour, correctly; teams colocating, incorrectly
Colocation, rack space, import and landed costs Own Anyone modelling on-prem from a cloud price list. Differs materially between India and the UK
Steady-state engineer hours (not the build) Own Founders and technical leads, systematically
Model management overhead — weights, upgrades, rollbacks Own Teams who have only ever run one model version
Absence of a service-level agreement Own Everyone, because it is priced as risk rather than as cash
Standby and redundancy capacity Own Teams whose first outage has not happened yet
Rate limits, quota negotiation, vendor management Buy Teams comparing against a self-host option with no ops line
Egress and data transfer Both Teams whose model and application sit in different regions

The calculator

Below is the model as runnable code. Paste it into a file and edit the INPUTS dictionary; every value in it is a placeholder, and the comments mark which inputs became volatile through 2026 and therefore need re-quoting quarterly rather than annually. The illustrative defaults are chosen to demonstrate the mechanics, not to describe your workload or anyone's market.

Two functions matter. elastic_crossover_tokens_per_day assumes you can buy capacity in arbitrarily small units — hourly rental, or a serverless platform. lumpy_crossover_tokens_per_day assumes you must commit whole GPUs for whole months, which is what a reservation or a purchase actually looks like. The gap between the two answers, on identical inputs, is one of the most useful things this model tells you.

"""Self-host vs API break-even: a model you re-run, not a number you quote.

Every value in INPUTS is a placeholder. Replace each one with a figure you
obtained yourself. Lines marked VOLATILE are the ones that moved through
2026 -- re-quote those every quarter, not every year.
"""

HOURS_PER_MONTH   = 730                      # 8760 / 12
SECONDS_PER_MONTH = HOURS_PER_MONTH * 3600
DAYS_PER_MONTH    = 30

INPUTS = {
    # --- what you would pay someone else, per million tokens, blended in+out
    "api_usd_per_mtok":       9.00,   # VOLATILE: re-read the provider's price page
    # --- what the hardware costs you
    "gpu_usd_per_hour":       3.20,   # VOLATILE: rental, or amortised purchase / 730
    "hidden_overhead_pct":    0.30,   # 0.20-0.40: power, setup, model management, no SLA
    # --- what the hardware gives you back
    "gpu_tokens_per_second": 2200,    # measure this on YOUR model, YOUR batch size
    "utilisation":            0.35,   # VOLATILE: served tokens / peak-sized capacity
    # --- what the humans cost
    "engineer_hours":        15.0,    # per month, steady-state, not the build
    "engineer_usd_per_hour": 30.0,    # your loaded internal rate
    # --- what the building costs (differs materially IN vs UK -- quote locally)
    "facility_usd_per_month": 0.0,    # VOLATILE: power, colo, duties. 0 for pure rental
    # --- how lumpy your commitment is
    "min_gpus":              1,       # smallest unit you can actually buy or reserve
}


def gpu_month_cost(gpu_usd_per_hour, hidden_overhead_pct):
    """Fully-loaded cost of one GPU for one month."""
    return gpu_usd_per_hour * HOURS_PER_MONTH * (1.0 + hidden_overhead_pct)


def gpu_month_tokens(gpu_tokens_per_second):
    """Tokens one GPU could serve in a month at 100% utilisation."""
    return gpu_tokens_per_second * SECONDS_PER_MONTH


def self_host_usd_per_mtok(cfg):
    """Your marginal self-host cost floor. Utilisation lives in the denominator."""
    per_token = gpu_month_cost(cfg["gpu_usd_per_hour"], cfg["hidden_overhead_pct"]) / (
        gpu_month_tokens(cfg["gpu_tokens_per_second"]) * cfg["utilisation"])
    return per_token * 1_000_000


def fixed_monthly(cfg):
    """Costs you pay whether you serve one token or a billion."""
    return cfg["engineer_hours"] * cfg["engineer_usd_per_hour"] + cfg["facility_usd_per_month"]


def elastic_crossover_tokens_per_day(cfg):
    """Crossover when capacity is divisible -- hourly rental or serverless."""
    margin = cfg["api_usd_per_mtok"] - self_host_usd_per_mtok(cfg)
    if margin <= 0:
        return None                      # no volume ever makes owning cheaper
    return fixed_monthly(cfg) / margin * 1_000_000 / DAYS_PER_MONTH


def lumpy_crossover_tokens_per_day(cfg):
    """Crossover when you must commit whole GPUs for a whole month."""
    committed = cfg["min_gpus"] * gpu_month_cost(
        cfg["gpu_usd_per_hour"], cfg["hidden_overhead_pct"]) + fixed_monthly(cfg)
    tokens = committed / cfg["api_usd_per_mtok"] * 1_000_000
    capacity = cfg["min_gpus"] * gpu_month_tokens(cfg["gpu_tokens_per_second"]) * cfg["utilisation"]
    if tokens > capacity:
        return None                      # the committed node cannot serve enough to pay for itself
    return tokens / DAYS_PER_MONTH


def describe(cfg, label="base case"):
    fmt = lambda v: "no crossover" if v is None else f"{v / 1e6:,.2f} M tokens/day"
    print(f"--- {label}")
    print(f"  self-host floor        : ${self_host_usd_per_mtok(cfg):,.2f} per million tokens")
    print(f"  API price              : ${cfg['api_usd_per_mtok']:,.2f} per million tokens")
    print(f"  fixed monthly cost     : ${fixed_monthly(cfg):,.2f}")
    print(f"  crossover, elastic     : {fmt(elastic_crossover_tokens_per_day(cfg))}")
    print(f"  crossover, whole-GPU   : {fmt(lumpy_crossover_tokens_per_day(cfg))}")


def variant(cfg, **overrides):
    new = dict(cfg)
    new.update(overrides)
    return new


if __name__ == "__main__":
    describe(INPUTS, "Chennai team, hourly rental, frontier API")
    describe(variant(INPUTS, engineer_usd_per_hour=95.0, gpu_usd_per_hour=4.10),
             "London team, hourly rental, frontier API")
    describe(variant(INPUTS, api_usd_per_mtok=0.50),
             "Chennai team vs a budget open-weight API")
    describe(variant(INPUTS, utilisation=0.05),
             "Chennai team, bursty traffic (5% utilisation)")

    print("--- sensitivity: elastic crossover, utilisation x GPU price")
    for util in (0.15, 0.35, 0.60):
        cells = []
        for price in (2.89, 4.50, 11.06):
            c = elastic_crossover_tokens_per_day(
                variant(INPUTS, utilisation=util, gpu_usd_per_hour=price,
                        engineer_usd_per_hour=55.0))
            cells.append("no crossover" if c is None else f"{c / 1e6:.2f} M/day")
        print(f"  utilisation {util:.0%}: " + " | ".join(cells))

The gpu_tokens_per_second input deserves a warning of its own. Do not take it from a vendor benchmark or from an article, including this one. Measure it on your model, at your batch size, with your prompt-length distribution, and re-measure it after any serving-stack upgrade. Our guide to cutting self-hosted serving cost with quantisation and batching covers how to move that number, and moving it is the one lever that improves your floor without any change in what you pay.

Every article here is written by a Verified Builder. Want your name on the next one?

AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.

Become a Verified Builder →

Two teams, one workload, two crossovers

To show the model working, here are two teams serving an identical workload on identical software. One is in Chennai, deploying into AWS Mumbai (ap-south-1); the other is in London, deploying into AWS London (eu-west-2). Every input below is illustrative — chosen to demonstrate the mechanics, not presented as market data for either country. The hardware rates sit inside a published span for H100 80GB rental of roughly 2.89 to 11.06 US dollars per hour across providers, with A100 80GB reported at roughly 1.09 to 5.07; the engineer rates are placeholders you must replace with your own loaded internal cost.

Illustrative input Chennai team (AWS Mumbai) London team (AWS London)
Blended frontier API price $9.00 per million tokens $9.00 per million tokens
GPU rental (illustrative, within published span) $3.20 per GPU-hour $4.10 per GPU-hour
Hidden overhead multiplier +30% +30%
Measured throughput (placeholder — measure your own) 2,200 tokens/sec/GPU 2,200 tokens/sec/GPU
Effective utilisation 35% 35%
Steady-state engineer time 15 h/month at $30/h (illustrative) 15 h/month at $95/h (illustrative)
Self-host floor (model output) $1.50 per million tokens $1.92 per million tokens
Crossover, elastic (hourly rental) ≈ 2.00 M tokens/day ≈ 6.71 M tokens/day
Crossover, whole reserved GPU ≈ 12.91 M tokens/day ≈ 19.69 M tokens/day

All inputs illustrative. Outputs computed by the calculator above from those inputs — they describe the model, not the market.

Three things fall out of that table. The first is that identical technology produces crossovers a factor of three apart purely because the fixed-cost line differs. The Chennai team's lower engineer cost means it clears its fixed disadvantage sooner, so it can rationally own at a volume where the London team should still buy. Neither answer is more correct; they are answers to different questions, and a team that borrowed the other's conclusion would be wrong in both directions.

The second is that both teams face the same underlying floor problem for the same reason. The floor is driven by hardware cost divided by utilisation, and hardware cost is the line that memory inflation moved. Re-quote the GPU rate upward and both floors rise together, both crossovers move outward, and the volume band in which owning makes sense narrows from the bottom.

The third, and the one most teams have never modelled, is the gap between the elastic and whole-GPU rows. On identical inputs the Chennai team's crossover is roughly 2.0 million tokens per day if it can buy capacity by the hour, and roughly 12.9 million if it must commit a whole node for a whole month. That is a six-fold difference produced entirely by commitment granularity. It is also the cleanest argument for reading the commitment-model guide alongside this one: the crossover you should use depends on the instrument you are actually able to buy.

Avoid

The naive calculation: take last month's API invoice, divide by an hourly GPU price found in a blog post, observe that the GPU looks cheaper, and conclude that self-hosting saves money. This has three defects. It compares a fully-loaded managed price against a bare compute price. It implicitly assumes the GPU is busy whenever it is paid for. And it uses a hardware price whose quote date is unknown — which, after the 2026 contract-price moves, is the single most likely source of error in the whole exercise.

Recommended

The sound calculation: take your own measured tokens per day from production logs, your own supplier's dated written quote, your own measured throughput at your own batch size, your own observed utilisation over a full week including nights and weekends, and your own loaded engineer rate. Run both the elastic and the whole-GPU crossover. Record the quote date next to the answer, and record the conditions that would invalidate it. A break-even with a quote date and an expiry condition is a decision; one without them is a number.

Which input moves the crossover most

Sensitivity analysis is the part of this exercise that repays the most effort, because it tells you which input to spend your time getting right. Run the calculator with one input varied at a time and a clear ordering emerges.

Effective utilisation dominates. It sits in the denominator of the per-token floor, so the relationship is inverse rather than linear: halving utilisation roughly doubles your floor. More importantly, it can flip the answer from "self-host at this volume" to "no crossover at any volume". With the illustrative Chennai inputs, dropping utilisation from 35 per cent to 5 per cent lifts the floor from roughly 1.50 to roughly 10.51 US dollars per million tokens — above the 9.00 US dollar API price — at which point no volume makes owning cheaper. The model returns "no crossover", and that is the correct output rather than an error.

Hardware capital cost comes second. It scales the floor linearly, and it is the input that memory inflation moved. It is also the input you can actually re-quote, which is what makes the re-forecasting discipline worth having.

Token volume comes third. Volume determines whether you clear your fixed costs; it cannot rescue a floor that is above the API price. This is the ordering most teams get backwards, because volume is the input they think about first.

The grid below shows the elastic crossover under low, expected and high assumptions on the two dominant inputs, holding the illustrative fixed cost at 15 engineer-hours a month at 55 US dollars an hour. Read the top-right cell carefully.

Effective utilisation GPU at $2.89/h (low) GPU at $4.50/h (expected) GPU at $11.06/h (high)
15% — bursty, peak-sized cluster 4.71 M tokens/day 6.75 M tokens/day No crossover at any volume
35% — realistic for steady product traffic 3.60 M tokens/day 3.99 M tokens/day 7.21 M tokens/day
60% — batch-heavy, well-packed 3.35 M tokens/day 3.54 M tokens/day 4.60 M tokens/day

Model outputs from illustrative inputs, computed by the calculator above. Not market data. Substitute your own quotes and re-run.

For calibration rather than for citation, it is worth knowing where published estimates land. Against frontier APIs, self-hosting break-even is variously placed at roughly 2 to 5 million tokens per day, or around 256 million tokens per month assuming 60 to 70 per cent sustained GPU utilisation. Against budget open-weight APIs priced around 0.14 to 0.50 US dollars per million tokens, one 2026 analysis places break-even for a single H100 near 5.7 billion tokens per month — a level most teams never reach. Every one of those is a single-source trade or aggregator figure, and the utilisation assumption embedded in the second one is exactly the assumption this section warns about. Trade reporting suggests real traffic is bursty enough that effective utilisation sits far below the 60 to 70 per cent used in back-of-envelope models, which pushes the true break-even higher than those anchors imply. Use them to sanity-check the order of magnitude your own model produces, and nothing more.

Watch out

High utilisation assumptions are the single biggest source of self-hosting regret. Teams model 60 to 70 per cent because that is the figure in the reference material, then discover that the cluster they sized for the Tuesday-afternoon peak is nearly idle at 03:00, at weekends, and during the three months before the product finds its traffic. Measure utilisation over a full week before you commit, treat the observed figure as the expected case rather than the pessimistic one, and run the "no crossover" scenario explicitly so you know what a bad month looks like.

The middle tier most teams skip

The framing throughout this article has been two-sided, and that is a simplification worth correcting before you act on it. There are three tiers, not two, and the middle one is where a great many teams belong without knowing it. Managed open-weight providers — Together.ai, Fireworks and similar — host open models and bill per use. They occupy a band that undercuts both proprietary frontier APIs and do-it-yourself self-hosting at light and medium volumes, because you get open weights and per-token billing without paying for idle hardware.

The reason this belongs in a re-forecasting guide rather than a decision guide is that it changes what your model is comparing against. If your api_usd_per_mtok input is a frontier price, your crossover is the volume at which owning beats the frontier — which is a question you may not need answered, because the managed open-weight tier probably beats both at your current volume. Re-run the model with the managed tier's price in that field and the crossover moves outward sharply, often past any volume you will reach. That is the honest answer, and it is only visible if the middle tier is in the comparison at all.

Tier Cost shape Operational load Where it wins
Frontier proprietary API Highest per token, no fixed cost, no idle Lowest — vendor management and quota negotiation only Light volume; workloads that need frontier capability; anything pre-product-market-fit
Managed open-weight API Middle per token, no fixed cost, no idle Low — model selection and version pinning Light to medium volume on open weights; bursty traffic; teams without platform engineers
Self-host (own or reserve GPUs) Lowest per token in principle, high fixed cost, you pay for idle Highest — serving stack, drivers, capacity, on-call, upgrades High and steady volume; hard residency or latency constraints; owned fine-tuned weights

If your traffic is bursty rather than steady, the serverless platforms are the practical route into the middle tier without a capacity commitment at all — our comparison of Modal, RunPod, Baseten and Replicate covers how they price and where each fits. And whichever tier you land in, the accounting question that follows is what a task actually costs you rather than what a token costs, which is why a re-forecast is only as useful as the volume figure feeding it. Our guide to forecasting the bill before launch covers building the three-scenario volume model that supplies the tokens per day input this calculator depends on.

When to re-run the model

The discipline this article is arguing for is small: a model in version control, and a written list of conditions that force it to be re-run. Calendar reviews alone are too slow, because 2026 demonstrated that a key input can move by most of a doubling inside a single quarter. Event triggers alone are too easy to ignore. Use both.

Re-run the full model when any of the following happens.

  • A hardware quote is refreshed or expires. Every written quote should carry a validity window; when it lapses, the model output lapses with it. This is the trigger that would have caught the 2026 move.
  • A supplier or cloud contract comes up for renewal. Renewal is the moment your protected or unprotected status in the two-tier market is actually decided.
  • Your served token volume moves by roughly a factor of two, in either direction. Growth is the obvious case. Decline matters more, because a committed cluster sized for last quarter's traffic silently destroys your utilisation and therefore your floor.
  • A provider changes published prices, on either side. A frontier price cut moves the crossover outward; a managed open-weight price change moves which tier you belong in.
  • A cost pass-through clause is invoked against you, or you learn that your provider's own supply is not covered by a long-term agreement. Given TrendForce's expectation that from Q3 2026 increases would shift toward customers without such agreements, this trigger deserves a named owner.
  • You upgrade the serving stack, change quantisation, or change batch configuration. Any of these moves measured throughput, which moves the floor without any price changing at all.
  • Effective utilisation drifts by more than about ten percentage points from the figure in the model. Instrument this so the drift is visible rather than discovered.
  • The annual budget cycle, unconditionally, even if none of the above fired.

Between full re-runs, refresh only the two volatile inputs — hardware cost and effective utilisation — quarterly, and record the date next to each. A model whose inputs each carry a quote date is auditable; one whose inputs are undated is folklore.

One more piece of housekeeping that pays for itself: keep the "no crossover" scenario in your reporting rather than deleting it. It is the scenario in which your committed hardware can never pay for itself at any volume, and knowing the conditions that produce it — low utilisation, an expensive quote, a cheap alternative tier — is what lets you spot it approaching instead of discovering it in a year-end review.

What to change in your contracts now

The re-forecasting discipline only helps if your contracts leave you room to act on what the model says. Four changes are worth pushing for, and none of them requires unusual leverage.

Cost pass-through caps. If a supplier contract permits component-cost pass-through, negotiate a cap on the magnitude and a floor on the notice period. An uncapped pass-through in a market where a supplier has told the public its capacity is essentially sold out is an open-ended liability, and the two-tier structure TrendForce described means smaller buyers are the ones most likely to see it exercised.

Quote validity windows, in writing. Ask for an explicit date until which a quote holds. This is not only commercial protection; it is what makes your break-even model auditable, because it gives every hardware input an expiry you can put in the trigger list.

Exit and downsizing terms. Model what happens if your volume halves. If a commitment cannot be reduced, transferred or exited without forfeiting the balance, then your downside case is worse than the model's "no crossover" scenario suggests, because you keep paying the fixed cost while the floor deteriorates. Prefer shorter terms and smaller units while a key input is volatile, and accept a somewhat worse headline rate for the optionality.

A named owner for the trigger list. Contracts do not re-forecast themselves. Give one person responsibility for the model, the quote dates and the triggers, and review it in the same meeting where the budget is reviewed.

None of this is exotic procurement. It is the ordinary consequence of accepting that the break-even is a variable. The teams that came through 2026 without an unpleasant surprise were not the ones that predicted memory prices — nobody in this article predicts memory prices, and you should be sceptical of anyone who does. They were the ones whose decision carried a quote date and an expiry condition, so that when the input moved they knew which conclusion had just stopped being true.

If you are the engineer who built that model, kept it current, and can show the decision it drove, that is unusually legible proof-of-work — the kind that reads clearly to the people hiring for platform and cost-engineering roles across India and the UK. A Verified Builder profile is where you put it.