Decide the workload before you decide the contract

Most GPU procurement conversations start in the wrong place: on a pricing page, arguing about whether a specialist provider at roughly $2.01 per GPU-hour for an H100 PCIe beats a hyperscaler at roughly $12.29 per GPU-hour for an H100 on-demand — both provider list prices as of September 2026. The gap is real. It is also the second question, not the first.

The first question is the shape of your demand curve. A commitment is a claim about future usage, and the price list only tells you what that claim is worth. Two teams paying the same published rate can differ by a factor of two in realised cost, purely because one reserved capacity it could keep busy and the other did not.

So this guide runs in a fixed order, and the order is the point. Classify the workload. Understand what each instrument guarantees. Compute the break-even arithmetic on your own numbers. Size the term to the floor of demand, not its mean. Check what your region will sell you. Then negotiate, and only then sign. Every figure below carries a date, because GPU rates moved sharply in both directions across 2025 and 2026.

Pro tip

GPU capacity is quoted in US dollars almost everywhere, including by Indian and UK resellers, so keep the whole model in one currency and convert once at the finance layer. Mixing rupees, pounds and dollars inside it is how a 6 per cent saving becomes a 3 per cent loss on a bad quarter for the exchange rate.

Step one: classify the workload

Four archetypes cover most AI infrastructure spend. Most teams run several at once, and buy a single instrument for all of them.

Steady inference baseline

A production endpoint that never drops below some floor of concurrent GPUs — the overnight trickle, the weekend minimum, the always-on internal tooling. That floor is the most valuable number in your cost model, because it is the only demand you can commit to with confidence. It is also the most commonly over-estimated: teams look at the mean, reserve to that, then discover the floor was half of it.

Spiky inference

Everything above the floor. A UK retailer whose traffic triples between six and nine in the evening; an Indian fintech whose load follows the payday cycle. Peaks are short, tall and predictable in shape but not in magnitude. This is the demand that punishes long commitments hardest, because capacity committed for the peak sits idle the other twenty-one hours of the day.

Batch and offline work

Embedding generation, nightly re-scoring, evaluation suites, dataset preprocessing. No user is waiting, the deadline is soft, and the work can be interrupted and resumed provided somebody built the harness. Almost every team has more of this than it realises, and is paying on-demand rates for it out of habit.

Training runs

Bursty by project and deadline-bound in a way batch work is not. A fine-tune wants sixteen GPUs for four days and then nothing for six weeks. That shape is a poor match for a multi-year commitment and a good one for short reservations, capacity blocks and interruptible capacity with a checkpoint-resume harness — mechanics covered in the spot and pre-emptible GPU checkpoint-resume playbook.

Map those four onto instruments before you look at any price. The rest of the guide is why.

Workload archetypeDemand shapePrimary instrumentOverflow tierWhat not to do
Steady inference baseline Flat floor, 24×7 Reserved, sized to the floor On-demand above the floor Running the SLA-bearing tier on spot alone
Spiky inference Low floor, tall short peaks On-demand, or serverless per-second Spot, where the peak is degradable Reserving to the peak "for safety"
Batch and offline Deferrable, no user-facing SLA Spot or pre-emptible On-demand backstop with a budget ceiling Paying on-demand rates out of habit
Training runs Project-bursty, high concurrency Short reservation or capacity block Spot with checkpoint-resume for sweeps Justifying a multi-year term with one project
Residency-constrained anything Any shape, pinned to a jurisdiction Whatever is genuinely available in-region In-region reserved, booked early Buying the cheapest quote in a region you cannot legally use
Sustained multi-year platform Flat, predictable, above ~60% utilisation Owned hardware, if you can get power Reserved or on-demand for overflow Buying on a forecast rather than a track record

Step two: what each instrument actually buys you

Here is the distinction that resolves more procurement arguments than any other, routinely missed because the vocabulary is sloppy. A price guarantee and a capacity guarantee are different products, frequently sold under the same word.

A committed-spend discount guarantees a rate. It does not guarantee a GPU will be free when you ask: you promised to spend, and in exchange you get a better price on whatever you manage to acquire. A capacity reservation guarantees hardware is set aside in a named zone for a defined window, and may or may not carry a discount — some cost the same as on-demand and simply remove the risk of being told no. On-demand guarantees neither.

Read the contract for which of the two you are buying, because the failure modes are opposite. A discount without capacity fails on the day you cannot get GPUs at any price. A capacity reservation without a discount fails as a slow bleed on margin. Both are survivable if you knew which you signed.

InstrumentPrice guaranteeCapacity guaranteeTypical discountInterruption riskBest-fit workload
On-demand Published rate while you hold the instance; the list rate itself can change None — you get capacity if it exists when you ask Baseline (0%) None once running Spiky peaks, short experiments, the backstop tier
Reserved / committed Yes, locked for the term Depends entirely on the contract — a spend commitment is not a capacity reservation Typically 30–40% below on-demand None, but you pay whether you use it or not Steady inference baseline, always-on platform services
Spot / pre-emptible No — market-clearing and variable, sometimes capped at the on-demand rate None, and it can be reclaimed with minutes of notice Typically 40–60% below on-demand High and correlated with demand peaks Batch, sweeps, evaluation runs, degradable overflow
Serverless per-second Published per-second rate, no commitment Platform-managed queue; usually no guaranteed concurrency Not comparable — a premium per busy second, zero for idle Cold starts and queueing rather than eviction Low duty-cycle, unpredictable, spiky workloads
Owned hardware Capex fixed at purchase; operating cost still floats Absolute — the cards are yours Not applicable; economics depend wholly on utilisation Your own failures, plus supply-chain and grid risk Sustained, predictable, multi-year, residency-constrained

Serverless does not sit on the same axis as the other four. You are not buying a GPU-hour; you are buying busy seconds and paying nothing for idle — structurally right for a low duty-cycle workload, structurally wrong for one that keeps a GPU saturated. Provider-level trade-offs are set out in the comparison of serverless GPU platforms; what matters here is that it belongs in the menu and is usually left out of it.

Step three: what a GPU-hour cost as of September 2026

The snapshot below is exactly that, with the source type named on every line. A list price and a tracker median are not the same kind of claim: one is a fact about a single seller, the other an observation across many sellers, most of whom you will never buy from.

None of it is a substitute for figures you have pulled yourself, so here is how to reproduce the table for your own shortlist rather than inherit ours. Ask each shortlisted provider for a written quote covering your specific accelerator tier, region and term, and ask explicitly what the rate does and does not include. Then check those quotes against a public GPU price-tracker aggregator, so you can see roughly where each one sits in the current spread rather than judging it in isolation. Then treat every published range, this one included, as a starting point for a negotiation rather than a number you can budget from. A quote you hold in writing is the only figure in this guide you can actually act on, and the only one you can refresh whenever you need to.

GPU / tierOn-demand (USD per GPU-hour)SpotReserved / committedSource type
H100, broad market Roughly $2–$12 across provider tiers; median near $3.37 Typically 40–60% below on-demand Typically 30–40% below on-demand Market tracker
H100 PCIe, one specialist provider About $2.01 Not quoted Not quoted Provider list price
H100 SXM5, same specialist provider About $3.94 About $2.91 Not quoted Provider list price
H100, one-year reserved contract About $2.35 as of March 2026, up from a low near $1.70 in October 2025 Market tracker
H100, AWS Around $12.29 Fell by as much as 88% from Jan 2024 to Sep 2025 in certain regions Provider list price / market analysis
H200 Median about $4.50; cheapest verified in-stock from about $3.00 From about $1.32 Six-month reservations from about $2.84 Market tracker
H200, AWS Capacity Blocks About $4.98 Provider list price
B200 / Blackwell Sources disagree: roughly $7–$10 in one August 2026 survey (Nebius about $7.15, Lambda about $6.99, RunPod from about $8.64); about $2.12–$6.04 on another tracker; Google Cloud a4-highgpu-8g about $16.11 Not reliably quoted 36-month reported as low as about $2.25; a 48-month reported around $3.13 Two market trackers plus provider list — conflicting

One line in that table carries a different date from the rest, and it matters more than the others. The $2.35 one-year reserved rate is a March 2026 contract rate, while the on-demand and spot figures beside it are September 2026. Every break-even number in this guide therefore mixes two dates, which is exactly the kind of quiet error this guide exists to prevent: re-run the arithmetic against a current written quote for your own term before it decides anything.

Watch out

The B200 row is not a range; it is a disagreement, and this guide does not reconcile it. One August 2026 survey put Blackwell at roughly $7 to $10 per GPU-hour across clouds; another tracker reported about $2.12 to $6.04; Google Cloud's a4-highgpu-8g listed at about $16.11 — a spread of more than seven times. The reserved figures conflict in direction too: a 36-month term reported as low as about $2.25 per GPU-hour against a 48-month commitment around $3.13, the wrong way round if longer terms are meant to be cheaper. Do not budget from any published Blackwell range, this one included. Get a written quote for your tier, region and term.

Two structural forces sit behind that mess, and what follows is our reading of the market rather than a published finding. Blackwell availability looked limited through 2026, so most published figures appear to describe scarce, tier-differentiated allocations rather than a liquid market. And the H100-to-H200 and B-series transition created a two-speed market, with H100 oversupply emerging as a risk at exactly the moment Blackwell was still rationed. That is why the same accelerator can list at $2.01 and $12.29 in the same month without either seller being dishonest: different interconnect, tier, support and scarcity.

Step four: the break-even arithmetic

Three formulas do almost all the work. Run them on your own numbers: the constants below are dated, the structure is not.

Reserved versus on-demand

A reservation bills every hour of its term whether the GPU is busy or idle; on-demand bills only the hours you consume. So the comparison is a single ratio — break_even_utilisation = discounted_rate / on_demand_rate — and the reservation wins when the fraction of paid hours you keep the GPU busy exceeds it. At the rates in the snapshot above — $2.35 reserved from March 2026 against $3.94 on-demand from September 2026, with the date mismatch noted there — break-even sits at 59.6 per cent, or about 435 of the 730 hours in an average month.

"""Break-even utilization for a committed (reserved) GPU contract.

Rates in USD per GPU-hour. GPU capacity is quoted in dollars almost
everywhere, including by Indian and UK resellers, so keep one currency
inside the model and convert once at the finance layer.
"""

HOURS_PER_MONTH = 730


def breakeven_utilization(on_demand_rate, committed_rate):
    """Fraction of PAID hours the reserved GPU must actually be busy
    before the commitment beats paying on demand.

    A reservation bills every hour of the term whether you use it or not:
        committed_rate * 1 paid hour   vs   on_demand_rate * u used hours
    """
    if on_demand_rate <= 0:
        raise ValueError("on_demand_rate must be positive")
    if committed_rate >= on_demand_rate:
        raise ValueError("no discount: do not commit")
    return committed_rate / on_demand_rate


def monthly_delta(hours_used, on_demand_rate, committed_rate, reserved_gpus=1):
    """Positive means the reservation saves money at this demand level."""
    on_demand_cost = hours_used * on_demand_rate
    committed_cost = reserved_gpus * HOURS_PER_MONTH * committed_rate
    return on_demand_cost - committed_cost


if __name__ == "__main__":
    # H100 SXM5 on-demand, September 2026; 1-year reserved, March 2026.
    ON_DEMAND = 3.94
    RESERVED_1Y = 2.35

    u = breakeven_utilization(ON_DEMAND, RESERVED_1Y)
    print(f"break-even utilization : {u:.1%}")                     # 59.6%
    print(f"break-even hours/month : {u * HOURS_PER_MONTH:.0f}")   # 435

    for hours in (300, 435, 600, 730):
        d = monthly_delta(hours, ON_DEMAND, RESERVED_1Y)
        verdict = "reserve" if d > 0 else "stay on demand"
        print(f"{hours:>4} h/month -> delta ${d:>8.2f}  {verdict}")

Break-even is therefore a property of the pair of rates, not of your workload; the workload only decides which side of the line you land on. Compute the ratio first, then ask honestly whether you will keep that capacity busy for that fraction of every billed hour across the whole term.

The blended base-plus-burst model

Nobody sensible buys one instrument. The realistic structure is a reserved floor, an on-demand tier above it, and spot underneath the deferrable work. Modelling that needs a demand curve rather than a single utilisation number, and the curve should come from your autoscaler.

"""Blended monthly cost for a base-reserved + burst-on-demand + spot mix.

demand_curve: concurrent GPUs, one entry per hour of the month (730
entries). Export it from your autoscaler or load-balancer metrics.
Do NOT model it as a sine wave -- the whole value of this function is
that it sees the real floor and the real peak.
"""


def blended_monthly_cost(demand_curve,
                         reserved_gpus,
                         reserved_rate,
                         on_demand_rate,
                         spot_rate=None,
                         spot_share=0.0,     # fraction of burst safe on spot
                         spot_waste=0.08):   # re-run overhead: MEASURE it
    hours = len(demand_curve)
    reserved_cost = reserved_gpus * hours * reserved_rate

    burst_gpu_hours = sum(max(0.0, d - reserved_gpus) for d in demand_curve)
    spot_gpu_hours = burst_gpu_hours * spot_share
    on_demand_gpu_hours = burst_gpu_hours - spot_gpu_hours

    on_demand_cost = on_demand_gpu_hours * on_demand_rate
    spot_cost = 0.0
    if spot_gpu_hours and spot_rate is not None:
        spot_cost = spot_gpu_hours * spot_rate * (1.0 + spot_waste)

    idle_reserved = sum(max(0.0, reserved_gpus - d) for d in demand_curve)

    return {
        "reserved_cost": round(reserved_cost, 2),
        "on_demand_cost": round(on_demand_cost, 2),
        "spot_cost": round(spot_cost, 2),
        "total": round(reserved_cost + on_demand_cost + spot_cost, 2),
        "idle_reserved_gpu_hours": round(idle_reserved, 1),
    }


def best_reservation_size(demand_curve, reserved_rate, on_demand_rate, **kw):
    """Sweep reservation sizes and return (cost, gpus) for the cheapest.

    The optimum sits near the FLOOR of the demand curve, not its mean.
    If the answer comes back above your observed floor, check whether
    spot_share is doing more work in the model than it will in reality.
    """
    ceiling = int(max(demand_curve)) + 1
    return min(
        (blended_monthly_cost(demand_curve, n, reserved_rate,
                              on_demand_rate, **kw)["total"], n)
        for n in range(0, ceiling + 1)
    )

Watch the idle_reserved_gpu_hours field. It tells you whether the reservation was sized correctly, and nobody reports it because it appears on no invoice. Idle reserved capacity is the most expensive capacity you own: a discounted rate for zero output.

Owning the hardware

Buying looks seductive when you divide a card price by a lot of hours — until you add the non-GPU costs. A single H100 card listed at roughly $25,000 to $35,000 during 2026 depending on form factor, and that is the card alone, before chassis, network fabric, power distribution, cooling and rack space. The multiplier covering all of that is the most important input in the model, and the only honest way to set it is a quote.

"""Owned-hardware payback against renting the same capacity.

Every price below is an INPUT, not a fact. Card prices, the non-GPU
multiplier and your colocation quote all move independently; refresh
them before this model is allowed to decide anything.
"""

HOURS_PER_YEAR = 8760


def owned_payback_years(gpus,
                        card_price,           # USD per card, ex-chassis
                        non_gpu_multiplier,   # chassis, NICs, switching, PDU,
                                              # cooling, rack -- get a quote
                        rent_rate,            # USD/GPU-hour you would else pay
                        utilization,          # sustained busy fraction, 0..1
                        annual_opex,          # colo power, cooling, space, hands
                        residual_fraction=0.0,  # resale value as a fraction of
                                                # ORIGINAL card cost, at exit
                        hold_years=3):
    gpu_capex = gpus * card_price
    capex = gpu_capex * non_gpu_multiplier
    residual = gpu_capex * residual_fraction

    rent_avoided_per_year = gpus * HOURS_PER_YEAR * utilization * rent_rate
    net_saving_per_year = rent_avoided_per_year - annual_opex
    if net_saving_per_year <= 0:
        return float("inf")      # never pays back at this utilization

    # Residual is credited only at exit, and only if you believe you can
    # actually sell the cards. See the note on the used-card market.
    net_capex = capex - residual
    years = net_capex / net_saving_per_year
    return years if years <= hold_years * 3 else float("inf")


if __name__ == "__main__":
    scenarios = [
        ("100% utilization, no residual",  1.00, 0.00),
        ("60% utilization, no residual",   0.60, 0.00),
        ("60% utilization, 15% residual",  0.60, 0.15),
        ("40% utilization, no residual",   0.40, 0.00),
    ]
    for label, util, residual in scenarios:
        yrs = owned_payback_years(
            gpus=8,
            card_price=30_000,        # mid of the ~$25k-$35k 2026 list band
            non_gpu_multiplier=1.6,   # ILLUSTRATIVE -- replace with a quote
            rent_rate=2.35,           # 1-year reserved H100, March 2026
            utilization=util,
            annual_opex=0.0,          # PLACEHOLDER: insert your colo quote.
                                      # Any real figure pushes payback out.
            residual_fraction=residual,
        )
        print(f"{label:<32} {yrs:5.2f} years (before opex)")

Run it and the numbers are sobering. An eight-card node at $30,000 a card with a 1.6 non-GPU multiplier is $384,000 of capital. Against a one-year reserved rent rate of $2.35 per GPU-hour, keeping all eight busy every hour of the year avoids about $164,700 of rent, so payback lands at 2.33 years — before a single pound or rupee of hosting cost. At 60 per cent sustained utilisation it stretches to 3.89 years; at 40 per cent, 5.83 years.

The rule of thumb most widely quoted for this decision puts the crossover at roughly 60 per cent sustained utilisation held for multiple years, and our arithmetic supports it. It also shows why that is a floor rather than a target: 60 per cent buys a payback just under four years, attractive only if you are confident about year four.

Watch out

Do not buy hardware into a collapsing resale market on the strength of a residual-value assumption. Used H100 cards that traded near $40,000 in late 2023 changed hands for as little as $6,000 by mid-2026 — a fall of roughly 85 per cent. A buy case leaning on selling the fleet at 30 or 40 per cent of cost in year three leans on a number the market has already refuted once. Model residual at zero.

Avoid

Sizing a reservation to your peak because it feels safer. The peak is by definition the demand that occurs least often, so peak-sized reserved capacity spends most of its life idle at a locked-in rate. The worked example below shows a case where it costs 30 per cent more than buying everything on demand, and 89 per cent more than a properly layered mix.

A worked example: steady baseline plus a spiky peak

Take a product team with users in Bengaluru and Manchester. Their inference fleet never drops below four H100s, so that is the floor. A predictable evening peak adds up to eight more for about four hours a day on weekdays, and a nightly embedding and evaluation batch consumes another 400 GPU-hours a month with no user waiting on it.

Monthly demand: 4 GPUs × 730 hours = 2,920 GPU-hours of baseline; 8 GPUs × 4 hours × 22 weekdays = 704 GPU-hours of burst; 400 GPU-hours of batch. Total 4,024 GPU-hours. Rates are the September 2026 specialist-provider list figures: $3.94 on-demand, $2.35 for a one-year reservation, $2.91 spot.

StrategyHow it is builtMonthly costVersus best
A — everything on demand 4,024 GPU-hours at $3.94 $15,854.56 +46%
B — reserve to the peak 12 GPUs reserved 24×7 (8,760 GPU-hours) at $2.35 $20,586.00 +89%
C — layered: reserved floor, on-demand burst, spot batch 2,920 h reserved at $2.35 = $6,862; 704 h on demand at $3.94 = $2,773.76; 400 h spot at $2.91 plus 8% re-run waste = $1,257.12 $10,892.88 winner

Strategy C saves $4,961.68 a month against pure on-demand, about 31 per cent, and $9,693 against the peak-sized reservation. The break-even formula predicted this: strategy B pays for 8,760 GPU-hours to serve 4,024, a utilisation of 45.9 per cent, comfortably below the 59.6 per cent break-even implied by those rates. The reserved floor in strategy C runs at effectively 100 per cent utilisation of the hours it pays for, which is where a commitment earns its discount.

State the assumptions out loud. The floor is measured, not a mean. The batch tier assumes a checkpoint-resume harness exists and that 8 per cent re-run overhead is measured rather than hoped for; without it that line reverts to on-demand, the batch tier costs $1,576.00 rather than $1,257.12, and the saving shrinks by about $319 a month. Note too that this provider's spot quote of $2.91 against $3.94 on-demand is only a 26 per cent discount, well short of the 40 to 60 per cent band spot is generally reported to deliver — precisely why you price your own quote. Nothing here includes storage, egress or the cost of operating three tiers; fold those in the way the guide to LLM unit economics and cost per task does.

Recommended

Build the layered structure even when the saving looks modest, because it is what makes the next decision cheap. Once the floor, the burst tier and the deferrable tier are separated in your deployment and your billing tags, re-pricing any one of them is a config change rather than a migration.

Every article here is written by a Verified Builder. Want your name on the next one?

AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.

Become a Verified Builder →

Term risk, and how to size a commitment

A 36-month lock at 2026 prices, in a market where a generation transition is under way, is not a cost optimisation. It is a directional bet on the price of compute, and deserves to be discussed as one.

The recent evidence cuts both ways. Anyone who locked a one-year H100 reservation near the October 2025 low of about $1.70 per GPU-hour looked clever by March 2026, when the same one-year rate had risen almost 40 per cent to around $2.35. Anyone who locked a long term earlier watched H100 spot prices fall by as much as 88 per cent between January 2024 and September 2025 in certain AWS regions, and the used-card market fall roughly 85 per cent from late 2023 to mid-2026. Prices rose on the reserved instrument and fell on the interruptible one over overlapping windows. Anybody telling you confidently which way the next three years go is guessing.

The response is not to forecast better; it is to size the bet to your confidence, using one rule: commit to the floor of your demand, not its mean, and never to its peak. The floor is the demand you would still have if your most optimistic product bet failed, your largest customer left and growth stopped this quarter — usually much smaller than the number in the plan.

Practically that means a ladder. Put the hard floor on the longest term offered, because it earns the deepest discount on the demand you are most confident about. Put the layer above it on six or twelve months, so it re-prices while the market is still moving. Everything above that goes on-demand, and everything deferrable on spot. If demand grows, the top of the on-demand tier converts into the next reservation at renewal.

One guard belongs in the contract rather than in a wiki page: ask what happens on a generation change — whether the commitment can migrate to newer silicon at an agreed conversion, or whether you are locked to one accelerator for the term. That clause is what turns a 36-month H100 commitment into a liability rather than a discount.

Regional reality for Indian and UK teams

Every number above assumes you can buy the instrument you chose. In practice availability, egress, residency and the electricity grid narrow the menu before price gets a vote — differently in Mumbai than in London.

Start with availability, the constraint people notice last. Our reading of the market through 2026 is that Blackwell allocation stayed limited, and that limited allocation goes by relationship and volume rather than to whoever visits the pricing page first. Smaller teams in Chennai or Leeds should assume the headline B200 rate is not available to them at their volume, and should validate stock before planning around it. H100 is the mirror image: oversupply emerging on the previous generation should mean a genuinely better negotiating position there, which is useful if your workload does not need the newest silicon.

Then egress and latency. A quote from a distant region that is a dollar an hour cheaper stops being cheaper once you count cross-region transfer and the latency your users feel on every request. For an Indian team serving Indian users, an in-region deployment usually wins on landed cost per served request even at a worse sticker rate; the same holds for a UK team serving UK users from London.

Residency is harder still, because it is not a trade-off. If personal data cannot leave a jurisdiction then capacity outside it is not cheap capacity, it is unavailable capacity, and it should be struck from the comparison. UK and EU obligations and India's DPDP framework each pin certain workloads to certain regions: filter the menu first, price second.

Finally, power. Global data-centre electricity demand was estimated to exceed 1,000 TWh in 2026, roughly double the 2023 baseline, and grid-connection waits ran up to seven years in some markets. That is why "we will just buy our own and colocate" is often not a decision your team gets to make on a normal planning horizon — the constraint is a queue at a substation, not a purchase order. The sequencing is covered in the guide to planning AI capacity in grid-constrained regions. For most Indian and UK teams below hyperscale, owned hardware is simply not on the menu.

Before you sign: questions and a refresh procedure

By this point the model has produced a recommendation. What remains is to interrogate the offer, and to build in the assumption that everything you just computed will be stale within two quarters.

The negotiation checklist

Ask all of these before signing, and get the answers in writing rather than on a call. Is this a price commitment, a capacity reservation, or both — and in which specific zones? What is the committed rate against the current list rate, and does the discount survive a list-price change in either direction? What happens to unused commitment: roll over, expire monthly, or expire at term end? Can I grow mid-term at the committed rate? Can the commitment migrate to newer silicon, and on what conversion? What is the interconnect and the per-GPU memory configuration, since "H100" describes at least two meaningfully different products? What are the egress terms? What is the exit path — early termination fee, transfer, or nothing? What availability has this provider delivered against similar commitments, and will they give a reference? And, because it most often produces a discount: what would you need from us to improve this rate?

Re-running this in six months

Every price here carries a date because every price here will eventually be wrong. Put a recurring quarterly entry in the calendar and work through five steps. Re-export the demand curve and recompute the actual floor, because the floor moves as the product does. Re-pull current on-demand, spot and reserved rates for your GPU class from both a market tracker and your provider's list, noting which is which. Recompute the break-even ratio. Re-run the blended model, checking idle_reserved_gpu_hours against the previous quarter. Then compare the realised invoice with what the model predicted, and treat any gap above a few per cent as a bug rather than noise: it is usually where forgotten capacity, unattributed burst or a quietly re-priced tier is hiding. The wider direction of travel is tracked in our reporting on AI inference cost economics.

None of this is glamorous, and that is rather the point. A team that can show a demand curve, a break-even calculation and a commitment sized to the floor has done something most teams never do, and it is unusually legible proof of engineering judgement. If you have run this model on a real fleet, put it on a Builder profile — the people hiring for platform and cost-engineering roles across Bengaluru, Chennai, London and Manchester want exactly that evidence.