Skip to content

Case study, modelled

The power that is paid for either way

We are on a committed power contract. We pay for it whether we draw it or not.

Modelled

A 12 MW envelope already paid for holds 6,048 to 6,480 accelerators where it held 5,112, with 432 to 864 beyond the requirement, modelled.

On a committed power contract the bill does not move, so the reduction changes what the contract buys. A 78-rack, 5,616-accelerator requirement that overshoots a 12 MW facility envelope by seven racks fits on both readings, modelled. That leaves 432 to 864 accelerators beyond it and 731 to 1,373 kW of headroom, inside megawatts already paid for.

Who it is for

A compute operator that owns its accelerators, sells the hours and holds its own electricity account on a committed quantity, short of capacity against its requirement. The signature sits with whoever owns the capacity plan and the revenue line, because on this contract the value lands on revenue.

GPU cloud and neocloud

The basis

Modelled, and on what basis

This is modelled. The operator is illustrative, with no customer behind it, and every rack, megawatt-hour and pound here is arithmetic over inputs you can change. One number is ours, and it is measured: up to 21% less GPU die power on NVIDIA H100 NVL over 48 hours. Managed and baseline arms ran under an equal cap, read from NVML die power. Die power is a lower bound on wall power. This rack is hardware other than H100 NVL, so the transfer to it is modelled until a baseline run on it. A three-week validation on your own fleet makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • GPU die power, sustained

    measured

    Up to 21% less

    NVIDIA H100 NVL, 48 hours, six managed against a baseline arm under an equal cap, NVML die power. Re-derived from the raw archive to within 0.4 points.

  • Tokens per watt

    derived

    Up to +22%, Balanced mode

    Same run, NVML die power against vLLM serving throughput.

  • Throughput

    derived

    0% change in the run

    Same run. The same accelerators produce the same tokens for less energy.

  • Reduction at the rack input

    modelled

    About 15% to 21%

    21% at the die times a 72% accelerator share, conservative; 21% carried whole to the rack input, stated basis.

  • Racks inside 10,000 kW

    modelled

    71 without; 84 to 90 with

    Whole racks: 10,000 kW over a managed rack draw of about 119 kW or about 111 kW.

  • Accelerators inside the envelope

    modelled

    5,112 without; 6,048 to 6,480 with

    Racks times 72.

  • Accelerators beyond the 5,616 requirement

    modelled

    432 to 864

    Six to twelve racks beyond a requirement the envelope did not hold unmanaged.

  • Headroom at the rack input

    modelled

    731 to 1,373 kW

    10,000 kW less the requirement's managed draw of about 9,270 kW or about 8,630 kW.

  • More accelerators inside the same envelope

    modelled

    About 18% to 27%

    Rack-count arithmetic, 6,048 or 6,480 against 5,112. Not a tokens-per-watt figure.

  • Energy avoided a year

    modelled

    About 4,400 to 6,100 MWh

    A 7,410 kW fleet mean times the rack-input reduction, credited to the 45% busy band only, over 8,760 hours.

  • Energy line below the commitment

    modelled

    Already paid for: the watts go to the 84 to 90 racks above

    Marginal cost below the commitment entered as zero, so the bill holds and the value lands on revenue.

  • Money at an illustrative blended 20 pence a kilowatt-hour

    modelled

    About GBP 880,000 to GBP 1,200,000 a year

    The invoice average, not the cost avoided on this contract. Where your contract bills the marginal kilowatt-hour, enter that cost from your billing-demand clause and this row carries it.

  • Draw to cover on a total runtime failure

    modelled

    920 kW at the 78-rack requirement; 1,760 to 2,600 kW at 84 to 90 racks

    Unmanaged racks at 140 kW against 10,000 kW: 10,920 kW at 78 racks, 11,760 or 12,600 kW at 84 or 90. Covered by an agreed curtailment right, reserved headroom, or a commitment set below the modelled figure.

  • Full-load hours on this site

    modelled

    5,944, falling by about 404

    64,912 MWh a year over a 10,920 kW peak; about 4,400 MWh removed, with the peak held unchanged in the model.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Contracted envelope

    12 MW, facility plane

    Reader-supplied: the committed quantity in your own agreement. The plane is yours to select.

  • Power usage effectiveness

    1.2

    Reader-supplied, never defaulted. It resolves the envelope to 10,000 kW at the rack input.

  • Rack draw, nominal

    140 kW at the rack input

    Rack specification, reader-supplied.

  • Accelerators per rack

    72

    Rack specification.

  • Accelerator share of rack draw

    72%

    Reader-supplied. The largest lever in the model.

  • Upstream cascade factor

    1, no credit

    Default. It gives our own physics argument nothing.

  • Requirement

    11 MW

    Reader-supplied.

  • Day shape

    45% busy and changing, 40% idle, 15% pinned at the power limit

    Reader-supplied in three bands that total 100. No default and no median fleet.

  • Longest continuous pinned run

    12 minutes

    Reader-supplied.

  • Binding limit

    Contractual, settled over a 30-minute interval

    Reader-supplied, no default. It decides whether pinned minutes set the limit.

  • Mean draw per rack

    95 kW

    Reader-supplied and illustrative. Left blank, the energy line is removed.

  • Electricity cost, two values

    Zero at the margin below the commitment; 20 pence a kilowatt-hour blended, illustrative

    Reader-supplied. The two columns show why the two figures differ.

  • Build stage

    Connected

    Reader-supplied. A reduction inside an envelope already held touches no connection queue.

  • Coefficient at the die

    21%

    The measured figure. Fixed, not an input.

The working

Two take-or-pays

The phrase covers two contracts pointing in opposite directions. On the sell side an operator is paid for committed compute whether the customer uses the hours or not. A reduction in drawn power lands in gross margin. On the buy side it pays for a committed quantity of electricity or network capacity whether it draws it or not. Below that floor the bill is fixed, so the watts the runtime frees become room to sell more compute inside power already bought. This page works the buy side.

The site as declared

The operator holds a 12 MW facility envelope and enters a power usage effectiveness of 1.2, which resolves it to 10,000 kW at the rack input. Its racks draw 140 kW nominal and carry 72 accelerators, which take 72% of rack draw. The requirement is 11 MW: on whole racks, 78 racks, 5,616 accelerators and 10,920 kW. The envelope holds 71 racks today, 5,112 accelerators. The operator is 504 accelerators short, seven racks, and needs about 8% off at the rack input. None of that contains a number of ours.

What fits with the runtime

Our coefficient sits at the accelerator die and the constraint at the rack input, so both readings travel together. Applied to the accelerator share alone, 21% at the die is about 15% at the rack input. Carried whole, on the stated basis, it stays at 21%. A managed rack then draws about 119 kW or about 111 kW, and the same 10,000 kW holds 84 or 90 racks. That is 6,048 to 6,480 accelerators, modelled, against a requirement of 5,616. The requirement fits on both readings, with 432 to 864 accelerators beyond it and 731 to 1,373 kW of headroom.

The margin is about 7 points on the conservative reading and about 13 on the stated basis. That is more than twenty times the measurement uncertainty carried to the rack input, and the fit survives at any accelerator share above about 40%. So the plane conversion holds the answer. The three-week validation on your own racks then measures the coefficient on this rack and on this fleet's day. The extra accelerators are hardware you still buy; the runtime creates the room.

The energy line, computed twice

On a committed contract the megawatts are already bought, so the energy line runs second. Credit goes to the busy-and-changing 45% of the day only; the idle and pinned bands get nothing. At a mean draw of 95 kW a rack the fleet averages 7,410 kW, and the reduction removes about 4,400 to 6,100 MWh a year, modelled. At the marginal cost this operator avoids below its commitment, entered as zero, that is GBP 0 a year. At an illustrative blended 20 pence a kilowatt-hour, the invoice average, it reads about GBP 880,000 to GBP 1,200,000 a year, modelled.

The gap between those columns is your contract, and only one of them is money. The capacity line and the energy line are alternative uses of the same watts and are never added. On this contract the choice is close to made, because the energy use of those watts is worth close to nothing. The value lands on the revenue line, as more work sold inside 12 MW already paid for. That is about 18% to 27% more accelerators, modelled: enough to close the operator's 504-accelerator shortfall, with room for 432 to 864 more as its sales grow.

Charging structures, and the one to check in Germany

Real charging structures put the marginal cost at or near zero across part or all of the range. The Georgia Power large-load schedule floors billing demand at the greatest of the contract minimum, 50% of contract capacity or 500 kW. It carries a ratchet over the preceding eleven months. In Great Britain the distribution capacity charge is levied on the agreed maximum import capacity. It falls only when that agreed figure is reduced, which is a contractual act.

The German network charge regulation rewards high full-load hours, energy divided by peak: sites above ten gigawatt-hours a year earn a reduced charge at 7,000, 7,500 or 8,000 hours. The runtime lowers energy and the model holds the peak, so full-load hours fall. This site sits at 5,944, below every threshold, and keeps its position. A site just above a threshold can cross it: one at 8,100 hours falls below 8,000, and its minimum share of the published charge moves from 10% to 15%. That site runs its own hours through the model, and its energy adviser confirms the position before the rack order.

Three routes for the same watts

A Great Britain reader will ask how the reduction reaches the capacity charge, and the answer is a trade. Route A banks it: reduce the agreed import capacity by the headroom, and the charge falls in proportion. You keep 5,616 accelerators with no room left. Route B fills it: keep the agreed capacity and put 432 more accelerators inside it on the conservative reading. The charge does not move. Route C holds it as reserve against the standing condition below. Choose one. Route A makes the runtime a condition of the connection agreement, so it comes after the mitigation.

The standing condition

Some accelerators fit only because the runtime reduces what they draw. If they are energised, removing the runtime puts the site over its commitment. The runtime fails open to full performance on a component fault, so the accelerators keep working. The site plans for it with one of three measures: an agreed curtailment right, reserved headroom, or a commitment set far enough below the modelled figure to absorb a total failure. Here 84 racks unmanaged draw 11,760 kW against 10,000 kW, an unplanned 1,760 kW, or 2,600 kW on the stated basis, and the plan protects the bill as well as the commitment.

What would change it

What moves the result, and how to check it on your own fleet.

Two fields move this result, and both are yours. The money line turns on the marginal cost you avoid below your commitment. Read the clause that sets your billing demand and your capacity charge, and enter that cost rather than the invoice average. The column that is yours then appears, in an afternoon. The capacity line turns on the accelerator share of rack draw, the largest lever in the model. Read it from one rack's power distribution unit beside the device telemetry over a normal day, and confirm it sits above about 40%. Then the day shape. The model credits the busy-and-changing hours, so your own day shape sets the energy line, and a week of your own telemetry gives the three bands. A pinned run longer than your 30-minute settlement interval turns the verdict conditional. A three-week validation on your own racks, with the rack plane read concurrently with NVML, settles whether the coefficient transfers.

Your own numbers

The same arithmetic, on your site.

Bring your contracted envelope with its plane and power usage effectiveness, your rack specification, your day shape and the date your commitment renews. We will run this arithmetic on your inputs and set up the validation on your own racks.