Skip to content

Case study, modelled

The heat that cannot leave

Our limit is cooling, not supply.

Modelled

47 to 50 racks under a design-day ceiling that holds 40 today, wherever pinned runs are shorter than the window the limit is measured over, modelled.

A site whose design-day ceiling holds 40 racks against a requirement of 45 is 360 accelerators short. Where that ceiling is a sustained window, a 15 to 21% reduction at the rack input lifts what fits to 47 to 50 racks wherever pinned runs are shorter than that window, modelled. The energy line runs under either answer, with the cooling-side saving still to add.

Who it is for

A site that owns its accelerators and controls what runs on them, limited by heat rejection on its design day. The head of infrastructure signs, with the facility engineer supplying the conversion from thermal limit to IT load.

GPU cloud and neocloud

The basis

A modelled scenario

This is modelled. The site is illustrative, with no customer behind it, and every rack, megawatt-hour and pound here is arithmetic over declared inputs you can change. One number is ours, and it is measured: up to 21% less GPU die power on NVIDIA H100 NVL over 48 continuous hours. Managed and baseline arms ran under an equal cap, read as NVML die power. Die power is a lower bound on wall power. Hardware other than H100 NVL, this rack included, is modelled until a baseline run on it. A three-week validation on your own fleet, in your design-day condition, makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • Less GPU die power

    measured

    Up to 21%

    NVIDIA H100 NVL, one continuous 48-hour window, managed and baseline arms under an equal cap, NVML die power.

  • Tokens per watt

    derived

    Up to +22%

    Same 48-hour run, Balanced mode, NVML die power against vLLM serving throughput.

  • Throughput change in the run

    derived

    0%

    Same 48-hour run, Balanced mode.

  • Reduction at the rack input

    modelled

    About 15% to 21%

    21% times the 72% share, conservative; 21% carried whole, stated basis.

  • Shortfall on the design day without the runtime

    modelled

    5 racks, 360 accelerators

    45 racks required; 40 held under 5,600 kW at 140 kW per rack.

  • Reduction the requirement needs

    modelled

    About 11%

    One minus 5,600 over 6,300.

  • Racks inside the ceiling, sustained-window answer

    modelled

    47 to 50 (40 without)

    5,600 kW over a managed rack draw of about 119 kW or about 111 kW, whole racks, holding in full once pinned runs are shorter than the window.

  • Accelerators inside the ceiling

    modelled

    3,384 to 3,600 (2,880 without)

    Racks times 72.

  • Racks beyond the requirement

    modelled

    2 to 5

    47 and 50 against 45 required.

  • Headroom under the ceiling

    modelled

    253 to 623 kW

    5,600 kW less the managed requirement draw on each reading.

  • Margin over the required reduction

    modelled

    About 4 to 10 points

    About 15% and 21% against about 11% required.

  • Accelerator share at which the fit still holds

    modelled

    About 53% or above

    About 11% over 21%.

  • Capacity effect, racks

    modelled

    About +18% to +25%

    47 and 50 against 40 unmanaged, on a limit measured over a sustained window.

  • Continuous demand avoided

    modelled

    About 323 to 449 kW

    A mean of about 4,275 kW times the reduction times the 50% busy share.

  • Energy avoided a year, IT side only

    modelled

    About 2,800 to 3,900 MWh

    Demand avoided over 8,760 hours at the rack input. The cooling-side saving adds to it at your plant's coefficient of performance.

  • Money a year at the illustrative tariff

    modelled

    About GBP 570,000 to GBP 790,000

    At an illustrative 20 pence a kilowatt-hour, yours to change.

  • Unplanned IT load on a total runtime failure

    modelled

    980 kW

    47 racks at 140 kW less the 5,600 kW ceiling, conservative reading.

  • Fail-open behaviour

    designed

    Full performance on a component fault

    Runtime design. The mitigation lives at the site.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • IT-load ceiling on the design-day condition

    5,600 kW at the rack input

    Reader-supplied. You convert your thermal limit to an IT-load ceiling.

  • Plane of the envelope

    IT load at the rack input

    Reader-selected, because you performed the conversion.

  • Rack draw at the rack input, nominal

    140 kW

    Rack specification, reader-supplied.

  • Provisioning basis

    Nominal

    Reader-selected. There is no default.

  • Accelerators per rack

    72

    Rack specification.

  • Accelerator share of rack draw

    72%

    Reader-supplied. The largest lever in the model.

  • Upstream cascade factor

    1, no credit

    Default. It gives our own physics argument no credit.

  • Requirement

    6,300 kW

    Reader-supplied.

  • Day shape, busy and changing

    50%

    Reader-supplied, entered for the design-day condition rather than an average day.

  • Day shape, idle or nearly idle

    30%

    Reader-supplied.

  • Day shape, pinned at the power limit

    20%

    Reader-supplied. Higher than most of the set, because a design day is a hot, busy day.

  • Longest continuous pinned run

    90 minutes

    Reader-supplied.

  • Binding limit

    Run twice: as a plant rating, and as a sustained condition over a 60-minute window

    Reader-supplied, no default. On this site it is the input that decides the answer.

  • Mean draw per rack

    95 kW

    Reader-supplied and illustrative. Left blank, the energy block goes.

  • Electricity cost

    20 pence a kilowatt-hour

    Reader-supplied and illustrative.

  • Build stage

    Connected

    Reader-supplied.

  • Coefficient at the die

    21%

    Measured. Fixed, not an input.

The working

The site and its limit

The site can import more power than it can reject the heat from. On a design-day ambient the plant derates, and the hall holds a lower load for a window of hours. ADAPT reduces the accelerator's own draw, and a watt not drawn at the rack is a watt of heat the plant does not reject. The arithmetic runs on the IT load your hall carries at the rack input on the day that binds you. Your facility engineer holds that number, from your plant's derate curve and commissioning data.

Declared for this site: a design-day ceiling of 5,600 kW of IT load at the rack input, a rack of 72 accelerators drawing 140 kW nominal, and a requirement of 6,300 kW. That requirement is 45 racks and 3,240 accelerators. The ceiling holds 40 racks and 2,880 accelerators today, so the site is 5 racks and 360 accelerators short on its design day, modelled. Closing the gap takes about an 11% reduction at the rack input, a figure with no coefficient of ours in it.

The question that decides the answer

Is your binding limit the rating of installed plant, or a sustained condition measured over a window? A chiller, pump, heat exchanger or coolant distribution unit is sized for the peak heat the racks can make. A sustained reduction relieves a plant rating only where that rating is enforced on an averaged basis, and some are. Your facility engineer knows which yours is.

The model runs both answers. Where your plant limit is enforced over a sustained window, here 60 minutes, the full rack arithmetic runs. Where it is an instantaneous rating, the value is the energy line, with the cooling-side saving your facility engineer adds on top. The engineer confirms which applies, and the instrument supplies no default for this input.

The reduction at your rack input

Our coefficient is at the accelerator die and your ceiling is at the rack input. Between them sit the accelerator share of rack draw, declared at 72%, and an upstream cascade factor defaulted to one. Two readings travel together, both modelled. The conservative reading applies the die figure to the accelerator share alone: about 15% at the rack input. The stated basis gives 21%, treating a watt not drawn at the die as a watt not converted, distributed or cooled. That is a physics argument, and the validation measures it.

What fits under the same ceiling

On the 60-minute window, a managed rack draws about 119 kW on the conservative reading and about 111 kW on the stated basis, modelled. The 5,600 kW ceiling then holds 47 to 50 racks, 3,384 to 3,600 accelerators, against 40 unmanaged. The requirement of 45 is met with 2 to 5 racks to spare, and 253 to 623 kW of headroom remains under the ceiling, modelled. The margin over the 11% needed is about 4 to 10 points, and the fit survives at any accelerator share of about 53% or above.

The declared design day is 50% busy and changing, 30% idle and 20% pinned, and the longest pinned run is 90 minutes. In pinned minutes the model takes the fleet at its full 6,300 kW against a 5,600 kW ceiling, with no reduction credited. The 47 to 50 racks hold across the busy-and-changing hours. On this declared design day the longest pinned run outlasts the 60-minute window, so your own longest run is the next figure to pull: under the window, the verdict clears. A longer window or a smaller pinned fraction clears it too.

The energy line, IT side only

The energy line runs under either answer. With a declared mean draw of 95 kW per rack, the 45-rack fleet averages about 4,275 kW at the rack input. The reduction applies to the busy-and-changing half of the day. The continuous demand avoided is about 323 to 449 kW, modelled: about 2,800 to 3,900 MWh a year over 8,760 hours. At an illustrative 20 pence a kilowatt-hour, yours to change, that is about GBP 570,000 to GBP 790,000 a year, modelled.

Those figures stop at the rack input. A watt not drawn at the rack is a watt of heat your plant does not have to reject, and the energy spent rejecting it adds to the line above. Your facility engineer can size that from your plant's coefficient of performance in about an hour. Same accelerators, same tokens, less energy: 0% throughput change in the run and up to +22% tokens per watt, both derived, in Balanced mode.

Running on the strength of the reduction

A site energising racks that fit under its ceiling on the strength of the reduction has made continued operation a condition of the configuration. The runtime is designed to fail open to full performance on a component fault. On the conservative reading, 47 racks unmanaged draw about 6,580 kW against a 5,600 kW ceiling. A total runtime failure therefore puts 980 kW above the ceiling, modelled, arriving over a time your hall's thermal mass sets. The site plan covers it with reserved thermal margin, an agreed curtailment right, or fewer racks energised.

Your facility engineer sets the balance: a curtailment right keeps the rack count, while reserved margin and fewer racks cover it by giving racks back, down to the 40 the ceiling holds today for full cover.

Making the figure yours

The mechanism is duty-cycle intelligence on bursty fleets, so a fleet average is the starting point for your own number. A three-week validation on your racks, measuring the rack plane alongside NVML on your load, settles the modelled half. Here it covers the design-day condition, or is extrapolated to it on terms agreed before it starts. Your facility engineer's hour on the cooling side settles the rest.

What would change it

What moves the result, and how to check it on your own fleet.

The basis of your binding limit moves the answer furthest. Entered as a sustained window, the capacity block returns 47 to 50 racks; entered as an instantaneous plant rating, the value is the energy line, with the cooling-side saving on top. After that, the length of your longest pinned run against the window decides conditional or clean. Ninety minutes against 60 is conditional, and anything under the window is clean. Check both on your own fleet. Ask your facility engineer whether the plant limit is enforced as an instantaneous rating or as a rolling average, and over what window. Then pull the longest continuous run at the power limit from your accelerator telemetry on a design day. Confirm the accelerator share of rack draw while you are there, because the fit holds at about 53% or above.

Your own numbers

The same arithmetic, on your site.

Bring the IT-load ceiling your facility engineer holds for your design day, your longest pinned run, the window your limit is measured over and your fleet. We will run this arithmetic on your numbers and your date.