Skip to content

Case study, modelled

Headroom put back to work

We would rather grow than save.

Modelled

48 racks become 54 inside the same 6,800 kW: 432 more accelerators selling compute, and each unit of work takes about 8 to 12 per cent less energy, modelled.

A 6,800 kW site holding 48 racks of 72 accelerators fits 54 racks with the runtime on, and 57 to 61 at the ceiling, modelled. The six added racks are 432 accelerators the operator buys. One rack in eight more runs on about 3 per cent more energy on the conservative reading, and on roughly the same energy on the stated basis, modelled. Energy per unit of work falls about 8 to 12 per cent.

Who it is for

An operator that owns its accelerators, sells compute from an envelope it cannot enlarge and has demand it cannot serve. The capacity planner makes the case and the CFO signs, because the decision is a hardware purchase before it is a software licence.

GPU cloud and neocloud

The basis

What this page is

This is modelled. The operator is illustrative, with no customer behind it, built from a rack specification, a power envelope and a day shape. Every rack, megawatt-hour and pound below is arithmetic on inputs you can change. One number is ours, and it is measured: up to 21% less GPU die power on NVIDIA H100 NVL over 48 hours. Managed and baseline arms ran under an equal cap, read as NVML die power. Die power is a lower bound on wall power. Other hardware is modelled until a baseline run on it. A three-week validation on your own fleet makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • GPU die power reduction

    measured

    Up to 21% less

    NVIDIA H100 NVL, one continuous 48-hour window, managed and baseline arms under an equal cap, NVML die power. A lower bound on wall power.

  • Tokens per watt

    derived

    Up to +22%

    Balanced mode, NVML die power against vLLM serving throughput on one model and serving stack, same 48-hour run.

  • Throughput change in the run

    derived

    0%

    Same 48-hour run, Balanced mode, throughput held.

  • Reduction at the rack input

    modelled

    About 15 per cent conservative to 21 per cent stated basis

    The die figure times a 72 per cent accelerator share times a cascade factor of one; the stated basis carries the die figure through.

  • Reduction the target needs

    modelled

    About 10 per cent

    54 racks at 140 kW draw 7,560 kW against 6,800 kW.

  • Margin over the required reduction at the target

    modelled

    About 5 points conservative to about 11 points stated basis

    The target survives at any accelerator share of about 48 per cent or more.

  • Accelerators added at the target

    modelled

    432, in 6 racks

    Arithmetic on your own rack figures. Hardware the operator buys.

  • Racks the envelope holds with the runtime

    modelled

    57 conservative to 61 stated basis

    4,104 to 4,392 accelerators, within the installed fleet, with no cooling change and no rack change.

  • Uplift on the installed fleet

    modelled

    One rack in eight at the target; about 19 to 27 per cent at the ceiling

    Rack-count arithmetic, 54, 57 and 61 against 48.

  • Margin at the ceiling

    modelled

    Under one point on either reading

    A third of a point at 57 racks conservative, so the target is set at 54.

  • Per-rack reduction on the energy line

    modelled

    About 8 per cent conservative to about 12 per cent stated basis

    The rack-input reduction times the 55 per cent busy share. Idle and pinned hours credited at zero.

  • Managed mean draw per rack

    modelled

    About 87 kW conservative to about 84 kW stated basis

    From the illustrative 95 kW mean.

  • Site consumption a year before absorption

    modelled

    About 40,000 MWh

    48 unmanaged racks at 95 kW mean over 8,760 hours.

  • Site consumption a year after absorption

    modelled

    About 41,200 MWh conservative to about 39,700 MWh stated basis

    54 managed racks over 8,760 hours.

  • Change in site consumption a year

    modelled

    About 1,300 MWh more conservative; roughly flat stated basis

    432 more accelerators drawing power: the outcome the operator chose.

  • Net change in the electricity bill, 54 racks against 48

    modelled

    About GBP 250,000 a year more conservative; roughly flat on the stated basis

    At an illustrative 20 pence a kilowatt-hour: 432 more accelerators powered, less the saving on the 48 racks already there. Set it against the gross margin those 432 accelerators sell.

  • Energy per rack-year

    modelled

    832 MWh before; 763 MWh conservative to 736 MWh stated basis after

    Site consumption over racks energised.

  • Energy per unit of work

    modelled

    About 8 to 12 per cent lower

    The per-rack reduction the day shape supports, seen from the other end.

  • Banking alternative at 48 racks

    modelled

    About 3,300 MWh and about GBP 660,000 a year, conservative

    Mutually exclusive with absorption; no accelerator added.

  • Draw to cover on a total runtime failure

    modelled

    760 kW over the envelope at 54 racks; 1,180 kW at 57; 1,740 kW at 61

    Unmanaged draw of the energised racks against 6,800 kW. Covered by an agreed curtailment right, reserved headroom, or fewer racks energised.

  • Worst-hour check

    modelled

    Pinned runs of 6 minutes against a 30-minute interval; the fit stands

    The envelope does not hold during a pinned run, and the runs are shorter than the settlement interval.

  • Same logic at ten times the scale

    modelled

    68 MW holds 485 racks, absorbs to 540, adds 3,960 accelerators, 7,600 kW to cover on a total failure

    Racks round at each scale separately; the ratios hold and the counts do not.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Power envelope

    6,800 kW, IT load at the rack input

    Reader-supplied; held flat by choice as well as by circumstance.

  • Rack draw, nominal

    140 kW at the rack input

    Rack specification; provisioning basis nominal.

  • Accelerators per rack

    72

    Rack specification.

  • Accelerator share of rack draw

    72 per cent

    Reader-supplied. The largest lever in the model.

  • Upstream cascade factor

    1, no credit

    Default. No credit for our own physics argument.

  • Installed fleet today

    48 racks, 3,456 accelerators

    Reader-supplied: what is energised now.

  • Absorption target

    54 racks, 3,888 accelerators

    Reader-supplied, chosen with margin rather than at the edge.

  • Day shape

    55 per cent busy and changing, 35 per cent idle or nearly idle, 10 per cent pinned at the power limit

    Reader-supplied. No default and no median fleet; the shape entered is the one before absorption.

  • Longest continuous pinned run

    6 minutes

    Reader-supplied.

  • Binding limit

    Contractual, measured over a 30-minute settlement interval

    Reader-supplied; no default.

  • Mean draw per rack

    95 kW

    Reader-supplied and illustrative. Left blank, the energy block goes.

  • Electricity cost

    20 pence a kilowatt-hour

    Reader-supplied and illustrative.

  • Build stage

    Connected

    Reader-supplied.

  • Operating mode

    Balanced

    The default preset, throughput and P99 latency held; the mode the coefficient was measured in.

  • Coefficient at the die

    21 per cent

    Measured on NVIDIA H100 NVL over 48 hours. Fixed, not an input.

The working

The operator and the decision

The site is at its envelope, and there is demand the operator cannot serve. A smaller bill is not what it wants. It wants more work out of the same building, in the currency NVIDIA has named: compute per megawatt inside a fixed power budget. Reducing what accelerators draw creates room, and this operator fills it with racks rather than banking it. That is a growth decision: consumption rises, work rises by more, and energy per unit of work falls. The scope is an installed fleet with no cooling change and no rack change.

The inputs and the plane conversion

The inputs are the operator's: a 6,800 kW envelope at the rack input, 48 racks of 72 accelerators at 140 kW, and a target of 54. Our coefficient is at the die and the constraint is at the rack input, so the conversion is written out. It is the die figure, times the accelerator share, times an upstream cascade factor defaulted to one. Both readings travel together. The conservative reading, 21 per cent times a 72 per cent share, is about 15 per cent at the rack input, modelled. The stated basis takes the die figure as a floor on wall watts, 21 per cent, our physics position.

What the envelope holds

Fifty-four racks draw 7,560 kW unmanaged, so the target needs a reduction of about 10 per cent to sit inside 6,800 kW. The conservative reading clears that by about 5 points and the stated basis by about 11, modelled. The target survives at any accelerator share of about 48 per cent or more.

Within the installed fleet, with no cooling change and no rack change, the envelope holds 57 racks on the conservative reading and 61 on the stated basis, modelled. That is about 19 to 27 per cent more than the installed 48. At 57 racks the conservative margin is a third of a point, so the plan sizes to 54 and keeps three racks of margin in hand. The 432 accelerators added at the target are hardware the operator buys.

The energy line

Everything here is modelled. The day shape puts 55 per cent of hours in the busy-and-changing band the coefficient is credited in. The per-rack reduction on the energy line is therefore about 8 per cent conservative and about 12 per cent on the stated basis. Before absorption, 48 unmanaged racks consume about 40,000 MWh a year. After it, 54 managed racks consume about 41,200 MWh on the conservative reading and about 39,700 MWh on the stated basis. That is about 1,300 MWh a year more, or roughly flat.

At the illustrative tariff that is about GBP 250,000 a year more on the conservative reading, and roughly flat on the stated basis, modelled, because 432 more accelerators draw power. Energy per unit of work falls about 8 to 12 per cent, the per-rack reduction seen from the other end. The freed watts have three mutually exclusive uses. Fill the envelope with racks. Bank it as energy: about 3,300 MWh and about GBP 660,000 a year at 48 racks, conservative. Or hold it as resilience margin. Choose one and plan against it.

The standing condition

Racks that fit only because the runtime is reducing what they draw make continued operation a standing condition of the configuration. The runtime is designed to fail open to full performance on a component fault, so the accelerators keep serving, and the site agrees its cover for a total failure. On a total runtime failure the target of 54 racks draws 760 kW over the envelope, modelled. At 57 racks it is 1,180 kW over, and at 61 racks 1,740 kW over.

Absorption turns headroom into racks, so the cover for a total runtime failure is agreed before the first order. Three covers exist at the site: an agreed curtailment right, reserved headroom, or fewer racks energised. Sizing to 54 rather than 57 keeps about 5 points of margin over the reduction the target needs.

The margin ladder, and the same logic at scale

Every step spends margin, modelled, on the conservative reading. Fifty racks need about 3 per cent and fit by about 12 points. Fifty-four need about 10 and fit by about 5. Fifty-seven fit by a third of a point, and 58 do not fit. On the stated basis 61 racks fit by under a point and 62 do not. The ratios hold at ten times the scale and the whole-rack counts do not. A 68 MW envelope holds 485 racks unmanaged and absorbs to 540. That adds 3,960 accelerators, modelled, for about 8,400 MWh a year more on the conservative reading, with 7,600 kW to cover on a total runtime failure, a site engineering decision.

Tokens per watt and the capacity effect

Per accelerator, the efficiency ratio is up to +22% tokens per watt in Balanced mode. It is derived from NVML die power and vLLM serving throughput on the same 48-hour H100 NVL run. On a fixed fleet the runtime holds throughput, 0% change in the run, derived, and cuts power. More output comes through the capacity released, and that is rack arithmetic, modelled: one rack in eight more at the target, about 19 to 27 per cent more at the ceiling. Size the target on work rather than racks.

The value, and what settles it

The value is the gross margin on 432 more accelerators of work, net of the electricity above, in an envelope that could not otherwise hold them. Your margin per accelerator-hour and your delivered cost complete it, and we run that arithmetic with you before the order.

Under absorption the cost of being wrong is racks in a hall the site cannot power, so the sequence matters. Run the model on today's day shape and on the one you expect afterwards. Validate for three weeks on your own racks, with the rack plane measured alongside NVML. Then order in stages and measure between them.

What would change it

What moves the result, and how to check it on your own fleet.

Two inputs move this result, and both are yours. The accelerator share of rack draw moves the fit. At 72 per cent the target clears by about 5 points, and it still clears at about 48 per cent. Read the share off your rack's own power breakdown, GPU trays against hosts, switches and fans. Or clamp the rack feed and compare it with the NVML sum. The day shape after absorption moves the energy line and, through the worst-hour gate, the fit verdict. Absorbed work that arrives in bursts behaves differently from work that arrives as a steady floor. Take your three bands from your own telemetry and run the model on today's shape and on the shape you expect. If the two differ materially, absorb in stages.

Your own numbers

The same arithmetic, on your site.

Bring your envelope, your rack specification, your installed and target rack counts and the date your next hardware order is due. We will run this arithmetic on your numbers before the order rather than beside it.