Skip to content

Case study, modelled

The fleet this is for

Are we a good candidate?

Modelled

A 7,200 kW envelope that holds 51 racks today holds 60 to 65 with the runtime, 3 to 8 racks beyond a 57-rack requirement, modelled.

An operator that owns its accelerators and its site, is sold out and runs a bursty fleet needs 57 racks inside a 7,200 kW envelope that holds 51 today. The requirement needs about a 10 per cent reduction at the rack input, and the runtime supplies about 15 to 21 per cent, modelled. So 60 to 65 racks fit, and the accelerators that fill them are hardware the operator still buys.

Who it is for

A GPU cloud operator that owns its accelerators, holds its own site and electricity account, and is sold out. The capacity planner brings it, the CFO signs it, and change control gives the clock-setting authority.

GPU cloud and neocloud

The basis

A modelled scenario

This is modelled. The operator is illustrative, with no customer behind it: we built it from a rack specification, a power envelope and a declared day shape. Every figure here is arithmetic over inputs you can change. One number is ours, and it is measured: up to 21% less GPU die power by NVML, on NVIDIA H100 NVL over 48 hours. Managed and baseline arms ran under the same cap. Die power is a lower bound on wall power. Hardware other than H100 NVL is modelled until a baseline run on it. A three-week validation on your fleet is what makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • Less GPU die power, sustained

    measured

    Up to 21%

    NVIDIA H100 NVL, one continuous 48-hour window, six managed against a baseline arm under the same cap, NVML die power.

  • Tokens per watt

    derived

    Up to +22%

    NVML die power and vLLM serving throughput on the same 48-hour run, Balanced mode.

  • Throughput change in the run

    derived

    0%

    Throughput and P99 latency held in Balanced mode on the same 48-hour run.

  • Reduction the requirement needs at the rack input

    modelled

    About 10%

    One minus 7,200 kW over 7,980 kW, the 57-rack requirement against the envelope.

  • Reduction at the rack input, conservative to stated basis

    modelled

    About 15% to 21%

    21% times a 72% accelerator share times a cascade factor of one; and 21% taken as a floor on wall watts.

  • Fit margin at the rack input

    modelled

    About 5 to 11 percentage points

    Reduction supplied less reduction needed, on each reading.

  • Racks inside 7,200 kW

    modelled

    51 today; 60 to 65 with the runtime

    Whole racks at 140 kW unmanaged, about 119 kW conservative, about 111 kW stated basis.

  • Accelerators inside 7,200 kW

    modelled

    3,672 today; 4,320 to 4,680 with the runtime

    Racks times 72.

  • Beyond the 57-rack requirement

    modelled

    3 to 8 racks, 216 to 576 accelerators

    Racks the envelope holds less the requirement. Hardware the operator must still buy.

  • Headroom inside the envelope

    modelled

    427 to 896 kW

    7,200 kW less the managed requirement draw of about 6,770 kW or about 6,300 kW.

  • Uplift on what the envelope holds today

    modelled

    About 18% to 27%

    60 or 65 racks against 51. Rack-count arithmetic, not a tokens-per-watt figure.

  • Energy avoided a year, if the headroom is banked

    modelled

    About 4,700 to 6,500 MWh

    A 5,415 kW mean draw times the reduction times the 65% busy share, over 8,760 hours. Idle and pinned bands credited at zero.

  • Value of the banked energy a year

    modelled

    About GBP 930,000 to GBP 1,300,000

    At the illustrative 20 pence a kilowatt-hour.

  • Accelerator share at which the fit still holds

    modelled

    About 47% or more

    Reduction needed over the die coefficient, against 72% declared here.

  • Draw to cover on a total runtime failure

    modelled

    780 kW at the requirement; 1,200 kW at the ceiling

    Unmanaged draw of 57 or 60 racks at 140 kW, less the 7,200 kW envelope. Covered by an agreed curtailment right, reserved headroom, or fewer racks energised.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Envelope, IT load at the rack input

    7,200 kW

    Reader-supplied. If it is total facility power, supply your own power usage effectiveness; there is no default.

  • Plane of the envelope

    IT load at the rack input

    Reader-selected. No plane is preselected.

  • Rack draw at the rack input, nominal

    140 kW

    Rack specification, reader-supplied.

  • Provisioning basis

    Nominal

    Reader-selected. There is no default.

  • Accelerators per rack

    72

    Rack specification.

  • Accelerator share of rack draw

    72 per cent

    Reader-supplied. The largest lever in the model.

  • Upstream cascade factor

    1, no credit

    Default. No credit for our own physics argument.

  • Requirement

    8 MW

    Reader-supplied.

  • Day shape, busy and changing

    65 per cent

    Reader-supplied. No default and no median fleet; this is condition four as a number.

  • Day shape, idle or nearly idle

    25 per cent

    Reader-supplied.

  • Day shape, pinned at the power limit

    10 per cent

    Reader-supplied.

  • Longest continuous pinned run

    6 minutes

    Reader-supplied.

  • Binding limit

    Contractual, measured over a 30-minute settlement interval

    Reader-supplied. This is condition seven as an input; there is no default.

  • Mean draw per rack

    95 kW

    Reader-supplied and illustrative. Left blank, the energy line goes.

  • Electricity cost

    20 pence a kilowatt-hour

    Reader-supplied and illustrative.

  • Build stage

    Connected

    Reader-supplied.

  • Coefficient at the die

    21 per cent

    Measured. Fixed; not an input.

The working

The seven conditions

The fit is strongest when seven things hold, and six can be checked without us. You own or contract the accelerators and decide what runs on them. You hold your own site or electrical envelope. You are commercially sold out. Your fleet is bursty rather than saturated. You hold the electricity account, or a tariff that charges for the marginal kilowatt-hour. You can give a service the authority to set accelerator clocks on your estate. And your binding limit is a contractual figure measured over a settlement interval, not the rating of installed plant. Four have their own studies where they read differently: a tenant in someone else's hall (The tenant and the landlord), an estate that sells no compute (The corporate estate), a committed power contract (The power that is paid for either way) and a limit set by cooling plant (The heat that cannot leave).

Four are public: your commercial model, site tenure, published availability and billing basis. One needs a week of your own telemetry, read only, as three day-shape bands: busy and changing, idle, and pinned. One is a question for your facility engineer, and one for your change-control and information-security functions. The mechanism is duty-cycle intelligence on bursty fleets. The majority of the saving is idle-floor reduction and most of the rest is clock-down at mid utilisation, so a week of read-only telemetry shows how much of your day it works on; a fully saturated fleet is a grid-flexibility conversation instead.

The site and the requirement

The operator holds a 7,200 kW envelope, IT load at the rack input, and needs 8 MW of racks. Each rack draws 140 kW nominal and carries 72 accelerators. The requirement rounds down to 57 whole racks, 4,104 accelerators, drawing 7,980 kW. Fitting that needs about a 10 per cent reduction at the rack input, modelled. Today the envelope holds 51 racks, 3,672 accelerators, so the operator is 6 racks and 432 accelerators short. That shortfall contains nothing of ours: it is your own requirement against your own envelope.

The step from the die to the rack input

Our coefficient is at the accelerator die; the constraint is at the rack input. Between them sit the accelerators' share of rack draw, declared at 72 per cent, and an upstream cascade factor defaulted to one. Both readings travel together. The conservative reading applies the die figure to the accelerator share alone: about 15 per cent, modelled. The stated basis treats a watt not drawn at the die as a watt not converted, distributed or cooled: 21 per cent, modelled. This page plans against the conservative reading.

What fits

On the conservative reading a managed rack draws about 119 kW. The 57-rack requirement draws about 6,770 kW and sits inside 7,200 kW with about 5 percentage points to spare, modelled. The envelope now holds 60 racks, 4,320 accelerators. The shortfall closes, and 3 racks and 216 accelerators sit beyond the requirement, with 427 kW of headroom. On the stated basis it holds 65 racks and 4,680 accelerators: 8 racks and 576 accelerators beyond the requirement, 896 kW of headroom and about 11 points to spare, modelled.

Two things travel with those figures. ADAPT, a per-GPU software runtime, creates the room, and the operator fills it with accelerators on its own order schedule. And the margin holds. About 5 points is about eighteen times our own measurement uncertainty at that plane, and the fit survives at any accelerator share of about 47 per cent or more. The arithmetic settles the plane conversion. The operator's own racks then measure the die-level figure on its own load shape.

The energy line, if the watts are banked

The freed watts have three uses that exclude each other: fill them with accelerators, bank them as energy, or hold them as margin against the standing condition below. The fleet's mean draw is 5,415 kW at a declared 95 kW per rack. Only the busy-and-changing share of the day, 65 per cent, carries the coefficient; the idle and pinned bands are credited at zero. Continuous demand avoided is about 530 kW conservative and about 740 kW on the stated basis, modelled.

Over 8,760 hours that is about 4,700 to 6,500 MWh a year, modelled. At an illustrative 20 pence a kilowatt-hour, your own tariff, it is about GBP 930,000 to GBP 1,300,000 a year, modelled. The energy line is conservative by construction: the model credits only busy-and-changing hours, while the measured run put the majority of recovered energy in the idle floor, so idle hours are upside the validation measures.

The standing condition

A tenancy that fits because the runtime is reducing rack draw keeps the runtime as a standing condition of the configuration, and the site plans its cover from the start. The runtime fails open on a component fault by design, so the accelerators keep serving. On the conservative reading, modelled, 60 energised racks draw 8,400 kW unmanaged against a 7,200 kW envelope: 1,200 kW to cover on a total runtime failure. The site covers it with an agreed curtailment right, reserved headroom, or fewer racks energised, and planning on 57 rather than 60 brings the draw to cover down to 780 kW.

How far the fit reaches

Plan on 57 to 59 racks on the conservative reading, with margin in hand, modelled. At 60 racks that reading still fits, by under a point; from 61 to 65 the room rests on the stated basis, and the validation on your own racks measures the reduction they deliver. At 66 racks, which need about 22 per cent, the route is a larger envelope.

The validation, and the commitment

Duty cycle sets the saving, so the three-week validation on your own racks, at your own plane, turns this arithmetic into your own rack count.

What would change it

What moves the result, and how to check it on your own fleet.

The day shape is the assumption that moves this page most, and it is your own input. Swept from 10 to 75 per cent busy and changing, the rack count stays at 60 on the conservative reading, because the model gates capacity on the worst hour. The energy line moves instead, from about 720 MWh a year to about 5,400 MWh, modelled. Two other inputs turn the capacity verdict. One is the accelerator share, where the fit survives at about 47 per cent or more. The other is the longest pinned run: at 45 minutes against a 30-minute settlement interval, the verdict turns conditional. Check all three on your own fleet. Take one week of read-only telemetry in busy-and-changing, idle and pinned bands. Add your rack specification, and one question to your facility engineer about the interval your limit is measured over.

Your own numbers

The same arithmetic, on your site.

Bring your envelope and the plane it is measured at, your rack specification, a week of your own telemetry as three day-shape bands and the date you need the racks energised. We will run this arithmetic on your inputs and set out the validation that makes the figure yours.