Skip to content

Case study, modelled

The fleet that is already paid for

The fleet is paid for. When do we retire it?

Modelled

About 14 to 16 months of deferred retirement on a 2,400-accelerator fleet, inside a range of 8 months to a little under 3 years set by the frontier, modelled.

A fully depreciated fleet of 2,400 accelerators runs inside a subscribed 1,200 kW envelope. Holding throughput while cutting die power lifts margin per watt by about 30 to 35 per cent, which buys about 14 to 16 months before retirement, modelled. The same reduction releases 151 kW to 252 kW, enough to stage the first replacement racks inside the same envelope.

Who it is for

A neocloud or GPU cloud operator running a fully depreciated inference fleet inside a subscribed envelope. The finance function signs the refresh budget, with the engineer beside them and the change-control reviewer holding the veto.

GPU cloud and neocloud

The basis

The modelled basis

This is modelled. Every rack, kilowatt, month and pound on this page is arithmetic over declared inputs you can change. One number is ours, and it is measured: up to 21 per cent less GPU die power on NVIDIA H100 NVL over 48 hours. Managed and baseline arms ran under an equal cap, read as NVML die power. Die power is a lower bound on wall power. The mechanism is duty-cycle intelligence on bursty fleets; on a saturated fleet it does not transfer. Hardware other than H100 NVL is modelled until a baseline run on it. A three-week validation on your own fleet makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • GPU die power reduction

    measured

    Up to 21 per cent

    NVIDIA H100 NVL, one continuous 48-hour window, six managed GPUs against a baseline arm under the same cap, NVML die power. A lower bound on wall power.

  • Tokens per watt

    derived

    Up to +22 per cent

    From NVML die power and vLLM serving throughput on the same 48-hour run, Balanced mode.

  • Throughput change in the run

    derived

    0 per cent

    Same 48-hour run, Balanced mode.

  • Margin per watt uplift

    modelled

    About 30 per cent (derived route) to about 35 per cent (arithmetic route)

    Energy at 30 per cent of operating margin. The derived route is the planning number; the two are never multiplied.

  • Deferred retirement at a 25 per cent frontier

    modelled

    About 14 months (derived) to 16 months (arithmetic)

    The logarithm of the uplift over the logarithm of one plus the frontier.

  • Deferred retirement across a 10 to 50 per cent frontier

    modelled

    About 8 months to a little under 3 years

    Derived route 8 to 33 months, arithmetic route 9 to 37 months. A fraction of one generation.

  • Die reduction a 12-month deferral needs, 25 per cent frontier

    modelled

    About 16 per cent, met with about 5 points to spare

    The uplift inverted. The margin is more than twelve times the re-derivation uncertainty.

  • Die reduction a 12-month deferral needs, 35 per cent frontier

    modelled

    About 21 per cent, the measured figure

    At that frontier the reduction we measured buys about 10 months on the derived route we plan against and about 12 on the arithmetic route.

  • Rack draw with the runtime

    modelled

    About 14 kW (conservative) to about 13 kW (stated basis), from 16 kW

    21 per cent applied to the 60 per cent share alone, or carried through as a floor on wall watts.

  • Headroom released inside 1,200 kW

    modelled

    151 kW to 252 kW

    Fleet draw falls from 1,200 kW to about 1,049 kW or 948 kW.

  • Racks the envelope holds

    modelled

    85 to 94, from 75

    Whole racks: 2,720 to 3,008 accelerators, about 13 to 25 per cent more.

  • Replacement racks staged inside the headroom

    modelled

    3 to 6

    At an illustrative 40 kW per replacement rack, with the cooling work priced first.

  • Energy cost avoided

    modelled

    About GBP 200,000 to GBP 330,000 a year

    About 795 to 1,325 MWh a year at a 60 per cent mean draw and 25 pence a kilowatt-hour.

  • Replacement capital deferred

    modelled

    GBP 20m, illustrative

    800 replacement accelerators, 2,400 at a 3 to 1 ratio, at GBP 25,000 each.

  • Value of the deferral, present value

    modelled

    About GBP 2,500,000 (14 months) to GBP 2,800,000 (16 months)

    At a 12 per cent cost of capital, in present value.

  • Deferred capital against the energy line over 14 months

    modelled

    About 11 times

    Energy over 14 months is about GBP 230,000 (conservative) to GBP 390,000 (stated basis).

  • The same fleet at ten times the scale

    modelled

    24,000 accelerators in 12 MW: 14 months, 1,512 kW to 2,520 kW headroom, about GBP 25m deferral value

    The capital half is scale-invariant; the site half moves with whole racks.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Accelerators in the fleet

    2,400

    Reader-supplied, illustrative.

  • Accelerator thermal design power

    300 W

    Reader-supplied, illustrative. Inside the 150 W to 350 W band of prior-generation NVIDIA datacentre parts.

  • Accelerators per rack

    32, in four eight-way air-cooled servers

    Reader-supplied, illustrative.

  • Racks in the fleet

    75

    Derived: 2,400 divided by 32.

  • Rack draw at the rack input, nominal

    16 kW

    Reader-supplied. The provisioning basis is your call.

  • Accelerator share of rack draw

    60 per cent

    Derived: 32 accelerators at 300 W is 9,600 W of a 16 kW rack. Lower than on a modern rack-scale system.

  • Site envelope, IT load at the rack input

    1,200 kW, fully subscribed

    Reader-supplied. If it is total facility power, your own power usage effectiveness is required.

  • Book value of the fleet

    Fully depreciated

    Reader-supplied.

  • Upstream cascade factor

    1, no credit

    Default. It gives our own physics argument no credit.

  • Frontier advance in output per watt

    25 per cent a year, illustrative; run from 10 to 50

    Your own lookup from vendor specifications.

  • Energy cost as a share of operating margin

    30 per cent

    Reader-supplied.

  • Throughput ratio, replacement to incumbent

    3 to 1

    Reader-supplied; the same lookup that fixes the frontier.

  • Capital per replacement accelerator

    GBP 25,000, illustrative

    Your own quotation replaces it the moment one exists.

  • Cost of capital

    12 per cent

    Reader-supplied. The deferral value is linear in it.

  • Electricity cost

    25 pence a kilowatt-hour

    Reader-supplied, United Kingdom commercial.

  • Fleet mean draw as a share of nominal

    60 per cent, illustrative

    Reader-supplied. The energy line needs your own measured mean draw.

  • Replacement rack draw at the rack input

    40 kW, illustrative

    Reader-supplied. It sizes the replacement racks the headroom can stage.

  • Operating mode

    Balanced

    The mode the coefficient was measured in. The retirement test holds revenue constant.

The working

The fleet in the model

A fleet of 2,400 prior-generation inference accelerators was bought three or four years ago and is fully depreciated. It runs inside a 1,200 kW envelope that is contracted, subscribed and full. Thirty-two accelerators sit in each of 75 air-cooled racks drawing 16 kW at the rack input. The accelerators account for 60 per cent of that draw. Every hour the fleet earns is close to free of capital cost. It also produces less output per watt than current silicon would in the same envelope. Somebody has to decide when it stops being worth running.

The retirement test

Where power is spare, a depreciated accelerator runs while it earns more than it costs. A machine earning a penny an hour passes that test forever. At a subscribed site the watt is the scarce input, so the test sharpens. The incumbent is retired when its margin per watt falls below what the replacement would earn in the same watt. Depreciation does not enter, because it is sunk. A fully depreciated fleet is the easier one to retire, which makes it the right subject for extending the window.

The runtime moves two terms and holds the third. Throughput is held, so revenue does not move. Energy cost and power fall together, because they are one physical quantity priced two ways. With energy at 30 per cent of operating margin, margin per watt rises by about 30 per cent on the derived route, modelled. That route carries the run's throughput behaviour inside it. The arithmetic route assumes throughput exactly held and gives about 35 per cent. We plan against the derived route and show the arithmetic route as the upper bound. They are one effect expressed twice and are never multiplied together.

The deferral in months

Margin per watt converts into time. At an illustrative frontier of 25 per cent a year, that lift buys about 14 months on the derived route and 16 on the arithmetic route, modelled. Across a frontier of 10 to 50 per cent a year the deferral runs from about 8 months to a little under 3 years, modelled. The refresh still happens, about 14 to 16 months later at the illustrative 25 per cent frontier, and the runtime moves onto the new silicon calibrated for it. One precondition is hard: the retirement must be economic. Where a tenant names the part, the contract sets the retirement date, and the same reduction still releases 151 kW to 252 kW inside the envelope, modelled.

Twelve months, and what it needs at the die

Inverted, the arithmetic says what a given deferral needs at the die. Twelve months at a 25 per cent frontier needs about 16 per cent. We measured 21, a margin of about 5 points and more than twelve times our re-derivation uncertainty, modelled. Twelve months at a 35 per cent frontier needs about 21 per cent, inside that uncertainty. At that frontier the reduction we measured buys about 10 months on the derived route we plan against and about 12 on the arithmetic route, modelled. One input separates those rows: the pace of the silicon you would refresh to. It is yours to look up in an afternoon.

What sets the saving on older silicon

Age does not set the saving; load shape does. Across our corpus the loaded power spreads show no trend by generation, and the saving rises with transition density rather than with a taller idle floor. Three things in your own telemetry decide it, in this order: the loaded power spread, the transition density, and the idle fraction with how far its floor can fall. A bursty fleet of any generation can carry all three.

The majority of the saving is idle-floor reduction, and most of the rest is clock-down at mid utilisation. The mid-utilisation channel returns more than twice as much per joule available. The load-bearing word in duty-cycle intelligence on bursty fleets is bursty.

The headroom and the refresh capital

Our coefficient sits at the accelerator die and your constraint sits at the rack input. Between them is your accelerator share, 60 per cent here. The conservative reading applies the 21 per cent to that share alone; the stated basis carries it through as a floor on wall watts. Both travel together. Fleet draw falls from 1,200 kW to about 1,049 kW or 948 kW, releasing 151 kW to 252 kW inside the same envelope. That is 85 to 94 racks, from 75, or 3 to 6 replacement racks at 40 kW each with the cooling work priced first, modelled. The capital argument needs none of this conversion.

The energy line is the smallest pool: about GBP 200,000 to GBP 330,000 a year at a 60 per cent mean draw and 25 pence a kilowatt-hour, modelled. Replacing 2,400 accelerators at a 3 to 1 throughput ratio takes 800 replacements, GBP 20m at an illustrative GBP 25,000 each. Deferring that spend by 14 to 16 months at the declared cost of capital is worth about GBP 2,500,000 to GBP 2,800,000 in present value, modelled. That is about 11 times the conservative energy line over the same window. Avoided capital moves spend from a capital line to an operating line, so ask your finance function early.

The read on your own fleet

Seven days of read-only power and utilisation telemetry at 1 Hz answers the three deciding variables at once. The idle fraction shows at the left edge, transition density in the time series, and loaded spread in the height of the right-hand cloud. Thirty minutes on one accelerator settles which lever holds, a clock lock or a power limit. One column decides the trial: a throughput counter, because without it a saving cannot be told from a throughput loss. Bring that log, and the three-week validation puts your number beside ours.

What would change it

What moves the result, and how to check it on your own fleet.

The pace of the output-per-watt frontier moves this result more than any other input. It alone carries the deferral from about 8 months to a little under 3 years. It is yours to check in an afternoon. Take your fleet's own part and the part you would refresh to. Read published throughput per watt for each on a like-for-like inference benchmark from the vendor's own specifications. Take the ratio and annualise it over the elapsed generations. The same lookup gives the throughput ratio that sizes the replacement capital. Behind it sits one question: whether the retirement is economic, or a tenant has named the part.

Your own numbers

The same arithmetic, on your site.

Bring your envelope, your fleet's part, the part you would refresh to and your refresh date. We will run this arithmetic on your inputs and set up the read-only week that shows which case you are in.