Skip to content

Case study, modelled

The fixed envelope

The envelope is fixed by the building, the connection and the funding round.

Modelled

Inside the same 2,000 kW envelope, 221 to 248 nodes instead of 196, and the 214-node funded scope fits, modelled.

A programme funded for 214 NVIDIA DGX H100 nodes needs about 8 per cent off at the node input to fit a 2,000 kW envelope that holds 196 nodes today. The conservative reading returns about 12 per cent and the stated basis 21 per cent, so the scope fits on both, modelled. The same envelope then holds 221 to 248 nodes: about 31 to 65 further awards a year at an illustrative median.

Who it is for

A publicly funded research compute programme. The facility or programme director who owns the funded scope signs, with the institution's information-security and research-governance function holding the gate.

Research computing

The basis

What this page is

This is modelled. Every node, accelerator-hour, megawatt-hour and pound here is arithmetic over inputs you can change. One number is ours, and it is measured: up to 21% less GPU die power on NVIDIA H100 NVL over 48 hours. Managed and baseline arms ran under an equal cap, read as NVML die power.

Die power is a lower bound on wall power. The node modelled here is an NVIDIA DGX H100, and hardware other than H100 NVL stays modelled until a baseline run on it. A three-week validation on your own nodes makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • Less GPU die power

    measured

    Up to 21%

    NVIDIA H100 NVL, one continuous 48-hour run, six GPUs managed against a baseline arm under the same cap, NVML GPU die power.

  • Tokens per watt

    derived

    Up to +22%

    NVML die power and vLLM serving throughput on the same 48-hour run, Balanced mode. The same tokens for less energy.

  • Throughput change in the run

    derived

    0%

    Same 48-hour run; the runtime holds throughput and cuts power.

  • Reduction required to fit the funded scope

    modelled

    About 8 per cent at the node input

    214 nodes at 10.2 kW nameplate draw about 2,183 kW against a 2,000 kW envelope.

  • Reduction at the node input, conservative reading

    modelled

    About 12 per cent

    21 per cent at the die times an accelerator share of about 55 per cent, cascade factor one.

  • Reduction at the node input, stated basis

    modelled

    21 per cent

    Die power taken as a floor on wall watts, a physics argument.

  • Nodes the envelope holds

    modelled

    196 unmanaged; 221 conservative; 248 stated basis

    2,000 kW over the managed node draw, rounded down to whole nodes.

  • Accelerators the envelope holds

    modelled

    1,568 unmanaged; 1,768 conservative; 1,984 stated basis

    Nodes times 8.

  • Racks at 4 nodes per rack

    modelled

    49 unmanaged; 55 and one node conservative; 62 stated basis

    Node count over 4.

  • Funded scope of 1,712 accelerators

    modelled

    Short by 144 unmanaged; met on both readings, with 56 or 272 beyond it

    Accelerators the envelope holds against the funded 1,712.

  • More accelerators inside the same envelope

    modelled

    About 13 to 27 per cent

    221 and 248 nodes against 196. Node-count arithmetic, not a tokens-per-watt figure.

  • Break-even accelerator share

    modelled

    40 per cent

    The required reduction over 21 per cent. The fit survives at any share at or above it.

  • Freed envelope, conservative reading

    modelled

    About 231 kW; about 69 kW spare after the 214 funded nodes (about 276 kW on the stated basis)

    2,000 kW less the managed draw of 196 nodes, and of 214 nodes.

  • Additional accelerator-hours a year

    modelled

    1,576,800 conservative; 3,279,744 stated basis

    200 or 416 additional accelerators at 7,884 available hours each.

  • Additional accelerator-hours over a four-year programme

    modelled

    6,307,200 conservative; 13,118,976 stated basis

    The annual figures times four.

  • Further awards a year

    modelled

    About 31 conservative; about 65 stated basis

    Additional hours over an illustrative median award of 50,000 accelerator-hours.

  • Avoided energy a year, if banked rather than filled

    modelled

    1,413 MWh conservative; 2,574 MWh stated basis

    196 nodes at an illustrative mean draw of 70 per cent of nameplate.

  • Avoided cost a year, if banked

    modelled

    About GBP 141,000 conservative; about GBP 257,000 stated basis

    At an illustrative 10 pence a kilowatt-hour, your own tariff.

  • The same logic at ten times the scale

    modelled

    20,000 kW holds 1,960 nodes unmanaged and 2,216 or 2,482 managed: about 16 million to 33 million more accelerator-hours a year

    Ratios are scale-invariant because the runtime acts per accelerator; whole-node rounding is done at each scale.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Envelope, IT load at the node input

    2,000 kW

    Reader-supplied. If it is total facility power, supply your own power usage effectiveness; there is no default.

  • Node type

    NVIDIA DGX H100, 8 x H100 SXM

    Published specification, from the NVIDIA DGX H100/H200 user guide.

  • Node draw, nameplate maximum

    10.2 kW

    Published specification. A ceiling: nameplate overstates operating draw.

  • Accelerators per node

    8

    Published specification.

  • Accelerator share of node draw

    About 55 per cent

    5,600 W of accelerators (8 at 700 W) over 10,200 W of node. Yours for your own nodes, and the largest lever in the model.

  • Upstream cascade factor

    1, no credit

    Default. It gives our own physics argument no credit.

  • Funded procurement

    214 nodes, 1,712 accelerators

    Reader-supplied. The scope the programme was funded for.

  • Nodes per rack

    4

    Reader-supplied. A floor-loading and cooling decision that changes only how the answer is reported.

  • Availability, for the hours conversion

    90 per cent

    Reader-supplied. Planned maintenance and outage allowance.

  • Hours per year

    8,760

    Fixed.

  • Median award, accelerator-hours per project

    50,000

    Reader-supplied and illustrative. It is your own allocation policy.

  • Mean operating draw, share of nameplate

    70 per cent

    Reader-supplied and illustrative. Used only in the energy line.

  • Electricity cost, energy line only

    10 pence a kilowatt-hour, illustrative

    Your own figure replaces it; the energy line scales linearly with it.

  • Coefficient at the die

    21 per cent

    Measured. Fixed; not an input.

The working

The envelope and the three things that fix it

A research compute facility runs inside an envelope fixed three times over. The building's electrical rating fixes it, so does the import capacity agreed for the site, and so does the funding round that paid for both. A commercial operator with a constrained site has options that cost money. This facility has none inside the programme period. The one thing it can change is what the compute draws.

Software that reduces that draw does not enlarge the envelope. It changes how much compute the envelope holds. For a funded programme that is counted in accelerator-hours allocated and projects served. On a fixed envelope the outcome is growth rather than saving.

Two questions before the arithmetic

First, establish in writing that a service may set accelerator clocks on your estate. ADAPT runs driver-adjacent as one lightweight service per accelerator and needs that authority on the host. On a shared research estate the information-security and research-governance function holds that gate. It will ask for install requirements, a privilege model and a data-protection position. Bring that function into the first conversation, so its answer arrives while the numbers are worked.

Then your load shape. The mechanism is duty-cycle intelligence on bursty fleets. Most of the saving is idle-floor reduction and most of the rest is clock-down at mid utilisation; on a fully saturated fleet it does not transfer. A queue kept full leaves little idle floor, where most of the saving comes from, and the mechanism still works in what remains: gaps between jobs, ramp and drain, maintenance windows, reservation shadows and mid-utilisation phases inside jobs. A three-week validation under your own queue measures how much.

From the die to the node input

Our coefficient is at the accelerator die and your envelope is bounded at the node input. Between them sits the share of node draw the accelerators account for. On an NVIDIA DGX H100 the published nameplate puts 5,600 W of accelerators inside a 10.2 kW node, about 55 per cent. The runtime does not touch the other 45 per cent.

Two readings follow, and they travel together. The conservative reading applies the die figure to the accelerator share alone and returns about 12 per cent at the node input, modelled. The stated basis treats die power as a floor on wall watts, because a watt not drawn at the die is not converted, distributed or cooled. It returns 21 per cent, modelled. The model carries both readings side by side, and the conservative one is the planning figure.

What the envelope holds

The programme in this model was funded for 214 nodes and 1,712 accelerators, inside 2,000 kW of IT load at the node input. At nameplate those nodes draw about 2,183 kW, so fitting them needs about 8 per cent off at the node input. Unmanaged, the envelope holds 196 nodes and 1,568 accelerators, 144 accelerators short of the funded scope.

Managed, the same envelope holds 221 nodes and 1,768 accelerators on the conservative reading, modelled. On the stated basis it holds 248 nodes and 1,984 accelerators. The funded scope fits on both, with 56 or 272 accelerators to spare, an uplift of about 13 to 27 per cent. That is node-count arithmetic, and the fit survives at any accelerator share of 40 per cent or above.

The freed envelope in allocated hours

A funded programme is measured on hours allocated and projects served. At 90 per cent availability each accelerator gives 7,884 hours a year. The 200 accelerators the conservative reading adds are 1,576,800 accelerator-hours a year; the 416 on the stated basis are 3,279,744, modelled. Over a four-year programme that is 6,307,200 to 13,118,976 hours. At an illustrative median award of 50,000 accelerator-hours, set by your own allocation policy, it is about 31 to 65 further awards a year.

Where your queue is oversubscribed, those hours become further awards, as many as your oversubscription absorbs. Where it is not, the freed envelope banks energy instead, the line below. Most of the conservative gain, 144 of the 200 accelerators, is funded scope the envelope could not hold; the last 56 sit at the edge of the conservative fit and are hardware beyond the funded scope, so plan the funded 214 nodes first.

One envelope, spent once

The freed envelope is spent once. Fill it with accelerators, the largest use where the queue is oversubscribed. Bank it as energy, the smallest use, and often a separately funded line at a public facility. Or hold it as resilience margin. If you fill it, your energy line does not fall and your compute rises. If you bank it, the extra nodes do not appear.

If nodes fit only because the runtime is reducing what they draw, continued operation becomes a standing condition of the configuration. The runtime fails open to full performance on a component fault, the right behaviour at the accelerator. The mitigation lives with the facility: reserved headroom, an agreed curtailment right, or a commitment set far enough below the modelled figure that margin absorbs a total failure.

The energy line, if the envelope is banked

Energy is computed from your mean operating draw, never from nameplate. At an illustrative 70 per cent of nameplate across the 196 nodes inside the envelope, the avoided energy is 1,413 MWh a year on the conservative reading, modelled. On the stated basis it is 2,574 MWh. At an illustrative 10 pence a kilowatt-hour, your own tariff, those lines are about GBP 141,000 and GBP 257,000 a year. It is the smallest line because a facility on a fixed envelope is short of compute, not electricity.

What would change it

What moves the result, and how to check it on your own fleet.

Your load shape moves the result most: a queue that runs flat out leaves less idle time to recover, and a three-week validation under your own queue measures how much remains. Inside the arithmetic the largest lever is the accelerator share of node draw. The fit holds at any share of 40 per cent or above, and you check yours by reading node-input draw against NVML die power on your own instruments. Two more checks follow. If your envelope is total facility power rather than IT load, supply your own power usage effectiveness. Plan the fit on the draw basis your envelope is allocated on, and enter your measured mean node draw for the energy line; the basis moves the required reduction by more than twenty points, so settle it first. Ask, too, whether your binding limit is a contracted figure over a settlement interval, which a sustained reduction relieves, or a plant rating, which it does not.

Your own numbers

The same arithmetic, on your site.

Bring your envelope at the node input, your funded node count and the date your next call closes. We will run this arithmetic on your inputs and set out the validation on your own nodes.