Skip to content

Case study, modelled

The corporate estate

We run our own accelerators. We are not a service provider.

Modelled

A hall 104 accelerators short of its funded programme powers all 1,720 of them, with 112 to 328 to spare, modelled.

A 3 MW corporate hall at a power usage effectiveness of 1.45 holds 202 DGX H100 nodes today, and the programme has bought 215. With the runtime the same envelope holds 229 to 256 nodes, so every accelerator bought is energised and 112 to 328 stand beyond the requirement, modelled. Where the platform has a queue, every accelerator energised goes straight to work on it.

Who it is for

A corporate estate running accelerators for internal demand, not for sale. The infrastructure or platform director who owns the programme signs, with finance settling which budget line moves and the third-party risk function clearing the supplier first.

Enterprise

The basis

A modelled scenario

This is modelled. The estate is illustrative, with no customer behind it, and every node, megawatt-hour and pound comes from one measured coefficient and arithmetic over declared inputs you can change. The coefficient is ours, and it is measured: up to 21% less GPU die power on NVIDIA H100 NVL over 48 hours. Managed and baseline arms ran under an equal cap, read as NVML die power. Die power is a lower bound on wall power. Hardware other than H100 NVL is modelled until a baseline run on it. A three-week validation on your own fleet makes the figure yours.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • GPU die power reduction

    measured

    Up to 21% less

    NVIDIA H100 NVL, one continuous 48-hour window, six managed against a baseline arm under the same cap, NVML GPU die power. A lower bound on wall power.

  • Tokens per watt

    derived

    Up to +22%

    The same 48-hour run, Balanced mode.

  • Throughput change

    derived

    0%

    In the same run, Balanced mode, throughput and P99 latency held.

  • Reduction at the node input

    modelled

    About 12 to 21 per cent

    21% times the 55 per cent share, conservative; 21% as a floor on wall watts, stated basis.

  • Reduction the requirement needs

    modelled

    About 6 per cent

    One minus 2,069 kW over about 2,193 kW, the draw of 215 nodes at nameplate.

  • Nodes the hall holds today

    modelled

    202 nodes, 1,616 accelerators

    2,069 kW over 10.2 kW, rounded down to whole nodes.

  • Shortfall today

    modelled

    104 accelerators, 13 nodes

    1,720 required less 1,616 held. No coefficient of ours in it.

  • Nodes the hall holds with the runtime

    modelled

    229 to 256 nodes, 1,832 to 2,048 accelerators

    2,069 kW over a managed node draw of about 9 kW (conservative) to about 8 kW (stated basis).

  • Accelerators beyond the requirement

    modelled

    112 to 328

    1,832 or 2,048 held less 1,720 required.

  • Uplift on what the hall holds today

    modelled

    About 13 to 27 per cent

    1,832 or 2,048 against 1,616.

  • Fit margin at the node input

    modelled

    About 6 to 15 percentage points

    Reduction supplied less the about 6 per cent required. The widest conservative margin in the set.

  • Headroom at the node input

    modelled

    129 to 336 kW

    2,069 kW less the managed requirement draw, conservative to stated basis.

  • Break-even accelerator share

    modelled

    About 27 per cent

    Below it the conservative reading no longer meets the requirement.

  • Additional accelerator-hours a year

    modelled

    About 1,700,000 to 3,400,000

    216 to 432 additional accelerators at 7,884 hours each, at an illustrative 90 per cent availability.

  • Energy avoided a year

    modelled

    About 540 to 990 MWh

    215 nodes at 7.14 kW mean draw, the reduction across the 35 per cent busy band, over 8,760 hours, if the headroom is banked.

  • Energy value a year

    modelled

    About GBP 110,000 to GBP 200,000

    At an illustrative 20 pence a kilowatt-hour. Never added to the capacity figure.

  • Worst-hour check

    modelled

    Pinned runs of 6 minutes against a 30-minute interval

    Pinned draw of about 2,193 kW exceeds 2,069 kW, but the runs are shorter than the settlement interval, so the fit stands.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Facility envelope

    3 MW at the facility plane

    Reader-supplied: the hall's electrical rating.

  • Power usage effectiveness

    1.45

    Reader-supplied, no default. An older air-cooled hall carries a worse ratio than a purpose-built facility.

  • Envelope at the node input

    About 2,069 kW

    Computed: 3 MW over 1.45.

  • Node type

    NVIDIA DGX H100, eight H100 SXM accelerators

    Published NVIDIA DGX H100/H200 user guide. The same architecture as the measured H100 NVL, not the same part.

  • Node draw, nameplate maximum

    10.2 kW

    Published specification: a ceiling, and labelled as one.

  • Accelerator share of node draw

    About 55 per cent (5,600 W of 10,200 W)

    Computed at 700 W per accelerator. The largest lever in the model.

  • Upstream cascade factor

    1, no credit

    Default. It gives our own physics argument nothing.

  • Requirement

    2,200 kW: 215 nodes, 1,720 accelerators

    Reader-supplied: the programme's funded scope, on whole nodes.

  • Day shape

    35 per cent busy and changing, 55 per cent idle or nearly idle, 10 per cent pinned at the power limit

    Reader-supplied and illustrative until your own telemetry replaces it.

  • Longest pinned run against the settlement interval

    6 minutes against 30 minutes

    Reader-supplied. The binding limit is a contractual interval figure.

  • Mean draw per node

    7.14 kW, 70 per cent of nameplate

    Reader-supplied and illustrative. Measure it rather than assume it.

  • Electricity cost

    20 pence a kilowatt-hour

    Reader-supplied and illustrative.

  • Availability

    90 per cent

    Illustrative, for the hours conversion only. Your own maintenance regime replaces it.

  • Build stage

    Connected

    Reader-supplied. The envelope is already held.

  • Coefficient at the die

    21 per cent

    Measured. Fixed, not an input.

The working

The hall and the programme

A large organisation runs accelerators for its own purposes. It sells no compute, so there is no revenue per accelerator-hour. The hall is a corporate data hall, air-cooled, with a worse power usage effectiveness than a purpose-built facility. The capital for the accelerators is already spent, and the programme wants more of them than the hall can power. The envelope cannot grow on the programme's timescale. Changing the building's electrical infrastructure is a capital project with an outage, in a building that runs the corporation's other systems. The hardware is bought; the room is what is missing.

The declared inputs

The hall's electrical rating is 3 MW at the facility plane. Its power usage effectiveness is 1.45, entered for an older air-cooled hall, which resolves the envelope to about 2,069 kW at the node input. The node is an NVIDIA DGX H100: eight H100 SXM accelerators at a published maximum of 10.2 kW. The accelerators account for about 55 per cent of that draw. The funded scope is 2,200 kW: 215 nodes, 1,720 accelerators. The day shape is 35 per cent busy and changing, 55 per cent idle and 10 per cent pinned, illustrative until yours replaces it.

The conversion from die to node input

Our coefficient is at the accelerator die and your constraint is at the node input. Between them is your accelerator share of node draw. The reduction at the node input is our die figure times your share and an upstream cascade factor we default to one. The conservative reading applies 21% to the 55 per cent share and gives about 12 per cent, assuming nothing about the rest of the node. The stated basis treats die power as a floor on wall watts and gives 21 per cent. Both are modelled.

The fit

Today the hall holds 202 nodes and 1,616 accelerators, against a programme of 215 nodes and 1,720. It is 104 accelerators short, 13 nodes, and that figure contains no coefficient of ours. Closing the gap needs about 6 per cent off the draw at the node input. The conservative reading supplies about 12 per cent, so the requirement fits with about 6 points to spare, the widest conservative margin in this set. The stated basis fits with about 15 points.

Inside the same envelope the managed hall holds 229 to 256 nodes and 1,832 to 2,048 accelerators, modelled. Every accelerator the programme bought is energised, and 112 to 328 more stand beyond the requirement. That is an uplift of about 13 to 27 per cent on today, with 129 to 336 kW of headroom left at the node input.

What the freed capacity is worth

For an estate that runs its own work, the count is the answer: every accelerator already bought is energised, and the same building holds about 13 to 27 per cent more compute, modelled. That is a growth statement for a programme measured on delivery. At an illustrative availability of 90 per cent, the 216 to 432 additional accelerators give about 1,700,000 to 3,400,000 accelerator-hours a year, modelled. The organisation values each hour on its own terms, and on a platform with a queue every one of them goes to work.

The energy line

Energy is computed from your mean operating draw, 7.14 kW per node, and only across the busy-and-changing 35 per cent of the day. The reduction avoids about 540 to 990 MWh a year, modelled. At an illustrative 20 pence a kilowatt-hour that is about GBP 110,000 to GBP 200,000 a year, modelled. If the room is banked instead, that is its value; for a programme with bought accelerators waiting, energising them comes first. The freed envelope has three exclusive uses: energise the waiting accelerators, bank the energy, or hold it as resilience margin. This page never adds them.

Which budget line

The licence lands on an operating line. The value arrives as capital not spent on electrical works in an occupied building. Inside a corporation those are two directors and often two organisations. Which line is hard to move is a question for the finance function in the first conversation rather than the fourth, because the answer decides who has to be in the room.

The accounting treatment is your adviser's to settle. The documents they will reach for are IAS 38 on intangible assets, the interpretations committee agenda decisions of March 2019 and March 2021 on cloud arrangements, FRS 102 paragraphs 18.3B and 18.4, and ASC 350-40. The question those documents turn on is whether the customer controls the software. A runtime installed on hardware you own is a different object from software reached over a network.

The questions that come first

Three questions decide more than any coefficient. Is the binding limit a contractual demand figure measured over a settlement interval, or the rating of a transformer, breaker or busway? A sustained reduction relieves the first and not the second. Is the platform oversubscribed? And can a service be given the authority to set accelerator clocks on the estate?

A corporate buyer also runs a supplier review of any agent taking privilege on production infrastructure. It runs in parallel with the technical evaluation and can add a month on its own. The reviewer will ask for install requirements in writing, a privilege model and a data-protection position for the telemetry. Next come a removal procedure that has been run and timed, and an update and rollback policy with a bounded blast radius per accelerator. Last come a named accountable person and a certification position. That pack belongs in the first conversation. The order of work is the supplier position, then the validation, then the numbers.

What would change it

What moves the result, and how to check it on your own fleet.

The hall's power usage effectiveness moves the arithmetic more than anything else, by more than ten points on its own, and it is entirely yours. Read it from a month of metered facility draw against metered IT load rather than from the design figure. Next is the accelerator share of node draw, which sets the conservative reading. Meter one node at its input while the accelerators report die power; the fit holds at any share at or above about 27 per cent. Whether the platform is oversubscribed decides what the freed hours are worth, and your scheduler's queue answers it. Each of these inputs moves the result further than the coefficient does, and each one can be read on site. A validation on your own fleet then measures the coefficient on your day shape.

Your own numbers

The same arithmetic, on your site.

Bring your hall's electrical rating and power usage effectiveness, the node type and the number the programme has bought, a typical day on the platform and the date those accelerators need to be live. We will run this arithmetic on your inputs and set out the validation that makes the figure yours.