Skip to content

Case study, modelled

The cluster being quoted

We are quoting a cluster and the architecture is still open.

Modelled

2,880 accelerators fit a 5 MW envelope with about 4 to 10 points to spare, or the same 5,600 kW carries 47 to 50 racks rather than 40, modelled.

A 40-rack, 2,880-accelerator inference cluster drawing 5,600 kW unreduced fits a 5 MW envelope once the runtime is assumed at design, modelled. The connection application shrinks by about 850 to 1,180 kW. Held the other way, the same 5,600 kW carries 7 to 10 more racks, about 18 to 25 per cent more accelerators.

Who it is for

An OEM, systems integrator or NVIDIA partner quoting a dedicated inference cluster. The partner's software-portfolio owner signs, with the end client's design authority and the capital signatory behind them.

Systems builders

The basis

A modelled scenario

This is modelled. Every rack, megawatt and pound is arithmetic over inputs you can change. One number is measured: up to 21% less GPU die power, NVML, on NVIDIA H100 NVL over 48 hours. Six managed GPUs ran against a baseline arm under the same cap. Die power is a lower bound on wall power. The quoted rack is a later generation than we measured, so every figure on it is modelled until a baseline run on that hardware. A three-week validation on your own fleet turns the coefficient into your number.

Key figures

Every figure, what kind of figure it is, and its basis.

One figure is measured. The derived figures come from the same run. Everything else is modelled on the inputs below.

  • GPU die power reduction

    measured

    Up to 21%

    NVIDIA H100 NVL, 48 hours, six managed GPUs against a baseline arm under the same cap, NVML die power.

  • Tokens per watt

    derived

    Up to +22%

    NVML die power and vLLM serving throughput on the same 48-hour run.

  • Throughput change in the run

    derived

    0%

    Same 48-hour run.

  • Reduction at the rack, conservative reading

    modelled

    About 15 per cent

    21% times a 72 per cent accelerator share times a cascade factor of one.

  • Reduction at the rack, stated basis

    modelled

    21 per cent

    The die figure treated as a floor on wall watts.

  • Reduction the requirement needs to fit

    modelled

    About 11 per cent

    One minus 5,000 over 5,600.

  • Provisioned envelope at the recorded convention

    modelled

    About 7,440 kW, about 33 per cent above nominal

    40 racks at 155 kW peak plus 20 per cent headroom, against 5,600 kW nominal.

  • Direction A: requirement draw with the runtime

    modelled

    About 4,750 to 4,420 kW

    40 racks at about 119 kW (conservative) to about 111 kW (stated basis).

  • Direction A: margin above the requirement

    modelled

    About 4 to 10 points

    Inside 5 MW, conservative to stated basis. Roughly fifteen times the re-derivation uncertainty at the conservative end.

  • Direction A: contracted capacity not requested

    modelled

    About 850 to 1,180 kW

    5,600 kW times the rack-input reduction, conservative to stated basis.

  • Direction A: commitment not posted

    modelled

    About GBP 200,000 to GBP 840,000

    Capacity not requested times the band proposed for Great Britain in the Ofgem Curate consultation of 29 July 2026.

  • Direction B: additional racks in the same 5,600 kW

    modelled

    7 to 10

    47 to 50 racks rather than 40, rounded down to whole racks.

  • Direction B: additional accelerators

    modelled

    504 to 720, about 18 to 25 per cent more

    Whole-rack arithmetic.

  • Energy avoided at a 40 per cent duty factor

    modelled

    About 3,000 to 4,100 MWh a year

    850 to 1,180 kW of continuous demand avoided in direction A, over 8,760 hours, times the duty factor.

  • Energy value at a 40 per cent duty factor

    modelled

    About GBP 740,000 to GBP 1,030,000 a year

    At about 25 pence a kilowatt-hour. It lands with whoever holds the electricity contract, a point to settle in the quote.

  • Heat not rejected per rack

    modelled

    About 21 to 29 kW

    Conservation of energy on a 140 kW rack, conservative to stated basis.

  • Capital not built, one-off, at an illustrative GBP 10m per MW

    modelled

    About GBP 8m to GBP 12m

    Megawatts not built times your own cost per megawatt. Deferred rather than avoided where demand is growing.

  • Smallest envelope the requirement fits

    modelled

    About 4,760 kW conservative; about 4,430 kW stated basis

    Requirement held at 2,880 accelerators. Keep at least a percentage point of reduction in hand above it.

  • At ten times the scale, capacity not requested

    modelled

    About 8 to 12 MW

    28,800 accelerators in 400 racks; every ratio holds. Direction B admits 71 to 106 additional racks.

Declared inputs

What the arithmetic runs on.

Each input is stated with its basis. Change any of them and the figures above move with it.

  • Accelerators per rack

    72

    NVIDIA published GB300 NVL72 configuration, 72 accelerators and 36 host processors. Reader-supplied for any other rack.

  • Rack draw at the rack input, nominal

    140 kW

    OEM system specifications, 2026, from a published 132 to 142 kW range. Reader-supplied.

  • Rack peak draw

    155 kW

    OEM system specifications. Used only to work the quoting convention.

  • Accelerator share of rack draw

    72 per cent

    About 101 kW of 140 kW from the same specifications. Reader-supplied, and the largest lever in the model.

  • Upstream cascade factor

    1

    Default. It gives our own physics argument no credit.

  • Client requirement

    2,880 accelerators

    Reader-supplied, from the client's workload. 40 whole racks; rack counts round down throughout.

  • Envelope offered on the required date

    5 MW, IT load at the rack input

    Reader-supplied. If it is total facility power, your power usage effectiveness is required.

  • Coefficient at the die

    21 per cent

    Measured on NVIDIA H100 NVL over 48 hours, NVML die power. Fixed, not an input.

  • Quoting convention

    Nameplate peak plus about 20 per cent headroom

    One recorded practitioner account of how a data hall is specified. Used only to separate what a sustained mean moves.

  • Connection commitment, proposed

    GBP 237,500 to GBP 712,500 per MW

    Ofgem, Curate: Demand Connections Reform, consultation of 29 July 2026, paragraph 4.15, as proposed for Great Britain.

  • Electricity cost, energy line only

    About 25 pence a kilowatt-hour

    A United Kingdom commercial figure. Reader-supplied.

  • Cost per megawatt of new capacity

    GBP 10m per MW, illustrative only

    The model sets no default. Your own capital plan replaces it.

The working

The moment the envelope is still a variable

An integrator, an OEM or an NVIDIA partner is quoting an inference cluster. The workload is understood. Racks, network, electrical design and the connection application are not yet fixed. Once the transformer is ordered and the application lodged, the envelope is set for the site's life. This is the one point where a software runtime changes a specification rather than a bill. At design stage the binding constraint is contracted capacity, not consumption. Every regime we examined rations what a site asks for rather than the energy it uses. Here a reduction converts into value as a smaller capacity request.

The rack, the inputs and the plane arithmetic

The reference rack is a declared input: the NVIDIA GB300 NVL72, 72 accelerators and 36 host processors, liquid-cooled. OEM specifications put it at 132 to 142 kW. The model uses 140 kW, with the accelerator dies at about 101 kW, a 72 per cent share. The client needs 2,880 accelerators: 40 whole racks drawing 5,600 kW unreduced, against a 5 MW envelope of IT load. Fitting needs a reduction of about 11 per cent, modelled.

Our coefficient sits at the accelerator die and your specification at the rack input, so the conversion is one multiplication. It is 21% times your accelerator share times an upstream cascade factor we default to one. Two readings travel together. The conservative reading gives about 15 per cent at the rack, modelled. The stated basis, which treats the die figure as a floor on wall watts, gives 21 per cent, modelled.

Where the reduction enters the specification

Busway, transformer and switchgear stay sized on nameplate peak plus margin, about 7,440 kW here, as they are today. That is the recorded convention, 155 kW peak plus about 20 per cent headroom, against 5,600 kW nominal, modelled. The reduction enters through the connection application, a declared demand built on a utilisation assumption, and a sustained-mean measurement is the evidence for that assumption.

Ofgem said in June 2026 that it was considering mandatory disclosure of intended ramp rates and utilisation factors. The design authority who owns the application converts that evidence into a smaller specification. Everything priced below rests on that row.

Direction A: the same 2,880 accelerators in a smaller envelope

Unreduced, a 5 MW envelope holds 35 racks, 2,520 accelerators, 5 racks short, modelled. With the runtime assumed, effective rack draw falls to about 119 kW on the conservative reading and about 111 kW on the stated basis. The 40 racks then draw about 4,750 to 4,420 kW and fit with about 4 to 10 points of margin, modelled. Our re-derivation delta at the die is under half a point; the conservative margin is roughly fifteen times that.

What gets smaller is the connection application. Contracted capacity not requested is about 850 to 1,180 kW, modelled. The Ofgem Curate consultation of 29 July 2026 proposed a connection commitment for data centres in Great Britain. Against it, that is about GBP 200,000 to GBP 840,000 not posted, modelled. At ten times the scale, 400 racks, the capacity not requested is about 8 to 12 MW, modelled, itself a small data centre.

Direction B: more accelerators in the same envelope

Hold the envelope at 5,600 kW and 47 racks fit on the conservative reading, 50 on the stated basis. That is 7 to 10 more racks and 504 to 720 more accelerators, about 18 to 25 per cent, modelled. Those racks are hardware the client or integrator still buys. The runtime removes the power constraint, so direction B is a larger quote, not a cheaper one. Per accelerator the run gave up to +22% tokens per watt, derived, at 0% throughput change, derived. Per quote, the token uplift is the rack count. The two directions spend the same watts once: choose one, quote it, and plan against it.

Energy, heat and capital, in order of size

The energy line is the smallest. Direction A avoids about 850 to 1,180 kW of continuous demand. At a 40 per cent duty factor that is about 3,000 to 4,100 MWh a year, worth about GBP 740,000 to GBP 1,030,000 at a United Kingdom commercial tariff, modelled. The figure uses a 40 per cent duty factor, because the saving comes from the idle and mid-utilisation hours that bursty load leaves; a baseline run on the client's own accelerators sets the real figure.

Heat follows the draw: about 21 to 29 kW per rack leaves the cooling loop, modelled. That feeds liquid-cooling sizing on your loop design. Capital is your field: megawatts not built times your own cost per megawatt. At an illustrative GBP 10m per MW, which is not a default, that is about GBP 8m to GBP 12m one-off, modelled. Where demand is growing it is deferred rather than avoided.

Who signs, and the standing condition

Three signatures decide this scenario. The partner's software-portfolio owner decides which third-party software the organisation carries, on a window of roughly two months. That owner asks for a support model and a qualification pack. The end client's design authority owns the diversity factor and the connection application. Whoever signs the capital authorisation fixes the envelope. Their first question is how the licence sits in the accounts, a question for their auditor, so bring the auditor in early.

The runtime installs driver-adjacent above NVML, one lightweight service per accelerator, with no hardware and no application code change: a software line in the quote. A cluster specified into an envelope on its strength keeps the runtime as a standing condition of the configuration. That condition belongs in the quote and the first conversation. A three-week validation on the client's own fleet, run before the application is lodged, gives the design authority a number they can put in writing.

What would change it

What moves the result, and how to check it on your own fleet.

The quoting basis moves the result more than every other input combined. Does the connection application rest on a declared utilisation assumption, or on nameplate maximum alone? Read your own application's utilisation basis to check it. The accelerator share is the next lever, defaulted at 72 per cent. One rack-input reading taken concurrently with NVML on one of your racks replaces it. Then the envelope itself. Hold the requirement at 2,880 accelerators and the conservative reading stops fitting below about 4,760 kW, the stated basis below about 4,430 kW, modelled. Plan with at least a percentage point in hand above either boundary. And the fleet must be bursty; a baseline run on your own accelerators settles that.

Your own numbers

The same arithmetic, on your site.

Bring the envelope you have been offered, the accelerator count the workload needs and the date the connection application is due. We will work both directions on your own inputs before the specification is fixed.