Proof and method
Attested, not asserted.
Every figure on this site comes from one measurement discipline. This page states it plainly, so an engineer can judge it before talking to us and an auditor can check it afterwards.
Two arms, not a before and after
A managed arm and a baseline arm ran side by side on the same hardware, under the same workload, at the same time. A before and after comparison can be moved by the weather or the workload mix. A paired run cannot.
An equal cap on both arms
Both arms were held at the same power cap, so the managed arm was never simply given a smaller ceiling. The cap value is released with the assessment rather than published.
A continuous window, not an excerpt
The headline result comes from a continuous 48 hour window inside a longer campaign. No selected minutes, no discarded runs.
Records that outlive the claim
Per GPU energy records are signed and hash linked, so any later edit breaks the chain. That is what makes the evidence auditable by someone who does not trust us.
Attested, not asserted. Signed per GPU energy records, hash linked and auditable by a third party, across 118 days of testing on six accelerator architectures.
The measured run
What was run, on what, for how long.
Compute is measured today, on NVIDIA H100 NVL. Every other domain is designed on the same method, with its depth shared with licensees under agreement.
- Hardware
- Production NVIDIA H100 NVL, eight GPUs in node, with GPUs 0 to 5 managed and the remainder held as the baseline arm.
- Held equal
- An equal power cap on both arms. Same node, same cooling, same workload, same window.
- Instrumented quantity
- GPU die power, sampled continuously, alongside throughput and latency.
- Headline window
- A continuous 48 hour window, reported whole.
- Campaign length
- 118 days of testing across six accelerator architectures.
- Result
- Up to 21% less GPU die power, Up to +22% tokens per watt as a derived figure, and no throughput change in the run.
- Re-derivation
- Re-derived in house from the raw archive, against the nearest archived reference, to within 0.4 percentage points.
- Access taken
- None to workloads, models or data. Driver adjacent, and removal restores driver defaults.
Measured against designed
Measured on NVIDIA H100 NVL over a continuous 48 hour window, inside 118 days of testing on six architectures. Every other figure carries its label.
Every figure on this site, inside the instrument too, carries the word for how it was reached, such as measured, derived or modelled, so your engineers see exactly what they are checking.
Up to 21% less GPU die power
Sustained over a continuous 48 hour window on eight H100 NVL GPUs, with GPUs 0 to 5 managed and a baseline arm held at an equal power cap.
Up to +22% tokens per watt
Computed from the measured run in balanced mode: the same throughput across the same 48 hour window, on less GPU die power.
Re-derived in house to within 0.4 percentage points against the nearest archived reference.
For the reviewer
What a third party can check without taking our word for anything.
Signed, hash linked, per GPU
- That both arms carried the same power cap for the whole window.
- That the managed and baseline GPUs sat in the same node and the same thermal environment.
- That the energy records are continuous, with no gaps around the reported figures.
- That the hash chain over the records is unbroken from first sample to last.
- That throughput and latency series come from the same window as the power series.
- That the derived figures follow arithmetically from the recorded ones.
Measured on production NVIDIA H100 NVL under an equal power cap on managed and baseline arms. Your fleet baseline confirms the figure for your site.
Customer studies
What a named customer study contains.
The case studies are modelled on declared inputs. Your own study, from a three week baseline on your fleet, carries all six of these, measured on your racks and published by name only with your written permission.
Written permission
The operator agrees to be named, or the study runs anonymised with the site type and scale stated instead.
The context
Hardware, workload mix, power envelope and what the site was constrained by before the run.
The paired arms
Which accelerators were managed, which were held as baseline, and that both sat in the same node and thermal environment.
The equal cap
Confirmation that one shared power cap held across both arms for the whole window, with the value released under the assessment.
Measured outcomes only
Power, throughput and latency from the same continuous window. Derived figures labelled derived. Nothing modelled presented as measured.
What the operator did with it
The released megawatts, and whether they became scheduled capacity, a deferred connection or a deferred build.
Start with the measured run above; your baseline puts your own figure beside it.
Then do it on your own fleet
