Hard questions
The questions a serious operator asks in the first ten minutes.
Answered plainly, in the order they usually arrive.
- Does it cost throughput?
- In the measured runs the managed GPUs completed the same work as the control GPUs while drawing less energy. Throughput and latency are the gate, not the trade: the service holds an SLA target and releases control if latency slips.
- What does ADAPT do to tokens per watt?
- Tokens per watt is the output a GPU serves for each watt it draws. In the NVIDIA H100 NVL run it rose by up to 22%, with a 0% change in throughput; both are derived figures. GPU die power fell by up to 21%, measured over 48 hours under an equal cap. More output comes from putting the released capacity to work.
- Is this just undervolting or a power cap?
- No. A cap is one static ceiling applied after the fact. This reads the live draw, recognises the workload signature and applies an operating point suited to that workload, then keeps adjusting as the workload changes.
- How is ADAPT different from DVFS?
- DVFS is the generic mechanism the vendor driver already uses to move frequency and voltage. ADAPT sits above that interface, reads the live draw, recognises the workload signature and decides when the vendor mechanism should move and by how much. It does not replace the driver; it governs it.
- Is ADAPT an energy efficiency tool?
- It makes the GPUs it manages more energy efficient. In the NVIDIA H100 NVL run the same work completed on up to 21% less GPU die power, measured over 48 hours under an equal cap. Where power is the constraint, that power is worth more as capacity than as a smaller bill.
- Does ADAPT improve PUE?
- It lowers what the whole facility draws. ADAPT cuts IT power, and the cooling that power needed comes off with it. PUE is facility power divided by IT power, so the ratio can hold level, or rise, while the site draws less. The headroom instrument takes your own PUE to show the cooling load the released power no longer adds.
- What happens if it fails?
- It fails open. On any fault the GPU returns to vendor default behaviour, and removing the service restores driver defaults. There is no state the fleet cannot get back to.
- What about vendor warranty and support?
- The service sits above the vendor interface and issues no out of specification requests. It does not modify firmware, does not reflash the board and does not run the part outside its rated envelope.
- What does the security team get?
- A written review, read before anything installs. It covers deployment shape, network posture, what the runtime reads and keeps, who can change an operating point, what you can observe, and failure and removal. Read the security review.
- How is the saving proven rather than asserted?
- Managed and control GPUs run side by side on your own fleet, with per GPU energy records that are hash linked so the record cannot be quietly edited. A third party can audit the result.
- What is the commercial risk?
- The baseline is complimentary and time boxed. It is a software install that can be removed, so the downside is three weeks and the upside is the capacity your own fleet releases, measured on your own hardware.
- What capex does it avoid?
- The power ADAPT releases can host capacity you would otherwise have to build. The headroom instrument prices that as capex avoided, on a build cost per megawatt you can change. Price it in the headroom instrument.
- Does it work on our accelerators?
- Compute results to date are measured on NVIDIA H100 NVL, with testing across six accelerator architectures. On other accelerators the figures are modelled, and the three week baseline measures them on your own fleet.
Measured on production NVIDIA H100 NVL under an equal power cap on managed and baseline arms. Your fleet baseline confirms the figure for your site.
Next
