Where the electricity goes
Every watt a data centre takes from the grid ends up in one of two places. Some reaches the IT equipment: servers, accelerators, storage and networking. The rest runs the building around them: cooling, fans, pumps, lighting, power conversion and distribution losses.
The total is large enough to matter. Data centres used around 415 TWh of electricity in 2024, about 1.5% of the world's total, according to the International Energy Agency's Energy and AI report (2025).
The two parts respond to different work. Better cooling and distribution shrink the overhead. Better use of the hardware shrinks what the equipment itself draws. Each side is graded by its own number, so a gain on one side can be invisible in the other side's figure.
Four places to measure a saving
A power saving is a number plus a place in the electrical chain. Move the instrument and the same saving becomes a different percentage.
At the accelerator die, you read what the chip itself draws, through the hardware vendor's own interface. At the rack input, you see everything in the rack: accelerators, host processors, memory, networking and conversion losses.
At a power distribution unit or a meter, you see the rack plus whatever else shares its circuit. On the bill, you see the meter plus the site overhead, the tariff and a month of time.
A percentage at one of those places is not a larger or smaller copy of a percentage at another. The same saving reads as a smaller percentage of the rack, and of the site, because each total includes equipment the saving never touches. Two numbers of your own, the accelerator share and your PUE, convert one percentage into the other.
So the first question for any efficiency figure is where it was measured. The second is what it was compared against.
Tools for reducing energy in data centres
The tools fall into two families, and they are easiest to compare by what they act on.
Building-side tools act on the overhead. They include cooling control and airflow management, warmer supply temperatures, liquid cooling, more efficient power conversion and distribution, and facility monitoring that finds waste. Their results show up in the facility's overhead.
Load-side tools act on the equipment. Placement and scheduling decide where work runs, so fewer machines sit idle. Consolidation retires or merges underused servers. Newer silicon does more work per watt. Per-accelerator power control sets what each GPU draws while the work runs.
The two families are not rivals. A site can work on its cooling and its load at the same time, and less heat from the load leaves less for the cooling to remove. The guide to reducing GPU power consumption goes further into the load-side levers.
What makes a data centre efficient
Lists of the most efficient data centres usually rank sites by PUE: total facility power divided by IT power. A value near 1.0 means little overhead.
PUE grades the building, not what the equipment produces for the power it takes. The PUE guide covers what a good value looks like and what the ratio cannot see.
For the load, the useful measure is output per unit of energy. For AI inference that is tokens per watt.
An efficient site does well on both numbers, and keeps its equipment on useful work rather than idling at a high draw. Neither number replaces the other.
How to read an efficiency claim
Take one sentence carrying one number and check it on its own. It should state four things: the value, the place it was measured, the hardware, and the window of time.
It should also say what kind of number it is. Measured means an instrument recorded it. Derived means it was calculated from measured inputs. Modelled means at least one input is an assumption. The weakest input sets the word, as the glossary sets out.
Then ask two questions. Whose instrument produced the figure? And did throughput and latency hold on the same run? A comparison arm running at the same time, on the same hardware, under the same power limit, is the design to look for.
Questions people ask
- What are the main tools for reducing energy in data centres?
- Building-side tools cut the overhead: cooling control, liquid cooling and more efficient power distribution. Load-side tools cut what the equipment draws: scheduling, consolidation, newer silicon and per-accelerator power control.
- How are data centres ranked for efficiency?
- Rankings usually use PUE, which grades the building's overhead. A full picture also needs output per unit of energy at the load, such as tokens per watt for AI inference.
Where ATHLAZ fits
ADAPT works on the load side
ADAPT, AI-Driven Adaptive Power Technology, is a per-GPU software runtime. It installs beside the driver and sets how each accelerator draws power while the work runs.
Measured: up to 21% less GPU die power on NVIDIA H100 NVL, over a continuous 48-hour window, with managed and baseline arms under an equal power cap. NVML GPU-die power is a lower bound on wall power. Derived from the same run: 0% throughput change.
The saving comes from duty cycle: the idle and part-loaded hours in a fleet's day. A week of your own per-GPU power readings shows how many of those hours you have, and the proof page sets out how our run was designed.
