1.5% of requests in reasoning mode doubles an AI service's electricity use
Once 1.5% of an AI service's requests switch to reasoning mode, its electricity use doubles. The ratio does all the work: a text prompt costs 0.3 Wh, a reasoning request roughly 20 Wh, 67 times more. A small minority of heavy requests is enough to match all the remaining traffic. The one peer-reviewed paper that touches this threshold puts it at 10%. The gap between the two figures is not an arithmetic dispute. It comes from the price you put on a reasoning request.
Figures as of June 30, 2026. No site value was changed for this page: both terms of the calculation were already published here.
The calculation, in the open
Two costs per request, one division. Nothing else feeds the threshold.
The site's reference figure for an average text prompt. The numbers published by Google, Epoch AI and OpenAI sit between 0.24 and 0.4 Wh.
The model writes out its intermediate steps before answering. No vendor publishes an official figure for its reasoning modes. Independent benchmarks place these requests between 15 and 30 Wh; the site uses 20 Wh.
A service where a share p of requests runs in reasoning mode averages (1 - p) × 0.3 + p × 20 Wh per request. Doubling means reaching 0.6 Wh. Solve it: p × 19.7 = 0.3, so p = 0.3 ÷ 19.7 = 1.5%. That is one request in 66.
Why does a reasoning request cost so much?
It writes before it answers. The model produces a chain of intermediate steps, often longer than the final answer. Every token generated triggers a full pass through the network. Cost tracks the number of tokens produced, not how hard the question looks.
A second effect piles on top, and it is less visible. A long request holds GPU memory for its whole duration. The server then runs fewer requests in parallel, so the machine's fixed cost spreads over less useful work. The Joule paper names both mechanisms together: more tokens out, less parallelism.
Held to the same prompt, the measured gap is brutal. Jegham and colleagues push one long prompt through thirty commercial models. o3 and DeepSeek-R1 pass 33 Wh, more than 70 times what GPT-4.1 nano draws on that same prompt. The question stays fixed. Only the model changes.
Why does the published threshold range from 1.5% to 10%?
Because the two camps are not measuring the same object. The threshold hangs on one parameter, the ratio between a reasoning request and a text prompt, and that ratio is not pinned down to better than a factor of six.
| Where the ratio comes from | Ratio used | Doubling share |
|---|---|---|
| Joule, April 2026 (derived) | 11 × | 10% |
| Site data | 67 × | 1.5% |
| How Hungry is AI?, May 2025 | 70 × | 1.4% |
Joule's 10% threshold comes from a Monte Carlo simulation on open-weight models, under production assumptions: batching, high utilisation, current hardware. The 70-times ratio comes from reconstructing energy use out of commercial API response times, on frontier models as they actually run. Each method has its weak spot. The simulation assumes an ideal stack that nobody guarantees. The reconstruction guesses at hardware it never sees.
The arithmetic, by contrast, holds. A doubling share and a consumption ratio say the same thing: the ratio equals 1 + 1 ÷ p. Joule's "10% of daily requests can more than double total energy consumption" therefore implies a ratio of exactly 11, the floor of its own "more than an order of magnitude". That is the most conservative reading available, and we keep it as the upper bound of the threshold. The same authors argue that the most widely quoted public estimates overstate energy use by a factor of 4 to 20. They work for a vendor with an interest in the low number, and that is worth saying.
What does this weigh across all generative AI?
The site uses 1.1 trillion text prompts a year for text generative AI as a whole. At 0.3 Wh each, that comes to 330 GWh a year. Move 16.8 billion of those requests into reasoning mode, 1.5% of the total, and the figure doubles.
The extra draw is another 330 GWh, about a year of electricity for 77,000 French households. That shift needs no new data centre, no new users, not one additional request. It needs one request in 66 to change mode.
That is why the threshold deserves a number. Agent and reasoning modes became the default setting on several assistants during 2026. The shift does not show up in request counts, the only figure vendors publish. It shows up in the mix, which nobody publishes.
What this threshold does not say
- Electricity in use, and nothing else. No model training, no server manufacturing, no network. The boundary matches the rest of the site.
- The 67-times ratio is treated here as a constant. In practice it moves with the model, with how long the reasoning runs, and with the hardware underneath.
- The threshold assumes the rest of the traffic stays text. A growing share of image or video lifts the starting average and pushes the threshold down.
- No vendor publishes the real mix of its traffic. This calculation says when consumption tips, not where services stand today.
Sources
- Oviedo et al. : « Energy use of AI inference, efficiency pathways, and test-time scaling », Joule (avril 2026, DOI 10.1016/j.joule.2026.102430 ; préprint arXiv 2509.20241) : médiane de 0,31 Wh par requête, intervalle interquartile 0,16 à 0,60 Wh
- « How Hungry is AI? » (arXiv 2505.09598)
- Google Cloud (21 août 2025) : impact environnemental de l'inférence IA, 0,24 Wh, 0,26 mL d'eau et 0,03 g CO₂e pour le prompt texte médian de Gemini Apps
- OpenAI / TechCrunch : ChatGPT ~2,5 milliards de prompts/jour (2026) ; part de l'IA générative ~80 % (Demandsage / Statista). Calcul : ChatGPT ~900 Md/an ÷ 0,8 ≈ 1 100 Md pour toute l'IA générative.
- ADEME : consommation des appareils ménagers
Cite this figure
At 20 Wh per reasoning request against 0.3 Wh per text prompt, 1.5% of requests in reasoning mode is enough to double an AI service's electricity use. The threshold rises to 10% under the most conservative ratio published to date (11 times, Joule, April 2026). Calculation and sources on this page: howmanyprompts.com.
One question this calculation does not settle: which uses earn a reasoning mode, and which do fine without one. The studio ghis.fr works on that with small and mid-sized companies. Here we measure, and we stop there.