Jevons Paradox in AI: Why Cheaper Intelligence Is Fuelling More Compute
August 1, 2026·Aperta Res Research·
Share
Editorial illustration.
AI is becoming dramatically cheaper and more efficient. Instead of reducing the industry’s appetite for computing power, those gains are driving demand, investment and energy use to new heights.
Andy Jassy told analysts on 30 July that Amazon now expects to spend about $220 billion in cash capital expenditure during 2026. Its previous estimate had been closer to $200 billion. The extra cost, he said, largely reflects the rising price of memory chips .
Then he delivered the line that helps explain what is happening across the rest of the technology industry.
“But even at that amount, we will still not have enough capacity to meet all the demand we have in 2026. And I believe this dynamic will also be true in 2027 too.”
Within nine days of one another, four of the world’s largest technology companies reported their second-quarter results. Together, they had spent $170.1 billion on property, equipment and data centres in just three months .
Each company can now buy far more computing power for every dollar than it could a year ago. Yet all four either raised their spending plans or reaffirmed plans to keep spending more.
The pattern is not new. William Stanley Jevons described it in 1865 while studying Britain’s coal consumption . His name has been attached to it ever since.
Cheaper Compute Means More Compute
William Stanley Jevons, who was 29 when he published The Coal Question. Period photographic portrait.
Jevons was interested in steam engines. Britain was making them much more efficient, and many people assumed that greater efficiency would reduce the country’s use of coal.
Jevons argued the opposite.
As steam power became cheaper, engines became economical in places where they had previously been too expensive. They spread through mines, mills, factories and railways. Each engine used coal more efficiently, but the number of engines grew so quickly that Britain ended up burning more coal overall .
He was right about coal. Today’s technology companies are betting that the same principle applies to computing power.
Satya Nadella made the comparison directly on 27 January 2025, after the release of a low-cost Chinese AI model wiped hundreds of billions of dollars from chip valuations and pulled down shares in nuclear power, engineering and construction companies .
“Jevons paradox strikes again! As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can’t get enough of.”
As a business argument, the logic is straightforward. Lower prices will attract new users and make new applications economical. Demand will grow faster than the cost of serving each request falls. The correct response to cheaper AI inference, therefore, is not to build less capacity. It is to build more.
Economists add an important qualification.
Lower prices causing higher demand is ordinary price theory. Jevons paradox applies only when the increase in demand is large enough to outweigh the efficiency saving. When the rebound falls below that threshold, efficiency still reduces total consumption .
Halve the Cost of a Task and Total Compute Can Still Rise
Suppose the cost of running an AI task falls by half. For total computing consumption to rise, users must run more than twice as many tasks.
If demand increases from one task to 2.6 tasks, for example, the amount spent or consumed rises from 1.00 to 1.30, even though each task costs only half as much. If demand rises by only 50%, total consumption falls instead.
Cost Per Task Halves, Total Compute Rises 30%
Jevons’s 1865 mechanism, drawn for compute. The two shaded areas are in exact proportion.
price
half
Demand
cost per task
tasks run
compute used beforecompute used after, a larger area
Halve the cost of a task and the number of tasks run more than doubles. The shaded rectangle is then the larger of the two despite being shorter, so total compute rises. Had demand risen by only half, consumption would have fallen instead. The paradox lives entirely in that threshold.
Diagram by Aperta Res, applying to compute the argument set out in the argument in The Coal Question, with the condition economists attach to it set out by NPR. The curve passes through (1, 1) and (2.6, 0.5), which fixes its elasticity and puts the two areas in a 1.00 to 1.30 ratio. That 2.6-fold response is a drawing choice, not a measurement: the substantive claim is only that quantity must more than double when price halves. Total consumption is drawn as price times quantity, which holds while the price of a task tracks the compute it consumes. Versions of this diagram captioned “demand increases by more than half” state the threshold incorrectly, since a half-again rise in demand against a halved price lowers total consumption.
That threshold is the entire paradox. Halving the cost of a task does not automatically increase total compute. The number of tasks must more than double.
Everything that follows is an attempt to measure whether that is happening.
Diagram note: The example applies the argument in The Coal Question to computing. The 2.6-fold increase is illustrative rather than measured. The substantive condition is simply that demand must more than double when the price per task is halved.
Four Companies, One Quarter, $170 Billion
Amazon spent $53.1 billion during the quarter . Alphabet spent $44.9 billion, twice what it had spent a year earlier . Microsoft spent $41 billion, an increase of 69.4%, bringing its expenditure for the full fiscal year to $145.3 billion . Meta spent $31.08 billion .
Together, the four companies spent $170.1 billion on buildings, chips, servers and power equipment in 90 days. That is roughly $1.9 billion every day.
What Four Companies Spent on Infrastructure in Three Months
Capital expenditure, April to June 2026, in billions of US dollars. The whole strip is the quarter. Reported figures, not guidance, ordered alphabetically.
Alphabet
$44.9B
Amazon
$53.1B
Meta
$31.1B
Microsoft
$41.0B
$0B
$50B
$100B
$150B
$170.1B
$170.1 billion(sum) of buildings, chips and power equipment in ninety days, about $1.9 billion a day. No single company carries it: the smallest of the four still spent $31.1 billion. Each told investors in the same week that it intends to spend more next year.
Sources: Amazon Q2 2026 earnings call, CNBC on Alphabet, The New York Times on Microsoft, and Meta’s results release. Microsoft’s figure is its fiscal fourth quarter, which covers the same three calendar months. Meta’s includes principal payments on finance leases. Amazon’s is cash capital expenditure as given by its finance chief on the call. The total and the daily rate are sums and arithmetic on four separately reported budgets, not a jointly managed programme.
The spending is not being driven by a single outlier. Even the smallest budget among the four came to more than $31 billion. Each company also told investors that it expected to spend more.
The effect on cash flow was immediate.
Alphabet generated $39.1 billion from operations but spent $44.9 billion, leaving it with negative free cash flow of $5.9 billion for the quarter .
Amazon’s free cash flow over the previous 12 months swung from an inflow of $18.2 billion a year earlier to an outflow of $7.6 billion. Its purchases of property and equipment had increased by $66.1 billion .
Meta was left with quarterly free cash flow of just $784 million .
Then the spending forecasts rose again.
Alphabet lifted its 2026 capital expenditure range from $180–190 billion to $195–205 billion. Analysts had expected about $188 billion. The company also warned that spending would rise significantly again in 2027 .
Its finance chief, Anat Ashkenazi, said the increase reflected “an acceleration in the delivery of capacity to meet growing demand” .
Meta issued full-year guidance of between $130 billion and $145 billion .
Investors did not respond to all four companies in the same way.
Microsoft’s Azure revenue grew 43%, compared with expectations of about 40%, while quarterly profit rose 31.6% . Amazon Web Services grew 37%, its fastest rate in 18 quarters .
Alphabet’s cloud revenue rose 82%, yet its shares fell about 7% the day after the earnings call . Meta reported a 14% decline in net income, while its operating margin dropped from 43% to 31% .
There was another complication in the numbers. Two companies reported profits that had been lifted substantially by non-operating investments.
Amazon recorded $53.4 billion in non-operating pre-tax income, largely because of its stake in Anthropic . Alphabet reported $99 billion in other income covering holdings that included Anthropic and SpaceX .
The spending is real. The question is how much of the eventual return will come from the businesses being built and how much will come from the rising value of investments surrounding them.
What a Single AI Task Costs
A frontier-scale model answering a typical query uses about 0.31 watt-hours of electricity, according to a peer-reviewed estimate published in April 2026. The study examined models with more than 200 billion parameters under production-serving conditions .
The middle half of queries used between 0.16 and 0.60 watt-hours.
Google remains the only major provider to have measured energy consumption directly across its own production fleet and published the result. In May 2025, the median Gemini text prompt used 0.24 watt-hours of electricity, produced 0.03 grams of carbon dioxide equivalent and consumed roughly five drops of water .
Only 58% of that energy went directly to the AI accelerator. Host processors and memory accounted for another 25%. Idle equipment and data-centre overhead consumed the rest.
Google has not updated the figure since then and has cautioned that it will change as new generations of models are introduced .
Many widely repeated estimates are considerably higher. The April 2026 study found that calculations based on non-production environments can overstate AI energy use by between four and 20 times .
Still, the energy cost of a query is not fixed.
Allow the same model to reason for 15 times longer and median consumption rises about thirteenfold, from 0.31 watt-hours to 3.91 watt-hours .
What a Single AI Task Costs in Electricity
Watt-hours per query on frontier-scale models, 2026. Dots are medians, bars are interquartile ranges. Logarithmic scale.
Reasoning query, 15x longerThe same model, reasoning 15x longer3.91 Wh2.15–7.05
Gemini prompt, May 2025Median Gemini prompt, measured May 20250.24 Wh
Letting a model reason for longer costs thirteen times the energy of a typical query. The unit is not a single number, and cheaper tokens move users up the ladder.
Sources: medians and interquartile ranges from Joule, April 2026, modelling frontier-scale models above 200 billion parameters on H100 nodes under production serving assumptions; the 13x rail is that paper’s own comparison of its own two scenarios. The hollow mark is Google’s in-situ measurement of its own median Gemini text prompt in May 2025, the only first-party production figure any large provider has published, and it describes an earlier generation of models. Separately, the French regulator Arcep and PEReN measured 22 models directly in May 2026 and found that switching on a reasoning mode raises energy use by up to 92%.
In May 2026, France’s telecommunications regulator commissioned direct measurements of 22 models on a national supercomputer. It found that enabling a reasoning mode increased energy consumption by as much as 92%. On code-generation tasks, however, performance improved by as much as 849% .
That trade-off matters. The direction of the industry will be determined by which number customers care about more: the additional energy consumed or the additional capability gained.
So far, the monetary cost of AI has fallen much faster than its energy cost.
When GPT-3 became publicly available in November 2021, it was the only model capable of scoring 42 on the MMLU benchmark. Access cost about $60 per million tokens .
By 2024, the cheapest model reaching the same score charged only $0.06 per million tokens.
That is a thousandfold decline in the price of reaching the same level of performance.
Holding capability constant is the clearest way to measure how quickly the price of intelligence is falling, because the models themselves keep changing. Epoch AI compared prices across six benchmarks and found annual declines ranging from ninefold to 900-fold, with a median decline of about 50-fold .
GPT-4, for example, cost roughly $40 per million tokens. Fourteen months later, Gemini 1.5 Flash exceeded its score on a benchmark of PhD-level science questions at an estimated price roughly 300 times lower .
The frontier tier is different.
The cheapest way to reproduce yesterday’s capabilities has collapsed in price. The cost of buying the best available capability has fallen much more slowly.
OpenAI’s most capable model in August 2026 was listed at $10 per million input tokens and $60 per million output tokens. The model it replaced had been listed at $12.50 and $75 respectively .
The Price of Intelligence, Then and Now
US dollars per million tokens, weighted three to one input to output. The first two lanes hold a benchmark score fixed. The third is the best model available, then and now. Logarithmic scale.
MMLU score of 42A score of 42 on the MMLU benchmarkabout 1,000x cheaper
$60.00
GPT-3, Nov 2021
$0.06
Llama 3.2 3BLlama 3.2 3B, 2024
GPT-4 science scoreGPT-4's score on PhD-level science questionsabout 300x cheaper
$40.00
GPT-4, Mar 2023
~$0.13 derived
Gemini 1.5 FlashGemini 1.5 Flash, 14 months later
The frontier tier itselfThe frontier tier itself, whatever it can currently doabout 1.2x cheaper
$28.13 to$22.50
gpt-5.5 to gpt-5.6-sol, Aug 2026
$0.01
$0.1
$1
$10
$100
Yesterday’s answer has become almost free. The best answer costs about what it always did. Buyers take the second deal, and take more of it.
Sources: Andreessen Horowitz for the MMLU 42 lane and Epoch AI for the GPT-4 science lane, whose $0.13 mark is $40 divided by the 300x Epoch reports, plotted at that exact ratio and printed rounded. The frontier lane is OpenAI’s published list prices as of August 2026, weighted three to one input to output to match the historical lanes; list prices exclude negotiated enterprise discounts. The two frontier marks nearly touch because the prices nearly match, and the models are not equivalent: gpt-5.6-sol is the more capable of the two, which is the reason its price did not fall.
A full generation of progress reduced the cost of the best available answer by about one-fifth. Over the same period, the cost of producing a fixed level of performance had fallen by orders of magnitude.
Yesterday’s answer is becoming almost free. The best answer available today still commands a premium.
Customers are buying both, but the greatest pressure on infrastructure comes from their appetite for the second.
The International Energy Agency describes the physical side of this shift in unusually direct terms :
“Measured per individual task, the energy efficiency of AI is improving at a rate unprecedented in energy history.”
The Unit Falls, but the Total Climbs
At Google’s developer conference in May, Sundar Pichai revealed how quickly the volume of AI activity had grown.
Google had processed 9.7 trillion tokens a month two years earlier. By the previous year’s keynote, that figure had reached roughly 480 trillion. In May 2026, it had passed 3.2 quadrillion tokens a month .
“I never imagined I’d say the word quadrillion in an I/O keynote. But here we are.”
Over two years, Google’s monthly token volume increased approximately 330-fold.
During the 12 months to May 2025, Google reported that the energy used by its median prompt had fallen by a factor of 33 . Over roughly the same period, the company’s monthly token volume rose about 49-fold .
Down Per Unit, Up in Total, in the Same Year
Change over roughly twelve months, as a multiplier. Logarithmic scale, anchored at no change.
per unit, fell
in total, rose
Energy per Gemini promptEnergy used by the median Gemini prompt33x less
Google reports a 33-fold fall
All data centre electricityElectricity used by all data centres+17%
IEA reports +17%
AI data centre electricityElectricity used by AI data centres+50%
IEA reports +50%
Google tokens per monthTokens processed by Google each month49x more
9.7T to 480T a month
÷10
×1
×10
×100
Detail, linear scale: the marked slice above, magnified
All data centre electricity+17%
AI data centre electricity+50%
0%
+10%
+20%
+30%
+40%
+50%
Left edge is ×1.00, meaning no change. Right edge is ×1.55.
Volume rose about fifty times while the energy per prompt fell about thirty-three times. Electricity grew far less than volume, which is efficiency doing real work, and it still grew.
Sources: Google Cloud for energy per prompt and Sundar Pichai’s I/O keynote figures for token volume, both covering May 2024 to May 2025. IEA for both electricity figures, covering calendar 2025. The periods are close but not identical. Every bar is a ratio computed from the figures printed beneath its label. Energy per prompt and total tokens measure different quantities and are not multiplied together anywhere.
That is Jevons paradox in its clearest form.
The amount of energy used by each prompt fell dramatically. The number of prompts and tokens grew even faster.
Global electricity use did not rise by anything close to 49 times, which shows that efficiency improvements were doing real work. Without them, the growth in demand would have been impossible to serve.
But electricity consumption still increased.
Total electricity demand from data centres grew 17% during 2025. Electricity use at AI-focused data centres rose 50% .
On these figures, the rebound passed the threshold required for Jevons paradox within a single year: energy use per unit fell, yet aggregate consumption climbed.
The IEA expects the trend to continue. It forecasts that global data-centre electricity demand will almost double between 2025 and 2030, rising from 485 terawatt-hours to about 950 terawatt-hours .
At that point, data centres would account for approximately 3% of global electricity demand. Electricity consumption specifically associated with AI is expected to triple over the same period.
Where the Limit Actually Bites
The immediate constraint is no longer the amount of money technology companies are willing to spend.
It is whether the necessary equipment and electricity can be supplied quickly enough.
Memory is the most urgent bottleneck. The IEA expects shortages of high-bandwidth memory to continue through the end of 2027 . Federal Reserve economists have also described memory availability as a binding constraint as AI models grow larger .
Both companies that raised their capital budgets in July pointed to the same problem. Amazon added another $20 billion to its 2026 forecast partly because memory had become more expensive .
Power is the slower and more difficult constraint.
Orders for gas turbines rose 70% in 2025 . With grid connections taking years in some locations, data-centre operators are increasingly considering or building their own electricity generation.
The IEA expects data centres to install between 15 and 27 gigawatts of on-site gas generation by 2030, most of it in the United States. It also expects them to deploy between 20 and 25 gigawatts of battery storage .
At that point, the comparison with coal is no longer merely rhetorical.
Efficiency gains inside the chip are turning into turbine orders at the substation.
The pressure is already appearing in electricity markets. Across PJM, the largest grid operator in the United States, capacity costs rose 398% in the first quarter of 2026. Transmission costs rose by about 5% .
Monitoring Analytics, PJM’s independent market monitor, identified data-centre load growth as “the primary reason for recent and expected capacity market conditions” .
It estimated that the two most recent capacity auctions had added $16.6 billion to customers’ bills .
Its conclusion was blunt :
“The price impacts on customers have been very large and are not reversible.”
Alphabet has already had to work around the shortage. Unable to finish its own infrastructure quickly enough, it told investors that it would rent capacity from third-party providers during the third quarter .
The IEA has reached a similar conclusion from the supply side. Bottlenecks throughout the equipment, power and construction chain have made the most aggressive short-term demand scenarios less likely, not because demand has disappeared, but because the industry may not be physically able to build fast enough .
The Plan to Leave the Planet
Some companies are now treating the limits of Earth-based infrastructure as a permanent problem rather than a temporary shortage.
In January 2026, SpaceX submitted an application to the Federal Communications Commission for a constellation of as many as one million satellites. The system would form the basis of an orbital AI data centre .
The scale of the proposal is the point.
SpaceX says the constellation could generate 120 gigawatts of power and support tens of millions of frontier-class graphics processors, potentially as many as 100 million. Building it would require at least a twentyfold increase in the company’s launch capacity .
For comparison, the IEA expects every data centre on Earth to consume about 950 terawatt-hours of electricity in 2030 . A system generating 120 gigawatts continuously would produce more energy than that over a year.
Elon Musk explained the logic at an event in Austin in March .
“Increasing power on Earth becomes harder over time and more expensive over time, but in space it becomes actually cheaper and easier over time.”
He has suggested that the economics could cross over within two or three years .
In June, Musk and SpaceX satellite engineering director Ian Dahl presented an early version of the hardware, known as the AI1 satellite. Musk characterised it as an extension of systems the company had already developed .
“There’s not some magic that’s necessary that doesn’t exist. A lot of this is technology we’ve already made for Starlink V3 satellites. Basically, we don’t think this is a super hard problem.”
The financial arithmetic is less reassuring.
Quilty Space estimates that a Starlink V3 satellite costs about $1 million. At that price, a constellation of one million satellites would require close to $1 trillion for the spacecraft alone. Ground systems could add another $100 billion .
Producing enough processors would require an industrial project of its own. SpaceX, Tesla and Intel have partnered on Terafab, a proposed 10-million-square-foot facility in Austin. It is expected to open in 2029 and could cost as much as $119 billion .
SpaceX is not alone.
Blue Origin filed plans in March for 51,600 data-centre satellites under Project Sunrise, with deployment expected to begin in late 2027. Google is pursuing Project Suncatcher in partnership with Planet Labs .
A CNBC review of the sector in June concluded that the economics remain challenging, even as launch prices fall . For now, the strongest argument for putting data centres in orbit is not that space has become cheap. It is that building on Earth is becoming more difficult.
Every one of these projects is a bet that the terrestrial constraint will not go away.
What the Buildout Costs the Average Worker
Most people will never buy an AI accelerator or reserve space in a data centre.
They will still encounter the infrastructure boom in two places: their electricity bill and the labour market.
PJM’s latest capacity auction is expected to add $6.3 billion to the electricity bills of households and businesses across 13 states and the District of Columbia within three years. The grid operator attributes much of the increase to demand from data centres .
Across the two previous auctions, data-centre load was associated with an estimated $16.6 billion increase .
The industry disputes the idea that data centres are solely or directly responsible for higher residential bills, and that argument deserves to be included.
A review by the consultancy E3 found no historical evidence that data centres had raised residential electricity prices under existing rate structures .
It attributed about half of the recent PJM capacity-price increase to data-centre demand. The remainder came from market-design changes, limited supply and the retirement of power plants .
The study also noted that Texas and Virginia had absorbed substantial data-centre growth while experiencing relatively modest rate increases. California and New York had seen larger increases even while electricity demand declined .
The review was promoted by the data-centre industry’s trade association, and one of the studies it relied on was funded by Amazon . Those facts do not invalidate its findings, but they should be considered alongside them.
The evidence on employment is less developed than the evidence on electricity.
Federal Reserve economists have found relatively strong productivity in industries with high exposure to AI, including information, finance and professional services .
However, productivity trends have remained broadly consistent when measured at different levels of the economy. The researchers interpret that pattern as a sign that improvements visible inside individual companies or tasks have not yet translated into a clear economy-wide acceleration .
They were similarly cautious about employment.
The earliest signs suggest that AI may be affecting younger workers through slower hiring rather than large-scale layoffs. Unemployment among people aged 20 to 24 is one of the indicators worth watching, even though overall unemployment remains moderate by historical standards .
The infrastructure boom is already large enough to appear in the national accounts.
Investment in AI-related data-centre construction, hardware and networking reached approximately 0.8% of US gross domestic product in the first quarter of 2026 .
Computing infrastructure more broadly represented about 1.5% of GDP. Between 2015 and 2022, the comparable average was roughly 0.7% .
The combined capital expenditure of five technology companies now exceeds global investment in oil and gas production .
A household in Ohio may never knowingly use the AI service being built. Through its electricity bill, however, it may still help finance the capacity required to run it.
What Jevons Paradox Does Not Promise
Jevons described what happened to coal. British coal consumption continued rising for decades after the publication of his book, just as he had predicted .
His argument said nothing about who would make money from burning it.
That distinction matters because Jevons paradox has increasingly become a defence of the technology industry’s capital spending.
If demand for computing power grows faster than its price falls, then almost any amount of new infrastructure may eventually find a use. But a use is not the same thing as a profitable return.
NPR has pointed out the missing step in Nadella’s argument: it assumes that profits will rise alongside demand. That outcome is far from guaranteed .
The final week of July suggested that investors had begun making the distinction.
Microsoft exceeded expectations on the number that most clearly connects its infrastructure spending to revenue: Azure growth. Investors rewarded it .
Alphabet reported 82% growth in cloud revenue but also went cash-flow negative for the quarter . Its shares fell about 7%.
Meta’s AI spending is embedded within an advertising business that already existed. The company reported a 14% decline in net income .
Two assumptions support the entire investment cycle.
The first is that the cost of a fixed level of AI capability will continue to collapse. Epoch AI has cautioned that some of the fastest declines it has measured are based on less than a year of data .
The second is that demand will continue growing even faster. Jassy says it will, and the order book extending into 2028 provides some evidence that he is right .
For 160 years, greater efficiency has repeatedly led to greater consumption.
Whether that pattern will also produce an adequate return on $170 billion of quarterly investment is a separate question.
The companies writing the cheques will discover the answer long before the economists do.
Achraf Rachidi
Independent researcher. Aperta Res was born from a simple frustration: too much noise, not enough signal. The goal is transparent, data-grounded analysis that cuts through complexity.
The central bank of central banks now rates AI's circular financing a threat to financial stability. Reported profits increasingly rest on paper markups and money that loops between the same handful of companies, even as their free cash flow heads toward zero.