0:00
Every conversation in the AI industry, or so it seems, is about GPUs and AI accelerators.
Every conversation in the AI industry, or so it seems, is about GPUs and AI accelerators.
Without the chips, without silicon, there is no AI.
But if we zoom out just a little bit, the next bottleneck isn’t HBM or GPUs, the next bottleneck is located in the much older, much slower and much more physical world of the electrical grid.
If you want to oversimplify it, a GPU cluster is an industrial machine that turns electricity into tokens.
Every training run, every inference request, eventually becomes power demand.
Which means, the AI race is also becoming a power race.
And this is where things get difficult, because AI companies are trying to grow at the speed of software and silicon, while the electric grid grows at the speed of substations, transmission lines, transformers, gas turbines, permits and utility planning.
In December 2025, SemiAnalysis forecast US AI power demand to grow to more than twenty-eight gigawatts by 2026, up from only roughly three gigawatts in 2023.
At the time our forecast looked aggressive, but it turned out to be pretty much on the money.
And the trajectory from here is even steeper: the expected new datacenter gross power demand is rising to 84 gigawatts by 2030.
That’s not just a large number on a chart, it’s a fundamental shift in how we have to think about AI infrastructure.
A single gigawatt-scale datacenter campus is not just another office building that needs a utility connection, it is closer to adding the power demand of a small city, all concentrated in one place, and on a timeline that the grid was never really designed to support.
An AI cloud can generate somewhere between ten to twelve billion dollars of revenue per gigawatt, per year.
That’s ten to twelve million dollars per megawatt, annually.
Which means that getting a two-hundred-megawatt AI cluster online six months earlier is worth roughly a billion dollars. A billion dollars. For six months.
Once you internalize that number, every strange decision in this industry suddenly starts to make sense.
Why a company would rent gas turbines instead of buying them.
Why it would build on a state border.
And why it would deliberately pay more per kilowatt-hour, forever, only to get it faster.
In this video we will take a closer look at AI energy demand to figure out why and how the industry is moving behind the meter.
And what that actually means.
Because AI is running out of power.
To understand what behind-the-meter power actually is and why it matters, we first have to understand what the grid normally does for a datacenter.
In the traditional world, the datacenter connects to the electric grid, the utility delivers power and the meter measures how much electricity the datacenter consumes.
The datacenter operator still needs backup systems, batteries, generators and redundant feeds, but the main job of producing and balancing electricity belongs to the grid.
It is a centralized system, built over decades, with utilities, power plants, substations, transmission lines and market operators all working together to keep supply and demand in balance.
Behind-the-meter changes that relationship.
Instead of waiting for the grid to deliver all the required power, the datacenter brings some or even most of the generation directly to the site.
That can mean gas turbines, reciprocating engines, fuel cells, batteries or hybrid systems that combine onsite generation with a limited grid connection.
The industry has a name for this: BYOG.
Bring Your Own Generation.
And the important point is that these sites won’t be disconnected from the grid forever.
In many cases, behind-the-meter is a bridge.
The datacenter needs power in 2027 or 2028, while the full grid interconnection might not arrive until 2030.
So instead of waiting, the developer builds the power plant next to the AI factory, and demotes it to backup equipment once the grid utility finally shows up.
And that name, AI factory, is important.
Because this is really what these new datacenters are becoming.
They are no longer just buildings full of servers that host websites and databases, they are industrial facilities designed to manufacture intelligence at scale.
And they are running at full speed almost all of the time.
For a long time, I struggled with the term “AI factory”, it seemed a bit too much like marketing cope. But that changed.
The raw material is electricity, the machines are GPUs and AI accelerators, and the output is model training, inference, tokens, code, videos, images and automation.
Once you look at it like that, it becomes much easier to understand why power availability is suddenly one of the most valuable inputs in the entire AI supply chain.
But there’s a problem with the grid.
No, the US grid is not failing. But it is full. Look at Texas.
Every month, tens of gigawatts of datacenter load requests pour into ERCOT, the Electric Reliability Council of Texas.
But in the twelve months leading to March 2026 only around two gigawatts of generation were approved.
That basically tells you all you need to know.
Demand arrives by the tens of gigawatts, but approvals happen by the gigawatt.
When we look at the entire country, our estimate is that roughly a terawatt of load requests have been submitted to US utilities and grid operators. One terawatt.
For reference, the entire US grid peaks somewhere in the range of seven to eight hundred gigawatts, which means, the requests now exceed the entire system.
Yes, you heard that right.
The power requests exceed the entire US grid.
But not all of those requests are actually real.
And that is a big problem.
What’s happening in the interconnection queue is a textbook prisoner’s dilemma.
If every developer submitted one honest request for one site they actually controlled, the queue would move quickly and everyone would get connected faster.
But nobody can afford to be the one honest player.
So, developers submit speculative requests to multiple utilities simultaneously, hedging across regions, hoping one comes through.
In October 2024, AEP Ohio was sitting on thirty-five gigawatts of load requests, and sixty-eight percent of them didn’t even have land control. No land. No datacenter.
Just a request for power.
Those phantom requests clog the queue for real projects.
Which makes everyone in the industry even more anxious.
Which makes them submit more speculative requests. The cycle feeds itself.
Meanwhile the grid is slow by design, and for good reasons.
Electricity supply and demand have to match almost perfectly, every second, or you get blackouts for millions of people, as the Iberian Peninsula found out in April 2025.
And every large new load or new power plant triggers deep engineering studies to make sure it won’t destabilize the network.
In some regions the grid topology now changes so fast that load studies go obsolete before they’re even finished.
As a result, the timeline from interconnection request to commercial operation now stretches to around five years for most generation types. Five years.
In a business where six months is worth a billion dollars.
We estimate the grid is adding around 15 gigawatts of net-new of so-called Effective Load Carrying Capability, or ELCC capacity, per year, with that number rising toward 20 gigawatts or more later this decade.
But datacenters are not the only source of new demand.
The same grid also has to serve factories, homes, EV charging, semiconductor fabs, industrial growth and normal economic expansion.
AI is competing for capacity inside a system that was already becoming tighter.
But notice that word: ELCC.
Effective load carrying capability.
It exists because nameplate capacity and useful capacity are not the same thing.
You can add a lot of solar, wind and batteries on paper, and those resources absolutely matter.
But a gigawatt of nameplate solar is not the same thing as a gigawatt of firm power available exactly when the grid is under stress.
Solar produces power when the sun shines. Wind depends on weather.
Batteries can shift energy across time, but only for as long as their duration allows.
A four-hour battery can solve a four-hour problem, but it does not solve a multi-day capacity problem.
So ELCC tries to answer a very simple but very important question: how much does this specific resource actually help the system serve load during the moments that matter for reliability?
A gas plant, a solar farm and a battery can all have the same nameplate capacity, but the grid will not treat them as equally reliable.
Their accredited capacity depends on region, weather, load patterns, reserve requirements and the exact stress events the system is trying to survive.
This is why datacenter power is much harder than just saying “build more renewables.
” Renewables are important, and they will remain important, but AI datacenters need firm power all day round.
Training does not only happen when solar output is high.
Inference demand does not stop at sunset.
And if a grid is short during the evening ramp, a datacenter that wants hundreds of megawatts of uninterrupted power becomes a very difficult load to serve.
The bottleneck is not only energy, it is firm capacity, in the right place, at the right time, with enough transmission and enough reserve margin behind it.
The best way to think about all of this is grid headroom.
Headroom is the spare accredited capacity a power market has left after it covers its own peak demand and required reserve margins.
In simple terms, it’ss the amount of room the grid has for large new loads before reliability starts to suffer.
If there is headroom, a datacenter can connect.
If there is no headroom, the utility can still study the project, negotiate with the developer and maybe even provide a tentative date, but the physical system does not really have spare capacity.
Our view is that available grid headroom is already approaching zero, and turns negative by 2027.
That does not mean that the entire US grid collapses in 2027.
What it means is that in more and more regions, the grid no longer has enough spare accredited capacity to comfortably add huge new datacenter loads while still meeting reliability requirements.
And once headroom goes negative, every new large load becomes a fight over who gets capacity, who pays for upgrades and who takes the risk if the timeline slips.
This affects everything, not just AI datacenters. It’s the entire industy.
And it’s something datacenter operators are already seeing.
A utility might initially tell a developer that it can serve a 500-megawatt load ramp in 2027.
But then the real constraints appear.
Main power transformers are delayed.
High-voltage breakers are delayed.
Substation upgrades take longer.
Transmission studies reveal network constraints.
New generation is not arriving fast enough.
And eventually the timeline moves from 2027 to 2029, or the promised load ramp gets revised down, or the developer is asked to post large security deposits and take-or-pay commitments to fund the generation needed to serve the site.
For a normal commercial customer, that might be painful but manageable.
For an AI lab, it can be existential.
If compute is the lifeblood of the business, then a two-year power delay is not just a construction delay, it is delayed training, delayed inference, delayed revenue, delayed product releases and delayed model progress.
In frontier AI, the value of compute can be so high that cheap power in 2030 may be worse than expensive power in 2027.
In 2024, xAI did something the datacenter industry did not think was possible.
They stood up a hundred-thousand-GPU cluster in just four months.
Construction started in June 2024 and training in September.
A lot of innovations went into that, but the energy strategy was the most impressive part, and it was almost insultingly simple: xAI didn’t ask the grid.
They generated onsite, using truck-mounted gas turbines and engines.
Not power plants in the traditional sense. Generation on wheels.
The specific choices are worth talking about in more detail, because they became the template that everyone else copied.
xAI used small modular sixteen-megawatt turbines from Solar Turbines, a Caterpillar subsidiary.
Sixteen megawatts is small by power plant standards, and that’s the point: it’s small enough to fit on a standard long-haul truck.
You drive it in, you set it down, you’re generating within weeks. Then the second move.
Elon didn’t even buy the turbines.
He rented them, from Solaris Energy Infrastructure, specifically to bypass equipment lead times.
Alongside those, he leased VoltaGrid’s fleet of truck-mounted gas engines, thirty-four systems at Colossus 1, built around Jenbacher high-speed engines.
xAI was renting the power plant.
Not because of the cost, but because buying one would have taken too long.
And then the third move, which is the one we keep coming back to.
When xAI needed permits for gigawatt-scale generation, they picked a site on the border between two states.
Two jurisdictions, two permitting authorities, two chances at a fast yes.
Tennessee couldn’t deliver on time. Mississippi could.
Site selection as regulatory arbitrage.
Today, xAI has more than five hundred megawatts of turbines deployed near its datacenters.
And one by one, everyone else has followed.
In October 2025, OpenAI and Oracle placed the largest order for onsite gas generation ever recorded: a 2.
3-gigawatt plant in Texas.
The market for onsite gas is now in triple-digit annual growth.
In our Datacenter Industry Model, we built a building-by-building tracker of every site deploying it, and the result actually surprised us: twelve different suppliers have each secured more than four hundred megawatts of US datacenter orders. Twelve.
In a market that barely existed in 2023.
And some of those names are genuinely strange.
Doosan Enerbility, the Korean industrial giant, timed its H-class turbine launch perfectly and booked a 1.
9-gigawatt order to serve xAI.
Wärtsilä, historically a ship engine manufacturer, worked out that the same engines that push cruise ships across oceans can power AI clusters, and has signed eight hundred megawatts of US datacenter contracts.
And then there’s Boom Supersonic.
Yes, the supersonic passenger jet company. They announced a 1.
2-gigawatt turbine contract with Crusoe, and they’re treating the margin from datacenter power generation as, essentially, another funding round for their Mach 2 airliner.
That is the state of this market.
A supersonic jet startup is financing itself by selling power generation to AI datacenters.
So, what are these companies actually buying?
Because “onsite gas” hides a lot of very different machines.
There are broadly three categories. First, gas turbines.
These run on the Brayton cycle: compress air, burn fuel in it, push the hot gas through a turbine.
Within turbines, the key differentiator is inlet temperature.
Higher temperature means higher efficiency and faster ramp, but higher cost and higher maintenance.
The most interesting category for datacenters is the aeroderivative.
And the name tells you exactly what it is: a turbine derived from aero engines.
It’s a jet engine bolted to the ground.
GE Vernova’s aeros come from GE jet engines.
Mitsubishi Power’s from Pratt & Whitney.
Siemens Energy’s from Rolls-Royce.
Because a jet engine is already designed to make enormous power in a package light enough to fly, adapting it for stationary use is almost easy.
Extend the shaft, bolt on a generator, add intake and exhaust mufflers, feed it gas.
That’s also why Boom Supersonic could pivot into this business so fast, most of their engineering carries straight over.
Aeros run about thirty to sixty megawatts per unit, ramp from cold to full output in five to ten minutes, and cost somewhere around seventeen hundred to two thousand dollars per kilowatt all-in.
Lead times: eighteen to thirty-six months, and climbing.
Industrial gas turbines, or IGTs, work on the same cycle but are designed from scratch for stationary use instead of adapted from aviation.
Lower inlet temperatures, simpler designs, cheaper to service, less efficient, slower to ramp with about twenty minutes.
Five to fifty megawatts, roughly fifteen hundred to eighteen hundred dollars per kilowatt.
Second, reciprocating engines, or RICE.
Sounds yummy, but you shouldn’t eat it.
These are car engines, scaled up to absurdity.
An eleven-megawatt engine can be over fourteen meters long.
High-speed engines run around fifteen hundred RPM and produce three to five megawatts; medium-speed engines run around seven hundred fifty RPM and produce seven to twenty megawatts, with lower mechanical stress and lower maintenance costs.
They ramp in about ten minutes, cost seventeen hundred to two thousand dollars per kilowatt, and handle heat, dust and dirty fuel better than turbines do.
They also run at much lower temperatures, six to seven hundred degrees Celsius, versus turbine inlet temperatures.
That dramatically reduces their need for exotic alloys, which is going to matter in a minute. Third, fuel cells.
Bloom Energy’s solid-oxide fuel cells were a niche product until very recently.
They generate power with no combustion at all: oxygen is electrochemically reduced to oxide ions, which flow through a ceramic electrolyte and combine with hydrogen stripped from methane.
Out comes water, CO2, and electricity.
No combustion means no meaningful air pollution beyond CO2, which means EPA permitting is dramatically simpler.
That’s why you see them installed near office buildings and population centers.
And installation is fast: precast pads, drop in modules, do the electrical work, done in weeks.
The catch is cost: three thousand to four thousand dollars per kilowatt, roughly double a turbine.
And the individual stacks only last five to six years before they have to be replaced, which accounts for about sixty-five percent of service cost.
Now, here’s the thing that tells you everything you need to know about the current market.
In principle, a developer should carefully select the optimal technology for their site. In practice?
Whoever has an open order book and a credible date wins the deal, almost regardless of the specs.
You can see it in the hardware.
Meta and Williams built a behind-the-meter plant in Ohio called Socrates South, and the equipment list is a patchwork: three Solar Titan 250 IGTs, nine Solar Titan 130s, three Siemens SGT-400s, and fifteen Caterpillar 3520 fast-start engines.
Four different product lines from three manufacturers.
Nobody designs a power plant that way on purpose.
That is the design pattern of “´we will deploy literally whatever we can get on time.
” And here’s what most people underestimate, and it’s the reason onsite power is more expensive than grid power.
The US electric grid delivers about 99. 93% uptime.
And it does that by being enormous: thousands of generators, hundreds of transmission lines, market mechanisms balancing all of it in real time.
If one plant trips, there are a thousand others.
When you go behind the meter, you have to reproduce that reliability with one power plant, serving one customer.
And the only way to do that is to overbuild.
Vendors typically insist on at least N+1, which means keeping enough spare generation to survive one unit failing without losing output.
Better still is N+1+1: enough spare to survive a failure and still take units offline for scheduled maintenance.
It’s the equivalent of driving with a spare tire and a repair kit.
But what does that mean physically?
Take a two-hundred-megawatt datacenter served by eleven-megawatt reciprocating engines.
You deploy twenty-six of them, for 286 megawatts of nameplate power.
Under normal operation, twenty-three run at about eighty percent load.
If one dies, the remaining twenty-two ramp to eighty-two percent and nothing happens.
Three engines stay free for maintenance rotation.
Or do it with thirty-megawatt aeroderivatives: nine units, 270 megawatts nameplate.
Seven run at ninety-five percent for best efficiency, the eighth starts when one trips, the ninth stays in reserve.
And in hot climates like the American Southwest, derating means you might need ten or eleven units instead of nine, because turbines simply produce less when the air is hot.
Crusoe’s Abilene site for Oracle and OpenAI runs a version of this: ten turbines, five GE Vernova LM2500XPRESS aeros and five Solar Titan 350s, for 360 megawatts of nameplate. Vantage is building a 1.
4-gigawatt campus in Shackelford County, Texas and deploying 2.
3 gigawatts of VoltaGrid systems to serve it.
That’s a sixty-four percent overbuild. Roughly 1. 4 to 1.
5x of that is standard over-provisioning for cooling and PUE, which you’d see on a grid-connected Texas site too.
The remaining ten to seventeen percent is pure redundancy. Pure insurance.
And that overbuild ratio is precisely why onsite gas power costs are, in most cases, structurally more expensive than power delivered by the grid.
That is important to remember, because it goes against the way this story usually gets told.
Behind-the-meter is not cheaper. It is earlier.
Companies are paying a premium, knowingly, to buy time.
And there’s one more problem.
AI training load is unpredictable, megawatt-scale surges and dips on a sub-second basis.
Power systems absorb that with inertia, the stabilizing effect of heavy spinning objects.
If frequency wanders too far from sixty hertz, breakers trip and equipment malfunctions.
These sites bolt on inertia: synchronous condensers, which are generators spun up as motors to absorb and supply reactive power for a few seconds.
Flywheels, which buffer real power for five to thirty seconds.
Or batteries, providing synthetic inertia through very fast inverter control.
VoltaGrid pairs its engine fleets with synchronous condensers.
Bergen bundles flywheels.
xAI, predictably, uses Tesla Megapacks.
Historically, datacenters were built around extremely high uptime.
The classic enterprise and cloud model was to connect to a strong grid substation, use redundant feeds, add backup generators and batteries, and design the entire system around three, four or five nines of availability.
That made sense for cloud regions, banking systems, enterprise workloads and internet infrastructure where downtime was extremely expensive and hard to tolerate.
But AI changes some of those assumptions.
Training workloads can often tolerate lower availability if the system is designed around checkpointing and recovery.
And honestly, large GPU clusters are unreliable enough on their own that the power plant is not the weakest link.
Inference systems can route around failed nodes, especially when traffic is distributed across many servers.
And some AI-specific datacenters are being built with lower redundancy targets than traditional cloud regions, because the priority is no longer perfect uptime at any cost.
This matters because redundancy is the most expensive part of behind-the-meter power.
And if the tenant accepts lower uptime, the economics change completely.
And we can see that style of thinking in practice.
At both Abilene and Memphis, the training clusters were built without diesel generator backup at all. Not reduced, absent.
That directly cuts capex.
The reasoning is that a training job doesn’t need five nines of uptime, and once the grid connection eventually arrives, the gas turbines themselves become the backup.
Which is also why fast-ramping equipment like aeros gets preferred: a turbine that can go cold-to-full in five minutes has a second career as an emergency generator.
That’s the bridge power model.
Start generating early, run a workload that tolerates it, then demote your power plant to backup when the utility finally arrives.
Everything so far has been our case for behind-the-meter.
Fast, flexible, available now, worth the premium.
When we published that case, we got some pushback.
One counterargument is actually strong, and we want to give the attention it deserves.
It isn’t about emissions.
It isn’t about cost of capital. It’s about people.
A power plant is not a product you buy, it is an organization you have to run.
The grid’s three nines of uptime aren’t delivered by hardware; they’re delivered by a mature, century-old labor system.
Plant operators, certified control room staff, high-voltage electricians, turbine field service engineers, millwrights, welders, instrumentation techs.
When a datacenter goes behind the meter, it isn’t just buying generators.
It is quietly signing up to staff and operate a power plant, twenty-four hours a day, indefinitely.
Let’s go back to Vantage in Shackelford County: 2.
3 gigawatts of high-speed engine systems.
If those are Jenbacher-class units in the four-to-five-megawatt range, you’re looking at something on the order of five hundred engines behind one fence.
Now apply a routine minor service interval of every two thousand operating hours.
Five hundred engines running continuously works out to more than two thousand service events a year. Roughly forty a week. Every week.
That’s not a maintenance contract, that is a standing industrial workforce.
And where is it standing?
In Shackelford County, Texas. Abilene. Rural Ohio.
These sites were chosen precisely because they’re empty.
Cheap land, fast permits, few neighbours to object.
The same emptiness means there is no local pool of gas engine mechanics or high-voltage electricians to hire from.
You are importing an entire craft workforce into a county that never had one.
And there’s a second problem on top.
That workforce doesn’t exist in surplus anywhere.
The entire gas turbine industry spent 2017 through 2022 at production lows of under ten gigawatts a year, down from 2001 where GE alone shipped more than sixty gigawatts.
An entire cohort of turbine technicians retired or left the trade during that bust.
You can rebuild an order book in a quarter, but you cannot rebuild a workforce in a quarter.
And that’s the detail that makes this argument genuinely hard to dismiss: the manufacturers are telling us this themselves, in their own guidance.
GE Vernova has promised to increase production to twenty-four gigawatts a year, which, notably, only returns them to their 2007 to 2016 levels.
Siemens Energy plans to scale from about twenty gigawatts to over thirty by the end of the decade.
And both have said they intend to do it without increasing factory footprint.
They are investing in staffing and shift utilization, not steel and concrete.
When a manufacturer tells you their expansion is about people rather than buildings, they are telling you their binding constraint is labor.
And it gets tighter further upstream.
Turbine blades and vanes are among the hardest things modern industry makes.
Single-crystal nickel alloys with rhenium, cobalt, tantalum, tungsten and yttrium, machined to super tight tolerances that represent a high-water mark of engineering competence.
Western production of those parts can be found with essentially four firms: Precision Castparts, Howmet Aerospace, Consolidated Precision Products and Doncasters. Four companies.
They are a fraction of the size of the customers they serve, they got hit by the turbine bust and the COVID aerospace slump simultaneously, and expanding means hiring specialized staff they’d have to train from scratch, for demand they quietly suspect could be a bubble they’d be left holding.
And that’s why you might be able to circumvent the grid, but not the workers.
Behind-the-meter doesn’t dodge the labor market.
On the contrary, it doubles down.
The same electricians and pipefitters who build the datacenters are the ones who build the power plants.
Same site, same schedule, same trades.
And they’re all simultaneously being bid for by fab construction, by transmission projects, by every other piece of the industrial buildout. So that’s the argument.
The obvious constraint is labor. It’s a good one.
And here’s our answer: First, that labor is a cost constraint rather than a timeline constraint, and behind-the-meter is bought for timeline.
At ten to twelve million dollars of annual revenue per megawatt, a gigawatt AI campus can outbid essentially any other employer in the United States for turbine technicians.
Even a generously staffed hundred-person operation is a rounding error against ten to twelve billion dollars a year of revenue capacity.
Yes, money does not conjure a skilled worker, but it absolutely solves “who will hire the site electrician.
” Second, the industry is already restructuring around exactly this problem.
That is what Energy-as-a-Service is.
Firms like VoltaGrid don’t just sell you engines, they sell you electric energy, power quality, guaranteed uptime, and a date.
They procure, design, build and operate, and they employ the technicians across a whole fleet of sites.
Which is to say: they share labor cost across many customers, the same way a utility does.
The datacenter doesn’t have to become a power company, it rents one.
And lastly, the labor burden is a function of technology choice, not of behind-the-meter itself.
Five hundred small engines are a labor nightmare.
Ten aeroderivatives are not.
That is a real part of why aeros and IGTs look so attractive despite being less efficient.
The headcount you need to run them is limited.
And OEM hot-swap programs convert the hardest skilled work, major turbine overhauls, from an onsite problem into a logistics problem handled back at a depot.
But there’s a part of this argument that still stands, because it’s the part that matters, and it does changes how we state our own thesis.
The labor constraint doesn’t really bother the datacenter.
It bites upstream, at those four casting houses and at heavy-duty turbine assembly, where no amount of AI capital expenditure just creates skilled workers out of thin air.
And that labor shortage is a hard physical limit on how fast the entire buildout can go.
It also means that behind-the-meter power stays structurally more expensive per megawatt-hour than grid power. Not temporarily. Structurally.
Labor per megawatt is simply worse when you run a control room for one gigawatt instead of thirty.
Which means our case for behind-the-meter only holds if we’re precise about what it is.
It is not an economic optimization.
If anyone sells it to you as cheaper power, the labor argument dismantles that claim completely.
It is a timing arbitrage, a very expensive way to buy eighteen months or maybe more.
But in a market where eighteen months is worth billions, that trade makes sense.
In a market where it isn’t, it doesn’t.
This is why Texas, and especially ERCOT, is becoming such an important testing ground.
ERCOT is an energy-only market with a different structure than capacity markets like PJM, and it is now being forced to deal with a wave of huge datacenter load requests.
The market is settling into hybrid structures that blend onsite generation with some continued grid access, and the central concept is very simple: how much can a site withdraw from the grid independent of its own generation?
Imagine a one-gigawatt AI campus.
The local grid might not be able to support the full gigawatt, at least not now or even anytime soon.
But maybe it can support 100 megawatts.
Under a hybrid structure, the site could use that 100-megawatt withdrawal limit from the grid and supply the rest with onsite generation.
As more generation comes online, the campus ramps.
If the site produces more power than it needs, it may be able to export some of that surplus.
If the grid is constrained, the site may have to reduce its withdrawal.
This is not just an engineering problem; it is a market design problem. Who gets to connect?
Who is allowed to withdraw power?
Who has to curtail during emergencies? Who pays for upgrades?
Can existing generation be redirected toward a private datacenter without hurting the rest of the grid?
And how should regulators treat new generation that is built primarily to serve a behind-the-meter AI campus but may still interact with the public grid?
We’ve written about structures like net-metering arrangements, bring-your-own-generation setups, withdrawal-limited private use networks and provisional controllable load resources.
The names are technical, but the underlying logic is straightforward.
The grid cannot always deliver the full requested load, so the datacenter brings its own generation and agrees to rules about how much it can pull from the system.
What makes ERCOT interesting is not just the market design.
It’s the institutional temperament.
In April 2025, ERCOT published a long-term load forecast, projecting up to 77.
9 gigawatts of potential datacenter load by 2030.
The prior year’s outlook had said 29. 6.
That’s more than a doubling, in a single revision.
Taken literally, it implies bolting an entire second ERCOT onto the existing system.
And then ERCOT did something unusual. They didn’t believe it.
In the May 2025 Capacity, Demand and Reserves report, they applied a deliberate haircut to their own numbers: generic requests discounted to 49.
8%, officer-attested requests to 55.
4%, and every in-service date pushed back by 180 days.
Their own analysts effectively said they would not plan for what developers claim until shovels actually move.
And it’s at this point where this stops being an infrastructure story and starts to become a political one.
In June 2025, New Jersey residents saw electricity rates jump roughly twenty percent, effectively overnight.
It became a big issue in that year’s elections.
And a lot of fingers got pointed at datacenters, including at a 300-megawatt Nebius facility being built for Microsoft in the state, which is a slightly awkward target given that more than eighty-five percent of its power is self-generated.
So: are AI datacenters making American households pay more for electricity?
The honest answer is that it depends almost entirely on which market you live in.
The 67 million residents of the PJM territory are set to see bills rise by an average of about fifteen percent in 2026 relative to the pre-AI-datacenter era.
Texas, which is absorbing an equivalent AI buildout, has seen prices roughly stable since 2023. Same technology. Same demand growth.
Completely different outcome.
Which tells you the variable isn’t AI.
To see why, we need to understand one obscure mechanism: the capacity market.
Your electric bill is roughly four different things.
Energy, the actual electrons, priced by real-time supply and demand.
Transmission and distribution, the poles and wires.
Assorted taxes and adders.
And in some markets, capacity.
Capacity is the strangest of the four.
It is money you pay power plants to sit idle, ready, for a peak event that might last a few hours a year.
It exists for good reason: New York City’s load can swing by two gigawatts in a single day against a peak of six to eight, and during a heat wave it can pull ten.
You need generation standing by for that, and standing by isn’t free.
ERCOT doesn’t have a capacity market at all. It’s energy-only.
When reserves get tight, real-time prices spike from a normal ten to fifty dollars per megawatt-hour toward a cap of five thousand.
That scarcity pricing is what pays for peakers and batteries: a handful of run-hours a year can be worth millions to a fifty-megawatt plant.
ERCOT doesn’t pay you for sitting idle, but it pays you a lot for delivering when demand it too high. PJM does it differently.
PJM runs an annual forward auction, the Base Residual Auction, held two years ahead of delivery, that sets a single price for capacity across the whole region.
And that price went from twenty-nine dollars per megawatt-day to two hundred seventy. A 9. 3x increase in one year.
And in some locations closer to four hundred fifty.
The subsequent auctions cleared at record prices too, and would have gone higher except that federal regulators imposed a price cap of three hundred twenty-nine dollars, a cap that has now been hit two years running.
But what does that mean for a household bill.
At $329 per megawatt-day, divided by hours in a day and applying a typical 40% load factor, gives us about $34 per megawatt-hour, or 3.
4 cents per kilowatt-hour.
Multiply by average PJM household consumption of 880 kilowatt-hours a month, and you get about thirty dollars.
And because those auctions have already cleared, this isn’t a forecast.
Households in PJM will pay twenty-five to thirty dollars more per month than they did in 2024.
Total transfer: roughly sixteen billion dollars per year for 2025, 2026, 2027, and 2028; versus two billion in 2024.
So, was it the datacenters?
PJM’s own independent market monitor ran the counterfactual, and on the surface the answer looks damning.
Strip all datacenters out of the forecast and peak load drops by 7,927 megawatts, cutting total capacity payments by $9. 33 billion.
A sixty-four percent reduction.
Count only datacenters already energized, and you still cut $7. 74 billion.
No other single factor came close.
But look carefully at how that sentence is constructed.
Remove datacenters from the forecast.
Because here is the thing about PJM’s capacity price: it is not set by a market, it is set by a simulation.
The clearing price is determined against something called the Variable Resource Requirement curve; an artificial supply-demand curve built from PJM’s own internal forecast model.
Not from what bidders think will happen, from what the central planner projected.
And that curve is extraordinarily sensitive: being wrong about datacenter load by a few gigawatts changes the curve’s shape near the clearing point and swings prices by billions.
So, the question isn’t “how much load will datacenters add,” but “how good is PJM at forecasting?
” Not very, it turns out.
Their own published data shows they can’t reliably forecast one year ahead.
In 2024, they cut their datacenter load forecast by 800 megawatts versus the prior year.
In 2025, they did it again, cutting 1.
1 gigawatts versus the forecast they had made just twelve months earlier.
Our Datacenter Industry Model tracks construction timelines for over five thousand individual facilities, quarter by quarter.
On that basis, we think PJM’s forecast is still too high.
Not because AI demand isn’t real, it obviously is, but because datacenters are chronically late. Construction slips. GPU deliveries slip.
New hardware platforms are buggy and take longer than expected to reach full utilization. The demand is coming.
It is just not coming when PJM says it is.
You can sanity-check this against a genuine market.
PJM Western Hub forward energy prices, where traders bet real money on real risk, are up twelve to twenty percent in the 2028 and 2030 windows.
Not nothing, but nowhere near a 9. 3x explosion.
And ERCOT forwards are up eleven to seventeen percent over the same period, which is to say: roughly the same as PJM Western Hub.
The energy markets in Texas and Pennsylvania basically agree with each other.
It is the capacity simulation that is the outlier.
And there’s more, on the supply side.
PJM’s offered capacity has fallen by about thirty-five gigawatts in four years.
Coal retirements were the largest driver, but close to twenty gigawatts of that disappearance came from PJM’s own methodology changes.
A single change in how they account for natural gas plants made fourteen gigawatts vanish overnight. Not one plant closed. The accounting changed.
Tighten supply on paper, inflate demand on paper, and run both through a curve that amplifies small errors into billion-dollar swings.
That is how you get to sixteen billion dollars.
And then, in January 2026, the whole thing got stress-tested.
Winter Storm Fern hit at January 23rd with the deepest cold from 26th to the 27th, and ended on February 2nd.
In the ERCOT region, with nothing priced in and a reduced demand forecast, the result was surprising.
Demand still ran below forecast.
No emergency procedures triggered.
Real-time prices peaked around three hundred dollars per megawatt-hour.
And PJM, the market that had just spent nine times more on capacity, explicitly to buy reliability, lost approximately twenty-one gigawatts of generation.
Fifteen percent of the entire fleet that had cleared the auction, knocked out by frozen equipment and fuel delivery failures.
Prices averaged seven hundred dollars per megawatt-hour system-wide, and Virginia’s datacenter-heavy Dominion zone hit eighteen hundred.
The Department of Energy had to issue emergency orders under Section 202(c) of the Federal Power Act, authorizing operators to bypass environmental limits and tap roughly thirty-five gigawatts of backup generation sitting at datacenters and industrial sites, capacity that, by the rules, wasn’t even eligible to bid into the auction in the first place. The 9.
3x increase was supposed to buy reliability. It did not.
The structural reason is almost embarrassingly simple.
In PJM, plants get paid whether or not they perform.
In ERCOT, plants earn their money during scarcity, which means only when they actually generate and deliver.
If peak hitsyou’re your plant isn’t supplying, you don’t make a single cent.
Guess which arrangement makes an operator winterize their equipment.
And there is one more detail in that storm worth looking at.
Those thirty-five gigawatts of backup generation at datacenters and industrial sites?
That is, in large part, the industry’s own behind-the-meter fleet.
And in an emergency, it functioned as a grid asset.
Neither market has systematically priced that in, which suggests the eventual relationship between AI datacenters and the grid may end up considerably less adversarial than the current politics imply.
So, to answer the question directly: yes, PJM households are paying more, and yes, datacenter load forecasts are the largest input driving that increase.
But the mechanism converting those forecasts into a nine-fold price spike is a market design, not a physical shortage of electricity.
Texas is absorbing the same buildout without it.
The problem is not that AI uses power.
The problem is a simulation that turns a forecasting error into sixteen billion dollars of household bills, and a capacity construct that pays for reliability it does not actually receive.
And that’s why energy is the part of the “AI story” that is easy to miss if you only look at semiconductors.
The AI supply chain is no longer just Nvidia, AMD, Intel, Qualcomm, Broadcom, TSMC, SK Hynix, Micron, switches and optics.
It’s also gas turbines, fuel cells, reciprocating engines, transformers, switchgear, high-voltage breakers, substations, power electronics, cooling systems, engineering firms, permitting specialists and grid interconnection experts.
AI is pulling in the heavy industrial economy.
That means the winners of the AI boom may not only be the companies with the best chips, but they may also be the companies that can deliver the equipment needed to power those chips.
If gas turbine lead times are three to four years, then turbine availability becomes a strategic asset.
If generator step-up transformers are delayed, transformer supply becomes an AI bottleneck.
If only a few regions can support massive load growth, then land near power, gas pipelines and transmission corridors becomes more valuable than land that merely looks good on a map.
The constraint keeps moving upstream into stranger and stranger places. From GPUs to HBM.
From HBM to advanced packaging.
From packaging to power.
From power to gas turbines.
From gas turbines to turbine blades.
From blade to skilled workers.
And from workers to yttrium, rhenium and single-crystal nickel.
And yttrium, incidentally, sits on China’s rare earth export control list.
The AI buildout now has a dependency on Chinese rare earth policy.
Not for chips, for power.
That is also why we’re seeing genuinely creative supply responses.
ProEnergy’s PE6000 program takes engine cores out of retired Boeing 747s and rebuilds them into working aeroderivative turbines with near-identical specs to a GE LM6000.
Old jumbo jets, converted into AI power plants.
And given that medium-speed engines are largely built by companies who have spent a century building ship engines, in the same factories, the obvious next question is when someone starts pulling engines out of decommissioned vessels.
This also changes how we should think about datacenter announcements.
A company can announce a five-gigawatt campus, but the real question is not whether the press release reads well, or even if they can get the actual racks.
The real question is whether the project has credible power. Is there a grid path?
Is there onsite generation? Is there a gas pipeline?
Are the turbines secured?
Are the transformers secured? Is the air permit filed?
And air permitting for onsite generation can take a year or more.
Even fast-permitting Texas has already delayed at least one gigawatt-scale Stargate facility.
Is the interconnection credible?
Is the local market actually able to absorb the load?
Without that, a giant AI campus is just a lot of useless silicon.
In that sense, power becomes the filter that decides which AI infrastructure projects are real.
Money matters, GPUs matter, land matters, cooling matters, but power is the gatekeeper.
A hyperscaler can have the balance sheet, an AI lab can have the demand, a chip vendor can have the accelerators and a developer can have the land.
But if the site cannot get electricity on the required timeline, the project does not become compute.
Of course, behind-the-meter power also creates a major environmental and political tension.
Many of these solutions involve gas, and that is uncomfortable for companies that have made aggressive climate commitments.
The AI industry wants fast power, reliable power, clean power and cheap power, but in the real world those goals often conflict.
Renewables and batteries will be part of the answer, but for firm 24/7 power at gigawatt scale, especially on a near-term timeline, dispatchable generation is still extremely difficult to avoid.
That does not mean gas is the final answer forever.
It means gas may become the bridge that allows AI infrastructure to grow while the grid, transmission system and cleaner firm resources try to catch up.
Some sites will use onsite gas temporarily.
Some will combine generation with batteries.
Some will use flexible load.
Some will eventually connect more deeply to the grid.
Some may move workloads to regions where power is cleaner, cheaper or easier to secure.
But the near-term pressure is so extreme that many buyers will choose certainty over elegance.
And this is where the AI boom starts to look less like a pure hard- and software story and more like an industrial strategy problem.
The question is not just whether a model can scale, the question is whether the physical world can scale with it.
Can we build enough substations?
Can we source enough transformers?
Can we permit enough generation? Can we move enough gas?
Can we expand enough transmission?
Can we train enough people to run all of it?
And can we do all of that while still protecting grid reliability and household power prices?
The most advanced software in the world now depends on some of the heaviest infrastructure in the world.
Concrete, steel, copper, fiber, water, gas, turbines, transformers and transmission lines.
And the people who know how to keep all of it running.
That is what makes this so fascinating.
The frontier of AI is not only moving forward through algorithms and chips, it is moving through power markets.
A model can only scale if the datacenter can scale, and the datacenter can only scale if the power system can scale.
The bottleneck is moving from silicon to energy, and that shift could define the next phase of the AI race.
If we’re right, behind-the-meter could power well over half of new US datacenters from 2028 onward, and the equipment market for datacenter behind-the-meter solutions could cross 50 gigawatts per year by 2029.
That would be a massive change.
It would mean that the largest AI players are not just cloud customers and chip buyers anymore.
They are becoming energy buyers, infrastructure developers and, in some cases, almost power companies by necessity.
Whether they can actually staff those power companies is still an open question.
The AI race is becoming a power race.
And the companies that understand that first might be the ones that define what comes next.
Let me know in the comments: do you think the labor constraint is the thing that actually caps this buildout, or is that just a problem that enough money solves?
And will it have lasting impact on the US labor marker?
If you enjoyed this breakdown, subscribe for more deep dives into semiconductors, AI infrastructure and the physical systems behind the AI boom.