AI model training is creating a new and increasingly difficult problem for data-centre operators and electricity networks: not just how much power AI consumes, but how quickly its electricity demand can rise and fall.

Large AI training clusters can contain thousands of GPUs operating in synchronised cycles. When those GPUs start, stop, communicate, checkpoint or switch between different stages of training, their electricity demand can change within milliseconds. In some cases, the resulting power swings can be far larger than what conventional data-centre infrastructure was designed to handle.

Recent reporting has highlighted cases where AI-related power fluctuations have reached levels around 50% above expected demand, putting pressure on batteries, turbines, cooling equipment and other electrical infrastructure.

The issue is becoming increasingly important as technology companies build ever-larger AI data centres that can consume electricity on the scale of major cities.

AI’s electricity problem is becoming a power-quality problem

The first phase of the AI energy debate focused mainly on total electricity consumption.

Training increasingly sophisticated models requires enormous numbers of GPUs, and those processors consume substantial amounts of electricity. But researchers and infrastructure companies are now highlighting a second problem: power volatility.

Traditional data centres generally maintain relatively stable electricity demand. AI training clusters can behave very differently.

Thousands of GPUs may perform the same computational operation simultaneously, causing the facility’s power demand to rise sharply. When that operation ends, demand can fall just as quickly.

Traditional data centre vs AI training centre

FactorTraditional data centreLarge AI training centre
WorkloadMore predictableHighly dynamic
Power demandRelatively stableRapidly fluctuating
GPU concentrationLowerExtremely high
Load changesGenerally gradualCan occur within milliseconds
Main concernTotal electricity useTotal use + power transients
Infrastructure stressMore predictableGreater volatility
Grid challengeCapacityCapacity + stability

Researchers behind the EasyRider project say large-scale AI training workloads can create rapid power swings during synchronous communication, start-up, shutdown and checkpointing. These changes can create voltage and frequency shifts and reactive-power transients capable of stressing transformers, converters and protection equipment.

What is causing the sudden power spikes?

AI model training is highly synchronised.

A large training cluster may contain thousands of GPUs working together on the same model. Instead of each processor behaving independently, many of them perform similar operations at almost exactly the same time.

That synchronisation creates a distinctive electrical pattern.

How an AI power spike happens

AI training begins
        ↓
Thousands of GPUs activate
        ↓
GPU power consumption rises rapidly
        ↓
Large electrical load appears
        ↓
Training step / communication phase ends
        ↓
GPU demand falls rapidly
        ↓
Power swings through data-centre equipment
        ↓
Stress on UPS, transformers,
converters and generators

The problem becomes more serious as AI clusters grow.

A small power fluctuation from one server is generally manageable.

A simultaneous fluctuation across thousands of GPUs can become a major electrical event.

The 50% power-spike problem

Recent reporting has highlighted AI data centres experiencing unexpected power spikes that can reach roughly 50% above the level infrastructure was designed to handle.

The exact magnitude varies by facility, workload and power architecture, so the 50% figure should not be interpreted as a universal characteristic of every AI data centre.

The broader point is that AI workloads can create much sharper changes in electricity demand than traditional data-centre workloads.

Key numbers

              AI POWER PRESSURE

       ~50%
 Potentially higher-than-expected
 power demand during spikes

        ↓

   MILLISECONDS
 Rapid GPU load changes

        ↓

  THOUSANDS
 GPUs operating together

        ↓

  ~1 GIGAWATT
 Potential scale of a very large
 AI training facility

        ↓

   HARDWARE STRESS
 Transformers • turbines • UPS
 converters • cooling systems

The volatility matters because electrical infrastructure is designed not only around average demand but also around how quickly demand changes.

Why sudden changes are more dangerous than high average demand

A data centre that continuously consumes 500 megawatts is an enormous electrical load.

But from a grid-engineering perspective, a stable 500-MW load can be easier to manage than a load that repeatedly jumps between significantly different levels.

Electricity networks need to maintain stable voltage and frequency.

Large and rapid changes can make that more difficult.

Average demand vs power volatility

SituationAverage demandVolatilityInfrastructure challenge
Traditional data centreHighLowManageable
AI inference workloadHighModerateIncreasing
AI training workloadVery highHighSignificant
Large synchronised AI clusterExtremely highVery highPotentially severe

This is why AI’s electricity footprint is becoming an engineering challenge rather than simply an energy-generation challenge.

GPUs are at the centre of the problem

Graphics processing units are the primary engines behind modern AI training.

A large AI cluster can contain tens of thousands or even hundreds of thousands of GPUs.

Each GPU may consume hundreds of watts under heavy workloads. Multiplied across thousands of processors, the electricity requirement becomes enormous.

More importantly, those processors can change their power consumption rapidly.

During some phases of training, GPU utilisation can rise sharply. During communication or synchronisation phases, the load can change again.

The resulting pattern is very different from many conventional enterprise workloads.

AI training can stress the hardware itself

Power fluctuations do not only affect the electricity grid.

They can also affect equipment inside the data centre.

According to the Uptime Institute, AI training workloads can produce short-duration GPU power spikes that place additional thermal and electrical stress on servers and their power-delivery components.

Repeated stress can potentially reduce the useful life of:

  • GPUs
  • Voltage regulators
  • Capacitors
  • UPS systems
  • Power converters
  • Transformers
  • Cooling equipment
  • Backup generators

The financial impact can therefore extend beyond electricity bills.

Turbines are facing an unexpected problem

One of the more unusual consequences is the effect on gas turbines and other on-site generation systems.

Large AI data centres are increasingly looking for dedicated power sources because grid connections can take years to secure.

Some facilities therefore rely on gas turbines or other generators to supply electricity directly.

The problem is that large thermal turbines are generally designed to operate relatively steadily.

AI workloads can introduce rapid load changes that force generators to repeatedly adjust their output.

Why turbine cycling matters

Stable industrial load
        ↓
Turbine operates steadily
        ↓
Lower mechanical stress


AI training load
        ↓
Rapid power increase
        ↓
Rapid power decrease
        ↓
Repeated cycling
        ↓
Mechanical oscillation
        ↓
Fatigue and component stress
        ↓
Potential equipment failure

Researchers have warned that rapid AI-related load fluctuations can create torsional interactions inside turbine-generator systems.

In extreme cases, these mechanical oscillations can damage turbine shafts.

That creates a major concern because large generators are expensive, difficult to replace and already facing long manufacturing backlogs.

The equipment replacement problem

The AI boom has created extraordinary demand for electricity-generation equipment.

Gas turbines from major suppliers are already heavily booked, with some delivery schedules extending toward the end of the decade.

That means an AI data-centre operator cannot necessarily replace damaged generation equipment quickly.

A turbine failure could therefore cause a much longer disruption than a conventional server failure.

Equipment at risk

EquipmentPotential impact from volatile AI loads
GPUsThermal and electrical stress
UPS batteriesFrequent charge/discharge cycles
Power convertersVoltage and current transients
TransformersThermal and electrical stress
TurbinesMechanical fatigue and oscillation
Cooling systemsRapid changes in heat load
Protection equipmentIncreased transient events
Grid infrastructureVoltage/frequency instability

Why UPS batteries are also under pressure

Uninterruptible power supply systems are designed to protect data centres from interruptions.

But rapid AI power fluctuations can make them work harder even when the grid itself has not failed.

If the power demand of a GPU cluster rises and falls quickly, energy-storage systems may repeatedly compensate for those changes.

Frequent cycling can reduce battery life.

This creates another hidden cost for AI infrastructure.

The operator may have sufficient electricity-generation capacity on paper but still face higher maintenance costs because batteries, converters and other components experience more aggressive operating conditions.

AI data centres are becoming as large as cities

The scale of AI infrastructure is another reason the problem matters.

Some AI facilities can require hundreds of megawatts of electricity, while the largest planned campuses are targeting power requirements approaching or exceeding 1 gigawatt.

For comparison, a 1-GW facility represents an enormous continuous electrical load.

At that scale, a sudden change in demand is no longer simply an internal data-centre problem.

It can become relevant to the wider electricity network.

AI data-centre power scale

Small enterprise data centre
        │
        ▼
     ~1-10 MW
        │
        ▼
Large cloud facility
        │
        ▼
    50+ MW
        │
        ▼
Large AI campus
        │
        ▼
  Hundreds of MW
        │
        ▼
Next-generation AI campus
        │
        ▼
    ~1 GW or more

Schneider Electric has highlighted the dramatic increase in AI rack densities, with conventional data-centre racks around the 10-kW range in some environments while high-density AI pods can exceed 140 kW per rack.

This means the electrical architecture of AI facilities is increasingly different from that of traditional data centres.

The grid was not designed for this type of demand

Electricity grids are built around a combination of predictable demand patterns and manageable fluctuations.

Residential electricity demand changes throughout the day.

Factories generally operate according to relatively predictable schedules.

Commercial buildings follow business hours.

AI training can introduce a different pattern.

A large cluster may suddenly increase its electricity consumption because thousands of processors begin a new computational stage.

That can create a load ramp much faster than conventional power systems were designed to accommodate.

The grid challenge

Traditional demand

Morning → gradual increase
Afternoon → relatively stable
Evening → gradual decline


AI training demand

       /\       /\        /\
      /  \     /  \      /  \
_____/    \___/    \____/    \____
      Rapid changes in load

The more AI campuses connect to the grid, the more important this issue becomes.

Blackout risk could increase

Power volatility can create challenges for grid stability.

If a large AI facility suddenly increases demand, the grid needs to supply the additional electricity while maintaining voltage and frequency.

If multiple large AI facilities experience similar demand changes at the same time, the effect could become more pronounced.

That does not mean every AI power spike will cause a blackout.

Modern power grids have multiple layers of protection.

But the rapid growth of AI loads is creating a new category of demand that utilities must increasingly account for in planning.

AI is forcing utilities to rethink power planning

Historically, utilities could forecast electricity demand using relatively predictable patterns.

AI makes the forecasting problem more complicated.

Utilities now need to know not only:

How much electricity will a data centre consume?

but also:

How quickly can its electricity consumption change?

That second question could become just as important as the first.

New questions for utilities

Old planning questionNew AI-era question
How much power does the facility need?How rapidly can demand change?
What is peak demand?How quickly can peak demand arrive?
How much generation is required?How much flexible generation is required?
Is grid capacity sufficient?Can the grid absorb rapid load swings?
Is the transformer large enough?Can the transformer handle transients?

This could influence where future AI data centres are built.

A location with abundant electricity may not necessarily be suitable if the local grid cannot handle the volatility.

Batteries could become part of the solution

Energy storage is emerging as one of the most promising ways to reduce the impact of AI power fluctuations.

Instead of allowing every GPU power swing to reach the grid, batteries can absorb some of the rapid changes.

How battery buffering works

AI GPUs
   │
   ▼
Rapid power fluctuations
   │
   ▼
Battery energy storage
   │
   ├── Absorbs sudden increases
   │
   └── Supplies power during sudden drops
   │
   ▼
Smoother electricity demand
   │
   ▼
Grid

The concept is straightforward.

The battery acts as a buffer between the highly dynamic AI workload and the relatively slower electricity grid.

This can reduce the rate at which power demand changes.

Researchers are already developing solutions

The EasyRider research project proposes a rack-level architecture designed specifically to mitigate power transients caused by AI workloads.

The system combines passive electrical components with actively controlled auxiliary energy storage.

Researchers tested the concept using a 400-volt DC prototype and found that it could filter rack-level power variations while avoiding changes to the AI training software.

The significance is that AI infrastructure may be able to solve part of the problem at the rack level rather than forcing utilities to redesign the entire grid.

Software could also help control electricity demand

Hardware is not the only solution.

AI training systems could potentially be designed to avoid synchronised power spikes.

For example, training workloads could be scheduled or staggered so that all GPUs do not change their power consumption at exactly the same moment.

GPU power caps can also reduce peak demand.

Research highlighted by MIT’s Lincoln Laboratory found that limiting GPU power could reduce energy consumption by around 12% to 15%, while increasing task completion time by roughly 3% in tested workloads.

That suggests operators may be able to trade a small amount of training speed for lower power consumption and potentially reduced hardware stress.

Potential mitigation strategies

SolutionMain benefitTrade-off
Battery storageSmooths power spikesAdds capital cost
GPU power capsReduces peak consumptionCan slow workloads
Workload schedulingReduces synchronisationMay reduce utilisation
Rack-level power managementControls local transientsRequires new infrastructure
Oversized electrical systemsAdds operating headroomHigher upfront cost
Better coolingHandles thermal spikesHigher infrastructure cost
Flexible generationResponds to demand changesFuel and maintenance costs

AI infrastructure costs could rise

Power volatility creates another cost layer for AI companies.

The economics of AI infrastructure are already dominated by expensive GPUs, networking equipment, cooling systems and electricity.

If rapid power fluctuations shorten equipment lifespans, operators may need to replace components more frequently.

That increases the total cost of ownership.

Hidden AI infrastructure costs

GPU purchase
     +
Electricity
     +
Cooling
     +
Networking
     +
Buildings
     +
Grid connection
     +
Backup generation
     +
Power-quality management
     +
Equipment degradation
     ↓
TOTAL AI INFRASTRUCTURE COST

This means the cheapest location for electricity may not necessarily be the cheapest location for AI computing.

Infrastructure resilience could become equally important.

The problem becomes bigger as AI models become larger

AI companies continue to scale model training.

Larger models require more GPUs.

More GPUs require more electricity.

More GPUs working synchronously can also create larger power fluctuations.

This produces a potentially difficult feedback loop.

The AI power-growth cycle

Larger AI models
       ↓
More GPUs
       ↓
Higher electricity demand
       ↓
Larger AI data centres
       ↓
Bigger synchronised workloads
       ↓
Larger power fluctuations
       ↓
Greater infrastructure stress
       ↓
Higher cost of AI computing

EPRI and Epoch AI previously estimated that training a leading AI model could require more than 4 GW of power by 2030, illustrating how rapidly the electricity requirements of frontier AI could grow.

This could affect where AI data centres are built

The traditional logic for data-centre location focused on factors such as:

  • Electricity availability
  • Fibre connectivity
  • Land
  • Cooling
  • Tax incentives
  • Proximity to users

AI adds another requirement:

Grid stability.

A region may have enough generation capacity but still lack the infrastructure needed to accommodate rapid AI load changes.

This could encourage developers to build AI campuses near:

  • Large power-generation facilities
  • Strong transmission networks
  • Battery-storage projects
  • Flexible generation
  • Renewable-energy hubs
  • Large industrial power systems

Renewable energy creates another challenge

Renewable energy can provide enormous amounts of electricity to AI data centres.

But solar and wind power are themselves variable.

That means AI power fluctuations and renewable-generation fluctuations could occur simultaneously.

For example, solar output falls in the evening while electricity demand may remain high.

An AI campus may therefore require storage or other flexible generation to balance both the data centre’s workload and renewable-energy variability.

The future AI power system

Solar + Wind
     │
     ▼
Renewable generation
     │
     ▼
Battery storage
     │
     ├──────────────┐
     ▼              ▼
AI Data Centre   Electricity Grid
     │
     ▼
Thousands of GPUs
     │
     ▼
Highly variable load
     │
     ▼
Power-management system

This makes batteries and flexible power systems increasingly important to the AI infrastructure industry.

AI power demand could become a new reliability issue

The electricity industry is now beginning to treat AI data centres differently from conventional large industrial customers.

The reason is scale and speed.

A factory may consume a lot of electricity, but its demand generally changes according to production schedules.

An AI training cluster can change its power consumption much faster.

That means utilities may need new technical requirements for large AI loads.

These could eventually include:

  • Mandatory power-quality standards
  • Maximum load-ramp limits
  • On-site energy storage
  • Flexible-load requirements
  • Dedicated transmission infrastructure
  • Power-factor requirements
  • Backup-generation standards

The industry is moving toward “power-aware AI”

The AI industry has traditionally optimised for:

Maximum performance per GPU.

It may increasingly need to optimise for:

Maximum performance per watt and per power transient.

That could change how AI systems are designed.

Companies may begin considering electricity behaviour as part of model-training architecture rather than treating power as an infrastructure problem.

The next AI optimisation metric

TODAY

Model performance
       +
Training speed
       +
Cost per token


NEXT PHASE

Model performance
       +
Training speed
       +
Cost per token
       +
Energy efficiency
       +
Peak power
       +
Power-ramp behaviour

This could become an important area of competition between AI infrastructure companies.

What this means for Nvidia and AI chipmakers

The issue could also influence future GPU architecture.

AI chipmakers have traditionally focused heavily on increasing compute performance while improving energy efficiency.

But future processors may also need more sophisticated power-management capabilities.

Possible developments include:

  • Dynamic power control
  • More granular power limits
  • Better workload scheduling
  • Improved voltage management
  • Lower transient loads
  • Integrated energy monitoring

The objective would be to prevent large numbers of processors from creating dangerous simultaneous power swings.

What happens next?

The problem is unlikely to disappear as AI development slows down because the opposite is happening.

AI companies are building increasingly large clusters, and data-centre operators are planning facilities with power requirements measured in hundreds of megawatts and, eventually, gigawatts.

That means power infrastructure will become an increasingly important constraint on AI expansion.

The immediate solutions are likely to involve a combination of better power electronics, batteries, smarter GPU controls, workload scheduling and stronger grid infrastructure.

AI power problem: key data

IndicatorFigure / finding
Potential reported AI power spikesUp to ~50% above expected levels
GPU load changesCan occur within milliseconds
Large AI facilitiesHundreds of MW
Next-generation AI campuses~1 GW or more
AI training power requirement projected for a leading model by 2030>4 GW
High-density AI rack power>140 kW in some deployments
Conventional rack example~10 kW
GPU power-capping energy reduction in MIT tests~12-15%
Training-time increase from power capping in tests~3%
EasyRider prototype400V DC

Conclusion

The AI industry’s electricity challenge is entering a new phase.

The problem is no longer simply that artificial intelligence consumes enormous amounts of power. The increasingly important issue is that AI training can consume electricity in rapid, highly synchronised bursts that traditional power infrastructure was not designed to handle.

Those fluctuations can put stress on GPUs, UPS systems, batteries, transformers, converters, cooling equipment and generators. At larger scales, they can also create challenges for electricity grids and increase the risk of voltage and frequency instability.

The problem is particularly significant because AI data centres are becoming enormous. Facilities consuming hundreds of megawatts are already being developed, while future AI campuses could approach or exceed 1 GW.

The industry is therefore likely to invest heavily in battery storage, power electronics, smarter workload scheduling and flexible generation.

For AI companies, the lesson is increasingly clear: building more powerful models is not only a computing challenge. It is also an electricity and infrastructure challenge.

The next generation of AI data centres will need to be designed not just around how many GPUs they can install, but around how intelligently those GPUs can use power without destabilising the equipment and grid supporting them.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.