On March 1, I put up this post that described an article I’d just read in the Wall Street Journal. The gist of the article was that the grid stability problem with data centers that we had been focusing on up to that point – the problem of insufficient supply causing blackouts – was still serious, but there was a much more serious problem. This one is related to too little demand for power, not too little supply.
The post described a phenomenon that had already occurred twice in northern Virginia, which has become the Saudi Arabia of data centers: a routine disturbance on the grid causes a lot of data centers to go to backup power at literally the same time, causing a sudden drop in power demand. This is a problem that has only recently been recognized, since it only happens when there are a number of “large loads” (which are usually data centers, but can include other facilities like electric arc furnaces) on the same section of the grid, all with backup power sources that they can switch to whenever their protective relays decide the grid is too dangerous at the moment.
The reason this sudden drop in load is so serious is that, unlike a sudden loss of generation (which of course happens regularly), it doesn’t seem to naturally self-correct – in fact, it naturally self-reinforces. If too much generation is lost, there will be a blackout, which of course means a lot of load will drop as well. With the help of various protective measures, the grid will quickly move to equilibrium at lower levels of supply and load.
However, if too much load drops at one time, that doesn’t naturally cause generation to drop – which is what would be needed. In fact, it might well cause the opposite to happen: The imbalance between supply and demand will increase as more and more large loads go to backup power. Where does this end? I’m no electrical engineer, but I don’t believe there’s necessarily a natural path by which stability will be restored, as there is in the case of too little supply.
When I wrote the post in March, there had been two such events in northern Virginia in the last couple of years (that area is part of the PJM grid); in both of them, less than 2,000 megawatts of load was suddenly lost. That’s a lot, but it was survivable. However, the article I quoted continued: “It didn’t cause an emergency, but I would say it caused concern,” said Mike Bryson, PJM’s senior vice president of operations. “What we’re worried about is, what if that happens for 3,000 megawatts or 5,000 megawatts?”
This week, we found out the answer to that question. To quote from Energy Central’s daily newsletter:
"...on Wednesday, a transmission line fault in Northern VA’s “Data Center Alley” triggered over 3 GW of load to suddenly disconnect from the grid. This massive drop occurred as data centers automatically shifted to backup power.
It could’ve been worse: It took around 10 minutes for Dominion to stabilize the grid, far longer than the typical millisecond-scale response time. But residents reported only minor problems. This means automatic systems (plus grid operator responses) likely kept things in check.
“The system, as far as we know, and as far as we've seen, responded incredibly well,” Kyle Thomas, VP of engineering and compliance services at Elevate Energy, told Energy Central. “That’s the really good side of this.”
Yes, but: The industry lacks good models and data to grasp when and where these events could occur next, Thomas said. These could inform more tools and practices to mitigate them.
To get ahead of future data center disconnections, NERC is developing a large-loads action plan and new reliability standards. Plus, CAISO and ERCOT are getting proactive with ride-through rules that require data centers to stay online in outages.
The takeaway? New rules aren’t enough—data center operators must collaborate with grid pros to adapt (or upgrade) their equipment to “meet the requirements of the grid,” Thomas told us.
When this problem was discovered to be so serious this spring, NERC (encouraged by FERC) took extraordinary steps – including developing a new fast-track development process for standards needed to address serious new threats. It was probably too early for NERC’s measures to have an effect on this week’s event, since it sounds like Dominion (which operates the Virginia portion of the PJM grid) did all the right things needed to keep the grid from falling off a cliff. On the other hand, the fact that fixing the problem took ten minutes, not the milliseconds it would normally take, shows that a lot more help is needed. Fortunately, help seems to be on the way.
Tom Alrich’s Blog, too is a reader-supported publication. You can view new posts for three months after they come out by becoming a free subscriber. You can also access all of my 1300 existing posts dating back to 2013, as well as support my work, by becoming a paid subscriber for $30 for one year (and if you feel so inclined, you can become a founding subscriber for $100). Whether free or paid, please subscribe.
If you would like to comment on what you have read here, I would love to hear from you. Please comment in my chat or email me at [email protected].