Wed, Sep 16

Your Estimated-Read Rate Is a Volume Metric — It Won't Tell You Which Estimates Are Wrong

Every meter-to-cash operation tracks an estimated-read rate. Two percent, four percent, whatever the number is this month — it goes in the report, and if it holds steady, nobody asks a follow-up question.

That number counts estimates. It says nothing about which ones are defensible.

The objection I'd expect from most utilities is: our VEE engine has been running for fifteen years and billing hasn't complained. Fair. But "billing hasn't complained" measures how many estimates got disputed, not how many were wrong. Most bad estimates are never disputed, because the customer has no way to know.

Here's what should concern anyone who owns estimate quality: the errors aren't spread evenly across your estimate population. They concentrate in three buckets — and each of those buckets is operationally identifiable before you fill the gap.

Multi-day gaps. A three-day communications outage on a weather-sensitive residential account in January isn't a bigger version of a 30-minute drop. It's a different problem. Benchmark work shows imputation quality degrading sharply as gaps approach 24 hours, and past that you aren't interpolating — you're guessing in a house style. Without weather or a neighbor profile as input, the information needed to recover that consumption isn't in the account's own history.

Cold-start meters. New installs, migrated accounts, meters that came over in an acquisition. Any method built on a rolling lookback window has nothing to look back at. It will still return a number, and that number will pass every validation rule you have — because plausibility and accuracy aren't the same test.

Storm-adjacent gaps. This is the one I'd put first. The meter stopped reporting because of the event. Fill that gap with a historical baseline and you systematically understate consumption during the exact interval that mattered — for the bill, and for the planners downstream who are about to size the system off that interval data.

The fix isn't a better algorithm. It's a routing layer: classify the gap before you fill it, and send each class to a method suited to it. Short random gaps don't need anything heavy. Medium gaps with good history are a forward/backward interpolation problem. Long gaps and cold-starts need external signal plus an explicit uncertainty flag, not a point estimate dressed up as a read. Storm-adjacent gaps need anomaly scoring to run before method selection, not after.

The objection I expect — and it's a real one — is auditability. Regulated billing needs logic you can defend line by line, and a rules engine gives you that. But routing by gap type is deterministic and fully loggable. You record which tier fired, on what criteria, at what confidence. That's a better audit trail than one estimation rule applied uniformly to cases that were never comparable.

Full breakdown of the three failure modes, the benchmark evidence, and the tiered architecture:
https://jamesaksanders.com/2026/06/21/when-ai-gap-filling-gets-the-easy-cases-right-and-the-hard-cases-wrong/

For those of you on the utility side — does your VEE configuration distinguish gap types today, or is it one estimation path for everything? And if you've tried to change it, what did the audit conversation look like?

1
1 reply