Hindcast evaluation of the county outage-risk model
Ten United States landfalling tropical cyclones between 2016 and 2023 were replayed through three candidate impact models, each storm withheld from the training pool before its own outages were predicted. Every model beats climatology by roughly a quarter of a CRPS, none of the three is distinguishable from the others on ten storms, and all three under-predict how many customers actually lose power. The model whose hindcasts are published to the dashboard, baseline B+, produces the sharpest point predictions in the set and the worst-calibrated intervals.
1. What this measures — and what it deliberately does not
The scores below were computed on the analysed hazard: each county's wind field comes from the reanalysed record of what the storm actually did, not from a forecast of it. The evaluation therefore isolates the impact model — the step that turns gust and duration into a distribution over the fraction of customers out — and excludes every source of forecast error ahead of it. A live Weather2Grid forecast inherits both, so the skill reported here is an upper bound on end-to-end forecast skill, never a substitute for it.
Each storm was scored under leave-one-event-out replay: the storm being predicted was removed from the fitting pool, so a model never sees the event it is asked to forecast. That is why every scorecard row shows a training pool of about 1,542 events out of 1,543 rather than the whole set.
These hindcasts run in verification mode with the release gate not passed and the degraded-mode flag set, the same status the live dashboard reports. The evaluation is research output about a research product; nothing here makes it operational guidance.
2. Data and event identity
The fitting set holds 45,743 county–event rows over 1,547 outage events; 45,740 rows and 1,543 events survive the evaluation filter. Scoring covers 1,838 county–events across the ten storms, each with an observed outage fraction to score against.
Storm identity is resolved before any scoring: an outage event carries its own opaque identifier, so it has to be matched to a named best track or the results cannot be attributed to a storm at all. Every requested storm matched a HURDAT2 track, four of them across two outage events that belong to the same system. The closest approach between event and track ranged from 2.5 km to 31.7 km, comfortably inside each event's match radius.
| Storm | HURDAT2 | Status | Outage events | Track distance | Match radius |
|---|---|---|---|---|---|
| Matthew 2016 | AL142016 | matched | 2 | 31.7 km | 539 km |
| Harvey 2017 | AL092017 | matched | 1 | 13.7 km | 428 km |
| Irma 2017 | AL112017 | matched | 1 | 14.1 km | 817 km |
| Florence 2018 | AL062018 | matched | 1 | 6.4 km | 465 km |
| Michael 2018 | AL142018 | matched | 1 | 2.7 km | 891 km |
| Isaias 2020 | AL092020 | matched | 2 | 8.6 km | 428 km |
| Laura 2020 | AL132020 | matched | 2 | 9.3 km | 483 km |
| Ida 2021 | AL092021 | matched | 2 | 2.5 km | 613 km |
| Ian 2022 | AL092022 | matched | 1 | 10.9 km | 928 km |
| Idalia 2023 | AL102023 | matched | 1 | 7.5 km | 539 km |
3. Scoring protocol
The primary metric is the continuous ranked probability score of the predicted outage fraction, averaged over the counties of a storm and then over storms with equal weight. Equal storm weight is the point: pooling counties instead would let Michael's 318 counties outvote Harvey's 66 and turn the headline number into a statement about geography rather than about the model. Pooled county diagnostics are reported alongside, and treated as diagnostics.
- Reference. CRPSS compares each model against a climatological distribution built from the same event pool. Zero means no better than climatology; negative means worse.
- Uncertainty. A 2,000-draw bootstrap resamples storms, not counties. Counties within a storm share a weather field and are nowhere near independent, so a county bootstrap would report an interval several times too narrow.
- Calibration. The PIT–KS statistic measures how far the probability integral transform of the observations departs from uniform; 90% interval coverage is the share of counties whose observed outage fell at or below the 90th-percentile forecast. A well-calibrated model returns about 0.90.
4. Skill: all three models beat climatology, none beats the others
| Model | CRPS | 95% interval | CRPSS | PIT–KS | 90% coverage | Bias | MAE | RMSE |
|---|---|---|---|---|---|---|---|---|
| baseline A | 0.163 | 0.129–0.206 | 0.251 | 0.234 | 67.6% | -0.138 | 0.196 | 0.309 |
| baseline B | 0.165 | 0.131–0.205 | 0.243 | 0.236 | 67.4% | -0.141 | 0.198 | 0.311 |
| baseline B+ (published) | 0.165 | 0.130–0.209 | 0.244 | 0.517 | 46.2% | -0.156 | 0.188 | 0.295 |
Skill against climatology is consistent across the set: baseline A scores 0.251 and baseline B+ scores 0.244, both meaningful improvements over a climatological forecast, neither distinguishable from the other. Per storm the ordering churns — baseline B+ takes the lowest CRPS in 5 of the ten storms and the highest in the other 5, which is what a set of models with no real separation looks like.
| Storm | Counties | Climatology | baseline A | baseline B | baseline B+ | Lowest CRPS |
|---|---|---|---|---|---|---|
| Matthew 2016 | 303 | 0.197 | 0.138 | 0.135 | 0.126 | baseline B+ |
| Harvey 2017 | 66 | 0.077 | 0.084 | 0.085 | 0.086 | baseline A |
| Irma 2017 | 235 | 0.353 | 0.307 | 0.298 | 0.317 | baseline B |
| Florence 2018 | 89 | 0.327 | 0.205 | 0.215 | 0.189 | baseline B+ |
| Michael 2018 | 318 | 0.229 | 0.178 | 0.182 | 0.168 | baseline B+ |
| Isaias 2020 | 182 | 0.213 | 0.182 | 0.188 | 0.199 | baseline A |
| Laura 2020 | 139 | 0.232 | 0.161 | 0.168 | 0.197 | baseline A |
| Ida 2021 | 156 | 0.165 | 0.093 | 0.094 | 0.122 | baseline A |
| Ian 2022 | 197 | 0.190 | 0.128 | 0.130 | 0.114 | baseline B+ |
| Idalia 2023 | 153 | 0.191 | 0.152 | 0.151 | 0.129 | baseline B+ |
5. Calibration: sharp, and overconfident
Skill and calibration part company here. baseline B+ carries the narrowest predictive distributions in the set — the lowest mean absolute error (0.188) and the lowest RMSE (0.295) of the three — but its intervals do not cover the outcome. Only 46% of counties fall at or below their 90th-percentile forecast, against a nominal 90%, and its PIT–KS statistic (0.517) is more than double baseline A's (0.234). The intervals it publishes are, in plain terms, too narrow to be believed.
6. Bias: the models under-predict, consistently
Averaged over storms, every model predicts a smaller outage fraction than was observed: -0.156 for baseline B+, -0.138 for baseline A. Pooled over counties the gap is starker still — a mean observed outage fraction of 0.280 against a mean prediction of 0.117. The single worst case is Irma 2017, where the published model predicted a mean outage fraction of 0.110 against 0.428 observed.
Under-prediction and narrow intervals compound rather than cancel: a distribution centred too low and too tight is wrong in the direction that matters most for staging crews. That is the first thing to fix, and it is measurable — both diagnostics on this page are the test that a fix has to move.
| Storm | CRPSS | PIT–KS | 90% coverage | Observed | Predicted | Bias | MAE | RMSE |
|---|---|---|---|---|---|---|---|---|
| Matthew 2016 | 0.362 | 0.452 | 54.5% | 0.244 | 0.129 | -0.114 | 0.150 | 0.253 |
| Harvey 2017 | -0.118 | 0.419 | 65.2% | 0.097 | 0.057 | -0.040 | 0.103 | 0.213 |
| Irma 2017 | 0.101 | 0.602 | 32.8% | 0.428 | 0.110 | -0.318 | 0.346 | 0.469 |
| Florence 2018 | 0.424 | 0.569 | 51.7% | 0.393 | 0.213 | -0.179 | 0.221 | 0.345 |
| Michael 2018 | 0.266 | 0.506 | 49.1% | 0.288 | 0.135 | -0.152 | 0.192 | 0.306 |
| Isaias 2020 | 0.067 | 0.647 | 29.1% | 0.273 | 0.078 | -0.195 | 0.218 | 0.304 |
| Laura 2020 | 0.154 | 0.592 | 36.0% | 0.287 | 0.077 | -0.210 | 0.219 | 0.339 |
| Ida 2021 | 0.261 | 0.641 | 36.5% | 0.208 | 0.079 | -0.128 | 0.140 | 0.244 |
| Ian 2022 | 0.403 | 0.490 | 56.9% | 0.242 | 0.136 | -0.106 | 0.134 | 0.222 |
| Idalia 2023 | 0.328 | 0.485 | 59.5% | 0.242 | 0.128 | -0.114 | 0.151 | 0.249 |
7. What this evaluation does not establish
- It is not forecast skill. Scoring against the analysed hazard removes weather forecast error entirely. The corresponding forecast-basis evaluation has not been run.
- Ten storms is a small sample. The bootstrap interval on the headline score spans roughly 0.08 CRPS. Differences between these models cannot be resolved at this sample size, and neither can a modest real improvement.
- All ten are tropical cyclones. Nothing here speaks to winter storms, derechos or ice, which the live product also runs on.
- Coverage is county-level, not customer-level. A county is one unit whatever its customer base, which is the right call for scoring the model and the wrong one for reading off expected system-wide impact.
- The observed outage record is a data product too, with its own reporting gaps. Counties absent from the record are absent from the score.
8. The published hindcasts
All ten baseline B+ hindcasts are published to the Weather2Grid archive dashboard. Open the run picker and choose a storm from the Hindcast verification group: each one carries the full predictive distribution per county alongside the observed outcome, so the observed and error map layers — which a live forecast cannot have, because its outcome has not happened yet — are available for every county in the storm.
| Storm | Cycle id | Counties | Window start | Window end | CRPSS |
|---|---|---|---|---|---|
| Matthew 2016 | 20160929T0200Z_hindcast-baseline-bplus-MATTHEW_2016 | 303 | 2016-09-29 | 2016-11-03 | 0.362 |
| Harvey 2017 | 20170823T2100Z_hindcast-baseline-bplus-HARVEY_2017 | 66 | 2017-08-23 | 2017-09-02 | -0.118 |
| Irma 2017 | 20170909T1100Z_hindcast-baseline-bplus-IRMA_2017 | 235 | 2017-09-09 | 2017-09-13 | 0.101 |
| Florence 2018 | 20180913T0800Z_hindcast-baseline-bplus-FLORENCE_2018 | 89 | 2018-09-13 | 2018-09-17 | 0.424 |
| Michael 2018 | 20181009T0300Z_hindcast-baseline-bplus-MICHAEL_2018 | 318 | 2018-10-09 | 2018-10-13 | 0.266 |
| Isaias 2020 | 20200802T1100Z_hindcast-baseline-bplus-ISAIAS_2020 | 182 | 2020-08-02 | 2020-08-05 | 0.067 |
| Laura 2020 | 20200826T1100Z_hindcast-baseline-bplus-LAURA_2020 | 139 | 2020-08-26 | 2020-08-31 | 0.154 |
| Ida 2021 | 20210829T0200Z_hindcast-baseline-bplus-IDA_2021 | 156 | 2021-08-29 | 2021-09-02 | 0.261 |
| Ian 2022 | 20220927T1800Z_hindcast-baseline-bplus-IAN_2022 | 197 | 2022-09-27 | 2022-10-05 | 0.403 |
| Idalia 2023 | 20230829T2100Z_hindcast-baseline-bplus-IDALIA_2023 | 153 | 2023-08-29 | 2023-09-01 | 0.328 |
9. Reproducibility
Every number on this page is generated from the evaluation bundle named below by
scripts/build_evaluation_page.py; the hindcast payloads on the archive dashboard
are written from the same bundle by scripts/publish_hindcast_evaluation.py.
Neither page is edited by hand. The bundle itself is published alongside the hindcasts, so
every table here can be recomputed from the artifacts it links.
- Evaluation run
- prism-37922961
- Generated (UTC)
- 2026-09-04T21:03:47.490497+00:00
- Code version
- 0.0.1
- Fitset
- /panfs/ccds02/nobackup/people/afahad/stormgrid/processed/fitset.parquet
- Fitset SHA-256
- 128a32e2ce42e8ac3dee66fe94d007a806b7484c493256797028649e65f1cbeb
- Evaluation content hash
- 1659a16c5812f6ac083f7c726585f751
- Best-track source
- hurdat2-1851-2025-02272026.txt
- Best-track SHA-256
- 1b9b0c7beed5b4505838658b1d30e159fc84330c60891a58cfcf43ae55c37202
- Manifest
- evaluation_manifest.json
- Scorecards
- baseline_a · baseline_b · baseline_bplus
- Score matrix
- evaluation_matrix.csv · model_summary.csv · event_identity_map.csv