Weather2Grid Hindcast evaluation
Weather2Grid technical note

Hindcast evaluation of the county outage-risk model

Evaluation run prism-37922961 Generated 2026-09-04 21:03:47 UTC Code version 0.0.1 Hazard basis: analysed

Ten United States landfalling tropical cyclones between 2016 and 2023 were replayed through three candidate impact models, each storm withheld from the training pool before its own outages were predicted. Every model beats climatology by roughly a quarter of a CRPS, none of the three is distinguishable from the others on ten storms, and all three under-predict how many customers actually lose power. The model whose hindcasts are published to the dashboard, baseline B+, produces the sharpest point predictions in the set and the worst-calibrated intervals.

10
Storms scored2016–2023, all matched to HURDAT2
1,838
County–eventseach with an observed outage fraction
0.24
CRPSS, published modelskill over climatology, 0 = no better
46%
90% interval coveragenominal is 90% — intervals are too narrow

1. What this measures — and what it deliberately does not

The scores below were computed on the analysed hazard: each county's wind field comes from the reanalysed record of what the storm actually did, not from a forecast of it. The evaluation therefore isolates the impact model — the step that turns gust and duration into a distribution over the fraction of customers out — and excludes every source of forecast error ahead of it. A live Weather2Grid forecast inherits both, so the skill reported here is an upper bound on end-to-end forecast skill, never a substitute for it.

Each storm was scored under leave-one-event-out replay: the storm being predicted was removed from the fitting pool, so a model never sees the event it is asked to forecast. That is why every scorecard row shows a training pool of about 1,542 events out of 1,543 rather than the whole set.

Status of the underlying product

These hindcasts run in verification mode with the release gate not passed and the degraded-mode flag set, the same status the live dashboard reports. The evaluation is research output about a research product; nothing here makes it operational guidance.

2. Data and event identity

The fitting set holds 45,743 county–event rows over 1,547 outage events; 45,740 rows and 1,543 events survive the evaluation filter. Scoring covers 1,838 county–events across the ten storms, each with an observed outage fraction to score against.

Storm identity is resolved before any scoring: an outage event carries its own opaque identifier, so it has to be matched to a named best track or the results cannot be attributed to a storm at all. Every requested storm matched a HURDAT2 track, four of them across two outage events that belong to the same system. The closest approach between event and track ranged from 2.5 km to 31.7 km, comfortably inside each event's match radius.

Table 4. Event identity. Every scored storm was matched to a HURDAT2 best track before scoring, so a storm's outage events, its hazard field and its name all refer to the same system. Distance is the closest approach between the outage event and the best track; the radius is the match tolerance that event was allowed.
StormHURDAT2StatusOutage eventsTrack distanceMatch radius
Matthew 2016AL142016matched231.7 km539 km
Harvey 2017AL092017matched113.7 km428 km
Irma 2017AL112017matched114.1 km817 km
Florence 2018AL062018matched16.4 km465 km
Michael 2018AL142018matched12.7 km891 km
Isaias 2020AL092020matched28.6 km428 km
Laura 2020AL132020matched29.3 km483 km
Ida 2021AL092021matched22.5 km613 km
Ian 2022AL092022matched110.9 km928 km
Idalia 2023AL102023matched17.5 km539 km

3. Scoring protocol

The primary metric is the continuous ranked probability score of the predicted outage fraction, averaged over the counties of a storm and then over storms with equal weight. Equal storm weight is the point: pooling counties instead would let Michael's 318 counties outvote Harvey's 66 and turn the headline number into a statement about geography rather than about the model. Pooled county diagnostics are reported alongside, and treated as diagnostics.

4. Skill: all three models beat climatology, none beats the others

baseline Abaseline Bbaseline B+
Primary CRPS by model with 95% bootstrap interval0.100.120.140.160.180.200.220.24climatology 0.218baseline Abaseline A: CRPS 0.163 (95% CI 0.129–0.206)0.163baseline Bbaseline B: CRPS 0.165 (95% CI 0.131–0.205)0.165baseline B+baseline B+: CRPS 0.165 (95% CI 0.130–0.209)0.165CRPS (outage fraction) — lower is better
Figure 1. Primary CRPS with its 95% bootstrap interval over storms. The three intervals overlap almost completely: with ten storms, a separation of 0.0015 CRPS between the best and the published model is far inside the noise. Choosing between these models on the strength of the headline number is not supported by this evidence.
Table 1. Model comparison. The primary score is the county CRPS averaged within a storm and then across storms, so a 318-county storm and a 66-county storm count the same. The interval is a 2,000-draw bootstrap over storms. PIT–KS and coverage are pooled county diagnostics.
ModelCRPS95% intervalCRPSSPIT–KS90% coverageBiasMAERMSE
baseline A0.1630.129–0.2060.2510.23467.6%-0.1380.1960.309
baseline B0.1650.131–0.2050.2430.23667.4%-0.1410.1980.311
baseline B+ (published)0.1650.130–0.2090.2440.51746.2%-0.1560.1880.295

Skill against climatology is consistent across the set: baseline A scores 0.251 and baseline B+ scores 0.244, both meaningful improvements over a climatological forecast, neither distinguishable from the other. Per storm the ordering churns — baseline B+ takes the lowest CRPS in 5 of the ten storms and the highest in the other 5, which is what a set of models with no real separation looks like.

baseline Abaseline Bbaseline B+
Continuous ranked probability skill score by storm and model-0.2-0.10.00.10.20.30.40.5Matthew 2016 · baseline A: 0.298Matthew 2016 · baseline B: 0.315Matthew 2016 · baseline B+: 0.362Matthew2016Harvey 2017 · baseline A: -0.093Harvey 2017 · baseline B: -0.097Harvey 2017 · baseline B+: -0.118Harvey2017Irma 2017 · baseline A: 0.129Irma 2017 · baseline B: 0.155Irma 2017 · baseline B+: 0.101Irma2017Florence 2018 · baseline A: 0.373Florence 2018 · baseline B: 0.342Florence 2018 · baseline B+: 0.424Florence2018Michael 2018 · baseline A: 0.225Michael 2018 · baseline B: 0.205Michael 2018 · baseline B+: 0.266Michael2018Isaias 2020 · baseline A: 0.144Isaias 2020 · baseline B: 0.118Isaias 2020 · baseline B+: 0.067Isaias2020Laura 2020 · baseline A: 0.306Laura 2020 · baseline B: 0.277Laura 2020 · baseline B+: 0.154Laura2020Ida 2021 · baseline A: 0.435Ida 2021 · baseline B: 0.430Ida 2021 · baseline B+: 0.261Ida2021Ian 2022 · baseline A: 0.326Ian 2022 · baseline B: 0.319Ian 2022 · baseline B+: 0.403Ian2022Idalia 2023 · baseline A: 0.205Idalia 2023 · baseline B: 0.211Idalia 2023 · baseline B+: 0.328Idalia2023CRPSS vs climatology
Figure 2. Skill over climatology by storm. Bars above zero beat the climatological reference for that storm. Harvey 2017 is the exception: every model scores worse than climatology there. Its observed outages were the mildest in the set (mean fraction 0.097), and the climatological distribution was already close to right.
Table 2. County CRPS by storm, one column per model, against the climatological reference built from the same event pool. Lower is better; the climatology column is the number each model has to beat for that storm.
StormCountiesClimatologybaseline Abaseline Bbaseline B+Lowest CRPS
Matthew 20163030.1970.1380.1350.126baseline B+
Harvey 2017660.0770.0840.0850.086baseline A
Irma 20172350.3530.3070.2980.317baseline B
Florence 2018890.3270.2050.2150.189baseline B+
Michael 20183180.2290.1780.1820.168baseline B+
Isaias 20201820.2130.1820.1880.199baseline A
Laura 20201390.2320.1610.1680.197baseline A
Ida 20211560.1650.0930.0940.122baseline A
Ian 20221970.1900.1280.1300.114baseline B+
Idalia 20231530.1910.1520.1510.129baseline B+

5. Calibration: sharp, and overconfident

Skill and calibration part company here. baseline B+ carries the narrowest predictive distributions in the set — the lowest mean absolute error (0.188) and the lowest RMSE (0.295) of the three — but its intervals do not cover the outcome. Only 46% of counties fall at or below their 90th-percentile forecast, against a nominal 90%, and its PIT–KS statistic (0.517) is more than double baseline A's (0.234). The intervals it publishes are, in plain terms, too narrow to be believed.

baseline Abaseline Bbaseline B+
Reliability diagram0.00.00.20.20.40.40.60.60.80.81.01.0perfect reliabilitybaseline A: forecast 18.6%, observed 32.0% (25 counties)baseline A: forecast 25.7%, observed 41.6% (370 counties)baseline A: forecast 34.2%, observed 50.1% (343 counties)baseline A: forecast 44.9%, observed 69.6% (230 counties)baseline A: forecast 54.7%, observed 72.5% (211 counties)baseline A: forecast 64.8%, observed 77.9% (154 counties)baseline A: forecast 75.1%, observed 83.2% (137 counties)baseline A: forecast 85.1%, observed 91.5% (129 counties)baseline A: forecast 97.0%, observed 96.7% (239 counties)baseline B: forecast 19.2%, observed 38.5% (13 counties)baseline B: forecast 25.8%, observed 40.5% (388 counties)baseline B: forecast 34.3%, observed 53.5% (357 counties)baseline B: forecast 44.6%, observed 66.7% (234 counties)baseline B: forecast 54.6%, observed 73.2% (205 counties)baseline B: forecast 64.8%, observed 78.6% (140 counties)baseline B: forecast 74.7%, observed 86.3% (139 counties)baseline B: forecast 84.8%, observed 89.1% (128 counties)baseline B: forecast 97.1%, observed 97.0% (234 counties)baseline B+: forecast 5.3%, observed 34.1% (446 counties)baseline B+: forecast 14.7%, observed 53.8% (353 counties)baseline B+: forecast 25.1%, observed 73.6% (235 counties)baseline B+: forecast 34.7%, observed 82.2% (146 counties)baseline B+: forecast 44.7%, observed 79.0% (100 counties)baseline B+: forecast 54.2%, observed 87.8% (74 counties)baseline B+: forecast 64.7%, observed 85.9% (78 counties)baseline B+: forecast 75.0%, observed 95.0% (80 counties)baseline B+: forecast 84.4%, observed 92.9% (70 counties)baseline B+: forecast 98.0%, observed 94.9% (256 counties)Forecast probabilityObserved frequency
Figure 3. Reliability of the forecast probability that a county loses more than 5% of its customers. A perfectly reliable model sits on the diagonal. All three models sit above it — events happen more often than they are forecast — and baseline B+'s curve departs furthest, especially in the low-probability bins where it assigns near-zero probability to counties that lost power roughly a third of the time.
baseline Abaseline Bbaseline B+
Share of counties whose observed outage fell at or below the 90th-percentile forecast0.000.250.500.751.00nominal 90%Matthew 2016 · baseline A: 73.9%Matthew 2016 · baseline B: 76.9%Matthew 2016 · baseline B+: 54.5%Matthew2016Harvey 2017 · baseline A: 86.4%Harvey 2017 · baseline B: 86.4%Harvey 2017 · baseline B+: 65.2%Harvey2017Irma 2017 · baseline A: 43.4%Irma 2017 · baseline B: 44.3%Irma 2017 · baseline B+: 32.8%Irma2017Florence 2018 · baseline A: 60.7%Florence 2018 · baseline B: 59.6%Florence 2018 · baseline B+: 51.7%Florence2018Michael 2018 · baseline A: 69.2%Michael 2018 · baseline B: 65.7%Michael 2018 · baseline B+: 49.1%Michael2018Isaias 2020 · baseline A: 59.9%Isaias 2020 · baseline B: 58.8%Isaias 2020 · baseline B+: 29.1%Isaias2020Laura 2020 · baseline A: 66.2%Laura 2020 · baseline B: 65.5%Laura 2020 · baseline B+: 36.0%Laura2020Ida 2021 · baseline A: 76.9%Ida 2021 · baseline B: 76.9%Ida 2021 · baseline B+: 36.5%Ida2021Ian 2022 · baseline A: 79.2%Ian 2022 · baseline B: 79.2%Ian 2022 · baseline B+: 56.9%Ian2022Idalia 2023 · baseline A: 70.6%Idalia 2023 · baseline B: 70.6%Idalia 2023 · baseline B+: 59.5%Idalia202390% interval coverage
Figure 4. 90% interval coverage by storm against the nominal 90% line. Every storm falls short for every model, and baseline B+ falls furthest short in every one of the ten. Coverage this far below nominal means an operator reading the p90 column as a reasonable worst case is reading something considerably milder than that.

6. Bias: the models under-predict, consistently

Averaged over storms, every model predicts a smaller outage fraction than was observed: -0.156 for baseline B+, -0.138 for baseline A. Pooled over counties the gap is starker still — a mean observed outage fraction of 0.280 against a mean prediction of 0.117. The single worst case is Irma 2017, where the published model predicted a mean outage fraction of 0.110 against 0.428 observed.

Under-prediction and narrow intervals compound rather than cancel: a distribution centred too low and too tight is wrong in the direction that matters most for staging crews. That is the first thing to fix, and it is measurable — both diagnostics on this page are the test that a fix has to move.

Table 3. Per-storm detail for baseline B+, the model whose hindcasts are published to the archive dashboard. Outage fractions are customer-weighted county means; bias is predicted minus observed.
StormCRPSSPIT–KS90% coverageObservedPredictedBiasMAERMSE
Matthew 20160.3620.45254.5%0.2440.129-0.1140.1500.253
Harvey 2017-0.1180.41965.2%0.0970.057-0.0400.1030.213
Irma 20170.1010.60232.8%0.4280.110-0.3180.3460.469
Florence 20180.4240.56951.7%0.3930.213-0.1790.2210.345
Michael 20180.2660.50649.1%0.2880.135-0.1520.1920.306
Isaias 20200.0670.64729.1%0.2730.078-0.1950.2180.304
Laura 20200.1540.59236.0%0.2870.077-0.2100.2190.339
Ida 20210.2610.64136.5%0.2080.079-0.1280.1400.244
Ian 20220.4030.49056.9%0.2420.136-0.1060.1340.222
Idalia 20230.3280.48559.5%0.2420.128-0.1140.1510.249

7. What this evaluation does not establish

8. The published hindcasts

All ten baseline B+ hindcasts are published to the Weather2Grid archive dashboard. Open the run picker and choose a storm from the Hindcast verification group: each one carries the full predictive distribution per county alongside the observed outcome, so the observed and error map layers — which a live forecast cannot have, because its outcome has not happened yet — are available for every county in the storm.

Table 5. The ten hindcast cycles published to the archive dashboard. Each carries the full predictive distribution per county plus the observed outcome, so the dashboard can draw the observed and error layers a live forecast cannot have.
StormCycle idCountiesWindow startWindow endCRPSS
Matthew 201620160929T0200Z_hindcast-baseline-bplus-MATTHEW_20163032016-09-292016-11-030.362
Harvey 201720170823T2100Z_hindcast-baseline-bplus-HARVEY_2017662017-08-232017-09-02-0.118
Irma 201720170909T1100Z_hindcast-baseline-bplus-IRMA_20172352017-09-092017-09-130.101
Florence 201820180913T0800Z_hindcast-baseline-bplus-FLORENCE_2018892018-09-132018-09-170.424
Michael 201820181009T0300Z_hindcast-baseline-bplus-MICHAEL_20183182018-10-092018-10-130.266
Isaias 202020200802T1100Z_hindcast-baseline-bplus-ISAIAS_20201822020-08-022020-08-050.067
Laura 202020200826T1100Z_hindcast-baseline-bplus-LAURA_20201392020-08-262020-08-310.154
Ida 202120210829T0200Z_hindcast-baseline-bplus-IDA_20211562021-08-292021-09-020.261
Ian 202220220927T1800Z_hindcast-baseline-bplus-IAN_20221972022-09-272022-10-050.403
Idalia 202320230829T2100Z_hindcast-baseline-bplus-IDALIA_20231532023-08-292023-09-010.328

9. Reproducibility

Every number on this page is generated from the evaluation bundle named below by scripts/build_evaluation_page.py; the hindcast payloads on the archive dashboard are written from the same bundle by scripts/publish_hindcast_evaluation.py. Neither page is edited by hand. The bundle itself is published alongside the hindcasts, so every table here can be recomputed from the artifacts it links.

Evaluation run
prism-37922961
Generated (UTC)
2026-09-04T21:03:47.490497+00:00
Code version
0.0.1
Fitset
/panfs/ccds02/nobackup/people/afahad/stormgrid/processed/fitset.parquet
Fitset SHA-256
128a32e2ce42e8ac3dee66fe94d007a806b7484c493256797028649e65f1cbeb
Evaluation content hash
1659a16c5812f6ac083f7c726585f751
Best-track source
hurdat2-1851-2025-02272026.txt
Best-track SHA-256
1b9b0c7beed5b4505838658b1d30e159fc84330c60891a58cfcf43ae55c37202
Manifest
evaluation_manifest.json
Scorecards
baseline_a · baseline_b · baseline_bplus
Score matrix
evaluation_matrix.csv · model_summary.csv · event_identity_map.csv