Hail Sentinel exists to answer one question for somebody who owns a roof: is the storm overhead putting hail on my address, and did the system know while it was happening? Anvil Gen 1 is the first release where we can answer that with a figure measured at the grain a person actually lives at — a storm, over a place, during an hour — rather than at the grain a radar engineer lives at.

Gen 1 is the first complete generation of the hail intelligence platform: four named models running in production under one naming scheme. Two are trained machine-learning models. Two are statistical, and they say so in their own class lines rather than leaving you to assume. What follows is what each one does and how it measures against the operational baselines the industry already runs on.

The four models

All four carry a version in their name, because a version is a promise: a generation you can cite, hold us to, and watch change. The scheme is anvil-JOB-VERSION — the middle word names what the model is for, never how it is built. That is why two of these names sit on statistical models without any sleight of hand, and it is also why each of those two prints its class line: the name tells you the job, the class line tells you the method, and neither is left to inference.

anvil-detect-6 — the detection model

Trained machine-learning model. It scores every storm object on every radar scan for whether that storm is putting hail on the ground right now. It is what stands behind an alert arriving while the sky is still dark, and it is the thing the accuracy figure below is about.

anvil-nowcast-2 — the address-level nowcast

Trained machine-learning model. It takes the storms the detection model is already scoring and carries hail probability forward onto one point on the map — an address, not a county. It is registered, it is serving, and business customers already receive it through the API.

anvil-today-3 — the same-day layer

It re-levels the day's hail outlook against what actually verifies. It is live in production, it cleared 4 of 4 hard gates, and it was re-certified on 1,734,784 out-of-sample cell-days. Its class line, verbatim from the artifact: Calibration layer. Not a trained machine-learning model. Read that line before reading any figure attached to this model — what it improves is calibration, not skill.

anvil-outlook-1 — the multi-day outlook

It carries the picture out across days 2 to 7, verified over 135 issuance dates. Its class line, verbatim: Calibrated statistical model. Not a trained machine-learning model. There is no trained machine-learning model behind this name, and nothing here should be read as announcing one.

Does it know while the hail is falling?

Every other detection figure we publish counts radar objects or radar scans. Neither is a thing anybody experiences. A storm is. So the headline for Gen 1 is measured over storm events: deduplicated ground reports clustered into the storm that produced them, with a storm counted as confirmed when anvil-detect-6 cleared its shipped bar inside the window the hail was actually falling in.

Higher is better
Storms carrying at least 2 independent ground reports, over the campaign window. Storms with no radar object at all are counted as misses rather than excused.

93.6% of the storms in the published cohort — 528 of 564 — were confirmed while the hail was falling, measured over storms that at least 2 independent observers on the ground reported. Corroboration buys cleaner truth: every storm in this cohort is one somebody stood under and wrote down.

Against the operational baselines

At track grain — one radar object's life — the balanced measure is the threat score, also called the critical success index. It charges misses and false alarms to the same account, so no system can lift it by alarming more often. Measured on corroborated truth across 1,800,879 storm tracks and 4,430 truth tracks, Anvil leads both external baselines:

1.34×
the threat score of ProbSevere v3
0.1268
Anvil threat score
0.0947
ProbSevere v3 threat score
Track grain · corroborated truth · coverage 100% against 49%
1.53×
the threat score of MESH
0.1268
Anvil threat score
0.0829
MESH threat score
Track grain · corroborated truth · same cohort, same window, same truth definition
The multiple is the claim. Both absolutes render inside the same card so neither ratio can be quoted apart from the numbers it came from.

Corroborated truth is a harder scope for every system in the comparison: it drops true positives from all three without dropping a single false alarm, so all three absolutes sit lower than they would elsewhere. The ratio is what survives that, which is why it is what we publish.

The threat score is the balanced measure. The catch rate is the one a roof cares about — and the objection to any catch rate is that a system can buy it by alarming more often. So hold Anvil to the alarm burden each baseline itself spends, and count again:

Anvil +21.1 points at the same alarm burden Higher is better Alarm burden
Anvil +8.5 points at the same alarm burden Higher is better Alarm burden
Share of hail-producing storm tracks each system caught, corroborated truth, with Anvil cut to the alarm burden the baseline spent. MESH over all storm tracks (4,430 truth tracks); ProbSevere v3 over tracks ProbSevere v3 scores (2,439 truth tracks), because it produces no value on the rest.

The season does not reduce to one pooled number either. Taken day by day across the storm days in the campaign, Anvil holds the higher threat score on most of them:

DAY BY DAY · CORROBORATED TRUTH
66 of 83
storm days Anvil scored higher than ProbSevere v3
DAY BY DAY · CORROBORATED TRUTH
78 of 83
storm days Anvil scored higher than MESH
Day-level threat score, corroborated truth, 83 storm days in the campaign window.

The address-level nowcast

anvil-nowcast-2 answers a narrower question than the detection model: given the storms on radar now, what is the hail probability at this point on the map over each horizon ahead. A horizon here is the distance ahead the model is scored at. It is not a promise about anybody's next hour, and we publish no figure that describes one.

The baseline it had to beat is the leg it replaced: Lagrangian persistence — each currently-observed storm keeps its present hail probability and advects along its own measured motion vector. Cut both arms at the threshold that reaches the same number of confirmed catches, then count what each one raised to get there.

1.34×
the alarms of persistence, at the same catches
161,698
alarms raised by persistence
120,920
alarms raised by anvil-nowcast-2
Lattice grain · 2026-07-14 to 2026-08-01, 20 storm days · both arms cut at the same catch count · day-clustered 95% interval 1.035–1.865
Fewer locations put on alarm for the same number of confirmed catches. Both arms are cut at the same catch count, so neither side can buy the ratio by alarming more.

The quiet days count too

A hail system that never misses and never shuts up is worth nothing. The product is a phone that stays silent until it should not be, so Gen 1 is measured on silence as well. The rates below are false-alarm rates in the strict sense — measured against days and regimes where the absence of hail is established rather than assumed, which is the only place the phrase is honest.

1,041 REAL-DATA CASES
0
clear-day false alarms across the nowcast battery
NOWCAST BATTERY · 25/25 GATES
0
winter false alarms in the same battery
SAME-DAY LAYER · 74 NULL DATES
15%
worst false-alarm rate on a null storm date, against a 30% bar
Discipline on quiet ground. The battery is an offline end-to-end run over real data, not a measurement of live production behaviour.

The same-day layer's contribution shows up in the same place. Over its re-certification set the day's published numbers move measurably closer to what the day actually does:

Calibration error down 46% Lower is better
Expected calibration error — how far the day's stated probabilities sit from the rate they verify at. Lower is better. What this layer improves is calibration, not skill.

Notes and scope

  • The storm cohort. 93.6% is confirmation over 564 storms carrying at least 2 independent ground reports. Requiring two observers selects for bigger storms over more populated ground, so it describes well-witnessed hail rather than all hail. Across all 1,343 storms in the window the same model at the same bar confirms 82.1%, and the 779 storms resting on a single report sit at 73.7% on their own. Storms with no radar object at all are counted among the misses: 4 of 36.
  • No baseline at storm grain. The comparison figures are at track grain because that is where a budget-matched measurement exists. At storm grain none was run, and printing the baselines' rates there would rank the systems on how often each one alarms.
  • Corroborated truth. Every comparison on this page is scored in that scope: a truth row is kept only when its ground truth attributes to a storm at least 2 people reported, applied identically to all three systems by a rule that reads no system's score.
  • The nowcast ratio is pooled. Anvil raised fewer alarms on 12 of the 20 storm days in that window; the pooled interval clears 1.0 because the winning days are the high-volume ones. The separation lives at T+30 and T+60.
  • Two of the four are statistical. anvil-today-3 — Calibration layer. Not a trained machine-learning model. anvil-outlook-1 — Calibrated statistical model. Not a trained machine-learning model. A figure from either is not evidence about a trained model.
  • Horizons are scored distances, not promises. Every horizon here is a distance ahead a model is graded at, after the fact. None of it is a statement about what reaches a person, or when, and we publish no figure that describes one.
  • Offline measurement. The detection and comparison figures are an offline re-scoring of the models across historical rows — the right way to obtain a sample this size, and not a measurement of live behaviour. Scope: Contiguous United States only.

Where to go next

If you want the data rather than the argument, the business API serves detections and address-level nowcasts directly; endpoints, payloads and rate limits are in the API documentation. The longer methodology write-up — operating curves, the size ladder and the sensitivity arms — is on How Anvil is measured.