Skip to content
Three pillars DETECT · FORECAST · OUTLOOK

Three questions, one model.

Anvil answers three questions about hail, and this page shows the working for each. On detection it outperforms the operational tools the weather industry runs on — ProbSevere v3 and MESH — on every column, at the bar each product ships. The multi-day outlook is measured separately and shares no data path with it. The early-warning pillar in between has no current benchmark, and it says so rather than reprinting an old one. Tap a card to jump to the proof.

Comparison covers public operational baselines; independent third-party benchmarking and the full research paper are planned for 2026.

Every storm watched, scored live, and segmented to the pixel — the three stages are laid out on How it works.

Pillar I Is it hailing right now?

Detect.

Every two minutes, Anvil gives every storm across the contiguous U.S. a calibrated hail risk score. The figures here count whole storms, the way you would: one storm that verifiably dropped damaging hail, caught or missed. Set beside the tools the industry runs on — the real-time storm model (ProbSevere v3) and the standard radar hail estimate (MESH) — Anvil catches more storms and a higher share of its alarms turn out to be real.

93.6% Anvil detection accuracy, storm by storm
564 Verified hail storms behind that figure

Detection accuracy is the share of verified hail storms Anvil confirmed while the hail was falling, over storms carrying 2 or more independent ground reports — definition, excluded tail and clustering rule in the methodology notes.

The mechanism ONE STORM · THEN EVERY STORM

One storm, start to finish.

A single real storm, rendered from the production radar output and coloured by the shader the app ships, driven by the probabilities Anvil produced for it at the time.

One storm · not a measurement Northeast Michigan · 27 July 2026
Looping animation of a single storm over Northeast Michigan. Its radar footprint grows and pulls away from grey as Anvil's hail probability rises to its peak, then fades back to grey as the storm collapses, with the path it has travelled and the path its motion projects drawn alongside.
Two questions, one strip each
Could it hail? nomaybelikely
How big? peaquartergolf ballbaseball
Drawn over the storm
  • the path it has already travelled
  • where its own motion projects it next

Two readings, one picture: how far a pixel sits from grey is Anvil’s chance of hail there, and the colour it warms to is the radar hail-size estimate. Bolder always means more confident.

Anvil calls a detection at 15.7% and this storm peaked at 53.1%. That bar is tied to the serving model’s own scale and is re-derived every time the model is retrained.

Its probability climbs as the storm organises and decays with it. 3 ground reports confirmed hail on this track — one storm, not a measurement. 28 of its 39 radar scans, looping — the quiet opening and ending are thinned so the active phase keeps its full cadence.

When a storm really hailed, Anvil knew.

Anvil’s detection accuracy this season was 93.6%: of the 564 storms that verifiably dropped hail an inch or bigger, it flagged 528 while the hail was still falling.

Confirmed while the hail was falling
ANVIL ONLY · ONE OPERATING POINT
Higher is better

564 storms carrying 2 or more ground reports, 2026-05-01 to 2026-07-30 · what this counts and excludes: methodology notes.

False alarms 2026-05-01 → 2026-07-30 · WEATHER THAT CANNOT MAKE HAIL

Quiet when hail is impossible.

Two tests of the same instinct: how often each tool fires in weather where hail is physically impossible, and whether a phone ever rings on a day when no hail fell anywhere.

Fire rate per million tracks, measured only in environments where hail is meteorologically impossible. No growth zone, no updraft, no instability. A model firing here is firing on nothing — no underreporting caveat applies.

How much more often each system fires where hail is possible than where it is not
25.1× Anvil
15.1× ProbSevere v3
6.6× MESH
50,709 storms in hail-capable weather against 377,277 where hail cannot form. A system that fires indiscriminately scores near 1× however loud or quiet it is overall. Regimes are classified from forecast-model fields that Anvil reads as inputs and MESH does not, so Anvil partly knows which regime it is in — correct behaviour for a model, but not a blind test for it the way it is for MESH.
Across 377,277 storms in weather where hail can’t form
Anvil 1,821/M
ProbSevere v3 1,235/M
MESH 4,750/M
MESH false-fires 2.6× as often as Anvil . Anvil fires 1.5× as often as ProbSevere v3, which sits at a tighter bar and scores barely half the storms — the same trade that buys Anvil its catch-rate lead .
No storm energy
hot and humid, but no thunderstorm fuel — nothing to grow ice
116,302 tracks in regime
Anvil 1,195 /M
ProbSevere 834 /M
MESH 1,135 /M
Tropical warm rain
air too warm aloft for ice to form — pure rain
260,975 tracks in regime
Anvil 2,100 /M
ProbSevere 1,414 /M
MESH 6,361 /M
Anvil false-fired 548 times in 260,975 storms. MESH fired 1,660 — 3.0× more.

Winter snow pellets: not shown — 820 tracks in this window and no system fired in any of them, so the regime separates nothing

On a day with no hail, your phone stays dark.

Over 52 audited days with no hail anywhere, 96.2% passed without a single warning. On clear days and during winter storms the count was zero — both hard release gates.

0.077 false warnings per quiet day, every one on a rainy day · measured at the gate that rings a phone, a separate instrument from the regime figure · release-board checks and cohort in the methodology notes.

CLEAR DAYS · 18 AUDITED
0
False warnings on a clear day
A hard release gate — the target is exactly zero
WINTER STORMS · 15 AUDITED
0
False warnings during a winter storm
A hard release gate — the target is exactly zero
RAIN, NO HAIL · 19 AUDITED
0.211
False warnings per rainy day with no hail
The gate allows 3.15 — this is the only category that fires at all
One number CRITICAL SUCCESS INDEX · THE NWS WARNING METRIC

One number, and it can’t be gamed from either side.

Catching more hail is easy if you shout constantly. Being right when you shout is easy if you almost never do. The critical success index — the number the National Weather Service scores warnings on — refuses both moves.

VS PROBSEVERE V3
1.34×
Anvil’s threat score, relative to the research baseline
0.127 against 0.095 — corroborated storms, same truth
VS RADAR HAIL SIZE (MESH)
1.53×
Anvil’s threat score, relative to the radar standard
0.127 against 0.083 — corroborated storms, same truth

All three absolutes are deflated by the same two mechanisms, identically for every system — both are in the methodology notes. The multiple is the comparison. Measured over 4,430 corroborated storms.

The whole trade, in one picture
19 SETTINGS · TWO FIXED BASELINES
ANVIL IS A DIAL. THE BASELINES ARE PLOTTED WHERE THEY OPERATE.
0% 20% 40% 60% 80% 100% 0% 20% 40% 60% 80% 100% 0.1 0.2 0.3 0.5 0.7 CSI Success ratio — alarms that were real further right = fewer false alarms → Probability of detection ↑ higher = more storms caught Anvil ProbSevere MESH

5 settings on Anvil's dial, including the shipped one, sit up AND to the right of both baselines at once.

Anvil 73.8% 13.3% 0.127 5.6× the setting the app ships at
ProbSevere v3 41.4% 10.9% 0.095 3.8× at its published 0.30 bar
Radar hail size (MESH) 56.7% 8.8% 0.083 6.4× at the 1-inch severe bar

Shaded bands are threat score — one number combining catches and false alarms, rising toward the top right. Dotted diagonals are alarms raised per storm that really hailed.

Up is more hail caught, right is fewer wasted alarms, the diagonals are alarm budgets · 4,430 corroborated storms, 2026-05-01 to 2026-07-30 · how to read it in full: methodology notes.

Give Anvil the same number of alarms.

Take the rival’s alarm count, give Anvil exactly that many, and see who catches more. Anvil’s bar isn’t chosen for this — it is derived, as whatever score lands on the rival’s 28,394th most confident storm.

Damaging hail caught, at the rival's own alarm count
1-INCH HAIL AND UP · CORROBORATED STORMS · SAME ALARMS
At the radar estimate’s alarm count 28,394 alarms · 4,284 damaging storms
Anvil 78.6%
Radar hail size (MESH) 57.5%

21.1 pp more damaging hail caught, for the same number of alarms

1.35× the threat score at that budget — 0.112 against 0.083, corroborated storms

Anvil’s scores tie at that cut, so it fires on 29,717 tracks — 4.7% above the budget. The margin is quoted with that overshoot attached, never without it.

At ProbSevere’s alarm count 16,795 alarms · 4,284 damaging storms
Anvil 66.6%
ProbSevere 41.9%

24.7 pp more damaging hail caught, for the same number of alarms

1.67× the threat score at that budget — 0.158 against 0.095, corroborated storms

At its own bar, on the storms ProbSevere covers, Anvil catches 79.2% from 12,647 alarms against 76.2% from 16,795 — a different operating point, detailed in the methodology notes.

Coverage 4,430 CORROBORATED STORMS · EVERY SIZE

It catches more of the hail that’s actually out there.

Every system is graded at the bar it actually ships. One definition of real hail, the same storms, counted once per storm. The gap widens as the hail gets worse.

If you don’t see hail at your spot, it probably fell nearby.

Most hail never gets officially reported. When Anvil fires and nobody filed anything, the hail was usually still real.

96%
Of 22,574 storm tracks that showed the radar’s own hail fingerprint, no report was ever filed.
That is the gap every raw “false alarm” number ignores. Rural hail, overnight hail, hail that melted before anyone drove out to see it — none of it makes the verification record. Treat 96% as an upper bound: the radar “fingerprint” is a semi-independent check — it shares some radar inputs with Anvil, so it confirms rather than fully proves. Anvil’s detection is validated primarily against human reports and agreement across independent sources; this radar gap is supporting evidence, not the headline number.
7,016
Report-matched storm tracks, 2026-05-01 to 2026-07-30
LSR + SPC, deduped, severe only (≥ 19.05 mm). For every reported hail track, the radar saw 3.2× more tracks carrying the radar’s own hail fingerprint.

Most of what we miss, nobody catches.

Of the damaging storms Anvil missed, 79.2% were missed by every other system too — a shared blind spot in the physics, not a gap only we have.

Every damaging storm in the window, by who caught it
4,284 STORMS · ONE PARTITION

20.1% of damaging hail this season was invisible to every instrument in the comparison.

  • Every system caught it 1,435 33.5%
  • Only some systems caught it 1,990 46.5%
  • No system caught it 859 20.1%

Of the storms exactly one system caught

  • Anvil 641
  • Radar hail size (MESH) 94
  • ProbSevere 92

Storms this system caught and the other two did not.

Of the storms Anvil missed, how many everyone missed

  • 79.2% Every damaging storm in the window 4,284 storms · a storm the baseline never scored counts as a miss for it
  • 63.7% Only storms all three systems looked at 2,354 storms · the like-for-like cohort, with the coverage gap removed

4,284 corroborated storms of 25.4 mm and up · all three together reach 80.0% · why the two shares differ: methodology notes.

Head to head 1,749,335 STORMS · 91 DAYS · ONE ALERT BUDGET

Same alert budget. More storms caught.

Both systems are pinned to the same false-alarm burden — the same number of quiet storms alarmed on — and only then compared. Whatever is left is skill.

At the identical alert burden Anvil catches +3.3 pp more of the verified hail storms than the baseline nationwide — 71.9% against 68.6%, with a 95% interval of +1.74 to +5.95 pp, clustered by day.

2.0×
the ranking mistakes
ProbSevere v3
pairs of storms put in the wrong order
Anvil
pairs of storms put in the wrong order
6,573 verified hail storms among 1,749,335 tracked, 91 consecutive days
Catch rate at a matched alert burden
5 CUTS OF ONE COHORT
EVERY ROW IS THE SAME STORMS, SCORED TWICE Higher is better
Anvil ProbSevere

per-track (national + 2 region strata + 3 elevation-band strata), 15km primary match radius + 25km robust match radius · 2026-05-01 → 2026-07-30 · a separate instrument, measured on all truth, never quoted alongside the figures above · west of the Rockies the two systems are level and the stratum is published in words, not drawn — methodology notes.

Day by day, and by size 83 STORM DAYS · 9,816 MATCHED REPORTS

It isn’t one good week — and it isn’t just “hail”.

It isn’t one good week.

A three-month average can be carried by a few enormous days, so every storm day was scored on its own storms. Anvil ranked the day’s hail better than the industry baseline on 66 of 83 days, and better than the standard radar estimate on 78. No ties.

Storm days won, head to head
83 DAYS · 100 ICONS
Higher is better

83 storm days, each scored independently · one measured cell narrows under this scope and is published in the methodology notes.

Not just whether. How big.

Anvil publishes a size band, measured against what people put on the ground. Across every matched report the band is 34% closer than the raw radar estimate, and on 121 of 126 hail days it was the better estimate.

How far off the size estimate lands
MILLIMETRES · LOWER IS BETTER
EVERY MATCHED REPORT, POOLED Lower is better

pooled matched-report size pairs, ungated (all admitted candidates), 3 severity brackets · events_v2, 126 hail-day events · millimetres, not catch rate, and not comparable to any storm-counting figure here · the bracket we lose and why: methodology notes.

Pillar II Will hail reach you in the next hour, or later today?

Forecast.

Detection tells you it’s hailing now. Forecast tells you it’s coming to your spot — in the next hour, and across the rest of the day. Anvil takes the storms it’s already tracking, projects them forward, and works out whether hail will actually reach a specific address before it gets there — address-specific tracking few weather tools offer.

Two instruments answer this pillar and neither is quoted through the other. The storm-projection nowcast below is address-grain and minutes-scale. Past the reach of a storm already on radar, the same day is carried by a separate same-day layer, measured on its own terms further down this page.

Same hail caught, fewer places warned.

Both methods are cut where they catch the same number of confirmed hail events, so the only question left is how many locations each had to warn to get there. Anvil warns 1.34× fewer, and the advantage is measurable at 30 and 60 minutes. At T+15 it is not — that range crosses parity, and the free tier serves that horizon.

Locations warned at matched catches
PROBABILITY FORECAST · ANVIL vs PERSISTENCE
120,920 Locations warned — Anvil
161,698 Locations warned — the standard method
T+30 min free tier
1.35× fewer
T+60 min premium
1.35× fewer
  • Ratio
  • 95% range · 20 storm days
  • Parity

20 storm days · 3,489,685 point-forecasts, storm-blind · the baseline is Lagrangian persistence, Anvil’s own previously shipped display · methodology notes.

Enough time to move the car.

A different artifact, on its own window: on locations that got a warning, how much notice actually arrived. The median was 33.3 minutes at the near horizon. The further out you look, the more time you get and the fewer storms it catches — which is why only the near horizon is allowed to ring a phone. On clear days it delivered no warnings at all, at every horizon.

Notice delivered, and what it cost
SEPARATE ARTIFACT · ANVIL ONLY, NO BASELINE
Median delivered notice in minutes, locations warned, and share of hail reports caught, at each forecast horizon.
Horizon Notice delivered median minutes, warned locations Locations warned across the graded events Share of hail caught at the graded bar
T+30 minutes out the only horizon that sends a notification 33.3 326.0 13.10%
T+60 minutes out panel content — never a push 63.3 217.0 8.30%
T+90 minutes out panel content — never a push 93.3 135.0 5.00%

2025-03-01..2026-06-10 · Anvil’s own delivered notice on warned locations, no baseline implied · what it measures, and why only the near horizon pushes: methodology notes.

Pillar III How likely is hail in the days ahead?

Outlook.

Radar can only see storms that already exist. The Outlook looks further out: it puts a daily hail chance on your area days ahead, from SPC outlook + HRRR convective fields + climatology. And the odds are meant to be read literally rather than as a mood: over 17,581 address-days in July 2026, on data the model had never trained on, hail turned up on 14.15% of the days its hail signal crossed the bar the app ships — 2,487 of 17,581, roughly one day in 7.

93/100 Real hail days flagged in the backtest
2d Median lead before a hail day
88% Showed a risk on the day itself

It carries publishable skill from day-of through day-3; beyond that the outlook is still produced, but at this refresh it no longer clears the bar we hold it to, and we say which leads and by how much rather than drawing them on the curve. That out-of-sample rate is bar-conditional — reliability given an alert, not an unconditional hit rate. Window, grain and method in the methodology notes.

Skillful through day-3, and it says where it stops.

A multi-day forecast has to tell hail days apart from quiet days, and the number it prints has to be worth acting on. It flagged 93 of 100 real hail days in the backtest and showed at least a marginal risk on 88% of them on the day itself, with the risk first appearing a median of 2 days beforehand and at most 6 days.

Skillful through day-3 SEPARATION SKILL BY LEAD · 95% CI BAND · N=100 EVENTS
Separation skill (0–1) 0.70 0.80 0.90 1.00 0.925 Day-of 0.850 +1d 0.814 +2d 0.743 +3d

Day-of AUC 0.925 (95% CI 0.889–0.962, n=100); false alarms stay at or under 2.8% across every plotted lead.

Backtest of 100 real hail days against matched controls, built from SPC outlook + HRRR convective fields + climatology · methodology notes.

Where the official outlook stops.

This is the same-day half of the Forecast pillar above: past the reach of a storm already on radar, but inside the day. It sits here because everything it claims is stated relative to the official day-1 outlook.

Anvil Today is not a rival to the official day-1 hail outlook — it is a layer that sits on top of it. Where the official outlook is drawn, Anvil passes it through exactly as issued. Its job starts where that outlook stops.

Anvil Today, on its own terms
A LAYER ON THE OFFICIAL OUTLOOK · NOT A RIVAL TO IT
0 Cells where Anvil overrode the official outlook Across 136,088 cells with a drawn contour, the largest disagreement was zero. Not “small” — zero.
1,598,696 Cell-days that got a number instead of silence Most of the country, most days, sits outside any drawn contour. That is not the same as “no chance of hail”, and Anvil puts a calibrated figure there.
0.45pp Average gap between the odds printed and what happened Down from 0.83 pp — 46% of the gap between the odds printed and the odds observed, removed.
15% Worst-case firing rate on days with no hail anywhere Against a 30% ceiling, across 74 confirmed hail-free days. Quiet days have to look quiet.

Squared forecast error down 28.4% against the uncalibrated layer, replicated at 29.0% on an earlier fold · clears 4 of 4 hard gates · methodology notes.

Put Anvil to work PERSONAL · BUSINESS

The proof is above. Here’s how to use it.

The same engine benchmarked on this page already runs live — pointed at the addresses, lots, and fleets that matter to you.

Full research paper coming 2026 — preview how Anvil works
Get free hail alerts