Detect. Forecast. Outlook.
Anvil is the hail intelligence engine behind Hail Sentinel — every storm in the country, re-scored every two minutes. It answers three questions no single instrument can: is it hailing right now, will hail reach your address in the next hour or later today, and how likely is hail in the days ahead.
Three questions, one model.
Anvil answers three questions about hail, and this page shows the working for each. On detection it outperforms the operational tools the weather industry runs on — ProbSevere v3 and MESH — on every column, at the bar each product ships. The multi-day outlook is measured separately and shares no data path with it. The early-warning pillar in between has no current benchmark, and it says so rather than reprinting an old one. Tap a card to jump to the proof.
Detect
Is it hailing right now?
Confirms almost every hail storm that people on the ground actually measured — while the hail is still falling, not the next morning.
Anvil detection accuracy 93.6% — 528 of 564 verified hail storms confirmed while the hail was falling
Forecast
Will hail reach you in the next hour, or later today?
Turns the storms it is already tracking into a warning for one address — and carries the risk on through the rest of the day.
A median 33.3 minutes of notice on warned locations at the T+30 horizon, measured at the gate that sends a notification
Outlook
How likely is hail in the days ahead?
Flags most hail days before they arrive — a median 2 days ahead.
Flagged 93 of 100 real hail days in the backtest, and showed at least a marginal risk on 88% of them on the day itself
Comparison covers public operational baselines; independent third-party benchmarking and the full research paper are planned for 2026.
Every storm watched, scored live, and segmented to the pixel — the three stages are laid out on How it works.
Detect.
Every two minutes, Anvil gives every storm across the contiguous U.S. a calibrated hail risk score. The figures here count whole storms, the way you would: one storm that verifiably dropped damaging hail, caught or missed. Set beside the tools the industry runs on — the real-time storm model (ProbSevere v3) and the standard radar hail estimate (MESH) — Anvil catches more storms and a higher share of its alarms turn out to be real.
Detection accuracy is the share of verified hail storms Anvil confirmed while the hail was falling, over storms carrying 2 or more independent ground reports — definition, excluded tail and clustering rule in the methodology notes.
One storm, start to finish.
A single real storm, rendered from the production radar output and coloured by the shader the app ships, driven by the probabilities Anvil produced for it at the time.
- the path it has already travelled
- where its own motion projects it next
Two readings, one picture: how far a pixel sits from grey is Anvil’s chance of hail there, and the colour it warms to is the radar hail-size estimate. Bolder always means more confident.
Anvil calls a detection at 15.7% and this storm peaked at 53.1%. That bar is tied to the serving model’s own scale and is re-derived every time the model is retrained.
When a storm really hailed, Anvil knew.
Anvil’s detection accuracy this season was 93.6%: of the 564 storms that verifiably dropped hail an inch or bigger, it flagged 528 while the hail was still falling.
564 storms carrying 2 or more ground reports, 2026-05-01 to 2026-07-30 · what this counts and excludes: methodology notes.
Quiet when hail is impossible.
Two tests of the same instinct: how often each tool fires in weather where hail is physically impossible, and whether a phone ever rings on a day when no hail fell anywhere.
Fire rate per million tracks, measured only in environments where hail is meteorologically impossible. No growth zone, no updraft, no instability. A model firing here is firing on nothing — no underreporting caveat applies.
Winter snow pellets: not shown — 820 tracks in this window and no system fired in any of them, so the regime separates nothing
On a day with no hail, your phone stays dark.
Over 52 audited days with no hail anywhere, 96.2% passed without a single warning. On clear days and during winter storms the count was zero — both hard release gates.
0.077 false warnings per quiet day, every one on a rainy day · measured at the gate that rings a phone, a separate instrument from the regime figure · release-board checks and cohort in the methodology notes.
One number, and it can’t be gamed from either side.
Catching more hail is easy if you shout constantly. Being right when you shout is easy if you almost never do. The critical success index — the number the National Weather Service scores warnings on — refuses both moves.
All three absolutes are deflated by the same two mechanisms, identically for every system — both are in the methodology notes. The multiple is the comparison. Measured over 4,430 corroborated storms.
5 settings on Anvil's dial, including the shipped one, sit up AND to the right of both baselines at once.
Shaded bands are threat score — one number combining catches and false alarms, rising toward the top right. Dotted diagonals are alarms raised per storm that really hailed.
Up is more hail caught, right is fewer wasted alarms, the diagonals are alarm budgets · 4,430 corroborated storms, 2026-05-01 to 2026-07-30 · how to read it in full: methodology notes.
Give Anvil the same number of alarms.
Take the rival’s alarm count, give Anvil exactly that many, and see who catches more. Anvil’s bar isn’t chosen for this — it is derived, as whatever score lands on the rival’s 28,394th most confident storm.
21.1 pp more damaging hail caught, for the same number of alarms
1.35× the threat score at that budget — 0.112 against 0.083, corroborated storms
Anvil’s scores tie at that cut, so it fires on 29,717 tracks — 4.7% above the budget. The margin is quoted with that overshoot attached, never without it.
24.7 pp more damaging hail caught, for the same number of alarms
1.67× the threat score at that budget — 0.158 against 0.095, corroborated storms
At its own bar, on the storms ProbSevere covers, Anvil catches 79.2% from 12,647 alarms against 76.2% from 16,795 — a different operating point, detailed in the methodology notes.
It catches more of the hail that’s actually out there.
Every system is graded at the bar it actually ships. One definition of real hail, the same storms, counted once per storm. The gap widens as the hail gets worse.
Share of real hail of each size that each tool caught. Coverage and the baselines’ selectivity are in the methodology notes.
If you don’t see hail at your spot, it probably fell nearby.
Most hail never gets officially reported. When Anvil fires and nobody filed anything, the hail was usually still real.
Most of what we miss, nobody catches.
Of the damaging storms Anvil missed, 79.2% were missed by every other system too — a shared blind spot in the physics, not a gap only we have.
20.1% of damaging hail this season was invisible to every instrument in the comparison.
- Every system caught it 1,435 33.5%
- Only some systems caught it 1,990 46.5%
- No system caught it 859 20.1%
Of the storms exactly one system caught
- Anvil 641
- Radar hail size (MESH) 94
- ProbSevere 92
Storms this system caught and the other two did not.
Of the storms Anvil missed, how many everyone missed
- 79.2% Every damaging storm in the window 4,284 storms · a storm the baseline never scored counts as a miss for it
- 63.7% Only storms all three systems looked at 2,354 storms · the like-for-like cohort, with the coverage gap removed
4,284 corroborated storms of 25.4 mm and up · all three together reach 80.0% · why the two shares differ: methodology notes.
Same alert budget. More storms caught.
Both systems are pinned to the same false-alarm burden — the same number of quiet storms alarmed on — and only then compared. Whatever is left is skill.
At the identical alert burden Anvil catches +3.3 pp more of the verified hail storms than the baseline nationwide — 71.9% against 68.6%, with a 95% interval of +1.74 to +5.95 pp, clustered by day.
per-track (national + 2 region strata + 3 elevation-band strata), 15km primary match radius + 25km robust match radius · 2026-05-01 → 2026-07-30 · a separate instrument, measured on all truth, never quoted alongside the figures above · west of the Rockies the two systems are level and the stratum is published in words, not drawn — methodology notes.
It isn’t one good week — and it isn’t just “hail”.
It isn’t one good week.
A three-month average can be carried by a few enormous days, so every storm day was scored on its own storms. Anvil ranked the day’s hail better than the industry baseline on 66 of 83 days, and better than the standard radar estimate on 78. No ties.
83 storm days, each scored independently · one measured cell narrows under this scope and is published in the methodology notes.
Not just whether. How big.
Anvil publishes a size band, measured against what people put on the ground. Across every matched report the band is 34% closer than the raw radar estimate, and on 121 of 126 hail days it was the better estimate.
pooled matched-report size pairs, ungated (all admitted candidates), 3 severity brackets · events_v2, 126 hail-day events · millimetres, not catch rate, and not comparable to any storm-counting figure here · the bracket we lose and why: methodology notes.
Forecast.
Detection tells you it’s hailing now. Forecast tells you it’s coming to your spot — in the next hour, and across the rest of the day. Anvil takes the storms it’s already tracking, projects them forward, and works out whether hail will actually reach a specific address before it gets there — address-specific tracking few weather tools offer.
Two instruments answer this pillar and neither is quoted through the other. The storm-projection nowcast below is address-grain and minutes-scale. Past the reach of a storm already on radar, the same day is carried by a separate same-day layer, measured on its own terms further down this page.
Same hail caught, fewer places warned.
Both methods are cut where they catch the same number of confirmed hail events, so the only question left is how many locations each had to warn to get there. Anvil warns 1.34× fewer, and the advantage is measurable at 30 and 60 minutes. At T+15 it is not — that range crosses parity, and the free tier serves that horizon.
- Ratio
- 95% range · 20 storm days
- Parity
20 storm days · 3,489,685 point-forecasts, storm-blind · the baseline is Lagrangian persistence, Anvil’s own previously shipped display · methodology notes.
Enough time to move the car.
A different artifact, on its own window: on locations that got a warning, how much notice actually arrived. The median was 33.3 minutes at the near horizon. The further out you look, the more time you get and the fewer storms it catches — which is why only the near horizon is allowed to ring a phone. On clear days it delivered no warnings at all, at every horizon.
| Horizon | Notice delivered median minutes, warned locations | Locations warned across the graded events | Share of hail caught at the graded bar |
|---|---|---|---|
| T+30 minutes out the only horizon that sends a notification | 33.3 | 326.0 | 13.10% |
| T+60 minutes out panel content — never a push | 63.3 | 217.0 | 8.30% |
| T+90 minutes out panel content — never a push | 93.3 | 135.0 | 5.00% |
2025-03-01..2026-06-10 · Anvil’s own delivered notice on warned locations, no baseline implied · what it measures, and why only the near horizon pushes: methodology notes.
Outlook.
Radar can only see storms that already exist. The Outlook looks further out: it puts a daily hail chance on your area days ahead, from SPC outlook + HRRR convective fields + climatology. And the odds are meant to be read literally rather than as a mood: over 17,581 address-days in July 2026, on data the model had never trained on, hail turned up on 14.15% of the days its hail signal crossed the bar the app ships — 2,487 of 17,581, roughly one day in 7.
It carries publishable skill from day-of through day-3; beyond that the outlook is still produced, but at this refresh it no longer clears the bar we hold it to, and we say which leads and by how much rather than drawing them on the curve. That out-of-sample rate is bar-conditional — reliability given an alert, not an unconditional hit rate. Window, grain and method in the methodology notes.
Skillful through day-3, and it says where it stops.
A multi-day forecast has to tell hail days apart from quiet days, and the number it prints has to be worth acting on. It flagged 93 of 100 real hail days in the backtest and showed at least a marginal risk on 88% of them on the day itself, with the risk first appearing a median of 2 days beforehand and at most 6 days.
Day-of AUC 0.925 (95% CI 0.889–0.962, n=100); false alarms stay at or under 2.8% across every plotted lead.
Backtest of 100 real hail days against matched controls, built from SPC outlook + HRRR convective fields + climatology · methodology notes.
Where the official outlook stops.
This is the same-day half of the Forecast pillar above: past the reach of a storm already on radar, but inside the day. It sits here because everything it claims is stated relative to the official day-1 outlook.
Anvil Today is not a rival to the official day-1 hail outlook — it is a layer that sits on top of it. Where the official outlook is drawn, Anvil passes it through exactly as issued. Its job starts where that outlook stops.
Squared forecast error down 28.4% against the uncalibrated layer, replicated at 29.0% on an earlier fold · clears 4 of 4 hard gates · methodology notes.
The proof is above. Here’s how to use it.
The same engine benchmarked on this page already runs live — pointed at the addresses, lots, and fleets that matter to you.
Get Anvil watching your address
Real-time alerts, 15–60 minute early warning, and a daily outlook for your home and vehicles.
Explore the personal appProtect your lot or fleet
Multi-location monitoring, calibrated hail data, and API access for the systems your operation runs on.
Explore business solutions