Most weather apps don't tell you how often they're wrong. We do.
Every three hours we snapshot what each of our five forecasts said. Then we check
every one of them, hour by hour, against an independent record of what the weather
actually did, which comes from none of the five forecasts being marked. The numbers
below are the result.
What we measure, and what good looks like
Four numbers per source. Each is colour-coded against what "good" actually means for that measure. They're deliberately different bars, because some things are much harder to forecast than others:
- Temperature error: the average gap between forecast and actual, in °C. Lower is better. Under 1° is good; the very best manage around half a degree. A score of 0.8° means the forecast was typically just under a degree off.
- Wind error: the average gap between forecast and actual wind speed, in km/h. Wind is locally gusty and harder to pin down, so under 5 km/h is good.
- Caught the rain: of the hours it actually rained, how many did the source call? A source "called it" when it gave the same rain chance the app itself treats as wet, currently 35%. That matters: until August 2026 this page graded at 50% while the app labelled a day wet at 25%, so it published a figure for a stricter app than the one you were using, and understated it badly. It now measures what you actually see. The bands moved up with it, so green still means the same thing. The bar is set where the real world sits rather than where we might like it: 50%+ is good. Catching more rain is trivially easy if you are willing to be wrong more often, so this number should always be read next to the false alarm rate below it. We deliberately do not headline the simpler "how many rain calls were right", because in a dry spell "no rain" is correct nearly every hour and every forecast scores about 99%, which tells you nothing. Scoring only the wet hours cannot be inflated by good weather. If a window contains no wet hours at all, we show a dash rather than a flattering number.
- Sky match: hour by hour, did the forecast's sky (fine / grey / wet) match what happened? This is the hardest call of the four: UK skies flip constantly and a category is either right or wrong, no partial credit. 70%+ is genuinely strong, and nobody on Earth gets 90% on this.
"Our blend" is the forecast the app actually shows you. It weighs the five sources by their proven track record at each measure and each time horizon, and re-weighs itself automatically as the scores evolve, so if one source goes off the boil, the blend leans away from it without anyone touching anything. The tick panel at the top tells you at a glance whether the blend is currently beating the individual forecasts.
What we use as ground truth
We pull observations from Open-Meteo's Historical Weather Archive, which blends ground-station readings with ECMWF's ERA5 reanalysis to give an hourly observation record for any point in the UK. We monitor a set of locations spread right across the UK, from the south coast to Edinburgh and Belfast, so the scores reflect performance nationally, not just one corner.
What we're comparing
These scores compare the forecast data each provider sends us through their API, which is what our blend is built from. Providers' own apps may add things their APIs don't send, so this is a comparison of the feeds rather than of finished apps.
What we don't do
We don't cherry-pick. We don't drop bad runs. We don't weight scores by what makes any one source look better. By default we show every day we've scored since tracking began; you can narrow the window with a ?days=N URL parameter.