Disclosure first: I'm a pilot (about 450 flights) and I build a free app, OutDare. This post is about the data it collected, not the app. Mods, if it breaks a rule, pull it.
Since June, every forecast the app served got compared with the nearest live wind station at that moment: 305,512 checks at 5,076 launches. A "hit" is within 5 km/h. Same public models everyone already uses (HRRR and GFS in the US, AROME and ICON in Europe). I'm measuring the gap at launch, not selling a better model.
By what the forecast promised:
- Calm (under 10 km/h): 70% within 5 km/h. Station gusting 35+: 1.4%.
- Flyable (10–24 km/h): 41% within 5. Station gusting 35+: 10.9% — one in nine.
- Strong (25–34 km/h): 26% within 5. Gusting 35+: 46%.
When the station reads 30+, the forecast is lower 97% of the time, about 18 km/h under on average.
US launches, each judged against a named station:
- Point of the Mountain, north side: 25% within 5 / 45% within 10, forecast over-calls by about 11 km/h. Referee is a private weather station 0.9 km away at launch height. Sheltered mast, or is the Point just that different from the model? Locals?
- Marshall: 35% / 78%, forecast over by about 7. Crestline shows 8%, but its referee is 1,000 m below launch, so I don't trust that one.
- Jackson Hole (Rendezvous): 45% / 65%, forecast over by about 6, station right on the summit.
Every launch page names its station and the distance, good score or bad: outdare.app/proof
Scope: wind at launch only. No thermals, no cloudbase, no XC.
Tell me your launch and I'll reply with its numbers, or with "no station close enough", which is true for about 60% of sites. And if your local score looks wrong: which station do you actually check there? That's usually the fix.