XIONIA · Forecast verification

Weather models compared with station observations

Live testing at Agios Stefanos: forecasts saved before the event are compared with actual weather station observations.

How to read the scoresAthens radiosonde reference

Read the comparison in three steps

Choose the same conditions

Compare the same variable, forecast lead time and period. Use a common sample where available.

Look at error and sample size

Lower error is better for the selected sample. A few dates or many forecasts of the same event do not establish a reliable winner.

Check what the result describes

Historical backtests, live tests and calibration answer different questions. A low score does not by itself activate a correction.

Hourly live test · original forecasts

Loading verification history…

Agios Stefanos bias correction and model weights v4 →

Historical forecast results · 2025–2026 →

Error by model — lower MAE is better
Model / providerComparisonsDaysMAERMSEMean bias

Forecasts are saved before the event. We compare the most recent available forecast received at least 24, 72 or 168 hours earlier. The station observation must fall within five minutes of the forecast valid time.

Rows may cover different dates or time steps, so they do not rank models using a common sample. The direct GFS fallback uses six-hour steps, while Open-Meteo provides an hourly series. The two GFS sources and the two IFS versions are not independent models when forming a blend.

Recent forecasts and observations

Athens local timeModel / providerForecastObservationError

What we retain for future checks

Large weather files are refreshed and cleaned up. Station forecasts and verification history are retained separately. The daily blend for the recommended profile is also stored, together with its version and applied weights. The scores above concern the original hourly forecasts; the daily blend is retained for a separate backtest against complete daily observations. A few successful forecasts do not trigger automatic retraining.

Snow verification requires independent observations of snowfall, precipitation type and snow accumulation in centimetres. Snow calculator scenarios are not observations.

What each metric means

MAE · average absolute error
The average distance between forecast and observation, ignoring the sign. A temperature MAE of 1°C means an average absolute error of 1°C; it is not a guarantee that every forecast is within ±1°C.
RMSE · more weight on large errors
Also measures forecast error, but penalises large misses more strongly. It uses the variable’s unit, such as °C, mm, km/h or hPa. Lower is better.
Bias · mean signed error
Forecast minus observation, averaged over the sample. Positive means higher forecasts; negative means lower forecasts. Opposite errors can cancel, so a bias near zero can coexist with a large MAE.
Sample · dates, times and forecasts
Days or distinct observation times describe the breadth of the sample. Multiple model runs can use the same observation. Counts are not interchangeable and missing data are not zero errors.
Common sample · a fairer comparison
Models are compared on matching cases within that table. “Available per model” can cover different weather and dates, so it does not establish a head-to-head ranking.
Lead time · how far ahead
+1 day is the next forecast day, not “this time tomorrow” in every report. Hourly live tests use the minimum time since forecast receipt; ensemble bands use hours since the model run. Read each selector’s wording.
CRPS · the ensemble distribution
Scores the ensemble’s range of possible outcomes against the observation. It uses the variable’s unit; lower is better. It is not an accuracy percentage.
Brier · event probability
The mean squared difference between the forecast event probability and whether the event happened. Lower is better; the score has no unit. This ensemble report uses temperature ≤0°C or precipitation ≥1 mm over six hours.
Weights and calibration
A weight is a model’s share of a blend, not its probability of being correct. A correction addresses a measured systematic error. Validation checks later data separately from training.
850 hPa · upper-air reference
A pressure level in the atmosphere, not a fixed elevation or a surface temperature. A radiosonde measures the atmospheric profile as a balloon ascends. The Hellinikon observation is a regional reference for Attica.