Forecasts are saved before the event. We compare the most recent available forecast received at least 24, 72 or 168 hours earlier. The station observation must fall within five minutes of the forecast valid time.
Rows may cover different dates or time steps, so they do not rank models using a common sample. The direct GFS fallback uses six-hour steps, while Open-Meteo provides an hourly series. The two GFS sources and the two IFS versions are not independent models when forming a blend.
Recent forecasts and observations
Athens local time
Model / provider
Forecast
Observation
Error
What we retain for future checks
Large weather files are refreshed and cleaned up. Station forecasts and verification history are retained separately. The daily blend for the recommended profile is also stored, together with its version and applied weights. The scores above concern the original hourly forecasts; the daily blend is retained for a separate backtest against complete daily observations. A few successful forecasts do not trigger automatic retraining.
Snow verification requires independent observations of snowfall, precipitation type and snow accumulation in centimetres. Snow calculator scenarios are not observations.
What each metric means
MAE · average absolute error
The average distance between forecast and observation, ignoring the sign. A temperature MAE of 1°C means an average absolute error of 1°C; it is not a guarantee that every forecast is within ±1°C.
RMSE · more weight on large errors
Also measures forecast error, but penalises large misses more strongly. It uses the variable’s unit, such as °C, mm, km/h or hPa. Lower is better.
Bias · mean signed error
Forecast minus observation, averaged over the sample. Positive means higher forecasts; negative means lower forecasts. Opposite errors can cancel, so a bias near zero can coexist with a large MAE.
Sample · dates, times and forecasts
Days or distinct observation times describe the breadth of the sample. Multiple model runs can use the same observation. Counts are not interchangeable and missing data are not zero errors.
Common sample · a fairer comparison
Models are compared on matching cases within that table. “Available per model” can cover different weather and dates, so it does not establish a head-to-head ranking.
Lead time · how far ahead
+1 day is the next forecast day, not “this time tomorrow” in every report. Hourly live tests use the minimum time since forecast receipt; ensemble bands use hours since the model run. Read each selector’s wording.
CRPS · the ensemble distribution
Scores the ensemble’s range of possible outcomes against the observation. It uses the variable’s unit; lower is better. It is not an accuracy percentage.
Brier · event probability
The mean squared difference between the forecast event probability and whether the event happened. Lower is better; the score has no unit. This ensemble report uses temperature ≤0°C or precipitation ≥1 mm over six hours.
Weights and calibration
A weight is a model’s share of a blend, not its probability of being correct. A correction addresses a measured systematic error. Validation checks later data separately from training.
850 hPa · upper-air reference
A pressure level in the atmosphere, not a fixed elevation or a surface temperature. A radiosonde measures the atmospheric profile as a balloon ascends. The Hellinikon observation is a regional reference for Attica.