XIONIA · Forecast verification

Historical weather model backtesting

Historical Open-Meteo forecasts for Agios Stefanos compared on common dates at 1-, 3- and 7-day lead times.

How to read the scoresAthens radiosonde reference

Read the comparison in three steps

Choose the same conditions

Compare the same variable, forecast lead time and period. Use a common sample where available.

Look at error and sample size

Lower error is better for the selected sample. A few dates or many forecasts of the same event do not establish a reliable winner.

Check what the result describes

Historical backtests, live tests and calibration answer different questions. A low score does not by itself activate a correction.

Open-Meteo · historical comparison, 2025–2026

Loading…

The same dates for every model
ModelDaysMAERMSEMean bias

A lower MAE means a smaller mean absolute error. Negative mean bias means the model predicted lower values than the station observed. The sample varies by variable and lead time, but all six models use the same dates within each table.

Forecasts come from the Open-Meteo Previous Runs archive, using a fixed lead time of 24, 72 or 168 hours for each forecast valid time. Daily maxima and minima are calculated from complete hourly samples. Precipitation is summed over the complete local calendar day in Athens, accounting for daylight saving time.

IFS 0.25° and IFS 9 km belong to the same model family and do not receive two independent votes in a blend. GFS here includes Open-Meteo processing; these results cannot be transferred directly to the six-hourly 0.5° fallback.

This period had already been examined in earlier work. This is a reproducible recheck, with no new API requests or automatic changes to model weights. Snow verification requires independent snowfall observations.

Comparison data (JSON) · Source documentation

What each metric means

MAE · average absolute error
The average distance between forecast and observation, ignoring the sign. A temperature MAE of 1°C means an average absolute error of 1°C; it is not a guarantee that every forecast is within ±1°C.
RMSE · more weight on large errors
Also measures forecast error, but penalises large misses more strongly. It uses the variable’s unit, such as °C, mm, km/h or hPa. Lower is better.
Bias · mean signed error
Forecast minus observation, averaged over the sample. Positive means higher forecasts; negative means lower forecasts. Opposite errors can cancel, so a bias near zero can coexist with a large MAE.
Sample · dates, times and forecasts
Days or distinct observation times describe the breadth of the sample. Multiple model runs can use the same observation. Counts are not interchangeable and missing data are not zero errors.
Common sample · a fairer comparison
Models are compared on matching cases within that table. “Available per model” can cover different weather and dates, so it does not establish a head-to-head ranking.
Lead time · how far ahead
+1 day is the next forecast day, not “this time tomorrow” in every report. Hourly live tests use the minimum time since forecast receipt; ensemble bands use hours since the model run. Read each selector’s wording.
CRPS · the ensemble distribution
Scores the ensemble’s range of possible outcomes against the observation. It uses the variable’s unit; lower is better. It is not an accuracy percentage.
Brier · event probability
The mean squared difference between the forecast event probability and whether the event happened. Lower is better; the score has no unit. This ensemble report uses temperature ≤0°C or precipitation ≥1 mm over six hours.
Weights and calibration
A weight is a model’s share of a blend, not its probability of being correct. A correction addresses a measured systematic error. Validation checks later data separately from training.
850 hPa · upper-air reference
A pressure level in the atmosphere, not a fixed elevation or a surface temperature. A radiosonde measures the atmospheric profile as a balloon ascends. The Hellinikon observation is a regional reference for Attica.