XIONIA · Forecast verification

Forecast calibration and model weights

How temperature and precipitation corrections and each model’s contribution to the forecast blend are evaluated.

How to read the scoresAthens radiosonde reference

Read the comparison in three steps

Choose the same conditions

Compare the same variable, forecast lead time and period. Use a common sample where available.

Look at error and sample size

Lower error is better for the selected sample. A few dates or many forecasts of the same event do not establish a reliable winner.

Check what the result describes

Historical backtests, live tests and calibration answer different questions. A low score does not by itself activate a correction.

Models receive different weights by variable, season and lead time when the change improves verification against the station. Maximum and minimum temperature are evaluated separately.

Open-Meteo · before and after in the historical test

1 September 2025–31 August 2026 · 328 common dates per variable and lead time. MAE is the mean absolute error; lower is better. For each date, coefficients are fitted and checked only against observations available before the corresponding forecast.

Lead timeVariableDaysPrevious blendNew methodMAE improvement
+1 dayMaximum °C3280.7050.705+0.00%
+1 dayMinimum °C3280.6170.617-0.01%
+1 dayPrecipitation mm3281.6291.628+0.04%
+3 daysMaximum °C3280.8970.841+6.24%
+3 daysMinimum °C3280.7130.696+2.38%
+3 daysPrecipitation mm3282.0562.056+0.00%
+7 daysMaximum °C3281.5871.541+2.92%
+7 daysMinimum °C3281.1471.141+0.57%
+7 daysPrecipitation mm3282.8112.811+0.00%

The largest improvement came from correctly matching the temperature correction to the three-day and seven-day lead times. Changes in model weights had a smaller effect. One-day minimum temperature is virtually unchanged, with a marginal increase in overall error. Not every season or forecast improves.

The period had already been examined in previous work. This is a retrospective recheck, not a new independent test. The tables do not guarantee the performance of future forecasts.

Archived Open-Meteo weights · data through 31 August 2026

Calculated using data through 31 August 2026. The blend retains 75% of the existing baseline, while 25% follows constrained weights derived from each model’s error. The two IFS versions share the IFS family contribution in the learned component.

SeasonLead timeVariableFinal contribution
Winter+1 dayMinimumGFS 27.3% · IFS 0.25° 23.0% · AIFS 23.0% · GEM 24.5% · IFS 9 km 2.1%
Winter+1 dayPrecipitationGFS 24.8% · IFS 0.25° 21.3% · AIFS 24.8% · GEM 26.3% · IFS 9 km 2.9%
Winter+3 daysMinimumGFS 27.3% · IFS 0.25° 23.0% · AIFS 23.0% · GEM 24.5% · IFS 9 km 2.1%
Summer+3 daysMinimumGFS 23.5% · IFS 0.25° 21.1% · AIGFS 5.6% · AIFS 22.5% · GEM 24.0% · IFS 9 km 3.4%
Summer+7 daysMinimumGFS 23.4% · IFS 0.25° 22.1% · AIGFS 5.0% · AIFS 22.7% · GEM 23.8% · IFS 9 km 3.2%

At least 45 common dates are required for fitting and 20 later dates for validation. The candidate blend must reduce validation MAE by at least 1%, without a material worsening of RMSE. Precipitation also requires at least ten wet days and rain detection that is no worse than the baseline.

For intermediate lead times, contributions change gradually between the tested horizons. Separate backtests have not been performed for every intermediate day. The existing blend applies when validation fails or a required model is missing. Learned weights are not applied to today, beyond seven days, to the direct NOAA fallback or to other locations.

Records for the next evaluation

When the server collector is active, it saves the original forecasts and the daily Xionia blend with their receipt time, version and weights. Personal browser settings do not change this shared record. Retraining requires a new evaluation and is not triggered automatically by a few successful forecasts.

Ongoing collection and verification → · Compare six models →

What each metric means

MAE · average absolute error
The average distance between forecast and observation, ignoring the sign. A temperature MAE of 1°C means an average absolute error of 1°C; it is not a guarantee that every forecast is within ±1°C.
RMSE · more weight on large errors
Also measures forecast error, but penalises large misses more strongly. It uses the variable’s unit, such as °C, mm, km/h or hPa. Lower is better.
Bias · mean signed error
Forecast minus observation, averaged over the sample. Positive means higher forecasts; negative means lower forecasts. Opposite errors can cancel, so a bias near zero can coexist with a large MAE.
Sample · dates, times and forecasts
Days or distinct observation times describe the breadth of the sample. Multiple model runs can use the same observation. Counts are not interchangeable and missing data are not zero errors.
Common sample · a fairer comparison
Models are compared on matching cases within that table. “Available per model” can cover different weather and dates, so it does not establish a head-to-head ranking.
Lead time · how far ahead
+1 day is the next forecast day, not “this time tomorrow” in every report. Hourly live tests use the minimum time since forecast receipt; ensemble bands use hours since the model run. Read each selector’s wording.
CRPS · the ensemble distribution
Scores the ensemble’s range of possible outcomes against the observation. It uses the variable’s unit; lower is better. It is not an accuracy percentage.
Brier · event probability
The mean squared difference between the forecast event probability and whether the event happened. Lower is better; the score has no unit. This ensemble report uses temperature ≤0°C or precipitation ≥1 mm over six hours.
Weights and calibration
A weight is a model’s share of a blend, not its probability of being correct. A correction addresses a measured systematic error. Validation checks later data separately from training.
850 hPa · upper-air reference
A pressure level in the atmosphere, not a fixed elevation or a surface temperature. A radiosonde measures the atmospheric profile as a balloon ascends. The Hellinikon observation is a regional reference for Attica.