Choose the same conditions
Compare the same variable, forecast lead time and period. Use a common sample where available.
XIONIA · Forecast verification
How temperature and precipitation corrections and each model’s contribution to the forecast blend are evaluated.
Compare the same variable, forecast lead time and period. Use a common sample where available.
Lower error is better for the selected sample. A few dates or many forecasts of the same event do not establish a reliable winner.
Historical backtests, live tests and calibration answer different questions. A low score does not by itself activate a correction.
Models receive different weights by variable, season and lead time when the change improves verification against the station. Maximum and minimum temperature are evaluated separately.
1 September 2025–31 August 2026 · 328 common dates per variable and lead time. MAE is the mean absolute error; lower is better. For each date, coefficients are fitted and checked only against observations available before the corresponding forecast.
| Lead time | Variable | Days | Previous blend | New method | MAE improvement |
|---|---|---|---|---|---|
| +1 day | Maximum °C | 328 | 0.705 | 0.705 | +0.00% |
| +1 day | Minimum °C | 328 | 0.617 | 0.617 | -0.01% |
| +1 day | Precipitation mm | 328 | 1.629 | 1.628 | +0.04% |
| +3 days | Maximum °C | 328 | 0.897 | 0.841 | +6.24% |
| +3 days | Minimum °C | 328 | 0.713 | 0.696 | +2.38% |
| +3 days | Precipitation mm | 328 | 2.056 | 2.056 | +0.00% |
| +7 days | Maximum °C | 328 | 1.587 | 1.541 | +2.92% |
| +7 days | Minimum °C | 328 | 1.147 | 1.141 | +0.57% |
| +7 days | Precipitation mm | 328 | 2.811 | 2.811 | +0.00% |
The largest improvement came from correctly matching the temperature correction to the three-day and seven-day lead times. Changes in model weights had a smaller effect. One-day minimum temperature is virtually unchanged, with a marginal increase in overall error. Not every season or forecast improves.
The period had already been examined in previous work. This is a retrospective recheck, not a new independent test. The tables do not guarantee the performance of future forecasts.
Calculated using data through 31 August 2026. The blend retains 75% of the existing baseline, while 25% follows constrained weights derived from each model’s error. The two IFS versions share the IFS family contribution in the learned component.
| Season | Lead time | Variable | Final contribution |
|---|---|---|---|
| Winter | +1 day | Minimum | GFS 27.3% · IFS 0.25° 23.0% · AIFS 23.0% · GEM 24.5% · IFS 9 km 2.1% |
| Winter | +1 day | Precipitation | GFS 24.8% · IFS 0.25° 21.3% · AIFS 24.8% · GEM 26.3% · IFS 9 km 2.9% |
| Winter | +3 days | Minimum | GFS 27.3% · IFS 0.25° 23.0% · AIFS 23.0% · GEM 24.5% · IFS 9 km 2.1% |
| Summer | +3 days | Minimum | GFS 23.5% · IFS 0.25° 21.1% · AIGFS 5.6% · AIFS 22.5% · GEM 24.0% · IFS 9 km 3.4% |
| Summer | +7 days | Minimum | GFS 23.4% · IFS 0.25° 22.1% · AIGFS 5.0% · AIFS 22.7% · GEM 23.8% · IFS 9 km 3.2% |
At least 45 common dates are required for fitting and 20 later dates for validation. The candidate blend must reduce validation MAE by at least 1%, without a material worsening of RMSE. Precipitation also requires at least ten wet days and rain detection that is no worse than the baseline.
For intermediate lead times, contributions change gradually between the tested horizons. Separate backtests have not been performed for every intermediate day. The existing blend applies when validation fails or a required model is missing. Learned weights are not applied to today, beyond seven days, to the direct NOAA fallback or to other locations.
When the server collector is active, it saves the original forecasts and the daily Xionia blend with their receipt time, version and weights. Personal browser settings do not change this shared record. Retraining requires a new evaluation and is not triggered automatically by a few successful forecasts.
Ongoing collection and verification → · Compare six models →