XIONIA · Forecast verification

Improving local forecasts with weather station data

From observations at Agios Stefanos and across the network to tested local corrections. See what is being collected, what is under evaluation and what has demonstrated an improvement on new forecasts.

Ελληνικά · English

Data collectionStation resultsWhen corrections are appliedUnderstanding the results

What this page shows

Record forecasts before the weather happens

We retain forecasts as issued and compare them later with checked observations. A count of recorded values is not an accuracy score.

Test corrections before applying them

A candidate correction is evaluated on later days that were not used to train it.

Require evidence of improvement

Enough distinct days and a consistent benefit are required. Zero active methods at the start means more data are needed.

Collection and status

Loading the saved report…

Observations
—

New values in the latest collection

Daily forecasts
—

New records in the latest capture

Forecasts at model times
—

New records in the latest capture

Methods under evaluation
—

Awaiting independent verification

Active methods
—

Have passed the required checks

The first three counts describe separate recent tasks. They are not archive totals or counts of distinct stations. One location can contribute several variables and forecast times.

Quality checks: why are some data held out?

We check observation times, sources and suitability. Some values are rejected; others are stored with flags and excluded from clean observation comparisons. The counts below do not represent distinct stations and cannot be added together to obtain an overall rejection rate.

  • Waiting for a quality report.

Up to two completed evaluation passes per day. Opening this page reads saved results; it does not start model training or download forecast model data.

Is collection continuing?

Waiting for an operational report.

A temporary lock skips one invocation without waiting. It does not stop the existing collector. A successful return can contain zero new records and does not necessarily mean a full evaluation pass finished.

Operational timestamps in Athens time · recorded after this update
TaskLatest successful returnLatest lock skipLatest outcome

How recent are the model runs?

Run age is measured from the model start time. The latest update attempt may have failed while a previous run remains available. This table describes data continuity, not forecast accuracy.

A complete step record describes saved metadata; it does not verify forecast field contents or the settings used by each forecast page. A failed latest attempt does not by itself remove the recorded active run. Times and waiting periods refer to the saved report.

Source-input counts describe the latest completed read per station within two hours, not a simultaneous network snapshot. Presence in these inputs does not prove that an optimization forecast was saved, matched or activated. GFS resolutions remain separate here; related model versions are not automatically independent evidence.

GFS 0.50°

Waiting for a report.

GFS 0.25°

Waiting for a report.

IFS

Waiting for a report.

AIFS

Waiting for a report.

AIGFS 0.25°

Waiting for a report.

GEM

Waiting for a report.

Progress by station

Waiting for a station report.

The list includes stations in the latest evaluation or with active forecast results. It is not a complete network catalogue.

Forecasts from active, verified methods

“Without new correction” is the reference forecast for the experiment. “With correction” is the value produced by the verified method; it is not an observed weather measurement.

Athens time · temperature in °C and daily rainfall in mm
Date / timeVariableWithout new correctionWith correction80% intervalProbability of ≥1 mm / day

There are no active forecast results yet.

Daily corrections are applied to the forecast when the method requirements are met. Results at available model times and uncertainty intervals are shown here; this does not mean they have been incorporated into every hourly chart on the site.

Terrain and distance from the sea

These indicators describe the station selected above. Terrain uses approximately 30 m data, excluding the sea and mapped permanent water. It describes the broader setting, rather than sensor exposure or proven microclimate similarity.

Waiting for a geographic report.

Waiting for station context.

Distance from coastline
—
Land relief within 2 km
—
Relative terrain position
—
Land in the 2 km area
—

A geographic classification hold affects only correction transfer from other stations. Data collection and independent local forecast training continue.

How should I read these indicators?

Distance refers to the sea, not lakes. Relief is the 90th–10th percentile difference of land surface heights within 2 km. Relative position compares median land surface-model height within 300 m with mean land height within 2 km. Positive means relatively higher ground; negative means lower ground. It does not measure a temperature inversion.

Classes and uncertainty margins are conservative rules for selecting candidate donor stations. Documented context does not activate a correction: elevation, separate sources, spatial exclusion and independent future-day tests still apply. A changed location or elevation invalidates the old documentation.

Sources: Copernicus GLO-30 / GLO-90 (surface models including buildings and vegetation), GSHHG 2.3.7 for coastlines and ESA WorldCover 2021 v200 to distinguish land and water. Processed on 30 September 2026. These datasets are not an on-site survey or a measurement of today's water levels.

produced using Copernicus WorldDEM-30 and WorldDEM-90 © DLR e.V. 2010-2014 and © Airbus Defence and Space GmbH 2014-2018 provided under COPERNICUS by the European Union and ESA; all rights reserved. GSHHG 2.3.7, Wessel & Smith (LGPL v3). © ESA WorldCover project 2021 / Contains modified Copernicus Sentinel data (2021) processed by ESA WorldCover consortium (CC BY 4.0).

What we improve and what is required

Local temperature

Small corrections by forecast lead time, using available information about time of day, wind and cloud cover. For example, we check whether certain conditions repeatedly produce a systematic error.

The next 1–6 hours

A recent forecast error may support a temporary correction that weakens over time. A recent observation and an available model time are required. Intermediate observations are not fabricated. Dew point is evaluated only when both the forecast and observation series exist.

Rain probability and amount

“How likely is at least 1 mm of rain during the day?” and “How much rain is expected?” are different questions. Probabilities, amounts, missed rain and false alarms are evaluated separately.

Transfer from comparable stations

This requires verified terrain and coastal context, comparable elevation and independent sources. Without those details, the geographic method remains on hold. The target station’s area is excluded from training.

Forecast uncertainty

An interval is shown only after independent checks of its coverage and width. Making an interval wider does not by itself make the method better.

Radiosondes

These verify the upper atmosphere separately from surface observations. The Elliniko reference does not automatically become a surface correction for Agios Stefanos. Inversions and the freezing level require reviewed full vertical profiles.

View verification against radiosondes →

When is a correction applied?

  1. Collection and training. At least 42 distinct days with suitable matched forecasts and observations are required.
  2. Separate initial validation. A 7-day gap is followed by at least 14 distinct validation days. The candidate must already improve the comparison.
  3. A new test with fixed coefficients. The method remains fixed for 35 days. At least 25 verified days and a consistent weekly benefit are required. Rain methods also need at least 10 wet days in the relevant samples.
  4. Activation and monitoring. Only methods that pass the checks are applied. If recent performance deteriorates sufficiently, the correction is withdrawn.

The process takes roughly three months of suitable new data, often longer. There is no guaranteed activation date or guaranteed improvement. Many records from one day do not replace distinct days.

Understanding the indicators

Collecting / under evaluation / active
Collecting: the suitable sample is still too small. Under evaluation: the method records new predictions without being applied yet. Active: it has passed independent verification.
Matched samples
Forecasts for which a suitable observation has been found. They are not necessarily distinct days, nor are they a count of stations.
MAE · mean absolute error
The average distance between forecast and observation. An MAE of 1°C means an average error of 1°C, not that every forecast is within ±1°C. Lower is better.
RMSE · more weight on large errors
Gives larger errors more influence and uses the same unit as the variable. Lower is better.
Brier score · probability quality
Checks forecast probabilities against whether an event occurred. Here the event is daily rainfall ≥1 mm. Lower is better; it is not an accuracy percentage.
80% interval
The aim is for around 8 out of 10 observed values to fall inside the interval across many new cases. This is not a guarantee for one particular day. “Under review” means the interval has not yet been validated.
Model run and age
A run is the start time of a simulation, not the file download time or the forecast’s valid time. Its reported age is measured when the report is generated.
Source wait / temporary lock
A source may not have published the new file yet. When tasks overlap, one may skip a capture to avoid conflicting with another. A single message does not prove a general outage; subsequent collection timestamps show whether work continues.