METAR & TAF 3D
METAR & TAF 3D
Aviation weather · meteogram by Kalynda
Kalynda

The Kalynda forecast model

Kalynda is not another numerical weather model. It takes the six that already exist and keeps what each one does best — ECMWF's medium-range accuracy, ICON's fine-scale detail over Europe, GFS's fast refresh, GEM's high-latitude behaviour, Météo-France's mesoscale nest, the Met Office's boundary-layer and coastal skill — then calibrates the result against the live METAR next to you. The single forecast that comes out is built to beat any one of them taken alone.

Why combine at all

Every global model solves the same physics with different grids, different starting analyses and different parameterisations. On a calm day they agree within a knot. On a changeable day they can disagree by 90° of wind direction and 10 knots of speed. Picking one model means inheriting its bad days. Averaging all of them is better — but a plain average lets a single badly-placed run drag the answer with it.

Step one — robust weighting

For each hour we take the six wind vectors and find their median. Each model is then weighted by how far it sits from that median, using a Tukey biweight: models near the middle keep full weight, models further out lose weight rapidly, and an extreme outlier is reduced to a token 5% rather than being deleted outright. Nothing is discarded — a lone model can still be right — but no single run can take the blend hostage.

Step two — live calibration

Robustness alone treats every model as equally credible. It is not. On a given day, over a given piece of terrain, some models simply track better. So we compare each model's current-hour forecast against the nearest reporting METAR station and score its error. Models matching the observation gain weight; models missing it lose weight, down to a floor of 35%. This is the same logic as classical Model Output Statistics, applied to a single live observation instead of a long archive — which is what makes it usable immediately.

Where it stops

Calibration only runs when the observing station is within about 55 km and the wind is above 4 knots. Beyond that distance local terrain dominates and the observation stops being evidence about the models. Below that speed, direction is unstable by nature and scoring models against it adds noise rather than skill. In both cases the blend falls back to robust weighting alone.

The other variables

Temperature, dew point, cloud, precipitation, visibility and pressure are blended the same way, but as scalars rather than vectors: median, biweight weights, weighted mean. Wind is handled as a vector because direction is circular — averaging 350° and 10° arithmetically gives 180°, which is precisely backwards.

What the confidence figure means

The percentage next to each forecast combines two things: how closely the models agree on direction, and how far apart their speeds are. High confidence means the models are telling the same story and you can plan against the forecast. Low confidence does not mean the forecast is wrong — it means the atmosphere is at a decision point and you should recheck closer to departure.

What we take from each model

Each centre has something it measurably does better, and the blend is weighted to use exactly that. ECMWF leads the medium range and has close to zero net wind bias, so it carries the most weight from day 3 onward and on pressure and dewpoint. DWD ICON has the lowest short-range errors against surface observations and, inside its European radar domain, assimilates radar every cycle — so it leads the first two days there, especially for precipitation and gusts. Météo-France AROME and the Met Office UKV run at 1.3 and 1.5 km with dedicated fog and low-cloud fields, so they carry visibility and cloud over France and the British Isles. NOAA's HRRR ingests radar every 15 minutes over the United States, so it leads precipitation and gusts there in the first hours. Environment Canada's GEM became the first operational hybrid physics-plus-machine-learning global system in May 2026, with gains peaking around day 5 and holding to day 10 — so its weight rises exactly in that window. Weights also change with the variable: a model can lead on wind and still be held back on visibility.

What we deliberately hold back

Published verification also tells you where not to trust a model, and we encode that too. ECMWF's own user guide says expectations for its visibility field should remain low and that it underestimates convective gusts, sometimes by a factor of two — so it carries little weight on both. GFS runs a documented dry bias in surface dewpoint over the United States and over-forecasts light rain while under-forecasting heavy rain, so its precipitation weight is reduced outside the HRRR domain. ARPEGE's grid stretches from 5 km over France to about 24 km at the antipodes, so its weight drops sharply away from Europe. UKV data arrives later than the rest, so it takes a freshness discount in the first hours. GEM carries a positive surface-pressure bias and under-predicts cloud. None of these models is dismissed — each is simply asked the questions it answers well.

Honest limits

Kalynda inherits everything its inputs get wrong. If all six models miss a sea breeze or a mountain wave, so does the blend. It has no radar assimilation, no nowcasting and no local terrain model. It is a better starting point than any single model, not a substitute for looking at the sky, reading the TAF, and making your own decision.

Model sources

ECMWF IFS · ECMWF AIFS · NOAA GFS · DWD ICON · CMC GEM · Météo-France ARPEGE · UK Met Office, all via Open-Meteo. Observations from NOAA Aviation Weather Center. Weights are grounded in published verification: ECMWF TM931/TM928 and the Forecast User Guide, DWD radar-assimilation documentation, Met Office RAL1/RAL3 (GMD), CNRM ARPEGE/AROME papers, ECCC GDPS/RDPS/HRDPS technical notes and NOAA NCEP GFS/HRRR documentation.

Open the live map