METAR & TAF 3D
Pilot guides
Live map

Winds Aloft and Multi-Model Forecasts

The METAR is one point at one time. The TAF is one aerodrome for a day. For everything else — the wind at 6,000 ft on your route, the freezing level tomorrow, whether the front arrives at 14:00 or 18:00 — you are relying on a numerical weather model, and there are half a dozen of them, run by different agencies, disagreeing with each other in ways that matter. This guide explains what they are, why they differ, and why the useful forecast is usually the one that takes all of them into account.

What a winds-aloft forecast is

Every global model computes the wind, temperature and humidity on a three-dimensional grid — typically 9 to 25 km horizontally, with 60 to 140 levels vertically — at every hour or three hours out to 10–16 days. A winds-aloft forecast is that grid sampled at your point and at aviation levels: 1,000 ft steps in the lowest layer, then 3,000, 6,000, 9,000, 12,000, 18,000 ft and the flight levels. The direction is true, the speed in knots, and the temperature in °C. Most pilots meet it as a table; on this site it is a panel on every airport page, drawn as barbs from the surface up to 13,000 ft in 1,000 ft steps so that the shear between levels, and the height at which the wind changes direction, is visible at a glance.

The models

ModelAgencyResolutionKnown for
IFSECMWF (Europe)9 km, 137 levelsBest medium-range skill on most verification metrics; the reference model
AIFSECMWF~25 kmMachine-learning model trained on IFS reanalysis; very fast, competitive at synoptic scale, weaker on fine detail
GFSNOAA (USA)13 km, 127 levelsFree and open; runs four times a day to 16 days; slightly behind IFS at 3–7 days
ICONDWD (Germany)13 km global, 6.5 km EuropeStrong over Europe, good on precipitation structure
GEMECCC (Canada)15 kmRobust global model, good over the Atlantic and Arctic
ARPEGEMétéo-FranceStretched grid, 7.5 km over FranceExcellent over western Europe and the Mediterranean
UMUK Met Office10 km globalConsistently among the top three for short range

They differ in resolution, in how they represent clouds and convection that are smaller than the grid, in the observations they assimilate, and in the time their runs start. Two models given the same atmosphere produce two forecasts, and the difference grows with lead time.

Why one model is not enough

Suppose IFS says the front reaches LFPG Paris at 15:00 with a wind shift to 300°/25 kt, and GFS says 18:00. If you plan on IFS alone you are certain of the wrong thing three hours out of six. If you look at both, you know the shift is coming this afternoon and that the timing is uncertain by three hours — which is the true state of knowledge, and the right basis for a fuel and alternate decision. The disagreement, called the spread, is itself a forecast: when seven models agree, the atmosphere is in a predictable regime and you can trust the number; when they scatter, the honest forecast is "uncertain", and the operational answer is margin.

Consensus and calibration

A simple average of models already beats most individual models over a season. It can be improved in two ways. First, by weighting: down-weight the outliers at each hour so that one model's spurious 40 kt gust does not drag the mean. Second, by calibration against reality: compare each model's forecast for the last few hours with what the METARs at the nearest stations actually reported, and trust the models that were right. The Kalynda blend shown on this site does both — a robust weighted mean across the models (a Tukey biweight, for readers who want the statistics) with weights that are then scaled by each model's recent skill at that location — and the result is verified continuously: its error against subsequent observations is measured per lead time and published on the model page. When you see the model comparison panel on an airport page, the per-model winds are the raw ingredients and the Kalynda line is the blend.

The 14-day limit

Models produce numbers out to 15 or 16 days. Beyond about day 10 the skill of a single deterministic run drops toward climatology; by day 14 only ECMWF retains meaningful information at the synoptic scale, and nothing does for a single point. This site stops at 14 days deliberately, and at the far end shows the consensus of the three models that still carry signal rather than the tail of one. A forecast presented as precise beyond its skill is worse than no forecast, because it will be believed.

Using the panel

  1. Read the surface METAR first; that is truth.
  2. Compare it with the lowest model level. If they disagree by 10 kt or 40°, the models have the boundary layer wrong today; trust them less at low level for the next few hours, and expect the blend to have already down-weighted the worst of them.
  3. Read the shear between 1,000 and 3,000 ft — that is the layer you climb and descend through.
  4. Look at the spread. Tight agreement, plan on the number; wide disagreement, plan on the range.
  5. For the forecast beyond the TAF, use the meteogram's hour-by-hour consensus, and revisit it at each new model cycle (00, 06, 12, 18 UTC, available about four to six hours later).

Frequently asked questions

Which weather model is the most accurate?

For medium-range synoptic forecasts the ECMWF IFS has led verification scores for most of the last decade, with the UK Met Office and GFS close behind at short range. For any single point and hour, though, the ranking changes from day to day — which is the argument for looking at several models rather than trusting one.

What is model spread?

The disagreement between models (or ensemble members) for the same time and place. Small spread means the atmosphere is in a predictable state and the forecast can be trusted; large spread means the outcome is genuinely uncertain, whatever any single model says. Spread is a forecast of forecast quality.

Why do forecasts beyond about 10 days stop being useful?

Because small errors in the initial state grow exponentially, and by 10–14 days they are as large as the difference between any two random days. Beyond that the models still produce numbers, but they carry little more information than the climatology. This site stops at 14 days for that reason.

Are winds aloft forecasts in true or magnetic degrees?

True, like every meteorological wind. Only tower, ATIS and runway numbers are magnetic. Temperatures aloft are in °C and heights are pressure altitudes or geometric heights depending on the product.

What does the Kalynda blend do?

It combines up to seven global models with weights that change hour by hour according to how well each model has been matching the live METAR observations at the nearest stations. A model that is currently wrong about the surface wind at a field is trusted less at that field for the coming hours; one that is currently right, more.

Read next