🌐 GEFS Ensemble

NOAA's 31-member Global Ensemble Forecast System · init Fri 8 AM ET · 8 days at 0.5° · mean | spread, rendered locally from NOAA's AWS open data
--
MSLP + 500 mb heights
GEFS mslp
Temperature (2 m)
GEFS t2m
PWAT (moisture)
GEFS pwat
Jet stream (250 mb)
GEFS jet
CAPE (instability)
GEFS cape
Winter precip type
GEFS ptype
All six products advance together through the shared lead times - one clock, six maps: synoptic pattern, temperature, moisture, jet, instability, winter precip. Thin annotations follow each map.

📆 Forecast verification lead-time picker · all leads vs the same observed hour · valid Fri 8 AM ET

MSLP + 500 mb heights · F072 from init AM ET · valid AM ET
RMS 2.03hPa · r 0.978 · ACC +0.98 · S/S 0.98 (r+0.36) · bias +0.65hPa
GEFS mslp F072 forecastmslp observed analysis
Temperature (2 m) · F072 from init AM ET · valid AM ET
RMS 2.01°F · r 0.996 · ACC +0.92 · S/S 0.80 (r+0.48) · bias -0.19°F
GEFS t2m F072 forecastt2m observed analysis
CAPE (instability) · F072 from init AM ET · valid AM ET
RMS 151.16J/kg · r 0.945 · ACC +0.82 · S/S 1.08 (r+0.71) · bias -1.85J/kg
GEFS cape F072 forecastcape observed analysis
PWAT (moisture) · F072 from init AM ET · valid AM ET
RMS 0.14in · r 0.98 · ACC +0.91 · S/S 1.08 (r+0.52) · bias -0.01in
GEFS pwat F072 forecastpwat observed analysis

Skill vs lead time - last 90 days (r & ACC)

Pattern correlation by valid day and leadAnomaly correlation by valid day and leadDay 3: r 0.83–1.00 over 90 daysDay 5: r 0.68–0.99 over 90 daysDay 7: r 0.50–0.99 over 90 daysDay 10: r 0.04–0.99 over 90 days
Pattern correlation r vs the GFS analysis, one point per valid day, one color per lead - the gap between the lines is the skill lost with lead time; the slope is day-to-day pattern variability. Scores are recomputed nightly from the archive and kept 90 days. ACC (second chart, red 0-line) is the anomaly correlation - fcst and obs vs the same-day climatology (mean of the last 5 years of GFS analyses): positive means the forecast beat climatology, the standard skill test, and points below the line are busts. S/S is the ensemble spread-skill ratio (spread / actual RMSE): ~1.0 means the 31 members know when they are uncertain, <1 means overconfident, and the parenthesised r links spread to where the error actually landed.
Each pair sets a lead's ensemble-mean forecast beside the observed analysis for the SAME valid hour - pick a lead to see how much skill the forecast loses with range (the grid frames above only survive 48 h, so these are re-fetches from NOAA's archive). Same levels and palette as the grid above, so the differences you see are the forecast's, not the map's. Truth: GFS 0.25-deg analysis (noaa-gfs-bdp-pds, no keys).
New to forecast verification? Start here

Every forecast is easy to check in hindsight - this card does it with numbers. The idea: take a real forecast from 3, 5, 7, or 10 days out (kept in NOAA's archive), wait for the weather to happen, then compare the two maps and score them. A score near the top means the forecast put the right weather in the right place.

ScorePlain-English question it answersGood
RMSOn average, how far off was the number? (hPa, degrees, ...)lower is better
rDo the map's highs and lows line up with what happened?1 = perfect
ACCDid it beat just predicting "normal weather for the date"? (vs a 5-year climatology)above 0; below 0 = worse than climatology
S/SDid the 31 members disagree as much as the forecast turned out to be wrong? (their spread vs the actual error)near 1.0; <1 = overconfident
biasDid it systematically run too warm/wet (+) or too cool/dry (−)?near 0

Rules of thumb from our own 77-day record (MSLP): day-3 forecasts are near-perfect (ACC 0.96); day-5 are good (0.88); day-7 are useful but rougher (0.72); day-10 average 0.55 - and in our record they dipped below 0 once, meaning on that day flipping a coin with climatology would have done better. Temperature holds up much better than rain-based fields at long range: pattern errors grow fastest in CAPE and PWAT, so trust a day-10 temperature outlook more than a day-10 storm forecast.

How to read a pair of maps: start with the forecast (left/top image), form a picture - where is the low, where is the warm tongue - then look at the observed map and ask what moved. Then check the badge: high ACC but visible differences means the big picture was right and details moved; ACC near 0 means the whole pattern was wrong. The S/S number tells you whether the ensemble was honestly uncertain about it.

Both chart panels show these scores for every valid day of the last ~3 months, one colored line per lead: the vertical gap between lines is skill lost with lead time; the red 0-line on the ACC chart is the bust line. Data: NOAA GEFS + GFS, no keys. Scores are computed nightly by this site from the public archive.

📅 Day high / low TMAX | TMIN ensemble means · East TN town table · init Fri 8 AM ET

--
GEFS day high/low

🧭 Deep dives

Full-screen animated panels with ensemble spread and (on the jet) 16-member axis spaghetti live on the Models page GEFS card. Storm tracks from these same members are on the NHC Tropical page.