↖ CPA Weather Lab

6  Chapter 6 — Datasets: The Full Archive Explained

The Datasets menu is where the IEM stops being a collection of tools and starts being a data warehouse. Everything in the previous five chapters — the networks, the current conditions displays, the Autoplot charts, the GIS services — ultimately draws from datasets that live here. This chapter maps those datasets: what each one actually contains, how far back it goes, how the data is structured, and when you should reach for one over another. Two of them — MOS and IEMRE — get extended treatment because they are genuinely complex and genuinely underexplained elsewhere.

The Datasets menu currently has twelve entries: Daily Climatology, Daily Observations, Dataset Documentation, IEM Reanalysis, Model Output Statistics, NEXRAD Mosaic, PIREP - Pilot Reports, Roads Mainpage, RADAR & Satellite, Rainfall Data, Sounding Archive, and Soil Moisture Satellite. NEXRAD Mosaic and RADAR & Satellite were covered in Chapters 5 and 6 respectively. Roads and PIREPs are specialized enough to note briefly and move on. That leaves the core analytical datasets, which are the heart of this chapter.


Daily Climatology: COOP Records and the Extremes Tool

mesonet.agron.iastate.edu/COOP/extremes.php

WarningData Quality Note — COOP Time of Observation Bias (TOBS)

COOP observers take readings at self-selected times — some at 7 AM, some at 5 PM. This creates a systematic bias: a 5 PM observer who records a 95°F peak on Tuesday will still see that peak on Wednesday’s max-min thermometer if a cold front moved in overnight. Wednesday gets logged as 95°F even though the actual high was 75°F. The 24-hour observation window double-counts the extreme.

If you are building consecutive-day heat wave analyses or comparing stations with different observation times, account for TOBS offsets before drawing conclusions. The full correction methodology is in Chapter 10.

The Daily Climatology entry in the Datasets menu points to the COOP extremes tool, which is the surface-level entry point to the IEM’s deepest historical climate dataset. The underlying data is the NWS Cooperative Observer network — the same observations covered in Chapter 2 — but through the Datasets lens, the question shifts from “how do I find a station” to “how do I work with a century of daily records.”

The COOP extremes tool at /COOP/extremes.php lets you select a state, a station, and a date, then browse the period-of-record statistics for that date: the all-time high and low temperatures, the highest and lowest precipitation totals, the record snowfall, and the years in which each record was set. For any station with a long record — Pennsylvania COOP stations routinely extend to the 1890s — this tool surfaces historical context that is otherwise buried in raw data files.

The companion download tool for normals is at mesonet.agron.iastate.edu/COOP/dl/normals.phtml, which outputs the 366-day climatological normals for any station in CSV format. These normals are computed by the IEM from the full period of record rather than the official 1991–2020 NCEI normals, which matters when you want a station-specific baseline that uses all available data rather than the most recent 30-year window. For stations with 80 or 100 years of record, the IEM’s period-of-record normals tend to be slightly more conservative (less influenced by recent warming) than the NCEI 1991–2020 normals — a meaningful difference when you are analyzing trend departures.

The Climodat application at mesonet.agron.iastate.edu/climodat/ builds on the same underlying COOP data and generates a series of pre-formatted climate reports — monthly summaries, annual rankings, first and last freeze dates, precipitation frequency statistics, and more — for any station in the network. Climodat is covered in Chapter 4 as an App, but it is worth noting here as the primary user interface to the COOP climatological dataset.


Daily Observations: The IEM Summary Dataset

mesonet.agron.iastate.edu/request/daily.phtml

The Daily Observations download portal is the IEM’s unified single-day-per-row summary product. Where the COOP download gives you raw cooperative observer reports and the ASOS download gives you hourly observations, this portal gives you a one-line-per-station-per-day summary computed by the IEM across multiple networks.

The variables available vary by network. For ASOS/AWOS stations, the daily summary includes high temperature, low temperature, average temperature, dew point high and low, peak wind gust and direction, average wind speed, precipitation total, snowfall (where reported), and sky condition. For COOP stations, the set is different: max/min temperature, precipitation, snowfall, snow depth, and observation time. The portal lets you select a network, a date range, and a set of stations, then download in CSV, Excel, or HTML.

The key distinction between this portal and the raw ASOS download at /request/download.phtml is aggregation level. The daily summary portal gives you IEM-computed daily statistics — a single row capturing the whole day. The ASOS download gives you individual hourly observations that you aggregate yourself. For most climatological analysis, the daily summary is faster and easier. For any analysis that requires sub-daily timing — hourly temperature evolution, precipitation start and end times, wind shift timing — go to the raw hourly download.


Dataset Documentation

mesonet.agron.iastate.edu/info/datasets/

This page is the IEM’s own reference documentation for its data products. It is written for developers and researchers rather than casual users, but it contains information that is genuinely hard to find elsewhere: the variable definitions, the units and encoding conventions, the processing pipeline for each dataset, and the known limitations. The Variables reference page at /info/variables.phtml is an essential companion — if you pull data from any IEM portal and see a column header you do not recognize, this is where you look it up.

Before building any serious analysis pipeline against IEM data, spend twenty minutes with the dataset documentation for the network you are using. The encoding choices are occasionally counterintuitive — wind direction in MOS, for example, is stored as the raw encoded value multiplied by 10 for the true degree value — and some variables have historical processing changes that affect how you interpret early versus recent data. The documentation catches these.


IEM Reanalysis (IEMRE)

mesonet.agron.iastate.edu/iemre/

IEMRE is one of those datasets where the IEM’s own description is so direct it is worth engaging with directly. The page says: “IEMRE is nothing special. It was cool maybe 15 years ago, but now days there are superior products available. Use those instead!” That is a remarkable thing for a data product to say about itself, and it tells you something important about both the product and the people behind it.

Here is what IEMRE actually is, why it exists despite that disclaimer, and when it is still the right choice.

What IEMRE Is

IEMRE is a gridded analysis product — meaning it takes point observations from across the IEM’s station networks and interpolates them onto a regular grid, producing spatially continuous fields of temperature, precipitation, dew point, wind speed, snowfall, soil temperature, and solar radiation. The current CONUS grid resolution is 0.125 × 0.125 degrees (roughly 8–14 km depending on latitude). The archive runs from 1950 to real-time, with hourly and daily temporal resolution. As of July 2025, non-CONUS domains have been added for South America, Europe, and Southeast Asia/China.

The honest version of why this exists: the IEM needed a gridded analysis product to drive agricultural modeling work — specifically the Daily Erosion Project, which requires spatially continuous precipitation and temperature fields to compute soil loss estimates across Iowa. Building IEMRE was solving that operational need. Once it existed, the IEM made it public.

The Data Flow

The intelligence in IEMRE is in how it sources each variable and handles the transition between real-time operational products and longer-term reanalysis products. For the CONUS domain, the precipitation chain works like this: hourly values are initially sourced from NCEP Stage IV (radar-estimated precipitation bias-corrected by gauges), then replaced with Stage IV adjusted by PRISM for dates more than eight days old, then replaced by ERA5-Land for dates before 1997. For temperature, the near-real-time source is RTMA (the NWS’s Real-Time Mesoscale Analysis) for post-2010 data, falling back to a cruder ASOS grid analysis for older data and ERA5-Land before 2010. Solar radiation as of January 2023 uses HRRR for the most recent eight days and ERA5-Land for anything older.

The practical implication of this layered sourcing is that IEMRE values for a given location will change retroactively as better analysis products become available. A precipitation estimate from last week may look different in six months once the Stage IV/PRISM adjustment cycle catches up. This is appropriate behavior — the system is improving its estimates with time — but it means you should not treat IEMRE values as fixed the way you would treat a station observation.

When to Use IEMRE

The disclaimer on the IEMRE page points to four alternatives: ERA5-Land, DayMet, GridMet, and PRISM. Those are all genuinely good products, and for research requiring defensible, peer-reviewed, stable gridded climate data, they are better choices. IEMRE’s advantages are narrow but real:

Real-time currency. IEMRE updates within about ten minutes of each hour. ERA5-Land has a five-day lag for near-real-time data, and DayMet and PRISM do not provide near-real-time at all. If you need gridded analysis for an event that happened yesterday or last week, IEMRE is the only free option.

IEM-native variables. IEMRE includes snowfall, snow depth, and soil temperature variables that are sourced from the IEM’s own network data. ERA5-Land does not include observed snowfall the way COOP observers report it.

Seamless API access. The IEMRE JSON API at /iemre/daily/{date}/{lat}/{lon}/json and /iemre/multiday/{date1}/{date2}/{lat}/{lon}/json is simple to call and returns clean JSON with no authentication required. If you want gridded climate data at a specific location and you are building something that needs to run programmatically, IEMRE’s API is considerably easier to work with than the NetCDF interfaces of ERA5-Land or PRISM.

Point sampling as spatial interpolation. The IEM uses IEMRE to backfill climate station values for the Climodat product — when a COOP station has missing data, the system samples from the IEMRE grid at that location as a fill value. This is documented and flagged in the output, but it means IEMRE and Climodat are tightly coupled in ways that matter for quality analysis.

Accessing IEMRE

The cleanest way to use IEMRE for a specific location is the JSON API. For a Pennsylvania example — say, the location of the Harrisburg airport at 40.19°N, 76.76°W — a daily query for 2024 looks like:

https://mesonet.agron.iastate.edu/iemre/multiday/2024-01-01/2024-12-31/40.19/-76.76/json

That returns a JSON object with daily values for all available IEMRE variables at the grid cell nearest to those coordinates. The output includes high_tmpf, low_tmpf, precip_in, snow_in, snowd_in, rstage4_precip_in (Stage IV precipitation), soil4t_f (4-inch soil temperature), and several others.

For raw gridded access, annual NetCDF files for the CONUS domain are available at mesonet.agron.iastate.edu/onsite/iemre/. These are the source files the IEM uses internally — xarray or any NetCDF-aware library can open them directly.


Model Output Statistics (MOS)

mesonet.agron.iastate.edu/mos/

MOS is one of the most useful datasets in the IEM archive for serious weather analysis, and also one of the most misunderstood. Most people who have heard of MOS think of it as “a forecast.” That is true but incomplete in ways that matter for how you use the archive.

What MOS Actually Is

Model Output Statistics is a statistical post-processing technique applied to numerical weather prediction model output. The NWS runs a deterministic forecast model — the GFS, the NAM — which produces gridded fields of atmospheric variables at regular intervals out to some forecast horizon. Those raw model fields have systematic biases: the model might run consistently too warm at certain stations, or too wet in complex terrain, or have poor skill at certain forecast lead times for certain variables. MOS corrects for those biases by developing statistical regression equations at each forecast station that relate the model’s large-scale output fields to the observed surface conditions at that station.

The equations are developed separately for each station, each variable, each season, and each forecast projection (the 6-hour, 12-hour, 24-hour, 48-hour forecast, and so on). The result is a station-specific bias correction that gives you a statistically calibrated point forecast at each location, derived from the model run but adjusted for that station’s historical relationship with the model.

The practical difference between a raw GFS temperature forecast and GFS-MOS at Harrisburg: the raw GFS is interpolating from a large-scale grid cell that may not capture the Susquehanna Valley’s cold-air pooling behavior in winter. The GFS-MOS equations learned from years of GFS runs and Harrisburg observations that the model needs to be adjusted downward by X degrees on calm clear nights in January, and the output reflects that. For temperature, precipitation type, and ceiling/visibility forecasts at ASOS locations, MOS has historically been competitive with or better than human forecasters at 12–24 hour lead times.

What the IEM Archives

The IEM’s MOS archive covers nine model configurations, with the following active archive dates confirmed as of March 2026:

Model Abbreviation Archive Start Status
AVN AVN 1 June 2000 Ended 16 Dec 2003
ETA ETA 24 Feb 2002 Ended 9 Dec 2008
GFS GFS 16 Dec 2003 Real-time
GFS LAMP LAV 12 Jul 2020 Real-time
GFS Extended MEX 12 Jul 2020 Real-time
NAM NAM 9 Dec 2008 Real-time
NBS NBS 7 Nov 2018 Real-time
NBE NBE 23 Jul 2020 Real-time

The GFS archive is the deepest operationally useful record, running continuously from December 2003 to present. The older AVN archive (predecessor to GFS) pushes the continuous record back to June 2000, meaning you have over two decades of consistent GFS-lineage MOS forecasts at every ASOS station in the country. That is a substantial verification dataset.

GFS vs. NAM vs. LAMP vs. NBS/NBE — what is each for?

The GFS MOS is the standard medium-range product, with projections from 6 hours out to 192 hours at 6-hour intervals for the first 72 hours and 12-hour intervals beyond. It is the workhorse.

The NAM MOS covers shorter projections — out to 60 hours — at the same 6-hour interval. NAM MOS tends to outperform GFS MOS at short ranges (6–24 hours) in complex terrain situations because the NAM runs at higher horizontal resolution over the CONUS and its mesoscale features are better resolved.

GFS LAMP (Local Analysis and Prediction System, abbreviated LAV in the IEM archive) is a very-short-range product, producing hourly projections out to 25 hours, updated every hour. It bridges the gap between the 6-hourly MOS guidance and the current observation. If you are doing hour-by-hour forecast verification or want the most frequently updated statistical guidance, LAMP is what you want.

NBS (National Blend of Models Statistics) and NBE (National Blend of Models Extended) are the successors to the GFS and MEX products respectively, blending multiple model inputs rather than relying on a single model. NBS runs from 2018 and represents the NWS’s current operational statistical guidance. It uses the same MOS framework but with a multi-model ensemble as input rather than a single deterministic run.

MEX (GFS Extended) covers projections beyond 72 hours, out to 192 hours (eight days), at 12-hour intervals. Useful for extended-range verification work.

Using the MOS Archive

The most intuitive interface is the Interactive MOS Tables at mesonet.agron.iastate.edu/mos/table.phtml. You enter a four-character station identifier (KMDT for Harrisburg, KIPT for Williamsport, etc.), select a model, select a variable, and choose a date range. The table shows you how the model’s forecast for that variable evolved across successive model runs for that period — each row is a different run time, each column is a different forecast projection. This is the interface to use when you want to see how a specific forecast evolved in time: did the GFS MOS temperature guidance for the day of a major winter storm flip dramatically 48 hours out? Did it nail the precipitation type, or was it wrong all the way through?

The raw ASCII MOS data is accessible through the text archive via the AFOS/AWIPS identifier system. Every MOS product is assigned a six-character identifier following the pattern <model><3-char station>. For GFS MOS at Harrisburg (MDT), that is GFSMDT. For NAM MOS at Harrisburg, NAMMDT. For LAMP at Harrisburg, LAVMDT. You pull these through the same text archive interface used for any NWS text product, covered in Chapter 7.

For programmatic access, the API endpoint at /api/1/mos.json accepts station, model, and runtime parameters and returns clean JSON. An example call pulling the most recent GFS MOS run for Harrisburg:

https://mesonet.agron.iastate.edu/api/1/mos.json?station=KMDT&model=GFS

The MOS data download form at /fe.phtml (linked from the MOS mainpage) lets you pull the raw MOS data for any station and time period in CSV format, with all variables for all forecast projections in a flat table. This is the format to use for bulk verification work.

Autoplot chart q=37 (MOS Forecasts vs. Actual Observations) gives you a visual monthly comparison of any MOS variable against the observed values at any ASOS station. For temperature verification, this is an excellent way to spot seasonal bias patterns — GFS MOS has historically run warm in winter at inland mid-Atlantic stations during cold-air damming events, for reasons that are directly legible in the chart when you look at January and February.

A Note on MOS Wind Direction Encoding

The MOS archive documentation flags one encoding issue worth knowing: wind direction (wdr in the database) is stored as the raw value from the text product, which is the true direction divided by 10. A wdr value of 27 means 270 degrees (west). This is documented in the MOS mainpage notes, but it is easy to miss if you are parsing the data without reading the documentation first.


Sounding Archive (RAOB)

mesonet.agron.iastate.edu/archive/raob/

The IEM’s sounding archive covers rawinsonde (RAOB) data for US and Canadian upper-air stations, twice-daily. The primary source is the Storm Prediction Center’s sounding archive, supplemented by additional backfill sources that extend some station records to the 1940s. The standard archive depth for most active stations runs from approximately the 1950s to real-time, with the twice-daily 00Z and 12Z launches from each upper-air station.

The download form at the archive URL accepts a station identifier (three or four characters — for Pittsburgh’s sounding site, that is PIT; for Wallops Island, WAL; for Albany, ALB), a start time, and an end time, and returns a CSV file with all available pressure levels. The columns are station, valid UTC time, pressure in millibars, height in meters, temperature in Celsius, dew point in Celsius, wind direction, wind speed in knots, and optional balloon bearing and range where available.

The JSON API endpoint at /json/raob.py offers the same data with additional query flexibility. You can pull a specific pressure level across a date range — useful for examining, say, 500 hPa temperature anomalies at Albany over a specific season — rather than downloading full sounding profiles and filtering yourself.

For mid-Atlantic upper-air work, the relevant stations are DIX (Upton/New York), IAD (Sterling/Washington DC), PIT (Pittsburgh), and WAL (Wallops Island, VA). The Pittsburgh and Wallops soundings together bracket central Pennsylvania, with Pittsburgh capturing the western approach flow and Wallops representing the coastal marine environment. The IAD sounding at Sterling is particularly useful for mid-Atlantic cold-air damming analysis — the 850 hPa temperature inversion structure at Sterling during CAD events has been studied and is directly readable in the IEM sounding archive.

Special soundings — launches outside the standard 00Z and 12Z schedule issued during severe weather events — are included in the archive where available.


Rainfall Data and MRMS

mesonet.agron.iastate.edu/rainfall/

The Rainfall Data entry in the Datasets menu points to the IEM’s GIS Rainfall portal, which provides access to gridded precipitation analysis products at hourly and daily time steps. The primary source is NCEP Stage IV, the NWS’s multi-sensor precipitation analysis that blends NEXRAD radar estimates with rain gauge observations and quality-controls the result.

Stage IV data is available from approximately 2002 onward. The IEM archives the hourly Stage IV grib files at /archive/data/YYYY/MM/DD/ under the stage4/ subdirectory, and provides derived daily and multi-day accumulation maps through the Autoplot interface (q=84 for precipitation maps, q=86 for daily gridded variable maps). For precipitation analysis during a specific event, the combination of Stage IV from the IEM and gauge observations from CoCoRaHS and ASOS gives you the most complete picture available from free sources.

The MRMS (Multi-Radar/Multi-Sensor) dataset is also archived. MRMS is the successor to Stage IV for real-time QPE — it runs at higher resolution (0.01 degree, roughly 1 km) and updates every 2 minutes rather than hourly. The IEM archives MRMS daily totals in NetCDF format at /onsite/mrms/, and the IEMRE system uses MRMS as a supplemental precipitation source. For event analysis where sub-hourly precision matters — a convective storm with sharp gradients — the MRMS archive is substantially more useful than Stage IV.


Soil Moisture Satellite

mesonet.agron.iastate.edu/smos/

The SMOS (Soil Moisture and Ocean Salinity) satellite dataset entry is a niche product for most weather enthusiasts — it provides passive microwave-derived soil moisture estimates from ESA’s SMOS satellite. The IEM archives and displays these for agricultural and hydrological research applications. For climate hobbyists, this is a dataset to know exists rather than one you will use regularly.


What to Watch Out For

The MOS archive has a backfill boundary. The IEM news post announcing MOS archiving in January 2009 noted that the archive was complete back to December 2008 at that point, with the possibility of further backfill. The GFS archive ultimately extends to December 2003. There is no MOS data at the IEM for anything before that, and the period between late 2003 and late 2008 may have occasional gaps depending on when different models entered the archive. For verification studies that need to go back before 2003, the NWS MDL maintains its own archives through a separate access pathway.

IEMRE values change retroactively. As noted above, the IEMRE pipeline replaces provisional estimates with better analysis products as they become available. If you download IEMRE data today for a date six months ago, the values may differ slightly from what you would have downloaded six months ago for that same date. This is by design but it matters for reproducibility — if you are building a research pipeline, either snapshot the data at a point in time or use a stable product like ERA5-Land where retroactive updates are versioned and documented.

RAOB data is not radar-verified. The sounding archive represents what the radiosonde actually measured as the balloon ascended. It is not quality-controlled against radar wind profilers or other independent measurements. Gross outliers do exist in the record — occasional instrument failures, balloon drift issues, or data transmission errors. For research applications involving the sounding archive, independent QC is good practice.

Daily summaries and COOP observation time. The IEM’s daily summary dataset uses a fixed 24-hour UTC period for its computations, which does not always match the observation time of the COOP reporter at that station. A COOP observer who takes their reading at 7 AM local time is reporting the 24-hour period ending at 7 AM, not the calendar day. This can create apparent discrepancies between the IEM daily summary and the raw COOP observation, particularly for precipitation and temperature extremes on days with overnight events. The COOP observation time (TOBS) is documented in the station metadata and is the key to resolving these discrepancies when they appear.



Workflow 1 — Looking Up Period-of-Record Daily Extremes for a COOP Station

The question: What is the all-time record high temperature for January 17th at your local COOP station? What year did it occur?

Start here: https://mesonet.agron.iastate.edu/COOP/extremes.php

The page presents a simple form: select a state and a date. The state dropdown defaults to Iowa — change it to Pennsylvania (or whichever state you are working in). Set the month and day to January 17. Submit.

What you get back is a table of every COOP station in that state with records for that date — station ID, station name, record high temperature, record low temperature, record precipitation, and the year each record was set. The table is sortable by any column, which is useful if you want to quickly scan for the most extreme values across the state rather than looking at a specific station.

To get the full annual record for a single station, click the station ID in the table. This takes you to a station-specific view showing every calendar day of the year with its record high, record low, and record precipitation, all drawn from the COOP archive for that station. For a station like Harrisburg with records going back to the late nineteenth century, this is a genuinely rich historical document — you can see the depth of the record and the years in which the extremes were set.

A few things to watch for: this tool covers the COOP network, not ASOS. The COOP record at a given location may extend back further than the ASOS record, but it is a once-daily observation rather than hourly. If you are looking for extreme hourly values — record hourly precipitation, record sustained wind — you will need Autoplot rather than this tool. Also note the “unofficial” label: these are IEM-computed records from the COOP archive, not the formally certified NWS climate records for the station.


Key URLs — Chapter 6

Page URL
Daily Climatology / COOP Extremes mesonet.agron.iastate.edu/COOP/extremes.php
COOP Normals Download mesonet.agron.iastate.edu/COOP/dl/normals.phtml
Daily Observations Download mesonet.agron.iastate.edu/request/daily.phtml
Dataset Documentation mesonet.agron.iastate.edu/info/datasets/
Variables Reference mesonet.agron.iastate.edu/info/variables.phtml
IEM Reanalysis (IEMRE) mesonet.agron.iastate.edu/iemre/
IEMRE CONUS NetCDF Files mesonet.agron.iastate.edu/onsite/iemre/
IEMRE Daily API mesonet.agron.iastate.edu/iemre/daily/{date}/{lat}/{lon}/json
IEMRE Multiday API mesonet.agron.iastate.edu/iemre/multiday/{date1}/{date2}/{lat}/{lon}/json
Model Output Statistics mesonet.agron.iastate.edu/mos/
MOS Interactive Table mesonet.agron.iastate.edu/mos/table.phtml
MOS API (JSON/CSV) mesonet.agron.iastate.edu/api/1/docs#/default/service_mos__fmt__get
MOS Raw Download mesonet.agron.iastate.edu/fe.phtml
MOS vs. Observations Chart (q=37) mesonet.agron.iastate.edu/plotting/auto/?q=37
Sounding Archive mesonet.agron.iastate.edu/archive/raob/
RAOB JSON API mesonet.agron.iastate.edu/json/raob.py?help=
Rainfall / Stage IV Portal mesonet.agron.iastate.edu/rainfall/
MRMS NetCDF Archive mesonet.agron.iastate.edu/onsite/mrms/
Stage IV Archive mesonet.agron.iastate.edu/onsite/stage4/
NEXRAD Mosaic Documentation mesonet.agron.iastate.edu/docs/nexrad_mosaic/
Soil Moisture Satellite (SMOS) mesonet.agron.iastate.edu/smos/
Climodat Reports mesonet.agron.iastate.edu/climodat/