↖ CPA Weather Lab

10  Chapter 10 — Understanding Data Quality

Every piece of data on the IEM carries an implicit question: how much should you trust it? The answer is not uniform. It depends on the network, the variable, the time period, and sometimes a specific quirk in how a particular station’s data has been ingested, processed, or stored. This chapter is about developing the instincts to ask that question productively — and knowing where on IEM to look for the answer.

This is not a catalog of everything that can go wrong. That would be both exhausting and discouraging. It is a practical orientation to the most consequential data quality issues you are likely to encounter as a serious weather enthusiast working with IEM data, with enough explanation of the underlying mechanics that you can recognize problems when you see them rather than waiting to be surprised.


The “Unofficial” Label

The word unofficial appears throughout the IEM — in chart titles, in download headers, in the daily features on the homepage. It is worth understanding precisely what it means and, equally importantly, what it does not mean.

IEM is not a National Weather Service office. It is not NCEI. It is an academic mesonet operated out of Iowa State University, and the data it archives, processes, and displays has not gone through the formal QC and certification pipeline that transforms raw observations into the official climate record. When IEM applies its own computed products — daily summaries, areal averages, derived statistics — those products are labeled unofficial because they are: they reflect IEM’s best processing of available data, not a federally certified observation.

What unofficial does not mean is wrong, or unreliable, or experimental. The raw observations feeding IEM’s archive are the same observations flowing to NCEI and NWS. The ASOS data ingested by IEM comes from the same NOAAPort satellite feed and NCEI data sources that official archives draw from. The processing IEM applies is documented, the methods are published in journal papers, and the results have been used in peer-reviewed research. Unofficial is a legal and institutional designation, not a quality judgment.

The practical implication is simple: if you are doing research that will be cited, published, or submitted to a formal climate record, cross-check your IEM-derived results against NCEI’s official products. For forum discussions, personal research, forecast verification, and operational awareness — which covers the vast majority of what weather enthusiasts do with this data — the unofficial label should not give you pause.


How IEM Handles Quality Control

IEM runs automated QC checks on incoming data from its various networks. The specific checks vary by network, but the general approach includes range testing (is this temperature physically plausible?), rate-of-change testing (did this sensor just jump 30 degrees in five minutes?), and spatial consistency testing (does this station’s reading agree with its neighbors within reasonable bounds?).

Beyond automated checks, IEM participates in the MADIS (Meteorological Assimilation Data Ingest System) QC framework for several networks. The MADIS QC page at https://mesonet.agron.iastate.edu/QC/madis/network.phtml lets you browse QC status by network, showing you which stations are currently flagged and for what reason.

The QC mainpage at https://mesonet.agron.iastate.edu/QC/ provides links to all the QC products IEM generates. The most immediately useful for a working weather enthusiast is the offline station list at https://mesonet.agron.iastate.edu/QC/offline.php, which shows stations that are currently not reporting or have been flagged as having data problems.

The Public QC Ticket System

IEM maintains a public trouble ticket log at https://mesonet.agron.iastate.edu/QC/tickets.phtml. This is worth knowing about for two reasons. First, it is a window into the kinds of problems the IEM team actually encounters and how they handle them — which gives you a more realistic picture of data reliability than any documentation page alone. Second, it is genuinely bidirectional: the IEM team reads submissions, and if you identify what looks like a data quality problem while working with a specific station, you can flag it there.

The tickets range from straightforward sensor outages to more complex ingestion problems — cases where data was received but processed incorrectly, or where a network partner changed their data format and the IEM ingestor had not yet caught up. Reading through a few months of tickets gives you a healthy sense of the kinds of edge cases that exist even in well-maintained systems.


ASOS Temperature Storage: The Whole-Degree Problem

This is the single most counterintuitive data quality issue in the ASOS record, and it trips up experienced analysts who do not know to look for it.

Here is the chain of events, as IEM’s own dataset documentation describes it. ASOS stations in the United States internally record temperature in whole degrees Fahrenheit. A sensor reading of 78°F is stored as 78. This is not ideal precision for climate analysis, but it is the base reality. The problem arises during transmission. METAR format transmits temperature in degrees Celsius. When the station converts 78°F to Celsius for the METAR, it gets 25.56°C, which it rounds to 26°C for transmission. Now IEM receives 26°C and converts it back to Fahrenheit: 26°C = 78.8°F, which rounds to 79°F. The station observed 78°F; the database records 79°F. A one-degree error has been introduced entirely by the round-trip unit conversion.

The fix is the METAR T-group — a supplemental field in METAR format that provides temperature to one-tenth of a degree Celsius precision. When the T-group is present, IEM can convert back to Fahrenheit with enough precision to recover the original whole-degree value reliably. IEM’s processing explicitly prioritizes data sources that include the T-group, and the NOAAPort feed — IEM’s primary ASOS source — does include it for most US stations most of the time.

What this means in practice: for most routine ASOS observations at most US airports, the temperature record in IEM is accurate to the original whole-degree Fahrenheit. But “most of the time” is doing real work in that sentence. In older archived data, in observations from certain auxiliary sources, and particularly in data from the MADIS high-frequency feed (which transfers temperature in Celsius without adequate precision), you may encounter the one-degree-off error. For a forum discussion about yesterday’s high temperature, this is irrelevant. For research computing climatological normals or extreme temperature statistics, it is something to be aware of and to validate against authoritative sources.

The one-minute ASOS data available through IEM carries an additional version of this problem: the NCEI one-minute data files are processed with a known Celsius precision issue in the FAA data transfer chain. IEM has documented this explicitly in its news archive. If you are pulling one-minute ASOS data for temperature analysis, read the dataset documentation at https://mesonet.agron.iastate.edu/info/datasets/metar.html before you begin.


COOP Data and Time of Observation

The NWS Cooperative Observer network has been taking daily temperature and precipitation readings since the 1880s at some stations. That longevity is exactly what makes COOP data irreplaceable for climatological research — and exactly why TOBS (Time of Observation) matters so much.

COOP observers take their readings once per day, at a time they choose. Some observe at 7 AM. Some at 5 PM. Some at midnight. The recorded high and low temperatures for a given calendar date represent the extremes during the 24-hour period ending at observation time, not the calendar midnight-to-midnight period. An observer who takes readings at 7 AM on the 15th is recording the high and low that occurred between 7 AM on the 14th and 7 AM on the 15th — and attributing those values to the 15th.

This creates what climate scientists call the time-of-observation bias. Morning observers (7 AM) are less likely to inadvertently count the same extreme twice across adjacent calendar days than afternoon observers (5 PM). The 5 PM observer who records an exceptionally hot afternoon might count that same hot afternoon in two consecutive daily records if the heat lingered past midnight. At a network scale, these observation time differences can inflate or suppress apparent monthly means by a fraction of a degree — enough to matter for trend analysis over long records.

IEM’s COOP pages note this directly. The COOP daily observations page warns that for observations made at 7 AM, the data covers a 24-hour period ending at that time, not the local calendar date. This is not IEM’s choice or problem to fix — it is the fundamental structure of how COOP data is collected. But it is something to know when you are comparing COOP daily records to ASOS records, which observe on a different and more consistent time basis, or when you are comparing COOP stations to each other across different observation time regimes.

For most forum-level discussions — “what was the daily high at this COOP station?” — TOBS is background noise. For research examining trends across long periods, or for comparing COOP against other networks, it requires either awareness or explicit correction.


Trace Precipitation

Across IEM’s download portals, you will encounter precipitation values stored as 0.0001 inches. This is not a measurement artifact or a rounding error. It is a deliberate encoding convention meaning Trace — precipitation was observed but too small to measure. In the COOP and ASOS networks, a Trace is officially defined as measurable precipitation below 0.005 inches (the minimum readable increment of a standard rain gauge).

When you download COOP or ASOS precipitation data from IEM and see 0.0001 in your spreadsheet, treat it as a qualitative flag — precipitation occurred — not as a quantitative measurement. Summing these values, computing monthly totals that include trace days, or doing statistics that treat 0.0001 as a real measurement will introduce errors that compound across long records.

The practical test: if you are computing wet-day frequency, a 0.0001 is correctly counted as a wet day. If you are computing seasonal totals, it should either be set to zero or excluded from the numeric sum. Most research-grade precipitation analyses use a threshold (often 0.01 inches) that naturally excludes trace amounts, but if you are working with raw IEM downloads, the convention is worth knowing explicitly.


Missing Data Codes

The specific value used to indicate missing data in IEM downloads varies by format and by download portal, and this inconsistency has caught many users off guard.

In CSV downloads from the ASOS download portal, missing temperature and dew point values typically appear as a blank field or as the string M. In the daily summary download, they may appear as None (Python-style null) or as an empty string depending on the output format selected. In JSON API responses, missing values are usually returned as null. In some older COOP archive formats, 999.9 or -99 appear as sentinel values for missing data — a legacy of pre-digital encoding conventions that occasionally survived into digitized records.

The safest approach when writing any code or script that processes IEM data: do not assume you know what the missing value sentinel will be. Check the first few rows of your download for unexpected values before computing statistics, and review the dataset documentation at https://mesonet.agron.iastate.edu/info/datasets/ for the specific format you are working with.


What a Real Data Quality Investigation Looks Like: MDT and CXY

Abstract discussion of data quality issues only goes so far. Here is what it looks like when you actually find something.

Harrisburg, Pennsylvania has two ASOS stations within a few miles of each other: KMDT (Harrisburg International Airport, in Middletown) and KCXY (Capital City Airport, in New Cumberland). Both have long continuous records. Both are official FAA ASOS installations. On paper they should agree reasonably well for most variables, given their proximity and shared climate regime.

A comparison of precipitation records between these two stations over a multi-year period reveals a persistent and substantial difference: KCXY runs approximately 11% wetter than KMDT on an annual basis. This is not noise. Eleven percent is well outside what you would expect from two stations separated by a few miles in flat valley terrain with no significant orographic difference between them.

What causes this? The most likely explanation is gauge siting and local exposure differences rather than instrument malfunction. ASOS tipping-bucket rain gauges are sensitive to wind effects at the orifice — a gauge in a more sheltered location will catch slightly more precipitation than one exposed to stronger winds, all else equal. Small differences in nearby obstructions, prevailing wind direction during precipitation events, and gauge height above grade can all contribute to systematic biases that are real, repeatable, and entirely invisible to automated QC systems that are checking for gross errors rather than subtle inter-station biases.

The lesson is not that one of these stations is wrong. Both stations are reporting what their gauges observe. The lesson is that gauge measurements are point observations of a physical process that has local variability — and that two official ASOS stations a few miles apart can have genuinely different precipitation climatologies that reflect their specific exposure rather than instrument error.

This kind of investigation — running a long-period comparison between nearby stations and looking for systematic divergence — is exactly the type of work IEM makes straightforward. Autoplot chart q=12 (precipitation summary comparisons), the daily summary download portal, and Climodat all provide the raw material. The MDT/CXY comparison illustrates both the power of having dense, long-record station networks and the caution required when interpreting any single station’s record in isolation.


Station Metadata: Moves, Equipment Changes, and Archive Gaps

A station’s observational record is not just its data — it is the history of its physical location, its equipment, and its observation practices. ASOS stations have been relocated. Their shelter configurations have changed. The transition from manual to automated observation happened at different times for different stations. COOP observers change, and with them, the specific protocols a station follows.

IEM’s station information page, accessible through the Station Locator at https://mesonet.agron.iastate.edu/sites/locate.php, provides metadata for every station in the network — location history, network membership, archive start and end dates, and any documented notes about the station’s history. For research involving long records, this page is where you start, not where you end up accidentally after finding something strange in your data.

When a COOP station moves even a fraction of a mile, its exposure to local drainage, cold air pooling, or orographic effects can change enough to introduce a detectable discontinuity in its temperature or precipitation record. These discontinuities are documented in the station metadata when they are known, but not all station moves are formally recorded in every archive. The variables reference at https://mesonet.agron.iastate.edu/info/variables.phtml is the companion document — it explains what each variable in IEM’s database actually represents, which is its own form of data quality documentation.


What to Watch Out For

A few practical rules that have come out of the issues covered in this chapter:

When computing temperature statistics from ASOS data, especially extremes, validate against the official CF6 climate summaries or NCEI monthly reports for the same station when the result seems surprising. A one-degree discrepancy from the whole-degree storage issue can occasionally push a computed extreme in or out of record territory.

When working with long COOP records, be aware that apparent trends may partially reflect changes in observation time rather than actual climate change. The TOBS adjustment is well-studied in the scientific literature, and the GHCN dataset applies it; raw COOP data as downloaded from IEM does not.

When comparing precipitation between nearby stations, an 11% systematic difference is not automatically a QC problem — it may be a real exposure difference worth understanding rather than correcting. But a 50% difference sustained over multiple years probably warrants a closer look at the gauge siting and maintenance history.

When you encounter an anomalous value in any IEM download, the QC ticket system at https://mesonet.agron.iastate.edu/QC/tickets.phtml is the right place to report it. The IEM team is small, they cannot actively audit every observation across every network, and user-flagged issues have led to real corrections in the archive.


Key URLs — Chapter 10

Resource URL
QC Mainpage https://mesonet.agron.iastate.edu/QC/
QC Trouble Tickets (public) https://mesonet.agron.iastate.edu/QC/tickets.phtml
Currently Offline Stations https://mesonet.agron.iastate.edu/QC/offline.php
MADIS Network QC https://mesonet.agron.iastate.edu/QC/madis/network.phtml
METAR/ASOS Dataset Documentation https://mesonet.agron.iastate.edu/info/datasets/metar.html
Dataset Documentation Index https://mesonet.agron.iastate.edu/info/datasets/
Variables Reference https://mesonet.agron.iastate.edu/info/variables.phtml
Station Locator / Metadata https://mesonet.agron.iastate.edu/sites/locate.php