↖ CPA Weather Lab
Weather Raccoon research course · 10 sessions

Rainfall truth → QC → independent validation → error-aware modeling

A companion guide to the TTS narrative for four papers. The purpose is not to memorize equations. It is to build the intellectual plumbing needed to decide whether a Pennsylvania precipitation model is actually better than PRISM without accidentally grading it with its own homework.

Course map

Paper 1 · Sieck et al. 2007

2 sessions. Establish what a gauge can and cannot tell us. Calibration, wind, leveling, malfunction, representativeness.

Paper 2 · Ośrodka et al. 2025

3 sessions. Turn uncertain measurements into a graded, auditable QC evidence system.

Paper 3 · Rivera-Giboyeaux & Weinbeck 2024

2 sessions. Build truly independent validation and investigate outliers physically.

Paper 4 · White & Nelson 2024

3 sessions. Predict an error distribution using terrain, storm, radar-quality, and observational features.

The master idea

The four papers form one chain. If the reference observation is questionable, QC must preserve that uncertainty. If QC uses radar, validation lineage changes. If validation leaks information, machine-learning skill is inflated. And if a model produces only one deterministic value, it conceals the uncertainty exposed by the first three papers.

Paper 1 · Sieck, Burges & Steiner (2007)

measurement errorwind undercatchdynamic calibrationpoint support

Big concept

A rain gauge is an instrument, not an oracle. Reliable point rainfall requires functioning hardware, calibration, maintenance, correct leveling, reasonable exposure, and critical evaluation. Even a perfect point observation does not automatically represent an 800 m or 4 km area.

What changes in our project

Never store only station,date,ppt. Keep raw amount, instrument/network metadata, observation window, QC evidence, correction status, local-consensus support, and lineage.

Two classes of disagreement

ClassMeaningExamplesImplication
Measurement errorThe point observation itself is biased or wrong.Calibration drift, clogging, wind undercatch, tilt, overflow, mechanical failure.May justify a QC penalty or correction.
Representativeness errorThe gauge is right at the point, but the grid-cell mean is different.Convective core misses gauge by 300 m, orographic gradient within cell.Do not “correct” the gauge merely to match the grid.

Interactive: why a small catch bias matters

Illustrative calculator, not a formula fitted by Sieck et al. It simply shows the arithmetic consequence of undercatch.

Observed gauge total
46.0 mm
Apparent model bias if model were perfect
+4.0 mm
Session 1 · What does a rain gauge actually measure?

Focus on physical measurement, malfunction, colocated evidence, and the distinction between point truth and grid support. Read the Goodwin Creek storm tables as a forensic exercise: dense networks make broken instruments visible.

Session 2 · Calibration, wind, and intense rainfall

Focus on rain-rate-dependent tipping-bucket calibration, seasonal/mechanical variability, aerodynamic undercatch, and why a sophisticated correction can fail if its input measurements and assumptions are themselves uncertain.

Paper basis: Sieck et al. (2007), especially sections on data QC, calibration, wind effects, and the Goodwin Creek storm comparisons.

Paper 2 · Ośrodka et al. (2025) RainGaugeQC

Big concept

QC should accumulate evidence, not simply delete inconvenient observations. Their modified system starts observations at QI = 1.0 and reduces the quality index as individual checks find problems.

Algorithm architecture

Abbrev.Purpose after modificationPA analogue
GECGross errorsImpossible totals, encoding failures
RCClimatological rangeNetwork-/season-aware plausibility
RCCFalse rainy/non-rainy eventsRadar/neighbor occurrence support
BSCBlocked sensorPersistent false zero detection
TCCGauge time series vs adjusted radarPersistent timing/response problems
BCBias detection/correction with adjusted radarSystematic low/high station behavior
SCCLocal outlier detection conditioned by radar variabilityBuddy check that respects convection

Equations worth keeping

Bias ratio: bias = ΣR / ΣG

The paper computes accumulated adjusted-radar precipitation ΣR and gauge precipitation ΣG over a recent period. Ratios far from 1 receive larger QI penalties.

Similarity rule: similar if 1.3 · min(ΣG, ΣR) + 7.0 > max(ΣG, ΣR)
RMSE: √[(1/n) Σ(Eᵢ − Oᵢ)²]

Interactive: simplified QI evidence accumulator

This demo mirrors the paper's philosophy but is not a complete reproduction of every operational branch.

Resulting illustrative QI
1.00
Session 1 · QC as an evidence engine

Understand why layered component flags plus a final QI are more useful than a binary keep/delete decision.

Session 2 · TCC, bias, SCC, and radar entanglement

Study how temporal correlation, accumulated bias, and spatial consistency catch different failure modes. Then confront the lineage consequence of using radar as QC evidence.

Session 3 · Pennsylvania implementation

Translate the paper into station-observation fields and product-specific validation eligibility masks. Calibrate thresholds locally rather than copying Polish values as if climate and networks were firmware-compatible.

Paper basis: Ośrodka et al. (2025), Table 2 and sections 3.3–3.6, verification against independent manual gauges, and case studies.

Paper 3 · Rivera-Giboyeaux & Weinbeck (2024)

Big concept

The Savannah River Site gauges were not part of HADS, the network used in the MRMS gauge-correction scheme. That gives the study a substantially cleaner independent validation set.

Lineage is a graph, not a network label

Reference gauge pathwayIndependent for MRMS gauge-corrected validation?Why
Gauge used in HADS correctionNoThe target helped build the product.
Gauge outside HADS, untouched by MRMS-based QCMuch strongerProduct did not directly assimilate that witness.
Gauge outside HADS, but MRMS used to QC/correct itCompromisedIndirect circularity has been introduced.

Interactive: paired validation

Adjust one held-out event. The quantity we care about is not “candidate RMSE” in isolation but candidate error minus PRISM error on the same witness.

PRISM absolute error
12 mm
Candidate absolute error
5 mm
Paired improvement
7 mm

Outliers are scientific cases

The paper connects some large discrepancies to convective/heavy-rain and hail classifications, but not every convective event fails. It discusses point-grid mismatch, fast evolution, below-beam microphysics, tipping-bucket undercatch, and cool-season beam overshoot as plausible mechanisms.

Session 1 · Independent validation genealogy

Define independence at physical station × time × product pathway. Build exclusion masks, not vague assurances that one network is “independent.”

Session 2 · Metrics, regimes, and outlier forensics

Use correlation and RMSE, but stratify by precipitation regime and investigate outliers physically instead of deleting them from the story.

Paper basis: Rivera-Giboyeaux & Weinbeck (2024), abstract, methods, event tables, regression/outlier analysis, hydrometeor-classification case studies, discussion and conclusions.

Paper 4 · White & Nelson (2024)

Big concept

Model the distribution of expected QPE error as a function of storm, terrain, radar quality, and location. Reliability is spatiotemporally variable, not a fixed property of “mountains.”

Quantile / pinball loss

L(y, ŷ) = (1/N) Σ [ α·max(yᵢ−ŷᵢ,0) + (1−α)·max(ŷᵢ−yᵢ,0) ]

Three models at α ≈ 0.05, 0.50, and 0.95 produce a median and an approximate 90% error interval. At α = 0.50 the loss is half the mean absolute error.

Interactive: asymmetric quantile penalty

0.50
Single-case pinball loss (illustrative units)

Features used after selection

FamilyExamples in the paperPotential PA extension
RainfallMedian intensity, intensity standard deviation, durationEvent total, intermittency, convective fraction
Radar qualityMinimum RQI, RQI variabilityBeam height/blockage, multi-radar agreement
TerrainElevation, aspectSlope, relief, terrain position, exposure, barrier geometry
StormArea, velocity, spatial/temporal variancePropagation vector, backbuilding, low-level inflow orientation
Location/timeLatitude, month/hour candidatesCounty, season, radar sector, PRISM facet classes

Validation warning

White & Nelson explicitly group samples by location in cross-validation and reserve a separate test set because random splitting of spatiotemporally correlated samples can inflate performance. Our Pennsylvania design should add event-block and geographic-block holdouts.

Session 1 · Predict the error, not just the rain

Reframe the target from deterministic correction to a conditional error distribution.

Session 2 · Quantile loss and blocked validation

Understand why quantiles handle heteroscedastic error, then focus on the much more important issue of preventing location/event leakage.

Session 3 · Feature importance without fairy tales

Intensity variability dominated; duration and RQI added skill; terrain was less dominant than expected. Treat importance as hypothesis generation and test interactions physically.

Paper basis: White & Nelson (2024), methods sections 2.2–2.3, Eq. 1, Figures 2–8, Table 2, discussion and conclusion.

Pennsylvania synthesis · What “better than PRISM” means

Truth layer

Independent, exact-window, well-QC’d physical gauges with explicit uncertainty and provenance.

QC layer

Component evidence plus final QI; never destroy raw data or the reason for a flag.

Lineage layer

Station × time × product pathway exclusion masks for PRISM, MRMS, Stage IV, and candidate models.

Validation layer

Station blocks + event blocks + geographic blocks + temporal holdout. Score identical cases.

Model layer

Precipitation estimate plus separate conditional error/uncertainty model.

Baseline layer

PRISM, RadarOnly, Stage IV, simple interpolation/CAI-like field, and a simple blend.

Minimum scorecard

Metric / testWhy it exists
Bias / mean errorDetect systematic high/low behavior.
MAERobust average magnitude of misses.
RMSEEmphasizes large misses; useful but extreme-sensitive.
CorrelationPattern agreement; never sufficient alone.
Wet/dry occurrenceFalse rain and missed rain.
Upper-tail errorHeavy-rain performance is its own problem.
Spatial-gradient/location errorWas the storm core put in the right place?
Event-total errorPreserves storm-scale hydrologic relevance.
Interval calibrationIf we claim 80% or 90% uncertainty intervals, do they actually cover at that rate?
Paired improvement vs PRISMDirect answer to the project question on identical withheld evidence.

Target product

Best estimate: 37 mm · 80% interval: 29–45 mm · independent gauge support: moderate · radar branch: reliable · local gradient: strong convective · PRISM expected uncertainty: elevated.

That is more useful than a naked 37.2 mm because it tells us what we think happened, how certain we are, and what evidence controls the confidence.

← CPA Weather Lab · Learn