Course map
2 sessions. Establish what a gauge can and cannot tell us. Calibration, wind, leveling, malfunction, representativeness.
3 sessions. Turn uncertain measurements into a graded, auditable QC evidence system.
2 sessions. Build truly independent validation and investigate outliers physically.
3 sessions. Predict an error distribution using terrain, storm, radar-quality, and observational features.
The master idea
The four papers form one chain. If the reference observation is questionable, QC must preserve that uncertainty. If QC uses radar, validation lineage changes. If validation leaks information, machine-learning skill is inflated. And if a model produces only one deterministic value, it conceals the uncertainty exposed by the first three papers.
Paper 1 · Sieck, Burges & Steiner (2007)
measurement errorwind undercatchdynamic calibrationpoint support
Big concept
A rain gauge is an instrument, not an oracle. Reliable point rainfall requires functioning hardware, calibration, maintenance, correct leveling, reasonable exposure, and critical evaluation. Even a perfect point observation does not automatically represent an 800 m or 4 km area.
What changes in our project
Never store only station,date,ppt. Keep raw amount, instrument/network metadata, observation window, QC evidence, correction status, local-consensus support, and lineage.
Two classes of disagreement
| Class | Meaning | Examples | Implication |
|---|---|---|---|
| Measurement error | The point observation itself is biased or wrong. | Calibration drift, clogging, wind undercatch, tilt, overflow, mechanical failure. | May justify a QC penalty or correction. |
| Representativeness error | The gauge is right at the point, but the grid-cell mean is different. | Convective core misses gauge by 300 m, orographic gradient within cell. | Do not “correct” the gauge merely to match the grid. |
Interactive: why a small catch bias matters
Illustrative calculator, not a formula fitted by Sieck et al. It simply shows the arithmetic consequence of undercatch.
Session 1 · What does a rain gauge actually measure?
Focus on physical measurement, malfunction, colocated evidence, and the distinction between point truth and grid support. Read the Goodwin Creek storm tables as a forensic exercise: dense networks make broken instruments visible.
Session 2 · Calibration, wind, and intense rainfall
Focus on rain-rate-dependent tipping-bucket calibration, seasonal/mechanical variability, aerodynamic undercatch, and why a sophisticated correction can fail if its input measurements and assumptions are themselves uncertain.
Paper 2 · Ośrodka et al. (2025) RainGaugeQC
Big concept
QC should accumulate evidence, not simply delete inconvenient observations. Their modified system starts observations at QI = 1.0 and reduces the quality index as individual checks find problems.
Algorithm architecture
| Abbrev. | Purpose after modification | PA analogue |
|---|---|---|
| GEC | Gross errors | Impossible totals, encoding failures |
| RC | Climatological range | Network-/season-aware plausibility |
| RCC | False rainy/non-rainy events | Radar/neighbor occurrence support |
| BSC | Blocked sensor | Persistent false zero detection |
| TCC | Gauge time series vs adjusted radar | Persistent timing/response problems |
| BC | Bias detection/correction with adjusted radar | Systematic low/high station behavior |
| SCC | Local outlier detection conditioned by radar variability | Buddy check that respects convection |
Equations worth keeping
The paper computes accumulated adjusted-radar precipitation ΣR and gauge precipitation ΣG over a recent period. Ratios far from 1 receive larger QI penalties.
Interactive: simplified QI evidence accumulator
This demo mirrors the paper's philosophy but is not a complete reproduction of every operational branch.
Session 1 · QC as an evidence engine
Understand why layered component flags plus a final QI are more useful than a binary keep/delete decision.
Session 2 · TCC, bias, SCC, and radar entanglement
Study how temporal correlation, accumulated bias, and spatial consistency catch different failure modes. Then confront the lineage consequence of using radar as QC evidence.
Session 3 · Pennsylvania implementation
Translate the paper into station-observation fields and product-specific validation eligibility masks. Calibrate thresholds locally rather than copying Polish values as if climate and networks were firmware-compatible.
Paper 3 · Rivera-Giboyeaux & Weinbeck (2024)
Big concept
The Savannah River Site gauges were not part of HADS, the network used in the MRMS gauge-correction scheme. That gives the study a substantially cleaner independent validation set.
Lineage is a graph, not a network label
| Reference gauge pathway | Independent for MRMS gauge-corrected validation? | Why |
|---|---|---|
| Gauge used in HADS correction | No | The target helped build the product. |
| Gauge outside HADS, untouched by MRMS-based QC | Much stronger | Product did not directly assimilate that witness. |
| Gauge outside HADS, but MRMS used to QC/correct it | Compromised | Indirect circularity has been introduced. |
Interactive: paired validation
Adjust one held-out event. The quantity we care about is not “candidate RMSE” in isolation but candidate error minus PRISM error on the same witness.
Outliers are scientific cases
The paper connects some large discrepancies to convective/heavy-rain and hail classifications, but not every convective event fails. It discusses point-grid mismatch, fast evolution, below-beam microphysics, tipping-bucket undercatch, and cool-season beam overshoot as plausible mechanisms.
Session 1 · Independent validation genealogy
Define independence at physical station × time × product pathway. Build exclusion masks, not vague assurances that one network is “independent.”
Session 2 · Metrics, regimes, and outlier forensics
Use correlation and RMSE, but stratify by precipitation regime and investigate outliers physically instead of deleting them from the story.
Paper 4 · White & Nelson (2024)
Big concept
Model the distribution of expected QPE error as a function of storm, terrain, radar quality, and location. Reliability is spatiotemporally variable, not a fixed property of “mountains.”
Quantile / pinball loss
Three models at α ≈ 0.05, 0.50, and 0.95 produce a median and an approximate 90% error interval. At α = 0.50 the loss is half the mean absolute error.
Interactive: asymmetric quantile penalty
Features used after selection
| Family | Examples in the paper | Potential PA extension |
|---|---|---|
| Rainfall | Median intensity, intensity standard deviation, duration | Event total, intermittency, convective fraction |
| Radar quality | Minimum RQI, RQI variability | Beam height/blockage, multi-radar agreement |
| Terrain | Elevation, aspect | Slope, relief, terrain position, exposure, barrier geometry |
| Storm | Area, velocity, spatial/temporal variance | Propagation vector, backbuilding, low-level inflow orientation |
| Location/time | Latitude, month/hour candidates | County, season, radar sector, PRISM facet classes |
Validation warning
White & Nelson explicitly group samples by location in cross-validation and reserve a separate test set because random splitting of spatiotemporally correlated samples can inflate performance. Our Pennsylvania design should add event-block and geographic-block holdouts.
Session 1 · Predict the error, not just the rain
Reframe the target from deterministic correction to a conditional error distribution.
Session 2 · Quantile loss and blocked validation
Understand why quantiles handle heteroscedastic error, then focus on the much more important issue of preventing location/event leakage.
Session 3 · Feature importance without fairy tales
Intensity variability dominated; duration and RQI added skill; terrain was less dominant than expected. Treat importance as hypothesis generation and test interactions physically.
Pennsylvania synthesis · What “better than PRISM” means
Independent, exact-window, well-QC’d physical gauges with explicit uncertainty and provenance.
Component evidence plus final QI; never destroy raw data or the reason for a flag.
Station × time × product pathway exclusion masks for PRISM, MRMS, Stage IV, and candidate models.
Station blocks + event blocks + geographic blocks + temporal holdout. Score identical cases.
Precipitation estimate plus separate conditional error/uncertainty model.
PRISM, RadarOnly, Stage IV, simple interpolation/CAI-like field, and a simple blend.
Minimum scorecard
| Metric / test | Why it exists |
|---|---|
| Bias / mean error | Detect systematic high/low behavior. |
| MAE | Robust average magnitude of misses. |
| RMSE | Emphasizes large misses; useful but extreme-sensitive. |
| Correlation | Pattern agreement; never sufficient alone. |
| Wet/dry occurrence | False rain and missed rain. |
| Upper-tail error | Heavy-rain performance is its own problem. |
| Spatial-gradient/location error | Was the storm core put in the right place? |
| Event-total error | Preserves storm-scale hydrologic relevance. |
| Interval calibration | If we claim 80% or 90% uncertainty intervals, do they actually cover at that rate? |
| Paired improvement vs PRISM | Direct answer to the project question on identical withheld evidence. |
Target product
Best estimate: 37 mm · 80% interval: 29–45 mm · independent gauge support: moderate · radar branch: reliable · local gradient: strong convective · PRISM expected uncertainty: elevated.
That is more useful than a naked 37.2 mm because it tells us what we think happened, how certain we are, and what evidence controls the confidence.