↖ CPA Weather Lab

Cumberland County
PRISM Compression Lab

From data to understanding — one county, many lessons

Interactive exploration of precipitation reconstruction using only a small fraction of the original 4 km grid.

87
PRISM cells in Cumberland
180
held-out test months
5
deterministic methods + random
933,300
deterministic residuals
15 y
test: 2011–2025

What this lab answers

How much of Cumberland's native PRISM grid can be removed while still reconstructing the hidden precipitation field?

Default view
K = 20
23.0% retained · 77.0% forgotten
Best strict whole-field RMSE @ K=20
6.63 mm
Spatial maximin + IDW
Terrain Ridge q90 MAE @ K=20
4.16 mm
Upper-tail skill beats spatial maximin
Archive EOF q90 MAE @ K=20
1.81 mm
Archive-aware upper benchmark

The headline

Cumberland is strongly compressible. At K=20, the strict spatial method keeps only 20 of 87 cells yet reconstructs the other 67 for all 180 held-out months with RMSE 6.63 mm and median monthly spatial correlation 0.926.

Spatial coverage is the dominant control on average error. Terrain becomes especially important for the upper tail and for explaining where the remaining errors live.

The surprise

The method that wins ordinary RMSE is not automatically the method that wins the wet end of the distribution. At K=20, terrain Ridge + IDW residual has worse whole-field RMSE than spatial maximin, but a better monthly q90 MAE.

That distinction is directly relevant to the future 800 m “holy grail”: predicting subgrid variability and extremes is a different target from predicting the average field.

Start exploring

Every control in this site is live. Nothing is decorative window dressing pretending to be software.

Theory & concepts

Use the controls while you learn. The point is to see the mathematics change the experiment, not merely stare at notation until morale improves.

K means retained cells

K is the number of native PRISM cells kept as observations. Everything else is hidden and predicted.

retention = K / 87
forget = 1 − K/87

Every hidden cell is scored

For every test month, the method predicts every non-retained cell. Retained cells do not get to congratulate themselves for predicting themselves.

scores = (87 − K) × 180

The split is temporal

Training and testing are separated by whole years so the future test period is never used to fit the model.

TRAIN 1895–1995
VALIDATE 1996–2010
TEST 2011–2025

Play with K

K=20 · retained · forgotten

hidden cell-month prediction errors will be scored.

For statewide county comparisons, retention percentage is the primary normalization. Raw K still matters operationally, but raw K alone is not fair across differently sized counties.

The reconstruction methods

RMSE

RMSE = √ mean[(prediction − truth)²]

Large misses are punished strongly.

MAE

MAE = mean[|prediction − truth|]

Typical absolute miss in millimeters.

Monthly spatial correlation

corr(hidden truth, hidden prediction)
computed month by month

Asks whether the within-county spatial pattern is right, not merely whether wet months are wet.

Experiment dashboard

Choose K and the performance metric. The charts are rebuilt from the actual Cumberland diagnostic package.

Retention
—
—
Best strict RMSE
—
—
Archive-aware RMSE
—
EOF + QR benchmark
Best strict q90 MAE
—
—

Compression curves

How performance changes as more cells are retained.

Monthly 90th percentile

Observed versus predicted upper spatial tail.

Error quantiles

Absolute cell-month error distribution for the selected q90 method and K.

Selected-K method table

MethodClassRMSEMAEq90 MAEq90 biasMonthly r

GIS Explorer

This local GIS is deliberately self-contained so it still works when Android opens the HTML as content://. Tap a cell and inspect the actual reconstruction diagnostics.

Cell RMSE
——
■ gold outline = retained sensor

Select a grid cell

Tap any cell on the map.

Layers

This is the analytical PRISM target-cell footprint, not a fabricated street map. A future served version can add your actual 30 m DEM/hillshade and MapLibre basemap without sacrificing this offline core.

Results & insights

The experiment did more than pick a winner. It separated three different sources of skill: spatial coverage, terrain physics, and learned historical spatial modes.

1 · Spatial coverage wins the average field

At K=20, spatial maximin + IDW reaches RMSE 6.63 mm and MAE 4.29 mm while keeping only 23% of Cumberland's cells.

That is the cleanest strict compression result.

2 · Terrain matters more in the wet tail

At K=20, terrain Ridge + IDW residual has q90 MAE 4.16 mm versus 5.10 mm for spatial maximin, despite a worse whole-field RMSE.

“Best model” depends on what part of the distribution you care about.

3 · Archive modes expose low dimensionality

Archive EOF + QR reaches RMSE 5.75 mm and q90 MAE 1.81 mm at K=20. It has an advantage because complete 1895–1995 spatial fields were allowed during training.

The field contains repeatable spatial structure that a relatively small number of observations can constrain.

The q90 warning

Spatial IDW systematically smooths peaks. At K=20 its monthly q90 bias is about −5.03 mm. Terrain Ridge reduces that to about −3.84 mm. At K=40 the archive EOF q90 bias is essentially zero.

This is exactly why a future 800 m downscaling system cannot be judged by ordinary RMSE alone.

The next scientific move

Run the county-size-normalized retention experiment, but keep the diagnostic anatomy. Later, use validation-only tuning for any hybrid weight. For the 800 m problem, start thinking in terms of distributional targets, multiscale spatial bases, terrain constraints, and independent station/gauge observations.

Full teaching lesson

A rewritten narrative built around the actual Cumberland experiment rather than a generic statistics lecture.

← CPA Weather Lab · Learn