Absolutely. This is the right moment for the lesson because now we have actual Cumberland behavior to teach from, rather than a stack of hypothetical equations wearing a lab coat.
Open the interactive guide if you have it handy. I’m going to teach it in the same order: Theory → Experiment → GIS → what it means for the 800 m holy grail.
1. What problem did we actually solve?
Forget the model names for a moment.
Cumberland County contains 87 native PRISM 4 km cells in our experiment.
Imagine those 87 cells are 87 weather observers spread across the county. We deliberately remove most of them and ask:
If I retain only a small number of these observations, how well can I reconstruct the precipitation values at every location I threw away?
That is the compression problem.
The symbol
So:
means:
and:
The county retention fraction is:
so at
or about 11.5% retained.
Therefore we have forgotten:
or 88.5% of Cumberland's grid.
That percentage, rather than raw K, is what eventually lets us compare counties of different sizes fairly.
2. What does one experiment actually look like?
This is crucial.
The test period is 2011 through 2025, which gives:
monthly fields.
At
For January 2011, the model predicts all 77.
For February 2011, it predicts all 77 again.
It continues through December 2025.
Therefore:
individual prediction-versus-observation comparisons.
So when you see:
RMSE = 8.78 mm
for spatial maximin at
It comes from 13,860 errors.
For every hidden cell
where:
-
= actual PRISM precipitation -
= predicted precipitation -
= prediction error
The retained cells are excluded from scoring.
That last piece matters enormously. Otherwise we could cheat by saying, “I predicted my retained cells perfectly,” which is roughly the statistical equivalent of grading your own exam after writing the answer key.
3. RMSE versus MAE
MAE asks:
It answers:
On average, how many millimeters away am I?
RMSE asks:
The squaring means large errors hurt much more.
Suppose five errors were:
MAE is:
But RMSE is:
That single 20 mm disaster wallops RMSE.
So when RMSE is much larger than MAE, that tells you:
Most predictions are decent, but there is a tail of ugly misses.
Our residual diagnostic package now lets us examine that tail directly.
4. The first method: spatial maximin
This one is beautifully simple.
Suppose we are allowed to retain 10 cells.
Instead of randomly choosing them, we try to spread them around Cumberland.
The algorithm begins with a cell and repeatedly asks:
Which unselected cell is farthest from the cells I've already selected?
That is maximin.
The name comes from maximizing the minimum distance.
For candidate cell
where
Then select:
In English:
Find each candidate's closest existing sensor. Choose the candidate whose closest sensor is still farthest away.
This progressively fills geographic holes.
Then ordinary inverse-distance weighting predicts the hidden cells.
For a target cell:
with approximately:
Nearby sensors get much more influence than distant ones.
And Cumberland loves this method.
At
At
At
Notice what just happened.
Going from:
doubles the retained cells.
But RMSE only goes:
That's practically a plateau.
That is our first glimpse of an elbow.
5. Why K≈20 looked so interesting
At
So we keep about 23% of Cumberland and throw away about 77%.
Yet spatial maximin gets:
and:
Even better, approximately 71.7% of all hidden cell-month errors are within 5 mm:
and about 88.7% are within 10 mm:
The 95th-percentile absolute error is:
Meaning 95% of those 12,060 predictions are off by less than about 14.7 mm.
That's much richer information than merely saying RMSE=6.63.
6. Why overall correlation can lie to us
At K=10 we got an overall correlation near:
It sounds miraculous.
But precipitation has enormous month-to-month variation.
A very wet month is wet almost everywhere. A dry month is dry almost everywhere.
A model can therefore get lots of correlation credit merely by recognizing:
July was wet.
without correctly reconstructing:
Which part of Cumberland was wetter than which other part?
So we calculated a second quantity.
For every month separately, correlate the predicted and observed values across Cumberland's hidden cells.
Then take the median across 180 months.
At
At
At
That tells us something much more meaningful:
With 20 strategically distributed observations, the internal monthly spatial pattern of Cumberland is usually reproduced extremely well.
7. Terrain K-means
Now we change the question.
Instead of saying:
Spread the observations geographically,
we say:
Preserve different kinds of terrain.
We represent cells with features such as elevation, slope, aspect-related quantities and surrounding relief.
Then K-means clustering partitions cells into terrain groups.
Conceptually, it minimizes:
where:
-
is the terrain feature vector for cell -
is cluster -
is its center
Then we retain representative cells from those terrain clusters.
This sounded promising.
But Cumberland taught us something:
terrain diversity alone does not guarantee good precipitation sampling.
At
Spatial maximin:
Terrain K-means:
Why?
Because the terrain scheme can select several very different terrain environments while accidentally leaving a geographic hole.
For current monthly precipitation, physical coverage matters enormously.
8. The hybrid experiment
We therefore built a hybrid selector.
Its distance says approximately:
We normalized both pieces first so that meters and arbitrary standardized terrain coordinates weren't fighting a barroom brawl over numerical scale.
This was a reasonable idea.
Cumberland politely rejected it.
At
versus:
So our naive 50/50 hybrid is not the answer.
That does not prove hybrid selection is useless.
It says:
Equal weighting is not automatically optimal.
And we absolutely must not tune that weight using 2011–2025 TEST data.
If we tune it later, we use the 1996–2010 validation period.
That distinction between validation and test is sacred if we want these numbers to mean anything.
9. Now the interesting beast: terrain Ridge + IDW
This one is much richer.
For every individual month, we take only the retained K cells.
Suppose K=20.
We know precipitation at those 20 cells.
We also know static predictors such as:
We fit:
But because we have many correlated terrain variables, we use ridge regression.
Ridge solves something like:
The first term wants predictions to fit observations.
The second penalizes enormous coefficients.
That penalty is what keeps a model with correlated terrain variables from becoming completely deranged.
Then we predict the terrain-based spatial trend everywhere.
But we don't stop there.
At each retained sensor:
We interpolate those residuals with IDW:
and finally:
That's why the method is:
terrain Ridge + IDW residual.
The regression explains broad systematic relationships.
The IDW residual catches local departures.
10. Why K=2 blew up
Now the ugly result makes sense.
At K=2 we're fitting a multivariate terrain/location relationship using two current precipitation observations.
That's not magic. That's two guys being asked to represent an entire county while eighteen predictor variables shout suggestions at them.
Ridge keeps the calculation mathematically stable.
It cannot make the relationship scientifically well determined.
Result:
versus ordinary terrain-selected IDW:
By K=20, however:
By K=40:
And at K=40 it actually slightly beats spatial maximin's:
in overall RMSE.
But this is where something much more interesting happens.
11. The q90 story
For each month we take the actual precipitation values across all hidden Cumberland cells.
Sort them.
Find the 90th percentile:
Then do the same using predictions.
The monthly error is:
Now watch this.
At K=20:
Spatial maximin:
Terrain Ridge:
At K=40:
Spatial maximin:
Terrain Ridge:
So something fascinating happens.
Spatial maximin is better at reproducing the whole field.
Terrain Ridge is better at reproducing the upper spatial tail.
Those are not contradictory statements.
They are different objectives.
12. Why IDW struggles with the upper tail
This gets directly to your future 800 m problem.
Ordinary IDW is fundamentally a smoothing operation.
A prediction is a weighted average:
with:
and:
Therefore the prediction normally lies between the retained values.
It has difficulty inventing a local extreme that none of its nearby sensors actually recorded.
That causes regression toward the middle.
And look at the Cumberland q90 bias.
At K=20, spatial maximin has:
At K=40:
Despite having twice as many retained cells.
The upper tail is systematically too low.
That's not merely “the model isn't accurate enough.”
It tells us something structural:
A convex spatial averaging method inherently smooths peaks.
And that should make your Weather Raccoon ears perk up.
Because predicting 800 m upper-tail variability from 4 km data will have exactly this problem.
13. Why terrain Ridge does better at q90
Regression is allowed to say:
This target cell has terrain characteristics that imply a value above the neighboring retained observations.
IDW alone has very limited ability to do that.
The terrain trend can generate contrast.
Then residual IDW corrects it locally.
At K=40 its q90 bias is still negative:
but that's substantially less negative than spatial maximin:
So terrain information is doing something useful to the upper tail even when it does not dominate average-field reconstruction.
That's one of the best scientific results from this pilot.
14. Now meet EOFs, our first actual “weird decomposition”
You joked about Fourier decomposition.
Well, surprise: we're already doing its weird cousin.
EOF means Empirical Orthogonal Function.
Meteorologists use this constantly. Mathematically it is closely related to principal component analysis and singular value decomposition.
Take the training archive.
Imagine a matrix:
After removing monthly climatology, we decompose:
This is the singular value decomposition, SVD.
The rows of
Think of them like spatial chords.
One pattern might broadly mean:
western Cumberland wetter, eastern Cumberland drier.
Another might represent:
ridge cells enhanced relative to valley cells.
Another could describe a north-south contrast.
Instead of describing 87 separate values, a monthly field may be approximated as:
where:
-
= normal field for that calendar month -
= learned spatial EOF pattern -
= coefficient describing how much of that pattern is present this month
This is conceptually very close to Fourier analysis.
Fourier says:
Let me express the signal using fixed sine and cosine waves.
EOF says:
Let me discover the spatial patterns that this dataset itself naturally uses.
That difference is huge.
15. What QR pivoting is doing
Once we have learned those modes, we ask:
Which physical PRISM cells give us the most information about the coefficients of those modes?
That is where QR pivoting enters.
Rather than choosing sensors because they are merely spread out geographically, QR tries to choose observations that make the learned spatial basis easiest to reconstruct.
Very loosely:
as informative and numerically well-conditioned as possible.
Then during the TEST period we observe only those sensor cells and solve for the mode amplitudes.
The full field follows from the modes.
And look what happens.
At K=20:
At K=40:
and at K=40:
Essentially zero.
That is remarkable.
But we must use the correct language.
This is archive-aware compression.
It is not the strict terrain reconstruction problem because it learned Cumberland's full spatial behavior from 1895–1995.
16. Why this result matters for your holy grail
This is where your Fourier instinct becomes genuinely interesting.
The archive result tells us:
The spatial precipitation field is highly low-dimensional.
In plain language:
Although Cumberland has 87 cells, those 87 numbers do not behave like 87 completely independent variables.
A relatively small collection of spatial patterns explains a lot of what they do.
That is exactly what dimensional decomposition discovers.
And your future problem is:
Can we infer fine 800 m spatial variability from coarse 4 km meteorological information, stations/gauges, and fine terrain without feeding the 800 m climate values into the predictor?
There may be a similar low-dimensional structure hiding inside each 4 km cell.
That is where combinations of:
become extremely interesting.
17. Go to the GIS tab mentally now
This is where aggregate statistics stop being enough.
At K=20, spatial maximin's worst hidden cell has roughly:
while some hidden cells are down near:
So “Cumberland RMSE = 6.63” hides enormous spatial variation.
And I ran another calculation from the per-cell diagnostic table.
For spatial maximin at K=10, the correlation between a hidden cell's RMSE and its distance from the nearest retained sensor is approximately:
That's strong.
At K=20:
At K=40:
So even once we have a decent network, distance to observational support remains a major driver of error.
That validates the importance of geographic spacing.
18. But terrain starts appearing in the residual pattern
Here's an even cooler clue.
At K=40, for spatial maximin, per-cell RMSE correlates approximately with:
Those are descriptive correlations, not causal proof, and the hidden population changes with K.
But the pattern is provocative:
Once geographic sampling becomes dense enough, terrain appears increasingly responsible for what spatial interpolation still cannot explain.
Now compare terrain Ridge at K=40.
Its correlation between RMSE and:
Basically nothing.
Yet RMSE versus nearest-sensor distance remains around:
That's a gorgeous result conceptually.
It suggests:
The terrain regression may be successfully removing much of the systematic terrain-related error, leaving spatial support distance as the remaining problem.
That is precisely the sort of thing the GIS was built to expose.
19. How I want you to use the interactive now
Here is the actual guided lab:
-
In Theory, set K to 10, then 20, then 40. Watch retained percentage rather than K. K=20 means about 23% kept and 77% forgotten.
-
Go to Experiment and select RMSE. Compare spatial maximin to terrain Ridge. Notice that maximin wins most of the useful K range, while Ridge catches it around K=40.
-
Change the metric to monthly q90 MAE. Now watch Ridge become much more competitive. This demonstrates that “best model” depends on what part of the distribution you care about.
-
Select Archive EOF + QR and study the q90 time series. Notice how much closer the predicted upper tail becomes once historical spatial modes are available.
-
Go to GIS, choose spatial maximin, K=20, and color by cell RMSE. Tap high-error cells. Then switch to nearest-sensor distance. You should see why spatial coverage matters.
-
Keep K=20 and switch to elevation, slope and ruggedness. Then switch methods to terrain Ridge. The question becomes: did the terrain-aware model reduce errors specifically where terrain is unusual?
-
Finally use K=40 and compare spatial maximin versus terrain Ridge. Average RMSE is now similar, but q90 and terrain-linked residual behavior are not. That is the clue that one scalar leaderboard is inadequate.
20. One subtle warning about comparing K values
There is a slightly nasty statistical detail.
At K=10, we score 77 hidden cells.
At K=20, we score 67.
At K=40, only 47.
And some cells that were targets at one K become sensors at another K.
Therefore the evaluation population changes.
That is why an error metric does not have to decrease monotonically as K increases.
It can occasionally go:
even though you added sensors.
You may have removed easy cells from the scoring set and left harder ones behind.
For practical compression, our current metric is correct:
How accurately can I reconstruct what I actually threw away?
But for pure method diagnosis, I also want a common fixed holdout set eventually.
Then every K and method would be tested on exactly the same cells.
That gives us a cleaner scientific comparison.
21. What Cumberland has actually taught us
Not “maximin wins.”
That's too simplistic.
The deeper lesson is:
Spatial coverage controls broad reconstruction accuracy. Terrain helps explain the hard residual structure and upper tail. Historical spatial modes reveal that the field is strongly compressible.
Those are three different pieces of the problem.
Spatial maximin says:
Don't leave holes.
Terrain Ridge says:
Geography alone smooths away physically meaningful contrast.
EOF says:
Much of the apparently huge spatial field actually lives in a much smaller mathematical space.
And q90 says:
A model that is excellent on average can still systematically erase the thing we care most about.
That final point is enormous for the Weather Raccoon holy grail.
Because if our goal eventually becomes:
then minimizing ordinary RMSE alone would be the wrong religion.
We need to model the shape of the subgrid distribution, not merely its mean.
And now, thanks to Cumberland, we have empirical evidence showing exactly why.