Local Effective Degrees of Freedom Explorer
This page explains one proposed organizing variable for the PRISM sparse-network experiment: local effective degrees of freedom, or local effective dimensionality.
The idea is simple: hundreds of PRISM cells do not necessarily contain hundreds of independent precipitation signals. The real scientific question is how many genuinely different spatial patterns are present.
How many values PRISM stores.
How many statistically distinct pieces of spatial information those values contain.
1. The central equation
Number of PRISM cells in the neighborhood.
Correlation between standardized precipitation anomalies at cells j and k.
Effective number of independent spatial modes.
Strong correlations make the denominator large, reducing Neff. Weak correlations make the denominator smaller, pushing Neff toward the physical cell count.
2. Interactive simplified neighborhood
This uses a simplified equal-correlation model where all off-diagonal cell pairs have the same correlation magnitude ρ. Real PRISM neighborhoods will have a full distribution of pairwise correlations.
3. Two extreme examples
All 40 cells behave identically
Forty stored values, approximately one independent signal.
All 40 cells behave independently
Forty stored values, approximately forty independent signals.
4. Eigenvalue interpretation
The same quantity can be written using the eigenvalues λ of the local correlation matrix:
For a correlation matrix, Σλj = n. If most variance is concentrated in only a few eigenmodes, the effective dimensionality is far below the number of physical cells.
20, 10, 5, 3, 1, 1, 0, …
Most of the information is concentrated in a handful of spatial patterns.
5. The Gram-matrix computational trick
Let Z be the standardized anomaly matrix, with time in rows and neighboring PRISM cells in columns. For 2003–2010 monthly data, T = 96.
The direct cell-by-cell matrix ZTZ is n × n. But the time-domain Gram matrix ZZT is only 96 × 96. This allows the Frobenius norm needed by the participation ratio to be computed without building a huge spatial correlation matrix.
Because temporal demeaning removes one degree of freedom, the strict sample rank is generally no greater than roughly T−1, so 96 monthly observations provide at most about 95 resolvable temporal dimensions.
6. Scale dependence: Neff(R)
One Neff value is not enough. Compute it at multiple radii such as 16, 32, 64, and 128 km. The rate at which effective dimensionality grows with radius is itself a measure of spatial decorrelation.
7. The money test: Neff versus the sensor-count knee
For each county, build an error-versus-sensor-count curve. Error should fall rapidly at low K and then flatten. The bend is the knee, where additional sensors stop buying much improvement.
8. What changes around 2002?
The useful question is no longer merely whether precipitation statistics changed. It becomes: did PRISM's effective spatial dimensionality change?
Estimate local dimensionality before the production transition.
Test whether the break is sharp, gradual, or absent.
Determine whether the later field contains more or fewer resolvable spatial modes.
9. Seasonality
Warm season
Localized convection may reduce correlation length and increase spatial fragmentation.
Cool season
Synoptic and orographic precipitation may be broader and more spatially coherent.
These are hypotheses to test, not assumptions to bake into the analysis.
10. A cleaner fingerprint model
This reframes the fingerprint program. Instead of directly regressing RMSE against dozens of county descriptors, ask whether physical and observational geography explain local dimensionality, and whether dimensionality predicts the scaling-curve knee.
11. Finite-sample bias
With only a finite number of months, unrelated cells still acquire accidental sample correlations. These inflate Σr² and bias Neff downward.
| Known synthetic rank | Illustrative estimate | Meaning |
|---|---|---|
| 5 | 4.8 | Small bias |
| 10 | 9.0 | Moderate compression |
| 20 | 15.7 | Larger finite-T bias |
| 40 | 24.1 | High ranks difficult to resolve |
These numbers are conceptual examples, not PRISM results. Actual calibration should use synthetic fields with known rank and realistic temporal structure, including AR(1)-matched or phase-randomized surrogates.
12. What Neff means, and what it does not
- Effective dimensionality of the local monthly PRISM precipitation field.
- Spatial redundancy versus complexity.
- Candidate predictor of sensor-count knee.
- A quantity to compare across scales, seasons, and production eras.
- The number of rain gauges the real atmosphere requires.
- The true dimensionality of precipitation independent of PRISM.
- Exact PRISM station-to-cell interpolation weights.
- A guaranteed one-for-one mapping between Neff and sensor count.
13. Rotating validation architecture
PRISM production era: pre-radar, transition, radar-informed.
Station-support era: stable observational epochs defined from station history.
Blocked temporal folds: whole years or contiguous blocks rotated through development and evaluation.
Cross-validation asks whether a design generalizes to withheld periods or station regimes. Block bootstrap asks how uncertain the resulting metrics are. Those are related, but not interchangeable.
14. The decisive falsifiable test
Effective dimensionality becomes a credible organizing variable for sensor requirement, seasonality, regime change, quota design, and optimizer behavior.
The theory fails cheaply, and the project avoids months of architecture built around a seductive but unsupported idea.
Bottom line
The physical number of PRISM cells is not the main quantity. The central question is how many independent spatial precipitation patterns those cells contain. Measure that locally, measure how it grows with spatial radius, test how it changes across seasons and production eras, and then ask whether it predicts where the sensor-count scaling curve reaches its knee.
If it does, effective dimensionality becomes the spine of the PRISM experiment.