↖ CPA Weather Lab
PRISM Terrain Prediction / Effective Information Content

Local Effective Degrees of Freedom Explorer

This page explains one proposed organizing variable for the PRISM sparse-network experiment: local effective degrees of freedom, or local effective dimensionality.

The idea is simple: hundreds of PRISM cells do not necessarily contain hundreds of independent precipitation signals. The real scientific question is how many genuinely different spatial patterns are present.

Physical grid-cell count

How many values PRISM stores.

Effective degrees of freedom

How many statistically distinct pieces of spatial information those values contain.

1. The central equation

Neff = n2 / Σj,k rjk2
n

Number of PRISM cells in the neighborhood.

rjk

Correlation between standardized precipitation anomalies at cells j and k.

Neff

Effective number of independent spatial modes.

Strong correlations make the denominator large, reducing Neff. Weak correlations make the denominator smaller, pushing Neff toward the physical cell count.

2. Interactive simplified neighborhood

This uses a simplified equal-correlation model where all off-diagonal cell pairs have the same correlation magnitude ρ. Real PRISM neighborhoods will have a full distribution of pairwise correlations.

Physical cells
40
Effective modes
1.54
Cells per effective mode
25.9

3. Two extreme examples

All 40 cells behave identically

rjk ≈ 1 for every pair
Σrjk2 ≈ 1600
Neff = 1600 / 1600 = 1

Forty stored values, approximately one independent signal.

All 40 cells behave independently

off-diagonal rjk ≈ 0
Σrjk2 ≈ 40
Neff = 1600 / 40 = 40

Forty stored values, approximately forty independent signals.

4. Eigenvalue interpretation

The same quantity can be written using the eigenvalues λ of the local correlation matrix:

Neff = (Σ λj)2 / Σ λj2

For a correlation matrix, Σλj = n. If most variance is concentrated in only a few eigenmodes, the effective dimensionality is far below the number of physical cells.

Conceptual eigenvalue spectrum

20, 10, 5, 3, 1, 1, 0, …

Most of the information is concentrated in a handful of spatial patterns.

5. The Gram-matrix computational trick

Let Z be the standardized anomaly matrix, with time in rows and neighboring PRISM cells in columns. For 2003–2010 monthly data, T = 96.

‖ZTZ‖F2 = ‖ZZT‖F2

The direct cell-by-cell matrix ZTZ is n × n. But the time-domain Gram matrix ZZT is only 96 × 96. This allows the Frobenius norm needed by the participation ratio to be computed without building a huge spatial correlation matrix.

Naive spatial matrix
200 × 200
Time-domain matrix
96 × 96

Because temporal demeaning removes one degree of freedom, the strict sample rank is generally no greater than roughly T−1, so 96 monthly observations provide at most about 95 resolvable temporal dimensions.

6. Scale dependence: Neff(R)

One Neff value is not enough. Compute it at multiple radii such as 16, 32, 64, and 128 km. The rate at which effective dimensionality grows with radius is itself a measure of spatial decorrelation.

7. The money test: Neff versus the sensor-count knee

For each county, build an error-versus-sensor-count curve. Error should fall rapidly at low K and then flatten. The bend is the knee, where additional sensors stop buying much improvement.

Neff → observed sensor-count knee → practical sensor requirement

8. What changes around 2002?

The useful question is no longer merely whether precipitation statistics changed. It becomes: did PRISM's effective spatial dimensionality change?

Pre-radar

Estimate local dimensionality before the production transition.

Transition

Test whether the break is sharp, gradual, or absent.

Radar era

Determine whether the later field contains more or fewer resolvable spatial modes.

Anti-circularity rule: if Neff itself is used to discover a breakpoint, the same series cannot be treated as independent proof that Neff changed at that breakpoint. Station-history or production-history evidence should provide independent segmentation.

9. Seasonality

Warm season

Localized convection may reduce correlation length and increase spatial fragmentation.

Hypothesis: Neff ↑ → required K ↑

Cool season

Synoptic and orographic precipitation may be broader and more spatially coherent.

Hypothesis: Neff ↓ → required K ↓

These are hypotheses to test, not assumptions to bake into the analysis.

10. A cleaner fingerprint model

Terrain / water / station support / climatology → Neff → sensor requirement → reconstruction error

This reframes the fingerprint program. Instead of directly regressing RMSE against dozens of county descriptors, ask whether physical and observational geography explain local dimensionality, and whether dimensionality predicts the scaling-curve knee.

11. Finite-sample bias

With only a finite number of months, unrelated cells still acquire accidental sample correlations. These inflate Σr² and bias Neff downward.

Known synthetic rankIllustrative estimateMeaning
54.8Small bias
109.0Moderate compression
2015.7Larger finite-T bias
4024.1High ranks difficult to resolve

These numbers are conceptual examples, not PRISM results. Actual calibration should use synthetic fields with known rank and realistic temporal structure, including AR(1)-matched or phase-randomized surrogates.

12. What Neff means, and what it does not

Legitimate interpretation
  • Effective dimensionality of the local monthly PRISM precipitation field.
  • Spatial redundancy versus complexity.
  • Candidate predictor of sensor-count knee.
  • A quantity to compare across scales, seasons, and production eras.
Not automatically justified
  • The number of rain gauges the real atmosphere requires.
  • The true dimensionality of precipitation independent of PRISM.
  • Exact PRISM station-to-cell interpolation weights.
  • A guaranteed one-for-one mapping between Neff and sensor count.
PRISM has already assimilated stations, terrain relationships, interpolation machinery, and later radar information. This measures PRISM's information structure. For reconstructing PRISM from subsets of PRISM cells, that is exactly the quantity of immediate interest.

13. Rotating validation architecture

Layer 1

PRISM production era: pre-radar, transition, radar-informed.

Layer 2

Station-support era: stable observational epochs defined from station history.

Layer 3

Blocked temporal folds: whole years or contiguous blocks rotated through development and evaluation.

Cross-validation asks whether a design generalizes to withheld periods or station regimes. Block bootstrap asks how uncertain the resulting metrics are. Those are related, but not interchangeable.

14. The decisive falsifiable test

Does local Neff predict the observed sensor-count knee?
If yes

Effective dimensionality becomes a credible organizing variable for sensor requirement, seasonality, regime change, quota design, and optimizer behavior.

If no

The theory fails cheaply, and the project avoids months of architecture built around a seductive but unsupported idea.

Bottom line

The physical number of PRISM cells is not the main quantity. The central question is how many independent spatial precipitation patterns those cells contain. Measure that locally, measure how it grows with spatial radius, test how it changes across seasons and production eras, and then ask whether it predicts where the sensor-count scaling curve reaches its knee.

If it does, effective dimensionality becomes the spine of the PRISM experiment.

← CPA Weather Lab · Learn