↖ CPA Weather Lab

The Kriging Emulator

Henriksen, Hubrechts, Møller, Knudsen & Pedersen — Journal of Hydrology, 2024 · interactive companion

This paper solves one specific bottleneck: kriging-based rain gauge network design doesn't scale. Testing whether a new gauge helps requires inverting a covariance matrix — do that for tens of thousands of candidate locations, hundreds of times over, and a national-scale optimization becomes a multi-month computation. Their fix is a cheap, fitted stand-in for kriging that runs 3000× faster.

291
existing DMI gauges
175
new gauges optimized
46,797
candidate grid cells
3000×
emulator speedup
125 days → ~1 hr
optimization runtime

1 · The pipeline

Three stages, each depending on the last:

Fit an ordinary kriging variogram
(with a coastal–inland distance term)
→
Fit a cheap emulator
to approximate kriging's output
→
Run a greedy + resubstitution heuristic
to place many new gauges

2 · What the emulator actually predicts

The emulator's job: given a candidate gauge location, predict how much the kriging standard deviation will drop at some nearby grid cell — without rerunning kriging. It takes three inputs: the distance x between the candidate and the cell, the cell's current uncertainty y, and the uncertainty z right at the candidate site itself. The fitted form is a log term (capturing behavior at zero distance) times an exponential decay (capturing falloff with distance):

gain(x, y, z) = a₁ · log(a₂z + 1) · exp (−cx · (z−y+b₂) / (z+b₁)ᵈ)

Drag the sliders below to see how the predicted "gain" curve — the drop in standard deviation as a function of distance from the new gauge — changes with the two uncertainty inputs. This uses illustrative parameter values in the same shape family as the paper's fitted model, not the exact fitted Danish coefficients.

4.0
12.0

Higher z (a candidate placed in a very under-observed area) produces a bigger, longer-reaching gain curve. Placing a gauge right where uncertainty is already low (small z) barely moves the needle anywhere.

3 · Why greedy placement alone isn't enough

Placing gauges one at a time, always picking whichever single location currently helps most, is fast — but it's short-sighted. The paper's toy example: three gauges going into a triangular region. A greedy algorithm plants the first gauge dead center (highest initial uncertainty), then two more near two of the corners. A more symmetric layout, spread toward all three corners, likely serves the whole triangle better — but the greedy process can never see that, because it never revisits its first choice.

Greedy: gauge 1 (bright) goes to the centroid — the point of highest uncertainty. Gauges 2 and 3 then go to two corners, leaving the third corner under-served.

Resubstitution fixes this cheaply: once all k gauges are placed, remove and reinsert them one at a time, in their original order, letting each one relocate given that the others are now fixed in place. The paper finds one round captures most of the benefit — the marginal gain of a 2nd through 7th round is small. The effect only really matters once several gauges are being placed near each other; for isolated placements, greedy is already close to optimal.

4 · Diminishing returns, at national scale

Running the full pipeline on Denmark — filling gaps around the already-dense urban gauge clusters — produces the marginal-return curve stakeholders actually want to see: the first new gauges close the biggest gaps, and each subsequent gauge buys less. That curve is the real deliverable — it lets DMI decide where the investment stops paying off, rather than being handed one fixed "you need N gauges" number.

Illustrative shape of the reported relationship: national average kriging standard deviation vs. number of new gauges added, showing steep early gains and a flattening tail.

← CPA Weather Lab · Learn