↖ CPA Weather Lab
READING GUIDE · HYDROLOGY / EXTREME PRECIPITATION

Where can a storm be moved to,
and still tell the truth?

How the SLAM test replaces expert judgment with a repeatable statistical procedure for drawing storm transposition domains — the geographic backbone of FEMA's next-generation flood risk models.

FitzGerald, Wright, Yan, Hendricks Dietrich & Sebastian (2025).
"An L-Moments-Based Hypothesis Test to Identify Homogeneous Storm Transposition Regions." SSRN preprint, not yet peer reviewed.

Borrowing storms from somewhere else

Stochastic storm transposition (SST) stretches a short rainfall record by borrowing real, gridded storms from a wider region and sliding them onto the watershed you actually care about.

That only works if the donor region — the transposition domain — genuinely behaves like the target watershed. Pick a domain that's too generous, and you're importing storms that don't belong. Pick one that's too narrow, and you throw away useful data.

Until now, that domain has mostly been drawn by hand: an expert weighs elevation, distance to moisture sources, and seasonal dewpoints, and sketches a boundary. Reasonable, but subjective — and FEMA needs this done, defensibly, for hundreds of watersheds nationwide as part of modernizing the National Flood Insurance Program.

Why the obvious statistical fix doesn't work

A well-established technique — regional L-moment homogeneity testing — already pools rainfall records across sites. It's even been used once before for storm transposition. So why not use it everywhere?

What that test asks

Are all these sites similar to each other, as a group? It also loses statistical power once you throw thousands of grid cells at it as "sites" — exactly the situation with modern high-resolution precipitation grids.

What SST actually needs

Is this candidate location similar to my one specific watershed? That's a region-of-influence problem — the same logic behind NOAA's Atlas 14 — not a group-homogeneity problem.

SLAM is that region-of-influence logic, rebuilt for gridded watershed-scale precipitation instead of point rain gauges.

Stop comparing totals. Compare shape.

SLAM's real contribution: instead of comparing how much rain falls on average, it compares the internal spatial texture of a typical extreme storm.

For a watershed, take every year's biggest rainfall event, keep the full grid-cell map (not just the average), and average those maps together across all years of record. The result is a composite annual maximum, or CAM — a single map of what a typical extreme storm looks like, spatially, in that watershed.

Flatten that map into a plain distribution of grid-cell values, and describe its shape with four L-moments — L-mean, L-CV, L-skewness, L-kurtosis. Now you're comparing storm texture, not just storm size.

Same total rain, different storm

Actual figure values from the paper — Kanawha's own composite storm vs. one example transposed location, 72-hour duration.

Kanawha watershed (actual)

L-mean68.04 mm
L-CV0.07
L-skewness+0.16
L-kurtosis0.22

Example transposed location

L-mean69.91 mm
L-CV0.05
L-skewness−0.06
L-kurtosis0.13

Nearly identical total intensity (L-mean) — but opposite-sign skewness. The watershed's storms concentrate rain into a smaller area with a long wet tail; the transposed location spreads it more evenly. A method comparing only totals would call these a match. SLAM doesn't.

What each L-moment is actually reading

Tap a card — these describe the spatial texture of one composite storm, cell to cell, not year-to-year variability.

From one watershed to a defensible boundary

Five stages, run in sequence. Step through them.

Peeling away false alarms, one at a time

Run several hundred thousand hypothesis tests, and thousands will come back "significant" by pure chance — the classic infinite-monkeys problem. SLAM controls this with false discovery rate (FDR) correction, but with a twist: it doesn't just run FDR once. It removes the single worst offending location, recomputes the threshold, and repeats — until every remaining location genuinely passes.

The chart below is an illustrative reconstruction of the shape the paper describes for the Kanawha basin (Figure 5) — not a pixel reproduction of their data.

Removed: 0 of ~120,000 candidate locations

— grey line: all candidate p-values, ranked smallest to largest. amber line: the adaptive FDR threshold. Locations above the threshold, at the far right, are removed each round — the threshold recalculates each time, letting the domain settle at its largest statistically defensible size. In the real Kanawha run, this took roughly 87,000 removals out of 120,000 candidates before the field stabilized.

Same method, four very different answers

The authors ran SLAM against four FEMA pilot watersheds, deliberately spanning very different terrain and storm climates. Switch between them.

A real fix, honestly stress-tested

The central idea — testing the spatial texture of a composite extreme storm instead of its basin-average total — is a clean, defensible reframing of a genuinely hard, previously subjective problem. The iterative FDR-peeling procedure is a real statistical contribution, not just an off-the-shelf correction. And releasing working code alongside a FEMA-facing method is a meaningfully higher bar than most hydrology papers clear.

Holds up
  • Objective, reproducible, and scalable in a way hand-drawn domains never were
  • Catches spatial-shape mismatches that basin-average methods miss entirely (Kanawha vs. its own transposed twin)
  • For Trinity, SLAM's domain fit the watershed's real rainfall record better than expert judgment did
  • Failure modes are reported openly, not smoothed over
Worth watching
  • Computationally heavy — days per watershed on a university HPC cluster
  • Validated on only four watersheds, however well-chosen
  • Breaks down precisely where FEMA needs it most: small, topographically extreme basins (Duwamish)
  • One unusually large storm can shift a domain boundary by ~150 miles with only ~43 years of record
← CPA Weather Lab · Learn