MADIS Pennsylvania Surface Network and Acquisition Study
Snapshot dated 22 August 2026 · fixed terrestrial surface observations · integrated mesonet, METAR, and RWIS
1. Current network makeup
The accompanying workbook contains every station row, identifier/abbreviation, name, coordinates, elevation, MADIS dataset, provider network, station type, variables observed, and observed modal time interval. The inventory combines seven public archive hours on 19 August 2026 (00, 01, 02, 03, 06, 12, and 18 UTC). A station row is unique by dataset, provider, identifier, and coordinates.
| Geography | Stations | Definition |
|---|---|---|
| PA | 821 | Inside Census state polygon |
| 0–25 km outside PA | 435 | Exclusive ring |
| 25–50 km | 388 | Exclusive ring |
| 50–75 km | 420 | Exclusive ring |
| 75–100 km | 498 | Exclusive ring |
| PA + 100 km | 2,562 | Cumulative footprint |
Dataset rows: 2,299 integrated mesonet, 128 METAR, and 135 RWIS. Largest provider labels are APRSWXNET (1,171), MesoWest (521), HADS (338), OHDOT (90), METAR (84), PADOT (75), RAWS (60), NJWxNet (54), DEOS (37), and NonFedAWOS (37). Most common observed cadences are 5 minutes (1,004 rows), 15 minutes (587), 60 minutes (370), and 10 minutes (334).
2. How far back can data be requested?
The public archive directory begins on 1 July 2001. That is directory coverage, not a guarantee that every dataset or provider exists continuously from that date; early years have fewer products and there are gaps. The attached surface CGI can return some historical point requests, but tests were inconsistent for arbitrary archive dates, so it should not be treated as the authoritative deep-archive extraction path. Use the archive directory for reproducible backfills and inventory each requested hour.
3–5. What can be filtered, and what is the file granularity?
- Deep archive: public files are organized by date, dataset, and hour. There is not one universal MADIS file: mesonet, METAR, RWIS, hydro, and other families have separate hourly files.
- Local MADIS API: the Fortran API filters by geographic domain, station/provider, variables, time boundaries, and QC after the netCDF file is local. It is a reader library, not a remote historical subset API.
- On-demand surface CGI: supports a latitude/longitude bounding box, state, station name/ID, provider, variables, and QC selection. The total backward-plus-forward window must be no more than 119 minutes. Output can be text/XML-style or CSV-like, depending on parameters.
- Multi-day requests: NOAA supplies individual hourly archive files. Archive scripts can loop over a day or date range, but there is no documented server-side daily/weekly bundle or deep-archive polygon package.
Useful CGI requirements include time=YYYYMMDD_HHMM (or 0 for current), minbck, minfwd, a selection mode (dfltrsel), then either latll/lonll/latur/lonur, state, or stanam. Provider, variable, and QC selectors are optional. Treat longitude signs carefully (west is negative).
6. Nationwide compressed download size
| Dataset | Mean/hour | Per day | Per year |
|---|---|---|---|
| Integrated mesonet | 34.35 MB | 0.824 GB | 301.12 GB |
| METAR | 0.87 MB | 0.021 GB | 7.59 GB |
| RWIS | 1.09 MB | 0.026 GB | 9.54 GB |
| Three-dataset surface total | 36.31 MB | 0.871 GB | 318.25 GB |
Measured from seven 2026-08-19 compressed files per dataset. Values are decimal MB/GB; archive volume varies by hour, season, provider participation, and schema.
7. Estimated retained regional size
| Scope | Hour MB | Day MB | Week MB | Month MB | Year GB |
|---|---|---|---|---|---|
| PA | 0.85 | 20.33 | 142.29 | 618.72 | 7.42 |
| OH | 0.61 | 14.70 | 102.92 | 447.53 | 5.37 |
| WV | 0.36 | 8.74 | 61.21 | 266.16 | 3.19 |
| VA | 0.26 | 6.18 | 43.24 | 188.03 | 2.26 |
| MD | 0.49 | 11.74 | 82.17 | 357.28 | 4.29 |
| DE | 0.31 | 7.51 | 52.57 | 228.57 | 2.74 |
| NJ | 0.41 | 9.80 | 68.62 | 298.38 | 3.58 |
| NY | 0.60 | 14.36 | 100.53 | 437.14 | 5.25 |
| CT | 0.20 | 4.75 | 33.22 | 144.43 | 1.73 |
| Nine-state total | 4.09 | 98.11 | 686.77 | 2,986.24 | 35.83 |
| PA + 25 km | 1.36 | 32.54 | 227.76 | 990.37 | 11.88 |
| PA + 50 km | 1.77 | 42.44 | 297.10 | 1,291.87 | 15.50 |
| PA + 75 km | 2.24 | 53.66 | 375.59 | 1,633.14 | 19.60 |
| PA + 100 km | 2.77 | 66.41 | 464.87 | 2,021.33 | 24.26 |
Method: national compressed bytes × the region's share of observations, summed across mesonet, METAR, and RWIS. Month = 365.25/12 days. Because regional recompression and final schema differ, plan with ±30%. Exact byte and GiB columns are in the workbook.
8. Cost- and time-effective acquisition design
Recommended: ephemeral Virginia cloud worker
- Rent a 4–8 vCPU, 8–16 GB RAM VM with 160–240 GB local NVMe in Ashburn, Virginia. Hetzner Cloud is a simple value option; AWS Spot in us-east-1 is a robust alternative for an interruption-safe queue.
- Generate the archive URL list and download only the required datasets. Run 4–8 concurrent hourly jobs, with retries, checksums, and an idempotent manifest.
- For each file: decompress to local NVMe, apply the PA polygon/buffer or state mask, normalize time/provider/station fields, write date-partitioned Parquet (or Zarr if multidimensional arrays are essential), verify output, then delete raw and temporary files.
- Upload retained partitions and manifests to Google Drive using
rclone. Do not mount Drive as the processing filesystem; its latency and API behavior make it a poor scratch disk. - Benchmark one day, then one week, before a multi-year run. Keep only 1–4 hours in the bounded work queue, so hundreds of GB of permanent cloud disk are unnecessary.
At 200 Mbit/s, 318 GB is about 3.5 hours of transfer; at 500 Mbit/s, about 1.4 hours; at 1 Gbit/s, about 0.7 hour, before decompression and filtering. A reasonable planning window is 4–8 VM-hours per archive year with 4–8 workers, but the one-day benchmark should set the real concurrency.
Cost model: VM-hours + temporary disk prorated + outbound traffic for retained output. NOAA-to-cloud ingress is normally free. Check the vendor's live price calculator because prices and included traffic vary by location and change over time. AWS Spot can be substantially cheaper but may be interrupted, so every hour must be restartable. Avoid long-lived object storage unless you intentionally keep raw files.
Your 4 TB Drive is ample for final output: about 165 years of PA+100 km at the current 24.26 GB/year estimate, or about 112 years of the nine-state footprint at 35.83 GB/year, before growth and overhead. It would hold only about 12.6 years of all three nationwide compressed surface datasets. Google documents a 750 GB per-user upload/copy limit per 24 hours, which is not restrictive for the retained regional outputs.
Method, limitations, and next decision
The best next step is to choose mesonet only versus mesonet + METAR + RWIS, the archive years, and whether every raw variable must be preserved. Those decisions determine whether the retained estimate is closer to the table above or materially smaller. Before production, run a stratified benchmark (winter/summer, weekday/weekend) and compare actual Parquet output to the ±30% planning band.
Primary sources
- NOAA MADIS data distribution overview
- NOAA MADIS API documentation and archive scripts
- NOAA MADIS public archive root
- Google Drive shared-drive limits
- Hetzner Cloud locations
- AWS EC2 Spot pricing
Prepared from NOAA files and the supplied MADIS API documentation/explorer materials. Snapshot timestamp: 2026-08-22 UTC.