← SEAS5 Precipitation Alerts
In development

HDX signal for SEAS5 — data hand-over

Where every input for the signal lives (exact blob paths), how each was processed, what it covers, the caveats to read before drawing a conclusion, and how to refresh it after a new issuance. Written September 2026 for the analyst taking the signal design forward. The repo copy is docs/dev-notes/hdx-signal-data.md.

1. The signal, as currently specified

Per country, per issuance, per target trimester: issue a signal when at least a fraction frac of the country's admin-1 units each satisfy all three of

  1. SEAS5 forecasts the trimester at a return period ≥ rp years. Dry and wet are two separate signals, one per tail.
  2. The forecast has skill ≥ r_min for that unit × issue month × trimester (detrended Pearson r of the hindcast against ERA5).
  3. The trimester is in season for that unit: its share of the unit's annual rainfall is ≥ season_share.
frac
0.60
fraction of the country's admin-1 units that must qualify
rp
5 yr
Weibull return period of the forecast in the unit's own hindcast (dry_rp / wet_rp)
r_min
0.30
minimum detrended Pearson r; the app's "moderate skill" cut
season_share
0.25
trimester's share of annual rain. 0.25 = above a flat climatology. The app uses the looser 0.15.

All four are free. Everything they threshold is a column in the input tables, so the whole space can be swept without touching a pipeline. Ethiopia, or any country with a bimodal climatology, can be run at admin-2 instead: the admin-2 table has identical columns.

The starting rule is implemented once, in src/hdx_signal.py (unit_condition, country_signal); see section 6.

2. Where the raw numbers come from

Nothing here touches rasters. Everything starts from the zonal statistics that the team's ds-raster-stats Databricks jobs write to the prod rasterstats Postgres DB (chd-rasterstats-prod.postgres.database.azure.com, reached with ocha_stratus.get_engine("prod")).

TableWhatKey columnsRecord
public.seas5 ECMWF SEAS5 ensemble-mean monthly precipitation, zonal mean per admin unit, one row per issued month × valid month pcode, iso3, adm_level, issued_date, valid_date, leadtime (0–6), mean + median/min/max/count/sum/std issued 1981-01 → 2026-09 (5th of each month; hindcast 1981–2016 + operational)
public.era5 ERA5 monthly precipitation, zonal mean per admin unit pcode, iso3, adm_level, valid_date, mean 1981-01 → 2026-08 (lands about the 6th of the next month)
public.polygon The admin units (COD boundaries via fieldmaps.io): name, level, area, pixel-coverage counts pcode, iso3, adm_level, name, seas5_n_intersect_raw_pixels, … admin 0/1 for 153 countries; admin 2 for a subset

Units are mm/day (monthly mean rate). leadtime 0 is the issue month itself; SEAS5's horizon is 7 monthly steps. Units smaller than one SEAS5 pixel (about 0.4°) get a NULL mean and are dropped downstream. The upstream job is documented in the team knowledge base at pipelines/raster-stats.md.

Admin-1 pcodes are the COD ones (ET01 Tigray, NE005 Tahoua). Admin-0 pcodes are mostly the ISO3 but not always: join on iso3 at country level and on pcode below it.

3. What this repo computes from them

src/skill.py:run_all_combinations runs once per admin unit, for all 12 issue months × 12 trimesters = 144 combinations:

  1. Aggregate to trimester means. SEAS5: the three valid months of one issuance, averaged. ERA5: the same three months, complete seasons only. For in-season trimesters (issue month inside the trimester, leads −1 and −2) the months already elapsed at issuance come from ERA5 and only the rest from SEAS5, each forecast month bias-corrected per calendar month. That is why in-season skill is high by construction.
  2. log1p, then normalise SEAS5 to ERA5's mean and standard deviation over the overlap years, so forecast and observation share a scale.
  3. Detrend both series (linear fit over the overlap, re-centred). The detrended variant is what the site, the HNRP tab and the signal tables use. The raw variant is in the blob too.
  4. Skill = Pearson r between the normalised forecast and ERA5 over the overlap years (n_years, typically 45–46; NaN below 10).
  5. Position of a forecast in its own hindcast: percentile (pct, share of hindcast forecasts ≤ it) and two Weibull return periods, dry_rp = (n+1)/rank with rank counted from the dry end and wet_rp from the wet end. The hindcast distribution is the years with both a forecast and an observation; the live forecast is ranked against it and past years are ranked in-sample. Both describe how unusual the forecast is, not how much rain is missing.

Only trimesters at signed lead −2 … 4 are kept (complete trimesters with at least one forecast month inside the horizon): 7 trimesters per issuance.

The in-season rule (src/season.py) is separate: the ERA5 monthly climatology per unit (mean of every January, every February, … over 1981–present), the trimester's share 3 × mean(3 months) / sum(12 months), thresholded.

4. Files in the blob

Storage account imb0chd0dev (the team's dev stage), container projects. All paths below are under ds-seas5-skill/processed/. Everything was rebuilt for the September 2026 issuance (2026-09-07; Ethiopia admin-2 and the climatologies on 2026-09-15; the signal tables on 2026-09-24, with forecast_mm / hist_mean_mm).

4a. Ready-made signal inputs (start here)

PathLevelRowsUnitsCountries
hdx_signal/signal_inputs_adm1.parquetadmin-1, all issuances 1981 → 2026-098,741,5352,283150
hdx_signal/signal_inputs_adm1_latest.parquetadmin-1, current issuance only15,9352,283150
hdx_signal/units_adm1.parquetone row per admin-1 unit with a climatology2,307152
hdx_signal/signal_inputs_adm2.parquetadmin-2, all issuances19,704,7955,14624
hdx_signal/signal_inputs_adm2_latest.parquetadmin-2, current issuance35,9205,14624
hdx_signal/units_adm2.parquetone row per admin-2 unit5,14624
hdx_signal/signal_inputs.parquet, …_latest.parquet, hdx_signal/units.parquetadmin-0, for comparison576,000 / 1,050 / 152150150

One row per unit × issuance (year, month) × trimester, built by pipeline/build_hdx_signal_inputs.py from the three products in 4b.

ColumnMeaning
pcode, iso3, name, adm_levelthe unit (name from public.polygon)
issued_year, issued_monththe SEAS5 issuance (first of the month)
trimesterJFM … DJF (src/constants.py:TRIMESTERS)
leadsigned months from issue month to the trimester's first month, −2 … 4. Negative = in season (blend of observed and forecast months)
season_yearthe year the trimester is anchored on (NDJ and DJF anchor on December's year)
pearson_r, n_yearsdetrended hindcast skill of this unit × issue month × trimester; NaN = no skill available
forecast_mean_log, obs_mean_logthe normalised, detrended trimester means in log1p(mm/day) space; obs_mean_log is NaN for the live forecast
in_sampleTrue for hindcast years (an observation exists), False for the live forecast
pctforecast percentile within the unit's hindcast forecasts (0 driest ever forecast, 100 wettest)
dry_rp, wet_rpWeibull return period of the forecast from the dry and the wet end; maximum n+1, about 47
forecast_mmthe row's forecast as a trimester total in mm: expm1(forecast_mean_log) (mm/day, normalised to ERA5 and detrended, the value the RP is computed from), clipped at 0, × the trimester's calendar days (TRIMESTER_DAYS, non-leap, 89–92; Feb trimesters are 1 day short in leap years, about 1 %)
hist_mean_mmthe "normal" to read forecast_mm against: expm1 of the mean of obs_mean_log over every observed year (1981→) of this unit × issue month × trimester, clipped at 0, × the same days. This is era5_mean of the skill file, so the number is identical to the "normal" the Forecast × HNRP tab shows next to the same forecast. A log-space mean, not the arithmetic mean of the mm values, which sits above the median for skewed rain and would read a median forecast (pct 50) as below normal. Constant across the years of a combo; NaN only where the combo has no hindcast (pct / RPs NaN too)
tri_share_annualthe trimester's share of the unit's annual ERA5 rainfall (climatology)
tri_mean_mm_daythe trimester's climatological mean, mm/day
in_season_flattri_share_annual ≥ 0.25, the starting rule
in_season_apptri_share_annual ≥ 0.15, what the app and site call "rainy"

units_adm*.parquet: pcode, iso3, name, adm_level, n_units_in_country, has_skill. Use n_units_in_country as the denominator for frac when a unit drops out for lack of data.

4b. The processed products the signal tables are built from

PathLevelContent
skill_stats_detrended.parquet
skill_stats_detrended_adm1.parquet
skill_stats_detrended_adm2.parquet
0 / 1 / 2 one row per unit × issue month × trimester: pearson_r, n_years, rmse, sigma, era5_mean, era5_std, lower_tercile_mm, current_forecast_year, current_forecast_mean, is_predictive, prob_lower_tercile, forecast_rp (dry), flood_rp (wet), prob_rp, forecast_percentile. Latest forecast only. 1.2 / 17 / 36 MB
paired_yearly_detrended.parquet
…_adm1.parquet
…_adm2.parquet
0 / 1 / 2 one row per unit × issue month × trimester × year: season_year, forecast_mean, obs_mean, hist_prob (log space). The full hindcast; everything per-year derives from it. 17 / 258 / 562 MB
skill_stats*.parquet, paired_yearly*.parquet without _detrended0 / 1 / 2the same, raw variant
monthly_clim.parquet
monthly_clim_adm1.parquet
monthly_clim_adm2.parquet
0 / 1 / 2 ERA5 monthly climatology: pcode, iso3, name, adm_level, month, mean_mm_day, 12 rows per unit

Producers, in order: pipeline/compute_skill.py (admin-0, about 30 min, also run by the monthly cron), compute_skill_adm1.py (about 40 min), compute_skill_adm2.py (scope = ADM2_ISO3S in that file), compute_monthly_clim.py --level N, build_hdx_signal_inputs.py --level N.

Other things under processed/ (raster/, enso_slides/, uga/, stories/, cma/, forecast_site.parquet, pop_adm3_worldpop.parquet) belong to other products of this repo. rainy_pairs.parquet and roc_auc_stats.parquet (May 2026) are orphans of an earlier iteration; nothing reads them.

5. Loading

import ocha_stratus as stratus

P = "ds-seas5-skill/processed/hdx_signal/"
latest = stratus.load_parquet_from_blob(P + "signal_inputs_adm1_latest.parquet", stage="dev")
units  = stratus.load_parquet_from_blob(P + "units_adm1.parquet", stage="dev")

The full-history admin-1 table is about 9 M rows (fine in pandas, about 1.5 GB). The admin-2 one is about 20 M rows: read it through Arrow with column pruning and a country filter.

import io, pyarrow.parquet as pq
cc = stratus.get_container_client(stage="dev", container_name="projects")
buf = io.BytesIO(cc.download_blob(P + "signal_inputs_adm2.parquet").readall())
eth = pq.read_table(buf, filters=[("iso3", "=", "ETH")],
                    columns=["pcode", "issued_year", "issued_month", "trimester", "lead",
                             "pearson_r", "dry_rp", "wet_rp", "tri_share_annual"]).to_pandas()

Credentials: the usual DSCI_AZ_BLOB_DEV_SAS (read) in the environment. The DB is only needed if you go back to public.seas5 / public.era5 yourself.

6. Evaluating the signal

from src.hdx_signal import country_signal, unit_condition

dry = country_signal(latest, direction="dry", frac=0.60, rp=5, r_min=0.30, season_share=0.25)
wet = country_signal(latest, direction="wet")            # same defaults
dry[dry["signal"]][["iso3", "trimester", "lead", "n_units", "n_qualifying", "frac_qualifying"]]

country_signal returns one row per country × issuance × trimester with n_units, n_qualifying, frac_qualifying, signal. Two denominators are offered: "all" (every unit with data, the literal "fraction of the admin-1s") and "in_season" (only units for which the trimester is in season, so a country's arid half cannot dilute its rainy half). Which one the signal should use is an open design choice. The same call on the full-history table backtests every issuance since 1981 (hindcast RPs are in-sample, see section 7).

To sweep, the four parameters are plain keyword arguments; the per-row condition unit_condition(df, direction, rp, r_min, season_share) is there for a different roll-up, for example population-weighted.

Smoke test on the current issuance, admin-1, starting parameters: dry Niger fires on JAS and ASO (8 of 8 units), Ethiopia on JAS (12 of 13) and ASO (11 of 13); wet fires nowhere in either. Read the first caveat below before taking that at face value.

7. Coverage and caveats

The September 2026 issuance sits at hindcast extremes almost everywhere: JAS and ASO (in-season, leads −2/−1) come out at percentile 0 / RP about 47 across the whole Sahel and much of East Africa. With the starting parameters that fires a dry signal for Niger, Ethiopia and most of the Sahel on JAS/ASO. We believe this is at least partly an artefact of the in-season blend (an ERA5/SEAS5 mismatch in July–August 2026), not an assessed drought. Treat the in-season leads of this issuance as suspect and do not quote them publicly. A full-forecast lead (0–4) signal is on much firmer ground, and backtests over earlier issuances are unaffected.

8. Refreshing after a new issuance

SEAS5 arrives on the 5th; ERA5 for the previous month around the 6th. The monthly cron (.github/workflows/monthly-refresh.yml, 7th 03:00 UTC) refreshes admin-0 only: the skill stats, then build_hdx_signal_inputs.py --level 0, so signal_inputs{,_latest}.parquet and units.parquet follow the country data automatically (verify_site_data.py --check-signal fails the run if they do not). The admin-1/2 tables are manual, from the repo (uv sync first; needs the blob write SAS and prod DB read credentials):

uv run python pipeline/compute_skill_adm1.py                 # ~40 min, 8 workers
uv run python pipeline/compute_skill_adm2.py                 # all ADM2_ISO3S
for L in 1 2; do uv run python pipeline/build_hdx_signal_inputs.py --level $L; done
# only if units were added or changed (new boundaries, new admin-2 country):
for L in 0 1 2; do uv run python pipeline/compute_monthly_clim.py --level $L; done

Adding a country at admin-2: add its ISO3 to ADM2_ISO3S in compute_skill_adm2.py, then compute_skill_adm2.py --iso3 XXX (merges into the existing files), compute_monthly_clim.py --level 2, build_hdx_signal_inputs.py --level 2.

Related: the app's Forecast × HNRP tab uses the same skill files with the looser 0.15 in-season rule. Method code: src/skill.py. Source: OCHA-DAP/ds-seas5-skill.