Where every input for the signal lives (exact blob paths), how each was
processed, what it covers, the caveats to read before drawing a conclusion, and how to
refresh it after a new issuance. Written September 2026 for the analyst taking the signal
design forward. The repo copy is
docs/dev-notes/hdx-signal-data.md.
Per country, per issuance, per target trimester: issue a signal when at least a
fraction frac of the country's admin-1 units each satisfy all three of
rp years.
Dry and wet are two separate signals, one per tail.r_min for that unit × issue month ×
trimester (detrended Pearson r of the hindcast against ERA5).season_share.dry_rp / wet_rp)All four are free. Everything they threshold is a column in the input tables, so the whole space can be swept without touching a pipeline. Ethiopia, or any country with a bimodal climatology, can be run at admin-2 instead: the admin-2 table has identical columns.
The starting rule is implemented once, in src/hdx_signal.py
(unit_condition, country_signal); see section 6.
Nothing here touches rasters. Everything starts from the zonal statistics
that the team's ds-raster-stats Databricks jobs write to the prod rasterstats
Postgres DB (chd-rasterstats-prod.postgres.database.azure.com, reached with
ocha_stratus.get_engine("prod")).
| Table | What | Key columns | Record |
|---|---|---|---|
public.seas5 |
ECMWF SEAS5 ensemble-mean monthly precipitation, zonal mean per admin unit, one row per issued month × valid month | pcode, iso3, adm_level, issued_date, valid_date, leadtime (0–6), mean + median/min/max/count/sum/std |
issued 1981-01 → 2026-09 (5th of each month; hindcast 1981–2016 + operational) |
public.era5 |
ERA5 monthly precipitation, zonal mean per admin unit | pcode, iso3, adm_level, valid_date, mean |
1981-01 → 2026-08 (lands about the 6th of the next month) |
public.polygon |
The admin units (COD boundaries via fieldmaps.io): name, level, area, pixel-coverage counts | pcode, iso3, adm_level, name, seas5_n_intersect_raw_pixels, … |
admin 0/1 for 153 countries; admin 2 for a subset |
Units are mm/day (monthly mean rate). leadtime 0 is the issue
month itself; SEAS5's horizon is 7 monthly steps. Units smaller than one SEAS5 pixel (about
0.4°) get a NULL mean and are dropped downstream. The upstream job is documented in the
team knowledge base at pipelines/raster-stats.md.
Admin-1 pcodes are the COD ones (ET01 Tigray, NE005 Tahoua).
Admin-0 pcodes are mostly the ISO3 but not always: join on iso3 at country
level and on pcode below it.
src/skill.py:run_all_combinations runs once per admin unit, for all 12 issue
months × 12 trimesters = 144 combinations:
n_years, typically 45–46; NaN below 10).pct,
share of hindcast forecasts ≤ it) and two Weibull return periods,
dry_rp = (n+1)/rank with rank counted from the dry end and wet_rp
from the wet end. The hindcast distribution is the years with both a forecast and an
observation; the live forecast is ranked against it and past years are ranked in-sample.
Both describe how unusual the forecast is, not how much rain is missing.Only trimesters at signed lead −2 … 4 are kept (complete trimesters with at least one forecast month inside the horizon): 7 trimesters per issuance.
The in-season rule (src/season.py) is separate: the ERA5 monthly climatology
per unit (mean of every January, every February, … over 1981–present), the trimester's share
3 × mean(3 months) / sum(12 months), thresholded.
Storage account imb0chd0dev (the team's dev stage), container
projects. All paths below are under
ds-seas5-skill/processed/. Everything was rebuilt for the
September 2026 issuance (2026-09-07; Ethiopia admin-2 and the climatologies on 2026-09-15; the
signal tables on 2026-09-24, with forecast_mm / hist_mean_mm).
| Path | Level | Rows | Units | Countries |
|---|---|---|---|---|
hdx_signal/signal_inputs_adm1.parquet | admin-1, all issuances 1981 → 2026-09 | 8,741,535 | 2,283 | 150 |
hdx_signal/signal_inputs_adm1_latest.parquet | admin-1, current issuance only | 15,935 | 2,283 | 150 |
hdx_signal/units_adm1.parquet | one row per admin-1 unit with a climatology | 2,307 | 152 | |
hdx_signal/signal_inputs_adm2.parquet | admin-2, all issuances | 19,704,795 | 5,146 | 24 |
hdx_signal/signal_inputs_adm2_latest.parquet | admin-2, current issuance | 35,920 | 5,146 | 24 |
hdx_signal/units_adm2.parquet | one row per admin-2 unit | 5,146 | 24 | |
hdx_signal/signal_inputs.parquet, …_latest.parquet, hdx_signal/units.parquet | admin-0, for comparison | 576,000 / 1,050 / 152 | 150 | 150 |
One row per unit × issuance (year, month) × trimester, built by
pipeline/build_hdx_signal_inputs.py from the three products in 4b.
| Column | Meaning |
|---|---|
pcode, iso3, name, adm_level | the unit (name from public.polygon) |
issued_year, issued_month | the SEAS5 issuance (first of the month) |
trimester | JFM … DJF (src/constants.py:TRIMESTERS) |
lead | signed months from issue month to the trimester's first month, −2 … 4. Negative = in season (blend of observed and forecast months) |
season_year | the year the trimester is anchored on (NDJ and DJF anchor on December's year) |
pearson_r, n_years | detrended hindcast skill of this unit × issue month × trimester; NaN = no skill available |
forecast_mean_log, obs_mean_log | the normalised, detrended trimester means in log1p(mm/day) space; obs_mean_log is NaN for the live forecast |
in_sample | True for hindcast years (an observation exists), False for the live forecast |
pct | forecast percentile within the unit's hindcast forecasts (0 driest ever forecast, 100 wettest) |
dry_rp, wet_rp | Weibull return period of the forecast from the dry and the wet end; maximum n+1, about 47 |
forecast_mm | the row's forecast as a trimester total in mm: expm1(forecast_mean_log) (mm/day, normalised to ERA5 and detrended, the value the RP is computed from), clipped at 0, × the trimester's calendar days (TRIMESTER_DAYS, non-leap, 89–92; Feb trimesters are 1 day short in leap years, about 1 %) |
hist_mean_mm | the "normal" to read forecast_mm against: expm1 of the mean of obs_mean_log over every observed year (1981→) of this unit × issue month × trimester, clipped at 0, × the same days. This is era5_mean of the skill file, so the number is identical to the "normal" the Forecast × HNRP tab shows next to the same forecast. A log-space mean, not the arithmetic mean of the mm values, which sits above the median for skewed rain and would read a median forecast (pct 50) as below normal. Constant across the years of a combo; NaN only where the combo has no hindcast (pct / RPs NaN too) |
tri_share_annual | the trimester's share of the unit's annual ERA5 rainfall (climatology) |
tri_mean_mm_day | the trimester's climatological mean, mm/day |
in_season_flat | tri_share_annual ≥ 0.25, the starting rule |
in_season_app | tri_share_annual ≥ 0.15, what the app and site call "rainy" |
units_adm*.parquet: pcode, iso3, name, adm_level, n_units_in_country,
has_skill. Use n_units_in_country as the denominator for frac
when a unit drops out for lack of data.
| Path | Level | Content |
|---|---|---|
skill_stats_detrended.parquetskill_stats_detrended_adm1.parquetskill_stats_detrended_adm2.parquet | 0 / 1 / 2 | one row per unit × issue month × trimester: pearson_r, n_years, rmse, sigma, era5_mean, era5_std, lower_tercile_mm, current_forecast_year, current_forecast_mean, is_predictive, prob_lower_tercile, forecast_rp (dry), flood_rp (wet), prob_rp, forecast_percentile. Latest forecast only. 1.2 / 17 / 36 MB |
paired_yearly_detrended.parquet…_adm1.parquet…_adm2.parquet | 0 / 1 / 2 | one row per unit × issue month × trimester × year: season_year, forecast_mean, obs_mean, hist_prob (log space). The full hindcast; everything per-year derives from it. 17 / 258 / 562 MB |
skill_stats*.parquet, paired_yearly*.parquet without _detrended | 0 / 1 / 2 | the same, raw variant |
monthly_clim.parquetmonthly_clim_adm1.parquetmonthly_clim_adm2.parquet | 0 / 1 / 2 | ERA5 monthly climatology: pcode, iso3, name, adm_level, month, mean_mm_day, 12 rows per unit |
Producers, in order: pipeline/compute_skill.py (admin-0, about 30 min, also
run by the monthly cron), compute_skill_adm1.py (about 40 min),
compute_skill_adm2.py (scope = ADM2_ISO3S in that file),
compute_monthly_clim.py --level N, build_hdx_signal_inputs.py --level N.
Other things under processed/ (raster/, enso_slides/,
uga/, stories/, cma/, forecast_site.parquet,
pop_adm3_worldpop.parquet) belong to other products of this repo.
rainy_pairs.parquet and roc_auc_stats.parquet (May 2026) are orphans
of an earlier iteration; nothing reads them.
import ocha_stratus as stratus P = "ds-seas5-skill/processed/hdx_signal/" latest = stratus.load_parquet_from_blob(P + "signal_inputs_adm1_latest.parquet", stage="dev") units = stratus.load_parquet_from_blob(P + "units_adm1.parquet", stage="dev")
The full-history admin-1 table is about 9 M rows (fine in pandas, about 1.5 GB). The admin-2 one is about 20 M rows: read it through Arrow with column pruning and a country filter.
import io, pyarrow.parquet as pq
cc = stratus.get_container_client(stage="dev", container_name="projects")
buf = io.BytesIO(cc.download_blob(P + "signal_inputs_adm2.parquet").readall())
eth = pq.read_table(buf, filters=[("iso3", "=", "ETH")],
columns=["pcode", "issued_year", "issued_month", "trimester", "lead",
"pearson_r", "dry_rp", "wet_rp", "tri_share_annual"]).to_pandas()
Credentials: the usual DSCI_AZ_BLOB_DEV_SAS (read) in the environment. The DB
is only needed if you go back to public.seas5 / public.era5
yourself.
from src.hdx_signal import country_signal, unit_condition dry = country_signal(latest, direction="dry", frac=0.60, rp=5, r_min=0.30, season_share=0.25) wet = country_signal(latest, direction="wet") # same defaults dry[dry["signal"]][["iso3", "trimester", "lead", "n_units", "n_qualifying", "frac_qualifying"]]
country_signal returns one row per country × issuance × trimester with
n_units, n_qualifying, frac_qualifying,
signal. Two denominators are offered: "all" (every unit with data,
the literal "fraction of the admin-1s") and "in_season" (only units for which
the trimester is in season, so a country's arid half cannot dilute its rainy half). Which
one the signal should use is an open design choice. The same call on the full-history table
backtests every issuance since 1981 (hindcast RPs are in-sample, see section 7).
To sweep, the four parameters are plain keyword arguments; the per-row condition
unit_condition(df, direction, rp, r_min, season_share) is there for a different
roll-up, for example population-weighted.
Smoke test on the current issuance, admin-1, starting parameters: dry Niger fires on JAS and ASO (8 of 8 units), Ethiopia on JAS (12 of 13) and ASO (11 of 13); wet fires nowhere in either. Read the first caveat below before taking that at face value.
public.polygon is present except Sint Maarten. 23 units have
no ERA5 rows either, hence no climatology (tri_share_annual NaN; they never
qualify). Admin-2 skill exists for 24 countries: the 23 HNRP/IPC-scope countries with
admin-2 boundaries in the DB plus Ethiopia (92 zones), added for this work. Kenya, for
instance, has no admin-2 in public.polygon; extending admin-2 means first
getting boundaries and zonal stats into the rasterstats DB, then adding the ISO3 to
ADM2_ISO3S and running compute_skill_adm2.py --iso3 XXX.The September 2026 issuance sits at hindcast extremes almost everywhere: JAS and ASO (in-season, leads −2/−1) come out at percentile 0 / RP about 47 across the whole Sahel and much of East Africa. With the starting parameters that fires a dry signal for Niger, Ethiopia and most of the Sahel on JAS/ASO. We believe this is at least partly an artefact of the in-season blend (an ERA5/SEAS5 mismatch in July–August 2026), not an assessed drought. Treat the in-season leads of this issuance as suspect and do not quote them publicly. A full-forecast lead (0–4) signal is on much firmer ground, and backtests over earlier issuances are unaffected.
lead ≥ 0
or at least report leads separately.build_hdx_signal_inputs.py:position_metrics if a backtest
needs it.issued_year on the _latest file and that its
in-season rows have in_sample == False.forecast_mm and hist_mean_mm (trimester totals, the
same values as the HNRP tab's "forecast vs normal") pair the anomaly with an amount;
tri_mean_mm_day is the raw ERA5 climatology in mm/day (arithmetic mean, so a
little above hist_mean_mm).n_units_in_country is in the units table.SEAS5 arrives on the 5th; ERA5 for the previous month around the 6th. The monthly cron
(.github/workflows/monthly-refresh.yml, 7th 03:00 UTC) refreshes admin-0 only:
the skill stats, then build_hdx_signal_inputs.py --level 0, so
signal_inputs{,_latest}.parquet and units.parquet follow the country
data automatically (verify_site_data.py --check-signal fails the run if they do
not). The admin-1/2 tables are manual, from the repo (uv sync first; needs the
blob write SAS and prod DB read credentials):
uv run python pipeline/compute_skill_adm1.py # ~40 min, 8 workers uv run python pipeline/compute_skill_adm2.py # all ADM2_ISO3S for L in 1 2; do uv run python pipeline/build_hdx_signal_inputs.py --level $L; done # only if units were added or changed (new boundaries, new admin-2 country): for L in 0 1 2; do uv run python pipeline/compute_monthly_clim.py --level $L; done
Adding a country at admin-2: add its ISO3 to ADM2_ISO3S in
compute_skill_adm2.py, then compute_skill_adm2.py --iso3 XXX (merges
into the existing files), compute_monthly_clim.py --level 2,
build_hdx_signal_inputs.py --level 2.