Somalia Riverine Flood Trigger / design comparison

Mixed models against one model per window

Both setups on GloFAS v5 and Google Flood Hub only, both on the five SWALIM gauges that still report, and both balanced so that each basin-season carries about the same probability of activating, at an overall return period of 1-in-3.2. September 2026.

What is held constant, so that only the mixing rule differs. GEOGloWS is excluded from both — its forecast archive begins July 2024, too short to fit thresholds on or to validate a debias against. Both use the five gauges that still report (Belet Weyne, Bulo Burti and Jowhar on the Shabelle; Luuq and Dollow on the Juba), so every activation is checkable against a live observation. Both are balanced: the search minimises the spread of the four basin-season activation counts, subject to the envelope landing between 1-in-3.0 and 1-in-3.6, and prefers non-unanimous rules wherever the result ties. Both land on 8 activations in 25 years, 1-in-3.2. Scored on the corrected benchmark (branch fix/corrected-benchmark-and-lag, return levels fitted strictly inside 1999–2023).
The algorithm, in two stages. 1. Select the best station-models. A unit is dropped if its signal trails the reference gauge by more than 3 days (a ≤ 6-day trigger cannot spare that), or if its rank correlation against the gauge falls below 0.50. What survives is then cut to everything within 0.10 ρ of the window's best — on a 25-year record, units that close are statistically interchangeable, so they all get a vote. For the mixed setup this runs model-blind across every (station, product) pair; for the one-model setup it runs within each product, so a product is judged on its own best stations rather than being crowded out of a window another product happens to dominate. 2. Calibrate the rule. Search the return-period threshold and the number of units required, per window, for the balanced and most accurate combination that keeps the envelope in the target band.

The two configurations

Each window's voting units are listed under its rule — a gauge paired with the product read at it, with its rank correlation against the reference gauge and its best lag in days. These are what survived selection; the count of candidates considered is shown beside each rule.

One asymmetry worth naming. Because the relative floor runs within each product for the one-model setup, that setup can reach units the mixed setup rejects as unfit. It bites in exactly one window: in Shabelle Deyr, Google's three gauges sit at ρ 0.62–0.64 against GloFAS at Belet Weyne's 0.83, so the mixed pool cuts all three (3 of 6 candidates) while the one-model setup is offered Google as a complete three-gauge option. It declined it — both setups adopt the same three GloFAS units there, so nothing on this page turns on the choice. The absolute ρ ≥ 0.50 floor applies either way, so the per-product floor can admit a mediocre unit but never a poor one.

loading…

1. Historical trigger record

One row per year, one block of columns per basin-season. The number is how many voting units were over their own thresholds at once; a cell is shaded when it reaches the threshold printed in its own column header, which is exactly when that window activates. Thresholds differ by column, so the same number can be shaded in one window and left plain in another. After the two trigger columns each block carries what actually happened: the gauge class, the CERF flood allocation and the EM-DAT people affected attributed to that basin-season, and whether the year is on the target list.

loading…
mixed models activates one model per window activates CERF allocation EM-DAT affected target year gauges 1-in-5 or rarer gauges 1-in-3 to 1-in-5

2. Summary statistics

Scored two ways. Against the gauge benchmark — the years in which two or more of a river's gauges reached a 1-in-5 level in that season — and against a constructed target list of years the framework arguably should have paid out in: . Both are computed per basin-season and then for the envelope. Lead time is the mean number of days between the first forecast alert at leads 1–6 and the reference gauge crossing its own 1-in-3 level; negative means the trigger fired after the river was already up.

loading…
score: <.40 .40 .55 .70 .85+ lead: ≤−8 d −8 0 +3 +8 d

The target years

loading…

CERF allocations naming neither river are counted under both basins. EM-DAT events are attributed by the river named in the location text and split to seasons by start month, so a nationwide event appears under both basins. This list is a proxy for “years the framework should have paid out”, not a validated impact record: it inherits EM-DAT's entry-criteria bias and the fact that a CERF allocation reflects an appeal as much as a flood.

Provenance. Thresholds are Weibull plotting positions on each product's own seasonal maxima, 1999–2023. Lead times use the Google GRRR (2016–2023) and GloFAS v4 (2003–2023) reforecast archives, GloFAS v5 being represented by v4 since it has no reforecast of its own. Impact records: EM-DAT (CRED) and CERF OneGMS. Related: the multi-source mechanism, the one-model-per-window mechanism and the balanced trigger.