Somalia Riverine Flood Trigger / data-source review

Somalia floods: data-source review

The evidence behind the trigger: the SWALIM gauge record, how each model performs against it, the stations selected, and the configuration that follows. Generated 2026-09-01 from the source data.

1-in-3.2 envelope activation rate (target 1-in-3)
7 of 8 severe years caught (1-in-5 or rarer at the gauge)
4 + 3 gauges monitored: Juba and Shabelle
7 days lead time | sources: GloFAS v5 + Google GRRR

1Key takeaways

2SWALIM thresholds

SWALIM publishes three risk levels per station: moderate, high and bank full. Each station's 1-in-3-year level is estimated from its own annual maxima and compared against them. Where the 1-in-3-year level sits above the moderate mark, that mark is crossed more often than once every three years.

Where the 1-in-3-year level sits against the official marks. The official Moderate band means very different things from gauge to gauge, which is why the trigger monitors the computed RP3 instead.
Where the 1-in-3-year level sits against the official marks. The official Moderate band means very different things from gauge to gauge, which is why the trigger monitors the computed RP3 instead.

3SWALIM events

An event is a spell in which the level sits at or above the baseline. The level must stay below it for at least 14 days before a new crossing counts separately, so a brief dip mid-flood does not split one flood in two. Readings stop rising once a gauge reaches bank full, so the largest events are understated.

What the two-gauge benchmark sees, and what it misses. A river-season counts as a flood only when two or more of that river's gauges cross their own level, which keeps one record from deciding the benchmark but has three consequences worth stating.
  • It agrees with the impact record on the big years. All five costliest years in EM-DAT and CERF terms — 2006, 2018, 2019, 2020 and 2023 — are severe under the rule.
  • 1999 to 2001 cannot be assessed. The Juba had no gauge reporting and the Shabelle only one, so no consensus is possible and those years read as quiet rather than as unknown. Nothing activates before 2005, so no year is wrongly scored as a miss, but the benchmark effectively starts in 2002.
  • 2021 is a genuine gap. EM-DAT records 400,000 people affected, yet only Bardheere on the Juba and Belet Weyne on the Shabelle crossed, so the rule reads no flood. Belet Weyne's peak that year sits exactly at bank full (8.30 m), where the gauge record is censored and the true level is unknown, so consensus can be understated in exactly the years that matter most.
  • 2016 is the reverse case. Three of the four Juba gauges and all three Shabelle gauges crossed their levels in Gu, but EM-DAT records nobody affected. Gauge levels and recorded impact are not the same measurement.
Seven gauges still reporting: four on the Juba, three on the Shabelle.
Seven gauges still reporting: four on the Juba, three on the Shabelle.
The two baselines side by side. Amber cells are years the official Moderate mark activated and the RP3 did not, which is where the two definitions disagree.
The two baselines side by side. Amber cells are years the official Moderate mark activated and the RP3 did not, which is where the two definitions disagree.

4SWALIM risk level crossings

How often each gauge crosses its RP3 level. Most gauges reach it in roughly one season-year in three, by construction.
How often each gauge crosses its RP3 level. Most gauges reach it in roughly one season-year in three, by construction.

5SWALIM against flood exposure

The gauge record matters because of who sits behind it: the districts along both rivers with the largest populations in flooded areas.
The gauge record matters because of who sits behind it: the districts along both rivers with the largest populations in flooded areas.

6Reanalysis performance

Each model's 1-in-3-year signal against the events recorded at the gauges, 2002 to 2023, within a 7-day window. POD is the share of events caught; FAR the share of alarms that were false.

Each model's 1-in-3-year signal against the floods the gauges actually recorded, within a 7-day window.
Each model's 1-in-3-year signal against the floods the gauges actually recorded, within a 7-day window.
StationEventsGloFAS v5GloFAS v4Google GRRRGEOGloWS v2
dollow31.00 / 0.400.67 / 0.601.00 / 0.500.67 / 0.60
luuq90.44 / 0.500.22 / 0.710.44 / 0.430.11 / 0.88
bardheere40.50 / 0.750.25 / 0.860.50 / 0.750.25 / 0.88
bualle70.29 / 0.670.43 / 0.570.43 / 0.500.43 / 0.57
belet_weyne100.60 / 0.140.10 / 0.880.50 / 0.290.30 / 0.62
bulo_burti90.44 / 0.430.11 / 0.880.33 / 0.570.44 / 0.50
jowhar100.20 / 0.710.10 / 0.880.10 / 0.860.10 / 0.86

7Model correlation

Daily tracking rather than event detection: the best-lag rank correlation between each model and the river's reference gauge, allowing for travel time.

Daily tracking rather than event detection: the best-lag rank correlation between each model and the river's reference gauge.
Daily tracking rather than event detection: the best-lag rank correlation between each model and the river's reference gauge.
RiverBestSecondThird
Juba (worse season)GloFAS v5 0.83GloFAS v4 0.82GEOGloWS v2 0.8Google GRRR 0.53
Shabelle (worse season)GloFAS v4 0.78GloFAS v5 0.74GEOGloWS v2 0.73Google GRRR 0.63

8Forecast correlation

The same test on the forecasts rather than the hindsight simulations, leads 1 to 7.

The same test run on the forecasts rather than the hindsight simulations. Accuracy holds across leads 1 to 7, which is what makes a one-week action window defensible.
The same test run on the forecasts rather than the hindsight simulations. Accuracy holds across leads 1 to 7, which is what makes a one-week action window defensible.

9Selected stations

RiverGauges used
Jubadollow, luuq, bardheere, bualle
Shabellebelet_weyne, bulo_burti, jowhar

Gauges whose records stop in 2008 or earlier can calibrate but cannot monitor. Upstream stations see the flood before the downstream reference gauge: Dollow leads Luuq by about five days in Deyr, and Belet Weyne leads by about six in Gu.

10Grid search

Every combination of station return period and number of gauges that must agree, scored against the gauge benchmark. The outlined cell is the one adopted.

Every combination of station return period and number of gauges that must agree, scored against the gauge benchmark. The outlined cell is what we adopted.
Every combination of station return period and number of gauges that must agree, scored against the gauge benchmark. The outlined cell is what we adopted.

11Best configuration

RiverSeasonForecastGauge thresholdGauges that must agreeWould have activated in
JubaGuGoogle GRRR1-in-53 of 42013, 2016, 2018, 2020
JubaDeyrGloFAS v51-in-43 of 42006, 2023
ShabelleGuGoogle GRRR1-in-62 of 32013, 2016, 2018, 2020
ShabelleDeyrGloFAS v51-in-42 of 32006, 2013, 2014, 2019, 2020, 2023

The full amount is released whenever the trigger is reached along either river in either season, so the combined rate across all four windows is what the budget must be sized on. As configured it would have released in 8 of 25 years, once every 3.2 years, catching 7 of the 8 severe years.

GEOGloWS is excluded from the adopted design. Its forecasts run below its own retrospective, so thresholds would have to be refitted on its forecast archive, which only begins in July 2024 and can neither be backtested at lead time nor validate a debias. Its per-gauge detection was also the weakest of the four models, and the SFDC bias correction lowers it further rather than fixing it.
Year by year, what the adopted configuration would have done, and what the envelope does when any of the four windows activates.
Year by year, what the adopted configuration would have done, and what the envelope does when any of the four windows activates.