Somalia Riverine Flood Trigger / trigger analysis

Trigger mechanism: one model per window

Proposed trigger for anticipatory action against riverine flooding on the Juba and Shabelle: a station consensus in which each river-season window runs on a single forecast source rather than a mixture, monitoring all seven points, with no threshold below 1-in-3. Calibrated against SWALIM river gauges. August 2026.

Provider set
Adopted: GloFAS + Google, one source per window
You are viewing the no-Google variant. The whole report below — mechanism table, return periods and backtest — is recalibrated from scratch with Google Flood Hub excluded, using identical rules (all seven points, one source per window, thresholds at or above 1-in-3, and the envelope judged jointly). This is a resilience scenario for the working group, not the proposed mechanism. Narrative sections that do not depend on the provider set (how the model is chosen, the threshold floor, the readiness band) read the same in both views.
You are viewing the all-providers configuration. With GEOGloWS allowed to compete it takes Juba Gu, and GloFAS v5 takes the other three windows. The envelope is unchanged at 1-in-3.2 (8 of 25 years) with the same 7 of 8 severe years, but it does not activate in 2013, so this configuration has no activation outside the benchmark. Dropping Google alone reproduces this same configuration exactly, which is why there is no separate no-Google view. The comparison sections further down show every candidate model regardless of the setting here: they are evidence about models, not statements about what was chosen.
ReviewTwelve notes are open on this document.

They are anchored beside the passage each one concerns. Nothing in the trigger design has been changed — the configuration, the numbers and the figures are exactly as generated. What has been corrected is text that contradicted the data, plus the broken provider switch; supporting evidence has been folded into collapsible blocks.

  1. R1 — the “7 of 8” headline rests on a lookahead in the benchmark
  2. R2 — every published lag is wrong; correlations understated
  3. R3 — three provider sets were really two
  4. R4 — the Juba leg never decides anything
  5. R5 — the per-window search does not beat an un-searched rule
  6. R6 — one model per window is defensible; mixing gains nothing
  7. R7 — the seven-point set is load-bearing, and two points are unverifiable
  8. R8 — excluding GEOGloWS introduces the one false activation
  9. R9 — the two lead bands overlap at lead 7
  10. R10 — five of sixteen activations have no readiness phase
  11. R11 — 1999–2001 unassessable but still in every denominator
  12. R12 — some benchmark levels rest on ten annual maxima
1-in-3.2overall action return period (either river)
1-in-3.2 / 1-in-4.3per river, Shabelle / Juba
4 + 3points monitored, Juba and Shabelle
7–12 dreadiness lead time (action at 1–7 days)

The mechanism at a glance

Review R4The Juba leg never decides anything.

Across 1999–2023 the envelope is identical to the Shabelle river on its own: there is no year in which the Juba activates and the Shabelle does not. Per-river rates are 1-in-4.3 (Juba) against 1-in-3.2 (Shabelle), so the equal-probability-per-river goal is not met. Both Gu windows also run on Google and fire in the same four years (2013, 2016, 2018, 2020), so the Gu half of the mechanism is effectively a single trigger counted twice.

Four river-season windows, each running on one forecast source rather than a mixture. Inside a window, every monitored point's flow is compared with its own return-period threshold and the window activates when enough points are over their thresholds on the same day. A season where the points cross on different days does not count. This rarely matters in practice: flood flows stay high for weeks, and in every past activation the points crossed within 0 to 5 days of each other, upstream gauges first (see "How far apart do the points cross?" below). At least two points must agree, so no single point releases the money, and never all of them, so one quiet point cannot block it.

All seven reporting-era points are monitored: four on the Juba (Luuq, Dollow, Bardheere, Bualle) and three on the Shabelle (Belet Weyne, Bulo Burti, Jowhar). No threshold, on either leg, sits below 1-in-3.

Review R7The seven-point set is load-bearing, and two of the points can no longer be checked.

Restricting to the five gauges that still report would leave the Juba with only Luuq and Dollow, so the only available rule is 2 of 2 — unanimous, which this design's own “never all of them” constraint forbids — and it adds a false activation in 2015, moving the envelope to 1-in-2.9. Bardheere and Bualle are therefore what make a non-unanimous Juba rule possible. The cost is that an activation driven by those two points can never again be verified against an observation: Bardheere's record ends 2023-11-30 and Bualle's 2024-03-14.

windowaction trigger (leads 1-7 d)leg RPreadiness (leads 7-12 d)readiness RP
Juba GuGoogle GRRR: 3 of 4 points over their 1-in-5-yr thresholds6.5 yrGloFAS v4 ens-median: 3 of 4 over 1-in-5-yr5.5 yr
Juba DeyrGloFAS v5: 3 of 4 points over their 1-in-4-yr thresholds13.0 yrGloFAS v4 ens-median: 3 of 4 over 1-in-4-yr4.4 yr
Shabelle GuGoogle GRRR: 2 of 3 points over their 1-in-6-yr thresholds6.5 yrGloFAS v4 ens-median: 2 of 3 over 1-in-5-yr5.5 yr
Shabelle DeyrGloFAS v5: 2 of 3 points over their 1-in-4-yr thresholds4.3 yrGloFAS v4 ens-median: 2 of 3 over 1-in-4-yr4.4 yr
The envelope. The full amount is released whenever any window activates, so the union is what the 1-in-3 target applies to. As configured it would have released in 8 of 25 years, once every 3.2 years, catching 7 of the 8 years in which two or more of a river's gauges recorded a 1-in-5 or rarer season. It activates once, in 2013, in a year the two-gauge benchmark does not record as a flood.

How the two rivers behave together

The Shabelle activates more often: 8 of 25 years (1-in-3.2) against the Juba's 6 (1-in-4.3). There is no year in which the Juba activates and the Shabelle does not; the Shabelle alone adds 2014 and 2019. Six activations fall in the same season on both rivers, and in all six the Juba reaches its thresholds first, by 1 to 24 days:

seasonJuba crossesShabelle crosseslead
Gu 201312 Apr6 MayJuba first, 24 d
Gu 20167 May12 MayJuba first, 5 d
Gu 201820 Apr5 MayJuba first, 15 d
Gu 202029 Apr30 AprJuba first, 1 d
Deyr 200629 Oct2 NovJuba first, 4 d
Deyr 202329 Oct9 NovJuba first, 11 d

So the Juba never adds an activation year (review note R4), but in every shared season it crosses first. Since either window releases the full amount, the Juba sets the response date in those six years; the Shabelle activates more often and adds the years the Juba misses.

Review R1The “7 of 8” headline rests on a lookahead in the benchmark.

Each gauge's return level is fitted on 2000–2026 while crossings are counted only within 1999–2023, so three years of post-window data enter the labels. That is what makes 2008 severe: Luuq's Deyr 1-in-5 level is 5.831 m fitted through 2026 but 5.869 m fitted through 2023, and the 2008 peak is 5.840 m — a 9 mm margin on a level that moves 38 mm purely from adding future years. Fit strictly to the backtest window and the severe set is 7, not 8, and this configuration catches 7 of 7. The much-discussed “miss in 2008” is an artifact of the asymmetry, since trigger thresholds are fitted strictly on 1999–2023. Corrected on branch fix/corrected-benchmark-and-lag (not merged, and this page is unchanged); the corrected numbers are used on the design comparison page.

Why GEOGloWS is not in the adopted design. Its forecasts run below its own retrospective — a median ratio of 0.89 across leads 1–7 — so a threshold fitted on the retrospective sits too high for the live forecast ever to reach at the intended rate. Quantified across the seven monitored points: a level set as 1-in-5 on the retrospective behaves like a median 1-in-13 on the forecast, and as rarely as 1-in-26 at Bardheere and Bualle in Gu. A window built that way would mostly sit silent.

The obvious answer is to debias, and that is what cannot be done yet: the GEOGloWS forecast archive begins in July 2024, roughly two years, which is neither enough to refit thresholds on the forecast climatology (this design asks about twelve years for a 1-in-3) nor enough to validate a flow-duration-curve correction against anything. The team's Nepal technical note is the cautionary precedent — SFDC did not rescue event detection there. The working group therefore leaned to excluding GEOGloWS on 2026-08-22 and adopted that on 2026-08-28. It stays in this view as a comparison, and stays a candidate for reinstatement once two or three Gu and Deyr seasons of its forecasts exist.
What the two-gauge benchmark sees, and what it misses. A river-season counts as a flood only when two or more of that river's gauges cross their own level, which keeps one record from deciding the benchmark but has three consequences worth stating.
  • It agrees with the impact record on the big riverine years. 2006, 2018, 2019, 2020 and 2023 are all severe under the rule. Ranked on raw EM-DAT totals, though, the third-costliest flood year is 2015 (916,000 affected), which the rule does not call a flood at all — correctly, since that event was in Galgaduud, Mudug and Nugal rather than on either river. The agreement holds only once EM-DAT is filtered to Juba and Shabelle events.
  • 1999 to 2001 cannot be assessed. The Juba had no gauge reporting and the Shabelle only one, so no consensus is possible and those years read as quiet rather than as unknown. Nothing activates before 2005, so no year is wrongly scored as a miss, but the benchmark effectively starts in 2002.
  • 2021 is a genuine gap. EM-DAT records 400,000 people affected, yet only Bardheere on the Juba and Belet Weyne on the Shabelle crossed, so the rule reads no flood. Belet Weyne's peak that year sits exactly at bank full (8.30 m), where the gauge record is censored and the true level is unknown, so consensus can be understated in exactly the years that matter most.
  • 2016 is the reverse case. Three of the four Juba gauges and all three Shabelle gauges crossed their levels in Gu, but EM-DAT records nobody affected. Gauge levels and recorded impact are not the same measurement.
  • Review R111999–2001 cannot be assessed, yet still count in every denominator.

    No window can reach two reporting gauges in those years, so they are scored as “no flood” rather than as unknown, while 25 years remains the denominator for every return period quoted on this page. The assessable record is effectively 22 years. The same mechanism silently affects Shabelle Deyr in 2000 and 2001, where Jowhar was the only reporting gauge and did cross its level, and Juba Gu in 2005–2006.

Review R12Some benchmark levels rest on very short gauge records.

The stated ceiling of “a quarter of the record” is applied to the forecast thresholds but not to the benchmark. Dollow has 10 Deyr annual maxima after 2000 and is still assigned a 1-in-5 level; Bualle has 16. Since a single gauge crossing decides whether the two-gauge rule is met, a level interpolated from ten values can decide whether a year counts as a flood at all.

What happens if Google Flood Hub goes away

Use the provider switch above. Removing Google returns the configuration to the all-providers one: GEOGloWS takes Juba Gu and GloFAS v5 takes the other three windows. The envelope stays at 1-in-3.2 (8 of 25 years) with the same 7 of 8 severe years, and it loses the 2013 activation, so on the calibration record dropping Google costs nothing measurable. The catch is that the replacement is GEOGloWS, the one model whose thresholds cannot yet be fitted on its own forecasts, so the fallback is weaker operationally than it looks on the backtest.

Review R3There were only ever two distinct provider sets, not three.

The “without Google” configuration is numerically identical to “all providers”, because when GEOGloWS is allowed to compete Google never wins a window. The third switch state has been removed, and the two competing switch controllers that were both live on the page (which corrupted the default view once you clicked) have been reduced to one.

Review R8Excluding GEOGloWS is a working-group decision; the 2013 activation is its price.

The decision was taken by the working group (leaning this way on Friday 2026-08-22, adopted 2026-08-28) and the reasoning is sound — see the callout above for the quantified version. The cost should be recorded alongside it: the all-providers configuration activates in 8 years with no activation outside the benchmark, whereas the adopted one activates in the same number of years but fires in 2013, a year the two-gauge benchmark does not call a flood. Trading one false activation for a model that can actually be operated is a defensible trade; it is just better made explicitly.

The seven monitored points
The seven points the trigger watches, four on the Juba and three on the Shabelle.
Where the design constraints come from. Three of the rules below are decisions rather than findings, and are recorded here so they are not mistaken for results:
  • One source per river-season window, never a mixture — requested by WFP, who raised concerns about mixing models inside a window. Tested against the alternative it costs nothing measurable (review note R6).
  • GEOGloWS excluded — working group, leaning that way 2026-08-22 and adopted 2026-08-28, because its forecasts run below its own retrospective and its two-year forecast archive is too short either to refit thresholds on or to validate a debias against.
  • The envelope activates no more often than 1-in-3 — directive of 2026-08-27, applied to the union of the four windows rather than to each window.

How the model was chosen, one per river and season

Review R5The per-window search does not beat a rule with no search in it — partly answered.

Applying one model, one return period and one vote count uniformly to all four windows: Google GRRR, 1-in-6, 3 of 4 reproduces the adopted result exactly (8 activations, 1-in-3.2, 7 of 8 severe, the 2013 false alarm). GloFAS v5, 1-in-5, 3 of 4 gives 9 activations at 1-in-2.9 with the same 7 of 8 and zero activations outside the benchmark. The roughly 41,000 admissible per-window combinations recover nothing a one-line rule does not. This is consistent with the page's own statement that many assignments tie, but it is stronger: the per-window freedom is not adding measurable value. Partly answered by the swap test added 1 September, which I have reproduced and which holds: at the adopted return periods and vote counts, neither model alone can carry all four windows without breaching the 1-in-3 ceiling — GloFAS v5 everywhere and Google everywhere both land at 1-in-2.6 — so the mixed assignment is doing real work, and the Deyr case for GloFAS v5 is clear (Google brings three false activations on the Juba and misses a Shabelle severe year). Those conclusions are unchanged on the corrected benchmark. What remains of this note is narrower: the swap test holds the thresholds fixed at the adopted values, so it shows the assignment is defensible given those settings, not that the configuration as a whole beats a simpler one. A uniform rule at different settings — Google GRRR, 1-in-6, 3 of 4 in every window — still reproduces the adopted envelope exactly. Both are true: the assignment is justified, while the tuning around it is not what earns the result.

One source per river and season, so four choices. Four steps:

Review R6One model per window comes from WFP, and it costs nothing measurable.

The constraint is a stakeholder requirement, not a data-driven one: WFP raised concerns about mixing models within a window. It is worth recording that it is cost-free on this record. Allowing every model to compete at every point inside a window (a mixed consensus over the same seven points) reaches 7 of 8 severe years at 1-in-3.7; the best single-model configuration reaches the same 7 of 8 at 1-in-3.2. Mixing buys no additional severe year, so the requirement can be honoured without giving anything up — and it buys a real operational simplification, since one model per window means one threshold transfer to validate rather than several.

  1. Build the candidate rules. For a window and a candidate model, put a threshold at every one of the river's points, taken from that model's own annual maxima (1-in-3 to 1-in-6). The window activates in a year when at least N points cross in that season.
  2. Score the envelope, not the window. The money is released when any of the four windows activates, so candidates are judged on that union: how often it activates, how many of the 8 severe years it catches, and how often it activates in a year with no recorded flood.
  3. Apply the constraints. Thresholds never below 1-in-3 nor above a quarter of the record; all seven points monitored; at least two must agree but never all of them; every window must activate at least twice in 25 years.
  4. Pick. Nearest 1-in-3 overall with the most severe years caught. Ties break on fewer no-flood activations, then on tracking correlation.
Best model per river and season on RP3-or-rarer events
Mean best-lag rank correlation between model and gauge inside observed RP3+ event windows (widened by 10 days), averaged over the window's gauges. The hit rate on the same events is printed in each bar. Models shown follow the provider switch; the marker is the adopted one. The lag search has the limitation flagged in review note R2.

The per-point tracking numbers behind the choice (ρ and best lag of each model against the river's reference gauge):

Review R2Every lag in the table below is wrong, and the correlations are understated.

The best-lag search shifts the gauge series by row position after filtering to season months, so any non-zero lag pulls rows from the adjacent year and always scores worse than lag 0. The result is that 46 of the 56 cells below report a lag of exactly 0, including Bualle against the Luuq reference gauge some 450 km upstream. Recomputed with a calendar-day shift, Belet Weyne against GloFAS v4 in Deyr peaks at +22 days with ρ 0.78, not at 0 days with ρ 0.67. Because tracking correlation is the final tie-break between models, that tie-break is currently running on a corrupted statistic. Corrected on branch fix/corrected-benchmark-and-lag (not merged, and this page is unchanged); the corrected numbers are used on the design comparison page.

Per-station selection detail (ρ / lag per model, all four windows)
loading…

Why the models perform the way they do

The main difference between the models is magnitude. Each reproduces the shape of the flood tail reasonably well, but over- or under-estimates its size by a roughly constant factor.

Model over observed discharge at return periods 3 to 6
Each model's return level divided by the observed return level, RP3 to RP6, at the four gauges with published discharge. The lines are close to flat: the bias is a constant scale factor and does not grow with severity. Models shown follow the provider switch.
One caveat on this figure. The observed rating curve caps at bank full, so where a model reads below 1 the true ratio may be closer to 1: the gauge cannot record the top of the largest floods. Ratios above 1 are unaffected.

How each model maps the RP3 events, gauge by gauge

The second question is whether a model is above its own 1-in-3 threshold when the gauge is above its own. That is the event the trigger has to catch.

GEOGloWS is not in the adopted set and is not shown here. Its numbers are under All providers.

Hit rate and false-alarm rate per gauge and model on RP3 events
Every observed 1-in-3 crossing at each gauge against each model's crossing of its own 1-in-3 threshold, counted as events and matched within 7 days. Left: the share of observed events the model caught. Right: the share of the model's own crossings with no observed event behind them.

The same comparison at the level the trigger operates: each window's vote rule, 3 of 4 points on the Juba and 2 of 3 on the Shabelle at the adopted return periods, with each model swapped in. POD counts severe years caught, the objective the envelope is sized on. FAR counts activations with no RP3 flood behind them:

Severe-year POD, false-alarm rate and F1 of each window rule per model
Each window's rule run on each model, 1999-2023, against the two-gauge benchmark. Gu 2023 is missed by every candidate; the envelope catches 2023 through the Deyr windows. The return periods were tuned for the adopted models; the other models are run at the same settings.

The same test on the forecasts the trigger would actually run on, at action leads of 1 to 7 days. GloFAS v5 has no reforecast, so GloFAS is represented by the v4 ensemble median, its only lead-time archive; GEOGloWS has no archive before July 2024 and cannot be scored:

The window swap test on the forecast archives
Same vote rules and return periods. Thresholds are fitted on each model's own reanalysis, as the trigger operates, and the leads-1-7 forecast series (ensemble median, best lead per day) is tested against them over the common years 2016-2023. Eight years give 1 to 3 severe events per window, so read direction, not decimals. This figure does not change with the provider switch.
Why Google carries Gu and GloFAS v5 carries Deyr.
  • On severe years, Gu is a tie. Both models catch 2018 and 2020 on the Juba and 2016 and 2020 on the Shabelle, and both miss Gu 2023.
  • Deyr on Google is worse on both counts: three false activations on the Juba and a missed severe year (2020) on the Shabelle. GloFAS v5's Deyr record is clean.
  • GloFAS v5 alone cannot meet the 1-in-3 ceiling. No v5-only configuration reaches 8 activations, and its nearest envelope fires once every 2.9 years. Some windows therefore have to run on Google, and Gu is where the swap costs nothing.
  • Google is the only model with forecast-side evidence. Its reforecast (2016-2023) shows event skill at leads 1 to 7 on the Shabelle; the only GloFAS lead-time archive (v4) scores near chance there. GloFAS v5 has no reforecast. It keeps Deyr because its thresholds carry over to the live v5 forecast, whose climatology matches the reanalysis they were fitted on.

GEOGloWS maps the RP3 events least well of the models shown. It reaches 0.30 at Belet Weyne against GloFAS v5's 0.60, 0.11 at Luuq against 0.44 for both v5 and Google, and 0.25 at Bardheere, and it pairs those misses with the worst false-alarm rates on the board: 0.88 at Luuq and Bardheere, 0.62 at Belet Weyne. It is not uniformly worst: it ties v5 at Bulo Burti (0.44) and reaches 0.67 at Dollow, where every model does well. These detection numbers, together with the threshold problem described above, are why it is left out of the adopted design.

GloFAS v5 is best or joint-best at six of the seven gauges and has the cleanest false-alarm record (0.14 at Belet Weyne, 0.40 at Dollow), which is why it carries both Deyr windows. Its weakest gauge is Bualle, where it catches 0.29 against 0.43 for the others. Google matches v5 at Dollow (1.00), Luuq (0.44) and Bardheere (0.50) and is the only model with a usable forecast archive on the Shabelle, which is why it carries Gu.

No model does well at Jowhar (0.10 to 0.20, false alarms 0.71 and up): off-takes and marsh hydraulics decouple the local level from upstream discharge. Requiring two gauges keeps it from deciding anything on its own.

Small samples. Each gauge has only 3 to 10 observed RP3 events in the scoring window, so a single event moves a rate by 0.1 or more. The ranking is consistent across gauges, but the individual values are not precise.

This shapes the mechanism in three ways. Thresholds are set on return periods, not on absolute discharge: with a flat tail ratio, matching frequency corrects the bias, so even a model running ten times too wet stays usable. A threshold is then only valid against the climatology it was fitted on, which is where GEOGloWS fails: its forecast climatology sits below its retrospective. Skill is flat across leads 1 to 7 days, because the signal is water already in the channel rather than forecast rainfall, so a one-week action window costs little accuracy. And no single gauge decides anything: Jowhar cannot be predicted from upstream discharge and Bardheere's record is suspect.

The seasonal-peak comparison is a static figure showing every candidate model. It is available under All providers.

Supporting evidence — seasonal peaks, model vs gauge

Seasonal peaks: model vs gauge

Carried over unchanged from the published multi-source study: this diagnostic compares models, so it does not depend on how many points vote or where the thresholds sit.

Daily rank correlation can flatter a model that merely tracks the seasonal cycle. The sharper diagnostic is one point per season-year: the model's seasonal peak against the observed seasonal peak at the reference gauge — does the model rank the years correctly? For display each model's peak is divided by its own 1-in-6 threshold (rank correlation is unaffected by the normalisation), so the quadrants read directly: top-right = the model ran over its own threshold in a year the gauge also ran high.

Scatter plots of model seasonal peaks over own threshold versus SWALIM observed seasonal peaks, by river and season, all three providers
Calibration record (reanalysis/retrospective, 1999–2023), at the reference station. Dark-edged points are selected pairs (the reference station carries two models in most windows). Vertical dashes: the gauge's Moderate (amber) and High (red) levels; horizontal dash: the model at its own 1-in-6 threshold.
windowpeak Spearman ρ — reanalysis (n = 21–25) peak ρ — reforecast ≤ 6 d
GEOGloWSGloFAS v5Google GRRR GloFAS v4 (proxy)Google GRRR
Juba Gu0.650.640.85 0.630.98
Juba Deyr0.650.590.74 0.640.75
Shabelle Gu0.680.800.85 0.640.74
Shabelle Deyr0.490.770.78 0.390.81
Scatter plots of reforecast seasonal peaks versus SWALIM observed peaks, by river and season
Operational record: the ≤ 6-day reforecast signal's seasonal peaks, normalised by a 1-in-6 threshold fitted on the reforecast's own seasonal maxima (GRRR's 8-season archive makes its own threshold estimate wide — treat its vertical placement, not its ranking, with caution).

Thresholds and calibration

Each monitored point gets its threshold from its own model's climatology: the Weibull plotting position of that point's seasonal maxima, so a model is judged on timing rather than on scale. 1-in-3 is the floor and a quarter of the record is the ceiling, which on 25 years allows 1-in-3 to 1-in-6.

Heat strip of activation years per trigger leg, per river versus severe years, and overall, 1999 to 2023
The adopted configuration's backtest. River rows compare activations with the severe (1-in-5) years of the two-gauge benchmark. The envelope catches 7 of the 8 severe years, missing only 2008, and activates once outside the benchmark, in 2013. The Juba river adds nothing to the envelope: every year it activates, the Shabelle activates too (review note R4).The all-providers configuration, same layout, with GEOGloWS carrying Juba Gu: the same envelope (8 years, 1-in-3.2) and the same 7 of 8 severe coverage, but no activation outside the benchmark — it does not fire in 2013. With 25-year records, one hit moves coverage by ~12–15 points — differences under ~0.15 are noise.

How far apart do the points cross?

A window activates only when enough points are above their thresholds on the same day, not just in the same season. How tight are the crossings in practice?

Across every activation, the first and last point to cross fall 0 to 5 days apart, and they cross in downstream order: Dollow, Luuq, Bardheere, Bualle on the Juba; Belet Weyne, Bulo Burti, Jowhar on the Shabelle. Shabelle Gu 2018 had all three within a single day, Juba Gu 2018 four points inside three days. Because high flows persist for weeks, same-day overlap is reached even where the first crossings are a week or more apart, as in Juba Deyr 2023 (11 days between Dollow and Bualle, yet four points above at once).

The exception is instructive. Juba Deyr 2015 had Dollow crossing on 24 October and the other three between 10 and 13 November, 20 days later: two separate pulses rather than one flood wave. Same-day overlap never reached three, so the window did not activate, which is the behaviour we want.

What a tolerance window would change. If crossings within 5 days of each other counted together instead of requiring the same day, Juba Deyr would add 2018, 2019, 2015 and 2004, and Juba Gu would add 2010. The envelope would move from 1-in-3.2 to 1-in-2.4 (8 activations to 11) and activations in years two gauges did not record a flood would go from one (2013) to three (2004, 2013, 2015). Severe-year coverage would not improve, staying at 7 of 8, because 2018 and 2019 are already covered by the Shabelle windows. Widening beyond 5 days changes nothing further: crossings are either within 5 days or 20 days apart. The same-day rule is therefore kept, and a tolerance is a lever for later if the activation rate is allowed to rise.

Calibrated on the reanalysis, checked on the forecasts

Refitting the same rule on the forecast archives themselves, 2003–2023, thresholds from 1-in-3 to 1-in-5. Out of the field: Google GRRR, whose archive is too short to carry a 1-in-3 threshold, so it can only be calibrated on the retrospective.

windowsourcegauge thresholdpoints that must agreewould have activated in
Juba GuGloFAS v41-in-32 of 42005, 2010, 2013, 2016, 2018, 2020, 2023
Juba DeyrGloFAS v41-in-53 of 42015, 2017, 2023
Shabelle GuGloFAS v41-in-42 of 32005, 2010, 2016, 2018
Shabelle DeyrGloFAS v41-in-52 of 32010, 2013, 2019, 2020

Envelope 1-in-2.2 (10 of 21 years), catching 5 of the 8 severe years. This is the closest rate the forecast record supports; nothing lands on 1-in-3.

Supporting evidence — the tuning surface

The tuning surface

The Nigeria-style grid search (per state there; per river × season here — the same recipe as P. Wairimu's notebook 08): every (per-pair RP threshold, N pairs required) cell is scored against the seasonal SWALIM benchmark. The adopted cell is not always the best single-window score — it is the best choice subject to the river-level constraints (near-equal per-river frequency across the Gu+Deyr union, severe-coverage-first, ≥ 3-yr overall enforced jointly across the rivers). With 4–5-pair pools the surface is small enough to read whole: the neighbors show what loosening a leg would cost in false activations, or tightening one in missed severe years.

Twelve heatmaps: for each of the four windows, POD, FAR and activation return period over the per-pair RP threshold by N-required grid, with the adopted cell outlined
POD and FAR are vs the seasonal 1-in-3 SWALIM benchmark; black box = the adopted configuration. Dark = high POD / low FAR / frequent activation. With 7–8 benchmark events per window, one hit moves POD by ~12 points — broad plateaus, not sharp optima, are the honest reading.
Supporting evidence — official SWALIM levels vs the fitted ones

How the official SWALIM levels compare with the fitted ones

The benchmark years above are defined by empirical return periods of the gauges' own seasonal maxima, not by SWALIM's published Moderate/High flood-risk levels — and the check below is why. Putting the official levels on the empirical RP scale (per active gauge and season, 1999–2023) shows they imply wildly inconsistent frequencies: "Moderate" ranges from a level the river reaches most years (Dollow Deyr, ~1-in-1.3; Jowhar Gu, ~1-in-1.6) to a genuinely rare one (Luuq Gu, ~1-in-4.6), and "High" from ~1-in-1.8 (Dollow) to ~1-in-11.5 (Luuq Gu). They are engineering levels of uncertain provenance, not a consistent severity scale — the same conclusion P. Wairimu's notebook 01 flagged when the published max_level at Belet Weyne came out below its bank-full level. Everything in this mechanism therefore runs on each gauge's own empirical RP3/RP5 levels (table som_ms_swalim_rp); the official levels remain useful only as familiar reference lines for readers of SWALIM bulletins.

Five panels, one per active SWALIM gauge, of empirical return-period curves of seasonal maximum levels for Gu and Deyr, with the official moderate and high flood-risk levels as horizontal dashed lines
Empirical Weibull curves of seasonal-maximum level (blue Gu, orange Deyr) per active gauge, with the official Moderate (amber) and High (red) levels as dashed lines. Where a dashed line crosses the curves far from the RP 3–5 band, the official label and the observed frequency disagree. Dollow's curve rests on only 8–9 seasons (gauge online 2015) — read it loosely.

The forecast-vs-own-reanalysis comparison is a static figure showing every candidate model. It is available under All providers.

Supporting evidence — forecast skill against each model's own reanalysis

Forecast skill against each model's own reanalysis

Carried over unchanged from the published multi-source study: this diagnostic compares models, so it does not depend on how many points vote or where the thresholds sit.

The assume-bias-for-all-sources rule, quantified: per model and lead time, how well the forecast reproduces the model's own reanalysis or retrospective — rank correlation (does it track itself?) and the median forecast/reanalysis ratio (is it biased against its own climatology?). Computed per station on flood-season days, median across the 18 stations.

Two panels: Spearman correlation of forecast versus own reanalysis by lead time, and median forecast-to-reanalysis ratio by lead time, for the three models
GloFAS v4 (2003–2023) and Google GRRR (2016–2023) forecasts are essentially self-consistent at action leads: rank ρ ≥ 0.99, ratio ≈ 1.00 — the licence for calibrating their thresholds on reanalysis. GloFAS drifts to ~0.90× by lead 12 (readiness band, shaded) — its 7–12-day thresholds should anticipate that. GEOGloWS runs 0.85–0.91× its own retrospective even at lead 1 (ρ ~0.82), on only ~2 years of archive (2024–26) — retrospective-fitted thresholds will under-fire on its live forecasts until refit on forecast climatology. No GloFAS v5 reforecast exists, so the operational v5 forecast's own consistency is untestable — the single most important gap.

Would it have worked operationally?

The calibration record is reanalysis — a model's afterwards-view of its own past. The operational test replays the historical forecasts: per pair and valid day, the most alarming ≤ 6-day signal (max over issue dates — with mixed providers in one window, a monitoring day combines each provider's latest forecast). Google pairs use their own reforecast (2016–2023); GloFAS v5 pairs the v4 reforecast (2003–2023, ensemble median, thresholds refit on v4's own climatology, since no v5 reforecast exists); GEOGloWS pairs the retrospective as a lead-0 stand-in (hindsight — flagged per window). That forecast skill barely decays from lead 1 to 7 — initial-condition-driven on these slow rivers — is P. Wairimu's notebook 03 finding, and is what makes both lead bands viable at all.

windowcalibrated onrun onreproduces its calibration years?
Juba GuGoogle GRRR reanalysisGoogle GRRR forecasts, leads 1-7yes, same archive
Juba DeyrGloFAS v5 reanalysisGloFAS v5 operationally; lead-time evidence from GloFAS v4cannot be shown directly: no reforecast for this version
Shabelle GuGoogle GRRR reanalysisGoogle GRRR forecasts, leads 1-7yes, same archive
Shabelle DeyrGloFAS v5 reanalysisGloFAS v5 operationally; lead-time evidence from GloFAS v4cannot be shown directly: no reforecast for this version
The caveat that matters most. At ≤ 6-day leads the GloFAS v4 reforecast misses Deyr 2023 on the Shabelle — the largest flood in the record — even though the v5 reanalysis flags it clearly. The operational forecast has been v5 since August 2026 and should behave like the v5 reanalysis here, but there is no v5 reforecast to prove it. Re-verifying both Deyr legs the moment EWDS publishes a v5 reforecast is the single most important follow-up.
What the calibration cannot settle. Google's reforecast starts in 2016, so a Google window cannot be calibrated on its own forecasts at a 1-in-3 threshold, which needs about 12 years: it is calibrated on the retrospective and only checked at lead time. GloFAS v5 has no reforecast at all, so its window inherits its lead-time evidence from v4. Readiness is not tuned to precede activation, so an action trigger may fire with no readiness phase ahead of it.

The readiness leg (7–12 days)

Readiness runs on GloFAS v4 ensemble-median forecasts at leads 7 to 12, the only archive covering that band, over the same full set of points, with thresholds refitted on that series. It releases only the mobilisation share and is held to the same floor: no threshold below 1-in-3.

windowreadiness rulereadiness yearscovers action years
Juba Gu3 of 4 points over 1-in-52010, 2016, 2018, 20203 of 4
Juba Deyr3 of 4 points over 1-in-42005, 2006, 2015, 2017, 20232 of 2
Shabelle Gu2 of 3 points over 1-in-52005, 2010, 2013, 20162 of 4
Shabelle Deyr2 of 3 points over 1-in-42007, 2013, 2014, 2019, 20204 of 6

Readiness carries each window's own rule, the same votes and the same return period, refitted on the 7-to-12-day series.

Review R9The two lead bands overlap, and “covers” does not check ordering.

Action runs at leads 1–7 and readiness at leads 7–12, both inclusive, so lead 7 sits in both bands and the same forecast day can serve either leg. The original specification was action under 6 days. Separately, the “covers action years” column is a season-year set intersection: it records that both legs activated in the same season, not that readiness came first.

It is not tuned to precede activation: an action trigger may activate with no readiness phase ahead of it, which is accepted (decision 2026-08-27). The last column reports how often readiness did lead an activation, as an observation rather than a requirement.

Review R10Five of sixteen window activations have no readiness phase ahead of them.

Counting across the four windows: Juba Gu 3 of 4 covered, Juba Deyr 2 of 2, Shabelle Gu 2 of 4, Shabelle Deyr 4 of 6. The uncovered activations include Shabelle Deyr 2023, the largest flood in the record, and Shabelle Gu 2018 and 2020. Under the staged model that would have been a design failure; under the 2026-08-27 decision it is accepted, but the working group should see the count before accepting it.

Return-period bookkeeping

leveltriggeractivations (1999-2023)RP
individualJuba Gu action4 — 2013, 2016, 2018, 20206.5 yr
individualJuba Deyr action2 — 2006, 202313.0 yr
individualShabelle Gu action4 — 2013, 2016, 2018, 20206.5 yr
individualShabelle Deyr action6 — 2006, 2013, 2014, 2019, 2020, 20234.3 yr
riverJuba (Gu or Deyr)64.3 yr
riverShabelle (Gu or Deyr)83.2 yr
overallaction, either river83.2 yr

Under all-in funding, where any activation releases the full envelope, the effective return period is the overall row: 1-in-3.2 with Google, 1-in-3.2 without. The individual windows are set rarer than that on purpose, because four windows each calibrated to 1-in-3 give a union of roughly 1-in-1.5.

Activations, impact and response, year by year

The two-historical-records rule: the trigger record against impact, not just the hazard benchmark. Everything is grouped river → season, reverse-chronological. The trigger column is the trigger's actual decision variable: the peak number of monitored points simultaneously over their thresholds that season-year — the cell fills red when it meets the requirement shown in the season header (= the trigger activates). RP is the empirical return period of that season's maximum level at the SWALIM reference gauge — 5yr = a 1-in-5-year level or higher, 3yr = 1-in-3 to 1-in-5 (Weibull on the gauge's own seasonal maxima, per the threshold check above — not the official flood-risk levels); EM-DAT (purple) is people affected and CERF (blue) the flood allocation (US$), shaded by magnitude, each attributed to the river(s) named in the event or allocation narrative.

loading…

Attribution & caveats. EM-DAT is a floor (entry-criteria bias), split to seasons by start month (Mar–Jun → Gu, Sep–Dec → Deyr); an event naming both rivers is counted under both rivers (not split — hover for deaths and the multi-river flag); events naming neither river (mostly flash floods elsewhere) are excluded. A CERF allocation naming neither river in its narrative is shown under both rivers and marked * (they are national riverine-flood responses). Impact columns are context, not a skill score: a formal skill-vs-impact statistic still requires choosing an impact threshold.

Supporting analysis — should impact tilt the probabilities?

Should impact tilt the probabilities?

Near-equal per-river probability is a design choice, not a law — a river or season carrying systematically more impact could get a lower relative threshold. Aggregating the river-attributed records (1999–2023) answers whether the data call for it:

Three horizontal bar panels showing each river-season window's share of EM-DAT people affected, EM-DAT deaths, and CERF flood funding, 1999 to 2023
Each window's share of total flood impact, per record (som_ms_impact_by_window). River totals are near-equal — Juba 47%, Shabelle 53% on the three-record average — but the seasons are not: Deyr carries ~65% of impact (and the two largest CERF flood allocations) against Gu's ~35%.

Re-running the calibration with the activation budget allocated by impact share instead of equally (table som_ms_impact_tilt) makes the conclusion concrete:

  • A river tilt changes nothing. The 47/53 split rounds to the same 5/6 activation allocation the joint calibration already adopted — the adopted mechanism's slight Shabelle lean is already impact-consistent.
  • A season tilt is possible but costs skill. Pushing Juba's budget toward Deyr (its higher-impact season) forces its Deyr leg to a lower threshold and its Gu leg higher: the tilted variant swaps documented Gu floods for extra Deyr activations, drops Juba's severe-year coverage from 0.50 to 0.38, and pushes the overall frequency to 1-in-2.9 — under the ≥ 3-year specification. The calibration already leans Deyr where the data support it (both Deyr legs run at lower pair-RP thresholds than the Gu legs); forcing more requires accepting misses the gauge record says are real.

The honest reading: at river level the impact data endorse the adopted near-equal split; at season level a deliberate Deyr tilt is a working-group lever, quantified and ready, but it trades verified gauge-flood coverage for impact-weighted frequency — a values decision, not a statistical one.

Open items before a trigger report