They are anchored beside the passage each one concerns. Nothing in the trigger design has been changed — the configuration, the numbers and the figures are exactly as generated. What has been corrected is text that contradicted the data, plus the broken provider switch; supporting evidence has been folded into collapsible blocks.
- R1 — the “7 of 8” headline rests on a lookahead in the benchmark
- R2 — every published lag is wrong; correlations understated
- R3 — three provider sets were really two
- R4 — the Juba leg never decides anything
- R5 — the per-window search does not beat an un-searched rule
- R6 — one model per window is defensible; mixing gains nothing
- R7 — the seven-point set is load-bearing, and two points are unverifiable
- R8 — excluding GEOGloWS introduces the one false activation
- R9 — the two lead bands overlap at lead 7
- R10 — five of sixteen activations have no readiness phase
- R11 — 1999–2001 unassessable but still in every denominator
- R12 — some benchmark levels rest on ten annual maxima
The mechanism at a glance
Across 1999–2023 the envelope is identical to the Shabelle river on its own: there is no year in which the Juba activates and the Shabelle does not. Per-river rates are 1-in-4.3 (Juba) against 1-in-3.2 (Shabelle), so the equal-probability-per-river goal is not met. Both Gu windows also run on Google and fire in the same four years (2013, 2016, 2018, 2020), so the Gu half of the mechanism is effectively a single trigger counted twice.
Four river-season windows, each running on one forecast source rather than a mixture. Inside a window, every monitored point's flow is compared with its own return-period threshold and the window activates when enough points are over their thresholds on the same day. A season where the points cross on different days does not count. This rarely matters in practice: flood flows stay high for weeks, and in every past activation the points crossed within 0 to 5 days of each other, upstream gauges first (see "How far apart do the points cross?" below). At least two points must agree, so no single point releases the money, and never all of them, so one quiet point cannot block it.
All seven reporting-era points are monitored: four on the Juba (Luuq, Dollow, Bardheere, Bualle) and three on the Shabelle (Belet Weyne, Bulo Burti, Jowhar). No threshold, on either leg, sits below 1-in-3.
Restricting to the five gauges that still report would leave the Juba with only Luuq and Dollow, so the only available rule is 2 of 2 — unanimous, which this design's own “never all of them” constraint forbids — and it adds a false activation in 2015, moving the envelope to 1-in-2.9. Bardheere and Bualle are therefore what make a non-unanimous Juba rule possible. The cost is that an activation driven by those two points can never again be verified against an observation: Bardheere's record ends 2023-11-30 and Bualle's 2024-03-14.
| window | action trigger (leads 1-7 d) | leg RP | readiness (leads 7-12 d) | readiness RP |
|---|---|---|---|---|
| Juba Gu | Google GRRR: 3 of 4 points over their 1-in-5-yr thresholds | 6.5 yr | GloFAS v4 ens-median: 3 of 4 over 1-in-5-yr | 5.5 yr |
| Juba Deyr | GloFAS v5: 3 of 4 points over their 1-in-4-yr thresholds | 13.0 yr | GloFAS v4 ens-median: 3 of 4 over 1-in-4-yr | 4.4 yr |
| Shabelle Gu | Google GRRR: 2 of 3 points over their 1-in-6-yr thresholds | 6.5 yr | GloFAS v4 ens-median: 2 of 3 over 1-in-5-yr | 5.5 yr |
| Shabelle Deyr | GloFAS v5: 2 of 3 points over their 1-in-4-yr thresholds | 4.3 yr | GloFAS v4 ens-median: 2 of 3 over 1-in-4-yr | 4.4 yr |
How the two rivers behave together
The Shabelle activates more often: 8 of 25 years (1-in-3.2) against the Juba's 6 (1-in-4.3). There is no year in which the Juba activates and the Shabelle does not; the Shabelle alone adds 2014 and 2019. Six activations fall in the same season on both rivers, and in all six the Juba reaches its thresholds first, by 1 to 24 days:
| season | Juba crosses | Shabelle crosses | lead |
|---|---|---|---|
| Gu 2013 | 12 Apr | 6 May | Juba first, 24 d |
| Gu 2016 | 7 May | 12 May | Juba first, 5 d |
| Gu 2018 | 20 Apr | 5 May | Juba first, 15 d |
| Gu 2020 | 29 Apr | 30 Apr | Juba first, 1 d |
| Deyr 2006 | 29 Oct | 2 Nov | Juba first, 4 d |
| Deyr 2023 | 29 Oct | 9 Nov | Juba first, 11 d |
So the Juba never adds an activation year (review note R4), but in every shared season it crosses first. Since either window releases the full amount, the Juba sets the response date in those six years; the Shabelle activates more often and adds the years the Juba misses.
Each gauge's return level is fitted on 2000–2026 while crossings are counted only within 1999–2023, so three years of post-window data enter the labels. That is what makes 2008 severe: Luuq's Deyr 1-in-5 level is 5.831 m fitted through 2026 but 5.869 m fitted through 2023, and the 2008 peak is 5.840 m — a 9 mm margin on a level that moves 38 mm purely from adding future years. Fit strictly to the backtest window and the severe set is 7, not 8, and this configuration catches 7 of 7. The much-discussed “miss in 2008” is an artifact of the asymmetry, since trigger thresholds are fitted strictly on 1999–2023. Corrected on branch fix/corrected-benchmark-and-lag (not merged, and this page is unchanged); the corrected numbers are used on the design comparison page.
The obvious answer is to debias, and that is what cannot be done yet: the GEOGloWS forecast archive begins in July 2024, roughly two years, which is neither enough to refit thresholds on the forecast climatology (this design asks about twelve years for a 1-in-3) nor enough to validate a flow-duration-curve correction against anything. The team's Nepal technical note is the cautionary precedent — SFDC did not rescue event detection there. The working group therefore leaned to excluding GEOGloWS on 2026-08-22 and adopted that on 2026-08-28. It stays in this view as a comparison, and stays a candidate for reinstatement once two or three Gu and Deyr seasons of its forecasts exist.
- It agrees with the impact record on the big riverine years. 2006, 2018, 2019, 2020 and 2023 are all severe under the rule. Ranked on raw EM-DAT totals, though, the third-costliest flood year is 2015 (916,000 affected), which the rule does not call a flood at all — correctly, since that event was in Galgaduud, Mudug and Nugal rather than on either river. The agreement holds only once EM-DAT is filtered to Juba and Shabelle events.
- 1999 to 2001 cannot be assessed. The Juba had no gauge reporting and the Shabelle only one, so no consensus is possible and those years read as quiet rather than as unknown. Nothing activates before 2005, so no year is wrongly scored as a miss, but the benchmark effectively starts in 2002.
- 2021 is a genuine gap. EM-DAT records 400,000 people affected, yet only Bardheere on the Juba and Belet Weyne on the Shabelle crossed, so the rule reads no flood. Belet Weyne's peak that year sits exactly at bank full (8.30 m), where the gauge record is censored and the true level is unknown, so consensus can be understated in exactly the years that matter most.
- 2016 is the reverse case. Three of the four Juba gauges and all three Shabelle gauges crossed their levels in Gu, but EM-DAT records nobody affected. Gauge levels and recorded impact are not the same measurement.
No window can reach two reporting gauges in those years, so they are scored as “no flood” rather than as unknown, while 25 years remains the denominator for every return period quoted on this page. The assessable record is effectively 22 years. The same mechanism silently affects Shabelle Deyr in 2000 and 2001, where Jowhar was the only reporting gauge and did cross its level, and Juba Gu in 2005–2006.
The stated ceiling of “a quarter of the record” is applied to the forecast thresholds but not to the benchmark. Dollow has 10 Deyr annual maxima after 2000 and is still assigned a 1-in-5 level; Bualle has 16. Since a single gauge crossing decides whether the two-gauge rule is met, a level interpolated from ten values can decide whether a year counts as a flood at all.
What happens if Google Flood Hub goes away
Use the provider switch above. Removing Google returns the configuration to the all-providers one: GEOGloWS takes Juba Gu and GloFAS v5 takes the other three windows. The envelope stays at 1-in-3.2 (8 of 25 years) with the same 7 of 8 severe years, and it loses the 2013 activation, so on the calibration record dropping Google costs nothing measurable. The catch is that the replacement is GEOGloWS, the one model whose thresholds cannot yet be fitted on its own forecasts, so the fallback is weaker operationally than it looks on the backtest.
The “without Google” configuration is numerically identical to “all providers”, because when GEOGloWS is allowed to compete Google never wins a window. The third switch state has been removed, and the two competing switch controllers that were both live on the page (which corrupted the default view once you clicked) have been reduced to one.
The decision was taken by the working group (leaning this way on Friday 2026-08-22, adopted 2026-08-28) and the reasoning is sound — see the callout above for the quantified version. The cost should be recorded alongside it: the all-providers configuration activates in 8 years with no activation outside the benchmark, whereas the adopted one activates in the same number of years but fires in 2013, a year the two-gauge benchmark does not call a flood. Trading one false activation for a model that can actually be operated is a defensible trade; it is just better made explicitly.
- One source per river-season window, never a mixture — requested by WFP, who raised concerns about mixing models inside a window. Tested against the alternative it costs nothing measurable (review note R6).
- GEOGloWS excluded — working group, leaning that way 2026-08-22 and adopted 2026-08-28, because its forecasts run below its own retrospective and its two-year forecast archive is too short either to refit thresholds on or to validate a debias against.
- The envelope activates no more often than 1-in-3 — directive of 2026-08-27, applied to the union of the four windows rather than to each window.
How the model was chosen, one per river and season
Applying one model, one return period and one vote count uniformly to all four windows: Google GRRR, 1-in-6, 3 of 4 reproduces the adopted result exactly (8 activations, 1-in-3.2, 7 of 8 severe, the 2013 false alarm). GloFAS v5, 1-in-5, 3 of 4 gives 9 activations at 1-in-2.9 with the same 7 of 8 and zero activations outside the benchmark. The roughly 41,000 admissible per-window combinations recover nothing a one-line rule does not. This is consistent with the page's own statement that many assignments tie, but it is stronger: the per-window freedom is not adding measurable value. Partly answered by the swap test added 1 September, which I have reproduced and which holds: at the adopted return periods and vote counts, neither model alone can carry all four windows without breaching the 1-in-3 ceiling — GloFAS v5 everywhere and Google everywhere both land at 1-in-2.6 — so the mixed assignment is doing real work, and the Deyr case for GloFAS v5 is clear (Google brings three false activations on the Juba and misses a Shabelle severe year). Those conclusions are unchanged on the corrected benchmark. What remains of this note is narrower: the swap test holds the thresholds fixed at the adopted values, so it shows the assignment is defensible given those settings, not that the configuration as a whole beats a simpler one. A uniform rule at different settings — Google GRRR, 1-in-6, 3 of 4 in every window — still reproduces the adopted envelope exactly. Both are true: the assignment is justified, while the tuning around it is not what earns the result.
One source per river and season, so four choices. Four steps:
The constraint is a stakeholder requirement, not a data-driven one: WFP raised concerns about mixing models within a window. It is worth recording that it is cost-free on this record. Allowing every model to compete at every point inside a window (a mixed consensus over the same seven points) reaches 7 of 8 severe years at 1-in-3.7; the best single-model configuration reaches the same 7 of 8 at 1-in-3.2. Mixing buys no additional severe year, so the requirement can be honoured without giving anything up — and it buys a real operational simplification, since one model per window means one threshold transfer to validate rather than several.
- Build the candidate rules. For a window and a candidate model, put a threshold at every one of the river's points, taken from that model's own annual maxima (1-in-3 to 1-in-6). The window activates in a year when at least N points cross in that season.
- Score the envelope, not the window. The money is released when any of the four windows activates, so candidates are judged on that union: how often it activates, how many of the 8 severe years it catches, and how often it activates in a year with no recorded flood.
- Apply the constraints. Thresholds never below 1-in-3 nor above a quarter of the record; all seven points monitored; at least two must agree but never all of them; every window must activate at least twice in 25 years.
- Pick. Nearest 1-in-3 overall with the most severe years caught. Ties break on fewer no-flood activations, then on tracking correlation.
The per-point tracking numbers behind the choice (ρ and best lag of each model against the river's reference gauge):
The best-lag search shifts the gauge series by row position after filtering to season months, so any non-zero lag pulls rows from the adjacent year and always scores worse than lag 0. The result is that 46 of the 56 cells below report a lag of exactly 0, including Bualle against the Luuq reference gauge some 450 km upstream. Recomputed with a calendar-day shift, Belet Weyne against GloFAS v4 in Deyr peaks at +22 days with ρ 0.78, not at 0 days with ρ 0.67. Because tracking correlation is the final tie-break between models, that tie-break is currently running on a corrupted statistic. Corrected on branch fix/corrected-benchmark-and-lag (not merged, and this page is unchanged); the corrected numbers are used on the design comparison page.
Per-station selection detail (ρ / lag per model, all four windows)
Why the models perform the way they do
The main difference between the models is magnitude. Each reproduces the shape of the flood tail reasonably well, but over- or under-estimates its size by a roughly constant factor.
- GEOGloWS runs 3.8 to 10 times too high on the Shabelle and at Luuq, at every return period. It ranks wet and dry years reasonably well, but its absolute discharge is far off.
- GloFAS v5 is the closest to the observations: 1.12 to 1.25 at Belet Weyne, 1.24 to 1.42 at Bulo Burti, 0.83 to 0.87 at Luuq. Its predecessor v4 ran 1.7 to 3.8 times too high, and v5 improves the daily rank correlation at the same time (Belet Weyne 0.65 to 0.77), so the gain is not just a rescaling. v4 is not shown as a candidate anywhere on this page. It appears only as the readiness leg's reforecast, the sole archive at leads 7 to 12 days.
- Google GRRR is close on the Shabelle (0.78 to 1.11) and runs about 0.4 times observed at Luuq.
- Bardheere disagrees with everything (0.14 to 1.53 across the candidates). That points to problems with the gauge record there, not with the models.
How each model maps the RP3 events, gauge by gauge
The second question is whether a model is above its own 1-in-3 threshold when the gauge is above its own. That is the event the trigger has to catch.
GEOGloWS is not in the adopted set and is not shown here. Its numbers are under All providers.
The same comparison at the level the trigger operates: each window's vote rule, 3 of 4 points on the Juba and 2 of 3 on the Shabelle at the adopted return periods, with each model swapped in. POD counts severe years caught, the objective the envelope is sized on. FAR counts activations with no RP3 flood behind them:
The same test on the forecasts the trigger would actually run on, at action leads of 1 to 7 days. GloFAS v5 has no reforecast, so GloFAS is represented by the v4 ensemble median, its only lead-time archive; GEOGloWS has no archive before July 2024 and cannot be scored:
- On severe years, Gu is a tie. Both models catch 2018 and 2020 on the Juba and 2016 and 2020 on the Shabelle, and both miss Gu 2023.
- Deyr on Google is worse on both counts: three false activations on the Juba and a missed severe year (2020) on the Shabelle. GloFAS v5's Deyr record is clean.
- GloFAS v5 alone cannot meet the 1-in-3 ceiling. No v5-only configuration reaches 8 activations, and its nearest envelope fires once every 2.9 years. Some windows therefore have to run on Google, and Gu is where the swap costs nothing.
- Google is the only model with forecast-side evidence. Its reforecast (2016-2023) shows event skill at leads 1 to 7 on the Shabelle; the only GloFAS lead-time archive (v4) scores near chance there. GloFAS v5 has no reforecast. It keeps Deyr because its thresholds carry over to the live v5 forecast, whose climatology matches the reanalysis they were fitted on.
GEOGloWS maps the RP3 events least well of the models shown. It reaches 0.30 at Belet Weyne against GloFAS v5's 0.60, 0.11 at Luuq against 0.44 for both v5 and Google, and 0.25 at Bardheere, and it pairs those misses with the worst false-alarm rates on the board: 0.88 at Luuq and Bardheere, 0.62 at Belet Weyne. It is not uniformly worst: it ties v5 at Bulo Burti (0.44) and reaches 0.67 at Dollow, where every model does well. These detection numbers, together with the threshold problem described above, are why it is left out of the adopted design.
GloFAS v5 is best or joint-best at six of the seven gauges and has the cleanest false-alarm record (0.14 at Belet Weyne, 0.40 at Dollow), which is why it carries both Deyr windows. Its weakest gauge is Bualle, where it catches 0.29 against 0.43 for the others. Google matches v5 at Dollow (1.00), Luuq (0.44) and Bardheere (0.50) and is the only model with a usable forecast archive on the Shabelle, which is why it carries Gu.
No model does well at Jowhar (0.10 to 0.20, false alarms 0.71 and up): off-takes and marsh hydraulics decouple the local level from upstream discharge. Requiring two gauges keeps it from deciding anything on its own.
This shapes the mechanism in three ways. Thresholds are set on return periods, not on absolute discharge: with a flat tail ratio, matching frequency corrects the bias, so even a model running ten times too wet stays usable. A threshold is then only valid against the climatology it was fitted on, which is where GEOGloWS fails: its forecast climatology sits below its retrospective. Skill is flat across leads 1 to 7 days, because the signal is water already in the channel rather than forecast rainfall, so a one-week action window costs little accuracy. And no single gauge decides anything: Jowhar cannot be predicted from upstream discharge and Bardheere's record is suspect.
The seasonal-peak comparison is a static figure showing every candidate model. It is available under All providers.
Supporting evidence — seasonal peaks, model vs gauge
Seasonal peaks: model vs gauge
Carried over unchanged from the published multi-source study: this diagnostic compares models, so it does not depend on how many points vote or where the thresholds sit.
Daily rank correlation can flatter a model that merely tracks the seasonal cycle. The sharper diagnostic is one point per season-year: the model's seasonal peak against the observed seasonal peak at the reference gauge — does the model rank the years correctly? For display each model's peak is divided by its own 1-in-6 threshold (rank correlation is unaffected by the normalisation), so the quadrants read directly: top-right = the model ran over its own threshold in a year the gauge also ran high.
| window | peak Spearman ρ — reanalysis (n = 21–25) | peak ρ — reforecast ≤ 6 d | |||
|---|---|---|---|---|---|
| GEOGloWS | GloFAS v5 | Google GRRR | GloFAS v4 (proxy) | Google GRRR | |
| Juba Gu | 0.65 | 0.64 | 0.85 | 0.63 | 0.98 |
| Juba Deyr | 0.65 | 0.59 | 0.74 | 0.64 | 0.75 |
| Shabelle Gu | 0.68 | 0.80 | 0.85 | 0.64 | 0.74 |
| Shabelle Deyr | 0.49 | 0.77 | 0.78 | 0.39 | 0.81 |
Thresholds and calibration
Each monitored point gets its threshold from its own model's climatology: the Weibull plotting position of that point's seasonal maxima, so a model is judged on timing rather than on scale. 1-in-3 is the floor and a quarter of the record is the ceiling, which on 25 years allows 1-in-3 to 1-in-6.
How far apart do the points cross?
A window activates only when enough points are above their thresholds on the same day, not just in the same season. How tight are the crossings in practice?
Across every activation, the first and last point to cross fall 0 to 5 days apart, and they cross in downstream order: Dollow, Luuq, Bardheere, Bualle on the Juba; Belet Weyne, Bulo Burti, Jowhar on the Shabelle. Shabelle Gu 2018 had all three within a single day, Juba Gu 2018 four points inside three days. Because high flows persist for weeks, same-day overlap is reached even where the first crossings are a week or more apart, as in Juba Deyr 2023 (11 days between Dollow and Bualle, yet four points above at once).
The exception is instructive. Juba Deyr 2015 had Dollow crossing on 24 October and the other three between 10 and 13 November, 20 days later: two separate pulses rather than one flood wave. Same-day overlap never reached three, so the window did not activate, which is the behaviour we want.
Calibrated on the reanalysis, checked on the forecasts
Refitting the same rule on the forecast archives themselves, 2003–2023, thresholds from 1-in-3 to 1-in-5. Out of the field: Google GRRR, whose archive is too short to carry a 1-in-3 threshold, so it can only be calibrated on the retrospective.
| window | source | gauge threshold | points that must agree | would have activated in |
|---|---|---|---|---|
| Juba Gu | GloFAS v4 | 1-in-3 | 2 of 4 | 2005, 2010, 2013, 2016, 2018, 2020, 2023 |
| Juba Deyr | GloFAS v4 | 1-in-5 | 3 of 4 | 2015, 2017, 2023 |
| Shabelle Gu | GloFAS v4 | 1-in-4 | 2 of 3 | 2005, 2010, 2016, 2018 |
| Shabelle Deyr | GloFAS v4 | 1-in-5 | 2 of 3 | 2010, 2013, 2019, 2020 |
Envelope 1-in-2.2 (10 of 21 years), catching 5 of the 8 severe years. This is the closest rate the forecast record supports; nothing lands on 1-in-3.
Supporting evidence — the tuning surface
The tuning surface
The Nigeria-style grid search (per state there; per river × season here — the same recipe as P. Wairimu's notebook 08): every (per-pair RP threshold, N pairs required) cell is scored against the seasonal SWALIM benchmark. The adopted cell is not always the best single-window score — it is the best choice subject to the river-level constraints (near-equal per-river frequency across the Gu+Deyr union, severe-coverage-first, ≥ 3-yr overall enforced jointly across the rivers). With 4–5-pair pools the surface is small enough to read whole: the neighbors show what loosening a leg would cost in false activations, or tightening one in missed severe years.
Supporting evidence — official SWALIM levels vs the fitted ones
How the official SWALIM levels compare with the fitted ones
The benchmark years above are defined by empirical return periods of the
gauges' own seasonal maxima, not by SWALIM's published Moderate/High
flood-risk levels — and the check below is why. Putting the official levels on the
empirical RP scale (per active gauge and season, 1999–2023) shows they imply wildly
inconsistent frequencies: "Moderate" ranges from a level the river reaches most
years (Dollow Deyr, ~1-in-1.3; Jowhar Gu, ~1-in-1.6) to a genuinely rare one (Luuq
Gu, ~1-in-4.6), and "High" from ~1-in-1.8 (Dollow) to ~1-in-11.5 (Luuq Gu). They
are engineering levels of uncertain provenance, not a consistent severity scale —
the same conclusion P. Wairimu's notebook 01 flagged when the published
max_level at Belet Weyne came out below its bank-full level.
Everything in this mechanism therefore runs on each gauge's own empirical RP3/RP5
levels (table som_ms_swalim_rp); the official levels remain useful
only as familiar reference lines for readers of SWALIM bulletins.
The forecast-vs-own-reanalysis comparison is a static figure showing every candidate model. It is available under All providers.
Supporting evidence — forecast skill against each model's own reanalysis
Forecast skill against each model's own reanalysis
Carried over unchanged from the published multi-source study: this diagnostic compares models, so it does not depend on how many points vote or where the thresholds sit.
The assume-bias-for-all-sources rule, quantified: per model and lead time, how well the forecast reproduces the model's own reanalysis or retrospective — rank correlation (does it track itself?) and the median forecast/reanalysis ratio (is it biased against its own climatology?). Computed per station on flood-season days, median across the 18 stations.
Would it have worked operationally?
The calibration record is reanalysis — a model's afterwards-view of its own past. The operational test replays the historical forecasts: per pair and valid day, the most alarming ≤ 6-day signal (max over issue dates — with mixed providers in one window, a monitoring day combines each provider's latest forecast). Google pairs use their own reforecast (2016–2023); GloFAS v5 pairs the v4 reforecast (2003–2023, ensemble median, thresholds refit on v4's own climatology, since no v5 reforecast exists); GEOGloWS pairs the retrospective as a lead-0 stand-in (hindsight — flagged per window). That forecast skill barely decays from lead 1 to 7 — initial-condition-driven on these slow rivers — is P. Wairimu's notebook 03 finding, and is what makes both lead bands viable at all.
| window | calibrated on | run on | reproduces its calibration years? |
|---|---|---|---|
| Juba Gu | Google GRRR reanalysis | Google GRRR forecasts, leads 1-7 | yes, same archive |
| Juba Deyr | GloFAS v5 reanalysis | GloFAS v5 operationally; lead-time evidence from GloFAS v4 | cannot be shown directly: no reforecast for this version |
| Shabelle Gu | Google GRRR reanalysis | Google GRRR forecasts, leads 1-7 | yes, same archive |
| Shabelle Deyr | GloFAS v5 reanalysis | GloFAS v5 operationally; lead-time evidence from GloFAS v4 | cannot be shown directly: no reforecast for this version |
The readiness leg (7–12 days)
Readiness runs on GloFAS v4 ensemble-median forecasts at leads 7 to 12, the only archive covering that band, over the same full set of points, with thresholds refitted on that series. It releases only the mobilisation share and is held to the same floor: no threshold below 1-in-3.
| window | readiness rule | readiness years | covers action years |
|---|---|---|---|
| Juba Gu | 3 of 4 points over 1-in-5 | 2010, 2016, 2018, 2020 | 3 of 4 |
| Juba Deyr | 3 of 4 points over 1-in-4 | 2005, 2006, 2015, 2017, 2023 | 2 of 2 |
| Shabelle Gu | 2 of 3 points over 1-in-5 | 2005, 2010, 2013, 2016 | 2 of 4 |
| Shabelle Deyr | 2 of 3 points over 1-in-4 | 2007, 2013, 2014, 2019, 2020 | 4 of 6 |
Readiness carries each window's own rule, the same votes and the same return period, refitted on the 7-to-12-day series.
Action runs at leads 1–7 and readiness at leads 7–12, both inclusive, so lead 7 sits in both bands and the same forecast day can serve either leg. The original specification was action under 6 days. Separately, the “covers action years” column is a season-year set intersection: it records that both legs activated in the same season, not that readiness came first.
Counting across the four windows: Juba Gu 3 of 4 covered, Juba Deyr 2 of 2, Shabelle Gu 2 of 4, Shabelle Deyr 4 of 6. The uncovered activations include Shabelle Deyr 2023, the largest flood in the record, and Shabelle Gu 2018 and 2020. Under the staged model that would have been a design failure; under the 2026-08-27 decision it is accepted, but the working group should see the count before accepting it.
Return-period bookkeeping
| level | trigger | activations (1999-2023) | RP |
|---|---|---|---|
| individual | Juba Gu action | 4 — 2013, 2016, 2018, 2020 | 6.5 yr |
| individual | Juba Deyr action | 2 — 2006, 2023 | 13.0 yr |
| individual | Shabelle Gu action | 4 — 2013, 2016, 2018, 2020 | 6.5 yr |
| individual | Shabelle Deyr action | 6 — 2006, 2013, 2014, 2019, 2020, 2023 | 4.3 yr |
| river | Juba (Gu or Deyr) | 6 | 4.3 yr |
| river | Shabelle (Gu or Deyr) | 8 | 3.2 yr |
| overall | action, either river | 8 | 3.2 yr |
Under all-in funding, where any activation releases the full envelope, the effective return period is the overall row: 1-in-3.2 with Google, 1-in-3.2 without. The individual windows are set rarer than that on purpose, because four windows each calibrated to 1-in-3 give a union of roughly 1-in-1.5.
Activations, impact and response, year by year
The two-historical-records rule: the trigger record against impact, not just the hazard benchmark. Everything is grouped river → season, reverse-chronological. The trigger column is the trigger's actual decision variable: the peak number of monitored points simultaneously over their thresholds that season-year — the cell fills red when it meets the requirement shown in the season header (= the trigger activates). RP is the empirical return period of that season's maximum level at the SWALIM reference gauge — 5yr = a 1-in-5-year level or higher, 3yr = 1-in-3 to 1-in-5 (Weibull on the gauge's own seasonal maxima, per the threshold check above — not the official flood-risk levels); EM-DAT (purple) is people affected and CERF (blue) the flood allocation (US$), shaded by magnitude, each attributed to the river(s) named in the event or allocation narrative.
| loading… |
Attribution & caveats. EM-DAT is a floor (entry-criteria bias), split to seasons by start month (Mar–Jun → Gu, Sep–Dec → Deyr); an event naming both rivers is counted under both rivers (not split — hover for deaths and the multi-river flag); events naming neither river (mostly flash floods elsewhere) are excluded. A CERF allocation naming neither river in its narrative is shown under both rivers and marked * (they are national riverine-flood responses). Impact columns are context, not a skill score: a formal skill-vs-impact statistic still requires choosing an impact threshold.
Supporting analysis — should impact tilt the probabilities?
Should impact tilt the probabilities?
Near-equal per-river probability is a design choice, not a law — a river or season carrying systematically more impact could get a lower relative threshold. Aggregating the river-attributed records (1999–2023) answers whether the data call for it:
som_ms_impact_by_window). River totals are near-equal — Juba 47%,
Shabelle 53% on the three-record average — but the seasons are not:
Deyr carries ~65% of impact (and the two largest CERF flood
allocations) against Gu's ~35%.Re-running the calibration with the activation budget allocated by impact share
instead of equally (table som_ms_impact_tilt) makes the conclusion
concrete:
- A river tilt changes nothing. The 47/53 split rounds to the same 5/6 activation allocation the joint calibration already adopted — the adopted mechanism's slight Shabelle lean is already impact-consistent.
- A season tilt is possible but costs skill. Pushing Juba's budget toward Deyr (its higher-impact season) forces its Deyr leg to a lower threshold and its Gu leg higher: the tilted variant swaps documented Gu floods for extra Deyr activations, drops Juba's severe-year coverage from 0.50 to 0.38, and pushes the overall frequency to 1-in-2.9 — under the ≥ 3-year specification. The calibration already leans Deyr where the data support it (both Deyr legs run at lower pair-RP thresholds than the Gu legs); forcing more requires accepting misses the gauge record says are real.
The honest reading: at river level the impact data endorse the adopted near-equal split; at season level a deliberate Deyr tilt is a working-group lever, quantified and ready, but it trades verified gauge-flood coverage for impact-weighted frequency — a values decision, not a statistical one.
Open items before a trigger report
- Impact years are still undefined. The trigger is scored against gauge levels, not against recorded humanitarian impact. Until impact years exist, severe-year coverage is a proxy.
- The reanalysis cannot pick the model. Many assignments tie, so the model per window rests on lead-time skill and on operational considerations, not on the backtest.
- Google cannot be calibrated on its own forecasts. Its reforecast starts in 2016, and a 1-in-3 threshold needs about 12 years, so a Google window is necessarily calibrated on the retrospective and only checked at lead time.
- Two Juba points can no longer be verified. Bardheere's gauge record ends 2023-11-30 and Bualle's 2024-03-14. Both can still be forecast at, but neither can be checked against observations from here on.
- Readiness coverage is uneven by season, which is accepted rather than solved: at 7 to 12 days GloFAS v4 leads Gu activations more reliably than Deyr ones, so a Deyr activation may arrive with no readiness phase.