Somalia Riverine Flood Trigger / supporting analysis

The ensemble agreement level

A point counts toward the trigger when the ensemble median crosses its level. The median is a choice: the rule could instead require 30%, 60% or 90% of the members. This note tests every agreement level the GloFAS v4 reforecast can resolve, on all four windows and both legs, to see whether any of them performs better than the median. September 2026.

1–11 of 11 agreement levels tested, on the 11-member v4 reforecast, 2003–2023
3 of 4 windows where the action leg gives identical results at every level
0 extra severe years caught at any level, on either leg
6 of 11 the median rule: 50% of members, rounded up to a whole member

The setup

Each agreement level k means: a point counts on a day when at least k of the 11 reforecast members put that day over the point's level, on one forecast issue. Everything else stays at the live rules: same points, the window's point count reached on the same issue and the same valid day (as the monitoring pipeline counts votes), same season months, action at leads 1–7 and readiness at leads 8–12, and the reanalysis levels the pipeline reads at each leg's return period (readiness capped at 1-in-5). The median rule is k = 6: the median crosses when at least half the members do, and half of 11 rounds up to 6. Each level is scored against the two-gauge benchmark on severe years (1-in-5, levels fitted 2000–2023), 2003–2023.

The sweep

Heatmap of severe-year F1 per window and leg across the eleven agreement levels
Severe-year F1 per window at each agreement level, readiness leg above, action below. The red box marks the median rule. On the action leg, three of the four windows give the same activations, detection and false alarms at every level from 1 of 11 to 11 of 11. With 3 to 5 severe events per window, any difference between these scores is one or two events.
What the sweep shows.

The years in detail

The table lists every year that activates at any agreement level on either leg, together with every severe year in the archive span. Raising the level can only remove crossings, never add them, so each year is summarised by the highest level at which it still activates: an entry of at every level means the activation survives even all 11 members being required, and no against a severe year means it was not caught even at 1 of 11. Most years fall at one of those two extremes; the handful in between are where the level choice has any effect at all.

windowyeargauge benchmarkreadiness activatesaction activates
Juba Deyr2005below 1-in-3up to 1 of 11at every level
2006severe (1-in-5)nono
2014severe (1-in-5)up to 1 of 11at every level
2015below 1-in-3up to 3 of 11at every level
2017flood (1-in-3)up to 7 of 11at every level
2023severe (1-in-5)at every levelat every level
Shabelle Deyr2006severe (1-in-5)nono
2010below 1-in-3noat every level
2013below 1-in-3up to 7 of 11at every level
2014severe (1-in-5)up to 3 of 11at every level
2019severe (1-in-5)at every levelat every level
2020severe (1-in-5)noat every level
2023severe (1-in-5)nono
Juba Gu2005flood (1-in-3)noat every level
2010flood (1-in-3)up to 1 of 11at every level
2013below 1-in-3up to 1 of 11no
2016flood (1-in-3)up to 1 of 11up to 6 of 11
2018severe (1-in-5)up to 1 of 11at every level
2020severe (1-in-5)up to 9 of 11at every level
2023severe (1-in-5)nono
Shabelle Gu2005flood (1-in-3)at every levelat every level
2010flood (1-in-3)up to 10 of 11at every level
2016severe (1-in-5)at every levelat every level
2018flood (1-in-3)up to 1 of 11at every level
2020severe (1-in-5)nono
2023severe (1-in-5)nono

The same question at full ensemble resolution

The reforecast's 11 members leave an opening: a 30% rule and the median sit one member apart there, and a real minority signal could hide between the steps. The operational archive closes that opening for one season. Gu 2024 was severe at the Shabelle gauges, all three points over their 1-in-5. In Gu, GloFAS carries only the readiness leg (Google carries action), and the operational v4 forecasts, 50 members issued daily, do not reach it under the median rule. At full resolution: across every issue at leads 8–12 of the season, not one of the 50 members crossed the 1-in-5 level at Belet Weyne (the highest member reached 509 m³/s against a 663 level) or at Bulo Burti (515 against 645). Jowhar alone shows a tail signal, 4 of 50 members over its 647 level for 6 May in the issue of 28 April, but one point cannot carry a 2-of-3 rule. An agreement level of one member in fifty would not have reached readiness for Gu 2024.

Deyr 2023, the largest flood in the record and the season the trigger is most often asked about, gives the same answer more sharply. The operational archive covers all three Shabelle points for every issue day from 1 October to 31 December. Across the whole season the highest value any of the 50 members produced at leads 1–7 was 766 m³/s at Belet Weyne and 764 at Bulo Burti, against 1-in-5 levels of 1,109; at Jowhar one member reached 1,109 against 1,159 and no other came near. At leads 8–12 the highest members were 791, 767 and 1,095. The member share over the level is 0% at every point on every day, on both legs. The forecast envelope peaked in the first week of October and fell through November while the river was reaching its highest recorded levels: SWALIM's alert came on 20 October, the second gauge crossed its 1-in-3 on 6 November and its 1-in-5 on 19 November, and on none of those days did a single member sit over the level. On the Juba the same model had Deyr 2023 as the largest season in 25 years at all four points, so the failure is specific to the Shabelle, and it is the model's, not the rule's.

Deyr 2023: the full 50-member operational ensemble at Belet Weyne, Bulo Burti and Jowhar stays under the 1-in-5 level all season
Deyr 2023 on the operational v4 ensemble at the three Shabelle points: shaded, the full 50-member range across leads 1–7 for each valid day; line, the ensemble median of the most alarming issue; dashed, the 1-in-5 level; dotted, SWALIM's alert (20 Oct) and the day the second gauge crossed its 1-in-3 (6 Nov) and 1-in-5 (19 Nov).

The Juba archive for the same season, complete for all four points from 1 October to 31 December, shows the other side. Under the median rule the action leg activates on 21 October, the day after SWALIM's alert and four days before the gauges crossed 1-in-3 and 1-in-5 together on 25 October, with all four points over their 1-in-4 levels on the same day. Every one of the 50 members was over the level at every point that day, so any agreement level from 10% to 50% activates on 21 October. The readiness leg (leads 8 to 12) is reached with the median on the issue of 23 October, for 30 October, and on the issue of 22 October at agreement levels of 2% to 10%, after the action leg rather than before it. The median peaks at 268% to 329% of the levels on 21 and 22 November. On the Juba the agreement level makes no difference to the date; on the Shabelle no level would have produced a flag.

What about Google for Gu?

The Google GRRR reforecast is a single trace per issue and lead: the archive has no ensemble members, so there is no member share to choose. The Gu legs count a point when the forecast value crosses its level, which is what the operational check on the trigger page replays. The live Flood Hub display does show uncertainty bounds, but they are not in the public archive; with API access they could be logged from now on, and this question asked of Google once a few seasons accumulate. Live, the Gu action leg reads Google's single forecast value and only the Gu readiness leg reads the GloFAS ensemble, where the member share applies as above.

What this means for monitoring

The median rule stands. On the action leg no level performs better than the median on any window; Juba Gu scores higher from 7 of 11 upwards only by dropping 2016, a flood year. On the readiness leg a lower level catches more severe years (at 1 of 11: Juba Gu 2018, Juba Deyr 2014, Shabelle Deyr 2014) and adds activations outside the benchmark; a higher level drops Juba Gu 2020, a severe year. One agreement level for both legs and all four windows is kept. The daily monitoring chart draws each point's ensemble median as a share of its level, so the operating rule is the median crossing the dashed line.

Caveats. The reforecast has 11 members, so agreement levels can only be tested in steps of 9%; the operational forecast has 50 members and reads finer, but the Gu 2024 and Deyr 2023 checks above reach the same conclusion at that resolution. These scores use the full 2003–2023 archive, so they differ slightly from the trigger page's swap heatmap, which scores only the years Google's archive also covers. The readiness rows use the same reanalysis levels as the adopted readiness leg, capped at 1-in-5, so the k = 6 row is the mechanism table's readiness record. And each window has 3 to 5 severe events: every difference on this page is a one-or-two-event difference.
Supporting analysis for the one-model-per-window trigger. Agreement levels tested on the GloFAS v4 reforecast (1 control + 10 perturbed members, twice-weekly issues, leads 1–12 days, 2003–2023), levels fitted on the v4 reanalysis seasonal maxima (Weibull, 1999–2023) at each leg's return period, benchmark from SWALIM gauge levels (two-gauge consensus). Juba points juba_07/juba_08 sit outside the reforecast box and are excluded, as in the operational backtest. P. Wairimu · September 2026.