The setup
Each agreement level k means: a point counts on a day when at least k of the 11 reforecast members put that day over the point's level, on one forecast issue. Everything else stays at the live rules: same points, the window's point count reached on the same issue and the same valid day (as the monitoring pipeline counts votes), same season months, action at leads 1–7 and readiness at leads 8–12, and the reanalysis levels the pipeline reads at each leg's return period (readiness capped at 1-in-5). The median rule is k = 6: the median crosses when at least half the members do, and half of 11 rounds up to 6. Each level is scored against the two-gauge benchmark on severe years (1-in-5, levels fitted 2000–2023), 2003–2023.
The sweep
- On the action leg no level catches an extra severe year. Detection does not change with k on any window: even one member of eleven, at the required number of points on the same issue and day, does not reach the levels in the missed years (Juba Gu 2023, Juba Deyr 2006, Shabelle Gu 2020 and 2023, Shabelle Deyr 2006 and 2023). Those misses come from the model's magnitudes.
- At leads 8–12 the level changes the record in both directions. The ensemble spread is wider at readiness leads. Below the median, 1 of 11 catches severe years the median misses: Juba Gu adds 2018, Juba Deyr and Shabelle Deyr add 2014, and Shabelle Gu adds 2018, a flood year; it also adds activations outside the benchmark (Juba Gu 2013, Juba Deyr 2005 and 2015, Shabelle Deyr keeps 2013). Above the median, 8 of 11 drops Juba Deyr 2017 and Shabelle Deyr 2013, 10 of 11 drops Juba Gu 2020, a severe year, and 11 of 11 drops Shabelle Gu 2010.
- At leads 1–7 the level barely matters. The ensemble is tight at short leads: most members cross together or not at all. The one exception is Juba Gu, where levels above the median drop the 2016 activation, a flood year but not a severe one on that window.
The years in detail
The table lists every year that activates at any agreement level on either leg, together with every severe year in the archive span. Raising the level can only remove crossings, never add them, so each year is summarised by the highest level at which it still activates: an entry of at every level means the activation survives even all 11 members being required, and no against a severe year means it was not caught even at 1 of 11. Most years fall at one of those two extremes; the handful in between are where the level choice has any effect at all.
| window | year | gauge benchmark | readiness activates | action activates |
|---|---|---|---|---|
| Juba Deyr | 2005 | below 1-in-3 | up to 1 of 11 | at every level |
| 2006 | severe (1-in-5) | no | no | |
| 2014 | severe (1-in-5) | up to 1 of 11 | at every level | |
| 2015 | below 1-in-3 | up to 3 of 11 | at every level | |
| 2017 | flood (1-in-3) | up to 7 of 11 | at every level | |
| 2023 | severe (1-in-5) | at every level | at every level | |
| Shabelle Deyr | 2006 | severe (1-in-5) | no | no |
| 2010 | below 1-in-3 | no | at every level | |
| 2013 | below 1-in-3 | up to 7 of 11 | at every level | |
| 2014 | severe (1-in-5) | up to 3 of 11 | at every level | |
| 2019 | severe (1-in-5) | at every level | at every level | |
| 2020 | severe (1-in-5) | no | at every level | |
| 2023 | severe (1-in-5) | no | no | |
| Juba Gu | 2005 | flood (1-in-3) | no | at every level |
| 2010 | flood (1-in-3) | up to 1 of 11 | at every level | |
| 2013 | below 1-in-3 | up to 1 of 11 | no | |
| 2016 | flood (1-in-3) | up to 1 of 11 | up to 6 of 11 | |
| 2018 | severe (1-in-5) | up to 1 of 11 | at every level | |
| 2020 | severe (1-in-5) | up to 9 of 11 | at every level | |
| 2023 | severe (1-in-5) | no | no | |
| Shabelle Gu | 2005 | flood (1-in-3) | at every level | at every level |
| 2010 | flood (1-in-3) | up to 10 of 11 | at every level | |
| 2016 | severe (1-in-5) | at every level | at every level | |
| 2018 | flood (1-in-3) | up to 1 of 11 | at every level | |
| 2020 | severe (1-in-5) | no | no | |
| 2023 | severe (1-in-5) | no | no |
The same question at full ensemble resolution
The reforecast's 11 members leave an opening: a 30% rule and the median sit one member apart there, and a real minority signal could hide between the steps. The operational archive closes that opening for one season. Gu 2024 was severe at the Shabelle gauges, all three points over their 1-in-5. In Gu, GloFAS carries only the readiness leg (Google carries action), and the operational v4 forecasts, 50 members issued daily, do not reach it under the median rule. At full resolution: across every issue at leads 8–12 of the season, not one of the 50 members crossed the 1-in-5 level at Belet Weyne (the highest member reached 509 m³/s against a 663 level) or at Bulo Burti (515 against 645). Jowhar alone shows a tail signal, 4 of 50 members over its 647 level for 6 May in the issue of 28 April, but one point cannot carry a 2-of-3 rule. An agreement level of one member in fifty would not have reached readiness for Gu 2024.
Deyr 2023, the largest flood in the record and the season the trigger is most often asked about, gives the same answer more sharply. The operational archive covers all three Shabelle points for every issue day from 1 October to 31 December. Across the whole season the highest value any of the 50 members produced at leads 1–7 was 766 m³/s at Belet Weyne and 764 at Bulo Burti, against 1-in-5 levels of 1,109; at Jowhar one member reached 1,109 against 1,159 and no other came near. At leads 8–12 the highest members were 791, 767 and 1,095. The member share over the level is 0% at every point on every day, on both legs. The forecast envelope peaked in the first week of October and fell through November while the river was reaching its highest recorded levels: SWALIM's alert came on 20 October, the second gauge crossed its 1-in-3 on 6 November and its 1-in-5 on 19 November, and on none of those days did a single member sit over the level. On the Juba the same model had Deyr 2023 as the largest season in 25 years at all four points, so the failure is specific to the Shabelle, and it is the model's, not the rule's.
The Juba archive for the same season, complete for all four points from 1 October to 31 December, shows the other side. Under the median rule the action leg activates on 21 October, the day after SWALIM's alert and four days before the gauges crossed 1-in-3 and 1-in-5 together on 25 October, with all four points over their 1-in-4 levels on the same day. Every one of the 50 members was over the level at every point that day, so any agreement level from 10% to 50% activates on 21 October. The readiness leg (leads 8 to 12) is reached with the median on the issue of 23 October, for 30 October, and on the issue of 22 October at agreement levels of 2% to 10%, after the action leg rather than before it. The median peaks at 268% to 329% of the levels on 21 and 22 November. On the Juba the agreement level makes no difference to the date; on the Shabelle no level would have produced a flag.
What about Google for Gu?
The Google GRRR reforecast is a single trace per issue and lead: the archive has no ensemble members, so there is no member share to choose. The Gu legs count a point when the forecast value crosses its level, which is what the operational check on the trigger page replays. The live Flood Hub display does show uncertainty bounds, but they are not in the public archive; with API access they could be logged from now on, and this question asked of Google once a few seasons accumulate. Live, the Gu action leg reads Google's single forecast value and only the Gu readiness leg reads the GloFAS ensemble, where the member share applies as above.
What this means for monitoring
The median rule stands. On the action leg no level performs better than the median on any window; Juba Gu scores higher from 7 of 11 upwards only by dropping 2016, a flood year. On the readiness leg a lower level catches more severe years (at 1 of 11: Juba Gu 2018, Juba Deyr 2014, Shabelle Deyr 2014) and adds activations outside the benchmark; a higher level drops Juba Gu 2020, a severe year. One agreement level for both legs and all four windows is kept. The daily monitoring chart draws each point's ensemble median as a share of its level, so the operating rule is the median crossing the dashed line.