1-in-3.2envelope return period
7 of 7severe years caught
noneactivations outside the benchmark
Gu = Deyrseasons exactly level at 1-in-6.5
What this design is
Three constraints, applied together:
- Only the five gauges that still report — Belet Weyne, Bulo
Burti and Jowhar on the Shabelle, Luuq and Dollow on the Juba. Every activation is
checkable against a live observation at the point the forecast speaks for.
- Any model combination. Every (station, product) pair competes
on merit; a station may carry more than one product and a window may mix them.
Selection is model-blind: a −3 day timing guard, then every pair within
0.10 ρ of the window's best.
- Balance is a constraint, not an outcome. The search minimises
the gap between the basins and between the seasons, subject to the envelope
staying at or rarer than 1-in-3 and to catching every severe year.
How the probability comes out
The seasons come out exactly level and the basins within one activation.
Gu and Deyr each release once every 6.5 years. The Shabelle runs at 1-in-4.3 against
the Juba's 1-in-5.2 — a difference of a single activation across 25 years. Both
basins do real work: the Juba releases on its own in and the
Shabelle in , so neither leg is redundant.
Exact basin parity is reachable, and it costs a severe year.
Searching the whole option space, the best configuration with the basins exactly
equal puts both at 5 activations, moves the envelope to 1-in-3.7 and catches
6 of 7 severe years instead of 7. Perfect parity is therefore
available but not free: it buys symmetry by giving up a flood the mechanism
currently catches. Within one activation is as balanced as this station set gets
without paying for it, which is why that is what is adopted here.
Year by year
The number is how many pairs in that window were over their own thresholds at once;
the cell is shaded when it reaches the requirement in the column header, which is
when that window activates. bench is what the gauges recorded that
season.
How it compares
| design | stations | models | Juba | Shabelle |
Gu | Deyr | envelope | severe | false |
| This page — balanced | 5 active | any mix |
1-in-5.2 | 1-in-4.3 | 1-in-6.5 | 1-in-6.5 |
1-in-3.2 | 7 of 7 | none |
| Multi-source (as published) | 5 active | any mix |
1-in-5.2 | 1-in-4.3 | 1-in-6.5 | 1-in-6.5 |
1-in-3.2 | 7 of 7 | none |
| One model per window | 7 | one per window |
1-in-4.3 | 1-in-3.2 | 1-in-6.5 |
1-in-4.3 | 1-in-3.2 | 7 of 7 | 2013 |
| Balanced, 25 Aug (superseded) | 18 points | any mix |
1-in-4.3 | 1-in-4.3 | 1-in-6.5 | 1-in-6.5 |
1-in-3.2 | 7 of 7 | 2013 |
Where the balance went, and how much of it comes back. The 25 August
configuration was exactly balanced on both axes, but it drew on all 18 registry
points including several with no gauge at all. Restricting to the five that still
report — so that an activation can always be checked against an observation
— is what cost the exact basin match. Applying the balance constraint
explicitly on that restricted set recovers the seasons completely and holds the
basins to a single activation apart, while removing the 2013 activation the
25 August version carried outside the benchmark. It is the same configuration the
published multi-source page already adopts, now with the balance stated as a design
requirement and its cost quantified rather than left implicit.
Caveats
- Shabelle Gu is unanimous. Its pool is four pairs and the rule
needs all four, so a single product failing on the right day silences the window.
That is forced by the pool size on this station set, not chosen.
- Juba Deyr is very rare at 1-in-26 on its own; the Juba's
activation rate is carried almost entirely by Gu.
- GEOGloWS units are scored on the retrospective at lead 0, since
its forecast archive begins July 2024. Any lead-time claim resting on them is
unproven, and its live forecasts run at 0.89× its own retrospective —
so a level set as 1-in-5 behaves like a median 1-in-13 unless the thresholds are
refit or the forecast is debiased.
- Scored on the corrected benchmark (branch
fix/corrected-benchmark-and-lag): return levels fitted strictly
within 1999–2023, which removes 2008 from the severe set.
Provenance. Selection from the station-vs-gauge correlation table
(P. Wairimu, notebook 06) restricted to the five reporting SNRFA gauges; thresholds
are Weibull plotting positions on each product's own seasonal maxima, 1999–2023;
the benchmark is the two-gauge consensus rule with return levels fitted strictly
inside the backtest window. Compare against the
design comparison, the
multi-source mechanism and the
one-model-per-window mechanism.