Comparative analysis of the 2025 and 2026 Adamawa action triggers, and exploration of GloFAS reforecast configuration options for the 2026 readiness trigger. Evaluation uses Floodscan SFED annual maxima (1998–2023) as ground truth and the Weibull plotting position formula for empirical return periods throughout.
| 2025 trigger | 2026 trigger | |
|---|---|---|
| Logic | GloFAS reanalysis ≥ 3,132 m³/s OR GRRR reanalysis ≥ 1,195 m³/s at Wuroboki (single-gauge) | ≥ 6 of 10 GRRR gauges simultaneously exceed 4-yr empirical RP |
| Reanalysis trigger years | 1998, 1999, 2003, 2012, 2019, 2022 | 1998, 1999, 2012, 2018, 2019, 2022 |
The two triggers differ in exactly one year each: 2025 fires in 2003 (a 3-yr Floodscan event, not a 4-yr event); 2026 fires in 2018 (a 4-yr Floodscan event). Both miss 2015 and 2023.
Evaluated against Floodscan event years: 3-yr = {1998, 1999, 2003, 2012, 2015, 2018, 2019, 2022, 2023}; 4-yr = {1999, 2012, 2015, 2018, 2022, 2023}; 5-yr = {1999, 2012, 2015, 2022, 2023}.
| Benchmark | 2025 POD / FPR / F1 | 2026 POD / FPR / F1 |
|---|---|---|
| 3-yr | 0.67 / 0.00 / 0.80 | 0.67 / 0.00 / 0.80 |
| 4-yr | 0.50 / 0.15 / 0.50 | 0.67 / 0.10 / 0.67 |
| 5-yr | 0.60 / 0.14 / 0.55 | 0.60 / 0.14 / 0.55 |
FPR = FP / (FP + TN). At 3-yr and 5-yr benchmarks the triggers are statistically identical. The 2026 trigger is clearly better at the 4-yr level — its designed threshold — because it fires in 2018 (a genuine 4-yr event) instead of 2003 (only a 3-yr event).
Ten gauges selected by Spearman ρ rank (daily correlation with Floodscan SFED, Aug–Nov, lag range −7 to +14 days). All are GRRR (Google) source. Trigger fires if ≥6 of 10 gauges simultaneously exceed their individual 4-yr empirical RP threshold on at least one wet-season day.
| Gauge ID | Lat | Lon | QV | best ρ | lag (d) | 4-yr threshold (m³/s) | n trigger yrs |
|---|---|---|---|---|---|---|---|
| hybas_1120842990 | 9.385 | 12.806 | 0.742 | −3 | 1,110.8 | 5/6 | |
| hybas_1120843610 | 9.369 | 12.852 | 0.742 | −3 | 1,101.5 | 5/6 | |
| hybas_1120845060 | 9.331 | 12.715 | 0.732 | −3 | 1,102.2 | 6/6 | |
| hybas_1120849600 | 9.244 | 12.577 | 0.728 | −3 | 1,113.7 | 6/6 | |
| hybas_1120848550 | 9.265 | 12.494 | 0.726 | −3 | 1,106.0 | 6/6 | |
| hybas_1121970280 | 9.315 | 12.435 | 0.723 | −3 | 1,110.1 | 6/6 | |
| hybas_1120842550 | 9.394 | 12.369 | ✓ | 0.719 | −3 | 1,114.0 | 6/6 |
| hybas_1120840700 | 9.444 | 12.356 | 0.719 | −3 | 1,112.6 | 6/6 | |
| hybas_1120840560 | 9.448 | 12.348 | 0.719 | −3 | 1,117.1 | 6/6 | |
| hybas_1120840690 | 9.444 | 12.340 | 0.711 | +2 | 142.7 | 4/6 |
QV = quality-verified by Google. 4-yr threshold = empirical Weibull RP from wet-season annual maxima, 1998–2023. n trigger yrs = number of the 6 reanalysis trigger years in which this gauge contributed (i.e. was one of the ≥6 simultaneously exceeding on the trigger day).
This gauge is a clear outlier: its 4-yr threshold is 142.7 m³/s, roughly 8× lower than the other nine gauges (~1,100 m³/s). It also has a positive lag (+2 days), meaning it peaks after the Floodscan SFED peak rather than before, and contributed to only 4 of 6 trigger years. It likely represents a smaller tributary rather than the main Benue channel. Despite the low threshold, it is not disqualifying in the ≥6-gauge logic since it only provides one vote out of ten.
The trigger uses empirical Weibull RP thresholds, not Google’s official levels. The Google model provides warning / danger / extreme danger thresholds (approximately 2-, 5-, and 20-yr RP) for gauges with quality-verified models:
| Gauge ID | Warning (m³/s) | Danger (m³/s) | Extreme danger (m³/s) | QV |
|---|---|---|---|---|
| hybas_1120842550 | 1,094 | 1,352 | 1,648 | ✓ |
| hybas_1120840690 | 125 | 163 | 205 | |
| hybas_1120840700 | 1,084 | 1,341 | 1,637 | |
| hybas_1120840560 | 1,086 | 1,346 | 1,646 | |
| hybas_1120848550 | 1,095 | 1,354 | 1,652 | |
| hybas_1120842990 | 1,078 | 1,324 | 1,604 | |
| hybas_1120843610 | 1,079 | 1,328 | 1,607 | |
| hybas_1120845060 | 1,088 | 1,344 | 1,639 | |
| hybas_1120849600 | 1,091 | 1,348 | 1,643 | |
| hybas_1121970280 | 1,092 | 1,347 | 1,640 |
The empirical 4-yr thresholds (used in the trigger) are very close to the Google warning levels for most gauges (~1,078–1,117 vs ~1,078–1,095 m³/s), consistent with warning level ≈ 2-yr RP and danger level ≈ 5-yr RP in Google’s framework.
Computed using the Weibull formula: RP = (n + 1) / k.
| Trigger | n | k | Weibull RP |
|---|---|---|---|
| 2026 action (1998–2023) | 26 | 6 | 4.5 yr |
| Readiness, option A (2003–2022) | 20 | 7 | 3.0 yr |
| Readiness, option B (2003–2022) | 20 | 6 | 3.5 yr |
| Combined action + readiness A (1998–2023) | 26 | 8 | 3.4 yr |
| Combined action + readiness B (1998–2023) | 26 | 7 | 3.9 yr |
The readiness trigger should have a shorter RP than the action trigger to reflect its role as a more sensitive, lower-cost signal. Options A and B both satisfy this (3.0 yr and 3.5 yr vs action’s 4.5 yr).
The readiness trigger uses the GloFAS reforecast ensemble mean at Wuroboki. A year fires if the ensemble mean exceeds the threshold at any leadtime ≤ MAX_LT in the wet season (Aug–Nov). Evaluated over 2003–2022 (n = 20 years). Action trigger years with reforecast coverage: 2012, 2018, 2019, 2022 (4 years) — all included in POD.
| Metric | Value |
|---|---|
| Fires in | 2003, 2008, 2012, 2014, 2016, 2019, 2022 |
| k | 7 |
| Weibull RP | 3.0 yr |
| POD (vs action years 2012, 2018, 2019, 2022) | 0.75 (hits 2012, 2019, 2022; misses 2018) |
| FP (non-action years) | 4 (2003, 2008, 2014, 2016) |
| FPR | 25% |
Lead times vs action trigger: 2012 +2d, 2019 +46d, 2022 +13d, 2018 not detected.
| Metric | Value |
|---|---|
| Fires in | 2003, 2014, 2016, 2019, 2022 |
| k | 6 |
| Weibull RP | 3.5 yr |
| POD (vs action years 2012, 2018, 2019, 2022) | 0.50 (hits 2019, 2022; misses 2012, 2018) |
| FP (non-action years) | 3 (2003, 2014, 2016) |
| FPR | 19% |
Lead times vs action trigger: 2012 −15d (fires after), 2019 +46d, 2022 +3d, 2018 not detected.
| Threshold | k | RP | POD | FPR | 2012 lead | 2019 lead | 2022 lead |
|---|---|---|---|---|---|---|---|
| 3,132, LT≤13d (option A) | 7 | 3.0 yr | 0.75 | 25% | +2d | +46d | +13d |
| 3,250, LT≤12d (option B) | 6 | 3.5 yr | 0.50 | 19% | −15d | +46d | +3d |
The key trade-off is whether to capture 2012. The maximum GloFAS ensemble mean (LT≤13d) before the Aug 23 action date in 2012 was only 3,142 m³/s at LT=1d — meaning any threshold above 3,142 m³/s fires after the action trigger in 2012. The +2d lead in option A is therefore operationally marginal (nowcast, not forecast). Both options fail to detect 2018; this is an inherent GloFAS forecasting limitation rather than a threshold choice. Both options provide strong early warning in 2019 (+46d) and adequate warning in 2022 (+3–13d depending on option).
Lead times vary substantially: 46 days in 2019, 2–13 days in 2022, and essentially zero in 2012. Stakeholders should not expect a consistent preparation window — in some years the readiness trigger fires weeks early, in others it fires at the same time as the action trigger or not at all. The 2019 result should not be used as the typical expectation.
In 2012, the GRRR multi-gauge action trigger (Aug 23) fired before or at the same time as any useful GloFAS signal. The flood came on quickly and was not forecast far in advance by GloFAS. This is a known limitation of single-basin ensemble forecasts for flash-type events on the upper Benue.
At the recommended threshold, the readiness trigger fires in approximately 1 of every 3–3.5 years, compared to the action trigger’s 1 in 4.5 years. Roughly 3–4 readiness activations per 20 years will not be followed by an action trigger activation. At 5% funding, this is financially acceptable, but stakeholders should be briefed that a readiness activation does not guarantee a full activation will follow.
2018 is a 4-yr Floodscan event and an action trigger year, but the GloFAS reforecast ensemble mean only exceeds the threshold at LT≥14d in that year. GloFAS appears to have underforecast the 2018 event at shorter leadtimes. The readiness trigger cannot reliably serve as an early warning for 2018-type events.
These are genuine 4–5-yr flood years that neither trigger configuration detects. Investigating why these years are not captured (different flood pathway, upstream conditions, gauge coverage) is a priority for improving future trigger design.
The GloFAS reforecast only covers 2003–2022 (20 years), and the GRRR multi-gauge reforecast (2016–2022, 7 years) is too short to draw robust conclusions about reforecast lead times at the action trigger level. All reforecast performance figures carry substantial uncertainty.
The 2025 documentation reported FAR of 18% (3-yr) and 10% (5-yr). These likely reflect a longer evaluation period extending before 1998 (where pre-Floodscan trigger activations were counted as false alarms), a different event year set derived from a different Floodscan pixel dataset, and/or rank-based rather than Weibull RP thresholds. The figures are not directly comparable to the metrics computed here.
| RP level | GloFAS threshold |
|---|---|
| 3-yr | 2,701 m³/s |
| 4-yr | 3,009 m³/s |
| 5-yr | 3,110 m³/s |
| ~6.5-yr | 3,334 m³/s |