Somalia Riverine Flood Trigger / trigger analysis

The multi-source trigger mechanism

Proposed trigger for anticipatory action against riverine flooding on the Juba and Shabelle: a multi-model station consensus across GloFAS, Google Flood Hub and GEOGloWS, calibrated against SWALIM river gauges. August 2026.

Provider set
Adopted mechanism — GloFAS v5 + Google GRRR + GEOGloWS
You are viewing the no-Google variant. The whole report below — mechanism table, return periods, backtest and the year-by-year table — is recalibrated from scratch with Google Flood Hub excluded, using identical rules (active gauges, −3 d timing guard, 0.10 ρ floor, joint basin calibration at ≥ 3 yr overall). This is a resilience scenario for the working group, not the proposed mechanism. Narrative sections that do not depend on the provider set (selection method, threshold philosophy, GEOGloWS debias, impact records) read the same in both views.
1-in-3.2 overall action return period (either basin)
1-in-4.3 / 5.2 per basin, Shabelle / Juba
5 gauges every forecast point sits on a live SWALIM gauge
7–12 d readiness lead time (action at ≤ 6 days)

The mechanism at a glance

The trigger is fully forecast-based. SWALIM's river-gauge record is the ground truth the mechanism is calibrated and validated against — it is not a trigger input. Each river basin (Juba, Shabelle) carries one specific trigger per flood season (Gu, rains arriving on the rivers March–May; Deyr, October–December), and each specific trigger is staged:

windowaction trigger (leads ≤ 6 d)leg RP readiness (leads 7–12 d)readiness RP
Juba Gu Google GRRR ×2 + GloFAS v5 ×2 + GEOGloWS: ≥ 3 of 5 pairs over their 1-in-6-yr thresholds 6.5 yrGloFAS ens-median: 2/2 stations over 1-in-2-yr3.7 yr
Juba Deyr GloFAS v5 ×2 + GEOGloWS ×2: ≥ 3 of 4 pairs over their 1-in-5-yr thresholds 26.0 yrGloFAS ens-median: 2/2 stations over 1-in-2-yr4.4 yr
Shabelle Gu Google GRRR ×3 + GloFAS v5: all 4 pairs over their 1-in-6-yr thresholds 13.0 yrGloFAS ens-median: ≥ 2/3 stations over 1-in-3-yr3.7 yr
Shabelle Deyr GloFAS v5 ×3 + GEOGloWS ×2: ≥ 3 of 5 pairs over their 1-in-5-yr thresholds 6.5 yrGloFAS ens-median: 3/3 stations over 1-in-3-yr7.3 yr

Backtested on 1999–2023, the Shabelle activates 6 of 25 years (1-in-4.3) and the Juba 5 (1-in-5.2) — near-equal by construction (the two basins are now calibrated jointly: activation counts may differ by at most one, and the either-basin union must stay at ≥ 3 years). At least one basin activates in 8 years — overall return period 3.2 years, meeting the ≥ 3-year specification. The activation years are 2006, 2010, 2014, 2016, 2018, 2019, 2020 and 2023 — every one at least a 1-in-3 flood year at the reference gauges (zero activations outside the moderate benchmark set). Per-basin thresholds are nominal; the final adjustment belongs to the framework working group.

Backtested on 1999–2023, both basins activate 6 of 25 years (1-in-4.3, exactly equal) and at least one activates in 7 — overall return period 3.7 years, still meeting the ≥ 3-year specification with room to spare. The activation years are 2005, 2010, 2016, 2018, 2019, 2020 and 2023, again all within the 1-in-3 gauge benchmark set. So the headline answer is yes, a workable mechanism survives without Google — but see what it costs, below.

What happens without Google Flood Hub

Google Flood Hub is the only provider here without a public inter-agency commitment behind it, so the working group asked what the mechanism looks like if it disappears. The switch at the top of this page recalibrates the entire report — not by deleting Google's votes from the adopted configuration, but by re-running selection, the RP × N grid and the joint basin calibration from scratch over GloFAS v5 + GEOGloWS only. Side by side:

adopted (three providers)variant (no Google)
Overall RP1-in-3.2 (8 yrs)1-in-3.7 (7 yrs) — still meets spec
Per basinShabelle 1-in-4.3, Juba 1-in-5.2both 1-in-4.3 (exactly equal)
Severe-year coverage Juba 4/8, Shabelle 6/7 Juba 5/8, Shabelle 4/7
False activations00
Smallest pool4 pairs (Shabelle Gu) 1 pair (Shabelle Gu — no consensus at all)
Windows changed both Gu legs rebuilt; Juba Deyr identical; Shabelle Deyr retuned (same pool, new RP/N)
The real cost is concentration, not frequency. The variant's return periods look fine — arguably tidier, since the basins come out exactly equal and the Juba even gains a severe year. Two things to weigh against that. (1) Shabelle Gu collapses to a single station-model: GloFAS v5 at Belet Weyne, the only pair clearing the timing guard once Google is gone (every GEOGloWS pair there trails the gauge by 4–10 days). A one-pair trigger has no consensus protection — a single model glitch is the whole decision, and the N-of-M design exists precisely to avoid that. (2) Shabelle severe coverage drops from 6/7 to 4/7, losing the 2006 and 2014 Deyr floods. Google's short-lead skill on the Shabelle is doing real work in the adopted mechanism. Read together: the variant is a viable contingency, not an equivalent design — and if it were ever adopted, Shabelle Gu would need either a relaxed timing guard (accepting GEOGloWS's lag) or an explicit single-source risk acceptance.
Two-panel map of Somalia showing the Juba and Shabelle rivers with the 18 monitoring stations, colored by the models that carry them in the Gu and Deyr windows; dual-model stations carry an outer ring
The station network. 18 registry points (GloFAS-grid-verified coordinates with per-provider ID mappings — P. Wairimu, src/constants.py) on the two rivers, of which only the five with a live SNRFA gauge can carry pairs; stars are the SWALIM reference gauges (Luuq, Belet Weyne) the mechanism is calibrated against; numbers are each station's best selection ρ. An outer ring means the station carries more than one model: in Gu, Dollow (all three) and Luuq on the Juba and Belet Weyne on the Shabelle; in Deyr, Dollow, Luuq, Belet Weyne and Bulo Burti (all GloFAS + GEOGloWS).

How the pairs were chosen

Every combination of the five active SWALIM gauges and the three providers competes on one metric: best-lag Spearman rank correlation between the provider's reanalysis/retrospective discharge and the SWALIM reference-gauge level (Belet Weyne for the Shabelle, Luuq for the Juba), per river and season — the correlations computed in P. Wairimu's notebook 06, against the SWALIM benchmark her notebook 01 established (post-2000 record, official Moderate/High risk levels) and her notebook 05 showed must be primary over FloodScan on the Shabelle. Pairs compete freely (decision 2026-08-25: the goal is the best forecasts on the right basin, and a station may carry two models), under three rules — none of which knows what a model is:

Scatter plots of correlation versus lag for every station-model pair, by river and season; same-station models connected by grey lines, selected pairs ring-marked and labeled, timing guard shown as a dotted line
The candidate pool. Positive lag = the model's signal precedes the reference gauge; the dotted red line is the timing guard; small pale-grey dots are pairs at points with no active gauge — ineligible regardless of score. A grey line connects one station's three models, and ring-marked, labeled dots are the pairs that clear the floor. Google dominates Gu, GloFAS v5 dominates Deyr; GEOGloWS clears the floor in every window except Shabelle Gu, where it trails the gauge by 4–10 days and no guard-passing pair exists.

Multi-model representation is earned, not injected: GEOGloWS clears the floor in three of the four windows — Juba Deyr (Dollow ρ 0.83 at +1 d and Luuq 0.82, ahead of both Google pairs), Juba Gu (Dollow 0.83, floor 0.77) and Shabelle Deyr (Belet Weyne 0.75 and Bulo Burti 0.74, floor 0.73). In Shabelle Gu it has no pair passing the timing guard (−4 to −10 d) and stays out — as does every GloFAS pair there except Belet Weyne (−3 d, exactly on the guard), which is why that window is Google-heavy. An earlier draft used an explicit diversity rule to the same effect; it proved unnecessary and was dropped. The full per-station numbers (greyed rows = no active gauge):

Per-station selection detail (ρ / lag per model, all four windows)
loading…
What GEOGloWS votes can and cannot prove. GEOGloWS has no historical forecast archive before July 2024, so its pairs are calibrated and backtested through the retrospective (the agreed reanalysis-as-stand-in convention) — hindsight, not lead time. Operationally its live forecasts do exist (leads to 14 days), but they run below its own retrospective climatology (quantified below) — so the retrospective-fitted thresholds here cannot be applied to its live forecasts as-is. Two debias routes for go-live: refit thresholds on its forecast climatology (needs years of archive), or map the live forecast onto the retrospective climatology by flow-duration-curve quantile mapping (SFDC — the GEOGloWS group's own bias-correction method, which the team already applied in the Nepal technical note, June 2026) fitted on the 2024+ overlap. Honest caveat from that note: SFDC did not rescue GEOGloWS's event detection in Nepal, where the forecast ran at ~half its retrospective — Somalia's bias is far smaller (0.85–0.91×) and GEOGloWS holds minority votes here, not the trigger basis, but the mapping must be validated before its votes count operationally. Until then, windows that include GEOGloWS pairs carry up to that many votes of unproven lead-time skill.

Seasonal peaks: model vs gauge

Daily rank correlation can flatter a model that merely tracks the seasonal cycle. The sharper diagnostic is one point per season-year: the model's seasonal peak against the observed seasonal peak at the reference gauge — does the model rank the years correctly? For display each model's peak is divided by its own 1-in-6 threshold (rank correlation is unaffected by the normalisation), so the quadrants read directly: top-right = the model ran over its own threshold in a year the gauge also ran high.

Scatter plots of model seasonal peaks over own threshold versus SWALIM observed seasonal peaks, by basin and season, all three providers
Calibration record (reanalysis/retrospective, 1999–2023), at the reference station. Dark-edged points are selected pairs (the reference station carries two models in most windows). Vertical dashes: the gauge's Moderate (amber) and High (red) levels; horizontal dash: the model at its own 1-in-6 threshold.
windowpeak Spearman ρ — reanalysis (n = 21–25) peak ρ — reforecast ≤ 6 d
GEOGloWSGloFAS v5Google GRRR GloFAS v4 (proxy)Google GRRR
Juba Gu0.650.640.85 0.630.98
Juba Deyr0.650.590.74 0.640.75
Shabelle Gu0.680.800.85 0.640.74
Shabelle Deyr0.490.770.78 0.390.81
Scatter plots of reforecast seasonal peaks versus SWALIM observed peaks, by basin and season
Operational record: the ≤ 6-day reforecast signal's seasonal peaks, normalised by a 1-in-6 threshold fitted on the reforecast's own seasonal maxima (GRRR's 8-season archive makes its own threshold estimate wide — treat its vertical placement, not its ranking, with caution).

Does the selection metric matter? Daily ρ vs peak ρ vs lead time

Nigeria's recipe: select on daily wet-season lag-optimised Spearman ρ, use annual-peak ρ and peak-timing offset as post-selection guards, and let lead time enter through the reforecast checks rather than the selection score. Here the question was put to the data directly: every candidate pair scored on both metrics, and the full pool → grid → basin calibration re-run under three selection rules — daily-ρ floor (adopted), peak-ρ floor, and their mean, each with the same timing guard.

Four scatter plots, one per window, of daily correlation versus seasonal-peak correlation for every candidate station-model pair; selected pairs dark-edged, timing-guard failures shown as crosses
Every candidate pair on both metrics. The two measure different things — within-window correlation between them ranges from +0.61 (Juba Gu) to −0.50 (Juba Deyr, where the pairs that best track daily levels are worst at ranking the extreme years). Crosses fail the −3 d timing guard; dark-edged points are the selected pool.
Result: near-identical, with one honest wobble. Recalibrating under a peak-ρ floor (or the mean of both metrics) reproduces the Shabelle exactly and leaves every return period unchanged at every level — but on the Juba it swaps one activation year, 2010 (daily-ρ selection) for 2013 (peak-ρ), both 1-in-3 gauge years (table som_ms_metric_robustness). With the active-gauge restriction the pools are 4–5 pairs, so a metric change that reshuffles one pair can move one marginal year — under the previous 8–12-pair pools the three metrics were exactly identical. Daily ρ stays as the primary metric because it is the far more stable estimator (~2,000 daily observations against ~22 annual peaks, where a single flood year moves peak-ρ by ~0.1); peak-ρ is published per pair as a diagnostic (figure above and the som_ms_pair_metrics table), and lead time is already in the design twice — the −3-day timing guard at selection, and the reforecast backtests at the operational lead bands. A composite score was considered and rejected: it would add a tunable weight to arbitrate a single marginal year.
An honest tension the working group should see. On the peak view, Google ranks seasonal peaks best in all four windows at the reference station — including the Deyr windows its daily tracking loses (in Juba Deyr its daily ρ is 0.58–0.60 against GloFAS v5's 0.85–0.87, and it misses the floor). Daily correlation and peak ranking disagree there. The adopted Deyr legs rest on all-pair daily tracking plus event coverage; whether a Google-heavier Deyr design would be more robust is a candidate refinement for the review, not a settled question.

Thresholds and calibration

Each selected pair gets its threshold from its own model's climatology: the Weibull plotting position of its seasonal maxima (1999–2023) at the target return period. Magnitude bias between models cancels out — each pair is equally likely to exceed its threshold by construction. The same principle carries into operations: we assume bias between reanalysis and forecast for every source (quantified in the next section), so a threshold is only ever applied to the product whose own historical values it was fitted on. To inspect the indicators themselves, year by year, use the indicator explorer. The basin-level search then chooses, per season, the per-pair return period and the consensus count N — with the two basins picked jointly: per-basin activation counts of 4–6 differing by at most one (near-equal probability), the either-basin union at ≥ 3 years overall (the specification, enforced rather than hoped for), maximum coverage of the basins' severe (1-in-5) benchmark years, fewest activations outside the moderate (1-in-3) benchmark set, and, among ties, the more frequent option.

Heat strip of activation years per trigger leg, per basin versus severe years, and overall, 1999 to 2023
The calibration backtest. Basin rows compare activations (amber/blue) with severe 1-in-5 years at the SWALIM reference gauge (blue = caught, red = missed). Shabelle catches 6 of 7 severe years (misses only 2021); Juba 4 of 8 (misses 2005, 2006, 2011, 2014 — though 2006 and 2014 are caught overall via the Shabelle). Neither basin activates outside the moderate set. The gauge restriction costs Juba severe coverage — its strongest candidates sat at now-dead gauges (Bardheere) and ungauged points.The no-Google variant's backtest, same layout. Juba catches 5 of 8 severe years (gaining 2005 over the adopted mechanism), Shabelle 4 of 7 — losing the 2006 and 2014 Deyr floods. Note the Shabelle Gu row: it is now driven by one station-model, so its activations and its silences are a single model's judgement. With 25-year records, one hit moves coverage by ~12–15 points — differences under ~0.15 are noise.

The tuning surface

The Nigeria-style grid search (per state there; per basin × season here — the same recipe as P. Wairimu's notebook 08): every (per-pair RP threshold, N pairs required) cell is scored against the seasonal SWALIM benchmark. The adopted cell is not always the best single-window score — it is the best choice subject to the basin-level constraints (near-equal per-basin frequency across the Gu+Deyr union, severe-coverage-first, ≥ 3-yr overall enforced jointly across the basins). With 4–5-pair pools the surface is small enough to read whole: the neighbors show what loosening a leg would cost in false activations, or tightening one in missed severe years.

Twelve heatmaps: for each of the four windows, POD, FAR and activation return period over the per-pair RP threshold by N-required grid, with the adopted cell outlined
POD and FAR are vs the seasonal 1-in-3 SWALIM benchmark; black box = the adopted configuration. Dark = high POD / low FAR / frequent activation. With 7–8 benchmark events per window, one hit moves POD by ~12 points — broad plateaus, not sharp optima, are the honest reading.

Are SWALIM's official flood-risk levels trustworthy?

The benchmark years above are defined by empirical return periods of the gauges' own seasonal maxima, not by SWALIM's published Moderate/High flood-risk levels — and the check below is why. Putting the official levels on the empirical RP scale (per active gauge and season, 1999–2023) shows they imply wildly inconsistent frequencies: "Moderate" ranges from a level the river reaches most years (Dollow Deyr, ~1-in-1.3; Jowhar Gu, ~1-in-1.6) to a genuinely rare one (Luuq Gu, ~1-in-4.6), and "High" from ~1-in-1.8 (Dollow) to ~1-in-11.5 (Luuq Gu). They are engineering levels of uncertain provenance, not a consistent severity scale — the same conclusion P. Wairimu's notebook 01 flagged when the published max_level at Belet Weyne came out below its bank-full level. Everything in this mechanism therefore runs on each gauge's own empirical RP3/RP5 levels (table som_ms_swalim_rp); the official levels remain useful only as familiar reference lines for readers of SWALIM bulletins.

Five panels, one per active SWALIM gauge, of empirical return-period curves of seasonal maximum levels for Gu and Deyr, with the official moderate and high flood-risk levels as horizontal dashed lines
Empirical Weibull curves of seasonal-maximum level (blue Gu, orange Deyr) per active gauge, with the official Moderate (amber) and High (red) levels as dashed lines. Where a dashed line crosses the curves far from the RP 3–5 band, the official label and the observed frequency disagree. Dollow's curve rests on only 8–9 seasons (gauge online 2015) — read it loosely.

Forecast skill against each model's own reanalysis

The assume-bias-for-all-sources rule, quantified: per model and lead time, how well the forecast reproduces the model's own reanalysis or retrospective — rank correlation (does it track itself?) and the median forecast/reanalysis ratio (is it biased against its own climatology?). Computed per station on flood-season days, median across the 18 stations.

Two panels: Spearman correlation of forecast versus own reanalysis by lead time, and median forecast-to-reanalysis ratio by lead time, for the three models
GloFAS v4 (2003–2023) and Google GRRR (2016–2023) forecasts are essentially self-consistent at action leads: rank ρ ≥ 0.99, ratio ≈ 1.00 — the licence for calibrating their thresholds on reanalysis. GloFAS drifts to ~0.90× by lead 12 (readiness band, shaded) — its 7–12-day thresholds should anticipate that. GEOGloWS runs 0.85–0.91× its own retrospective even at lead 1 (ρ ~0.82), on only ~2 years of archive (2024–26) — retrospective-fitted thresholds will under-fire on its live forecasts until refit on forecast climatology. No GloFAS v5 reforecast exists, so the operational v5 forecast's own consistency is untestable — the single most important gap.

Would it have worked operationally?

The calibration record is reanalysis — a model's afterwards-view of its own past. The operational test replays the historical forecasts: per pair and valid day, the most alarming ≤ 6-day signal (max over issue dates — with mixed providers in one window, a monitoring day combines each provider's latest forecast). Google pairs use their own reforecast (2016–2023); GloFAS v5 pairs the v4 reforecast (2003–2023, ensemble median, thresholds refit on v4's own climatology, since no v5 reforecast exists); GEOGloWS pairs the retrospective as a lead-0 stand-in (hindsight — flagged per window). That forecast skill barely decays from lead 1 to 7 — initial-condition-driven on these slow rivers — is P. Wairimu's notebook 03 finding, and is what makes both lead bands viable at all.

windowarchivestand-in pairs reproduces calibration years?notes
Juba Gu2016–20231 2/3 — 2018, 2020; misses 2016 near-clean
Juba Deyr2003–20232 1/1 — catches 2023 (25 d after the Luuq Moderate crossing) over-fires: adds 2014, 2015 (~1-in-7 vs the calibration's 1-in-26) — see below
Shabelle Gu2016–20230 1/2 — catches 2016 (−3 d); misses 2020 fragile: the all-4-pairs rule leaves no slack on forecasts — see below
Shabelle Deyr2003–20232 1/4 — catches 2019 (+11 d); misses Deyr 2023, 2014, 2006; adds 2013 (+21 d) and 2020 (−1 d) the weak window; see the caveat below
The caveat that matters most. At ≤ 6-day leads the GloFAS v4 reforecast misses Deyr 2023 on the Shabelle — the largest flood in the record — even though the v5 reanalysis flags it clearly. The operational forecast has been v5 since August 2026 and should behave like the v5 reanalysis here, but there is no v5 reforecast to prove it. Re-verifying both Deyr legs the moment EWDS publishes a v5 reforecast is the single most important follow-up.
Two operational tuning items the calibration cannot settle. (1) Juba Deyr over-fires on the proxy: at its 1-in-5 pair thresholds the v4-reforecast evaluation activates 3 times in 21 years (~1-in-7) against the calibration's single year (1-in-26) — partly the two hindsight stand-in votes, partly v4-proxy threshold transfer. (2) Shabelle Gu's all-4-pairs consensus is fragile on forecasts: one pair short on the right valid day and it stays silent (it missed 2020 operationally). Both windows' N and RP deserve a working-group pass with operational data before go-live — the basin-level RP arithmetic leaves room to trade N against RP without moving the basin frequencies.
Bar chart of days between the first alert and the SWALIM Moderate crossing, per activation event
Per-event anticipation: days between the alert date (the N-th pair's earliest alerting forecast on the first consensus day) and the season's first SWALIM Moderate crossing at the reference gauge. Range −25 to +21 days. Stakeholders should not expect a consistent warning window — the readiness leg exists to buy mobilisation time.

The readiness leg (7–12 days)

GloFAS ensemble-median forecasts at leads 7–12 days (the leads-8-12 reforecast was downloaded specifically for this band), over each window's distinct stations — just 2 on the Juba, 3 on the Shabelle now — thresholds on GloFAS v4's own climatology, at a lower bar: readiness releases only the mobilisation share and may activate more often than action. It fully covers the action years for Juba Gu, Juba Deyr and Shabelle Gu at 1-in-3.7–4.4 frequencies — unchanged in the no-Google variant, since readiness was always GloFAS-only; only the required station counts relax (1-of-2 on the Juba, 1-of-1 on Shabelle Gu) to match the variant's smaller action pools. Shabelle Deyr readiness covers 2 of 4 action years — at 7–12 days GloFAS v4 cannot see Deyr-2023-type events at all. Individually the readiness legs sit at 1-in-3.7–7.3, but their union runs ~1-in-1.8 — the working group should see that number explicitly when deciding the mobilisation share.

Return-period bookkeeping

leveltriggeractivations (1999–2023)RP
individualJuba Gu action 4 — 2010, 2016, 2018, 2020 6.5 yr
individualJuba Deyr action1 — 202326.0 yr
individualShabelle Gu action 2 — 2016, 2020 13.0 yr
individualShabelle Deyr action 4 — 2006, 2014, 2019, 2023 6.5 yr
basinJuba (Gu or Deyr) 55.2 yr
basinShabelle (Gu or Deyr)64.3 yr
overallaction, either basin 83.2 yr

Under all-in funding (the full envelope on any activation, the working assumption in the evidence deck) the effective RP equals the overall RP, 3.2 years. A split structure would raise the effective RP above it; that is a framework-team decision.

Activations, impact and response, year by year

The two-historical-records rule: the trigger record against impact, not just the hazard benchmark. Everything is grouped basin → season, reverse-chronological. The trigger column is the trigger's actual decision variable: the peak number of (station, model) pairs simultaneously over their thresholds that season-year — the cell fills red when it meets the requirement shown in the season header (= the trigger activates). RP is the empirical return period of that season's maximum level at the SWALIM reference gauge — 5yr = a 1-in-5-year level or higher, 3yr = 1-in-3 to 1-in-5 (Weibull on the gauge's own seasonal maxima, per the threshold check above — not the official flood-risk levels); EM-DAT (purple) is people affected and CERF (blue) the flood allocation (US$), shaded by magnitude, each attributed to the basin(s) named in the event or allocation narrative.

loading…

Attribution & caveats. EM-DAT is a floor (entry-criteria bias), split to seasons by start month (Mar–Jun → Gu, Sep–Dec → Deyr); an event naming both rivers is counted under both basins (not split — hover for deaths and the multi-basin flag); events naming neither river (mostly flash floods elsewhere) are excluded. A CERF allocation naming neither river in its narrative is shown under both basins and marked * (they are national riverine-flood responses). Impact columns are context, not a skill score: a formal skill-vs-impact statistic still requires choosing an impact threshold.

Should impact tilt the probabilities?

Near-equal per-basin probability is a design choice, not a law — a basin or season carrying systematically more impact could get a lower relative threshold. Aggregating the basin-attributed records (1999–2023) answers whether the data call for it:

Three horizontal bar panels showing each basin-season window's share of EM-DAT people affected, EM-DAT deaths, and CERF flood funding, 1999 to 2023
Each window's share of total flood impact, per record (som_ms_impact_by_window). Basin totals are near-equal — Juba 47%, Shabelle 53% on the three-record average — but the seasons are not: Deyr carries ~65% of impact (and the two largest CERF flood allocations) against Gu's ~35%.

Re-running the calibration with the activation budget allocated by impact share instead of equally (table som_ms_impact_tilt) makes the conclusion concrete:

The honest reading: at basin level the impact data endorse the adopted near-equal split; at season level a deliberate Deyr tilt is a working-group lever, quantified and ready, but it trades verified gauge-flood coverage for impact-weighted frequency — a values decision, not a statistical one.

Open items before a trigger report

  1. GloFAS v5 reforecast — re-verify both Deyr legs and the readiness band when EWDS publishes one.
  2. Operational tuning of Juba Deyr and Shabelle Gu — the proxy over-firing (and the 1-in-26 calibration leg it sits under) and the all-4-pairs fragility above; trade N against RP within the basin budgets using operational data.
  3. Provider-dependency decision — the no-Google variant (switch at the top of this page) shows the mechanism survives Google's loss at 1-in-3.7 overall, but Shabelle Gu collapses to a single station-model and Shabelle severe coverage falls to 4/7. If the working group wants a provider-independence guarantee, that is a design constraint to set now, not after a provider change.
  4. Gauge network watch — the candidate set is the five SNRFA gauges still reporting. If SWALIM revives Bardheere (data stop Nov 2023) or Bualle (Mar 2024), the Juba pool widens and its severe-year coverage likely recovers; conversely a further gauge loss shrinks a pool that has no slack. Season tilt (Deyr-ward) is quantified in som_ms_impact_tilt if the working group wants it.
  5. GEOGloWS operational debias — before its votes count live, either SFDC-map its forecasts onto the retrospective climatology (the method from the team's Nepal technical note, fitted on the 2024+ overlap) or refit its thresholds on forecast climatology as the archive grows; validate against whichever Deyr/Gu seasons the archive then covers.
  6. Version pinning for operations — action thresholds here are v5-climatology (GloFAS pairs), the readiness/backtest thresholds v4-climatology; refit on one consistent operational product before go-live.
  7. Formal impact-record cross-check — the year-by-year table above is descriptive; a formal statistic requires the working group to fix an impact threshold ("what counts as a year that should have triggered").
  8. Final threshold adjustment to the working group's exact return-period targets, and the funding-split decision.

Provenance. Produced by notebook 11 in ds-aa-som-floods (August 2026), building directly on the evidence notebooks 01–10 by Pauline Wairimu: SWALIM thresholds and flood events (01), model-vs-gauge trust (02), lead-time skill (03), GloFAS v4/v5 bias (04), flood benchmarks (05), the correlation screen (06), station selection (07), trigger grids (08–09) and the staged-mechanism exploration (10) — summarised in her "Somalia Floods" slide deck (team Drive) and the evidence deck in the repo. Ground truth: FAO SWALIM SNRFA river levels. Providers: GloFAS v4/v5 (Copernicus EWDS), Google Flood Hub (GRRR), GEOGloWS v2 (ECMWF). Impact records: EM-DAT (CRED), CERF OneGMS. Calibration record 1999–2023; all return periods are Weibull plotting positions. Configuration tables live in blob under ds-aa-som-floods/processed/workflow/som_ms_*.