Every forecast the system makes is recorded and later graded against what the river actually did — this page is that report card. It currently scores 82 physics crest predictions, 1044 empirical likelihood forecasts, and 155 recession countdowns (0 still maturing). Updated daily; treat it as experimental.
Where the truth comes from: USGS keeps revising its provisional gauge readings for months. The gauge archives these grades are checked against were last re-pulled from USGS on 2026-09-10 05:20 UTC (trailing 120 days, with approval status), and every settled grade whose readings are not yet USGS-approved is re-graded nightly; a grade that moves keeps its first verdict on the record. Last pass (2026-09-10 06:05 UTC): rise: 499 checked, 0 verdicts moved last night, 172 re-graded so far, 0 final; recession: 1872 checked, 162 verdicts moved last night, 778 re-graded so far, 0 final; physics: 321 checked, 0 verdicts moved last night, 107 re-graded so far, 192 final; ponca: 736 checked, 0 verdicts moved last night, 678 re-graded so far, 0 final; empirical: 1154 checked, 0 verdicts moved last night, 0 re-graded so far, 0 final. USGS is retiring the data service these gauge feeds use; its replacement has been running alongside it since 2026-09-04: 135 hourly checks, 27247 readings compared between the two services with 0 value differences, and 23145 readings compared against what the live feeds stored with 10 differences. 30 readings were seen by the new service before the old one and 0 the other way round (timing, not disagreement). Right now 1 gauge series returns no values at all (07055875/00060) — see the data-health table below.
Predicts the crest a gauge will reach from rainfall + antecedent moisture. Since 2026-09-08 every hourly run is a claim — a rise, or no rise — and both are graded: detection (rises the engine called) and false alarms side by side, then how close the crest and its timing were when a called rise happened.
Sample: 82 records · 1 events · 2 on the new grading · last 45 days · truth as of 2026-09-10 05:20 UTC · events counted on v2 records only (fixed 18-h horizon, from 2026-09-04)
New grading from 2026-09-04: each prediction is checked over a fixed 18-hour window after it was issued (before, the window closed 2 h after the predicted peak, so a later crest was never seen). It is graded on three separate questions: did the gauge rise at all, how far off was the crest when it did, and how far off was the timing. Predictions whose crest was not above the current level are no longer archived as claims.
| Predictor | Rise predictions records / events | Gauge rose records / events | Crest error given a rise | Timing given a rise median |err| / signed | Within ±20% (rose) |
|---|---|---|---|---|---|
| Richland Creek | 2 / 1 | 0.0% / 0.0% | — | — / — | n/a for stage |
Claims (from 2026-09-08): every hourly run is archived as a claim — a rise, or no rise (quiet / rain too light) — and graded 18 h later against the gauge. Detection = rises the engine called, out of every rise that happened (a no-rise claim followed by a rise is a miss). False alarms = rise predictions the gauge did not follow. Events = episodes of consecutive hourly runs.
| Predictor | Rise claims | No-rise claims graded / issued | Detection (POD) records / events | False alarms (FAR) records / events | Missed rises | No-rise verified |
|---|---|---|---|---|---|---|
| Richland Creek | 2 | 0 / 44 | — / — | 100.0% / 100.0% | 0 | — |
| Predictor | Scored | Avg error (MAE) | Bias (mean / median) | Within ±20% |
|---|---|---|---|---|
| Cossatot River | 33 | 0.41 ft | 0.41 / 0.28 ft | 67.0% (n=33) |
| Richland Creek | 24 | 0.79 ft | 0.79 / 0.46 ft | 17.0% (n=24) |
| Hailstone (upper Buffalo) | 25 | 44.67 cfs | 44.67 / 18.7 cfs | 44.0% (n=25) |
Bias is the mean / median signed error (predicted − actual). A large gap between them means a few outlier events — often a single flash-flood onset — are dragging the mean; the median is the typical miss.
Cossatot River
Richland Creek
Hailstone (upper Buffalo)
Each hourly forecast is a calibrated probability that the gauge rises to the next level within the promised window (12 h on the fast creeks, 30 h on the slow ones). Scored as probabilities: Brier = mean squared error (0 is perfect); skill (BSS) = improvement over always forecasting the long-run rate (>0 is skill); obs/forecast = how many rises happened per rise forecast (1.0 is calibrated); reliability = what actually happened when the engine said X%. Probabilistic scoring began 2026-09-02; earlier records are not comparable and are not shown.
Sample: 1044 records · 2 events · 23 claims / 1021 no-rise calls · last 45 days · truth as of 2026-09-10 05:20 UTC · events = claim episodes (issuances >= 8 % within 3 h of each other); non-claims do not form episodes
"No rise indicated" calls (forecast below 8%) are graded separately on their miss rate: 1021 such hourly calls in the window, 0.0% of which were followed by a rise to the target level. Real claims (8% and above): 23, of which 0 verified.
| Basin | Forecasts hourly / episodes | Rises | Mean forecast | Observed | Brier lower is better | Skill (BSS) >0 beats climatology | Obs / forecast 1.0 = calibrated | AUC |
|---|---|---|---|---|---|---|---|---|
| Big Piney Creek | 146 / 0 | 0 | 0.9% | 0.0% | 0.00019 | — (no rises in window) | — | — |
| Buffalo at Boxley | 175 / 0 | 0 | 0.0% | 0.0% | 0.0 | — (no rises in window) | — | — |
| Cossatot River | 175 / 0 | 0 | 0.2% | 0.0% | 1e-05 | — (no rises in window) | — | — |
| Hailstone (upper Buffalo) | 175 / 0 | 0 | 0.0% | 0.0% | 0.0 | — (no rises in window) | — | — |
| Illinois River (Hwy 16) | 42 / 2 | 0 | 10.8% | 0.0% | 0.01384 | — (no rises in window) | — | — |
| Mulberry River | 156 / 0 | 0 | 0.1% | 0.0% | 1e-05 | — (no rises in window) | — | — |
| Richland Creek | 175 / 0 | 0 | 0.0% | 0.0% | 0.0 | — (no rises in window) | — | — |
All basins pooled: 1044 forecasts, 0 rises, Brier 0.00059, skill None, obs/forecast None, AUC None.
| Forecast bin | n | Mean forecast | Observed |
|---|---|---|---|
| 0-2% | 979 | 0.1% | 0.0% |
| 2-7% | 42 | 4.3% | 0.0% |
| 7-15% | 21 | 15.0% | 0.0% |
| 15-25% | 2 | 15.9% | 0.0% |
| Max forecast bin | Episodes | Mean forecast | Rose |
|---|---|---|---|
| 15-25% | 2 | 15.9% | 0.0% |
| Level | n | Rises | Brier | Obs / forecast | AUC |
|---|---|---|---|---|---|
| LOW | 1002 | 0 | 3e-05 | None | None |
| MEDIUM | 1044 | 0 | 0.00057 | None | None |
| HIGH | 1044 | 0 | 0.0 | None | None |
Predicts how long until a falling river drops to each threshold. HIT% = reached the level near the predicted time; MAE = average timing error (hours). Newly launched — most predictions are still maturing.
Sample: 161 records · 5 events · last 45 days · truth as of 2026-09-10 05:20 UTC · events = recession episodes per gauge/target
| Gauge | Target | Graded records | HIT% records | Episodes reached | Never reached | Timing MAE records / last call | Bias last call |
|---|---|---|---|---|---|---|---|
| big_piney/below_longpool | low_floatable | 8 | 100.0% | 1 (100.0%) | 0.0% | 13.3h / 9.1h | 9.1h |
| big_piney/below_longpool | too_low | 33 | 100.0% | 1 (100.0%) | 0.0% | 22.5h / 8.5h | 8.5h |
| buffalo/harriet | low_floatable | 65 | 100.0% | 1 (100.0%) | 0.0% | 10.5h / 5.4h | 5.4h |
| buffalo/st_joe | low_floatable | 49 | 100.0% | 2 (100.0%) | 0.0% | 12.8h / 4.6h | 2.2h |
An episode is one run of hourly countdowns for a gauge and target (a new episode starts after a 3-hour gap); it counts as reached if any countdown in it was. "Last call" is the timing error of the final countdown issued before the crossing.
When rain arms it, predicts how big the Ponca gauge will get — a class (Fizzle/Moderate/High/Flood), a typical-peak band, and a flood-risk %. Scored per armed EVENT since 2026-09-02: the first call and the most-informed ('mature', most rain so far) call of each event are graded against the crest the event actually reached; the flood-risk % is scored as a probability (Brier, 0 is perfect; 'shown' is the number on the card, 'raw' the analog match alone). After the crest the card hides the % unless a second, higher crest is coming, so post-peak calls are graded only on whether a higher crest came. Events, not 15-minute re-issues, are the sample size.
Sample: 304 records · 5 events · last 45 days · truth as of 2026-09-10 05:20 UTC
Last 45 days — 5 armed event(s), 304 fifteen-minute calls, 0 flood(s).
| Events | Class right (first call) | Class right (mature call) | Within one class | Crest in likely band | Flood-risk Brier: shown / raw |
|---|---|---|---|---|---|
| 5 | 100% | 100% | 100% | 0% | 0.0065 / 0.0065 (n=5) |
All time since 2026-06-08 — 12 event(s), 2 flood(s).
| Events | Class right (first call) | Class right (mature call) | Within one class | Crest in likely band | Flood-risk Brier: shown / raw |
|---|---|---|---|---|---|
| 12 | 83% | 92% | 100% | 33% | 0.0217 / 0.0217 (n=10) |
Reliability of the shown flood-risk % on rising calls (304 fifteen-minute rows — rows within one event are near-duplicates, so this shows calibration shape, not sample size): 0-10%: said 1% → 0% flooded (n=232), 10-25%: said 14% → 0% flooded (n=72).
| Event (UTC) | Calls | First call | Mature call | Actual crest | Actual class |
|---|---|---|---|---|---|
| 2026-08-26T22:02 | 56 | Fizzle · 0% flood @ 0.45″, 3 cfs | Fizzle · 0% flood @ 0.9″, 10 cfs | 10 cfs | Fizzle |
| 2026-08-24T07:02 | 60 | Fizzle · 0% flood @ 0.3″, 0 cfs | Fizzle · 0% flood @ 0.6″, 3 cfs | 3 cfs | Fizzle |
| 2026-08-08T01:02 | 80 | Fizzle · 11% flood @ 0.63″, 42 cfs | Fizzle · 0% flood @ 1.66″, 78 cfs | 33 cfs | Fizzle |
| 2026-08-07T01:02 | 56 | Fizzle · 0% flood @ 1.19″, 9 cfs | Fizzle · 17% flood @ 2.34″, 45 cfs | 33 cfs | Fizzle |
| 2026-07-29T22:02 | 52 | Fizzle · 0% flood @ 0.42″, 29 cfs | Fizzle · 6% flood @ 0.93″, 29 cfs | 14 cfs | Fizzle |
| 2026-07-15T02:02 | 120 | Fizzle · 0% flood @ 0.28″, 51 cfs | Fizzle · 0% flood @ 1.85″, 53 cfs | 28 cfs | Fizzle |
| 2026-07-11T20:02 | 82 | Fizzle · 11% flood @ 0.62″, 54 cfs | Fizzle · 17% flood @ 1.38″, 61 cfs | 30 cfs | Fizzle |
| 2026-06-27T13:02 | 52 | Moderate · 22% flood @ 0.78″, 213 cfs | Moderate · 39% flood @ 0.84″, 705 cfs | 580 cfs | Moderate |
While rain is falling, the Ponca card reads the last 40 minutes of weather radar, moves the picture forward, and says roughly how much more rain lands on the creeks above Ponca in the next two hours (or that the rain is about done). Each radar frame's estimate is graded against the rain the next two hours actually delivered, using the card's own hourly rain accounting. Bias = estimate minus actual (negative means radar under-called); 'more coming' = the card said 0.1 in or more was on the way; 'about done' = it said little was left. The 'past storms like this went on to a real rise' perspective line is graded against whether the event rose at all.
Last 45 days — 303 radar frames graded across 5 rain event(s) (0 not gradable).
| Amount claim | Frames | Error (MAE) | Bias | Actual ÷ estimate (median) |
|---|---|---|---|---|
| All frames | 303 | 0.026″ | 0.003″ | 0.39 |
| When it said ≥0.1″ was coming | 28 | 0.155″ | 0.111″ | 0.39 |
| Call | Times made | Verified | Caught |
|---|---|---|---|
| “More rain lined up” (≥0.1″ coming) | 28 | 57% | 57% |
| “About done” (little left) | 264 | 97% | — |
Verified = the call was right when made; caught = share of the real ≥0.1″ two-hour periods the card had called in advance.
| Estimate said | Frames | Mean estimate | Mean actual | ≥0.1″ actually fell |
|---|---|---|---|---|
| 0″ | 218 | 0.0″ | 0.004″ | 1% |
| 0-0.1″ | 57 | 0.032″ | 0.056″ | 18% |
| 0.1-0.25″ | 18 | 0.156″ | 0.12″ | 50% |
| 0.25-0.5″ | 8 | 0.344″ | 0.171″ | 62% |
| >=0.5″ | 2 | 0.72″ | 0.176″ | 100% |
By storm motion (estimates ≥0.1″): <8 mph (near-stationary): n=4, bias 0.248″, actual÷est 0.39; 8-20 mph: n=11, bias 0.072″, actual÷est 0.69; >=20 mph: n=13, bias 0.103″, actual÷est 0.33. Slow-moving storms grow over the watershed and deliver more than radar advection suggests; fast movers pass through and deliver less.
“Past storms like this went on to a real rise” line (12 lines in 4 events): “more often than not” (63%): said 10 times, the event rose 0% of them; “about half the time” (52%): said 2 times, the event rose 0% of them.
All time since 2026-07-11 — 494 radar frames graded across 7 rain event(s) (0 not gradable).
| Amount claim | Frames | Error (MAE) | Bias | Actual ÷ estimate (median) |
|---|---|---|---|---|
| All frames | 494 | 0.026″ | 0.002″ | 0.63 |
| When it said ≥0.1″ was coming | 45 | 0.16″ | 0.105″ | 0.54 |
| Call | Times made | Verified | Caught |
|---|---|---|---|
| “More rain lined up” (≥0.1″ coming) | 45 | 62% | 58% |
| “About done” (little left) | 424 | 98% | — |
Verified = the call was right when made; caught = share of the real ≥0.1″ two-hour periods the card had called in advance.
| Estimate said | Frames | Mean estimate | Mean actual | ≥0.1″ actually fell |
|---|---|---|---|---|
| 0″ | 346 | 0.0″ | 0.005″ | 1% |
| 0-0.1″ | 103 | 0.033″ | 0.056″ | 17% |
| 0.1-0.25″ | 25 | 0.156″ | 0.13″ | 52% |
| 0.25-0.5″ | 12 | 0.338″ | 0.203″ | 58% |
| >=0.5″ | 8 | 0.688″ | 0.379″ | 100% |
By storm motion (estimates ≥0.1″): <8 mph (near-stationary): n=9, bias 0.15″, actual÷est 0.56; 8-20 mph: n=21, bias 0.09″, actual÷est 0.69; >=20 mph: n=15, bias 0.1″, actual÷est 0.33. Slow-moving storms grow over the watershed and deliver more than radar advection suggests; fast movers pass through and deliver less.
“Past storms like this went on to a real rise” line (24 lines in 6 events): “more often than not” (63%): said 20 times, the event rose 0% of them; “about half the time” (53%): said 4 times, the event rose 0% of them.
For each Buffalo mainstem gauge, predicts a coming rise (slight / moderate / large) from local rain + upstream propagation, with a timing window. Graded on whether the gauge actually rose, and within the predicted window. Predictions issued in the last 45 days, with the all-time totals beside them; predictions only fire during rain events, so this fills in over time.
Sample: 243 records · 16 events · last 45 days · truth as of 2026-09-10 05:20 UTC
Predictions issued in the last 45 days: 243 graded records, 16 events (10 rose). All-time: 499 graded records, 40 events (21 rose, 52%).
| Gauge | Graded records | Rise happened (records) | On-time | No-rise records | Events | Rise happened (events) | Band held (rose) | Chance skill (Brier) |
|---|---|---|---|---|---|---|---|---|
| boxley | 23 | 70% | 0% | 7 | 3 (2 rose) | 67% | — | — |
| ponca | 49 | 24% | 100% | 37 | 4 (2 rose) | 50% | — | — |
| pruitt | 23 | 100% | 0% | 0 | 2 (2 rose) | 100% | — | — |
| st_joe | 69 | 30% | 0% | 48 | 3 (2 rose) | 67% | — | — |
| harriet | 79 | 24% | 0% | 60 | 4 (2 rose) | 50% | — | — |
Since 2026-09-03 the local-rain layer serves a rise probability and a typical rise band (25th–75th percentile, cfs) instead of a size word. "Band held" = the share of real rises that landed inside the band (target about 50%); Brier = mean squared error of the probability (0 is perfect, 0.25 is a coin flip).
A "rise" here means peak >= 1.25 x v0 and peak - v0 >= clamp(0.10 x v0, 10, 30) cfs — the same rule the engine uses to decide a predicted rise has arrived. Records graded before 2026-09-03 used a fixed +30 cfs floor, which read some real low-water rises as misses; the 45-day window mixes both vintages until mid-October. An event = a run of records less than 3 h apart.
Typical actual rise by predicted category: slight: ~27.2 cfs (n=53), moderate: ~28.2 cfs (n=36), large: ~12.2 cfs (n=2). (Categories come from rainfall, so this is how they map to real gauge rises — calibration that accrues over time.)
| When (UTC) | Gauge | Predicted | Window | Outcome | Actual rise |
|---|---|---|---|---|---|
| 2026-08-27T12:24 | ponca | slight | 0.0-6.0h | no_rise | — |
| 2026-08-27T11:24 | ponca | slight | 0.0-7.0h | no_rise | — |
| 2026-08-27T10:24 | ponca | slight | 0.0-8.0h | no_rise | — |
| 2026-08-27T09:24 | ponca | slight | 0.0-9.0h | no_rise | — |
| 2026-08-27T08:24 | ponca | slight | 0.0-10.0h | no_rise | — |
| 2026-08-27T07:24 | ponca | slight | 0.0-11.0h | no_rise | — |
Retired 2026-09-02. The Ponca card's own downstream model duplicated the wave tracker's reaches with weaker validation, so the card now narrates the wave tracker's bands and arrival windows for Pruitt, St. Joe and Harriet, and those are graded in the Routed Waves section below. The four events this section had graded included two synthetic host tests, which have been removed.
Retired — graded under Routed Waves below.
When an upstream gauge crests, the wave tracker predicts the crest size and arrival window at each downstream gauge. Every confirmed wave is graded after its window closes: did a real rise arrive, was the crest inside the predicted band, and did it arrive inside the window. Waves are rare events — this fills in slowly.
Sample: 1 records · 1 events · last 90 days · truth as of 2026-09-10 05:20 UTC · one grade per confirmed wave-target
Waves graded: 1 (significant: 1); crest inside the likely band 1/1; inside the wide band 1/1; arrived inside the window 0/1; median predicted/actual crest 1.0×. Alerts issued: 0; floods with no alert: 0.
| Wave | Upstream crest | Predicted band | Outcome | Actual peak | Lag (h) |
|---|---|---|---|---|---|
| st_joe → harriet | 445.0 cfs | 430–496 cfs (med 467) | significant | 467.0 cfs | 17.5 |
The rain-triggered WATCH/WARNING words on the calibrated creeks (Upper Buffalo, Richland Main, Upper Cossatot, Upper Big Piney, Mulberry), graded against the creek's own gauge: an alert verifies when the creek reached its floatable floor within 24 h; a rise with rain and no alert is a miss. Same recipe as the leave-one-year-out numbers on the Watersheds page.
Sample: 22980 records · 2 events · last 90 days · truth as of 2026-09-10 05:20 UTC · records = 15-min Tier-1 decisions; events = alert episodes (+ rises with no alert)
Alert episodes: 2 (verified 1, false alarms 1, still open 0); rises with no alert: 2; median lead from first WARNING 3.2 h.
| Drainage | Started | Word | Start bucket | Rain / trigger | Outcome | Creek peak | Lead (h) |
|---|---|---|---|---|---|---|---|
| Upper Cossatot | 2026-08-08 01:17Z | WATCH | BONE-DRY | 1.26" / 1.30" | false alarm | 20 / 200 cfs | — |
| Upper Cossatot | 2026-07-12 20:17Z | WARNING | DRY | 1.52" / 1.20" | verified | 253 / 200 cfs | 3.2 |
Rises with no alert: Upper Cossatot 2026-07-15 18:17Z (rain 0.77" of 1.20"); hailstone 2026-06-27 19:47Z (rain 0.86" of 1.15")
A neural network projects the level hour by hour from 48 h of 1 km radar rainfall (Ponca 6 h ahead, Boxley 4 h). Every hourly issuance is graded against what the gauge actually did: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.
0 hourly issuance(s) logged and maturing. Each settles ~8 hours after it is made; skill numbers become meaningful once the river actually moves.
The Cossatot pixel net projects the level hour by hour from 48 h of 1 km radar rainfall (6 h ahead); crossing probabilities come from its companion gradient-boosted ensemble. Graded the same way as the Buffalo net: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.
0 hourly issuance(s) logged and maturing. Each settles ~8 hours after it is made; skill numbers become meaningful once the river actually moves.
Every grade above is checked against a USGS gauge feed. A single parameter can go silent while its site keeps reporting (USGS stops publishing discharge when the stage falls below its rating's lowest measured point); this table shows the age of the newest USGS reading per gauge and parameter, and whether a stage-derived stand-in is being served in its place.
21 series checked at 2026-09-10 06:30 UTC — 1 silent, 0 lagging.
| Gauge | Parameter | Last USGS reading | Age | Series state | Stand-in |
|---|---|---|---|---|---|
| Richland 07055875 | discharge | 2026-09-04 16:00 | 129.5 h | empty | 0 cfs (stage below the rating floor) |
What the scorecard grades today, and what is still being wired into the loop:
| Predictor family | Status | Notes |
|---|---|---|
| Physics rise predictors | graded | Cossatot, Richland, Hailstone — predicted crest height/flow vs. the actual peak; rise calls graded by confidence label (HIGH / MEDIUM / LOW, eager-plus-labels 2026-09-09) against the rate each label quoted. |
| Empirical forecast engine | graded | 6 basins — 'likelihood of rise to tier X' vs. the tier the gauge actually reached. |
| Recession countdowns | graded | 9 gauges — maturing; the longest horizons settle ~7-10 days after they're issued. |
| Ponca AI Rainfall Event Analysis | graded | Scored per armed EVENT (first call + mature call vs the event crest), flood-risk % as a probability; post-peak rows only on 'did a higher crest come' (ponca_analog_eval.py v2, 2026-09-02). |
| Radar nowcast (Ponca card) | graded | The card's 'roughly another X in falls in the next 2 hours' radar estimate and its 'about done' / 'more coming' calls, graded per radar frame against the areal rain the next two MRMS hours actually delivered (radar_eval.py, 2026-09-03). |
| Downstream propagation forecast | retired | Retired 2026-09-02: the Ponca card now narrates the wave router's own bands, which the Routed Waves section grades (wave_eval.py). propagation_eval.py kept read-only. |
| Buffalo per-gauge rise predictions | graded | Per-gauge rise nowcasts recorded + graded vs the gauge's actual rise (buffalo_predictions_archive.py). |
| Routed waves (wave_router) | graded | F4 tracked waves — each confirmed wave-target's frozen crest band + arrival window graded vs what the target gauge actually did (wave_eval.py). |
| Buffalo flood_risk labels | not yet logged | Lower-priority companion in buffalo_output; still ungraded. |
| Q-bucket / watershed alerts (/watersheds, Signal, Facebook mirror) | graded | Tier-1 alert episodes from creeks/logs/qbucket_shadow.log graded against the creek's own reading reaching its floatable floor within 24 h; rises with rain and no alert counted as misses (qbucket_eval.py, G4 2026-09-08). Tier-2 (reference-gauge) rows have no gauge of their own and stay ungraded. |
| Neural Net Predictions (pixel-v5) | graded | Ponca 6 h / Boxley 4 h hour-by-hour level tables — every hourly issuance graded against the observed level at each horizon (neural_eval.py). |
| Neural Net Predictions (Cossatot) | graded | Cossatot 6 h level tables + GBM-ensemble crossing probs — every hourly issuance graded at each horizon (neural_cossatot_eval.py). |