← Back to Home

Prediction Scorecard

How accurate Creek Intelligence's forecasts have been lately. Experimental — self-grading, updated daily.
Updated 2026-09-10 06:30 UTC · scoring window: last 45 days

Every forecast the system makes is recorded and later graded against what the river actually did — this page is that report card. It currently scores 82 physics crest predictions, 1044 empirical likelihood forecasts, and 155 recession countdowns (0 still maturing). Updated daily; treat it as experimental.

Where the truth comes from: USGS keeps revising its provisional gauge readings for months. The gauge archives these grades are checked against were last re-pulled from USGS on 2026-09-10 05:20 UTC (trailing 120 days, with approval status), and every settled grade whose readings are not yet USGS-approved is re-graded nightly; a grade that moves keeps its first verdict on the record. Last pass (2026-09-10 06:05 UTC): rise: 499 checked, 0 verdicts moved last night, 172 re-graded so far, 0 final; recession: 1872 checked, 162 verdicts moved last night, 778 re-graded so far, 0 final; physics: 321 checked, 0 verdicts moved last night, 107 re-graded so far, 192 final; ponca: 736 checked, 0 verdicts moved last night, 678 re-graded so far, 0 final; empirical: 1154 checked, 0 verdicts moved last night, 0 re-graded so far, 0 final. USGS is retiring the data service these gauge feeds use; its replacement has been running alongside it since 2026-09-04: 135 hourly checks, 27247 readings compared between the two services with 0 value differences, and 23145 readings compared against what the live feeds stored with 10 differences. 30 readings were seen by the new service before the old one and 0 the other way round (timing, not disagreement). Right now 1 gauge series returns no values at all (07055875/00060) — see the data-health table below.

⚠️ Needs attention

Physics Rise Predictors

Predicts the crest a gauge will reach from rainfall + antecedent moisture. Since 2026-09-08 every hourly run is a claim — a rise, or no rise — and both are graded: detection (rises the engine called) and false alarms side by side, then how close the crest and its timing were when a called rise happened.

Sample: 82 records · 1 events · 2 on the new grading · last 45 days · truth as of 2026-09-10 05:20 UTC · events counted on v2 records only (fixed 18-h horizon, from 2026-09-04)

New grading from 2026-09-04: each prediction is checked over a fixed 18-hour window after it was issued (before, the window closed 2 h after the predicted peak, so a later crest was never seen). It is graded on three separate questions: did the gauge rise at all, how far off was the crest when it did, and how far off was the timing. Predictions whose crest was not above the current level are no longer archived as claims.

PredictorRise predictions
records / events
Gauge rose
records / events
Crest error given a riseTiming given a rise
median |err| / signed
Within ±20% (rose)
Richland Creek2 / 10.0% / 0.0%— / —n/a for stage

Claims (from 2026-09-08): every hourly run is archived as a claim — a rise, or no rise (quiet / rain too light) — and graded 18 h later against the gauge. Detection = rises the engine called, out of every rise that happened (a no-rise claim followed by a rise is a miss). False alarms = rise predictions the gauge did not follow. Events = episodes of consecutive hourly runs.

PredictorRise claimsNo-rise claims
graded / issued
Detection (POD)
records / events
False alarms (FAR)
records / events
Missed risesNo-rise verified
Richland Creek20 / 44— / —100.0% / 100.0%0
Legacy grading (window closed 2 h after the predicted peak; records issued before 2026-09-04)
PredictorScoredAvg error (MAE)Bias (mean / median)Within ±20%
Cossatot River330.41 ft0.41 / 0.28 ft67.0% (n=33)
Richland Creek240.79 ft0.79 / 0.46 ft17.0% (n=24)
Hailstone (upper Buffalo)2544.67 cfs44.67 / 18.7 cfs44.0% (n=25)

Bias is the mean / median signed error (predicted − actual). A large gap between them means a few outlier events — often a single flash-flood onset — are dragging the mean; the median is the typical miss.

Worst recent misses

Cossatot River

Richland Creek

Hailstone (upper Buffalo)

Empirical Forecast Engine

Each hourly forecast is a calibrated probability that the gauge rises to the next level within the promised window (12 h on the fast creeks, 30 h on the slow ones). Scored as probabilities: Brier = mean squared error (0 is perfect); skill (BSS) = improvement over always forecasting the long-run rate (>0 is skill); obs/forecast = how many rises happened per rise forecast (1.0 is calibrated); reliability = what actually happened when the engine said X%. Probabilistic scoring began 2026-09-02; earlier records are not comparable and are not shown.

Sample: 1044 records · 2 events · 23 claims / 1021 no-rise calls · last 45 days · truth as of 2026-09-10 05:20 UTC · events = claim episodes (issuances >= 8 % within 3 h of each other); non-claims do not form episodes

"No rise indicated" calls (forecast below 8%) are graded separately on their miss rate: 1021 such hourly calls in the window, 0.0% of which were followed by a rise to the target level. Real claims (8% and above): 23, of which 0 verified.

BasinForecasts
hourly / episodes
RisesMean forecastObservedBrier
lower is better
Skill (BSS)
>0 beats climatology
Obs / forecast
1.0 = calibrated
AUC
Big Piney Creek146 / 000.9%0.0%0.00019— (no rises in window)
Buffalo at Boxley175 / 000.0%0.0%0.0— (no rises in window)
Cossatot River175 / 000.2%0.0%1e-05— (no rises in window)
Hailstone (upper Buffalo)175 / 000.0%0.0%0.0— (no rises in window)
Illinois River (Hwy 16)42 / 2010.8%0.0%0.01384— (no rises in window)
Mulberry River156 / 000.1%0.0%1e-05— (no rises in window)
Richland Creek175 / 000.0%0.0%0.0— (no rises in window)

All basins pooled: 1044 forecasts, 0 rises, Brier 0.00059, skill None, obs/forecast None, AUC None.

Reliability (when the engine said X%, how often did the rise come?)

Forecast binnMean forecastObserved
0-2%9790.1%0.0%
2-7%424.3%0.0%
7-15%2115.0%0.0%
15-25%215.9%0.0%
Per storm episode (highest probability issued)
Max forecast binEpisodesMean forecastRose
15-25%215.9%0.0%
By target level
LevelnRisesBrierObs / forecastAUC
LOW100203e-05NoneNone
MEDIUM104400.00057NoneNone
HIGH104400.0NoneNone

Recession Countdowns

Predicts how long until a falling river drops to each threshold. HIT% = reached the level near the predicted time; MAE = average timing error (hours). Newly launched — most predictions are still maturing.

Sample: 161 records · 5 events · last 45 days · truth as of 2026-09-10 05:20 UTC · events = recession episodes per gauge/target

GaugeTargetGraded
records
HIT%
records
Episodes
reached
Never reachedTiming MAE
records / last call
Bias
last call
big_piney/below_longpoollow_floatable8100.0%1 (100.0%)0.0%13.3h / 9.1h9.1h
big_piney/below_longpooltoo_low33100.0%1 (100.0%)0.0%22.5h / 8.5h8.5h
buffalo/harrietlow_floatable65100.0%1 (100.0%)0.0%10.5h / 5.4h5.4h
buffalo/st_joelow_floatable49100.0%2 (100.0%)0.0%12.8h / 4.6h2.2h

An episode is one run of hourly countdowns for a gauge and target (a new episode starts after a 3-hour gap); it counts as reached if any countdown in it was. "Last call" is the timing error of the final countdown issued before the crossing.

Ponca AI Rainfall Event Analysis

When rain arms it, predicts how big the Ponca gauge will get — a class (Fizzle/Moderate/High/Flood), a typical-peak band, and a flood-risk %. Scored per armed EVENT since 2026-09-02: the first call and the most-informed ('mature', most rain so far) call of each event are graded against the crest the event actually reached; the flood-risk % is scored as a probability (Brier, 0 is perfect; 'shown' is the number on the card, 'raw' the analog match alone). After the crest the card hides the % unless a second, higher crest is coming, so post-peak calls are graded only on whether a higher crest came. Events, not 15-minute re-issues, are the sample size.

Sample: 304 records · 5 events · last 45 days · truth as of 2026-09-10 05:20 UTC

Last 45 days — 5 armed event(s), 304 fifteen-minute calls, 0 flood(s).

EventsClass right (first call)Class right (mature call)Within one classCrest in likely bandFlood-risk Brier: shown / raw
5100%100%100%0%0.0065 / 0.0065 (n=5)

All time since 2026-06-08 — 12 event(s), 2 flood(s).

EventsClass right (first call)Class right (mature call)Within one classCrest in likely bandFlood-risk Brier: shown / raw
1283%92%100%33%0.0217 / 0.0217 (n=10)

Reliability of the shown flood-risk % on rising calls (304 fifteen-minute rows — rows within one event are near-duplicates, so this shows calibration shape, not sample size): 0-10%: said 1% → 0% flooded (n=232), 10-25%: said 14% → 0% flooded (n=72).

Recent events — first call, mature call, actual

Event (UTC)CallsFirst callMature callActual crestActual class
2026-08-26T22:0256Fizzle · 0% flood @ 0.45″, 3 cfsFizzle · 0% flood @ 0.9″, 10 cfs10 cfsFizzle
2026-08-24T07:0260Fizzle · 0% flood @ 0.3″, 0 cfsFizzle · 0% flood @ 0.6″, 3 cfs3 cfsFizzle
2026-08-08T01:0280Fizzle · 11% flood @ 0.63″, 42 cfsFizzle · 0% flood @ 1.66″, 78 cfs33 cfsFizzle
2026-08-07T01:0256Fizzle · 0% flood @ 1.19″, 9 cfsFizzle · 17% flood @ 2.34″, 45 cfs33 cfsFizzle
2026-07-29T22:0252Fizzle · 0% flood @ 0.42″, 29 cfsFizzle · 6% flood @ 0.93″, 29 cfs14 cfsFizzle
2026-07-15T02:02120Fizzle · 0% flood @ 0.28″, 51 cfsFizzle · 0% flood @ 1.85″, 53 cfs28 cfsFizzle
2026-07-11T20:0282Fizzle · 11% flood @ 0.62″, 54 cfsFizzle · 17% flood @ 1.38″, 61 cfs30 cfsFizzle
2026-06-27T13:0252Moderate · 22% flood @ 0.78″, 213 cfsModerate · 39% flood @ 0.84″, 705 cfs580 cfsModerate

Radar Nowcast (Ponca card)

While rain is falling, the Ponca card reads the last 40 minutes of weather radar, moves the picture forward, and says roughly how much more rain lands on the creeks above Ponca in the next two hours (or that the rain is about done). Each radar frame's estimate is graded against the rain the next two hours actually delivered, using the card's own hourly rain accounting. Bias = estimate minus actual (negative means radar under-called); 'more coming' = the card said 0.1 in or more was on the way; 'about done' = it said little was left. The 'past storms like this went on to a real rise' perspective line is graded against whether the event rose at all.

Last 45 days — 303 radar frames graded across 5 rain event(s) (0 not gradable).

Amount claimFramesError (MAE)BiasActual ÷ estimate (median)
All frames3030.026″0.003″0.39
When it said ≥0.1″ was coming280.155″0.111″0.39
CallTimes madeVerifiedCaught
“More rain lined up” (≥0.1″ coming)2857%57%
“About done” (little left)26497%

Verified = the call was right when made; caught = share of the real ≥0.1″ two-hour periods the card had called in advance.

Estimate saidFramesMean estimateMean actual≥0.1″ actually fell
0″2180.0″0.004″1%
0-0.1″570.032″0.056″18%
0.1-0.25″180.156″0.12″50%
0.25-0.5″80.344″0.171″62%
>=0.5″20.72″0.176″100%

By storm motion (estimates ≥0.1″): <8 mph (near-stationary): n=4, bias 0.248″, actual÷est 0.39; 8-20 mph: n=11, bias 0.072″, actual÷est 0.69; >=20 mph: n=13, bias 0.103″, actual÷est 0.33. Slow-moving storms grow over the watershed and deliver more than radar advection suggests; fast movers pass through and deliver less.

“Past storms like this went on to a real rise” line (12 lines in 4 events): “more often than not” (63%): said 10 times, the event rose 0% of them; “about half the time” (52%): said 2 times, the event rose 0% of them.

All time since 2026-07-11 — 494 radar frames graded across 7 rain event(s) (0 not gradable).

Amount claimFramesError (MAE)BiasActual ÷ estimate (median)
All frames4940.026″0.002″0.63
When it said ≥0.1″ was coming450.16″0.105″0.54
CallTimes madeVerifiedCaught
“More rain lined up” (≥0.1″ coming)4562%58%
“About done” (little left)42498%

Verified = the call was right when made; caught = share of the real ≥0.1″ two-hour periods the card had called in advance.

Estimate saidFramesMean estimateMean actual≥0.1″ actually fell
0″3460.0″0.005″1%
0-0.1″1030.033″0.056″17%
0.1-0.25″250.156″0.13″52%
0.25-0.5″120.338″0.203″58%
>=0.5″80.688″0.379″100%

By storm motion (estimates ≥0.1″): <8 mph (near-stationary): n=9, bias 0.15″, actual÷est 0.56; 8-20 mph: n=21, bias 0.09″, actual÷est 0.69; >=20 mph: n=15, bias 0.1″, actual÷est 0.33. Slow-moving storms grow over the watershed and deliver more than radar advection suggests; fast movers pass through and deliver less.

“Past storms like this went on to a real rise” line (24 lines in 6 events): “more often than not” (63%): said 20 times, the event rose 0% of them; “about half the time” (53%): said 4 times, the event rose 0% of them.

Buffalo Rise Engine (per-gauge nowcast)

For each Buffalo mainstem gauge, predicts a coming rise (slight / moderate / large) from local rain + upstream propagation, with a timing window. Graded on whether the gauge actually rose, and within the predicted window. Predictions issued in the last 45 days, with the all-time totals beside them; predictions only fire during rain events, so this fills in over time.

Sample: 243 records · 16 events · last 45 days · truth as of 2026-09-10 05:20 UTC

Predictions issued in the last 45 days: 243 graded records, 16 events (10 rose). All-time: 499 graded records, 40 events (21 rose, 52%).

GaugeGraded recordsRise happened (records)On-timeNo-rise recordsEventsRise happened (events)Band held (rose)Chance skill (Brier)
boxley2370%0%73 (2 rose)67%
ponca4924%100%374 (2 rose)50%
pruitt23100%0%02 (2 rose)100%
st_joe6930%0%483 (2 rose)67%
harriet7924%0%604 (2 rose)50%

Since 2026-09-03 the local-rain layer serves a rise probability and a typical rise band (25th–75th percentile, cfs) instead of a size word. "Band held" = the share of real rises that landed inside the band (target about 50%); Brier = mean squared error of the probability (0 is perfect, 0.25 is a coin flip).

A "rise" here means peak >= 1.25 x v0 and peak - v0 >= clamp(0.10 x v0, 10, 30) cfs — the same rule the engine uses to decide a predicted rise has arrived. Records graded before 2026-09-03 used a fixed +30 cfs floor, which read some real low-water rises as misses; the 45-day window mixes both vintages until mid-October. An event = a run of records less than 3 h apart.

Typical actual rise by predicted category: slight: ~27.2 cfs (n=53), moderate: ~28.2 cfs (n=36), large: ~12.2 cfs (n=2). (Categories come from rainfall, so this is how they map to real gauge rises — calibration that accrues over time.)

Recent rise predictions — predicted vs. actual

When (UTC)GaugePredictedWindowOutcomeActual rise
2026-08-27T12:24poncaslight0.0-6.0hno_rise
2026-08-27T11:24poncaslight0.0-7.0hno_rise
2026-08-27T10:24poncaslight0.0-8.0hno_rise
2026-08-27T09:24poncaslight0.0-9.0hno_rise
2026-08-27T08:24poncaslight0.0-10.0hno_rise
2026-08-27T07:24poncaslight0.0-11.0hno_rise

Downstream Propagation (Ponca → Pruitt → St. Joe)

Retired 2026-09-02. The Ponca card's own downstream model duplicated the wave tracker's reaches with weaker validation, so the card now narrates the wave tracker's bands and arrival windows for Pruitt, St. Joe and Harriet, and those are graded in the Routed Waves section below. The four events this section had graded included two synthetic host tests, which have been removed.

Retired — graded under Routed Waves below.

Routed Waves (wave tracker)

When an upstream gauge crests, the wave tracker predicts the crest size and arrival window at each downstream gauge. Every confirmed wave is graded after its window closes: did a real rise arrive, was the crest inside the predicted band, and did it arrive inside the window. Waves are rare events — this fills in slowly.

Sample: 1 records · 1 events · last 90 days · truth as of 2026-09-10 05:20 UTC · one grade per confirmed wave-target

Waves graded: 1 (significant: 1); crest inside the likely band 1/1; inside the wide band 1/1; arrived inside the window 0/1; median predicted/actual crest 1.0×. Alerts issued: 0; floods with no alert: 0.

WaveUpstream crestPredicted bandOutcomeActual peakLag (h)
st_joe → harriet445.0 cfs430–496 cfs (med 467)significant467.0 cfs17.5

Watershed Alerts (Q-bucket triggers)

The rain-triggered WATCH/WARNING words on the calibrated creeks (Upper Buffalo, Richland Main, Upper Cossatot, Upper Big Piney, Mulberry), graded against the creek's own gauge: an alert verifies when the creek reached its floatable floor within 24 h; a rise with rain and no alert is a miss. Same recipe as the leave-one-year-out numbers on the Watersheds page.

Sample: 22980 records · 2 events · last 90 days · truth as of 2026-09-10 05:20 UTC · records = 15-min Tier-1 decisions; events = alert episodes (+ rises with no alert)

Alert episodes: 2 (verified 1, false alarms 1, still open 0); rises with no alert: 2; median lead from first WARNING 3.2 h.

DrainageStartedWordStart bucketRain / triggerOutcomeCreek peakLead (h)
Upper Cossatot2026-08-08 01:17ZWATCHBONE-DRY1.26" / 1.30"false alarm20 / 200 cfs
Upper Cossatot2026-07-12 20:17ZWARNINGDRY1.52" / 1.20"verified253 / 200 cfs3.2

Rises with no alert: Upper Cossatot 2026-07-15 18:17Z (rain 0.77" of 1.20"); hailstone 2026-06-27 19:47Z (rain 0.86" of 1.15")

Neural Net Predictions (Ponca & Boxley)

A neural network projects the level hour by hour from 48 h of 1 km radar rainfall (Ponca 6 h ahead, Boxley 4 h). Every hourly issuance is graded against what the gauge actually did: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.

0 hourly issuance(s) logged and maturing. Each settles ~8 hours after it is made; skill numbers become meaningful once the river actually moves.

Neural Net Predictions (Cossatot)

The Cossatot pixel net projects the level hour by hour from 48 h of 1 km radar rainfall (6 h ahead); crossing probabilities come from its companion gradient-boosted ensemble. Graded the same way as the Buffalo net: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.

0 hourly issuance(s) logged and maturing. Each settles ~8 hours after it is made; skill numbers become meaningful once the river actually moves.

Data health

Every grade above is checked against a USGS gauge feed. A single parameter can go silent while its site keeps reporting (USGS stops publishing discharge when the stage falls below its rating's lowest measured point); this table shows the age of the newest USGS reading per gauge and parameter, and whether a stage-derived stand-in is being served in its place.

21 series checked at 2026-09-10 06:30 UTC — 1 silent, 0 lagging.

GaugeParameterLast USGS readingAgeSeries stateStand-in
Richland 07055875discharge2026-09-04 16:00129.5 hempty0 cfs (stage below the rating floor)

Coverage

What the scorecard grades today, and what is still being wired into the loop:

Predictor familyStatusNotes
Physics rise predictorsgradedCossatot, Richland, Hailstone — predicted crest height/flow vs. the actual peak; rise calls graded by confidence label (HIGH / MEDIUM / LOW, eager-plus-labels 2026-09-09) against the rate each label quoted.
Empirical forecast enginegraded6 basins — 'likelihood of rise to tier X' vs. the tier the gauge actually reached.
Recession countdownsgraded9 gauges — maturing; the longest horizons settle ~7-10 days after they're issued.
Ponca AI Rainfall Event AnalysisgradedScored per armed EVENT (first call + mature call vs the event crest), flood-risk % as a probability; post-peak rows only on 'did a higher crest come' (ponca_analog_eval.py v2, 2026-09-02).
Radar nowcast (Ponca card)gradedThe card's 'roughly another X in falls in the next 2 hours' radar estimate and its 'about done' / 'more coming' calls, graded per radar frame against the areal rain the next two MRMS hours actually delivered (radar_eval.py, 2026-09-03).
Downstream propagation forecastretiredRetired 2026-09-02: the Ponca card now narrates the wave router's own bands, which the Routed Waves section grades (wave_eval.py). propagation_eval.py kept read-only.
Buffalo per-gauge rise predictionsgradedPer-gauge rise nowcasts recorded + graded vs the gauge's actual rise (buffalo_predictions_archive.py).
Routed waves (wave_router)gradedF4 tracked waves — each confirmed wave-target's frozen crest band + arrival window graded vs what the target gauge actually did (wave_eval.py).
Buffalo flood_risk labelsnot yet loggedLower-priority companion in buffalo_output; still ungraded.
Q-bucket / watershed alerts (/watersheds, Signal, Facebook mirror)gradedTier-1 alert episodes from creeks/logs/qbucket_shadow.log graded against the creek's own reading reaching its floatable floor within 24 h; rises with rain and no alert counted as misses (qbucket_eval.py, G4 2026-09-08). Tier-2 (reference-gauge) rows have no gauge of their own and stay ungraded.
Neural Net Predictions (pixel-v5)gradedPonca 6 h / Boxley 4 h hour-by-hour level tables — every hourly issuance graded against the observed level at each horizon (neural_eval.py).
Neural Net Predictions (Cossatot)gradedCossatot 6 h level tables + GBM-ensemble crossing probs — every hourly issuance graded at each horizon (neural_cossatot_eval.py).
Gauges Watersheds Cossatot Intelligence Richland Intelligence Mulberry Intelligence Big Piney Intelligence Illinois River Intel Hailstone Intelligence Buffalo Intelligence Buffalo Study
Guide Changelog Scorecard Page Suggestions and Corrections
Home
♥ Support this project
DISCLAIMER: This site provides creek condition estimates for informational purposes only. Gauge data, radar estimates, and forecasts may be delayed, inaccurate, or unavailable. Always exercise independent judgment. Whitewater kayaking is inherently dangerous — water conditions can change rapidly. This site and its maintainers assume no responsibility for decisions made based on information displayed here.