← Back to Home
Arkansas Creek Monitor — Changelog
Version 2026.9.9.2 — September 9, 2026
Richland, Hailstone and Cossatot rise predictions are now issued eagerly, and every rise call carries a confidence label that says how often calls like it have come true.
- Richland, Hailstone and Cossatot Physics Predictor cards: the engine now runs three copies of itself every hour — a cautious one, a middle one and an eager one. The eager copy makes the call (so you will see more rise predictions than before, and earlier), and each rise call is labelled by how many of the copies agree: HIGH (all three), MEDIUM (middle and eager), or LOW (eager only). Three small chips under the label show which copies called a rise; hover one for the crest it predicted.
- Each label quotes its own track record, from a replay of every hour since 2014–2015 in which the engine was never allowed to see the year it was being tested on: roughly 7 in 10 HIGH calls have verified, 4 in 10 MEDIUM, 3 in 10 LOW. From June through September the card quotes the summer figure instead, which is lower — LOW summer calls have verified only about 2 in 10 times on Richland and Hailstone. We show it anyway rather than hide the call; treat a summer LOW as "rain fell upstream, worth watching" and nothing more. Experimental: these are the replay numbers, and the scorecard will show whether the live calls keep to them.
- Scorecard: a new "Rise calls by confidence" table under the physics Claims table — how many HIGH / MEDIUM / LOW rise calls were issued and graded, how many verified, and the live rate beside the rate the card quoted.
Not user-facing, recorded for completeness: the served physics fit on all three basins is the eager (τ 0.75) G6 refit, the τ 0.6 and τ 0.5 fits ride along in engine_v2.dials; physics_serve.rise_confidence(); per-dial peaks and the quoted rate on every archive record (quiet runs too); score_all._physics_claims.by_confidence; the nightly scorer had been crashing on the physics claims' event view since 2026-09-09 morning (the scorecard page was a day stale) — fixed and re-pushed; evidence in research/physics_confidence/REPORT.md.
Version 2026.9.9.1 — September 9, 2026
A radar hole is now "unknown", not "dry", from the radar feed to the page; a dead rain bucket shows N/A instead of 0.00; grading no longer counts readings USGS has withdrawn.
- Watersheds page: when the radar mosaic could not see part of a basin for an hour, that hour is now recorded as unknown rather than as no rain. A Q-bucket drainage's detail says "radar gap N h" and the storm's rain event stays open through the gap instead of closing on false dry hours; a page-level "Radar gap" line appears whenever a rolling window is missing hours or radar coverage. The data-age line gains "MRMS hour" (the actual radar hour behind the rain columns) and the "Radar (QPE)" age now tracks the radar feed itself, so a dead feed no longer looks fresh.
- Buffalo page: a small amber line reports when the newest rainfall hour has radar holes over the basin.
- Gauge cards: a USGS tipping-bucket rain sensor that reads 0.00 for hours while radar shows heavy rain over it is now reported as "N/A" (sensor dead) rather than "0.00 in". Experimental — the rule needs 6 h of zeros against ≥ 0.5 in of radar.
- Scorecard: the "Where the truth comes from" paragraph now also says how many readings only the new USGS service had seen (timing, not disagreement) and names any gauge parameter USGS served nothing for in the nightly re-pull.
Not user-facing, recorded for completeness: one shared missing-radar rule (basinlib/radar_gap.py) in all shared-QPE readers (Illinois's reader became a basinlib shim), day-files keyed on the MRMS hour with a 6-hour backfill of missed cycles, the shared fetcher replaces an hour built from a shifted fallback frame once the top-of-hour frame appears; one shared truth reader (basinlib/truth.py) for every settler (withdrawn and stage-derived readings are never graded on); nightly truth refresh marks withdrawals only inside the span USGS served and reports parameters served nothing; recession re-grading marks a countdown whose revised start was already at its target as invalidated (the first nightly pass will flip ~168 "hit at 0.25 h" grades); physics/empirical re-grading skips windows that end after the truth pull; study day-file gauge names from the study config; the dashboard's stale-data thresholds now come from the shared data_age_thresholds.yaml.
Version 2026.9.8.7 — September 8, 2026
Neural Net Predictions: the level table and the crossing alert can no longer contradict each other, the Cossatot page shows real feet, and a USGS outage no longer blanks the page.
- Buffalo (Ponca/Boxley) and Cossatot Neural Net Predictions: whenever the crossing model is at or above its watch point, the hour-by-hour table now shows the model's full projected move instead of the usual dry-spell nudge toward "no change". Replaying the June 22 flood through the previous version showed the table topping out near 1,400 cfs beside an active flood alert while the river went past 8,000 cfs; the new rule serves the model's 8,000+ at that hour. Validated on twelve held-out years: the table now agrees with the alert as often as the model itself can, without losing to "no change" on ordinary hours. Experimental.
- Below 10 cfs with no rain in the last two days the table simply holds the current level, and its likely range is never narrower than the gauge's own reporting step. On held-out years the model was worse than "no change" there at every horizon; the old rows (for example 1.69–1.74 cfs around 1.73) measured rounding, not the river.
- Cossatot feet are now converted with USGS's current rating for the gauge (refreshed monthly), so the page no longer shows six rows of 1.92 ft under a "Currently 1.83 ft" line. On that rating the 240-cfs floatable line is 3.46 ft and the 1,270-cfs high-water line 5.38 ft; the tier labels keep the site's round 3.4 / 5.5 ft for now.
- When USGS does not answer (it happened 77 times since July), the page keeps the previous hour's projection with an orange "Gauge feed down" banner instead of a blank "could not run this hour" page, and the outage is recorded for the scorecard.
- Every hour's small print now lists what the model noticed about its inputs (a gauge that had not posted yet, a missing radar hour, re-ordered quantiles), whether or not anything is degraded. Boxley's crossing line now says its chance was never validated for that gauge (no percentage will show) instead of "signal below the validated watch level"; the Ponca 800-cfs line is labelled as rising through the middle of the optimal range; Boxley's table turns red at the dashboard's 1,500-cfs Flood level. A "Model raw" column appears when the nudge changes a number by more than half.
- Scorecard: the neural sections now grade the September serving tables separately from the July model (earlier models are collapsed below), skip hours issued below 10 cfs as "not scorable", and count hours with no gauge data.
Not user-facing, recorded for completeness: rain regime now five bands (no cliff at 10 mm), Cossatot shrink factors smoothed across horizons, Boxley posting lag logged hourly ahead of a cron move, missing radar hours re-read for 48 hours, clock-skew guard, notes persisted in the ledgers; research/neural_g7/REPORT.md.
Version 2026.9.8.6 — September 8, 2026
Empirical Forecast card: the rain columns say what they are, the Illinois headline stops flip-flopping, and the archive pages grade every hour in plain words.
- On every Empirical Forecast card (Cossatot, Richland, Hailstone, Mulberry, Big Piney, Illinois), the Details table now reads "Level / Chance / Rain before past rises (p25 / p50 / p75) / n". The coloured "above typical" pill and the "Soil-stratified data unavailable" note are gone: those rain columns were never a probability, only the rain that fell before past rises to each level. They are shown only where at least 10 past rises are on record; the Chance column is the actual forecast for every level.
- The card shows a warning line when the radar feed is stale or when part of the basin had no radar coverage in the window it is summing, instead of silently reporting less rain.
- Illinois River (Hwy 16): when the gauge is hovering just under a level (as it has been all week at 216–237 cfs against the 250-cfs MEDIUM line), the card now stays on "Gauge is already within reach of MEDIUM" until the river clearly drops away, instead of flipping between that, "~15 % chance" and "No rise indicated" every hour on a few cfs of gauge noise. Experimental.
- Empirical archive pages (
/<creek>/empirical_predictions/<date>) print what the engine actually issued each hour — the label and the exact chance — and grade it in words: a claim that verified or did not, "no rise indicated — correct negative", a missed rise, or "not a forecast" for hours when the gauge was already within reach. The last 45 days were re-rendered.
- The "~X %" number on the card and the claim line the scorecard grades now use the same rule (an exact 8 %), so a forecast printed at 8 % or more is always a graded claim and one printed below never is.
Not user-facing, recorded for completeness: empirical_predict.py 2026.9.8.6 — rain_context replaces the v1 percentile_band/p25/p50/p75/n fields, the legacy WATCH headline path is deleted, MRMS missing-radar sentinels are counted per basin-hour, per-window gap status, flags.stale_qpe/qpe_age_h, hysteresis input from the previous issue; calibrated.py CLAIM_P, display_pct, per-tier capped, table-driven hysteresis (Illinois only, gated on the v2 rig: OOS BSS 0.303 → 0.374); empirical_archive.py study-archive day-file branch for Boxley/Hailstone, coverage test at both window ends, new .md renderer; resettle.py first-settles never-settled records; score_all.py claim-only episodes; hailstone_analyze.py reads the per-basin v2 archive and takes --date. ARCHITECTURE §4.8 + §16 item 51; research/empirical_g8/.
Version 2026.9.8.5 — September 8, 2026
Ponca AI Rainfall Event Analysis card: its history now matches the way it was trained, and a stale write-up says so.
- The card's "Now: X cfs (…)" label uses the same tier words as the Ponca gauge card beside it (Too Low / Low but Floatable / Optimal / Flood) instead of its own Low / Moderate / High.
- If the written note cannot be refreshed (the AI writer is busy or down), the note is shown greyed with the time it was written and a line saying the flow and rain figures below are current. Before, an hours-old note could sit under a fresh "% chance" with no hint.
- The class of rise the note builds toward ("a flood" / "a strong, high rise") now agrees with the "% chance > 1600 cfs" shown under it; the two could disagree before.
- Experimental: the comparison against past storms now uses the river's true 72-hour pre-storm baseline and a true 6-hour trend (it had been using a 30-hour baseline and a 90-minute trend), and while an inbound wave keeps the card up after the rain has stopped, the storm's rain total stays in the comparison instead of being reset to zero. The historical library itself is unchanged.
Not user-facing, recorded for completeness: ponca_analog.py 2026.9.8.5 — per-series history (80 h of flow readings), storm_reference() frozen storm frame with cum_source, trend_by_time(), dry hours by MRMS-hour time, tied router quantiles handled in p_exceed, class_word(), qwen ValueError caught + qwen_ok/narrative_at/narrative_age_min/narrative_stale in the outlook, shared 7-min qwen budget; storm-LOO and season-replay gates in research/ponca_analog_g9/. DMZPi render_ponca_outlook_card() stale-note rendering.
Version 2026.9.8.4 — September 8, 2026
The three physics rise predictors (Cossatot, Richland, Hailstone) were re-fit on an honest basis, no longer under-call a rise while the gauge feed lags, show feet that match USGS's own rating, and are now graded on every hourly run — including the ones that said "no rise".
- Why the engines were re-fit. The September 2 fit of the physics engine had a flaw in how it was scored: the fitting rig could see rain that fell after each prediction hour, which the live engine never can. Its advertised detection rate ("78% of rises") was measured that way. Scored honestly the shipped parameters caught about 40% of rises, and over the last 45 days of drought showers the engines issued no rise prediction at all. All three were re-fit on rain up to the prediction hour only. The confidence line on each card now quotes the honest number ("rain to the issue hour only"): roughly half of rises detected, with about one call in three not verifying. Experimental: these are drought-summer numbers; the first real wet-season storms will tell more.
- No more under-calling while the gauge lags. USGS readings reach the site 35 to 100 minutes late. The engine used to subtract the response it assumed the gauge had already seen, so a lagging gauge cost up to a third of the predicted rise, and a storm the engine thought had already arrived could print "too light to drive a rise" right before a crest. The engine now reckons from the gauge reading's own time.
- Feet that match USGS. Cossatot and Richland cards convert the engine's flow to feet. The old conversion stretched a low-water quirk up the whole rating (a 500-cfs Richland call read 4.08 ft "Optimal" where USGS says 3.50 ft "Low but Floatable"). The site now pulls USGS's own shift-adjusted rating weekly and applies the low-water shift the way USGS does; displayed stages match USGS to within 0.02 ft at 3–6 ft.
- Card wording. Band lines now add up to the headline rise (they used to sum to about 1.4× it); "arriving" means within the next hour, otherwise a clock time is shown — and that clock time (and the peak time on the Cossatot and Richland cards) is now Central time, not UTC labelled CT; the ground line says what it is based on ("by pre-storm baseflow 3.2 cfs; about 1.4″ of rain soaks in before runoff starts"); the quiet card warns when the prediction is stale; radar gaps in the last 12 hours are noted on the card.
- Every run is graded. The scorecard's physics section gains a Claims table: each hourly run is archived as a rise or no-rise claim and checked 18 hours later, so detection (rises the engine called), false alarms and missed rises appear side by side, and the per-day predictions pages list the no-rise claims with their verdicts. The Hailstone page's rainfall rows are relabelled: the top row is the whole Boxley catchment (154 km²) and Smith Creek is shown as part of it rather than counted twice.
Not user-facing, recorded for completeness: basinlib/physics_engine.py (18-h timeline, gauge lag, kernel-end hold, crest recession, NaN radar slots, model floor built but gated off — it raised false alarms to 45%), physics_serve.py, rating.py + new fetch_usgs_rating.py (weekly exsa cron), predictions_archive.py / hailstone_predictions_archive.py grade v3 (claims, level at issue from the card, cfs at settle), score_all.py claims block + engine filter, resettle.py back-fill, Hailstone reader on ZoneInfo dates with the radar-gap flag and MRMS-hour dedup, refit engine_v2 blocks (v2.1-g6, quantile 0.6 / 0.5 / 0.6). Fresh-eyes review session G6; evidence research/physics_serve_path/REPORT.md; ARCHITECTURE §4.3–4.5, §16 item 49.
Version 2026.9.8.3 — September 8, 2026
The Buffalo wave tracker stops guessing outside its experience: at drought base flow it now uses the plain lag table instead of a multi-day arrival window, and its "still rising" windows use the gauge readings it was actually fitted on.
- Arrival windows at low water. The stage-aware lag model was fitted on rivers running at 12–95 cfs or more. With the lower Buffalo sitting at 2–36 cfs this week it was extrapolating: the first wave of the season would have carried a Pruitt-to-St. Joe arrival window of 36–54 hours (the longest lag ever measured on that reach is 30). The tracker now checks the model's fitted range first and falls back to the typical-lag table when the river is below it; the "Incoming Wave" card and the Upstream line say so ("timing from the typical-lag table — the river is below the lag model's fitted range").
- "Still rising" windows. While an upstream gauge is still climbing, the window depends on how fast it has moved in the last reading. On the four gauges whose readings arrive in hourly batches (Boxley, St. Joe, Harriet, Bear Creek) the tracker was reading that rate as zero on nearly every update, and it only saw one reading per hour instead of four. It now uses every reading and the rate the model was fitted on. Experimental: on the historical record this changes the early window little overall (it helps most at Boxley); the honest number is now in the scorecard.
- Local-rain card wording. Once the expected window has passed while rain continues, the card no longer counts down to "0–0.5 hr"; it says the rise is overdue and that the gauge is being watched (and the grader no longer scores that as a timed miss). At Pruitt, St. Joe and Harriet the card now also says when the claim will be withdrawn if the river has not started to move — the tracker drops a no-show early on purpose, and the card used to promise a window it would not keep.
- Scorecard. The Routed Waves section now grades the early "still rising" forecasts as their own family (did the rise come, and how much notice did the first forecast give), shows which arrival-window source each graded wave used, and the Buffalo rise table gains a fixed 24-hour rise rate beside the window-based one.
- Also: a wave whose upstream gauge stops reporting mid-rise now expires after six hours instead of being re-issued forever; a downstream crest that arrives before the earliest measured lag is no longer credited to the wave; a small incoming wave no longer silences the local-rain card; the wave line's confidence word now reflects the width of the predicted range.
Not user-facing, recorded for completeness: reach_models.json lag_model.regression.support per reach; wave_router.lag_quantiles/remaining_rise/_prior_reading/_readings_seq/_value_at, DEAD_FEED_MARGIN_MIN, ARRIVAL_EARLY_SLACK_H, window_early; buffalo_assemble readings feed, wave_inbound_significant, wave_band_confidence, timing_state, no_move_cut_h, last_rain_obs/consumed_at; wave_eval families + reopen_count keys; buffalo_predictions_archive overdue skip + rose_within_24h; mid kept after the v8d2/v8f re-run (one extra late rise per lower gauge per decade for +0.2–1.8 h fizzle hold); simulator lagQuantiles port + bundle, golden 2,200/0. Fresh-eyes session G3; evidence research/buffalo_g3/REPORT.md; ARCHITECTURE §4.6 / §16 item 48.
Version 2026.9.8.1 — September 8, 2026
Watershed alerts now follow the creek gauge, say plainly how often they verify, and reach Facebook exactly as they reach Signal.
- The gauge comes first on the calibrated creeks. Until now a rain WARNING on Upper Buffalo, Richland Main, Upper Cossatot or Upper Big Piney stayed on the page and the home cards for the drainage's lag plus two hours no matter what the creek did — on July 12 the Cossatot page said WARNING for twelve hours, of which the creek was actually up for under two. Now, once a calibrated creek is running (or has already come up and is receding), the rain word clears and the creek's own reading is the status. FLOOD, which is the gauge's word, follows the reading with a one-hour confirmation after it drops back below the flood level. Rain WATCH/WARNING holds on the other drainages are unchanged.
- Honest wording. "High confidence in boatable conditions" is gone: a WARNING now reads "Rain past the trigger — creek expected to come up", a WATCH "Rain nearing the trigger — conditions developing", and each calibrated creek's card, Signal message and Watersheds row carries its own record — for example "about 6 in 10 WARNINGs verified, typical lead 3 h; WATCH 5 in 10" — from an eleven-to-twelve-year leave-one-year-out test. The Watersheds page has a new table of those numbers, and the WATCH level (75% of the trigger) is scored for the first time: it catches more rises but roughly half of WATCHes on the drier creeks do not verify. Neighbouring drainages that borrow a gauge say "Heuristic trigger, not calibrated".
- Signal and Facebook now say the same thing. The Facebook mirror had drifted from the Signal rules since the September 2 refit (a Tier-1 FLOOD would have posted "0.0 inches in 12hr, 0% of trigger"). ScriptPi now records every alert transition in a ledger and Facebook replays it word for word — same wording, same timing (hourly), same rules.
- A second storm pages again, and floods get an all-clear. After a WARNING, a second storm arriving inside the twelve-hour cooldown was silent on Signal and Facebook even though the page showed it; it now pages as a new alert once the first storm's word has cleared. When a creek drops back below its flood level after a FLOOD, one "Flood over" update goes out (Signal and Facebook).
- Scorecard: watershed alerts are graded.
/scorecard gains a "Watershed Alerts (Q-bucket triggers)" section: every rain alert on a calibrated creek is checked against the creek's own gauge (did it reach its floatable floor within 24 hours, and how far ahead did the WARNING come), and rises that had rain but no alert are listed as misses. Experimental — two months of dry-season data so far (one verified WARNING, one false WATCH, two missed rises).
- Small fixes: a held alert's home-card text is no longer cut off mid-sentence; the Watersheds page no longer colours the 1–24 hour rain cells of a calibrated creek against its storm-total threshold;
/admin shows the real trigger arithmetic for gauge-driven rows.
Not user-facing, recorded for completeness: resolve_status() / compute_alert_transitions() / notify_events() + alert_state.json event_ledger → predictor_output.json alert_events (alert_rules 3.1, FLOOD_DEBOUNCE_H 1.0); qbucket_shadow.py creek_ran/event_peak_q_cfs, skill_text/skill, America/Chicago season + day-file names; qbucket_config.json 3.1.2026-09-08 (thresholds byte-identical, loyo.watch added by recal_v3.py); Mulberry's Tier-1 record logged every cycle; antecedent_precip.json keys before 09-04 re-keyed +1 day (WP5 side effect); prediction_eval/qbucket_eval.py + score_all qbucket block (coverage → graded); creek-social ledger replay with a _meta.last_event_id cursor; Tier-2 reachability table (two months: no Boxley-referenced drainage has come within 60% of its dry-season threshold) recorded in ARCHITECTURE §4.2 — multipliers unchanged by decision. Fresh-eyes review session G4; evidence research/qbucket_v31/.
Version 2026.9.7.4 — September 7, 2026
Richland Creek's flow feed has been silent since September 4 and the site now says so — and keeps predicting honestly instead of quietly switching to an older, less reliable method.
- What happened. USGS stops publishing a creek's flow (cfs) when the water drops below the lowest point its rating table covers. Richland Creek has sat at about 0.22 ft since September 4 with no flow figure at all, while the stage reading kept coming in. Nothing on the site noticed: the Richland physics predictor fell back to its older engine (and showed higher confidence while doing it), the Richland watershed rows were about to fall back to a plain rain-only trigger overnight, and every health check reported green.
- Richland flow now reads 0 cfs, labelled. When USGS is not publishing flow and the stage is below the rating floor, the site takes the flow as 0 cfs and says so: the Richland page shows "0.0 cfs (est. from stage; USGS not publishing flow)". The physics predictor runs its normal engine on it (ground state "VERY DRY"), the Richland watershed trigger keeps its dry-ground threshold, and the Buffalo page's Richland tributary row shows 0 instead of a grey "Unknown".
- Every physics card names its engine. The Cossatot, Richland and Hailstone physics cards now carry an "Engine:" line — the normal offline-fit engine, or "v1 (fallback — discharge feed silent since …)" when the predictor had to use the older method. A fallback prediction is always marked low confidence now; it used to inherit a "HIGH — Well-calibrated" label it had not earned.
- Watershed rows say when they are on the rain-only fallback. On the Watersheds page, a drainage that normally uses its gauge's pre-storm flow to set the rain threshold but cannot (flow feed silent or stale) shows a "⚠ fallback" tag beside its trigger and the reason in its detail text. Experimental: the fallback rows still alert on rain, just with the older threshold.
- Scorecard: a Data health table and fewer blind spots.
/scorecard gains a table of every gauge feed the grades depend on, per parameter, with the age of the newest USGS reading and whether a stand-in value is being used. A scorer that fails now says so instead of showing "0 records"; the empirical skill score is blank ("no rises in window") during dry spells instead of showing a perfect 1.0; the physics attention flag no longer fires on rows graded the old way; and the weekly Signal digest's Ponca line is back after a silent month.
Not user-facing, recorded for completeness: basinlib/qsynth.py (stage-derived discharge in a separate discharge_synth day-file block, real USGS readings untouched), series_state on every fetched series, per-parameter data_age_alert, usgs_shadow empty-series counts, engine/q_source on archive records, Q-bucket 6-h Q age rule and clamped-pre-event-Q refusal, legacy_fallback trigger method, ZoneInfo day keys in the basinlib/Illinois/study gauge fetchers (the nightly 00:00–01:00 CDT "yesterday's reading" hour on four basin pages should be gone from tonight), data_health block in scorecard.json. Fresh-eyes review session G1; details ARCHITECTURE §4.2/§4.3/§4.13/§4.14, §16 item 45.
Version 2026.9.7.3 — September 7, 2026
Buffalo rain totals corrected: since early July the Buffalo page's "last hour" (and 3/6/12/24-hour) radar rainfall had been counting one extra hour about half the time.
- A timing quirk in the hourly rainfall reader meant that, on roughly 45% of updates, the previous hour's rain was added into the current hour's total, and every longer window was one hour too long. Rain totals on the Buffalo page, the rainfall map's 1-hour view, the rain-driven rise predictions, the Ponca AI card's storm total and the radar-nowcast scorecard all read a little high (the Ponca card's storm totals about 1.8× on the summer's showers). Fixed on the reader: every window now counts exactly the hours it names.
- Radar nowcast re-graded. The Ponca card's "roughly another X inches in the next 2 hours" radar estimate had been graded against the inflated rain, which made the radar look like it under-called. Against the corrected record the radar estimate is a ceiling, not a floor (about half of what it projects actually falls, in every storm-motion regime), and "about done" verifies 98% of the time. The scorecard's radar block now shows the corrected grades. The card's wording is unchanged for now, pending review.
- New notice on the Buffalo page. If the radar rainfall feed falls more than 90 minutes behind, a yellow notice now says so and explains that rain totals, the rain map and rain-driven predictions are frozen at the last good hour while gauge readings continue. Nothing else on the page changes when the feed is healthy.
- Everything else was checked and left alone: replaying the summer with the corrected rain against what was actually served moved no alert threshold (the local-rain, rise-size, wave and Ponca products all read slightly less exposure on clean rain; no rises were missed or gained).
Not user-facing, recorded for completeness: buffalo_read_qpe.py windows keyed on the MRMS hour anchored at the newest snapshot (window_rule + newest_mrms_hour_utc in current.json); buffalo_assemble.py exports qpe_stale/qpe_age_min/snapshots_last_24h and skips the magnitude band while stale; radar_eval.py --truth-file/--resettle with truth_source per row (494 frames re-settled on the F7 archive; the ledger-truth rows kept as radar_settled.jsonl.bak.20260907_1953_g2_ledger_truth); boundary test test_g2_boundary_windows in the F6 suite (30/30); evidence research/buffalo_windows/REPORT.md.
Version 2026.9.7.2 — September 7, 2026
The Buffalo historical archive now refreshes itself monthly.
- On the 2nd of each month the four gauge archives under Historical Levels & Storm Archive pull the latest USGS readings (including any revisions to recent provisional data), extend the radar and daily rainfall records, and rebuild — so "this date in history" and the storm-rise catalog keep filling in through the current season without anyone touching them. The "Built" stamp at the bottom of each page shows the last refresh.
Not user-facing, recorded for completeness: harness cron monthly_refresh.sh (USGS full re-pull, nClimGrid last 6 months, MRMS zone rain extended from NOAA S3 with the F7 catchment cell lists, build_archive.py, scp as the new agenthost-hist DMZPi identity, public-URL verify, failure email). No Flask change.
Version 2026.9.7.1 — September 7, 2026
New: a historical levels and storm archive for the four main-stem Buffalo gauges — Ponca, Pruitt, St. Joe (Hwy 65) and Harriet (Hwy 14) — reaching back to 1939 at St. Joe.
- This date in history. Pick any calendar date and see the river level on that day in every year of record, ranked lowest to highest, for one gauge or all four at once. Prompted by the outfitters' "September 6 over the years" look-backs — now anyone can do that for any date.
- Days in each level, by year. How many days a year each gauge spent Very Low, Low, Moderate, High and in Flood, using the National Park Service's own river-level categories for that gauge.
- A typical year. The seasonal curve of the river (median and the middle 50% / 80% of years) with any single year overlaid, so you can see how this year compares.
- Low-water season. When the river usually drops under 200 cfs and how long it stays there.
- Every storm rise on record, searchable by year, month and level reached: how much rain caused it, where in the watershed it fell, how high the river crested, and how long it stayed High or in Flood. Rainfall comes from NOAA radar (2015 on), the NOAA/NCEP Stage IV analysis (1997–2015) and NOAA's nClimGrid daily grid (1951 on). Rises before the 15-minute gauge record (St. Joe 1939–1991) are detected from daily means and marked "daily-resolution era".
- Find it under Buffalo Intelligence → Historical Levels & Storm Archive, or directly at
/buffalo/historical/. The old Ponca historical report now lives inside this archive (its old address forwards).
- Experimental in the sense that the storm-rise detection is automatic; if an event looks wrong, use the Page Suggestions link. Recent USGS data are provisional and may be revised.
Not user-facing, recorded for completeness: new research/buffalo_historical/ on the harness (full USGS period of record for the four gauges, hourly Stage IV 1997–2015 and daily nClimGrid 1951–present crops of the Buffalo basin, NPS threshold provenance, event pipeline build_archive.py); DMZPi gets buffalo_historical/ (shared JS shell + per-gauge JSON) and three new routes; /ponca/historical/ routes are 301s; pages registered for the sitemap/llms.txt.
Version 2026.9.6.1 — September 6, 2026
No visible change: the gauge-corrected rain product was back-tested against every rain-driven predictor, and only one piece (the Richland physics engine) earned a possible switch.
- The new gauge-corrected radar rain archive (October 2020 to now) was replayed through the Buffalo alerts, the creek rain triggers, the physics rise engines, the neural nets and a research Boxley predictor, each against an identical radar-only twin on the same storms. Nothing on the site changed. The Richland physics engine is the one case that improves without a catch-rate cost (fewer false rises, especially in summer); the Cossatot and Hailstone physics engines would trade a few points of catch rate for far fewer false alarms, which is a judgment call rather than a clear win; the other predictors gain little once the live hour has to come from the radar-only product (the corrected product arrives about an hour late). A switch for Richland is designed but not deployed.
Not user-facing, recorded for completeness: research/mrms_pass2/backtests/ (REPORT.md, PLAN.md, per-family RESULT.md, SNAPSHOT_TREE_DESIGN.md); archive extended with a Mulberry/Big Piney window; no Pi file changed; qwen3.8:27b stopped for two GPU training blocks on the inference host and reloaded. Model-review fault F7, work package 7b item 2.5.
Version 2026.9.5.3 — September 5, 2026
No visible change: the Buffalo rise-size band was re-checked on the new catchment zones and kept as is.
- After the Buffalo zones became true gauge catchments (2026.9.4.7), the rise-size band and rise-probability model were refit on twelve years of the new zone rain and compared, gauge by gauge, with the version that is live. The live version held up as well or better at three of the five gauges on the last two years of storms, so nothing was replaced. The rain-trigger thresholds behind the Buffalo local-rain alerts were re-measured the same way and also hold.
- Background finding, no page change: the radar rain record has a step at the October 2020 archive change (today's radar-only product reads about 11 percent higher against rain gauges than the pre-2020 record did, while the gauge-corrected product matches the gauges within 2 percent). This now informs how future models are trained; a gauge-corrected rain archive is being built.
Not user-facing, recorded for completeness: magnitude_hourly.json unchanged (same-basis refit gated against the shipped artifact on the same rows, research/f7_catchments/wp7b/); buffalo_config.yaml unchanged (research/f7_catchments/wp7b/threshold_rules_live_vs_cand.csv); era ratio research/era_ratio/REPORT.md; replay rigs re-synced to the live reach models with the catchment zone rain as default; legacy per-HUC12 archive retired; MultiSensor Pass2/Pass1 archive fetch running (research/mrms_pass2/). Model-review fault F7, work package 7b (items 2.1, 2.2, 2.4, 2.6); plan research/f7_data_layer/WP7B_PLAN.md.
Version 2026.9.5.2 — September 5, 2026
The Cossatot empirical forecast now reads rain over exactly the ground above the Vandervoort gauge.
- The empirical "chance of rising" card for the Cossatot had been averaging radar rain over a basin outline that included about 10 percent of ground below the gauge. It now uses the same 232-cell upstream area the physics predictor uses (89.6 square miles, matching USGS), and its probability table was rebuilt on that rain. Twelve years of storms show the two outlines almost never disagree (no storm differs by more than 10 percent), so the card's numbers barely move; this is a consistency fix, not a new model. Experimental, as before.
- Checked and deliberately left alone: the Boxley/Hailstone and Richland empirical rain outlines, and the Q-bucket drainage areas, all sit within noise of their gauge catchments. Nothing else on the site changed.
Not user-facing, recorded for completeness: empirical_forecast/data/shared_qpe_indices.json (basins.cossatot = cossatot/shared_qpe_indices.json zone upstream, 232 cells; other basins unchanged) + the Cossatot block of calibrated_table.json (version 2.2026-09-05, rebuilt on the F2 v2 rig from the Cossatot pixel archive; out-of-sample skill equal to or slightly above the old mask at every tier). Also this session: the live radar-grid convention was confirmed against the pixel archives (only the legacy per-HUC12 research archive was offset by one column, now retired), the Buffalo replay rigs were re-synced to the live reach models with the catchment zone rain as their default, and the MultiSensor Pass2 archive fetch was started. Model-review fault F7, work package 7b (item 2.3); plan research/f7_data_layer/WP7B_PLAN.md.
Version 2026.9.5.1 — September 5, 2026
Mulberry and Big Piney recession countdowns were running fast, sometimes by half in winter; they are rebuilt from each river's own record and now show a typical range.
- The "Time to Too Low" and "Time to Low but Floatable" countdowns on the Mulberry and Big Piney section cards were built from a curve that only counted the hours the river was actively dropping. On twelve years of record the real wait to reach "Too Low" was typically 1.4 to 1.8 times what the card said (Mulberry both sections, Big Piney below Longpool), and in winter and spring at Big Piney it was about double; the scorecard had been recording a third or more of those countdowns as "never reached" because the river was still above the line when the card's window closed. The countdown is now looked up directly from what the river did every time it sat at this level in the cool (December–May) or warm (June–November) half of the year: the middle of that history is the number shown, and a small "typically X–Y" line under it gives the range that covered about half of the past cases. Checked on held-out years, the new numbers are within a few percent of the real wait at every countdown on both rivers, winter included.
- The section cards now say plainly that the countdowns assume no significant new rain. Small showers are already inside the numbers; a real storm restarts the clock.
- One known limit, seen in this June's replay: in the first days after a large flood crest the river drains slower than its typical history, so the countdowns issued right after a big rise still read short until the recession settles in. Experimental: an antecedent-aware version is a possible follow-up.
- Richland and Cossatot were checked the same way and were not biased; their countdowns are unchanged.
Not user-facing, recorded for completeness: tables blocks for 07252000 / 07257006 in recession_eval/recession_curves.json (v3; builder research/recession_eval/build_tables_stage.py, cool/warm season pools, LOYO + holdout in VALIDATE_TABLES_STAGE.md, live-ledger replay in ledger_replay_mbp/REPLAY.md); basinlib/assemble.compute_recession table path with curve fallback (serves both assemblers); DMZPi per-section renderer range line + caption (one restart); simulator snapshot refreshed, bundle unchanged, golden 2,200 / 0. Both gauges' stage ratings checked for drift (≤ 0.2 ft at the served thresholds) — tables stay in feet. Plan and decisions in MULBERRY_BIGPINEY_RECESSION_PLAN.md §7; ARCHITECTURE §16 item 42.
Version 2026.9.4.7 — September 4, 2026
The Buffalo page's gauge zones are now the gauges' real catchments, so "rain over the Boxley zone" means rain that actually flows past the Boxley gauge.
- The Buffalo dashboard had grouped the radar-rain cells by whole HUC12 sub-watersheds. That left the Boxley zone at 58 percent of Boxley's true drainage, put part of the ground below the Ponca gauge into Ponca's zone, and gave Bear Creek 14 percent too much. Each zone is now the USGS-delineated catchment between one gauge and the next (437 of 3,540 cells changed zone; Boxley's zone grew from 87 to 154 cells, Ponca's shrank from 300 to 147, Pruitt's grew from 120 to 197). Every catchment now matches the USGS drainage area within about 1.5 percent (Boxley 3.8). Two HUC12s on the map carry a new gauge label: Smith Creek–Buffalo River now reads Boxley, and Whiteley Creek–Buffalo River reads Pruitt, because most of each drains to that gauge.
- Before switching, the whole engine was replayed on twelve years of rain re-cut to the new zones. Rise rates, flood detection and lead times held at every gauge; the wave-router reach models were refit on the new rain (they barely moved and ship with the switch so fit and serve agree); the rise-probability and size bands keep their coverage on the new rain at every gauge and are unchanged. Zone rainfall totals barely move (Boxley −2 percent, Pruitt +2), but roughly one storm in six over the headwater zones shifts by more than 10 percent, so the zone-based numbers on the page will differ slightly from what the old zones would have shown. Experimental.
Not user-facing, recorded for completeness: buffalo_dashboard/shared_qpe_indices.json v2 (NLDI basins rasterized on true cell centers; Bear Creek WBD-divide patch; gauge_zone_v1 kept on each cell), buffalo_config.yaml zone_summary areas + USGS drainage_km2 + two huc12_zones labels, reach_models.json = the v2 refit. Evidence: research/f7_catchments/ (REPORT.md, history/, refit_replay/, refit_wp4/). Side findings recorded there: the per-HUC12 research archive samples MRMS one column west of the live grid convention; the shipped magnitude artifact is already stale against the live engine's longer timing windows (a same-basis refit is queued). Model-review fault F7, work package 7a; plan research/f7_data_layer/F7_PLAN.md.
Version 2026.9.4.6 — September 4, 2026
Groundwork for the USGS data-service retirement: one shared gauge-fetch module with the replacement service running in shadow, checked every hour against the live feeds.
- USGS is retiring the data service every gauge feed on this site uses (degrading from August 2026, gone in early 2027). The site now has a single fetch module that speaks both the old service and USGS's new API, and an hourly check that pulls the last three hours through both and compares them with each other and with what the live feeds actually stored. The scorecard's "Where the truth comes from" line reports the running tally. The first check: 216 readings compared between the two services with no differences, 203 compared against the live feeds with no differences. Nothing switches until a clean week is on the record.
- The nightly truth refresh now pulls through the shared module too.
Not user-facing, recorded for completeness: basinlib/usgs.py (fetch(sites, codes, period_hours|start/end, backend="iv"|"ogc") → the fetchers' exact record shape; OGC adapter renders UTC times as the legacy local-offset strings, maps approval status to A/P; cursor paging; precip_sensor_status() for the dead-tipping-bucket case; --selftest, --live), prediction_eval/usgs_shadow.py (cron :42; usgs_shadow_state.json + usgs_shadow.jsonl), truth_refresh.pull() via the module, score_all truth block usgs_shadow, DMZPi sentence. The five live fetchers still call the legacy host directly; they move to the module after the shadow week. Model-review fault F7, work package 6; plan research/f7_data_layer/F7_PLAN.md.
Version 2026.9.4.5 — September 4, 2026
Radar-rain bookkeeping fixes under the hood: missing-radar hours are now counted instead of read as dry, rolling rain totals say how many hours they actually cover, and every "Central date" in the pipeline is a real Central date.
- Nothing changes on the pages today. When a radar site drops out of the national mosaic, the affected cells used to arrive as "0.00 inches" with no trace; each hourly radar snapshot now records which cells had no radar, and each basin's hourly rain record carries the share of its area that was missing, with an "insufficient" flag above 10 percent. The 1- to 24-hour rain totals now say how many hours they cover and which hours are missing, instead of silently stretching across a gap.
- All the places that computed a "Central date" with a fixed six-hour offset (one hour off all summer) now use the real America/Chicago clock, so a prediction or a day's rainfall issued in the first hour after midnight lands on the right day.
Not user-facing, recorded for completeness: shared_qpe/fetch_shared_qpe.py (to_inches counts sentinels before clipping; snapshots gain n_missing/frac_missing/missing_idx; collect_recent_snapshots selects by wall-clock hour with a previous-hour fallback; windows gain hours_covered/hours_expected/window_end/missing_hours/frac_missing_max; sums bit-identical on complete hours), backfill_shared_qpe.py prefers the HH:00 key, the basinlib/Cossatot/Richland readers write frac_missing + insufficient, buffalo_read_qpe counts missing cells for real, and ZoneInfo replaces UTC−6 in the Cossatot/Richland predictors, four assemblers, accumulate_daily_precip.py, and the three prediction archives (the physics gauge reader walks both day-file keys). Not changed: the Hailstone reader/predictor pair (consistent with each other) and the 7-day antecedent (display only since engine v2). Model-review fault F7, work package 5; plan research/f7_data_layer/F7_PLAN.md.
Version 2026.9.4.4 — September 4, 2026
Every scorecard section now reports the same sample line, the physics crest grades get an honest window, and countdowns and forecasts are counted by event as well as by hour.
- Under each section header the page now states its sample the same way: how many records, how many distinct events they belong to, the window applied, and the date of the gauge data they were checked against. Hourly re-issues of the same call are no longer the only sample size on the page.
- Physics crest predictions (Richland, Cossatot, Hailstone) are now graded over a fixed 18-hour window after they are issued instead of a window that closed 2 hours after the predicted peak, which could never see a later crest. Each prediction is graded on three separate questions: did the gauge rise at all, how far off was the crest when it did, and how far off was the timing. Predictions whose crest was not above the current level are no longer archived as claims. The first rows on the new basis appear 18 hours after the next rain-driven prediction; the old grades stay under "Legacy grading". Experimental.
- Recession countdowns add an "Episodes" column: one run of hourly countdowns for a gauge and target counts once, and the timing error of the last countdown before the crossing is shown beside the all-records average. Harriet's 105 summer countdowns to floatable were two crossings, 0.2 hours off on the last call.
- The empirical engine's "No rise indicated" calls (below 8 percent) are graded on their own miss rate, stated under the table, separately from the real claims.
Not user-facing, recorded for completeness: prediction_eval/schema.py (episodes, event rates, Brier/reliability/BSS, coverage, log error, the standard header block, selftest); physics settlers v2 (grade_version, horizon_h 18, level_at_issue_*, rose, crest_err_ft / crest_log_err, timing_err_h, falling-limb claims skipped at archive time; grade_v2() reused by resettle.py); score_all _physics_v2, recession episodes, empirical non-claims, schema blocks on every family, v2-aware physics flags; DMZPi _sc_schema_line + section changes. Model-review fault F7, work package 4; plan research/f7_data_layer/F7_PLAN.md.
Version 2026.9.4.3 — September 4, 2026
The scorecard now re-grades itself as USGS revises its gauge data, and says where its truth comes from.
- USGS publishes gauge readings as provisional and keeps revising them for months; at Ponca, 94 percent of this year's readings have changed since they were first reported, with the July and August low flows cut by more than half. Every grade on the scorecard had been settled once against the first value seen and never revisited. Each night the system now re-pulls the last 120 days from USGS, with each reading's approval status, updates its gauge archives in place (keeping the value it first saw), and re-grades every settled prediction whose readings are not yet approved. A grade that moves keeps its first verdict on the record, and a grade whose readings are all approved is marked final and left alone.
- The first pass re-graded 172 Buffalo rise records (56 changed verdict), 778 recession countdowns (165 changed verdict, mostly timing shifts at Ponca and Harriet), 48 physics crests, and 678 Ponca event-analysis rows (21 changed class); no probabilistic empirical grade moved. The scorecard's header now carries a "Where the truth comes from" line with the last refresh time and, per family, how many grades were checked, moved, re-graded so far, and final.
- The Hailstone crest grades now read Boxley's flow from the same refreshed archive instead of a 24-hour snapshot, so they can be re-graded like the rest.
Not user-facing, recorded for completeness: prediction_eval/truth_refresh.py (00:20, NWIS IV with qualifiers into the six truth archives under each fetcher's flock: value/value_first/qual, removed readings marked X, per-file truth_refresh block) and prediction_eval/resettle.py (01:05, per-family re-grade with v0 from the same archive; outcome_first/actual_first/first_actual_peak_*, resettled_at, resettle_n, truth {archive, vintage, qual mix, final}; physics and empirical day files re-pushed); score_all truth block; hailstone settler _read_study_readings; tarball backups prediction_eval/{ledgers,truth_archives}.bak.20260904_0914_f7_wp3.tar.gz. Waves and the neural graders are not resettled (documented). Model-review fault F7, work package 3; plan research/f7_data_layer/F7_PLAN.md.
Version 2026.9.4.2 — September 4, 2026
Scorecard bookkeeping fixes: every countdown now gets graded, the Buffalo Rise table uses the same 45-day window as the rest of the page, and the physics crest grades no longer close before the gauge data has arrived.
- Long recession countdowns (St. Joe, Harriet, Mulberry, Big Piney) could mature after the grader had stopped looking at that day's records, so some were never graded. The grader now revisits every record until it is settled.
- The Buffalo Rise table had been showing all-time numbers while every other section showed the last 45 days. It now shows the last 45 days with the all-time totals stated underneath.
- USGS gauge readings arrive 35 to 99 minutes behind real time, and the physics crest grades (Richland, Cossatot, Hailstone) were being closed 15 to 30 minutes after their window ended, so the true peak sometimes had not landed yet. The grader now waits for the readings to reach the end of the window (up to 3 hours) and marks the rare grade that still had to close short. The 64 past grades that had closed short (41 Richland, 4 Cossatot, 19 Hailstone) were re-graded on the complete window; the original grade is kept on the record.
- The "Within ±20%" column in the physics table now shows how many predictions it was computed on, since dry-bed and zero-flow rows are excluded from that test.
Not user-facing, recorded for completeness: recession_archive.settle_ledger scans all ledger files; buffalo_predictions_archive.score(cutoff=) + score_all all_time block; wave_eval.score windows on expire_after; basinlib/predictions_archive.py and hailstone_predictions_archive.py gain TRUTH_SLACK_MIN/TRUTH_WAIT_HOURS and store truth_last_reading / truth_truncated; score_physics adds n_within_eligible, n_truncated, n_resettled; one-shot resettle_physics_truncated.py (records keep first_actual_peak_*, resettled_at, resettle_reason). Model-review fault F7, work package 2; plan research/f7_data_layer/F7_PLAN.md.
Version 2026.9.4.1 — September 4, 2026
The scorecard now grades the Boxley gauge, which it had silently skipped since June.
- The Prediction Scorecard's Buffalo Rise table has always listed Boxley with nothing graded. The reason was a data gap, not the river: the archive the scorecard checks predictions against recorded Boxley's stage but never its flow, so every Boxley call was filed as "could not decide". Flow is now archived for Boxley (and Richland Creek), the gap was filled back to late February from USGS, and the 57 Boxley rise predictions and 77 Boxley recession countdowns issued since June have been graded. Boxley's rise row now shows 34 rises seen and 23 not, over 6 rain events (5 of which rose); its recession countdowns landed within about 3 hours of the predicted time on the median. Experimental, like the rest of the page.
- 28 long recession countdowns at St. Joe, Harriet, Mulberry and Big Piney that had matured but never been graded were graded in the same pass; a permanent fix for why they were missed follows in the next update.
Not user-facing, recorded for completeness: buffalo_study/study_config.yaml adds USGS 00060 for 07055646 and 07055875 (no extra USGS calls — the batched fetcher already asked for it); NWIS backfill of 36,272 readings into 190 study day files under the fetcher's flock after a tarball backup; censored Boxley ledger rows re-opened and settled with the current rise rule (research/f7_data_layer/deploy/). This is work package 1 of model-review fault F7 (data layer and scoring); the review measured how much USGS's later revisions to provisional data change settled grades (Ponca is the outlier — 94 % of its 2026 readings have since been re-rated) and the plan is in research/f7_data_layer/F7_PLAN.md.
Version 2026.9.3.8 — September 4, 2026
Buffalo local-rain rise predictions now say how likely a rise is and how big it would typically be, in cfs, instead of "slight / moderate / large".
- The size word on the Buffalo gauge cards was never a size: at the first hour a storm triggered a prediction it read "slight" 90 to 97 percent of the time, and across twelve years a first-fire "slight" preceded rises of 400 cfs at Boxley and 2,000 at Harriet as often as small ones. Each card now shows a chance of a rise (for example "62% chance of a rise in 13–27 hr") and, if it rises, a typical range in cfs with the median. The range is the 25th to 75th percentile, so about half of real rises land inside it, and the confidence word now follows the probability.
- The numbers come from a model fit hour by hour on the same records the scorecard grades, using the rain over the gauge's own catchment, the river's level and rate, how long since the storm triggered, the season, and the next gauge upstream. On twelve held-out years the range covered 48 to 50 percent of real rises at every gauge, and the probability is far better calibrated than the old word (mean squared error 0.17 to 0.20 against 0.25 for the word). Replayed on this summer's 341 graded predictions it held 74 percent of the rises; on the August 8 St. Joe downburst all 17 records stayed in band while the chance fell from 94 to 29 percent as the storm ended.
- The scorecard's Buffalo Rise table adds "Band held" and a probability skill score. Experimental: the ranges are wide on purpose, especially at St. Joe and Harriet where most big rises arrive as tracked waves rather than local rain.
Not user-facing, recorded for completeness: buffalo_dashboard/magnitude_serve.py + magnitude_hourly.json (per-gauge logistic rise probability and p25/p50/p75 quantile regression on log(rise + 10), with a no-upstream variant; stdlib), predict_local_rainfall emits rise_prob / rise_p25/p50/p75_cfs, the ledger stores them and the grader adds mag_in_band + Brier, DMZPi card and scorecard (one restart), simulator engine port + golden local family (2,200 cases). The category word stays in the JSON for the ledger and the wave-compat merge. Fault F6 work package WP3; evidence research/buffalo_f6/wp3b_hourly/REPORT.md (the first-fire event model in wp3_magnitude/ is the recorded comparator).
Version 2026.9.3.7 — September 4, 2026
Buffalo wave arrival windows now account for how much water is in the river, and the local-rain timing windows were re-measured against twelve years of graded predictions.
- The "incoming wave" arrival window used to depend only on how big the upstream crest was. A wave rides pools and takes longer at low water, so the window now also uses the downstream gauge's base flow and the predicted crest size. On held-out storms the typical arrival-time error on the long reaches drops by about half an hour to an hour and a quarter (Pruitt to St. Joe 3.3 to 2.9 hours, Ponca to St. Joe 4.3 to 3.3, Pruitt to Harriet 4.4 to 3.5); short reaches are unchanged.
- While the upstream gauge is still rising, the window is no longer pinned to the moment of the running peak. It now opens at the earliest plausible arrival and closes at the latest plausible one given how much more rise is expected. On held-out storms the actual crest fell inside the window 61 percent of the time, up from 30 percent; the price is a wider window while the upstream rise is still in progress.
- The local-rain "rise in X–Y hr" windows on Pruitt, St. Joe and Harriet were about half the time the river actually takes, measured on the same basis the scorecard grades them. They are re-measured: Pruitt 13 to 27 hours, St. Joe 19 to 36, Harriet 21 to 33 (Boxley and Ponca move by an hour or two). The dry-ground stretch stays on top, and the point at which a prediction is dropped for lack of movement is unchanged, so false alarms are not held longer.
- Small model corrections on the tributary reaches: the Richland and Bear Creek crest models no longer swing by a factor of ten on a rain-distribution ratio that the training data never reached, and Bear Creek drops that term entirely, which tested better on held-out years. Experimental.
Not user-facing, recorded for completeness: wave_router.py lag_quantiles (stage-conditioned quantile regression with the U-band table as fallback), provisional-window policy C with per-launch-gauge remaining-rise models embedded in reach_models.json, per-reach ratio guards (TRIB_RATIO_CLIP 1.5, TRIB_RATIO_MIN_RAIN 0.10), Bear Creek no-ratio crest model, buffalo_config.yaml local windows; 48 unit tests; router gate and rise-ledger replay gate under research/buffalo_f6/; simulator engine/bundle rebuilt (golden 2,200). Fault F6 work packages WP4 and WP6; evidence research/buffalo_f6/wp4_waves/REPORT.md.
Version 2026.9.3.6 — September 3, 2026
The Buffalo wave tracker now follows a storm's second pulse and no longer confuses an earlier crest for the wave it is watching.
- When an upstream gauge crests, falls a little, and then climbs to a higher crest in the same storm, the downstream "incoming wave" card used to go quiet after the first arrival and say nothing about the bigger second pulse. It now re-opens the prediction for the higher crest. In twelve years of replayed storms this lifts flood watch detection at Ponca from 75 to 80 percent and at Pruitt from 92 to 96 percent, with no change in how often a watch turns out to be right.
- While the upstream gauge is still rising, a crest at the downstream gauge is no longer counted as the wave "arriving"; that crest belongs to an earlier pulse, and the real arrival is still ahead. Experimental.
Not user-facing, recorded for completeness: wave_router.py re-open logic (REOPEN_FRAC 1.10), provisional-arrival guard, trough-based base_cfs backfill, per-cycle notes logged by the assembler; 34 unit tests; replay gate research/buffalo_f6/fidelity/scripts/gate_router.py. Fault F6 work package WP5.
Version 2026.9.3.5 — September 3, 2026
Buffalo rise predictions are now graded on a scale-aware rule, the scorecard counts events as well as hours, and the gauge arrows finally see slow low-water rises.
- A predicted rise on the Buffalo page used to count as verified only if the gauge climbed at least 25% and at least 30 cfs. At summer base flows that second bar hid real rises: on August 24–25 St. Joe climbed from 55 to 84 cfs, exactly the "slight rise" that was predicted, and 35 of 36 records were scored as misses. The bar now scales with the river: 25% plus a floor that runs from 10 cfs at low water up to 30 cfs, the same rule the engine uses to decide a predicted rise has arrived.
- The scorecard's Buffalo Rise table now shows events beside hourly records. A 35-hour false alarm used to count as 35 misses and a 5-hour hit as 5 hits; the new columns count each storm once. Records graded before today keep their old verdicts, so the 45-day window mixes the two rules until mid-October, and the page says so.
- The RISING / FALLING arrows on the Buffalo gauge cards were tuned for big water and read STABLE through every slow rise this summer (St. Joe showed STABLE on all 41 records where a real rise followed). Each gauge now has its own floor, scaled to the current flow, and needs the movement to hold over two hours before it flips the arrow. On twelve years of quiet days that flags a false rise less than 0.6% of the time; on the slow low-water rises the old rule missed, it catches 86–100% of them. The arrow can still flicker between RISING and STABLE on a very slow creep. Experimental.
Not user-facing, recorded for completeness: basinlib/rise_rule.py (one rule for grader and engine), prediction_eval/buffalo_predictions_archive.py (settle + event-level score, rise_floor_cfs in every outcome), buffalo_assemble.py (consumption/movement bar on the shared rule; compute_trend per-gauge trend_floor with 2-h persistence, fit in research/buffalo_f6/wp2_trend/), buffalo_config.yaml trend floors, DMZPi scorecard columns (one restart). Fault F6 work package WP2.
Version 2026.9.3.4 — September 3, 2026
Buffalo recession countdowns were running fast, sometimes by half; they are rebuilt from the river's own record and now show a typical range.
- The "Time to Low but Floatable" and "Time to Too Low" countdowns on the Buffalo page (Boxley through Harriet, and the Hailstone card) were built from a curve that only counted the hours the river was actively dropping and skipped every flat or bumpy hour. On twelve years of record the real wait to reach the threshold was typically 1.3 to 2.3 times what the card said, worst at St. Joe and Harriet. The countdown is now looked up directly from what the river did every time it was at this level in this season: the middle of that history is the number shown, and a small "typically X–Y" line under it gives the range that covered about half of the past cases. Checked on held-out years, the new numbers are within a few percent of the real wait at all five gauges.
- The countdowns now say plainly that they assume no significant new rain. Small showers are already inside the numbers; a real storm restarts the clock.
- The confidence word under the countdowns is more often "medium" or "low" than before. That is honest: it now reflects a range that was measured, not one that was guessed narrow.
- The Buffalo Simulator's recession math was rebuilt to match, and a check against the live code found the simulator had missed August's drought-window change on the rise cards; that is fixed too.
Not user-facing, recorded for completeness: recession_curve.transit_time_table() + tables in recession_curves.json (Buffalo five; build_tables.py, LOYO in research/recession_eval/VALIDATE_TABLES.md); compute_recession table path in buffalo_assemble.py / hailstone_assemble.py; recession_archive.py stores the served band and settles on its upper end + 48 h; DMZPi range line + wording (one restart); simulator bundle/engine + golden family (2,200 cases). The Buffalo replay rig was corrected to replay the live wave router (baseline v8; it had been grading the retired propagation layer since July). Fault F6 of the 2026-09-01 model review, work packages WP0 + WP1; evidence in research/buffalo_f6/.
Version 2026.9.3.3 — September 3, 2026
The Cossatot neural net was retrained without the two gauges that stopped reporting, and its alerts now come in two levels.
- The Cossatot network and its companion crossing model were trained in July with two neighbouring gauges, Board Camp Creek and the Mountain Fork, that went silent in early August. Both were retrained today on the Cossatot's own gauge and the radar alone, so the Degraded banner from this morning is gone and the page no longer depends on data that is not arriving. On twelve held-out years the new hour-by-hour level table beats "no change" by a wider margin than the old one on ordinary hours, and never contradicts itself between its low, middle and high estimates. It is a little less sharp than the old one in the first three hours of a rise, and we say so on the scorecard rather than hide it.
- The crossing model lost some precision without those gauges, and the page reflects it honestly: on held-out years it still catches about 9 in 10 floatable rises, but roughly 1 alert episode in 4 does not verify, where it was about 1 in 8 with the gauges. Alerts now come in two levels, WATCH (the level that catches 9 in 10 rises) and ALERT (the level where about 6 in 7 episodes verify), and a percentage is shown only when one of them is up. Replayed on the one floatable rise since the pages launched, July 15, the new model raised a watch four hours before the river reached floatable and an alert an hour later, with no false alarms anywhere else this summer. Experimental.
Not user-facing, recorded for completeness: neural_cossatot/final_export.npz (cossatot-pixel-v6b, Cossatot-only, serving tables from its own LOYO OOF), gbm_export.npz (115-feature no-aux ensemble, isotonic + watch/alert operating points), gbm_infer.py v2 tiers, DMZPi Cossatot renderer (one restart). Evidence in research/neural_f5/REPORT.md §3–4.
Version 2026.9.3.2 — September 3, 2026
The neural-net pages now show calibrated crossing chances, and only when the model has earned a watch; the Cossatot page says plainly which of its inputs are missing.
- The Buffalo neural-net page (Ponca and Boxley) used to print the network's raw output as a "chance of crossing". Checked against twelve held-out years, hours it scored 50–70% actually crossed about 3% of the time. The number is now calibrated to observed frequency, and it appears only once the model's signal passes a level that was validated on real events: a WATCH badge (the level that catches about 9 in 10 crossings) or an ALERT badge (the level where about 4 in 5 alerts verify). Below that the page says "no watch" instead of quoting a percentage.
- The hour-by-hour level tables on both neural pages are steadier in dry spells. The network's predicted change is now scaled back toward "no change" when little rain has fallen in the last two days, and the likely range is sized so it contains the truth about half the time, in dry and wet weather alike. Over twelve held-out years this turns a table that was slightly worse than "no change" on quiet hours into one that beats it at every horizon, while keeping most of its skill when the river is moving.
- The Cossatot neural page no longer colours its level table green or red on its own. Green (floatable) and red (high water) now follow the companion crossing model, which is the part that was validated on real events; when the table climbs to a threshold the crossing model does not agree with, the page says so. Two neighbouring gauges the Cossatot network was trained with, Board Camp Creek and the Mountain Fork, stopped reporting in early August. The page now carries a Degraded banner naming them, instead of silently feeding zeros while claiming to use them. A Cossatot-only model is being trained to remove the dependency. Experimental.
- The scorecard's two neural sections now grade the network against the honest baseline of "the level doesn't change": average error of both side by side, skill on all hours and on hours when the river actually moved, how often the likely range contained the truth on calm versus moving hours, and how well the crossing chances matched what happened.
Not user-facing, recorded for completeness: serving tables (isotonic calibration, regime-stratified shrinkage, band scales, watch/alert thresholds) live inside final_export.npz and are applied by neural_infer.py v2; producers v2 (unrounded ledger with raw and served values, --dry-run, degraded/aux status); graders rebuilt on prediction_eval/neural_eval_lib.py (persistence columns, all-hours skill, no rounding slack, Brier, tier episodes, settle retry); shared_qpe/fetch_shared_qpe.py now fetches the top-of-hour MRMS frame (HH:00) instead of the newest 2-minute file (HH:02) — every basin's hourly radar snapshot is now the same hour the models were trained on; DMZPi neural pages, scorecard sections and the neural math page (one restart). Fault F5 from the 2026-09-01 model review; evidence in research/neural_f5/REPORT.md.
Version 2026.9.3.1 — September 3, 2026
The Ponca card's radar estimate is now graded on the scorecard, and its wording follows what the grades showed.
- New scorecard section, "Radar Nowcast (Ponca card)": every time the card said "roughly another X inches falls in the next two hours" or "this rain is about done", it is now checked against the rain that actually fell, frame by frame, back to July. So far: when it said more was coming it was right 78% of the time, but it only called about 44% of the real bursts in advance; "about done" verified 93% of the time.
- The wording changed to match. Slow-moving storms and storms raining hard right now have delivered more than the radar picture suggested, so the card now says "at least" in those cases; fast movers delivered about half, so it says "up to". When storm cells are breaking up or changing shape rather than moving cleanly, it no longer quotes a number at all. And when radar shows little behind a band that is still raining, it no longer says the rain is done, because in two cases out of three more fell anyway.
- The "past storms like this went on to a real rise" line was wrong every time this summer, because it compared today's storm with storms from wetter ground and cooler months. It now compares only with storms that started from a similar creek level in the same season, and stays quiet when there aren't enough of those. Experimental.
Not user-facing, recorded for completeness: prediction_eval/radar_eval.py (new family, folded into score_all.py), ponca_analog.py 2026.9.3.1 (radar_summary logs motion spread + lead coverage; radar_facts regime wording; projected_context baseline+season cohort), ponca_analog_library.npz v3.1 (adds month + baseline arrays, k-NN data unchanged), DMZPi scorecard section (one restart). Evidence in research/ponca_analog_v2/REPORT.md §6.
Version 2026.9.2.6 — September 2, 2026
The Ponca "AI Rainfall Event Analysis" card was rebuilt: its flood chance is now calibrated, it goes quiet after the crest, and the downstream card now speaks for the wave tracker.
- The "chance >1600 cfs" number used to throw away most past floods before it looked, because it only kept storms whose rain stopped soon after the moment being matched. On June 22 it read 0% with two thirds of an inch down and Ponca at 96 cfs, three hours before an 8,700 cfs crest. Replayed over eleven years of storms, the old number caught 0.2% of floods while the gauge was still asleep; the new one catches about one in five, and its early-hours average now matches how often those storms actually flooded.
- The matching itself was retuned so that a bone-dry August creek at 9 cfs no longer looks like a wet April creek at 300 cfs, and it now knows how long the rain has been stopped. On this summer's drought storms it stays near zero, as it should.
- After Ponca has crested, the percentage disappears and the card says what it crested at. The one exception: if the wave tracker sees a second, higher push coming down from Boxley, the card shows the chance that push tops 1600 cfs and says so in plain words. The old card held "85% chance" through an entire recession.
- The "River Rise & Propagation" card for Pruitt, St. Joe and now Harriet reads straight from the wave tracker: the same crest range, arrival window (in local time) and flood watch/warning the Buffalo page already shows, so the two can no longer disagree. A wave still building upstream is labelled "at least".
- The scorecard's Ponca section now counts events, not 15-minute re-issues, grades the first call and the most-informed call of each event against the crest that actually came, and grades post-crest calls only on whether a higher crest followed. Two synthetic test events that had been polluting the old "Downstream Propagation" grades were removed and that section is retired in favor of Routed Waves. Experimental.
Not user-facing, recorded for completeness: buffalo_dashboard/ponca_analog.py rewrite (episode state, spell-start fix, uncapped k-NN, max-of-calibrated-sources probability, router-narrated propagation, --dry-run/--no-history/--selftest), ponca_analog_library.npz v3 (log-scaled features + trailing dry hours), prediction_eval/ponca_analog_eval.py v2, score_all.py (windowed event-level Ponca scorer, propagation family retired, swallowed exceptions now logged), ledgers purged of 3 synthetic rows and 28 mis-keyed rows (backups kept), SJ_MODEL archived under research/ponca_analog_v2/retired/, DMZPi two card renderers + two scorecard sections (one restart). Fault F4 from the 2026-09-01 model review; evidence in research/ponca_analog_v2/REPORT.md.
Version 2026.9.2.5 — September 2, 2026
Watershed triggers refit; FLOOD now means the gauge, not the rain; Big Piney and Mulberry join the calibrated creeks.
- FLOOD on the watersheds page, in Signal and on Facebook now means a calibrated creek's own gauge is at or above its flood level. Before, FLOOD was "twice the rain trigger", which in practice was reached after the creek had already come up and was then hidden as "already running". Rain alone now tops out at WARNING.
- Upper Big Piney is now a calibrated creek in its own right, and Haw Creek, Hurricane Creek and Spirits Creek borrow the Big Piney and Mulberry gauges as their ground-dryness read, the same way the Buffalo, Richland and Cossatot families already do. Only the two Kiamichi creeks still run on fixed thresholds.
- Every calibrated threshold was refit on eleven and a half years of storms, this time counting a catch only if the trigger fired at least an hour before the creek came up, and checked one year at a time. The Cossatot's thresholds now change with the leaf season (its summer false-alarm rate was the worst of the set). The page footer explains the new rules.
- A quirk that produced a false WARNING on Falling Water on July 15 is fixed: a neighboring drainage no longer has its threshold collapse the moment the reference creek's rain stops.
- Held alerts now say what is being held and until when, instead of showing a WARNING next to "0.00 inches of rain".
Not user-facing, recorded for completeness: creeks/qbucket_config.json (v3, five creeks, per-bucket thresholds by leaf season where adopted), qbucket_shadow.py v3 (4-day lookback, rolling discharge store, area-weighted Cossatot rain, gauge FLOOD, last-event memory), assemble_and_push.py (Tier-2 freeze, episode-max re-escalation, running-creek episodes, gauge FLOOD messages), drainages.yaml q_ref for haw/hurricane/spirits, DMZPi wording (one restart), the July checkup reminder retired. Fault F3 from the 2026-09-01 model review; report in research/q_bucket_triggers/reports/v3/.
Version 2026.9.2.4 — September 2, 2026
The "chance of rising" forecasts on the Mulberry, Big Piney, Illinois, Cossatot, Richland and Hailstone cards were rebuilt, and the scorecard now grades them honestly.
- The card now promises one clear window and means it: "within the next 12 h" on the fast creeks and "within the next 30 h" on the slow ones. Before, the card said one window while the probability had been built for a shorter one, which understated the day-ahead chance on the Mulberry by about half.
- The probability now starts from where the gauge is sitting, not just from the last week of rain. Out of sample across all basins that lifts the engine's skill by about 40% and brings forecast and observed rises into balance.
- A gauge hovering just under a level no longer gets a "chance of rising" at all. The card says "already within reach" instead, because a single blip over the line from that position is not a forecastable rise (the Illinois River did that 61 times in July and August).
- The scorecard's empirical section used to report that the engine "over-warned 100%" because it graded every hourly "No rise indicated" line as a rise call. That was a scoring artifact, not the engine. It now scores the forecasts as probabilities: how often the rise came when the engine said 10%, 30% or 50%, and whether it beats simply quoting the long-run rate. The section is empty until the first new forecasts settle, 12 to 30 hours from now.
Not user-facing, recorded for completeness: calibrated_table.json v2 (gauge-quartile strata, horizon_h, margins), calibrated.py v2, empirical_archive.py v2 records and day-file settling, score_all.py probabilistic empirical scorer + flags, digest line, DMZPi scorecard render (one restart), rolling regime factor retired, harness shadow-logger cron disabled. Fault F2 from the 2026-09-01 model review; report in research/empirical_recalibration/v2/.
Version 2026.9.2.3 — September 2, 2026
The Cossatot, Richland and upper-Buffalo (Boxley) rise predictors have been rebuilt and refit on twelve years of storms.
- The old engine only looked at the last twelve hours of rain and, on the fast creeks, threw a storm's rain away once it was a few hours old. That is why it so often stayed silent while a rise was already on its way: on Boxley it issued nothing for about four out of five real rises. The new engine follows every hour of rain through the basin for 36 hours, so a rise that is still in transit shows up on the card, and rain whose effect has already arrived is not counted twice.
- Every parameter (how much each part of the basin responds, how long it takes, how much ground soaks up first) now comes from a fit on all storms from 2014 to 2026, including the ones that fizzled, checked by holding each year out in turn. Replayed on identical hours, the new engine catches 72 to 80% of rises on Richland and Boxley where the old one caught 18 to 32%, with fewer false alarms in every basin, and its peak timing is typically within an hour instead of five to seven.
- The math now runs in cubic feet per second. The Cossatot and Richland cards still show feet, converted with a rating rebuilt every run from the last 45 days of USGS readings, so the site follows USGS's own recalibrations (Richland's 4.0 ft has meant anywhere from 700 to 1,000 cfs over the years).
- Expect the cards to call a floatable-level rise more often than before, sometimes wrongly: about half of those calls verify, versus an old engine that almost never made one. The confidence shown on each card now reflects the measured record ("medium" everywhere today), and crest size remains the least certain number.
Not user-facing, recorded for completeness: basinlib/physics_engine.py, physics_serve.py, rating.py; engine_v2 blocks in the three calibration files with fit provenance; v1 path retained as automatic fallback; fit, validation and last-45-day replay in research/physics_recal/REPORT.md. Fixes 3 and 6 of fault F1 from the 2026-09-01 model review. ScriptPi-only, no Flask restart.
Version 2026.9.2.2 — September 2, 2026
The nightly AI analysis for the Cossatot, Richland and Hailstone predictors now grades, but no longer tunes.
- Each night a local AI model reviews the day's rain and gauge trace against what the physics predictor forecast, and writes the daily analysis pages linked from those creeks' cards. Until now it also nudged the predictor's coefficients a little every night based on that one day. The tuning is switched off: the model's suggestions are still shown on the daily page under "Coefficient Adjustments (suggested, not applied)", but the predictor's calibration is now set from the long storm record instead of drifting night to night.
- The nightly grade now looks at the predictions that were actually issued during the day (the highest one), rather than whatever happened to be on the card at midnight. A day with rain and a rise but no forecast is now graded as a miss instead of being skipped. Expect the daily analysis pages to be blunter about false alarms and misses from here on.
Not user-facing, recorded for completeness: APPLY_CALIBRATION_UPDATES = False in the three *_analyze.py; select_grading_prediction() reads predictions/YYYY-MM-DD.json; event records gain suggested_changes, graded_source, graded_issued_at. Second fix from the 2026-09-01 model review (fault F1). ScriptPi-only.
Version 2026.9.2.1 — September 2, 2026
Rise predictions on the Cossatot, Richland Creek and the upper Buffalo (Boxley) now read the creek's own baseflow before deciding how much rain runs off.
- The physics rise predictors used to judge ground wetness from the last seven days of radar rain. That tier turned out to carry almost no information: on twelve years of storms it barely separated rises from fizzles, while how low the creek was sitting before the rain began separated them strongly. Every one of the last 45 days' predicted rises on these three creeks was a false alarm on parched summer ground.
- Each prediction now starts from the creek's pre-storm flow and the time of year (leaf-on vs leaf-off). Dry ground gets a "soak-in" allowance before any runoff is predicted — on a bone-dry Boxley basin about the first inch and a half of a storm disappears into the ground — while already-wet winter ground gets a bigger response than before.
- The "Conditions" line on the Cossatot, Richland and Hailstone prediction cards now shows the baseflow-based class (VERY DRY / DRY / NORMAL / WET / SATURATED) alongside the 7-day rain total.
- Experimental, as before. Replaying the change over the last 45 days of live predictions removed all 15 of Hailstone's false alarms, half of Richland's and a third of the Cossatot's. The size of a predicted crest remains the least certain part of these forecasts and is next on the list.
Not user-facing, recorded for completeness: new shared module basinlib/runoff.py; baseflow_runoff blocks in the three calibration files; fit and leave-one-year-out validation in research/physics_baseflow/REPORT.md; first fix from the 2026-09-01 system-wide model review (reviews/2026-09-01_model_review/). ScriptPi-only, no Flask restart.
Version 2026.8.31.2 — August 31, 2026
The original dark map style is back.
- Earlier today we swapped the background maps to a temporary substitute after our map provider started requiring an API key. We've since registered for a (free) key, so the maps are back on the original CARTO dark style — the slightly deeper black look the site has always had, now with proper OpenStreetMap/CARTO credits in the corner. If you still see an "API KEY REQUIRED" watermark or the lighter gray substitute, force-refresh the page; map tiles cache aggressively in the browser.
Not user-facing, recorded for completeness: all 6 tile references moved to CARTO's new keyed URL form (rastertiles/dark_all/…?key=…); the key is a public, domain-scoped client-side key recorded in ARCHITECTURE.md §10, gotcha #16 updated, mirrors and research staging copies synced. Same-day follow-up to 2026.8.31.1.
Version 2026.8.31.1 — August 31, 2026
Fixed the "API KEY REQUIRED" watermark on the maps.
- The background street/terrain maps behind the watershed views (the basin dashboard maps, Buffalo TV, and the simulator) recently started showing an "API KEY REQUIRED" watermark. That was our map-tile provider changing its free-usage rules, not anything broken on our end — none of the gauge data or predictions were affected. The maps now use a different dark basemap (Esri's Dark Gray Canvas) that doesn't require a key. It looks very slightly lighter than the old one, but everything else — the rivers, watershed outlines, gauge dots, and rain shading — is unchanged.
Not user-facing, recorded for completeness: all 6 CARTO dark_all tile references (3 in dashboard.py, plus buffalo_tv.html, buffalo_development_tv.html, buffalo_simulator.html) swapped to Esri World_Dark_Gray_Base + _Reference (key-free, {z}/{y}/{x}); attributions updated; Flask restarted; harness mirrors and research staging copies patched. ARCHITECTURE.md gotcha #16, same version.
Version 2026.8.16.1 — August 16, 2026
The AI writer behind the nightly creek analyses and the Buffalo rainfall card got an upgrade.
- The plain-English narratives on this site — the Ponca "AI Rainfall Event Analysis" card on the Buffalo page during rain events, and the nightly analysis archives for Cossatot, Richland and Hailstone — are written by a language model that runs entirely on our own hardware. That model moved from Qwen 3.6 to Qwen 3.8 today. In head-to-head testing on our other systems, 3.8 was noticeably better at sticking to the actual numbers (fewer made-up details, more honest counts) and about 1.5× faster. Nothing about what gets analyzed or when changes — same data, same triggers, same layout; you should just see slightly tighter, more accurate write-ups. Experimental as always: the AI narrates, the arithmetic decides.
Not user-facing, recorded for completeness: OLLAMA_MODEL constant flipped qwen3.6:27b → qwen3.8:27b in ScriptPi ponca_analog.py, cossatot_analyze.py, richland_analyze.py, hailstone_analyze.py (all .bak'd); DMZPi dashboard.py Hailstone analysis-archive caption updated (Flask restarted). Follows the harness fleet's 2026-08-15 blind cookoff; the inference host serves a single resident model (OLLAMA_MAX_LOADED_MODELS=1), so this also stops the creek callers from evicting the fleet's resident model on the first rain event. ARCHITECTURE.md §9 + item 31 updated to the same version.
Version 2026.8.12.1 — August 12, 2026
New "Data Sources" page — and the site now speaks fluent AI.
- There's a new Data Sources & Methodology page (linked from the home page) that explains in plain English where every number on this site comes from — USGS gauges, NOAA radar rainfall — what we compute ourselves, how every prediction gets graded nightly on the public scorecard, and how to credit the data if you reuse it.
- Behind the scenes, the site now introduces itself properly to AI assistants (ChatGPT, Claude, Gemini, Perplexity and friends). If you ask an AI about Arkansas creek or Buffalo River conditions, it can now find this site's pages directly, understand what each one is, and see machine-readable tags saying exactly which government data each dashboard is built on. The goal: when people ask an AI "is the Buffalo runnable?", the answer comes from verified local data with timestamps — not from a guess.
Not user-facing, recorded for completeness: new Flask routes /robots.txt, /sitemap.xml, /llms.txt, /data-sources/ driven by a single AGENT_PAGES registry in dashboard.py; schema.org JSON-LD (WebSite/Organization on the landing page, Dataset with USGS/NOAA isBasedOn on all 7 basin dashboards); <meta name="description"> added to those heads. Hidden/dev pages (/admin/, /buffalo/development/, unlaunched simulator) deliberately excluded from all discovery surfaces.
Version 2026.8.9.2 — August 9, 2026
Buffalo rise predictions now understand drought soils.
- During last week's storms, the Buffalo dashboard correctly flagged a coming rise at St. Joe — then gave up on it about 14 hours before the river actually jumped from 55 to 445 cfs early Saturday morning. The rain was real and the call was right; what fooled the system was six weeks of drought. Bone-dry ground soaks up the first rounds of rain and releases the rise many hours later than normal soils do, and the early signs of that rise are a slow creep rather than a clear jump.
- Rise predictions on the Buffalo pages now stretch their time window (about 50% longer) and watch for smaller early movements whenever a storm lands on drought-dry ground. Spring and normal-condition predictions are completely unchanged — we replayed 12 years of history to confirm that.
- What you'll see: after a summer storm on dry ground, a rise prediction card may stay up for a day or more, marked "drought-stretched window." That's intentional — in drought, patience is accuracy. Tested against last week's storm: with this change, the prediction would have still been on the page when the river rose. Experimental, like all rise predictions.
Not user-facing, recorded for completeness: local_dry_stretch config + frozen-at-anchor stretch and a relaxed dry-tier movement bar in predict_local_rainfall (ScriptPi); full study in research/local_rain_lag_stretch/; replay baseline promoted to v7.
Version 2026.8.9.1 — August 9, 2026
Buffalo TV: river-section levels between St. Joe and Harriet recalibrated.
- On Buffalo TV's map, each stretch of river colors itself from a nearby gauge using its own cutoffs. The four stretches between St. Joe and the Harriet gauge (St. Joe → Gilbert → Maumee → Spring Creek → Harriet) had draft cutoffs that disagreed where they met at Maumee — the map could show a floatable river just above Maumee and "Too Low" just below it in the same water. All four are now scaled consistently by drainage area, matching the already-tuned stretches around them, so the color changes gradually along the river instead of flipping at Maumee.
- Practical effect: Maumee → Harriet no longer reads "Too Low" when that water is marginally floatable, and St. Joe → Maumee is more honest about skinny water (it previously claimed floatable down to a bare 40 cfs on the St. Joe gauge).
- These per-stretch cutoffs remain a working draft — if you paddle the Gilbert-to-Harriet corridor and the colors don't match what your boat tells you, we'd love to hear it.
Not user-facing, recorded for completeness: thresholds updated in SEG_PINNED (build_river_geometry.py), the deployed and repo copies of buffalo_tv_geo.json, and the geo cache-buster bumped to 20260809a in the TV, development-TV, and simulator pages.
Version 2026.7.25.2 — July 25, 2026
The two Illinois River kayak parks are now two distinct parks.
- The Siloam Springs Kayak Park (Fisher Ford Road, AR) and the WOKA Whitewater Park (Watts, OK) were listed as one combined entry — they're actually two different parks on two different stretches of the Illinois. They're now separate: the Siloam Springs Kayak Park is rated Class II (matching American Whitewater), and WOKA shows as Class II–III (a strong II on an easy day, pushing III at big water — American Whitewater doesn't rate the park itself).
- The Upper Illinois Water Trail now reads simply Class I–II for the trail as a whole.
- On the Illinois watershed map, hovering the access-point dots for both parks now shows each park's rating.
- Live flow levels for WOKA aren't available yet — the park keeps that data internal for now. We're working on getting access; when we do, WOKA will get its own card.
Version 2026.7.25.1 — July 25, 2026
Siloam Springs Kayak Park difficulty rating corrected.
- The Siloam Springs Kayak Park was listed as Class II–III; per American Whitewater it's a Class II run. Its rating now reads Class II everywhere it appears, and the Upper Illinois Water Trail description now says "Class I–II (II at the kayak parks)."
Version 2026.7.17.1 — July 17, 2026
The Buffalo wave tracker learned from its first real test — and now grades itself.
- On July 15 an isolated thunderstorm hit Richland Creek and the wave tracker predicted a much bigger rise at St. Joe than actually arrived (the water did come — about 30 minutes after the tracker gave up watching for it). We rebuilt the model behind those predictions: it now looks at where the rain actually fell and how much water the creek pulse is really carrying, so a lone storm cell on one tributary no longer reads like a basin-wide event. Experimental, as always.
- Wave "inbound" cards now stay up longer when the river is low — slow summer water genuinely takes longer to travel, and the tracker's patience now matches that.
- The Prediction Scorecard has a new Routed Waves section: every tracked wave is now graded after the fact — did a real rise arrive, was it inside the predicted range, and did it show up inside the predicted window. Waves are rare, so this table fills in storm by storm.
Not user-facing, recorded for completeness: trib reach models refit on real zone QPE with a 48h outcome window (richland_to_st_joe adds a volume/spikiness feature, both trib reaches add a rain-distribution ratio); wave_router gains volume tracking, a QPE-outage ratio guard, and low-base-flow expiry margins; new prediction_eval/wave_eval.py closes the wave record→settle→score loop; simulator engine/bundle rebuilt and golden-gated (1800/0).
Version 2026.7.14.1 — July 14, 2026
The Cossatot gets its own Neural Net Predictions page.
- New page: Neural Net Predictions for the Cossatot, linked from a violet card on the Cossatot dashboard. Every hour, a neural network trained on 12 years of storms looks at the last 48 hours of 1 km radar rainfall over the watershed — plus the river's own recent behavior and two neighboring gauges — and projects the Vandervoort gauge hour by hour for the next 6 hours, in feet and cfs, with a likely range around each number. Experimental.
- The page also shows the chance the river rises past floatable (3.4 ft) or past high water (5.5 ft) within the next 6 hours. In 12 years of replayed history, the floatable-rise alert caught about 9 in 10 rises with about 9 in 10 of its alerts being real — the strongest numbers of any predictor on the site so far.
- Like everything here, it uses only rain that has already fallen — if more rain comes, the numbers go up, not down. It re-issues every hour as new radar arrives.
- The Prediction Scorecard now grades this predictor too, hour by hour, the same way it grades the Buffalo neural net.
Not user-facing, recorded for completeness: the model is a port of the Buffalo pixel-net architecture (CNN + GRU over a 48×96 1-km window, three gauges multi-task); crossing probabilities come from a companion gradient-boosted ensemble validated at 90.7% recall / 88.0% precision (floatable tier, bootstrap-confirmed) — the full experiment log lives in the operator harness (COSSATOT_NEURAL_PLAN.md).
Version 2026.7.13.3 — July 13, 2026
Easier-to-read text on the Buffalo, Mulberry, and Big Piney cards.
- Fixed hard-to-read text on the Buffalo page's gauge cards — most visibly the green "Falling — typically settles near…" note, which was nearly invisible on a yellow "Low but Floatable" card. Status notes are now white, and colored details (countdowns, alerts, badges) automatically switch to white whenever they'd blend into the card's own color — for example, a red "Time to Too Low" countdown no longer disappears on a green Optimal card.
- The same fix applies to the recession countdowns on the Mulberry and Big Piney section cards.
- The shaded info boxes on yellow cards are slightly darker now, so all their text stands out better.
Version 2026.7.13.1 — July 13, 2026
Tune-up after the radar feature's first storm.
- Reviewed every AI analysis the Ponca card issued through the weekend rain. The radar held up well — its "rain estimates for the next two hours" were essentially unbiased, and its "this rain is about done" calls verified 93% of the time, usually well ahead of the hourly rain totals.
- Fixed the one miss: a brief passing shower early Saturday morning was described as likely to push the river toward runnable. The analysis now recognizes a short rain band with nothing behind it and says so — "not enough to change the river much."
Not user-facing, recorded for completeness: the card's AI calls now wait longer before giving up when the inference server is busy with the nightly Buffalo study analysis (the cause of a few carried-over narratives at midnight).
Version 2026.7.11.2 — July 11, 2026
The Ponca AI analysis now tells a continuous story instead of a series of hot takes.
- Each 15-minute update of the "AI Rainfall Event Analysis" card now knows what its previous note said, so the story carries forward instead of swinging between takes — when the picture changes, the card says so plainly ("the extra rain never materialized") rather than silently reversing itself.
- The note now usually ends with when to check back — "keep an eye on it as this next batch of rain falls" — instead of leaving you guessing.
- When radar shows meaningfully more rain inbound, the analysis adds historical perspective: what storms that ended up with that much total rain went on to do. Only appears when it actually changes the odds; still observation-based, and the official predictions remain untouched.
Version 2026.7.11.1 — July 11, 2026
The Ponca AI rain analysis now watches live radar while it's raining.
- The "AI Rainfall Event Analysis" card on the Buffalo page used to see only rain totals that update once an hour — so when a line of storms was bearing down on the Ponca creeks, the card couldn't know yet, and when the rain was wrapping up, it kept hedging. Now, while rain is falling, the analysis also reads the last 40 minutes of weather radar over the watershed: how hard it's raining right now, whether the storms are sweeping through or parked in place, and whether more rain is lined up behind. The written outlook can now say "this one's still building" or "this is about done" — often well before the hourly totals catch up. Experimental: radar is observation of rain actually falling, not a forecast.
- The official rise predictions and flood-risk numbers are unchanged — they still run only on rain that has already been measured. Radar informs the written analysis only.
Not user-facing, recorded for completeness: new ponca_radar_context.py cron on ScriptPi reading MRMS PrecipRate (2-minute radar rain-rate mosaics); radar context logged alongside each analysis for later self-grading on the scorecard pipeline; zero downloads in dry weather.
Version 2026.7.10.4 — July 10, 2026
New "Understand the Math" page: how the neural net actually works.
- "Neural net" joins the math links on the Buffalo page. For the technically curious: what a neural network really is (no mystique — weighted sums, a squash function, and 138,610 learned numbers), how ours reads 48 hours of radar like a movie while watching the gauges respond, how it was trained and honestly validated on twelve years of storms, and the full lab notebook — all nine experiments, including the five that failed and what each failure taught us. Ends with the exact numbers the shipped model carries and the reasons not to over-trust it. Linked from the Buffalo header and from the bottom of the Neural Net Predictions page.
Version 2026.7.10.3 — July 10, 2026
The Neural Net Predictions now grade themselves on the public scorecard.
- New section on the Prediction Scorecard page: every hourly Neural Net prediction is now recorded and, once the hours it predicted have passed, graded against what the gauges actually did — average error, how often the likely range contained the truth, and skill versus simply assuming the level won't change. Numbers will fill in over the coming days and become meaningful once the river actually moves. Same honesty rule as every other predictor on the page: the grades are computed automatically, good or bad.
Version 2026.7.10.2 — July 10, 2026
The Neural Net Predictions page is now linked from the Buffalo River dashboard.
- New button on the Buffalo page: next to Buffalo TV you'll now find "🧠 Neural Net Predictions" — the experimental hour-by-hour level projections for Ponca and Boxley released earlier today (see 2026.7.10.1 below). Same page, now one tap away instead of hidden.
Version 2026.7.10.1 — July 10, 2026
New (experimental): hour-by-hour Neural Net Predictions for Ponca and Boxley.
- A new experimental page predicts the next few hours at the two headwater gauges. At
/buffalo/neural-net-predictions/, a small neural network — trained on 12 years of storms — watches the last 48 hours of radar rainfall over every square kilometer above the Ponca and Boxley gauges, plus the gauges' own recent behavior, and projects the level hour by hour: six hours ahead at Ponca, four at Boxley (its smaller basin gives the radar less lead time). Each hour shows the model's best estimate and a likely range, with high-water and flood coloring.
- These projections use only rain that has already fallen — never a forecast. If more rain comes, the numbers become a floor, not a ceiling; the page refreshes every hour as new radar arrives.
- It's experimental and deliberately tucked away (not yet linked from the main Buffalo page). Trust the main dashboard and your own judgment first, and expect this page to evolve.
Not user-facing, recorded for completeness: the model (pixel-v5) trains on a dedicated GPU machine but serves from the existing prediction Pi as a 530 KB pure-numpy engine, parity-tested against the trained network; hourly cron, isolated from all other site systems; honest "stale/degraded/offline" banners when inputs are missing.
Version 2026.7.9.1 — July 9, 2026
New: Buffalo TV — a full-screen, self-updating view of the whole river.
- A "Buffalo TV" button is now on the Buffalo River page. Up top, alongside the "Understand the math" links, there's a new Buffalo TV button. It opens a full-screen, self-refreshing display of all seven gauges and current watershed conditions at a glance — the mainstem colored by whether each stretch is floatable right now, live rainfall shaded on the map, and any inbound rise or flood forecasts. It's built to live on a wall-mounted screen at an outfitter or visitor center, or to run full-screen on your own monitor while you plan a trip; leave it up and it keeps itself current. (Experimental — always verify conditions before you go, and it's not an official NPS or USGS product.)
Not user-facing, recorded for completeness: the long-soft-launched /buffalo/tv/ kiosk page (disk-read, no restart) is now linked from the Buffalo header via an inline-styled pill button in dashboard.py. The /buffalo/simulator/ page was renamed "Buffalo National River Simulator" (test-tube icon removed, EXPERIMENTAL WHAT-IF TOOL banner kept) but remains unlinked and noindex — staged for a later launch.
Version 2026.7.7.4 — July 7, 2026
Ground-dryness alerting extended to 13 more drainages.
- The neighboring drainages now borrow a calibrated gauge for their dryness read. This morning's update gave Upper Buffalo, Richland Main, and Upper Cossatot thresholds that adjust to how dry the ground is, read from their own gauges. Tonight the surrounding drainages join in: the Buffalo-family creeks (Adkins, Boen Gulf, Beech), the Little Buffalo family (EFLB, West Fork Shop, Thomas), the Kings family (Upper Kings, Osage), the Richland tributaries (Falling Water, Big and Long Devils, Bobtail), and Baker Creek each read the nearest calibrated gauge — Boxley, Richland, or the Cossatot — as a regional ground-moisture indicator, with the threshold scaled up for each creek's smaller, flashier watershed. The Detail column shows the reasoning, e.g. "43% of 1.90" trigger (Cossatot DAMP × 2.0)".
- These are watches, not forecasts. Unlike the three gauge-calibrated drainages, these thresholds rest on a neighboring gauge plus local paddling knowledge — the reference gauge can be dry while a storm parks on one hollow. Treat them as a smarter heads-up, and expect the multipliers to get tuned as real storms test them. (Experimental.)
- Spirits Creek, the Big Piney family, and the Kiamichi creeks keep the previous fixed-threshold system — no suitable reference gauge nearby yet.
Not user-facing, recorded for completeness: q_ref: {parent, multiplier} blocks on 13 drainages in drainages.yaml; assembler scales the parent's Q-bucket threshold and keeps each drainage's own rolling rain window; per-cycle legacy fallback; Tier-2 decisions dual-logged with tier: 2; multipliers = Dave's May 2026 numbers for six creeks, pattern-derived (confirmed) for seven.
Version 2026.7.7.3 — July 7, 2026
Watershed alerts now know how dry the ground is — using the creek itself as the moisture gauge.
- Smarter alert thresholds for Upper Buffalo, Richland Main, and Upper Cossatot. How much rain it takes to bring a creek up depends enormously on the season: 1.5" on saturated February ground can mean a great run, while the same storm in July soaks into dry dirt and the creek barely moves. The Watersheds page previously used one fixed rain trigger with a small seasonal adjustment — which is why last night's 1.5" storm fired a Richland watch even though the creek was dead low and barely budged. These three drainages now read their own USGS gauge as a ground-dryness indicator: a dead-low Richland demands over 2" of storm rain before alerting, while an already-flowing Richland alerts on as little as 0.55". The thresholds were calibrated against 11+ years of radar rainfall and gauge history, and in a wet-season backtest this method caught 11 of 11 boatable rises (the old scheme caught 6) with roughly 3–4 hours more warning. (Experimental — first dry-season deployment; thresholds will be reviewed as real storms come through.)
- The Trigger column now shows what the current threshold is and why. For these three drainages you'll see the rain amount adjust as the creek's flow changes, marked "/ event" (rain accumulated over the current storm) instead of a fixed time window, with the reasoning spelled out in the Detail column — e.g. "No active rain event — Q 25 cfs (DRY: needs 1.05")".
- The other drainages (no usable gauge on the creek) keep the existing fixed-threshold system for now.
Not user-facing, recorded for completeness: v2 Q-bucket engine (qbucket_shadow.py, hardcoded calibration) promoted from a clean 12-day shadow run to authoritative in assemble_and_push.py with per-cycle legacy fallback; legacy status still computed and dual-logged to logs/qbucket_shadow.log; suppress-while-already-running maps to NO_ALERT; post-promotion checkup reminder 2026-07-25; research in research/q_bucket_triggers/ and research/q_bucket_recent_backtest/.
Version 2026.7.5.5 — July 5, 2026
Smarter Buffalo flood warnings: creek-driven floods, cloudburst detection, and all-time-record context.
- Warnings when the creeks flood first. Some Buffalo floods start on Richland Creek or Bear Creek — one hollow gets hammered while the rest of the basin stays dry, and the main river doesn't react until the creek's water arrives. The wave forecaster now launches an arrival forecast the moment either creek starts a major rise. In a March 2024 storm of exactly this shape, the Buffalo at St. Joe jumped from 326 to 14,500 cfs nine hours after Richland crested — warnings for that kind of event now begin while the creek is still rising. (Experimental.)
- Cloudburst detector. Gauge cards now flag "Locally Intense Rain" when a single sub-watershed takes 2"+ in 24 hours — the concentrated storms that basin-wide rainfall averages hide. In twelve years of radar data, about four out of five of these flags were followed by a real rise downstream within 36 hours. (Experimental.)
- All-time-record context on wave forecasts. When a big wave is inbound, the forecast now shows how the predicted crest compares to the all-time record at that gauge (records reach back to 1915 at St. Joe), with a clear callout if a wave approaches record territory. Forecast bands also no longer display above just-past-record levels.
Not user-facing, recorded for completeness: trib-launch reach models (Richland→St. Joe, Bear Creek→Harriet; LOYO-banded, fit on 268/300 paired events) added to the wave router as pure data — no router code changed; per-zone max-HUC12 rainfall stats added to the data feed; the Buffalo Simulator engine/bundle updated in lockstep (golden gate 1,800/1,800); research in research/buffalo_sim/envelope/ and research/trib_launch_waves/.
Version 2026.7.2.10 — July 2, 2026
Buffalo page: recession countdowns now know what season it is.
- Seasonal settle levels. The level a river falls toward swings enormously with the calendar — St. Joe typically holds above ~600 cfs in March but drops to ~50 by September. Countdowns used to chase one flat year-round floor, promising spring drops the river wouldn't deliver (in an 11½-year backtest, targets below the month's typical floor were reached within 10 days only 5–16% of the time). Countdowns to below-the-floor levels now say "typically settles near ~X cfs this time of year" instead of a number that won't come true.
- And an honest null result, published on the math page: the paddler folklore that water drops faster in summer is TRUE — we measured 1.2–2.3× faster in matched conditions (the trees really are drinking the river). But adding a calendar rule to the countdowns made them worse in backtest, because the system already measures each recession's actual speed live and adapts. The folklore is confirmed; the fix was already built in.
Not user-facing, recorded for completeness: monthly p15 floors (circularly smoothed) for Pruitt/St. Joe/Harriet in config recession_floor_monthly (Boxley/Ponca floors never bind their thresholds); seasonal rate multipliers fit and backtested in two variants (result-scaling and anchor-recentering), both regressed vs the live k_now anchor — not shipped; full analysis in research/buffalo_replay outputs/RECESSION_SEASONALITY_REPORT.md.
Version 2026.7.2.9 — July 2, 2026
Buffalo page: the Ponca AI card explained, and recession math corrections.
- New math page: "The Ponca AI card, explained" (in the Understand-the-math menu). The full anatomy of the orange card: the 394-storm analog library and its four outcome classes, the three honesty rules on the matching, the physics floors that override the statistics, and — most importantly — what the AI actually does: every number is computed by transparent math first; the language model only writes the sentences and cannot change the call.
- Recession page corrected. The per-gauge fall-rate table had a lookup bug: Ponca showed the same rate at high and low water, and St. Joe/Harriet showed nothing. Real numbers now — e.g., Boxley sheds ~89% of its flow per day from high water but only ~40% near floatable; St. Joe and Harriet now display their rates plus the baseflow they settle toward.
Not user-facing, recorded for completeness: recession page now imports the deployed curve module directly for lookups (raw-Q bands, edge fallback) so it cannot mis-render again. Separately measured this session, pending review before any model change: paddler folklore confirmed — summer recessions run 1.2–2.3× faster than winter in matched flow bands (evapotranspiration), and lower-gauge baseflow swings seasonally (~600 cfs March vs ~50 September at St. Joe against a flat 76 in config); plan at RECESSION_SEASONALITY_PLAN.md.
Version 2026.7.2.8 — July 2, 2026
Buffalo page: the full math behind every prediction, published.
- A new "Understand the math" menu at the top of the Buffalo dashboard links five explanation pages. Nothing is hidden — actual equations, coefficients, thresholds, and measured hit rates, because transparency on public safety builds trust:
- Horizons — how far ahead the system can see at each gauge (shipped earlier today).
- Wave tracker — the real per-reach equations that turn an upstream crest into a downstream forecast, travel-time tables by wave size, how watches/warnings are defined, and how the tracker corrects itself as a wave passes each gauge.
- Local rain — the exact trigger thresholds per gauge, how the river's own baseflow tunes them for ground wetness (with the fizzle-rate proof), and the four rules that make a prediction stand down.
- Recession — the measured decay curves behind "time to floatable" (how fast each gauge actually falls at every flow level) and when a countdown pauses.
- Confidence — every high/medium/low label is a measured historical hit rate, and this page shows the actual table, sample sizes included.
Not user-facing, recorded for completeness: pages generated from deployed model artifacts by build_math_pages.py (they cannot drift from the running code; rebuild + scp on model changes, no restart); confidence labels read from the live CONFIDENCE_TABLE; horizons page corrected same-day to baseline v6 + baseflow ground-wetness copy after an accuracy audit.
Version 2026.7.2.7 — July 2, 2026
New page: "How far ahead can this system see?" — the honest math behind Buffalo predictions.
- A new page at /buffalo/horizons/ (linked from the top of the Buffalo dashboard) lays out, gauge by gauge, what warning time this system can honestly give and how often it's right — measured by replaying our prediction code against 11½ years of river history. Boxley and Ponca are rain-driven gauges: expect roughly 5 hours from the flood-making storm's first signal, physics-capped. Pruitt, St. Joe, and Harriet ride the wave trackers: typically 6–10 hours of watch before high water, with 91–98% of floods caught.
- It also says what we can miss — each gauge card names its blind spot plainly, because knowing the limits of a forecast is part of the forecast.
- The page includes the river-chain travel times (how long a surge takes from each gauge to the next) and a plain-English explanation of what a watch and a warning actually mean here.
Not user-facing, recorded for completeness: page content generated from the replay baseline artifacts by build_horizon_page.py (regenerated on model changes; the route serves a static fragment — scp updates, no restart); buffalo_horizons.json published alongside as machine-readable model facts (groundwork for a future what-if simulator). Lead metrics use causal-run semantics (the unbroken prediction streak into the crossing) after operator review caught the prior-pulse inflation in the draft.
Version 2026.7.2.6 — July 2, 2026
Buffalo page: the river itself now tells the system how wet the ground is.
- Ground wetness is now read from the river, not just the rain gauge. Rise predictions scale their triggers by how primed the ground is. That used to come only from the past week's rainfall — but a week of no rain after a soaking month leaves the ground far wetter than the rain total suggests. The system now reads each gauge's pre-storm baseflow against 11½ years of history for that month: high baseflow for the season means saturated ground and a hair-trigger river. In replay this caught 5–9% more real rises with no extra false alarms — the storms the old "dry" label caused it to shrug off.
- You'll see it in the prediction reasons: "on wet ground (river at 92nd pct)" instead of a rain-total figure, when the river data supports it.
Not user-facing, recorded for completeness: soil-moisture ladder rung 1 (pre-storm baseflow percentile, month-conditioned, 6–12h lagged against self-wetting; rain-antecedent remains the fallback); NLDAS-2/SPoRT-LIS rungs skipped (no data access, and the river signal may well be the better sensor anyway); replay baseline promoted to v6.
Version 2026.7.2.5 — July 2, 2026
Buffalo page: recession countdowns stay on screen unless a rise is genuinely imminent.
- "Time to floatable / too low" countdowns no longer vanish because of far-away predictions. The countdowns used to hide whenever any rise was predicted anywhere upstream — even a minor blip two gauges up with the water still 14 hours away. In an 11½-year replay, 97–99% of those blackouts saw no rise within six hours. Countdowns now pause only when a rise is expected at your gauge soon: local rain building there, or a tracked wave arriving within about 3 hours — and when a wave is further out, you'll see both the countdown and the incoming-wave forecast together, which is the honest picture.
Not user-facing, recorded for completeness: post-audit follow-on — arrival-aware suppression gate (rise_imminent_at) replacing has_upstream_rise; measured on the full replay before deployment (~104 restored countdown-hours/yr, bad-shows +2/yr all during hours the system had no signal anyway).
Version 2026.7.2.4 — July 2, 2026
Buffalo page: steadier behavior during radar and data outages.
- Predictions can no longer get stuck during a radar outage. If the rainfall feed goes down mid-storm, frozen rain totals used to keep "rise expected" cards alive indefinitely; the system now recognizes stale data and lets predictions expire on their normal schedule.
- Missed radar hours heal themselves. If the rainfall reader misses an hour, it now recovers it automatically from the archive instead of losing it forever — so rain totals and ground-wetness readings stay accurate through hiccups.
- The Ponca analysis card stands down gracefully when its data source goes stale, instead of freezing its last message on the page.
Not user-facing, recorded for completeness: fix package F6 (audit finding #14 + #15a) — wall-clock QPE windows, snapshot backfill, staleness surfaced to consumers, MRMS sentinel counter, clock-skew guards, fsync on state files, flood-threshold boundary classification; 19 synthetic outage/gap/skew tests. This completes the 2026-07-01 fresh-eyes audit: all six fix packages (F1–F6) are now live.
Version 2026.7.2.3 — July 2, 2026
Buffalo page: smarter local rain triggers, honest timing windows, and Bear Creek finally counts.
- Rise predictions now key off rain intensity, not stale accumulation. The old trigger couldn't tell 2 inches in 3 hours from 2 inches dribbled over a day. Predictions now fire on real short-burst intensity (with a separate very-heavy-soaker trigger on the lower river), which roughly doubles how often a fired prediction verifies. Thresholds were tuned per gauge against 11½ years of history.
- Each gauge now predicts on its own honest clock. Timing windows are measured from how each basin actually responds — Boxley peaks 5–10 hours after rain starts, Harriet's middle-basin storms take 9–20 — and the window is set when the event starts and counts down, instead of resetting every 15 minutes.
- Rain over Bear Creek now counts toward Harriet. Over a third of the drainage between Grinder Ferry and Harriet drains through Bear Creek, and its rain previously fed nothing. It's now part of Harriet's local rise and flood-risk signal, and the Richland/Bear Creek signal gauges show real trend arrows instead of "—".
- The wave tracker and local predictions now divide the work. When a wave is already tracked inbound (the usual case on the lower river), the gauge card shows the wave forecast — arrival window and size — instead of a vaguer local-rain guess. Local rain cards now appear mainly for the storms the wave tracker can't see coming: rain concentrated over the middle basin. Their confidence labels are freshly calibrated, so a "medium" or "low" tag means exactly that.
Not user-facing, recorded for completeness: fix package F5 complete (findings #5, #8, #10 from the 2026-07-01 audit) — per-gauge intensity/accumulation trigger ladders (leave-one-year-out fit), measured first-fire→peak windows frozen at event anchor, tributary-zone blending, router-jurisdiction handoff, scale-aware fizzle cuts, recalibrated confidence table; five candidate iterations against the replay gate before deployment; replay baseline promoted to v5. Remaining audit package: F6 (robustness).
Version 2026.7.2.2 — July 2, 2026
Buffalo page: rise predictions now stand down once the rise they predicted has happened.
- No more phantom "rise coming" cards on a falling river. Rise predictions used to run on a rain timer — after a storm, a card could keep saying "rise expected in 2–6 hours" for hours after the river had already crested and was clearly dropping. Predictions now watch the gauge itself: once the predicted rise has arrived and the river turns down, the card clears. If fresh rain keeps falling, the system can still call a genuine second rise — but it has to be a real new rise from the river's current level, not an echo of the last one.
- What you'll notice: rise prediction cards disappear sooner after a storm peaks, and the ones you do see are right more often — in an 11½-year replay, predictions verified better at every gauge, with Boxley improving the most. No flood detection was lost anywhere.
- Recession countdowns come back sooner. Because stale predictions no longer linger, the "time to floatable/too low" countdowns — which pause while a rise is expected — resume hours earlier after a storm.
Not user-facing, recorded for completeness: fix package F5.1 from the 2026-07-01 audit (prediction-consumed event state in local_rain_state; post-peak prediction-hours −88% in replay, flood POD byte-identical); gated on a full 2014–2026 replay before deployment; replay baseline promoted to v4. F5.2–.4 (intensity/accumulation split, measured windows, tributary zones) pending.
Version 2026.7.2.1 — July 2, 2026
Buffalo page: the river now tracks flood waves as they travel downstream — with predicted size and arrival time.
- New "Incoming Wave" section on the Buffalo gauge cards (experimental). When a surge crests at an upstream gauge, downstream cards now show what's coming: a predicted flow range (for example "~8,400–11,900 cfs") and an arrival window in local time. The forecast appears while the upstream gauge is still rising (marked "still growing"), firms up when it crests, and stays on screen until the wave actually arrives — it no longer vanishes the moment the upstream gauge starts dropping, which used to happen hours before the water reached you.
- Smarter wave alerts. The alert banner now fires based on how big the wave will be at the gauge you care about, not whether some upstream gauge happened to cross its own flood line. A "watch" means the predicted crest reaches at least 60% of that gauge's flood threshold; a "warning" means even the conservative end of the prediction floods. In an 11½-year replay, this caught 96–98% of St. Joe and Harriet floods with a typical 10 hours of notice — roughly double the old lead time — and near-miss waves that used to be invisible now get a watch.
- Lower Buffalo paddlers and NPS gravel-bar users benefit most: Harriet, which used to get under 5 hours of warning (and sometimes none at all), now typically sees a wave coming 10+ hours out, tracked all the way down the chain from Ponca and Pruitt.
Not user-facing, recorded for completeness: fix package F4 from the 2026-07-01 audit — per-reach crest and travel-time models fit on an 11.6-year event library (leave-one-year-out validated), a stateful wave tracker replacing the trend-gated single-hop mechanism, target-anchored alert tiers tuned on the full replay (all deploy gates passed vs baseline v3), Richland/Bear Creek gauges now feed the lower-reach predictions. Packages F5–F6 pending.
Version 2026.7.1.9 — July 1, 2026
Buffalo page: rise predictions now judge storms against how wet the ground was before the rain.
- "Antecedent" now means what it says. The Buffalo page's rise predictions and flood-risk levels scale their rain thresholds by ground wetness — dry ground soaks up rain that saturated ground sheds. A long-standing bug let the storm being measured count toward its own "prior wetness," so the system almost always judged storms as falling on wet ground and could lower its own bar mid-storm. It now measures the week of rain before the last 24 hours.
- What you'll notice: somewhat fewer rise predictions overall — mostly the ones that didn't pan out. In an 11½-year replay, predictions that a rise was coming verified more often at every gauge, and "large rise" calls at Harriet that actually approached flood went from 79% to 90%. Dry-ground conditions (common in late summer) are now genuinely recognized for the first time, so a modest storm after a dry stretch is less likely to trigger a rise call the river shrugs off.
- The heavy-rain flood-risk banner gets rarer and more meaningful — roughly 3 episodes a year instead of 9, because it no longer fires on storms that made their own ground look wet.
Not user-facing, recorded for completeness: fix package F3 from the 2026-07-01 audit (pre-storm antecedent, precip_7day_total_in preserved for the propagation model's fitted definition, all 8 zones exported incl. Richland); gated on a full 2014–2026 replay before deployment (St. Joe flood detection unchanged, no Boxley/Ponca regression); replay-rig baseline promoted to v3. Packages F4–F6 pending.
Version 2026.7.1.8 — July 1, 2026
Buffalo page: the downstream "river rise" card now stays up until the wave actually arrives — and won't inflate its numbers late in an event.
- The "AI River Rise & Propagation" card no longer disappears before the surge reaches St. Joe / Grinder Ferry. It used to stand down once rain ended — often 6+ hours before the wave arrived downstream. It now stays up until the water has had time to reach Grinder Ferry, through the whole arrival window.
- Its downstream crest estimates hold steady through an event instead of drifting. The card now locks in the storm's rainfall picture and the pre-storm ground wetness when an event starts, the same way its prediction model was calibrated. Before, re-reading decaying rainfall mid-event could quietly inflate the St. Joe estimate — on a replay of the November 2024 record flood, the old readings would have called ~164,000 cfs at Grinder Ferry when ~71,000 actually came; the corrected inputs call ~60,000–85,000. Experimental, as always.
- No more multi-thousand-cfs "crests" from drizzle. On ordinary wet-season days with a bit of rain and normal flows, the card now says plainly that downstream will see little change, instead of quoting a precise crest from a model being used outside its comfort zone.
- Honest wording when the upper river has crested: the card now says Boxley "crested and is now dropping — its surge is already on the way downstream" instead of claiming it's still climbing.
Not user-facing, recorded for completeness: this is fix package F2 from the 2026-07-01 fresh-eyes audit (event-frozen SJ_MODEL features + pre-storm antecedent + wave-in-transit hold + fit-domain gate); validated over 173 historical Ponca events (antecedent error vs fit: +1.63″ → +0.17″ median on big events; late-event regime flips 38% → 3.5%) plus a live end-to-end synthetic flood test. Packages F3–F6 pending.
Version 2026.7.1.7 — July 1, 2026
Buffalo page: an earlier flood-risk banner, and confidence labels you can take at face value.
- A new heavy-rain flood-risk banner (amber = watch, red = warning) now appears on the Buffalo page when a lot of rain has fallen on a gauge's watershed but no gauge has flooded yet — the early-warning step before the red "gauge is in Flood" banners you've seen during events. A display bug had kept this banner from ever appearing; expect it a handful of times a year, during real storms. Experimental: it's rain-based, so it will sometimes warn on a storm the river ends up shrugging off.
- The small per-gauge "Flood Risk: WATCH" note now shows in amber (and WARNING in red) as intended — WATCH had been rendering in red.
- Rise-prediction confidence labels are now honest. The "(high/medium/low)" next to each Buffalo rise prediction is now calibrated against 11½ years of actual gauge outcomes — "high" now means predictions like this one verified at least ~65% of the time historically. The old rule could label weak signals "high"; some predictions that used to read "high confidence" will now honestly read "low."
Not user-facing, recorded for completeness: rise predictions are now graded against the timing window of the driver that set the category (grader + replay-rig change, removes ~6 pts of unearned timing credit from the accuracy baseline); antecedent tier boundaries are read from config instead of being hardcoded; this ships fix package F1 from the 2026-07-01 fresh-eyes audit of the Buffalo prediction chain (research/buffalo_replay/outputs/FRESH_EYES_AUDIT.md) — packages F2–F6 pending.
Version 2026.7.1.4 — July 1, 2026
The site is faster, and the Illinois River forecast card is now complete.
- Pages load noticeably faster. The home page, changelog, and guide now respond several times quicker, and the site stays responsive when many people check it at once during a rain event.
- The Illinois River page's Empirical Forecast card now shows the full detail table — per-tier chance, rainfall band, and historical rain percentiles — like every other basin. (Its "rise to floatable" row honestly reads insufficient data: the Illinois at Hwy 16 has dropped below 150 cfs only a handful of times in 12 years.) Experimental, as always.
- The Illinois empirical forecast now records its predictions for grading. Its accuracy will start appearing on the Scorecard page as forecasts settle over the coming days — closing a gap where the newest basin was the only one not checking its own work.
Not user-facing, recorded for completeness: USGS gauge fetching was consolidated into batched requests (~93% fewer API calls per 15-min cycle, with per-site retry fallback preserved); the per-basin script clones for Mulberry, Big Piney, Cossatot, and Richland now run through one shared library (/home/dave/basinlib/ shims — equivalence-verified before deploy) with a retry hardening all four inherit; a git baseline of both Pis' code/config now lives on the operator harness (creekintelligence/mirrors/); the dashboard gained an mtime-keyed file cache, explicit Cache-Control headers, and a threaded gunicorn worker (-w 1 --threads 4). See ARCHITECTURE 2026.7.1.4.
Version 2026.7.1.2 — July 1, 2026
The Buffalo River Watershed Study has wrapped up after 122 days — a month longer than planned — and its page is now a finished, browsable archive.
- The Buffalo Study page (linked from the tools menu) now opens with a short "Study concluded" summary of what the 122-day experiment learned: how the same inch of rain moves the river much more when the ground is already wet, how the upstream-to-downstream cascade is timed (about 10 hours St. Joe→Harriet), each gauge's personality, and the record June 22 flash flood that anchored the extreme end. (The study ran an unusually wet June a full month past its planned 90 days — that wet stretch delivered its most valuable data.)
- The daily archive and the final hypothesis document are still fully browsable — now clearly marked as a frozen record rather than a live-updating one.
- Nothing you rely on for live conditions changes. The study's findings were folded into the everyday forecasts (watershed alerts, rise likelihoods, downstream propagation), and those keep running exactly as before.
Not user-facing, recorded for completeness: the study's nightly Claude Opus narrative stream was retired on 2026-07-01 (cron study_daily_analysis.py --no-opus), ending the project's only per-night external API cost; the local qwen3.6:27b stream continues a no-cost nightly pulse on ScriptPi for periodic hand-review. /study/ pages are frozen at Day 122 (June 30, 2026); the concluded banner + per-page frozen-archive notes are served from dashboard.py (study_index/study_hypothesis/study_daily). See ARCHITECTURE §4.7 + §16 item 16 (version 2026.7.1.2).
Version 2026.7.1.1 — July 1, 2026
Arkansas Creek Intelligence has its own web address now: arcreekintel.com.
- The site now lives at its own dedicated domain — arcreekintel.com. Every page is reachable there: the home creek list, the individual creek and watershed pages, the historical reports, and the scorecard — all of it, exactly as before, just at a shorter address that's all about the creeks.
- Your existing bookmarks still work. The old web address keeps pointing to the very same site during the transition, so nothing you've saved will break. Going forward, arcreekintel.com is the one to save and share.
Version 2026.6.30.1 — June 30, 2026
New page: the Illinois River and the Upper Illinois Water Trail.
- There's now a dedicated page for the Illinois River at /illinois/. It tracks the Hwy 16 gauge near Siloam Springs for the 15.5-mile Upper Illinois Water Trail (Chamber Springs → WOKA), with float levels — too low under 150 cfs, low-but-floatable 150–250, optimal 250–2,500, above-recommended over 2,500. (Those bands come from the local paddling community; the AGFC's official 200–1,000 cfs range is noted on the page too.)
- The Upper Illinois Water Trail also shows as a live creek card on the home page and the gauges list, with its current float level — so it's an official entry in the creek lineup, not just a standalone page.
- Upstream gauges give you a head start. Because the Illinois drains a big watershed, the page also watches the gauges on Osage Creek (at Elm Springs and Cave Springs) and at Savoy — they typically come up several hours before that water reaches the trail, so a bump upstream is your early warning. The downstream Watts, OK gauge (below WOKA) gets its own card and 24-hour graph as well.
- A watershed rainfall map and forecast. Rainfall is broken out sub-basin by sub-basin, above and below Hwy 16, each with a 24-hour rain-forecast column, plus an experimental "chance the river rises" forecast for the run. (Experimental.)
- A "⚠ heavy rain below Hwy 16" caution for the rare storm that soaks the lower river toward WOKA without showing up on the Hwy 16 decision gauge.
- A new historical page at /illinois/historical/ — how much of the year the river spends in each float range (it sits in "optimal" about 62% of the time), the biggest floods on record (April 2017 peaked at 148,750 cfs), and a catalog of all 350 storms that have pushed it past 1,000 cfs over the last 23 years, with the rainfall behind each recent one.
Not user-facing, recorded for completeness: new ScriptPi basin /home/dave/illinois/ (config + assembler + multi-gauge fetch / read-qpe / weather+QPF, cloned from Big Piney with Buffalo's signal_for multi-gauge pattern grafted on), cron every 15 min → illinois_output.json. Full history collected to the harness research/illinois_*: 8 USGS gauges' complete records + hourly MRMS per-HUC12 QPE 2014→present over the 21 HUC12s above Watts. Empirical illinois_below block trained on the above-Hwy-16 basin (calibrated_table.json + lookup index + empirical_config.yaml). Upstream→Hwy-16 propagation lags from storm-pulse cross-correlation (Savoy 6–8 h, Osage 8–13 h, Mud 11–16 h; Watts +3–5 h downstream). New DMZPi routes /illinois/ + /illinois/historical/. See ARCHITECTURE §4.9c (version 2026.6.30.1).
Version 2026.6.29.7 — June 29, 2026
The rainfall forecast on the five creek pages now gives you an honest percentage instead of a vague "rise likely."
- On the Cossatot, Richland, Hailstone, Mulberry, and Big Piney pages, the Empirical Forecast now reads like "~30% chance of rising to floatable in the next 6–12 hours" instead of a yes/no "WATCH — rise likely." The old version cried wolf: across 2,314 real rainstorms back to 2014, its "rise likely" was a false alarm 60% of the time. The new number is calibrated against what the creek actually did — and it holds up every year.
- When the new forecast says a number, the creek does it about that often:

- Why it's better: the old engine only knew how much rain usually fell before a creek came up — it never learned how often that same rain just soaked in and did nothing. The new one knows both, so a heavy rain that usually drains is reported as the modest chance it really is, not a false alarm. It also quietly leans a little higher when the ground is already saturated and lower in a dry spell. (Experimental — it's an honest probability, not a guarantee: a 30% morning can still come up, and an 80% one can still fizzle.)
Not user-facing, recorded for completeness: new calibrated engine (empirical_forecast/calibrated.py + calibrated_table.json — denominator-corrected P(rise | rain, soil) over the full 2014-2026 MRMS+USGS record) feeds the headline in empirical_predict.py, with a per-basin warm rolling factor off the local settled archives. Boxley (no page) unchanged; empirical stays off Facebook per standing choice (Watersheds only). Full research + validation in the harness research/empirical_recalibration/ (REPORT.md). ARCHITECTURE §4.13/§16-11 updated.
Version 2026.6.29.6 — June 29, 2026
Added a privacy policy page.
- There's now a privacy policy at /privacy. The short version: you can use the site without an account or handing us any personal information; we run no ads, no analytics of our own, and never sell anyone's data. It's there so our (minimal) information practices are documented in one place — and so the Facebook page can link to a real policy.
Version 2026.6.29.5 — June 29, 2026
Our forecast "report card" got clearer and more honest, and a couple of predictions were tuned based on what it's been telling us.
- The Scorecard page (
/scorecard) is easier to read and harder to mislead. The recession-countdown table now hides the dozens of barely-started rows so you only see the gauges with enough data to mean something, and each predictor's "bias" now shows its average and its typical (median) miss side by side — so a single freak event (like the June 22 flash flood) no longer makes an otherwise-solid predictor look broken. (Experimental, self-grading page.)
- The Buffalo "AI Rainfall Event Analysis" card now stands down as the river falls. Once the Ponca gauge is clearly past its crest and dropping, the card eases its flood-risk number and wording back down instead of staying pinned near "Flood" through the recession — so it stops over-warning about water that's already on its way out. (Experimental.)
- Big Piney's "Above Longpool" run no longer shows a "time to too low" countdown that never panned out. That upper Class II+ stretch settles so close to its too-low mark that the countdown to it was unreliable, so we've switched it off for that section — the "low but floatable" and "optimal" countdowns there are unchanged.
Not user-facing, recorded for completeness: this came out of the weekly scorecard / Signal-digest review. score_all.py now emits a per-predictor median error, min-n-gates the recession rows (MIN_RECESSION_GRADED), and adds a recession timing-bias flag; ponca_analog.py decays the dominant-driver flood-floor toward the raw k-NN once Ponca is past crest (raw values still logged for backtest); big_piney_assemble.py had recession_baseline support restored to its ported compute_recession (the port had dropped it) and above_longpool was given recession_baseline: 2.0. The empirical engine's over-warning (also visible on the scorecard) is a deeper rebuild blocked on an offline tool — tracked, not yet changed. See ARCHITECTURE §4.13 / §16-11 (version 2026.6.29.5).
Version 2026.6.29.4 — June 29, 2026
Creek Intelligence is now on Facebook — follow Arkansas Creek Intelligence for automatic creek alerts and a weekend forecast.
- We launched an Arkansas Creek Intelligence Facebook page, and it updates itself. When watersheds start coming up after rain, an alert bot posts a watershed alert to the page automatically — every creek crossing into Watch, Warning, or Flood that hour bundled into a single post, so you get one clean heads-up instead of a stream of notifications. It only speaks up when something new is developing or getting worse, so the page stays quiet between events. (Experimental — our Signal alert group is still the fastest, real-time channel.)
- New Weekend Creek Forecast, every Friday at 3 PM. A single post for weekend paddlers: the creeks running optimal right now that should still be runnable through the weekend (based on how fast each one is dropping), a short list of marginal "catch-it-early" runs, and a heads-up on any rain headed our way in the next 24 hours. (Experimental.)
- Why a Facebook page: it puts the same intelligence the dashboard already computes in front of paddlers where they already are — the alerts come to you, no page to keep checking.
Not user-facing, recorded for completeness: new creek-social toolset on the harness server — creek_social.py mirrors the watershed alerts to Facebook (fully decoupled from ScriptPi: it reads the same predictor_output.json the dashboard already produces, replays the alert engine's new/escalation logic against its own state, and posts via the Facebook Graph API), weekend_forecast.py (Friday cron; recession-based "holds through the weekend" filter), and announce.py (auto-posts a short feature announcement to the page whenever we ship a user-facing change — a new shipping convention now in CLAUDE.md §8 + ARCHITECTURE §4.2). Posts are deliberately link-free and bot-signed.
Version 2026.6.29.2 — June 29, 2026
The Buffalo Study pages now correctly report the June 22 flash flood — it had been mistakenly logged as a "data gap"
- On the experimental Buffalo Study pages (
/study/), June 22 is now described accurately as the biggest event of the study so far: a dry-ground flash flood that set records up in the headwaters — Boxley, Ponca, and Pruitt all hit study-record levels — then faded to just below flood stage by the time it reached St. Joe / Grinder Ferry. The running write-up had previously called this stretch a "data gap" and guessed it was only a moderate event, because the night-of analysis that captured the flood never made it into the rolling summary. Both the gap and the wrong guess are now corrected, with the real numbers.
- Why it matters: it's the clearest example yet of a pattern these pages track — an upstream-only flood loses most of its punch crossing the dry middle of the watershed before it reaches the lower river (the same idea behind today's earlier propagation-card fix). Still experimental research pages, not the live gauges.
Not user-facing, recorded for completeness: two fixes to the nightly Buffalo Study analyzer (study_daily_analysis.py) plus a data backfill. (1) The transfer-ratio validator was QPE-gated — its 2,000 cfs/in cap had unconditionally false-rejected the model's correct 06-22 headwater numbers (Ponca 2,756 cfs/in matched the deterministic truth card), silently dropping that night's calibration; the cap now applies only on low-rainfall days, where the frame-override artifact it guards against actually occurs. (2) The Opus output cap was raised 40000→56000 after the daily+thinking+hypothesis rewrite pinned at the cap for five straight nights (06-21..06-25, incl. the flood), freezing the rolling hypothesis. (3) The 06-22 study-record flood was backfilled into both the qwen knowledge.md calibration tables and the Opus hypothesis.md (now live on /study/), from the truth card + the existing night-of Opus daily; logged in analysis/curation_audit.md. The qwen-only (--no-opus) cutover stays deferred — even post-fix, qwen still drops a gauge on the complex flood night. See ARCHITECTURE §4.7 + §16 item 16.
Version 2026.6.29.1 — June 29, 2026
More accurate downstream forecasts: the St. Joe / Grinder Ferry propagation card no longer overshoots when the rain falls up high
- The Buffalo page's "River Rise & Propagation Analysis" card now looks at where the rain actually fell before estimating how big St. Joe / Grinder Ferry will get. During the big June 22 flood the rain landed almost entirely up in the headwaters (around Boxley and Ponca), and the flood wave simply rolled downstream and faded — St. Joe ended up cresting right around its flood stage. The old card had assumed the rain fell across the whole watershed and forecast St. Joe two-to-three times too high — it read like a catastrophic flood when the river actually came up to about flood stage. That's the gap this release closes.
- What you'll notice: when a rise is concentrated up high and the middle of the watershed stays dry, the card now expects St. Joe to come up roughly in line with Ponca — the wave attenuates as it travels — instead of several times higher. When rain falls across the whole basin it still calls the big amplified crest, because that's when St. Joe genuinely does run several times above Ponca. Pruitt's forecast is unchanged (it was already accurate). Arrival-time estimates are a little tighter too.
- Still experimental — it's a downstream estimate given as a range, meant to help with gravel-bar timing; always check it against the live gauges.
Not user-facing, recorded for completeness: the St. Joe magnitude in ponca_analog.py build_propagation() was a fixed basin-wide amplification (≈2.9× Ponca) applied unconditionally; on an upper-concentrated event the wave routes through the dry 1,342 km² intervening basin and attenuates (~0.9×), so the constant over-predicted 2–4× (6-22: card 19–33k, actual 7.9k). Recalibrated against 152 historical Ponca events (2014–2026) from the local buffalo_huc_qpe + buffalo_gauges archives (research/buffalo_propagation_calibration/, with REPORT.md + figures). St. Joe is now a rainfall-distribution + antecedent conditioned log-linear model (SJ_MODEL) reading the upper-vs-intervening qpe_24hr ratio and the intervening 7-day antecedent — all already in buffalo_output.json — with a crest band floored at the routed wave (0.85×) and the reach's current flow. Leave-one-out CV cut the upper-concentrated bias from +46% to ≈0; out-of-sample on 6-22 it predicts St. Joe 8.5k–11.9k–16.9k (lower bound ≈ actual). Pruitt kept at ~1.13× (rain-insensitive, 11-yr confirmed); peak-to-peak lags re-derived. ScriptPi-only — render_propagation_card() reads only the kept crest-band fields, so no DMZPi change. See ARCHITECTURE §4.6 + §16 item 15.
Version 2026.6.24.8 — June 24, 2026
New: a 24-hour gauge chart on the Cossatot, Richland, and Hailstone pages
- Each of those three pages now shows a 24-hour hydrograph right under the big gauge reading — a simple line chart of the gauge's last 24 hours, so you can see at a glance whether it spiked and how it's receding, the way you would on the USGS site.
- The runnable range is shaded right into the chart (red = too low, yellow = low but floatable, green = optimal, blue = above recommended), so you can watch the river climb up into the good range and fall back out of it — not just a bare line. The peak is labeled and "now" is marked with a dot.
- It's a lightweight built-in chart (no third-party widgets) and fills in live as new readings come in every 15 minutes. When a creek is sitting below every threshold, it just reads all-red ("too low") with the trend visible.
Not user-facing, recorded for completeness: each physics-basin assembler now emits gauge.readings_24h (a [[epoch, value], …] array of the last 24 h, ~96 pts at 15-min) via a build_readings_24h() helper — Cossatot/Richland stitch yesterday+today's daily height files to span midnight, Hailstone uses its already-rolling cfs buffer from creeks/gauge_data.json. dashboard.py gained render_hydrograph(), an inline-SVG renderer (tier bands + line + peak + "now" dot, fully self-contained, no chart library), placed under the gauge card on all three pages. Display-only, no DMZPi compute (same contract as the HTML tables and Leaflet maps). See ARCHITECTURE §5.4.
Version 2026.6.24.7 — June 24, 2026
The Scorecard now flags what needs attention — and reports its own drift weekly
- The Scorecard page now opens with a "Needs attention" box that calls out, in plain terms, where the predictors are drifting — e.g. "the rise-likelihood engine over-warns on Boxley / Mulberry / Richland," "the heaviest-rain band is less reliable than the band below it," "two crest predictors are off-calibration." The page tells you what to look at instead of making you hunt for it.
- A weekly summary now goes out automatically (behind the scenes, to the operator) so the system surfaces its own drift without anyone opening the page — closing the loop on self-checking without manual review.
Not user-facing, recorded for completeness: score_all.py now computes a flags array (min-n-guarded: empirical over-warn ≥70% no-rise at n≥20, band inversion above_p75 vs. p50_to_p75, physics within-±20% <50% at n≥15, Ponca override-vs-raw flood-Brier gap, recession never-reached ≥60% once graded) into scorecard.json; /scorecard renders the "Needs attention" box from it. New prediction_eval/scorecard_digest.py (weekly cron Mon 07:30 Central) reads the scorecard, formats a per-family accuracy snapshot + the flags, and sends a Signal DM via the existing signal_config.yaml (same signal-cli plumbing as scp_alert / data_age_alert). This is Phase 3 of the prediction-logging initiative — the self-checking / feedback half. The harness-consolidation refactor was deliberately deferred (a risky change to working cron code with no user benefit). See ARCHITECTURE §4.13 + §16-14 + §7.2.
Version 2026.6.24.6 — June 24, 2026
Scorecard now also tracks the downstream-propagation forecast (recording started)
- The Scorecard page has a new "Downstream Propagation" section. When Ponca rises, the system predicts whether — and how big — a bump reaches Pruitt and then St. Joe / Grinder Ferry, and how many hours after Ponca peaks. Those predictions are now recorded and will be graded: did a noticeable bump actually come, was it the right size, and did it arrive in the predicted window?
- Nothing to show yet — like the Buffalo rise engine, this only logs during an actual Ponca rise, and a downstream crest can take most of a day to play out, so the first graded results appear after the next event. With this, every live forecast the system makes is now on the self-grading loop.
- Experimental, like the rest of the Scorecard.
Not user-facing, recorded for completeness: ponca_analog.py now folds the per-reach propagation block (lag window, crest band, bump/no-bump flag, current flow) into each ponca_analog_history.jsonl record — the record step the predictions had been missing. New prediction_eval/propagation_eval.py settles each matured event two-stage: it finds Ponca's actual peak from buffalo_study/data/gauges, then for Pruitt (07055680) and St. Joe (07056000) grades the bump/no-bump call (crest ≥ 1.25× baseline & +50 cfs), the magnitude-band hit, and the timing window relative to Ponca's peak — using the cycle nearest Ponca's peak as the representative prediction. A --selftest (10 checks) validates the logic. score_all.py folds the settle in and adds a propagation block; /scorecard renders it; Coverage flips the propagation forecast to "graded." This completes Phase 2 of the prediction-logging initiative — every live predictor family is now recorded and graded. See ARCHITECTURE §4.13 + §16-14.
Version 2026.6.24.5 — June 24, 2026
The Ponca "AI Analysis" cards now read like a person — and the numbers make sense
- Rewrote both Ponca AI cards (the rainfall outlook and the downstream propagation card on the Buffalo page) to read in plain language for a casual paddler — and, more importantly, fixed the numbers.
- No more peaks below the current level. The rainfall card had been quoting a "typical peak" from past storms even when the river was already past it (e.g. "peak ~364 cfs" while sitting at 415). It now floors the peak by what the river has actually done — so it says things like "it's near the top for this storm, ~431, and shouldn't climb much more."
- Propagation numbers fixed. The downstream card was predicting crests below where Pruitt and St. Joe already sat (e.g. "St. Joe will crest at 948–1,638" while it was already at 1,820). It now models the added bump the Ponca rise puts on top of each downstream gauge's current flow — and when that bump is small (a light rain), it plainly says "no noticeable change downstream" instead of quoting a fake crest window.
- Plainer, less robotic. Out: "no deterministic watershed rainfall alerts or upstream triggers firing." In: a short, plain read of whether the creek's coming up, how high, and when. The safety behavior is unchanged — when heavy rain hits the feeder creeks or the upstream river floods, the card still leads with that and leans into the bigger outcome.
Not user-facing, recorded for completeness: in buffalo_dashboard/ponca_analog.py, both qwen system prompts were slimmed and rewritten for a casual-paddler audience (no jargon, fact-fed, model reasons rather than recites) with one hard rule — never state a peak/crest at or below the current flow. build_propagation() now models est_crest = reach_current + (Ponca rise above baseline) × amplification, floored at the reach's current value (was Ponca_flow × amp absolute, which fell below a downstream gauge's own flow whenever Ponca was small relative to that reach's drainage), plus a significant / added_bump_cfs gate so small pulses read "minor bump, no noticeable change." A qualitative 6-hour trend word (holding steady / creeping up / rising steadily / rising fast / easing down / dropping fast) keeps the model from overstating a slow creep. The deterministic override floors (upstream Boxley wave, heavy-rain, watershed-feeder alerts) are unchanged — only their narration. See ARCHITECTURE §4.6.
Version 2026.6.24.4 — June 24, 2026
Scorecard now tracks the Buffalo rise engine (recording started)
- The Scorecard page has a new "Buffalo Rise Engine" section. Behind the Buffalo dashboard, the system makes a per-gauge "a rise is coming" nowcast (slight / moderate / large, with a timing window) from local rain and upstream propagation. Until now those predictions were computed, shown indirectly, and then thrown away every 15 minutes — never checked. They're now recorded and will be graded on whether each gauge actually rose, and whether it rose within the predicted window.
- Nothing to show yet — the rise engine only fires during rain events, and it's dry right now, so the section reads "recording started, no events captured yet." The first results appear after the next storm. (This was the largest prediction surface that had been completely untracked.)
- Experimental, like the rest of the Scorecard.
Not user-facing, recorded for completeness: new prediction_eval/buffalo_predictions_archive.py (modeled on recession_archive.py) — --record on a 15-min cron (:09,:24,:39,:54, debounced) logs each active predictions[gauge].combined_category from buffalo_output.json (with timing window + current value) to buffalo_ledger/<date>.jsonl; --settle grades vs. the gauge's actual discharge rise over [generated_at, timing_high + 18 h] (rose if peak ≥ 1.25× start AND ≥ +30 cfs → records rise magnitude, hours-to-peak, timing-in-window; else no_rise / censored); --score aggregates. A --selftest (14 synthetic checks) validates the logic since production is dry. score_all.py folds the settle in and adds a buffalo_rise block; /scorecard renders it; Coverage splits Buffalo into "rise predictions: graded" and "flood_risk + propagation_alerts: not yet logged." Magnitude is captured per category (categories derive from rainfall, not gauge rise) so the rainfall→rise mapping calibrates over time. ARCHITECTURE §4.13 + §16-14 + §7.1.
Version 2026.6.24.3 — June 24, 2026
Scorecard now grades the Ponca "AI Rainfall Event Analysis"
- The Scorecard page now scores the Ponca AI Rainfall Event Analysis (the orange card on the Buffalo page). Every call it makes during a rain event is now recorded and later graded against what the Ponca gauge actually did — so you can see how often it gets the size of an event right. Until now it was logged but never scored.
- What it shows: how often the predicted class (Fizzle / Moderate / High / Flood) matched the actual crest (exactly, and within one level), how close the "typical peak" estimate landed, and how well-calibrated the flood-risk % has been. On the ~4 weeks logged so far it gets the class within one level 95% of the time, and its predicted peak range has captured the actual crest about half the time — though it tends to over-call Flood when the river actually tops out at High.
- An honest self-check it surfaced: during the long recession after a flood, the card's deterministic "flood floor" stays elevated (e.g. 85%) while the river is already falling — so on the logged sample the raw analog has been a bit better-calibrated for flood risk than the floored version. Catching things like that is the whole point of this page; it points at a future tuning. (Small sample — directional only.)
- Experimental, like the rest of the Scorecard; the numbers will firm up as more events are logged.
Not user-facing, recorded for completeness: new prediction_eval/ponca_analog_eval.py reads the producer's append-only ponca_analog_history.jsonl (read-only) and grades each armed call against Ponca 07055660's max discharge over a forward 36 h window (from buffalo_study/data/gauges), writing ponca_analog_settled.jsonl (idempotent, keyed by generated_at). It grades the k-NN class/peak (modal vs. actual class, median_peak_cfs MAE, p25–p75 band hit) and the flood Brier for flood_risk_pct (override-floored) vs. raw_flood_risk_pct (raw k-NN, on the common sample). score_all.py calls settle() inline (no extra cron) and adds a ponca_analog block to scorecard.json; the /scorecard route renders it; the Coverage table flips Ponca from "logged, not graded" → "graded." See ARCHITECTURE §4.13 + §16-14.
Version 2026.6.24.2 — June 24, 2026
Retired the experimental neural-net (LSTM) predictor
- The old LSTM forecast is fully retired. It was an experimental neural-net gauge predictor for Richland and Cossatot; its on-page cards were pulled back in April (too many false alarms), and it has now been fully decommissioned behind the scenes. Nothing paddler-facing changes — the physics and empirical forecasts you actually see are unaffected.
- On the new Scorecard page, the LSTM no longer appears in the "Coverage" list. It briefly showed there as "not yet graded" right after the page launched; since it is retired, it has been removed.
Not user-facing, recorded for completeness: the :14 nn_predict.py inference cron was commented out, the dead nn_prediction read/embed was removed from cossatot_assemble.py + richland_assemble.py (the key no longer appears in their output JSON), the stale "LSTM forecast" wording was dropped from the Richland page's social/meta description, and all LSTM-only artifacts (nn_predict.py, nn_predict_debug.py, nn_alert.py, nn_alert_state.json, nn_output.json, both model_epoch100.pt weight files) were moved to /home/dave/creeks/retired_lstm/ (reversible — a README there documents how to resurrect). The HUC12 masks under /home/dave/models/ were kept — they are read by the live Richland physics QPE reader and the map-polygon builder, not just the LSTM. Full record in ARCHITECTURE.md §4.10 + §14 Gotcha #10. Side benefit: one less hourly job on the 2 GB ScriptPi.
Version 2026.6.24.1 — June 24, 2026
New: a Prediction Scorecard page — see how accurate the site's forecasts have actually been
- There's a new "Scorecard" page (linked from the home page and the footer of every page) that grades the site's own forecasts against what the rivers actually did. Every prediction the system makes is now recorded and checked later — so you can see how much to trust it, instead of just taking its word.
- It covers three kinds of forecast: the crest predictors (Cossatot, Richland, Hailstone — how close the predicted peak was to the real one), the "likelihood of a rise" headlines (how often a forecast rise actually showed up vs. how often the rain fizzled), and the recession countdowns (how close the "time to drop to X" estimates land — these are brand-new, so that section is still filling in over the coming weeks).
- It's honest about its own blind spots. Right now the scorecard openly shows that the "likelihood of a rise" engine over-warns — on several rivers a forecast rise doesn't actually materialize most of the time. Showing that plainly is the point; it's the first step toward fixing it. A "Coverage" section at the bottom lists which forecasts are graded and which aren't graded yet.
- Experimental. This is a behind-the-scenes accountability tool we're now exposing publicly. The numbers will shift as more predictions settle, and some forecast types aren't graded yet.
Not user-facing, recorded for completeness: a new ScriptPi subsystem /home/dave/prediction_eval/ (score_all.py, cron 01:30 daily) reads the settled prediction archives for three already-self-recording families — physics (<basin>/predictions/), empirical (<gauge>_empirical_predictions/), and recession (recession_eval/ledger/) — and writes one display-ready scorecard.json, SCP'd to DMZPi. Physics is scored on peak MAE / bias / within-±20%; empirical on hit-rate (rose to ≥ called tier = verified + missed_higher) vs. no-rise rate (no_change), broken out by basin, confidence, and rainfall percentile band (which surfaces the known selection bias — the above_p75 band currently shows a higher no-rise rate than p50_to_p75); recession reuses recession_archive.score(). The DMZPi /scorecard route renders the JSON read-only (no compute on DMZPi; validated via the venv test_client before the gunicorn restart). This is Phase 1 of a system-wide "log every prediction, grade it later" initiative — audit + plan in creekintelligence/research/prediction_audit/. The still-ungraded predictors (Ponca analog, downstream propagation, the Buffalo per-gauge rise engine, the retired LSTM) are listed as not-yet-graded in the page's Coverage table and are Phase 2.
Version 2026.6.23.1 — June 23, 2026
Recession countdowns are now realistic — and honest about how sure we are
- The "Time to …" countdowns now reflect how rivers actually recede. They used to take the current (fast) drop and extrapolate it in a straight line, which badly over-shot — on the flashier gauges it could say a creek would hit "Too Low" in 2 days when history says it's more like a week-plus. Every gauge now uses a per-creek recession curve built from years of USGS history: the drop is fast up high and slows way down near baseflow, just like the real river, and the estimate is anchored to how fast the current recession is actually falling. For example, Pruitt's "Time to Too Low" went from ~2 days to ~7 days — much closer to reality. This is on every recession card now (Cossatot, Richland, Hailstone, the Buffalo mainstem gauges, and Mulberry / Big Piney's float sections). Still an estimate — it assumes no new rain.
- St. Joe and Harriet no longer show a "Time to Too Low." Those two lower Buffalo gauges almost never actually drop to the Too-Low mark — they recede into the low hundreds of cfs and just sit there (historically they reach Too Low only ~2% of the time), so that countdown was simply wrong. It's now hidden for them; they still show "Time to Low but Floatable." The upper gauges, which genuinely do bleed all the way down, are unchanged.
- The "Confidence" label now means what you'd expect. It used to reflect how much recent data we had — so it could say "high" on a fuzzy, days-out guess. Now it reflects the actual uncertainty of the number: a tight, near-term call reads high; a wide, multi-day one reads low. (A 7-day recession forecast is a coin-flip on whether it even stays dry, so it'll honestly say low.)
- The countdowns read consistently everywhere. The Buffalo mainstem gauge cards and the Mulberry / Big Piney float-section cards now use the same "Time to Optimal → Time to Low but Floatable → Time to Too Low" wording and order as the standalone creek pages.
Not user-facing, recorded for completeness: the timing model is a per-gauge master recession curve — a flow-dependent decay rate k(Q) learned from USGS history, integrated to a transit time and event-anchored to the current observed rate (clamped 0.5–2×) — in /home/dave/recession_eval/recession_curve.py + recession_curves.json (9 gauge curves), imported by every *_assemble.py compute_recession() with the old single-exponential kept as a fallback. The St. Joe/Harriet suppression uses a config recession_baseline (empirical floor, ~p10) as the decay asymptote so any threshold below it returns no countdown. Confidence is now derived from each prediction's relative spread + horizon (recession_curve.confidence_level). A prediction-evaluation framework (recession_archive.py, on cron) now records every recession countdown to a ledger and settles it against actuals to score accuracy over time. Threshold keys were standardized to too_low/low_floatable/optimal across all configs/assemblers. Full design, history pulls, and analysis live in creekintelligence/research/recession_eval/.
Version 2026.6.22.2 — June 22, 2026
Clearer recession countdowns + gauges stop briefly greying out
- The "Recession Countdown" cards now read in plain falling order — Time to Optimal → Time to Low but Floatable → Time to Too Low. On the Cossatot, Richland, and Hailstone (Upper Buffalo) pages, each countdown is labeled "Time to …" and laid out left-to-right in the order the creek actually drops through the levels, color-coded to match the level legend (green Optimal, yellow Low but Floatable, red Too Low). When a creek is running above the recommended range and falling, the card now leads with Time to Optimal — when the high water is expected to settle back into the good range — then the lower levels after. (Previously the above-recommended case showed only the two lower levels, and on the Hailstone card the order and colors were off.)
- Hailstone's recession card sizing is fixed — the countdown numbers had been rendering in an odd, unstyled font; they now match the big, readable style the other creeks use.
- The "Confidence" line under each countdown is clearer — centered and bold, and it now reads "Confidence in countdown timing" so it's obvious the High/Medium/Low refers to how trustworthy the time estimates are (how cleanly the creek is falling), not the level itself.
- Gauges no longer briefly turn grey when a single USGS reading drops out. Every so often the USGS feed returns nothing for one gauge for a cycle, which used to blank that gauge (grey, "Unknown") for ~15 minutes until the next update — and it rotated across gauges (Boxley, Ponca, Pruitt…). The system now retries, and if a reading is still missing it holds the last good value (still timestamped, so you can see its age) rather than going blank — unless the gauge has truly been silent for over 90 minutes, in which case it still shows as stale.
Not user-facing, recorded for completeness: the three height/CFS recession cards were consolidated into one shared render_recession_card() helper in dashboard.py (they had drifted into near-duplicate blocks edited in parallel). Hailstone's recession model now also emits hours_to_high (time to fall to the 2000-cfs Above-Recommended boundary) in hailstone_output.json. The gauge-fetch resilience lives in creeks/fetch_gauges.py: fetch_stream_readings() retries transient empty USGS responses, and main() carries forward the previous good reading (new carried_forward flag, 90-minute cap) so one failed pull no longer greys a gauge.
Version 2026.6.22.1 — June 22, 2026
Ponca rainfall card now leads with the trusted signals + new downstream propagation card
- The "Ponca Gauge — AI Rainfall Event Analysis" card now leads with the deterministic signals, not the historical analog. During a developing event it reads, in order of reliability: the watershed rainfall alerts (the hand-calibrated tripwires on the specific drainages above Ponca — Upper Buffalo, Beech Creek, Adkins, Boen Gulf), the upstream Boxley gauge, and the observed rainfall + antecedent wetness. The historical-analog comparison is now a clearly-labeled secondary note — it leads only on ordinary, low-level rain days (its accurate range) and can no longer talk the card down to "won't develop / Fizzle" while those upstream signals are firing. The card also now says the river is in flood when it's already there (instead of "headed to flood"), and pivots to the live concern — the continuing climb and downstream propagation. It refreshes every 15 minutes (was hourly). Prompted by the June 22 event, where the old analog-led card read "won't develop into a meaningful rise," then "could go either way," while the river was on its way to a major flood.
- New "Ponca Gauge — AI River Rise & Propagation Analysis" card (
/buffalo/) — appears above the mainstem gauges only while the river is rising. It narrates how a rise at Ponca rolls downstream: roughly when the surge will crest at Pruitt (~5 h behind Ponca, about Ponca's level) and at St. Joe / Grinder Ferry (~14 h behind, typically several times higher), and how high each is expected to get — as a range, compared to that spot's own flood stage. Travel time shifts with how wet the ground is (dry ground = slower). Built for NPS rangers and campers weighing the gravel bars downstream — it answers "how high will Grinders get, and how long before it spikes." Experimental — timing is approximate and magnitude is a range; verify against the gauges. Runs on the same local AI model (no external calls, no cost).
Not user-facing, recorded for completeness: the rainfall card now reads the live creek-alert state (read-only) and the Buffalo physics propagation/flood signals already in buffalo_output.json; the propagation card's lag/amplification constants come from the 12-year USGS archive cross-checked against the Buffalo Study's calibrated 2026 event log, and are antecedent-conditional. Both narratives are produced by the local qwen3.6:27b model — the propagation one is a separate focused prompt embedded in ponca_outlook.json.
Version 2026.5.24.1 — May 24, 2026
Ponca Gauge historical event reporting + experimental AI rainfall event analysis
- New Ponca Gauge Historical Event Reporting page (
/ponca/historical/) — a deep history of how the Buffalo at Ponca has behaved, using the National Park Service flow buckets for the Ponca→Pruitt commercial-float section (Very Low <100, Low 100–200, Moderate 200–900, High 900–1600, Flood ≥1600 cfs; outfitters may launch 100–1600). It shows days per year in each stage and a storm-by-storm catalog back through the radar-rainfall record (Feb 2015→present, with earlier flow-only events on an extended page). Each storm lists the rainfall that drove it (basin-average over the four HUC12s above Ponca, plus a per-HUC breakdown and the wettest single sub-watershed), the starting flow, and every ≥4-hour stage the river held on the way up and down. Linked from the Buffalo page below the 3-day forecast.
- New "Ponca Gauge — AI Rainfall Event Analysis" card on the Buffalo page (
/buffalo/) — appears above the mainstem gauges only when rain is actively falling over the Ponca headwaters (≥0.25″ accumulated). It compares the unfolding storm to ~166 similar past storms and writes a short, plain-English outlook — the most likely peak stage, and whether there's meaningful risk of blowing past the 1,600 cfs launch ceiling — in qualitative terms ("watching… / likely Moderate / high confidence of Flood"). It updates each hour as the storm builds, switches to an "already crested, now receding" read once past the peak, and disappears after about a day of dry weather. Runs on a local AI model (no external API calls, no cost). Experimental — it's a directional read of upstream rainfall vs. history, not a forecast; always verify against the gauge.
- Navigation buttons added to all historical report pages — the Cossatot, Richland, Hailstone, and Ponca historical pages now carry the standard navigation footer (previously they had no way back to the rest of the site).
- Fixed the American Whitewater link on the historical pages (the Ponca-section link had pointed to the wrong run).
Not user-facing, recorded for completeness: built a local archive of MRMS radar rainfall for all 37 Buffalo HUC12 sub-watersheds (2014→present, two source archives) plus full USGS history for all 7 Buffalo gauges, which back the historical report and the analog evaluator. The evaluator (buffalo_dashboard/ponca_analog.py, hourly cron :20, flock) runs a numpy k-NN match against a shipped analog library and calls the local model only when its call changes (~10 inferences per event), writing ponca_outlook.json for the dashboard to render.
Version 2026.5.21.1 — May 21, 2026
PayPal-based donation support
- New
/support/ page — explains what donations fund and embeds a PayPal donate button (merchant KZM3SRM3W34US). Dark-themed to match the rest of the site. Includes a fallback text link for cases where the PayPal button image fails to load, and a required disclosure that donations are personal gifts (not tax-deductible — Arkansas Creek Intelligence by Druidnetworks is not a 501(c)(3)).
- New
/support/thanks/ page — optional return destination after a completed donation. Short thank-you, link back to the gauges.
- Unobtrusive "♥ Support this project" footer link added to
/, /gauges/, /watersheds, /cossatot/, /richland/, /hailstone/, /mulberry/, and /buffalo/. Deliberately styled in muted gray near the existing disclaimer — available without nagging. Not added to /admin/, /suggest, /study/*, /changelog, or /guide/.
- No new dependencies, no environment variables, no database changes — the donate form posts directly to
paypal.com; PayPal handles payment, receipt email, and merchant deposit. Removing the routes removes the donation surface entirely.
Version 2026.5.20.1 — May 20, 2026
Hailstone reaches feature parity + nightly AI analysis moves to local inference
- Hailstone Intelligence is now a full peer of Cossatot and Richland — the Upper Buffalo run page gained a physics predictor, recession countdown, nightly AI analysis, and per-prediction + per-event archives. Previously the page was composer-only (gauge state + empirical forecast from the buffalo dashboard); it now models rainfall → CFS response for the Boxley drainage (156 cells, ~156 km²) on its own hourly cycle. The same gauge (USGS 07055646) drives both the Boxley card on the main /gauges/ table and the Hailstone-specific tier mapping (500/700/2000 CFS).
- Hailstone gauge tile now shows feet below CFS — the creek system now fetches both flow and gauge height for 07055646 (previously CFS-only). The Hailstone card mirrors Cossatot/Richland's primary-value/secondary-value layout: CFS as the headline value, current stage in feet on a small subdued line below the tier label.
- Nightly AI Analysis Archives moved to local inference — Cossatot's and Richland's nightly calibration analyzers used to call Claude Sonnet via the public API. Both now run against a locally-hosted qwen3.6:27b model on the inference host, using server-default parameters. Hailstone's new analyzer uses the same local model. No external API calls are required for any of the three nightly reviews.
- Nightly analyzer scope expanded to include the empirical forecast — all three nightly reviews now grade two distinct models per day: the physics predictor (which they may also adjust calibration on) and the empirical forecast headline (which they comment on but don't tune — the empirical engine's percentile tables are regenerated separately). Each daily markdown now has an "Empirical Forecast" section showing the headline that was issued, the settled outcome, and the AI's assessment of whether the headline panned out.
- Physics predictor baseline fix — the RISING branch in both
cossatot_predict.py and richland_predict.py was double-counting in-transit water during active rainfall, producing fake peaks of up to +2.0 ft on Cossatot dawn predictions and a similar pattern on Richland. The baseline now holds at the current gauge value until band contributions finish arriving, then transitions to the recession decay. Validation on the worst archived errors: Cossatot 2026-05-20 06:09 prediction went from 5.28 ft → 3.33 ft against an actual peak of 3.29 ft. A retrospective supplement at /cossatot/analysis/2026-05-20-Opus-Supplement documents the diagnosis and the per-prediction replay improvements.
- Recession countdown now suppresses during active rainfall — when measurable rain is falling near the gauge (>0.05"/hr), the recession card displays "Suppressed — active rainfall in the basin invalidates decay extrapolation" instead of the countdown values. The underlying
recession_k / recession_h_base parameters are still computed in case the physics predictor needs them, but the dashboard hides time-to-tier estimates that the exponential-decay assumption can't honor mid-event.
- Standardized card layout across the three intelligence pages — Cossatot, Richland, and Hailstone now follow the same eight-card structure: Gauge → Recession (conditional) → Rainfall → Map → Physics Predictor → Physics Predictor Archives → Empirical Forecast → Empirical Forecast Archives → Nightly AI Analysis Archive. All archive page titles match their card-link names exactly.
- Color-coded tier breakpoints on the gauge card — the right side of each gauge card now lists all four tiers in their semantic colors (Too Low red, Low but Floatable yellow, Optimal green, Above Recommended blue) with the exact thresholds for each. Replaces the old "Tier breakpoints: <3.0 ft Too Low · 3.0–3.4 …" inline strip that ran together at the bottom of the card.
- "Too Low" gauge tile flipped from gray to red — the colored tile inside the gauge card now uses the same red as the recession card and the tier list when a creek is below its low threshold, matching the Buffalo dashboard's existing scheme. The gray was a holdover from earlier when "Too Low" was treated as a passive state.
- Predictor status text standardized on rainfall language — both non-active states use the same wording on all three pages: "No upstream rainfall to drive a prediction" when QPE is quiet across the watershed, "Upstream rainfall too light to drive a rise" when bands have rainfall but the contributions sum below the rise threshold.
- Subtitle text format unified — all three intelligence pages now use
USGS {gauge id} · {river} near {town}, AR, matching the USGS station's official name. Previously Richland listed zones in the subtitle and Hailstone listed the run name.
- Hailstone gauge precip line added — the per-page Gauge precip total (sourced from the USGS 00045 collocated rain gauge on 07055646) now appears in the Hailstone gauge card next to the trend/time lines, matching Cossatot and Richland.
- "Phase 2" header label removed from Cossatot — the Cossatot meta line used to read "Data age: N min · Phase 2"; now reads just "Data age: N min" like the others. The Phase 2 work shipped weeks ago.
Underlying schema, infrastructure, and unit-bug cleanups (not user-facing but recorded for completeness): new hailstone_predict.py + hailstone_calibration.json + hailstone_predictions_archive.py, three new cron entries (:17 Hailstone QPE consumer, :21 Hailstone physics predictor, 23:55 Hailstone nightly analyzer), fixed a 25.4× unit error in the new Hailstone QPE consumer (had been treating already-in-inches shared QPE values as mm and dividing again — antecedent moisture was reading DRY when the basin was actually NORMAL/WET), exposed recession_k/recession_h_base on Richland's gauge output (latent gap that left the physics predictor falling back to a flat baseline instead of an exponential decay), and added a per-supplement file-naming convention (YYYY-MM-DD-suffix.md) for retrospective analyses written outside the nightly cadence.
Version 2026.5.15.1 — May 15, 2026
Stale-Data Warning Banners Across All Mobile Dashboards
- Home page (
/) freshness pill — new always-on indicator at the top of the page with three states: green "Live — updated N min ago" when data is fresh, yellow "Data may be stale — last update was N minutes ago" when past the 30-min threshold, and red "Data feed unavailable — ScriptPi push may be down" when predictor_output.json can't be loaded at all (previously the page silently rendered as "All Quiet" during push failures, hiding the outage)
- New stale banners on five sub-dashboards —
/cossatot/, /richland/, /hailstone/, /mulberry/, and /buffalo/ now render the same yellow top-of-page warning banner that /gauges/ and /watersheds have had. Fires only when each page's underlying JSON is past its 30-min _stale threshold, so the banner only appears when something is actually wrong
- New shared helpers —
render_stale_banner(data, expected_interval_min=15) and render_freshness_pill(data, expected_interval_min=15) live alongside the other freshness utilities (_format_age, _age_color). The banner markup was previously duplicated inline on /gauges/ and /watersheds; those two pages now call the helper (no visual change)
- Motivation — a recent ScriptPi push failure went unnoticed because the home page and sub-dashboards had no top-of-page indicator that data was stale. Only the two pages that had banners (gauges/watersheds) made the failure visible
Version 2026.4.21.1 — April 21, 2026
Creek Page Title & Buffalo Card Standardization
- White page titles across creek dashboards — Cossatot, Richland, and Buffalo h1s now render in white (
#ffffff) to match the landing page style (previously cyan #4fc3f7)
- Cossatot title shortened — "Cossatot River — Watershed Intelligence" → "Cossatot River Intelligence"
- Buffalo title renamed — "Buffalo River Dashboard" → "Buffalo River Intelligence"
- Buffalo gauge cards now color-coded by level — full-card backgrounds match the landing-page scheme: blue (Above Recommended), green (Optimal), yellow (Low but Floatable), red (Too Low), neutral gray (Unknown). Replaces the small inline level badge with a full-width spelled-out level line.
- Sub-section contrast tweak — trend arrow and gauge-sub blocks (Rise Prediction, Recession, Flood Risk) on Buffalo cards now sit on a semi-transparent dark background so their semantic colors remain legible against the colored card.
Version 2026.4.15.1 — April 15, 2026
Buffalo River Dashboard
- New
/buffalo/ dashboard — Buffalo River Dashboard page with 5 mainstem gauge cards (Boxley, Ponca, Pruitt, St. Joe, Harriet), 2 tributary signal separator cards (Richland Creek, Bear Creek), interactive Leaflet.js HUC12 rainfall map, zone rainfall table, flood/propagation alert banners, rise prediction and recession countdown sections per gauge
- New
buffalo_output.json data source — pushed by ScriptPi every 15 min, contains 7 gauge readings, 37 HUC12 QPE values, 8-zone rainfall aggregates, per-gauge rise predictions, recession countdowns, propagation alerts, and flood risk levels
- Leaflet.js HUC12 choropleth map — 37 sub-watershed polygons colored by rainfall intensity, 3 toggle modes (QPE 1hr, QPE 24hr, QPF 24hr), 7 gauge circle markers colored by level, CartoDB dark tiles, tooltips with zone name and rainfall values, graceful fallback if GeoJSON not deployed
- GeoJSON static file — simplified
buffalo_huc12_simple.geojson (136KB, 37 features) served from new /static/ directory
- Contextual recession messages — gauge cards always show recession status: "Already below thresholds" (Too Low), "Rising" (rising trend), "Holding steady" (stable, not receding), "Beyond forecast window" (receding but thresholds far off), or countdown timers with confidence
- buffalo_only gauge filtering — 4 new creeks.yaml entries (Pruitt, St. Joe, Harriet, Bear Creek with
buffalo_only: true) filtered from /gauges/ table and landing page cards/counts
- Buffalo QPF drainages hidden from
/watersheds/ — "Buffalo QPF" family drainages (pruitt_zone, stjoe_zone, bear_creek_zone, harriet_zone) filtered from the watersheds page
- QPE/QPF explainer — subtitle below rainfall map: "QPE = Precipitation Estimate (observed) · QPF = Precipitation Forecast"
- Home button in nav footer — gray "Home" button added to the shared
nav_button_footer() bottom row alongside Guide, Changelog, and Suggestions
- Buffalo River link — added to landing page tool grid and
nav_button_footer() blue button grid
- Centered text across mobile pages — all non-card text (titles, subtitles, data age, section labels, rainfall metadata) centered on Cossatot, Richland, and Buffalo dashboard pages
- Watersheds "Back to Home" repositioned — moved above the page title
- Dashboard grew from ~2,975 to ~3,565 lines
Version 2026.4.7.1 — April 7, 2026
Richland Creek Dashboard & Neural Net Predictors
- New
/richland/ dashboard — Richland Creek Watershed Intelligence page (USGS 07055875) with gauge status card, recession countdown, two-zone watershed rainfall table (Upper Richland + Falling Water + Combined), physics predictor card, and LSTM neural net forecast card (8 horizons)
- New
/richland/analysis/ placeholder routes — "Nightly analysis coming soon" index page; date routes return 404
- New
richland_output.json data source — pushed by ScriptPi every 15 min, contains gauge data, two-zone rainfall, physics prediction, and neural net prediction
- Neural net predictor card on
/cossatot/ — 12-horizon LSTM forecast with peak CFS, forecast timeline table, threshold crossings, and data quality indicators
- Shared
render_nn_card() helper — renders identical card layout for both rivers; shows peak forecast, per-hour timeline with peak row highlighted, CFS-to-level threshold crossings, and QPE/ASOS data quality footer; handles null (initializing), error, stale (>90 min), and active states
- Physics card renamed on
/cossatot/ — "Rise Prediction" / "Prediction" headers changed to "Physics Predictor" with gear emoji across all card states (active, quiet, no data)
- Landing page updated — added "Richland Intelligence" to the tool navigation grid; renamed "Cossatot Predictor" to "Cossatot Intelligence"
- Nav footer updated — added Richland Intelligence link to the shared
nav_button_footer() used across all pages; renamed Cossatot link to match
- Shared constants refactored —
COSSATOT_LEVEL_COLORS renamed to LEVEL_COLORS, COSSATOT_TREND_SYMBOLS renamed to TREND_SYMBOLS (used by both river dashboards)
- New constants —
NN_THRESHOLDS (CFS-to-level mappings per basin: Richland 500/1000/1500, Cossatot 150/730/1800), NN_LEVEL_COLORS (gray/green/blue for Too Low/Optimal/High), is_nn_stale() helper
- New architecture document —
DMZPI_ARCH_4_7_26.md supersedes DMZPI_ARCH_3_11_26.md
- Dashboard grew from ~2,460 to ~2,975 lines
Version 2026.3.10.1 — March 10, 2026
Landing Page & Route Restructure
- New landing page at / — mobile-first dark theme with site title, orientation blurb, live status summary, and dynamic creek condition cards
- Creek cards show all Optimal gauges (green) or, if none, all Low but Floatable gauges (yellow); "All Quiet" message when no creeks are runnable
- Status summary line shows gauge and watershed condition counts with color-coded text
- Tool navigation grid — four main buttons (Gauges, Watersheds, Cossatot Predictor, Buffalo Study) plus secondary links (Guide, Changelog, Suggest a Creek)
- Gauge table moved from / to /gauges/ — all table logic unchanged
- Temporary legacy link on landing page points to /gauges/ for returning users
- New /guide/ placeholder page — "Coming Soon" with back link to home
- Back link audit — watersheds, changelog, study, and suggest pages link back to /gauges/; error pages and Cossatot nav link to / (landing)
- Last updated timestamp on landing page converted from UTC to Central time
Version 2026.3.9.1 — March 9, 2026
Four-Color Unification
- Unified color language across creek levels, prediction status, and recent rain: Red (nothing) / Yellow (maybe) / Green (go) / Blue (lots)
- Prediction status expanded to four tiers: No Alert (red), Watch (yellow), Warning (green), Flood (blue)
- Recent Rain column now shows 7-day precipitation total in inches with color-coded background, replacing category labels (MOIST/SEMI-DRY/DROUGHT)
- Recent Rain multipliers updated: <0.25" = 1.4x trigger, <0.75" = 1.2x, <1.50" = 1.0x, ≥1.50" = 0.9x
- FLOOD status triggers at 200% of effective trigger threshold — indicates exceptional rainfall
- Micro-creek lag display: Drainages with 0-1 hour lag now show "NOW" instead of numeric range
- Signal alerts updated with lag-aware messaging (micro-creeks show "NOW", mainstem rivers show hours)
Version 2026.3.6.1 — March 6, 2026
Drainage Trigger & Timing Calibration
- Adkins: 2.0" / 4hr → 2.5" / 6hr
- Boen Gulf: 2.0" / 4hr → 2.5" / 6hr
- Upper Buffalo: window 6hr → 12hr, lag 4-6hr → 6-8hr
- Beech Creek: 1.5" → 2.0"
- Upper Kings: 1.5" → 2.0"
- Osage: 1.5" / 4hr → 2.5" / 6hr
- Richland Main: window 6hr → 12hr
- Falling Water: 1.5" / 6hr → 1.75" / 12hr
- Upper Cossatot: window 6hr → 12hr
- Upper Big Piney: window 6hr → 24hr, lag 10-12hr → 12-16hr
- EFLB: 1.5" → 2.0"
- Pine Creek OK: window 6hr → 12hr
Version 2026.3.5.3 — March 5, 2026
Cosmetic Updates
- Renamed "DRY" status to "QUIET" on the Watersheds page and in alert bar logic (same red styling, new CSS class .st-quiet).
- Renamed "Conditions" column to "Recent Rain" and changed from styled badge spans to full-cell background coloring (matching the Status column style).
- Updated status_colors/status_text_colors dicts to use "QUIET" key.
Version 2026.3.5.2 — March 5, 2026
Sticky Status Hold
- WATCH and WARNING statuses now hold for the duration of a drainage's lag time plus a 2-hour buffer before clearing, preventing premature status downgrade before water reaches the gauge.
Antecedent Dryness System
- New "Conditions" column on Watersheds page showing MOIST, SEMI-DRY, or DROUGHT based on recent rainfall history.
- Trigger thresholds automatically increase during dry conditions: +15% for SEMI-DRY, +30% for DROUGHT.
- Dryness is computed from rolling 7-day and 30-day precipitation totals per drainage.
- Trigger column on Watersheds page now shows the effective (adjusted) trigger value.
Cossatot Drainage Update
- Upper Cossatot trigger raised from 1.00" to 1.25" in 6 hours.
- Upper Cossatot lag time updated from 3-6 hours to 8-10 hours based on observed March 5 event.
YAML Sync
- Creek and drainage definitions now automatically sync from ScriptPi to DMZPi every 15 minutes, ensuring single source of truth.
Changelog
- Added this changelog page, accessible from the footer of both dashboard pages.
Version 2026.3.5.1 — March 5, 2026
Baseline version. All prior changes consolidated.
WATCH / WARNING Terminology
- Renamed "TRIGGER" status to "WARNING" to align with NWS conventions.
- WATCH threshold raised from 50% to 75% of trigger value to reduce false positives.
Color Scheme Standardization
- Watersheds page: RED = DRY (no go), YELLOW = WATCH (maybe), GREEN = WARNING (go time).
- Main creek page: Watershed radar column and alert bar colors match the same scheme.
- Alert bar is now green for WARNING, yellow for WATCH-only.
Drainage Trigger Updates
- Bobtail Creek: 1.5" → 2.0" in 6hr
- Long Devils Fork: 1.5" / 4hr → 2.5" / 6hr
- Big Devils Fork: 1.5" / 4hr → 2.5" / 6hr
- West Fork Shop Creek: 2.0" / 4hr → 2.5" / 6hr
- Thomas Creek: 2.0" / 4hr → 2.5" / 6hr