← Back to Home

Arkansas Creek Monitor — Changelog

Version 2026.9.9.2 — September 9, 2026

Richland, Hailstone and Cossatot rise predictions are now issued eagerly, and every rise call carries a confidence label that says how often calls like it have come true.

Not user-facing, recorded for completeness: the served physics fit on all three basins is the eager (τ 0.75) G6 refit, the τ 0.6 and τ 0.5 fits ride along in engine_v2.dials; physics_serve.rise_confidence(); per-dial peaks and the quoted rate on every archive record (quiet runs too); score_all._physics_claims.by_confidence; the nightly scorer had been crashing on the physics claims' event view since 2026-09-09 morning (the scorecard page was a day stale) — fixed and re-pushed; evidence in research/physics_confidence/REPORT.md.

Version 2026.9.9.1 — September 9, 2026

A radar hole is now "unknown", not "dry", from the radar feed to the page; a dead rain bucket shows N/A instead of 0.00; grading no longer counts readings USGS has withdrawn.

Not user-facing, recorded for completeness: one shared missing-radar rule (basinlib/radar_gap.py) in all shared-QPE readers (Illinois's reader became a basinlib shim), day-files keyed on the MRMS hour with a 6-hour backfill of missed cycles, the shared fetcher replaces an hour built from a shifted fallback frame once the top-of-hour frame appears; one shared truth reader (basinlib/truth.py) for every settler (withdrawn and stage-derived readings are never graded on); nightly truth refresh marks withdrawals only inside the span USGS served and reports parameters served nothing; recession re-grading marks a countdown whose revised start was already at its target as invalidated (the first nightly pass will flip ~168 "hit at 0.25 h" grades); physics/empirical re-grading skips windows that end after the truth pull; study day-file gauge names from the study config; the dashboard's stale-data thresholds now come from the shared data_age_thresholds.yaml.

Version 2026.9.8.7 — September 8, 2026

Neural Net Predictions: the level table and the crossing alert can no longer contradict each other, the Cossatot page shows real feet, and a USGS outage no longer blanks the page.

Not user-facing, recorded for completeness: rain regime now five bands (no cliff at 10 mm), Cossatot shrink factors smoothed across horizons, Boxley posting lag logged hourly ahead of a cron move, missing radar hours re-read for 48 hours, clock-skew guard, notes persisted in the ledgers; research/neural_g7/REPORT.md.

Version 2026.9.8.6 — September 8, 2026

Empirical Forecast card: the rain columns say what they are, the Illinois headline stops flip-flopping, and the archive pages grade every hour in plain words.

Not user-facing, recorded for completeness: empirical_predict.py 2026.9.8.6 — rain_context replaces the v1 percentile_band/p25/p50/p75/n fields, the legacy WATCH headline path is deleted, MRMS missing-radar sentinels are counted per basin-hour, per-window gap status, flags.stale_qpe/qpe_age_h, hysteresis input from the previous issue; calibrated.py CLAIM_P, display_pct, per-tier capped, table-driven hysteresis (Illinois only, gated on the v2 rig: OOS BSS 0.303 → 0.374); empirical_archive.py study-archive day-file branch for Boxley/Hailstone, coverage test at both window ends, new .md renderer; resettle.py first-settles never-settled records; score_all.py claim-only episodes; hailstone_analyze.py reads the per-basin v2 archive and takes --date. ARCHITECTURE §4.8 + §16 item 51; research/empirical_g8/.

Version 2026.9.8.5 — September 8, 2026

Ponca AI Rainfall Event Analysis card: its history now matches the way it was trained, and a stale write-up says so.

Not user-facing, recorded for completeness: ponca_analog.py 2026.9.8.5 — per-series history (80 h of flow readings), storm_reference() frozen storm frame with cum_source, trend_by_time(), dry hours by MRMS-hour time, tied router quantiles handled in p_exceed, class_word(), qwen ValueError caught + qwen_ok/narrative_at/narrative_age_min/narrative_stale in the outlook, shared 7-min qwen budget; storm-LOO and season-replay gates in research/ponca_analog_g9/. DMZPi render_ponca_outlook_card() stale-note rendering.

Version 2026.9.8.4 — September 8, 2026

The three physics rise predictors (Cossatot, Richland, Hailstone) were re-fit on an honest basis, no longer under-call a rise while the gauge feed lags, show feet that match USGS's own rating, and are now graded on every hourly run — including the ones that said "no rise".

Not user-facing, recorded for completeness: basinlib/physics_engine.py (18-h timeline, gauge lag, kernel-end hold, crest recession, NaN radar slots, model floor built but gated off — it raised false alarms to 45%), physics_serve.py, rating.py + new fetch_usgs_rating.py (weekly exsa cron), predictions_archive.py / hailstone_predictions_archive.py grade v3 (claims, level at issue from the card, cfs at settle), score_all.py claims block + engine filter, resettle.py back-fill, Hailstone reader on ZoneInfo dates with the radar-gap flag and MRMS-hour dedup, refit engine_v2 blocks (v2.1-g6, quantile 0.6 / 0.5 / 0.6). Fresh-eyes review session G6; evidence research/physics_serve_path/REPORT.md; ARCHITECTURE §4.3–4.5, §16 item 49.

Version 2026.9.8.3 — September 8, 2026

The Buffalo wave tracker stops guessing outside its experience: at drought base flow it now uses the plain lag table instead of a multi-day arrival window, and its "still rising" windows use the gauge readings it was actually fitted on.

Not user-facing, recorded for completeness: reach_models.json lag_model.regression.support per reach; wave_router.lag_quantiles/remaining_rise/_prior_reading/_readings_seq/_value_at, DEAD_FEED_MARGIN_MIN, ARRIVAL_EARLY_SLACK_H, window_early; buffalo_assemble readings feed, wave_inbound_significant, wave_band_confidence, timing_state, no_move_cut_h, last_rain_obs/consumed_at; wave_eval families + reopen_count keys; buffalo_predictions_archive overdue skip + rose_within_24h; mid kept after the v8d2/v8f re-run (one extra late rise per lower gauge per decade for +0.2–1.8 h fizzle hold); simulator lagQuantiles port + bundle, golden 2,200/0. Fresh-eyes session G3; evidence research/buffalo_g3/REPORT.md; ARCHITECTURE §4.6 / §16 item 48.

Version 2026.9.8.1 — September 8, 2026

Watershed alerts now follow the creek gauge, say plainly how often they verify, and reach Facebook exactly as they reach Signal.

Not user-facing, recorded for completeness: resolve_status() / compute_alert_transitions() / notify_events() + alert_state.json event_ledgerpredictor_output.json alert_events (alert_rules 3.1, FLOOD_DEBOUNCE_H 1.0); qbucket_shadow.py creek_ran/event_peak_q_cfs, skill_text/skill, America/Chicago season + day-file names; qbucket_config.json 3.1.2026-09-08 (thresholds byte-identical, loyo.watch added by recal_v3.py); Mulberry's Tier-1 record logged every cycle; antecedent_precip.json keys before 09-04 re-keyed +1 day (WP5 side effect); prediction_eval/qbucket_eval.py + score_all qbucket block (coverage → graded); creek-social ledger replay with a _meta.last_event_id cursor; Tier-2 reachability table (two months: no Boxley-referenced drainage has come within 60% of its dry-season threshold) recorded in ARCHITECTURE §4.2 — multipliers unchanged by decision. Fresh-eyes review session G4; evidence research/qbucket_v31/.

Version 2026.9.7.4 — September 7, 2026

Richland Creek's flow feed has been silent since September 4 and the site now says so — and keeps predicting honestly instead of quietly switching to an older, less reliable method.

Not user-facing, recorded for completeness: basinlib/qsynth.py (stage-derived discharge in a separate discharge_synth day-file block, real USGS readings untouched), series_state on every fetched series, per-parameter data_age_alert, usgs_shadow empty-series counts, engine/q_source on archive records, Q-bucket 6-h Q age rule and clamped-pre-event-Q refusal, legacy_fallback trigger method, ZoneInfo day keys in the basinlib/Illinois/study gauge fetchers (the nightly 00:00–01:00 CDT "yesterday's reading" hour on four basin pages should be gone from tonight), data_health block in scorecard.json. Fresh-eyes review session G1; details ARCHITECTURE §4.2/§4.3/§4.13/§4.14, §16 item 45.

Version 2026.9.7.3 — September 7, 2026

Buffalo rain totals corrected: since early July the Buffalo page's "last hour" (and 3/6/12/24-hour) radar rainfall had been counting one extra hour about half the time.

Not user-facing, recorded for completeness: buffalo_read_qpe.py windows keyed on the MRMS hour anchored at the newest snapshot (window_rule + newest_mrms_hour_utc in current.json); buffalo_assemble.py exports qpe_stale/qpe_age_min/snapshots_last_24h and skips the magnitude band while stale; radar_eval.py --truth-file/--resettle with truth_source per row (494 frames re-settled on the F7 archive; the ledger-truth rows kept as radar_settled.jsonl.bak.20260907_1953_g2_ledger_truth); boundary test test_g2_boundary_windows in the F6 suite (30/30); evidence research/buffalo_windows/REPORT.md.

Version 2026.9.7.2 — September 7, 2026

The Buffalo historical archive now refreshes itself monthly.

Not user-facing, recorded for completeness: harness cron monthly_refresh.sh (USGS full re-pull, nClimGrid last 6 months, MRMS zone rain extended from NOAA S3 with the F7 catchment cell lists, build_archive.py, scp as the new agenthost-hist DMZPi identity, public-URL verify, failure email). No Flask change.

Version 2026.9.7.1 — September 7, 2026

New: a historical levels and storm archive for the four main-stem Buffalo gauges — Ponca, Pruitt, St. Joe (Hwy 65) and Harriet (Hwy 14) — reaching back to 1939 at St. Joe.

Not user-facing, recorded for completeness: new research/buffalo_historical/ on the harness (full USGS period of record for the four gauges, hourly Stage IV 1997–2015 and daily nClimGrid 1951–present crops of the Buffalo basin, NPS threshold provenance, event pipeline build_archive.py); DMZPi gets buffalo_historical/ (shared JS shell + per-gauge JSON) and three new routes; /ponca/historical/ routes are 301s; pages registered for the sitemap/llms.txt.

Version 2026.9.6.1 — September 6, 2026

No visible change: the gauge-corrected rain product was back-tested against every rain-driven predictor, and only one piece (the Richland physics engine) earned a possible switch.

Not user-facing, recorded for completeness: research/mrms_pass2/backtests/ (REPORT.md, PLAN.md, per-family RESULT.md, SNAPSHOT_TREE_DESIGN.md); archive extended with a Mulberry/Big Piney window; no Pi file changed; qwen3.8:27b stopped for two GPU training blocks on the inference host and reloaded. Model-review fault F7, work package 7b item 2.5.

Version 2026.9.5.3 — September 5, 2026

No visible change: the Buffalo rise-size band was re-checked on the new catchment zones and kept as is.

Not user-facing, recorded for completeness: magnitude_hourly.json unchanged (same-basis refit gated against the shipped artifact on the same rows, research/f7_catchments/wp7b/); buffalo_config.yaml unchanged (research/f7_catchments/wp7b/threshold_rules_live_vs_cand.csv); era ratio research/era_ratio/REPORT.md; replay rigs re-synced to the live reach models with the catchment zone rain as default; legacy per-HUC12 archive retired; MultiSensor Pass2/Pass1 archive fetch running (research/mrms_pass2/). Model-review fault F7, work package 7b (items 2.1, 2.2, 2.4, 2.6); plan research/f7_data_layer/WP7B_PLAN.md.

Version 2026.9.5.2 — September 5, 2026

The Cossatot empirical forecast now reads rain over exactly the ground above the Vandervoort gauge.

Not user-facing, recorded for completeness: empirical_forecast/data/shared_qpe_indices.json (basins.cossatot = cossatot/shared_qpe_indices.json zone upstream, 232 cells; other basins unchanged) + the Cossatot block of calibrated_table.json (version 2.2026-09-05, rebuilt on the F2 v2 rig from the Cossatot pixel archive; out-of-sample skill equal to or slightly above the old mask at every tier). Also this session: the live radar-grid convention was confirmed against the pixel archives (only the legacy per-HUC12 research archive was offset by one column, now retired), the Buffalo replay rigs were re-synced to the live reach models with the catchment zone rain as their default, and the MultiSensor Pass2 archive fetch was started. Model-review fault F7, work package 7b (item 2.3); plan research/f7_data_layer/WP7B_PLAN.md.

Version 2026.9.5.1 — September 5, 2026

Mulberry and Big Piney recession countdowns were running fast, sometimes by half in winter; they are rebuilt from each river's own record and now show a typical range.

Not user-facing, recorded for completeness: tables blocks for 07252000 / 07257006 in recession_eval/recession_curves.json (v3; builder research/recession_eval/build_tables_stage.py, cool/warm season pools, LOYO + holdout in VALIDATE_TABLES_STAGE.md, live-ledger replay in ledger_replay_mbp/REPLAY.md); basinlib/assemble.compute_recession table path with curve fallback (serves both assemblers); DMZPi per-section renderer range line + caption (one restart); simulator snapshot refreshed, bundle unchanged, golden 2,200 / 0. Both gauges' stage ratings checked for drift (≤ 0.2 ft at the served thresholds) — tables stay in feet. Plan and decisions in MULBERRY_BIGPINEY_RECESSION_PLAN.md §7; ARCHITECTURE §16 item 42.

Version 2026.9.4.7 — September 4, 2026

The Buffalo page's gauge zones are now the gauges' real catchments, so "rain over the Boxley zone" means rain that actually flows past the Boxley gauge.

Not user-facing, recorded for completeness: buffalo_dashboard/shared_qpe_indices.json v2 (NLDI basins rasterized on true cell centers; Bear Creek WBD-divide patch; gauge_zone_v1 kept on each cell), buffalo_config.yaml zone_summary areas + USGS drainage_km2 + two huc12_zones labels, reach_models.json = the v2 refit. Evidence: research/f7_catchments/ (REPORT.md, history/, refit_replay/, refit_wp4/). Side findings recorded there: the per-HUC12 research archive samples MRMS one column west of the live grid convention; the shipped magnitude artifact is already stale against the live engine's longer timing windows (a same-basis refit is queued). Model-review fault F7, work package 7a; plan research/f7_data_layer/F7_PLAN.md.

Version 2026.9.4.6 — September 4, 2026

Groundwork for the USGS data-service retirement: one shared gauge-fetch module with the replacement service running in shadow, checked every hour against the live feeds.

Not user-facing, recorded for completeness: basinlib/usgs.py (fetch(sites, codes, period_hours|start/end, backend="iv"|"ogc") → the fetchers' exact record shape; OGC adapter renders UTC times as the legacy local-offset strings, maps approval status to A/P; cursor paging; precip_sensor_status() for the dead-tipping-bucket case; --selftest, --live), prediction_eval/usgs_shadow.py (cron :42; usgs_shadow_state.json + usgs_shadow.jsonl), truth_refresh.pull() via the module, score_all truth block usgs_shadow, DMZPi sentence. The five live fetchers still call the legacy host directly; they move to the module after the shadow week. Model-review fault F7, work package 6; plan research/f7_data_layer/F7_PLAN.md.

Version 2026.9.4.5 — September 4, 2026

Radar-rain bookkeeping fixes under the hood: missing-radar hours are now counted instead of read as dry, rolling rain totals say how many hours they actually cover, and every "Central date" in the pipeline is a real Central date.

Not user-facing, recorded for completeness: shared_qpe/fetch_shared_qpe.py (to_inches counts sentinels before clipping; snapshots gain n_missing/frac_missing/missing_idx; collect_recent_snapshots selects by wall-clock hour with a previous-hour fallback; windows gain hours_covered/hours_expected/window_end/missing_hours/frac_missing_max; sums bit-identical on complete hours), backfill_shared_qpe.py prefers the HH:00 key, the basinlib/Cossatot/Richland readers write frac_missing + insufficient, buffalo_read_qpe counts missing cells for real, and ZoneInfo replaces UTC−6 in the Cossatot/Richland predictors, four assemblers, accumulate_daily_precip.py, and the three prediction archives (the physics gauge reader walks both day-file keys). Not changed: the Hailstone reader/predictor pair (consistent with each other) and the 7-day antecedent (display only since engine v2). Model-review fault F7, work package 5; plan research/f7_data_layer/F7_PLAN.md.

Version 2026.9.4.4 — September 4, 2026

Every scorecard section now reports the same sample line, the physics crest grades get an honest window, and countdowns and forecasts are counted by event as well as by hour.

Not user-facing, recorded for completeness: prediction_eval/schema.py (episodes, event rates, Brier/reliability/BSS, coverage, log error, the standard header block, selftest); physics settlers v2 (grade_version, horizon_h 18, level_at_issue_*, rose, crest_err_ft / crest_log_err, timing_err_h, falling-limb claims skipped at archive time; grade_v2() reused by resettle.py); score_all _physics_v2, recession episodes, empirical non-claims, schema blocks on every family, v2-aware physics flags; DMZPi _sc_schema_line + section changes. Model-review fault F7, work package 4; plan research/f7_data_layer/F7_PLAN.md.

Version 2026.9.4.3 — September 4, 2026

The scorecard now re-grades itself as USGS revises its gauge data, and says where its truth comes from.

Not user-facing, recorded for completeness: prediction_eval/truth_refresh.py (00:20, NWIS IV with qualifiers into the six truth archives under each fetcher's flock: value/value_first/qual, removed readings marked X, per-file truth_refresh block) and prediction_eval/resettle.py (01:05, per-family re-grade with v0 from the same archive; outcome_first/actual_first/first_actual_peak_*, resettled_at, resettle_n, truth {archive, vintage, qual mix, final}; physics and empirical day files re-pushed); score_all truth block; hailstone settler _read_study_readings; tarball backups prediction_eval/{ledgers,truth_archives}.bak.20260904_0914_f7_wp3.tar.gz. Waves and the neural graders are not resettled (documented). Model-review fault F7, work package 3; plan research/f7_data_layer/F7_PLAN.md.

Version 2026.9.4.2 — September 4, 2026

Scorecard bookkeeping fixes: every countdown now gets graded, the Buffalo Rise table uses the same 45-day window as the rest of the page, and the physics crest grades no longer close before the gauge data has arrived.

Not user-facing, recorded for completeness: recession_archive.settle_ledger scans all ledger files; buffalo_predictions_archive.score(cutoff=) + score_all all_time block; wave_eval.score windows on expire_after; basinlib/predictions_archive.py and hailstone_predictions_archive.py gain TRUTH_SLACK_MIN/TRUTH_WAIT_HOURS and store truth_last_reading / truth_truncated; score_physics adds n_within_eligible, n_truncated, n_resettled; one-shot resettle_physics_truncated.py (records keep first_actual_peak_*, resettled_at, resettle_reason). Model-review fault F7, work package 2; plan research/f7_data_layer/F7_PLAN.md.

Version 2026.9.4.1 — September 4, 2026

The scorecard now grades the Boxley gauge, which it had silently skipped since June.

Not user-facing, recorded for completeness: buffalo_study/study_config.yaml adds USGS 00060 for 07055646 and 07055875 (no extra USGS calls — the batched fetcher already asked for it); NWIS backfill of 36,272 readings into 190 study day files under the fetcher's flock after a tarball backup; censored Boxley ledger rows re-opened and settled with the current rise rule (research/f7_data_layer/deploy/). This is work package 1 of model-review fault F7 (data layer and scoring); the review measured how much USGS's later revisions to provisional data change settled grades (Ponca is the outlier — 94 % of its 2026 readings have since been re-rated) and the plan is in research/f7_data_layer/F7_PLAN.md.

Version 2026.9.3.8 — September 4, 2026

Buffalo local-rain rise predictions now say how likely a rise is and how big it would typically be, in cfs, instead of "slight / moderate / large".

Not user-facing, recorded for completeness: buffalo_dashboard/magnitude_serve.py + magnitude_hourly.json (per-gauge logistic rise probability and p25/p50/p75 quantile regression on log(rise + 10), with a no-upstream variant; stdlib), predict_local_rainfall emits rise_prob / rise_p25/p50/p75_cfs, the ledger stores them and the grader adds mag_in_band + Brier, DMZPi card and scorecard (one restart), simulator engine port + golden local family (2,200 cases). The category word stays in the JSON for the ledger and the wave-compat merge. Fault F6 work package WP3; evidence research/buffalo_f6/wp3b_hourly/REPORT.md (the first-fire event model in wp3_magnitude/ is the recorded comparator).

Version 2026.9.3.7 — September 4, 2026

Buffalo wave arrival windows now account for how much water is in the river, and the local-rain timing windows were re-measured against twelve years of graded predictions.

Not user-facing, recorded for completeness: wave_router.py lag_quantiles (stage-conditioned quantile regression with the U-band table as fallback), provisional-window policy C with per-launch-gauge remaining-rise models embedded in reach_models.json, per-reach ratio guards (TRIB_RATIO_CLIP 1.5, TRIB_RATIO_MIN_RAIN 0.10), Bear Creek no-ratio crest model, buffalo_config.yaml local windows; 48 unit tests; router gate and rise-ledger replay gate under research/buffalo_f6/; simulator engine/bundle rebuilt (golden 2,200). Fault F6 work packages WP4 and WP6; evidence research/buffalo_f6/wp4_waves/REPORT.md.

Version 2026.9.3.6 — September 3, 2026

The Buffalo wave tracker now follows a storm's second pulse and no longer confuses an earlier crest for the wave it is watching.

Not user-facing, recorded for completeness: wave_router.py re-open logic (REOPEN_FRAC 1.10), provisional-arrival guard, trough-based base_cfs backfill, per-cycle notes logged by the assembler; 34 unit tests; replay gate research/buffalo_f6/fidelity/scripts/gate_router.py. Fault F6 work package WP5.

Version 2026.9.3.5 — September 3, 2026

Buffalo rise predictions are now graded on a scale-aware rule, the scorecard counts events as well as hours, and the gauge arrows finally see slow low-water rises.

Not user-facing, recorded for completeness: basinlib/rise_rule.py (one rule for grader and engine), prediction_eval/buffalo_predictions_archive.py (settle + event-level score, rise_floor_cfs in every outcome), buffalo_assemble.py (consumption/movement bar on the shared rule; compute_trend per-gauge trend_floor with 2-h persistence, fit in research/buffalo_f6/wp2_trend/), buffalo_config.yaml trend floors, DMZPi scorecard columns (one restart). Fault F6 work package WP2.

Version 2026.9.3.4 — September 3, 2026

Buffalo recession countdowns were running fast, sometimes by half; they are rebuilt from the river's own record and now show a typical range.

Not user-facing, recorded for completeness: recession_curve.transit_time_table() + tables in recession_curves.json (Buffalo five; build_tables.py, LOYO in research/recession_eval/VALIDATE_TABLES.md); compute_recession table path in buffalo_assemble.py / hailstone_assemble.py; recession_archive.py stores the served band and settles on its upper end + 48 h; DMZPi range line + wording (one restart); simulator bundle/engine + golden family (2,200 cases). The Buffalo replay rig was corrected to replay the live wave router (baseline v8; it had been grading the retired propagation layer since July). Fault F6 of the 2026-09-01 model review, work packages WP0 + WP1; evidence in research/buffalo_f6/.

Version 2026.9.3.3 — September 3, 2026

The Cossatot neural net was retrained without the two gauges that stopped reporting, and its alerts now come in two levels.

Not user-facing, recorded for completeness: neural_cossatot/final_export.npz (cossatot-pixel-v6b, Cossatot-only, serving tables from its own LOYO OOF), gbm_export.npz (115-feature no-aux ensemble, isotonic + watch/alert operating points), gbm_infer.py v2 tiers, DMZPi Cossatot renderer (one restart). Evidence in research/neural_f5/REPORT.md §3–4.

Version 2026.9.3.2 — September 3, 2026

The neural-net pages now show calibrated crossing chances, and only when the model has earned a watch; the Cossatot page says plainly which of its inputs are missing.

Not user-facing, recorded for completeness: serving tables (isotonic calibration, regime-stratified shrinkage, band scales, watch/alert thresholds) live inside final_export.npz and are applied by neural_infer.py v2; producers v2 (unrounded ledger with raw and served values, --dry-run, degraded/aux status); graders rebuilt on prediction_eval/neural_eval_lib.py (persistence columns, all-hours skill, no rounding slack, Brier, tier episodes, settle retry); shared_qpe/fetch_shared_qpe.py now fetches the top-of-hour MRMS frame (HH:00) instead of the newest 2-minute file (HH:02) — every basin's hourly radar snapshot is now the same hour the models were trained on; DMZPi neural pages, scorecard sections and the neural math page (one restart). Fault F5 from the 2026-09-01 model review; evidence in research/neural_f5/REPORT.md.

Version 2026.9.3.1 — September 3, 2026

The Ponca card's radar estimate is now graded on the scorecard, and its wording follows what the grades showed.

Not user-facing, recorded for completeness: prediction_eval/radar_eval.py (new family, folded into score_all.py), ponca_analog.py 2026.9.3.1 (radar_summary logs motion spread + lead coverage; radar_facts regime wording; projected_context baseline+season cohort), ponca_analog_library.npz v3.1 (adds month + baseline arrays, k-NN data unchanged), DMZPi scorecard section (one restart). Evidence in research/ponca_analog_v2/REPORT.md §6.

Version 2026.9.2.6 — September 2, 2026

The Ponca "AI Rainfall Event Analysis" card was rebuilt: its flood chance is now calibrated, it goes quiet after the crest, and the downstream card now speaks for the wave tracker.

Not user-facing, recorded for completeness: buffalo_dashboard/ponca_analog.py rewrite (episode state, spell-start fix, uncapped k-NN, max-of-calibrated-sources probability, router-narrated propagation, --dry-run/--no-history/--selftest), ponca_analog_library.npz v3 (log-scaled features + trailing dry hours), prediction_eval/ponca_analog_eval.py v2, score_all.py (windowed event-level Ponca scorer, propagation family retired, swallowed exceptions now logged), ledgers purged of 3 synthetic rows and 28 mis-keyed rows (backups kept), SJ_MODEL archived under research/ponca_analog_v2/retired/, DMZPi two card renderers + two scorecard sections (one restart). Fault F4 from the 2026-09-01 model review; evidence in research/ponca_analog_v2/REPORT.md.

Version 2026.9.2.5 — September 2, 2026

Watershed triggers refit; FLOOD now means the gauge, not the rain; Big Piney and Mulberry join the calibrated creeks.

Not user-facing, recorded for completeness: creeks/qbucket_config.json (v3, five creeks, per-bucket thresholds by leaf season where adopted), qbucket_shadow.py v3 (4-day lookback, rolling discharge store, area-weighted Cossatot rain, gauge FLOOD, last-event memory), assemble_and_push.py (Tier-2 freeze, episode-max re-escalation, running-creek episodes, gauge FLOOD messages), drainages.yaml q_ref for haw/hurricane/spirits, DMZPi wording (one restart), the July checkup reminder retired. Fault F3 from the 2026-09-01 model review; report in research/q_bucket_triggers/reports/v3/.

Version 2026.9.2.4 — September 2, 2026

The "chance of rising" forecasts on the Mulberry, Big Piney, Illinois, Cossatot, Richland and Hailstone cards were rebuilt, and the scorecard now grades them honestly.

Not user-facing, recorded for completeness: calibrated_table.json v2 (gauge-quartile strata, horizon_h, margins), calibrated.py v2, empirical_archive.py v2 records and day-file settling, score_all.py probabilistic empirical scorer + flags, digest line, DMZPi scorecard render (one restart), rolling regime factor retired, harness shadow-logger cron disabled. Fault F2 from the 2026-09-01 model review; report in research/empirical_recalibration/v2/.

Version 2026.9.2.3 — September 2, 2026

The Cossatot, Richland and upper-Buffalo (Boxley) rise predictors have been rebuilt and refit on twelve years of storms.

Not user-facing, recorded for completeness: basinlib/physics_engine.py, physics_serve.py, rating.py; engine_v2 blocks in the three calibration files with fit provenance; v1 path retained as automatic fallback; fit, validation and last-45-day replay in research/physics_recal/REPORT.md. Fixes 3 and 6 of fault F1 from the 2026-09-01 model review. ScriptPi-only, no Flask restart.

Version 2026.9.2.2 — September 2, 2026

The nightly AI analysis for the Cossatot, Richland and Hailstone predictors now grades, but no longer tunes.

Not user-facing, recorded for completeness: APPLY_CALIBRATION_UPDATES = False in the three *_analyze.py; select_grading_prediction() reads predictions/YYYY-MM-DD.json; event records gain suggested_changes, graded_source, graded_issued_at. Second fix from the 2026-09-01 model review (fault F1). ScriptPi-only.

Version 2026.9.2.1 — September 2, 2026

Rise predictions on the Cossatot, Richland Creek and the upper Buffalo (Boxley) now read the creek's own baseflow before deciding how much rain runs off.

Not user-facing, recorded for completeness: new shared module basinlib/runoff.py; baseflow_runoff blocks in the three calibration files; fit and leave-one-year-out validation in research/physics_baseflow/REPORT.md; first fix from the 2026-09-01 system-wide model review (reviews/2026-09-01_model_review/). ScriptPi-only, no Flask restart.

Version 2026.8.31.2 — August 31, 2026

The original dark map style is back.

Not user-facing, recorded for completeness: all 6 tile references moved to CARTO's new keyed URL form (rastertiles/dark_all/…?key=…); the key is a public, domain-scoped client-side key recorded in ARCHITECTURE.md §10, gotcha #16 updated, mirrors and research staging copies synced. Same-day follow-up to 2026.8.31.1.

Version 2026.8.31.1 — August 31, 2026

Fixed the "API KEY REQUIRED" watermark on the maps.

Not user-facing, recorded for completeness: all 6 CARTO dark_all tile references (3 in dashboard.py, plus buffalo_tv.html, buffalo_development_tv.html, buffalo_simulator.html) swapped to Esri World_Dark_Gray_Base + _Reference (key-free, {z}/{y}/{x}); attributions updated; Flask restarted; harness mirrors and research staging copies patched. ARCHITECTURE.md gotcha #16, same version.

Version 2026.8.16.1 — August 16, 2026

The AI writer behind the nightly creek analyses and the Buffalo rainfall card got an upgrade.

Not user-facing, recorded for completeness: OLLAMA_MODEL constant flipped qwen3.6:27bqwen3.8:27b in ScriptPi ponca_analog.py, cossatot_analyze.py, richland_analyze.py, hailstone_analyze.py (all .bak'd); DMZPi dashboard.py Hailstone analysis-archive caption updated (Flask restarted). Follows the harness fleet's 2026-08-15 blind cookoff; the inference host serves a single resident model (OLLAMA_MAX_LOADED_MODELS=1), so this also stops the creek callers from evicting the fleet's resident model on the first rain event. ARCHITECTURE.md §9 + item 31 updated to the same version.

Version 2026.8.12.1 — August 12, 2026

New "Data Sources" page — and the site now speaks fluent AI.

Not user-facing, recorded for completeness: new Flask routes /robots.txt, /sitemap.xml, /llms.txt, /data-sources/ driven by a single AGENT_PAGES registry in dashboard.py; schema.org JSON-LD (WebSite/Organization on the landing page, Dataset with USGS/NOAA isBasedOn on all 7 basin dashboards); <meta name="description"> added to those heads. Hidden/dev pages (/admin/, /buffalo/development/, unlaunched simulator) deliberately excluded from all discovery surfaces.

Version 2026.8.9.2 — August 9, 2026

Buffalo rise predictions now understand drought soils.

Not user-facing, recorded for completeness: local_dry_stretch config + frozen-at-anchor stretch and a relaxed dry-tier movement bar in predict_local_rainfall (ScriptPi); full study in research/local_rain_lag_stretch/; replay baseline promoted to v7.

Version 2026.8.9.1 — August 9, 2026

Buffalo TV: river-section levels between St. Joe and Harriet recalibrated.

Not user-facing, recorded for completeness: thresholds updated in SEG_PINNED (build_river_geometry.py), the deployed and repo copies of buffalo_tv_geo.json, and the geo cache-buster bumped to 20260809a in the TV, development-TV, and simulator pages.

Version 2026.7.25.2 — July 25, 2026

The two Illinois River kayak parks are now two distinct parks.

Version 2026.7.25.1 — July 25, 2026

Siloam Springs Kayak Park difficulty rating corrected.

Version 2026.7.17.1 — July 17, 2026

The Buffalo wave tracker learned from its first real test — and now grades itself.

Not user-facing, recorded for completeness: trib reach models refit on real zone QPE with a 48h outcome window (richland_to_st_joe adds a volume/spikiness feature, both trib reaches add a rain-distribution ratio); wave_router gains volume tracking, a QPE-outage ratio guard, and low-base-flow expiry margins; new prediction_eval/wave_eval.py closes the wave record→settle→score loop; simulator engine/bundle rebuilt and golden-gated (1800/0).

Version 2026.7.14.1 — July 14, 2026

The Cossatot gets its own Neural Net Predictions page.

Not user-facing, recorded for completeness: the model is a port of the Buffalo pixel-net architecture (CNN + GRU over a 48×96 1-km window, three gauges multi-task); crossing probabilities come from a companion gradient-boosted ensemble validated at 90.7% recall / 88.0% precision (floatable tier, bootstrap-confirmed) — the full experiment log lives in the operator harness (COSSATOT_NEURAL_PLAN.md).

Version 2026.7.13.3 — July 13, 2026

Easier-to-read text on the Buffalo, Mulberry, and Big Piney cards.

Version 2026.7.13.1 — July 13, 2026

Tune-up after the radar feature's first storm.

Not user-facing, recorded for completeness: the card's AI calls now wait longer before giving up when the inference server is busy with the nightly Buffalo study analysis (the cause of a few carried-over narratives at midnight).

Version 2026.7.11.2 — July 11, 2026

The Ponca AI analysis now tells a continuous story instead of a series of hot takes.

Version 2026.7.11.1 — July 11, 2026

The Ponca AI rain analysis now watches live radar while it's raining.

Not user-facing, recorded for completeness: new ponca_radar_context.py cron on ScriptPi reading MRMS PrecipRate (2-minute radar rain-rate mosaics); radar context logged alongside each analysis for later self-grading on the scorecard pipeline; zero downloads in dry weather.

Version 2026.7.10.4 — July 10, 2026

New "Understand the Math" page: how the neural net actually works.

Version 2026.7.10.3 — July 10, 2026

The Neural Net Predictions now grade themselves on the public scorecard.

Version 2026.7.10.2 — July 10, 2026

The Neural Net Predictions page is now linked from the Buffalo River dashboard.

Version 2026.7.10.1 — July 10, 2026

New (experimental): hour-by-hour Neural Net Predictions for Ponca and Boxley.

Not user-facing, recorded for completeness: the model (pixel-v5) trains on a dedicated GPU machine but serves from the existing prediction Pi as a 530 KB pure-numpy engine, parity-tested against the trained network; hourly cron, isolated from all other site systems; honest "stale/degraded/offline" banners when inputs are missing.

Version 2026.7.9.1 — July 9, 2026

New: Buffalo TV — a full-screen, self-updating view of the whole river.

Not user-facing, recorded for completeness: the long-soft-launched /buffalo/tv/ kiosk page (disk-read, no restart) is now linked from the Buffalo header via an inline-styled pill button in dashboard.py. The /buffalo/simulator/ page was renamed "Buffalo National River Simulator" (test-tube icon removed, EXPERIMENTAL WHAT-IF TOOL banner kept) but remains unlinked and noindex — staged for a later launch.

Version 2026.7.7.4 — July 7, 2026

Ground-dryness alerting extended to 13 more drainages.

Not user-facing, recorded for completeness: q_ref: {parent, multiplier} blocks on 13 drainages in drainages.yaml; assembler scales the parent's Q-bucket threshold and keeps each drainage's own rolling rain window; per-cycle legacy fallback; Tier-2 decisions dual-logged with tier: 2; multipliers = Dave's May 2026 numbers for six creeks, pattern-derived (confirmed) for seven.

Version 2026.7.7.3 — July 7, 2026

Watershed alerts now know how dry the ground is — using the creek itself as the moisture gauge.

Not user-facing, recorded for completeness: v2 Q-bucket engine (qbucket_shadow.py, hardcoded calibration) promoted from a clean 12-day shadow run to authoritative in assemble_and_push.py with per-cycle legacy fallback; legacy status still computed and dual-logged to logs/qbucket_shadow.log; suppress-while-already-running maps to NO_ALERT; post-promotion checkup reminder 2026-07-25; research in research/q_bucket_triggers/ and research/q_bucket_recent_backtest/.

Version 2026.7.5.5 — July 5, 2026

Smarter Buffalo flood warnings: creek-driven floods, cloudburst detection, and all-time-record context.

Not user-facing, recorded for completeness: trib-launch reach models (Richland→St. Joe, Bear Creek→Harriet; LOYO-banded, fit on 268/300 paired events) added to the wave router as pure data — no router code changed; per-zone max-HUC12 rainfall stats added to the data feed; the Buffalo Simulator engine/bundle updated in lockstep (golden gate 1,800/1,800); research in research/buffalo_sim/envelope/ and research/trib_launch_waves/.

Version 2026.7.2.10 — July 2, 2026

Buffalo page: recession countdowns now know what season it is.

Not user-facing, recorded for completeness: monthly p15 floors (circularly smoothed) for Pruitt/St. Joe/Harriet in config recession_floor_monthly (Boxley/Ponca floors never bind their thresholds); seasonal rate multipliers fit and backtested in two variants (result-scaling and anchor-recentering), both regressed vs the live k_now anchor — not shipped; full analysis in research/buffalo_replay outputs/RECESSION_SEASONALITY_REPORT.md.

Version 2026.7.2.9 — July 2, 2026

Buffalo page: the Ponca AI card explained, and recession math corrections.

Not user-facing, recorded for completeness: recession page now imports the deployed curve module directly for lookups (raw-Q bands, edge fallback) so it cannot mis-render again. Separately measured this session, pending review before any model change: paddler folklore confirmed — summer recessions run 1.2–2.3× faster than winter in matched flow bands (evapotranspiration), and lower-gauge baseflow swings seasonally (~600 cfs March vs ~50 September at St. Joe against a flat 76 in config); plan at RECESSION_SEASONALITY_PLAN.md.

Version 2026.7.2.8 — July 2, 2026

Buffalo page: the full math behind every prediction, published.

Not user-facing, recorded for completeness: pages generated from deployed model artifacts by build_math_pages.py (they cannot drift from the running code; rebuild + scp on model changes, no restart); confidence labels read from the live CONFIDENCE_TABLE; horizons page corrected same-day to baseline v6 + baseflow ground-wetness copy after an accuracy audit.

Version 2026.7.2.7 — July 2, 2026

New page: "How far ahead can this system see?" — the honest math behind Buffalo predictions.

Not user-facing, recorded for completeness: page content generated from the replay baseline artifacts by build_horizon_page.py (regenerated on model changes; the route serves a static fragment — scp updates, no restart); buffalo_horizons.json published alongside as machine-readable model facts (groundwork for a future what-if simulator). Lead metrics use causal-run semantics (the unbroken prediction streak into the crossing) after operator review caught the prior-pulse inflation in the draft.

Version 2026.7.2.6 — July 2, 2026

Buffalo page: the river itself now tells the system how wet the ground is.

Not user-facing, recorded for completeness: soil-moisture ladder rung 1 (pre-storm baseflow percentile, month-conditioned, 6–12h lagged against self-wetting; rain-antecedent remains the fallback); NLDAS-2/SPoRT-LIS rungs skipped (no data access, and the river signal may well be the better sensor anyway); replay baseline promoted to v6.

Version 2026.7.2.5 — July 2, 2026

Buffalo page: recession countdowns stay on screen unless a rise is genuinely imminent.

Not user-facing, recorded for completeness: post-audit follow-on — arrival-aware suppression gate (rise_imminent_at) replacing has_upstream_rise; measured on the full replay before deployment (~104 restored countdown-hours/yr, bad-shows +2/yr all during hours the system had no signal anyway).

Version 2026.7.2.4 — July 2, 2026

Buffalo page: steadier behavior during radar and data outages.

Not user-facing, recorded for completeness: fix package F6 (audit finding #14 + #15a) — wall-clock QPE windows, snapshot backfill, staleness surfaced to consumers, MRMS sentinel counter, clock-skew guards, fsync on state files, flood-threshold boundary classification; 19 synthetic outage/gap/skew tests. This completes the 2026-07-01 fresh-eyes audit: all six fix packages (F1–F6) are now live.

Version 2026.7.2.3 — July 2, 2026

Buffalo page: smarter local rain triggers, honest timing windows, and Bear Creek finally counts.

Not user-facing, recorded for completeness: fix package F5 complete (findings #5, #8, #10 from the 2026-07-01 audit) — per-gauge intensity/accumulation trigger ladders (leave-one-year-out fit), measured first-fire→peak windows frozen at event anchor, tributary-zone blending, router-jurisdiction handoff, scale-aware fizzle cuts, recalibrated confidence table; five candidate iterations against the replay gate before deployment; replay baseline promoted to v5. Remaining audit package: F6 (robustness).

Version 2026.7.2.2 — July 2, 2026

Buffalo page: rise predictions now stand down once the rise they predicted has happened.

Not user-facing, recorded for completeness: fix package F5.1 from the 2026-07-01 audit (prediction-consumed event state in local_rain_state; post-peak prediction-hours −88% in replay, flood POD byte-identical); gated on a full 2014–2026 replay before deployment; replay baseline promoted to v4. F5.2–.4 (intensity/accumulation split, measured windows, tributary zones) pending.

Version 2026.7.2.1 — July 2, 2026

Buffalo page: the river now tracks flood waves as they travel downstream — with predicted size and arrival time.

Not user-facing, recorded for completeness: fix package F4 from the 2026-07-01 audit — per-reach crest and travel-time models fit on an 11.6-year event library (leave-one-year-out validated), a stateful wave tracker replacing the trend-gated single-hop mechanism, target-anchored alert tiers tuned on the full replay (all deploy gates passed vs baseline v3), Richland/Bear Creek gauges now feed the lower-reach predictions. Packages F5–F6 pending.

Version 2026.7.1.9 — July 1, 2026

Buffalo page: rise predictions now judge storms against how wet the ground was before the rain.

Not user-facing, recorded for completeness: fix package F3 from the 2026-07-01 audit (pre-storm antecedent, precip_7day_total_in preserved for the propagation model's fitted definition, all 8 zones exported incl. Richland); gated on a full 2014–2026 replay before deployment (St. Joe flood detection unchanged, no Boxley/Ponca regression); replay-rig baseline promoted to v3. Packages F4–F6 pending.

Version 2026.7.1.8 — July 1, 2026

Buffalo page: the downstream "river rise" card now stays up until the wave actually arrives — and won't inflate its numbers late in an event.

Not user-facing, recorded for completeness: this is fix package F2 from the 2026-07-01 fresh-eyes audit (event-frozen SJ_MODEL features + pre-storm antecedent + wave-in-transit hold + fit-domain gate); validated over 173 historical Ponca events (antecedent error vs fit: +1.63″ → +0.17″ median on big events; late-event regime flips 38% → 3.5%) plus a live end-to-end synthetic flood test. Packages F3–F6 pending.

Version 2026.7.1.7 — July 1, 2026

Buffalo page: an earlier flood-risk banner, and confidence labels you can take at face value.

Not user-facing, recorded for completeness: rise predictions are now graded against the timing window of the driver that set the category (grader + replay-rig change, removes ~6 pts of unearned timing credit from the accuracy baseline); antecedent tier boundaries are read from config instead of being hardcoded; this ships fix package F1 from the 2026-07-01 fresh-eyes audit of the Buffalo prediction chain (research/buffalo_replay/outputs/FRESH_EYES_AUDIT.md) — packages F2–F6 pending.

Version 2026.7.1.4 — July 1, 2026

The site is faster, and the Illinois River forecast card is now complete.

Not user-facing, recorded for completeness: USGS gauge fetching was consolidated into batched requests (~93% fewer API calls per 15-min cycle, with per-site retry fallback preserved); the per-basin script clones for Mulberry, Big Piney, Cossatot, and Richland now run through one shared library (/home/dave/basinlib/ shims — equivalence-verified before deploy) with a retry hardening all four inherit; a git baseline of both Pis' code/config now lives on the operator harness (creekintelligence/mirrors/); the dashboard gained an mtime-keyed file cache, explicit Cache-Control headers, and a threaded gunicorn worker (-w 1 --threads 4). See ARCHITECTURE 2026.7.1.4.

Version 2026.7.1.2 — July 1, 2026

The Buffalo River Watershed Study has wrapped up after 122 days — a month longer than planned — and its page is now a finished, browsable archive.

Not user-facing, recorded for completeness: the study's nightly Claude Opus narrative stream was retired on 2026-07-01 (cron study_daily_analysis.py --no-opus), ending the project's only per-night external API cost; the local qwen3.6:27b stream continues a no-cost nightly pulse on ScriptPi for periodic hand-review. /study/ pages are frozen at Day 122 (June 30, 2026); the concluded banner + per-page frozen-archive notes are served from dashboard.py (study_index/study_hypothesis/study_daily). See ARCHITECTURE §4.7 + §16 item 16 (version 2026.7.1.2).

Version 2026.7.1.1 — July 1, 2026

Arkansas Creek Intelligence has its own web address now: arcreekintel.com.

Version 2026.6.30.1 — June 30, 2026

New page: the Illinois River and the Upper Illinois Water Trail.

Not user-facing, recorded for completeness: new ScriptPi basin /home/dave/illinois/ (config + assembler + multi-gauge fetch / read-qpe / weather+QPF, cloned from Big Piney with Buffalo's signal_for multi-gauge pattern grafted on), cron every 15 min → illinois_output.json. Full history collected to the harness research/illinois_*: 8 USGS gauges' complete records + hourly MRMS per-HUC12 QPE 2014→present over the 21 HUC12s above Watts. Empirical illinois_below block trained on the above-Hwy-16 basin (calibrated_table.json + lookup index + empirical_config.yaml). Upstream→Hwy-16 propagation lags from storm-pulse cross-correlation (Savoy 6–8 h, Osage 8–13 h, Mud 11–16 h; Watts +3–5 h downstream). New DMZPi routes /illinois/ + /illinois/historical/. See ARCHITECTURE §4.9c (version 2026.6.30.1).

Version 2026.6.29.7 — June 29, 2026

The rainfall forecast on the five creek pages now gives you an honest percentage instead of a vague "rise likely."

When the new forecast says a percentage, the creek rose that often: said under-10% rose 4%, 10-25% rose 17%, 25-50% rose 49%, over-50% rose 80%

Not user-facing, recorded for completeness: new calibrated engine (empirical_forecast/calibrated.py + calibrated_table.json — denominator-corrected P(rise | rain, soil) over the full 2014-2026 MRMS+USGS record) feeds the headline in empirical_predict.py, with a per-basin warm rolling factor off the local settled archives. Boxley (no page) unchanged; empirical stays off Facebook per standing choice (Watersheds only). Full research + validation in the harness research/empirical_recalibration/ (REPORT.md). ARCHITECTURE §4.13/§16-11 updated.

Version 2026.6.29.6 — June 29, 2026

Added a privacy policy page.

Version 2026.6.29.5 — June 29, 2026

Our forecast "report card" got clearer and more honest, and a couple of predictions were tuned based on what it's been telling us.

Not user-facing, recorded for completeness: this came out of the weekly scorecard / Signal-digest review. score_all.py now emits a per-predictor median error, min-n-gates the recession rows (MIN_RECESSION_GRADED), and adds a recession timing-bias flag; ponca_analog.py decays the dominant-driver flood-floor toward the raw k-NN once Ponca is past crest (raw values still logged for backtest); big_piney_assemble.py had recession_baseline support restored to its ported compute_recession (the port had dropped it) and above_longpool was given recession_baseline: 2.0. The empirical engine's over-warning (also visible on the scorecard) is a deeper rebuild blocked on an offline tool — tracked, not yet changed. See ARCHITECTURE §4.13 / §16-11 (version 2026.6.29.5).

Version 2026.6.29.4 — June 29, 2026

Creek Intelligence is now on Facebook — follow Arkansas Creek Intelligence for automatic creek alerts and a weekend forecast.

Not user-facing, recorded for completeness: new creek-social toolset on the harness server — creek_social.py mirrors the watershed alerts to Facebook (fully decoupled from ScriptPi: it reads the same predictor_output.json the dashboard already produces, replays the alert engine's new/escalation logic against its own state, and posts via the Facebook Graph API), weekend_forecast.py (Friday cron; recession-based "holds through the weekend" filter), and announce.py (auto-posts a short feature announcement to the page whenever we ship a user-facing change — a new shipping convention now in CLAUDE.md §8 + ARCHITECTURE §4.2). Posts are deliberately link-free and bot-signed.

Version 2026.6.29.2 — June 29, 2026

The Buffalo Study pages now correctly report the June 22 flash flood — it had been mistakenly logged as a "data gap"

Not user-facing, recorded for completeness: two fixes to the nightly Buffalo Study analyzer (study_daily_analysis.py) plus a data backfill. (1) The transfer-ratio validator was QPE-gated — its 2,000 cfs/in cap had unconditionally false-rejected the model's correct 06-22 headwater numbers (Ponca 2,756 cfs/in matched the deterministic truth card), silently dropping that night's calibration; the cap now applies only on low-rainfall days, where the frame-override artifact it guards against actually occurs. (2) The Opus output cap was raised 40000→56000 after the daily+thinking+hypothesis rewrite pinned at the cap for five straight nights (06-21..06-25, incl. the flood), freezing the rolling hypothesis. (3) The 06-22 study-record flood was backfilled into both the qwen knowledge.md calibration tables and the Opus hypothesis.md (now live on /study/), from the truth card + the existing night-of Opus daily; logged in analysis/curation_audit.md. The qwen-only (--no-opus) cutover stays deferred — even post-fix, qwen still drops a gauge on the complex flood night. See ARCHITECTURE §4.7 + §16 item 16.

Version 2026.6.29.1 — June 29, 2026

More accurate downstream forecasts: the St. Joe / Grinder Ferry propagation card no longer overshoots when the rain falls up high

Not user-facing, recorded for completeness: the St. Joe magnitude in ponca_analog.py build_propagation() was a fixed basin-wide amplification (≈2.9× Ponca) applied unconditionally; on an upper-concentrated event the wave routes through the dry 1,342 km² intervening basin and attenuates (~0.9×), so the constant over-predicted 2–4× (6-22: card 19–33k, actual 7.9k). Recalibrated against 152 historical Ponca events (2014–2026) from the local buffalo_huc_qpe + buffalo_gauges archives (research/buffalo_propagation_calibration/, with REPORT.md + figures). St. Joe is now a rainfall-distribution + antecedent conditioned log-linear model (SJ_MODEL) reading the upper-vs-intervening qpe_24hr ratio and the intervening 7-day antecedent — all already in buffalo_output.json — with a crest band floored at the routed wave (0.85×) and the reach's current flow. Leave-one-out CV cut the upper-concentrated bias from +46% to ≈0; out-of-sample on 6-22 it predicts St. Joe 8.5k–11.9k–16.9k (lower bound ≈ actual). Pruitt kept at ~1.13× (rain-insensitive, 11-yr confirmed); peak-to-peak lags re-derived. ScriptPi-only — render_propagation_card() reads only the kept crest-band fields, so no DMZPi change. See ARCHITECTURE §4.6 + §16 item 15.


Version 2026.6.24.8 — June 24, 2026

New: a 24-hour gauge chart on the Cossatot, Richland, and Hailstone pages

Not user-facing, recorded for completeness: each physics-basin assembler now emits gauge.readings_24h (a [[epoch, value], …] array of the last 24 h, ~96 pts at 15-min) via a build_readings_24h() helper — Cossatot/Richland stitch yesterday+today's daily height files to span midnight, Hailstone uses its already-rolling cfs buffer from creeks/gauge_data.json. dashboard.py gained render_hydrograph(), an inline-SVG renderer (tier bands + line + peak + "now" dot, fully self-contained, no chart library), placed under the gauge card on all three pages. Display-only, no DMZPi compute (same contract as the HTML tables and Leaflet maps). See ARCHITECTURE §5.4.


Version 2026.6.24.7 — June 24, 2026

The Scorecard now flags what needs attention — and reports its own drift weekly

Not user-facing, recorded for completeness: score_all.py now computes a flags array (min-n-guarded: empirical over-warn ≥70% no-rise at n≥20, band inversion above_p75 vs. p50_to_p75, physics within-±20% <50% at n≥15, Ponca override-vs-raw flood-Brier gap, recession never-reached ≥60% once graded) into scorecard.json; /scorecard renders the "Needs attention" box from it. New prediction_eval/scorecard_digest.py (weekly cron Mon 07:30 Central) reads the scorecard, formats a per-family accuracy snapshot + the flags, and sends a Signal DM via the existing signal_config.yaml (same signal-cli plumbing as scp_alert / data_age_alert). This is Phase 3 of the prediction-logging initiative — the self-checking / feedback half. The harness-consolidation refactor was deliberately deferred (a risky change to working cron code with no user benefit). See ARCHITECTURE §4.13 + §16-14 + §7.2.


Version 2026.6.24.6 — June 24, 2026

Scorecard now also tracks the downstream-propagation forecast (recording started)

Not user-facing, recorded for completeness: ponca_analog.py now folds the per-reach propagation block (lag window, crest band, bump/no-bump flag, current flow) into each ponca_analog_history.jsonl record — the record step the predictions had been missing. New prediction_eval/propagation_eval.py settles each matured event two-stage: it finds Ponca's actual peak from buffalo_study/data/gauges, then for Pruitt (07055680) and St. Joe (07056000) grades the bump/no-bump call (crest ≥ 1.25× baseline & +50 cfs), the magnitude-band hit, and the timing window relative to Ponca's peak — using the cycle nearest Ponca's peak as the representative prediction. A --selftest (10 checks) validates the logic. score_all.py folds the settle in and adds a propagation block; /scorecard renders it; Coverage flips the propagation forecast to "graded." This completes Phase 2 of the prediction-logging initiative — every live predictor family is now recorded and graded. See ARCHITECTURE §4.13 + §16-14.


Version 2026.6.24.5 — June 24, 2026

The Ponca "AI Analysis" cards now read like a person — and the numbers make sense

Not user-facing, recorded for completeness: in buffalo_dashboard/ponca_analog.py, both qwen system prompts were slimmed and rewritten for a casual-paddler audience (no jargon, fact-fed, model reasons rather than recites) with one hard rule — never state a peak/crest at or below the current flow. build_propagation() now models est_crest = reach_current + (Ponca rise above baseline) × amplification, floored at the reach's current value (was Ponca_flow × amp absolute, which fell below a downstream gauge's own flow whenever Ponca was small relative to that reach's drainage), plus a significant / added_bump_cfs gate so small pulses read "minor bump, no noticeable change." A qualitative 6-hour trend word (holding steady / creeping up / rising steadily / rising fast / easing down / dropping fast) keeps the model from overstating a slow creep. The deterministic override floors (upstream Boxley wave, heavy-rain, watershed-feeder alerts) are unchanged — only their narration. See ARCHITECTURE §4.6.


Version 2026.6.24.4 — June 24, 2026

Scorecard now tracks the Buffalo rise engine (recording started)

Not user-facing, recorded for completeness: new prediction_eval/buffalo_predictions_archive.py (modeled on recession_archive.py) — --record on a 15-min cron (:09,:24,:39,:54, debounced) logs each active predictions[gauge].combined_category from buffalo_output.json (with timing window + current value) to buffalo_ledger/<date>.jsonl; --settle grades vs. the gauge's actual discharge rise over [generated_at, timing_high + 18 h] (rose if peak ≥ 1.25× start AND ≥ +30 cfs → records rise magnitude, hours-to-peak, timing-in-window; else no_rise / censored); --score aggregates. A --selftest (14 synthetic checks) validates the logic since production is dry. score_all.py folds the settle in and adds a buffalo_rise block; /scorecard renders it; Coverage splits Buffalo into "rise predictions: graded" and "flood_risk + propagation_alerts: not yet logged." Magnitude is captured per category (categories derive from rainfall, not gauge rise) so the rainfall→rise mapping calibrates over time. ARCHITECTURE §4.13 + §16-14 + §7.1.


Version 2026.6.24.3 — June 24, 2026

Scorecard now grades the Ponca "AI Rainfall Event Analysis"

Not user-facing, recorded for completeness: new prediction_eval/ponca_analog_eval.py reads the producer's append-only ponca_analog_history.jsonl (read-only) and grades each armed call against Ponca 07055660's max discharge over a forward 36 h window (from buffalo_study/data/gauges), writing ponca_analog_settled.jsonl (idempotent, keyed by generated_at). It grades the k-NN class/peak (modal vs. actual class, median_peak_cfs MAE, p25–p75 band hit) and the flood Brier for flood_risk_pct (override-floored) vs. raw_flood_risk_pct (raw k-NN, on the common sample). score_all.py calls settle() inline (no extra cron) and adds a ponca_analog block to scorecard.json; the /scorecard route renders it; the Coverage table flips Ponca from "logged, not graded" → "graded." See ARCHITECTURE §4.13 + §16-14.


Version 2026.6.24.2 — June 24, 2026

Retired the experimental neural-net (LSTM) predictor

Not user-facing, recorded for completeness: the :14 nn_predict.py inference cron was commented out, the dead nn_prediction read/embed was removed from cossatot_assemble.py + richland_assemble.py (the key no longer appears in their output JSON), the stale "LSTM forecast" wording was dropped from the Richland page's social/meta description, and all LSTM-only artifacts (nn_predict.py, nn_predict_debug.py, nn_alert.py, nn_alert_state.json, nn_output.json, both model_epoch100.pt weight files) were moved to /home/dave/creeks/retired_lstm/ (reversible — a README there documents how to resurrect). The HUC12 masks under /home/dave/models/ were kept — they are read by the live Richland physics QPE reader and the map-polygon builder, not just the LSTM. Full record in ARCHITECTURE.md §4.10 + §14 Gotcha #10. Side benefit: one less hourly job on the 2 GB ScriptPi.


Version 2026.6.24.1 — June 24, 2026

New: a Prediction Scorecard page — see how accurate the site's forecasts have actually been

Not user-facing, recorded for completeness: a new ScriptPi subsystem /home/dave/prediction_eval/ (score_all.py, cron 01:30 daily) reads the settled prediction archives for three already-self-recording families — physics (<basin>/predictions/), empirical (<gauge>_empirical_predictions/), and recession (recession_eval/ledger/) — and writes one display-ready scorecard.json, SCP'd to DMZPi. Physics is scored on peak MAE / bias / within-±20%; empirical on hit-rate (rose to ≥ called tier = verified + missed_higher) vs. no-rise rate (no_change), broken out by basin, confidence, and rainfall percentile band (which surfaces the known selection bias — the above_p75 band currently shows a higher no-rise rate than p50_to_p75); recession reuses recession_archive.score(). The DMZPi /scorecard route renders the JSON read-only (no compute on DMZPi; validated via the venv test_client before the gunicorn restart). This is Phase 1 of a system-wide "log every prediction, grade it later" initiative — audit + plan in creekintelligence/research/prediction_audit/. The still-ungraded predictors (Ponca analog, downstream propagation, the Buffalo per-gauge rise engine, the retired LSTM) are listed as not-yet-graded in the page's Coverage table and are Phase 2.


Version 2026.6.23.1 — June 23, 2026

Recession countdowns are now realistic — and honest about how sure we are

Not user-facing, recorded for completeness: the timing model is a per-gauge master recession curve — a flow-dependent decay rate k(Q) learned from USGS history, integrated to a transit time and event-anchored to the current observed rate (clamped 0.5–2×) — in /home/dave/recession_eval/recession_curve.py + recession_curves.json (9 gauge curves), imported by every *_assemble.py compute_recession() with the old single-exponential kept as a fallback. The St. Joe/Harriet suppression uses a config recession_baseline (empirical floor, ~p10) as the decay asymptote so any threshold below it returns no countdown. Confidence is now derived from each prediction's relative spread + horizon (recession_curve.confidence_level). A prediction-evaluation framework (recession_archive.py, on cron) now records every recession countdown to a ledger and settles it against actuals to score accuracy over time. Threshold keys were standardized to too_low/low_floatable/optimal across all configs/assemblers. Full design, history pulls, and analysis live in creekintelligence/research/recession_eval/.


Version 2026.6.22.2 — June 22, 2026

Clearer recession countdowns + gauges stop briefly greying out

Not user-facing, recorded for completeness: the three height/CFS recession cards were consolidated into one shared render_recession_card() helper in dashboard.py (they had drifted into near-duplicate blocks edited in parallel). Hailstone's recession model now also emits hours_to_high (time to fall to the 2000-cfs Above-Recommended boundary) in hailstone_output.json. The gauge-fetch resilience lives in creeks/fetch_gauges.py: fetch_stream_readings() retries transient empty USGS responses, and main() carries forward the previous good reading (new carried_forward flag, 90-minute cap) so one failed pull no longer greys a gauge.


Version 2026.6.22.1 — June 22, 2026

Ponca rainfall card now leads with the trusted signals + new downstream propagation card

Not user-facing, recorded for completeness: the rainfall card now reads the live creek-alert state (read-only) and the Buffalo physics propagation/flood signals already in buffalo_output.json; the propagation card's lag/amplification constants come from the 12-year USGS archive cross-checked against the Buffalo Study's calibrated 2026 event log, and are antecedent-conditional. Both narratives are produced by the local qwen3.6:27b model — the propagation one is a separate focused prompt embedded in ponca_outlook.json.


Version 2026.5.24.1 — May 24, 2026

Ponca Gauge historical event reporting + experimental AI rainfall event analysis

Not user-facing, recorded for completeness: built a local archive of MRMS radar rainfall for all 37 Buffalo HUC12 sub-watersheds (2014→present, two source archives) plus full USGS history for all 7 Buffalo gauges, which back the historical report and the analog evaluator. The evaluator (buffalo_dashboard/ponca_analog.py, hourly cron :20, flock) runs a numpy k-NN match against a shipped analog library and calls the local model only when its call changes (~10 inferences per event), writing ponca_outlook.json for the dashboard to render.


Version 2026.5.21.1 — May 21, 2026

PayPal-based donation support


Version 2026.5.20.1 — May 20, 2026

Hailstone reaches feature parity + nightly AI analysis moves to local inference

Underlying schema, infrastructure, and unit-bug cleanups (not user-facing but recorded for completeness): new hailstone_predict.py + hailstone_calibration.json + hailstone_predictions_archive.py, three new cron entries (:17 Hailstone QPE consumer, :21 Hailstone physics predictor, 23:55 Hailstone nightly analyzer), fixed a 25.4× unit error in the new Hailstone QPE consumer (had been treating already-in-inches shared QPE values as mm and dividing again — antecedent moisture was reading DRY when the basin was actually NORMAL/WET), exposed recession_k/recession_h_base on Richland's gauge output (latent gap that left the physics predictor falling back to a flat baseline instead of an exponential decay), and added a per-supplement file-naming convention (YYYY-MM-DD-suffix.md) for retrospective analyses written outside the nightly cadence.


Version 2026.5.15.1 — May 15, 2026

Stale-Data Warning Banners Across All Mobile Dashboards


Version 2026.4.21.1 — April 21, 2026

Creek Page Title & Buffalo Card Standardization


Version 2026.4.15.1 — April 15, 2026

Buffalo River Dashboard


Version 2026.4.7.1 — April 7, 2026

Richland Creek Dashboard & Neural Net Predictors


Version 2026.3.10.1 — March 10, 2026

Landing Page & Route Restructure - New landing page at / — mobile-first dark theme with site title, orientation blurb, live status summary, and dynamic creek condition cards - Creek cards show all Optimal gauges (green) or, if none, all Low but Floatable gauges (yellow); "All Quiet" message when no creeks are runnable - Status summary line shows gauge and watershed condition counts with color-coded text - Tool navigation grid — four main buttons (Gauges, Watersheds, Cossatot Predictor, Buffalo Study) plus secondary links (Guide, Changelog, Suggest a Creek) - Gauge table moved from / to /gauges/ — all table logic unchanged - Temporary legacy link on landing page points to /gauges/ for returning users - New /guide/ placeholder page — "Coming Soon" with back link to home - Back link audit — watersheds, changelog, study, and suggest pages link back to /gauges/; error pages and Cossatot nav link to / (landing) - Last updated timestamp on landing page converted from UTC to Central time


Version 2026.3.9.1 — March 9, 2026

Four-Color Unification - Unified color language across creek levels, prediction status, and recent rain: Red (nothing) / Yellow (maybe) / Green (go) / Blue (lots) - Prediction status expanded to four tiers: No Alert (red), Watch (yellow), Warning (green), Flood (blue) - Recent Rain column now shows 7-day precipitation total in inches with color-coded background, replacing category labels (MOIST/SEMI-DRY/DROUGHT) - Recent Rain multipliers updated: <0.25" = 1.4x trigger, <0.75" = 1.2x, <1.50" = 1.0x, ≥1.50" = 0.9x - FLOOD status triggers at 200% of effective trigger threshold — indicates exceptional rainfall - Micro-creek lag display: Drainages with 0-1 hour lag now show "NOW" instead of numeric range - Signal alerts updated with lag-aware messaging (micro-creeks show "NOW", mainstem rivers show hours)


Version 2026.3.6.1 — March 6, 2026

Drainage Trigger & Timing Calibration - Adkins: 2.0" / 4hr → 2.5" / 6hr - Boen Gulf: 2.0" / 4hr → 2.5" / 6hr - Upper Buffalo: window 6hr → 12hr, lag 4-6hr → 6-8hr - Beech Creek: 1.5" → 2.0" - Upper Kings: 1.5" → 2.0" - Osage: 1.5" / 4hr → 2.5" / 6hr - Richland Main: window 6hr → 12hr - Falling Water: 1.5" / 6hr → 1.75" / 12hr - Upper Cossatot: window 6hr → 12hr - Upper Big Piney: window 6hr → 24hr, lag 10-12hr → 12-16hr - EFLB: 1.5" → 2.0" - Pine Creek OK: window 6hr → 12hr


Version 2026.3.5.3 — March 5, 2026

Cosmetic Updates - Renamed "DRY" status to "QUIET" on the Watersheds page and in alert bar logic (same red styling, new CSS class .st-quiet). - Renamed "Conditions" column to "Recent Rain" and changed from styled badge spans to full-cell background coloring (matching the Status column style). - Updated status_colors/status_text_colors dicts to use "QUIET" key.


Version 2026.3.5.2 — March 5, 2026

Sticky Status Hold - WATCH and WARNING statuses now hold for the duration of a drainage's lag time plus a 2-hour buffer before clearing, preventing premature status downgrade before water reaches the gauge.

Antecedent Dryness System - New "Conditions" column on Watersheds page showing MOIST, SEMI-DRY, or DROUGHT based on recent rainfall history. - Trigger thresholds automatically increase during dry conditions: +15% for SEMI-DRY, +30% for DROUGHT. - Dryness is computed from rolling 7-day and 30-day precipitation totals per drainage. - Trigger column on Watersheds page now shows the effective (adjusted) trigger value.

Cossatot Drainage Update - Upper Cossatot trigger raised from 1.00" to 1.25" in 6 hours. - Upper Cossatot lag time updated from 3-6 hours to 8-10 hours based on observed March 5 event.

YAML Sync - Creek and drainage definitions now automatically sync from ScriptPi to DMZPi every 15 minutes, ensuring single source of truth.

Changelog - Added this changelog page, accessible from the footer of both dashboard pages.


Version 2026.3.5.1 — March 5, 2026

Baseline version. All prior changes consolidated.

WATCH / WARNING Terminology - Renamed "TRIGGER" status to "WARNING" to align with NWS conventions. - WATCH threshold raised from 50% to 75% of trigger value to reduce false positives.

Color Scheme Standardization - Watersheds page: RED = DRY (no go), YELLOW = WATCH (maybe), GREEN = WARNING (go time). - Main creek page: Watershed radar column and alert bar colors match the same scheme. - Alert bar is now green for WARNING, yellow for WATCH-only.

Drainage Trigger Updates - Bobtail Creek: 1.5" → 2.0" in 6hr - Long Devils Fork: 1.5" / 4hr → 2.5" / 6hr - Big Devils Fork: 1.5" / 4hr → 2.5" / 6hr - West Fork Shop Creek: 2.0" / 4hr → 2.5" / 6hr - Thomas Creek: 2.0" / 4hr → 2.5" / 6hr