Get Full Government Meeting Transcripts, Videos, & Alerts Forever!
Get email alerts on the Hydrology Reach Dataset topic
No spam. Unsubscribe anytime.
Upstream Tech presents statewide reach-scale daily historical flow datasets to CEF work group
Summary
Upstream Tech presented new statewide reach-scale daily flow datasets to the California Environmental Flows Work Group, describing paired unimpaired and actual daily flow predictions for about 20 years at NHDPlus reach scale.
Get email alerts on the Hydrology Reach Dataset topic
No spam. Unsubscribe anytime.
Upstream Tech presented new statewide reach-scale daily flow datasets to the California Environmental Flows Work Group during a virtual meeting, describing two products: daily unimpaired-flow predictions and daily actual-flow predictions covering about 20 years at reach scale on NHDPlus flow-lines.
The datasets — one intended to represent natural (unimpaired) flows and one representing observed/managed (actual) flows — were trained using machine-learning approaches and validated against held-out gauges in California, presenters said. Mustafa Alkurdy, a machine-learning modeling engineer at Upstream Tech, said the project produced reach-scale predictions on thousands of NHD flow-line segments and that the team downscaled subbasin predictions to reach-level outputs.
Why it matters: Managers, researchers and practitioners told the group the new products should help fill observational gaps where gauges are missing, support applications such as functional-flow calculations and inform planning for restoration, reservoir operations and water-allocation studies.
How the datasets were produced and validated Upstream Tech described a multi-step approach: (1) train base models on many gauged sites to learn basic hydrologic behavior; (2) route predictions through subbasins tuned to local flow dynamics; and (3) downscale routed predictions to NHD reach units. Laura Reed, of Upstream Tech’s hydro-forecast team, said, “we've actually delivered the 2 final datasets of daily unimpaired flow predictions and … actual flow predictions for about 20 years, and those are statewide.”
The team trained the actual-flows model on roughly 1,770 gauged sites across the conterminous U.S. and the unimpaired model on about 500 reference (unimpaired) gauged sites, Alkurdy said. For evaluation they reserved a set of California gauges the models had not seen during training and compared predicted hydrographs to observations at those locations.
Scope and data footprint Presenters said the prediction grid began with about 80,000 NHD flow-lines and, after filtering, produced predictions for roughly 30,000–50,000 reach segments for the full 20-year period. Results and maps are planned for public hosting; presenters said TNC (The Nature Conservancy) will host an interactive map and that a final report and a peer-reviewed publication are expected.
Performance: where the models perform best and where they are weaker Upstream Tech reported stronger performance in snowmelt-driven basins and large basins, and generally better results for base flows than for flashy high flows. Reed and Alkurdy showed validation summaries using metrics including bias, Nash–Sutcliffe efficiency (NSE), Pearson correlation and normalized RMSE. In discussion Pauline (participant) asked about NSE and RMSE; presenters explained NSE ranges up to 1 (perfect) and that RMSE is in flow units and is normalized by the site mean to allow comparison across sites.
Areas of lower confidence include small, dry, highly intermittent basins (common in southern California), and locations downstream of regulated dams where reservoir operations strongly alter flows. Alkurdy said the actual-flows model performs well below gauges because the model uses upstream gauge observations to route downstream predictions, producing narrow confidence bounds where observations are available. By contrast, locations directly below dam outlets — where operational releases are not represented in the training inputs — showed wider uncertainty and larger errors.
Representation of human impacts Presenters said human-impairment signals (dams, canals, impervious area) were represented through available datasets (for example EPA StreamCat and other catchment-scale variables), but the models do not include operational rules or time-varying diversion schedules. As Mustafa Alkurdy put it, the training data can help the model learn patterns produced by regulated and diverted systems, but the team cannot subtract explicit diversion volumes or reservoir release schedules unless those data are available. Presenters noted the model tends to widen its confidence bounds where upstream regulation is known to exist.
Use cases and next steps Participants asked whether the unimpaired dataset could feed mechanistic reservoir-operations models; presenters agreed that unimpaired inflow predictions could be used as input to reservoir-scenario modeling and then routed downstream to produce scenario-based downstream flow time series. Reed said the team plans to make the datasets public and to finalize the project report while TNC hosts the interactive map.
Availability Meeting participants were told the data and final report will be posted publicly; a participant noted that Ron (not on the call) had indicated the data would be accessible by the end of 2025 and that the project team would share availability with the work group when ready.
Ending Presenters encouraged follow-up questions and provided contact emails and said they would share slides with the group. Upstream Tech’s work was presented as a complement to existing gauges and to be used in concert with local mechanistic modeling where operational reservoir or diversion data exist.

