WDG JournalResearch

New research

The latest papers on water and machine learning

Peer-reviewed articles and preprints on floods, drought, groundwater, rainfall and water quality that use AI, from the past 21 days. Every title links to the original.

Updated 10 October 2026 · 60 papers

9 October 2026

8 October 2026

International Journal of Digital Earth

Parameter-efficient fine-tuning of SAM for SAR flood mapping: comparative performance and encoder–decoder representation analysis

Ziming Wang, Yigong Hu, Jeff Neal, Peter M. Atkinson and 1 other

Read abstract

Synthetic aperture radar (SAR) enables flood mapping under cloud and adverse weather, but segmentation remains challenging across heterogeneous environments. Deep-learning models can improve flood delineation, yet their dependence on task-specific training data limits transfer across regions and events. This motivates adapting pretrained vision foundation models to SAR despite the domain gap from RGB imagery. Among these models, we use the Segment Anything Model (SAM) to compare five parameter-efficient fine-tuning (PEFT) strategies, comprising Adapter, Prefix tuning, Prompt tuning, BitFit and Low-Rank Adaptation (LoRA), under a consistent prompting protocol across four held-out test scenes and repeated training configurations. All PEFT strategies outperform unadapted SAM, with flood-class IoU ranging from 32.26% to 57.94% and F1 scores from 46.72% to 69.17%. A hybrid LoRA-BitFit configuration achieves the highest overall IoU and F1, while LoRA favours recall and BitFit precision. Encoder–decoder representation analysis further shows that LoRA and LoRA-BitFit strengthen flood/non-flood separability in the image encoder, whereas LoRA-BitFit maintains high attention-weight separability across several later mask-decoder stages. These results show that SAR adaptation depends not simply on the number of trainable parameters, but on which model components are modified and how adaptation is distributed across the encoder–decoder architecture.

Abstract reproduced under CC BY · doi:10.1080/17538947.2026.2744033

Land

A Pilot Machine-Learning Framework for Cumulative Antecedent Rainfall Threshold in Landslide Assessment Along Jalan Simpang Pulai–Blue Valley, Malaysia

Rabeah Adawiyah Hashim, Faizah Che Ros, Shahrum Shah Abdullah, Aniza Ibrahim and 4 others

Read abstract

Rainfall-induced landslides along Malaysian highway cut slopes are commonly evaluated using empirical thresholds, with limited consideration of cumulative antecedent rainfall. This study evaluates a machine-learning rainfall threshold for five documented shallow-slope failures along Federal Route FT185 (Jalan Simpang Pulai–Blue Valley), Perak, Malaysia, using 520 site-days of daily rainfall records. Cumulative rainfall was calculated over seven windows (1, 3, 5, 7, 14, 21 and 30 days), together with a short-burst intensity feature and a wet-day count and six classifiers (logistic regression, Random Forest, Gradient Boosting, k-nearest neighbours, a support vector machine and XGBoost) were trained under class-weighted, leave-one-site-out cross-validation. Random Forest achieved the best discrimination (ROC-AUC = 0.817), with 80% sensitivity and 90% specificity. Antecedent rainfall consistently ranked above same-day rainfall, indicating the importance of longer-duration rainfall accumulation in the present dataset. An interpretable two-feature logistic regression boundary (14-day antecedent rainfall vs. same-day rainfall) and Random Forest partial-dependence saturation points (approximately 203, 330 and 447 mm for the 14-, 21- and 30-day windows) were derived and combined into a two-tier watch/warning rainfall-probability threshold. The results provide preliminary, site-specific rainfall thresholds for the investigated corridor within a pilot methodological framework requiring prospective validation with additional independent failure events before operational application.

Abstract reproduced under CC BY · doi:10.3390/land15101892

Water

Hierarchical Rainfall-Intensity-Aware Hourly Precipitation Merging Based on Tree-Model Routing and LSTM Conditional Regression: Spatiotemporal Generalization and Hydrological Utility

Xinlin Zhang, Jinbao Liu

Read abstract

In mountainous regions with uneven gauge coverage, satellite precipitation errors are amplified by nonlinear rainfall–runoff processes, while conventional hourly merging often relies on a single continuous regression that inadequately handles zero inflation and intensity heterogeneity. We propose a hierarchical intensity-aware framework comprising a wet/dry gate, a frequency-matched four-class intensity router, and a shared long short-term memory (LSTM) network with class-conditional outputs. In the upper Fujiang River basin, GPM, CMORPH, ERA5-Land and topographic variables were used as predictors; models were trained on 2010–2013, with 2014 retained for temporally held-out validation across point and areal scales, spatial cross-validation and streamflow simulation. Progressive ablation shows that intensity stratification drives the main point-scale gain (Kling–Gupta efficiency (KGE), 0.22 → 0.43), while temporal modeling improves held-out catchment-event performance (KGE 0.75). Oracle diagnosis identifies intensity routing as the main remaining bottleneck, and attribution and source-ablation analyses show that data-source value varies with prediction stage and evaluation scale. Hydrologically, merged precipitation raises the overall Nash–Sutcliffe efficiency (NSE) from −0.13 (GPM) and −0.21 (CMORPH) to 0.40 and reduces absolute peak bias from 52–55% to 32%. Routing discrimination remains the principal residual limitation under the present architecture.

Abstract reproduced under CC BY · doi:10.3390/w18192479

Water

Spatio-Temporal Statistical Assessment of Water Quality Dynamics in the Kelani River, Sri Lanka: A Data-Driven Framework for Sustainable River Basin Management

Sandeepa Samarasinghe, Anuradha P. Hewaarachchi, D. M. P. Dissanayaka, Shameen Jinadasa and 2 others

Read abstract

The Kelani River is one of the most important freshwater resources in Sri Lanka, supplying drinking water, supporting aquatic ecosystems, and sustaining industrial and agricultural activities. However, increasing anthropogenic pressures have intensified the need for comprehensive long-term assessments of river water quality. This study evaluated the spatio-temporal variability of six physicochemical water-quality parameters (pH, temperature, turbidity, chemical oxygen demand, dissolved oxygen, and chloride) using monthly observations from twelve monitoring locations covering January 2007 to May 2022, representing the most up-to-date long-term CEA monitoring record available to the authors at the time of data acquisition. An integrated statistical framework comprising correlation analysis, Granger causality testing, vector autoregressive (VAR) modelling, ARIMA and seasonal ARIMA (SARIMA) forecasting, K-means time-series clustering, Pettitt temporal homogeneity testing, changepoint detection, and regression kriging was applied to investigate temporal dynamics, predictive relationships, temporal stability, and spatial variability. VAR, ARIMA, and SARIMA were selected as interpretable baseline models for assessing lagged dependence and seasonal temporal structure in the monthly water-quality series; comparison with machine-learning and hybrid models was beyond the scope of this study. The best-performing imputation method varied among parameter–location series; for pH, the minimum RMSE values of the selected methods ranged from 0.07 to 0.27. Temporal forecasting performance was also strongly parameter- and location-dependent. For DO, test-set RMSE ranged from 0.69 to 1.33 mgL−1, whereas chloride forecast RMSE ranged from 5.60 to 838.56 mgL−1, with the largest error observed at Victoria Bridge, reflecting the pronounced variability of the downstream chloride series. After Bonferroni correction across the 360 directional parameter–site comparisons, only two Granger-predictive relationships remained statistically significant: COD → chloride at Maha Oya and temperature → pH at Kaduwela Bridge. The corresponding VAR models had R2 values of 0.20 and 0.25, respectively, indicating modest explanatory power, while ARIMA and SARIMA models provided satisfactory forecasting performance for several parameters. The Pettitt homogeneity assessment identified nine candidate shifts at the unadjusted 5% significance level, but none remained statistically significant after Bonferroni correction across the 72 parameter–location tests. Regression kriging provided an exploratory representation of spatial variability in water-quality characteristics along the river. Overall, the proposed framework provides a robust statistical approach for long-term river water-quality assessment and supports evidence-based monitoring, pollution management, and sustainable river basin management, contributing to Sustainable Development Goal 6.

Abstract reproduced under CC BY · doi:10.3390/w18192480

Remote Sensing

PLUME: Staged Fusion of FY-4A AGRI and Rain-Gauge Observations for Bidirectional Refinement of Short-Duration Heavy-Rainfall Warnings

Xiang Lin, Yunying Li

Read abstract

Local short-duration heavy-rainfall warning requires early assessment of hazardous rainfall within a specified neighborhood. Geostationary meteorological satellites provide frequent, spatially continuous observations of cloud systems and storm evolution, but only indirectly reflect surface rainfall. Rain gauges measure surface accumulation directly, but only at irregular locations. Deep learning fusion methods commonly convert gauge observations to gridded representations before combining them with satellite observations. This treatment merges gauge measurements and local gauge support into a single representation, limiting how local surface conditions inform warning decisions and increasing the likelihood that local heavy-rainfall risk is underestimated or overestimated. We present PLUME (Precipitation Nowcasting with Late-Stage Residual Updates from Rain-Gauge Measurements for Short-Duration Heavy-Rainfall Events), a staged fusion model with gridded and point-origin gauge pathways. PLUME first combines a gauge-derived gridded rainfall background with FY-4A AGRI observations to form a base rainfall representation. The point-origin pathway uses local gauge features and observation support to update this representation through signed residuals, which the station decoder then maps to neighborhood-event probabilities. This staged design preserves the complementary roles of continuous spatial context, local rainfall measurements, and gauge availability. We evaluated PLUME for 0–3 h local short-duration heavy-rainfall warning over central and eastern China using an independent test set from May to September 2023. Across Barnes analysis, inverse distance weighting, and kriging, PLUME consistently improved the critical success index (CSI) over matched satellite–grid baselines. Relative CSI gains were 2.6–3.3% under complete gauge input and 3.6–5.9% under simulated gauge missingness. PLUME also reduced the associated CSI loss by 45.8–67.4%. Analysis of the variant with frozen batch normalization (BN) statistics showed that, under all three backgrounds and both input conditions, the largest positive probability updates were concentrated among baseline misses and converted some to hits, whereas the largest negative probability updates were concentrated among baseline false alarms and converted some to correct negatives. The dominant conditional CSI contribution shifted from false-alarm suppression at 0–1 h to missed-event recovery at 1–3 h. By using gridded and point-origin gauge information at different fusion stages, PLUME improves local heavy-rainfall warning through lead-dependent bidirectional refinement.

Abstract reproduced under CC BY · doi:10.3390/rs18193440

Zenodo

FUSING DEEP LEARNING AND REMOTE SENSING DATA TO ESTIMATE NON-OPTICALLY ACTIVE WATER QUALITY VARIABLES

Vasiliki Thomopoulou, Panagiotis Kossieris, Georgios Bariamis, Konstantinos Peroulis and 2 others

Read abstract

The efficient monitoring of inland water bodies such as rivers and lakes is crucial not only for ensuring human health and ecosystem balance, but also for the effective management of water resources. Two water quality variables of high importance and interest are Dissolved Oxygen (DO) and pH. The former variable is an ecosystem health index as it is connected to hypoxia events, while the latter one is linked to biochemical processes and the acidity of water. Currently, conventional water quality monitoring methods (e.g., via in-situ point sampling) are accurate for the particular point, but costly and time-consuming. On the other hand, satellite remote sensing provides a remedy, as it offers high spatial and temporal coverage. Most of the water quality applications of remote sensing data have focused on optically active parameters like Chlorophyll-a and Colored Dissolved Organic Matter (CDOM), but the estimation of non-optically active variables such as pH and DO poses a significant challenge. This study introduces a novel approach, that fuses Earth Observation (EO) and in-situ data, in a deep learning context, to provide estimations of non-optically active variables, characterized by high accuracy, high spatial coverage and low latency. The approach also considers the data imbalance challenge, which is common in environmental datasets. The approach is applied to Lake Yliki in central Greece. The Boeotikos Kifissos river basin, which lies upstream of Lake Yliki, is dominated by intense agricultural activity. The lake is also part of Athens’ water distribution system, making its efficient monitoring even more important.

Abstract reproduced under CC BY · doi:10.5281/zenodo.23243307

Sustainability

Multidimensional Vegetation Drought Response to Meteorological and Soil Moisture Deficits Across the Yellow River Basin

Yanfang Wang, Tao Jin, Shihui Liu, Xu Zhang and 4 others

Read abstract

Drought poses substantial threats to terrestrial ecosystems, yet whether vegetation vulnerability patterns systematically diverge between soil moisture and meteorological droughts—and how long-term hydrological adaptation governs ecosystem susceptibility to rare extreme droughts—remains poorly understood. Here, we systematically quantify vegetation responses to soil moisture drought (Standardized Soil Moisture Index, SSI) and meteorological drought (3-month Standardized Precipitation Evapotranspiration Index, SPEI-3) across the Yellow River Basin from 1982 to 2022, examining resistance, resilience, and peak loss alongside their spatial differentiation and temporal lags. SSI and SPEI-3 identified droughts exhibit distinct event characteristics: meteorological droughts occur more frequently with shorter durations, whereas soil moisture droughts persist 2–3 times longer once initiated, consistent with the buffering and memory effects of soil water storage. Vegetation response patterns diverge substantially between drought definitions, with SSI-derived resistance showing stronger associations with drought duration, whereas SPEI-3 conditions elevate the importance of drought-period climate. Wetter regions and deciduous-mixed forests exhibit the highest standardized peak losses despite comparable or higher resistance, consistent with the hypothesis that vegetation adapted to long-term water-abundant conditions may be more susceptible to rare extremes. Random Forest and SHapley Additive exPlanations analyses demonstrate that resistance is associated with drought attributes and drought-period climatic conditions, resilience is strongly associated with drought-period temperature and post-drought climate conditions, and peak loss shows the strongest background environmental dependence. Temporal lags between peak vegetation loss and drought intensity vary spatially, with stronger heterogeneity under SSI than SPEI-3. Our findings demonstrate that comprehensive vulnerability assessment requires integrating complementary drought metrics and multidimensional response frameworks, providing critical insights for improving drought monitoring, ecological risk assessment, and sustainable ecosystem management under climate change.

Abstract reproduced under CC BY · doi:10.3390/su181910234

Water

Improving National Water Model Evapotranspiration Forecasts Using Machine Learning Post-Processing: A California Case Study

Abin Raj Chapagain, Iman Maghami, Daniel P. Ames, Sujan Chandra Mondol and 2 others

Read abstract

Evapotranspiration (ET) forecasts from the U.S. National Water Model (NWM) carry systematic biases that are most apparent in water-stressed, intensively managed regions. This study develops a machine learning (ML) post-processing framework that bias-corrects 1–7 day accumulated ET (ACCET) forecasts from the NWM medium-range configuration using NWM meteorological forcings, the raw NWM-ACCET forecast, and auxiliary spatial and temporal predictors. Three tree-based models—Random Forest (RF), XGBoost, and LightGBM—were trained on AmeriFlux eddy-covariance observations at four California model-development sites and evaluated using nested five-fold temporally blocked cross-validation with a ±7-day purge window; two additional sites were withheld for independent off-site validation. All three models substantially outperformed the raw NWM forecasts. Pooled out-of-fold performance across the four model-development sites yielded R2 values of 0.832–0.854 and mean-normalized root mean squared error (NRMSE) values of 0.427–0.460, compared with R2 = 0.086 and NRMSE = 1.206 for the raw NWM forecasts. LightGBM and XGBoost were favored over RF in this four-site out-of-fold evaluation. Independent off-site validation yielded R2 values of 0.796–0.874 and NRMSE values of 0.380–0.459, with RF showing the most robust overall performance. Converting ACCET to lead-time-specific increments revealed substantially stronger lead-time degradation than was apparent from the accumulated metrics. Feature-importance and SHapley Additive exPlanations (SHAP) analyses identified wind, lead time, seasonality, and radiation as dominant predictors. The framework offers a practical route to more reliable operational ET forecasts for water management, agriculture, and drought monitoring, with potential for extension to climatically similar regions.

Abstract reproduced under CC BY · doi:10.3390/w18192476

Figshare

Explainable machine learning for spatial modelling of groundwater nitrate contamination: a multi-year analysis in the Duero basin

Manuel Rodríguez del Rosario, Victor Gómez-Escalonilla, Héctor Aguilera, Pedro Martínez‐Santos and 1 other

Read abstract

Groundwater nitrate contamination is a major water-quality concern, with implications for drinking water supplies and groundwater-dependent ecosystems. This study assesses the Duero River basin (Spain), where groundwater nitrate concentrations have increased in recent decades due to intensive agriculture and livestock farming. The probability of elevated groundwater nitrate concentrations was predicted using machine learning, formulated as a binary classification problem with a threshold of 37.5 mg/L. Monitoring data were combined with spatially derived environmental and anthropogenic predictors. Several tree-based algorithms were evaluated, with Random Forest selected for its consistent predictive performance. Model performance was assessed through repeated nested cross-validation and the use of metrics suitable for imbalanced datasets, achieving an AUC of 0.79-0.88 and an F1-score of 0.56-0.70 for the minority class across the three analysed hydrological years. SHAP analysis identified precipitation, diffuse agricultural pressures, distance to surface water bodies, NDVI, and soil properties as the main contributing predictors. Multi-year probability maps showed substantial spatial coherence with officially designated Nitrate Vulnerable Zones (46.9-64.9%), while identifying additional high-probability areas outside them (12.7-23.6%). The integration of explainable machine learning, multi-year spatial modelling, and regulatory comparison provides a framework to support groundwater monitoring, review vulnerable-zone delineation, and prioritise mitigation measures.

Abstract reproduced under CC BY · doi:10.6084/m9.figshare.34244894.v1

Journal of Global Ecology and Environment

Chemical Engineering Strategies for Wastewater Treatment in Agro-industries

S. Lakshmi, V. Yuvasree, R. Lavanya

Read abstract

Agro-industries, including sugar and distillery complexes, dairies, palm and olive oil mills, starch and tapioca units, breweries, slaughterhouses, poultry processors and fruit and vegetable canneries, generate effluents that are seasonal, hydraulically erratic and extraordinarily concentrated, with chemical oxygen demand (COD) frequently between 2,000 and 200,000 mg L-1 and nutrient loads far above those of municipal sewage. Conventional end-of-pipe disposal wastes both water and the nitrogen, phosphorus and potassium these streams carry, while land application of untreated effluent degrades soil structure, salinises the root zone and depresses crop performance. This review organises the treatment of agro-industrial wastewater as a sequence of chemical engineering unit operations rather than a catalogue of technologies. Effluent characteristics are first compiled by sector and translated into design-relevant descriptors - COD/BOD5 ratio, C:N:P stoichiometry, alkalinity, suspended and colloidal fractions, oil and grease, phenolics, colour and salinity - that determine process selection. Preliminary and physicochemical operations (equalisation, coagulation–flocculation, dissolved air flotation, chemical precipitation), biological reactor engineering (activated sludge and its sequencing-batch and attached-growth variants; upflow anaerobic sludge blanket, expanded granular sludge bed and anaerobic membrane bioreactors), advanced oxidation and electrochemical processes, adsorption on agro-residue-derived biochar, and pressure- and osmotically driven membrane separations are each examined in terms of governing kinetics, mass-transfer limitations, loading rates, energy demand and residual management. Particular attention is given to resource recovery - biomethane, struvite, ammonia, volatile fatty acids, polyphenols and reclaimed water - and to the agronomic quality of reclaimed effluent, evaluated through electrical conductivity, sodium adsorption ratio, residual sodium carbonate, trace-element burden and pathogen indicators. The review argues that no single unit operation is adequate for these matrices and that robust performance comes from deliberately sequenced trains in which an anaerobic bioreactor handles the bulk organic load and energy recovery, a polishing step addresses recalcitrant colour and micropollutants, and a nutrient-recovery loop diverts N and P to a fertiliser product rather than to a receiving water. Optimisation approaches (response surface methodology, artificial neural networks, life-cycle and techno-economic assessment) and the principal barriers to adoption in small and medium-sized agro-enterprises are discussed, and research priorities are identified for low-energy, low-sludge, recovery-oriented process trains suited to tropical and subtropical agricultural economies.

Abstract reproduced under CC BY · doi:10.56557/jogee/2026/v22i411207

Journal of Flood Risk Management

Rethinking Potential Flood Damage Assessment: A Data‐Driven Alternative to Varnes‐Based Models

Federica Zambrini, Daniele Fabrizio Bignami, Daniele Bocchiola, Enrico Maria Nava and 1 other

Read abstract

Reducing flood damage is a key challenge for both administrators and communities, particularly in the face of increasing climate‐related impacts. A fundamental prerequisite for effective damage reduction is a detailed understanding of the spatial distribution of territorial susceptibility to flood damage. In this context, we present an innovative, data‐driven methodology for mapping susceptibility to flood damage to private assets, leveraging machine learning techniques to overcome some of the limitations and uncertainties associated with traditional approaches based on the superposition of hazard, exposure, and vulnerability layers. The proposed methodology is applied to a case study in the Tuscany region of Italy, using approximately 11,000 claims related to flood events that occurred between 2013 and 2023. The claims are analyzed in relation to 15 predisposing factors representing both territorial and socio‐environmental characteristics of the study area. The introduction of negative samples, corresponding to locations where residential damage is assumed to be negligible or absent, allows the model to distinguish susceptible from non‐susceptible areas. The susceptible areas are subsequently classified into damage magnitude classes based on citizens' self‐reported assessments. The intrinsic variability and non‐linearity of the process, together with the uncertainty associated with citizen‐reported data and potential perception and reporting biases, are addressed through a second‐stage ensemble of independently trained models. The final outputs include a majority‐vote susceptibility map and a map of prediction stability. The model performs well in identifying the occurrence of reported damage, while the discrimination between damage magnitude classes remains more challenging, highlighting the intrinsic uncertainty associated with citizen‐reported damage data. The resulting maps are intended as practical tools to support flood damage reduction efforts and should be viewed as complementary to existing flood risk maps used in land‐use planning. The proposed framework provides a transferable approach for using observed damage records to characterize spatial patterns of potential damage, while its application to new geographical contexts requires site‐specific calibration and validation.

Abstract reproduced under CC BY · doi:10.1111/jfr3.70259

Zenodo

How Much Do Earth-Observation Foundation Models Help Flood Mapping When Labels Are Scarce? Frozen, LoRA and Full Fine-Tuning of Prithvi-EO 2.0 on Sen1Floods11

Mert Semih Sarıyerli

Read abstract

Earth-observation foundation models are expected to reduce the number of labels needed for tasks such as flood mapping. We test this expectation in a controlled setting on the Sen1Floods11 benchmark. Two U-Nets (trained from scratch and ImageNet-initialized) are compared with Prithvi-EO 2.0, a 300M-parameter masked-autoencoder foundation model, under three fine-tuning regimes: frozen encoder, low-rank adaptation (LoRA) and full fine-tuning. All models see the same six Sentinel-2 bands and the same budget of optimizer steps, with 1% to 100% of the training labels (3 to 252 chips), three seeds per setting and 84 runs in total. The U-Nets lead at every label fraction: 0.829 versus 0.767 water IoU for Prithvi with LoRA at full labels, and 0.809 versus 0.704 with only three labeled chips. The gap therefore does not close as labels become scarce; it widens. Among the fine-tuning regimes, LoRA matches full fine-tuning (0.767 versus 0.770) with about 140 times fewer trainable parameters, and a frozen encoder trails both by about seven points. Pretraining also does not reduce the accuracy drop on the held-out Bolivia region. A pixel-level error analysis locates the foundation model's shortfall at water edges and on water bodies smaller than 0.1 km^2, consistent with its 16-pixel patch resolution. For Sentinel-2 flood mapping on this benchmark, a small well-tuned CNN remains the better choice, even with a handful of labeled scenes.

Abstract reproduced under CC BY · doi:10.5281/zenodo.23105699

Environmental Research Communications

Transferability assessment of U-net model for rapid large-scale flood prediction over unseen catchments

Gianmarco Guglielmo, Jean‐Michel Tucny, Andrea Montessori, Michele La Rocca and 1 other

Read abstract

Assessing river flood hazard in ungauged catchments remains one of the most pressing challenges in hydrology. This study assesses a deep learning framework designed to achieve cross-basin transferability, in order to predict river flood hazard maps in unseen catchments. A U-net model is trained on data from a single basin under multiple synthetic rainfall events and evaluated on a different, unseen basin using a one-to-one basin transfer setup. Focus is placed on the beneficial effect of a physically augmented training in data-scarce scenarios. Results demonstrate the potential of the proposed framework to transfer flood hazard prediction from data-rich to data-scarce regions.

Abstract reproduced under CC BY · doi:10.1088/2515-7620/aeb236

7 October 2026

Water

A Magnitude–Variability Framework for Dynamic River Water Quality Characterization

Yike Chen, Yanbing Chi, Penghong Li

Read abstract

Continuous river water-quality monitoring increasingly captures short-term fluctuations that are poorly represented by conventional static assessments, yet frameworks for translating high-frequency observations into interpretable dynamic water-quality states remain limited. In this study, we developed a state-based framework for dynamic water-quality characterization using continuous river monitoring data. Ecological water quality departure (EWQD) was first used to quantify deviations from background conditions, after which two complementary descriptors, exposure level (EWQDL) and exposure variability (EWQDV), were introduced to characterize the magnitude and temporal instability of water-quality departures. Their joint distribution was then used to construct a two-dimensional dynamic state space, from which four representative departure states were identified to describe distinct combinations of departure intensity and temporal variability. Machine-learning models were further employed to identify the environmental controls associated with transitions among these dynamic states. The results showed that 34.7% of monitoring days exhibited within-day transitions in water-quality states, demonstrating substantial short-term variability. Simulated conventional grab sampling failed to identify 18.8% of unfavorable water-quality states revealed by continuous monitoring, illustrating the information loss that can arise from discrete observations. The EWQDL–EWQDV state space distinguished persistent, transient, stable, and highly unstable departure regimes that could not be differentiated from departure magnitude alone. Hydrometeorological conditions, particularly antecedent rainfall and temperature, were the dominant controls on dynamic state transitions, whereas land use primarily regulated the background susceptibility of river systems to water-quality departures. These findings demonstrate that river water quality is better represented as a time-varying state process than as a static condition. The proposed framework provides a systematic approach for translating continuous monitoring records into interpretable dynamic states and offers a basis for more representative water-quality characterization and adaptive monitoring strategies.

Abstract reproduced under CC BY · doi:10.3390/w18192474

Journal of Hydrology Regional Studies

Climate warming and intensified precipitation extremes triggered regime shifts in Lake Ayakkum, East Kunlun Mountains

Yongxiao Zhu, Junqiang Yao, Yang Wang

Read abstract

Study Region Ayakkum Lake on the Kumukuli Basin. Study Focus This study utilizes the meteorological and lake-level data to analyze climatic and hydrological change in the Ayakkum Lake region during 1961–2023, elucidating the driver mechanisms through which climate variablity and extreme weather events have influenced water level change of Ayakkum Lake. New Hydrological Insights for the Region The results indicate that the Mann–Kendall abrupt change analysis reveals a more pronounced warming–wetting and extreme trend in the Kumukuli Basin and its surrounding areas during 1995–2023. The climatic warming–wetting have led to a significant rise in the water level of Ayakkum Lake at a rate of 0.209 m·year⁻¹ (p < 0.01) during 1995–2009, with the rate accelerating to 0.438 m·year⁻¹ (p < 0.01) during 2010–2023. Random Forest analysis identified extreme precipitation (26.26% contribution) and temperature-driven glacial ablation (45.58% contribution) as the primary drivers of lake water level (LWL) rise, jointly explaining 71.84% of the increase. Between 1995–2009, the rise in LWL was predominantly driven by temperature-induced snow and ice melt. In contrast, the marked acceleration of LWL after 2010 was primarily attributed to increased extreme precipitation. These findings provide a significant scientific basis for water resource management in Lake Ayakkum.

Abstract reproduced under CC BY · doi:10.1016/j.ejrh.2026.104062

Journal of Hydrology Regional Studies

An exposure-aware graph neural network for peak flood depth prediction in New Jersey urban watersheds

Ahmad Ibrahim, Firas Gerges, Viravid Na Nagara, Abul Hassan Fatemi and 2 others

Read abstract

Study region This study evaluates flood prediction for three northern New Jersey watersheds. The municipalities, Englewood, Irvington, and Paterson, New Jersey, are susceptible to flooding, but vary in watershed size and extent of riverine influence. Study focus This study develops a spatio-temporal neural network that accurately infers overland and riverine flow dynamics from rainfall. Instead of conventional gridded approaches, this study utilizes unstructured graph representation of hydraulic computation points to preserve physical connectivity. A novel, importance-weighted loss function increases emphasis on high-priority areas while mitigating high-value dominance from channel regions. Incorporating synthetic rainfall patterns exposes the model to irregular storm sequences, bridging historical events and normalized design storms. New hydrological insights for the region The results indicate that the FloodMAGNet effectively predicts peak flood depths aligned with risk assessment objectives. Training with synthetic rainfall sequences established strong predictive capacity for non-monotonic events, demonstrated by an R 2 of 0.975 and a 3.3 cm RMSE on high-priority nodes for an observed hyetograph in Englewood. Evaluation of design storms in Irvington yielded an RMSE of 8.3 cm across the full domain and 5.8 cm in prioritized areas. Furthermore, evaluating the larger, river-dominated watershed in Paterson resulted in a baseline RMSE of 29.3 cm, which improved to 14.4 cm after loss function recalibration, demonstrating the framework’s methodological generalizability to distinct topographies and hydrologic responses.

Abstract reproduced under CC BY · doi:10.1016/j.ejrh.2026.104041

Science Advances

Predicting regional gray swans via translocation: AI weather models and Dubai’s unprecedented 2024 rainfall

Y. Qiang Sun, Pedram Hassanzadeh, Tiffany A. Shaw, Hamid A. Pahlavan and 1 other

Read abstract

Artificial intelligence (AI) models have transformed weather forecasting, but their skill for unprecedented weather extremes is unclear. Here, we analyze GraphCast, AIFS, and FuXi forecasts of the unprecedented 2024 Dubai storm, which had twice the training set's highest rainfall in that region. GraphCast and AIFS accurately forecast this event up to 8 days ahead. FuXi forecasts the event but underestimates the rainfall. Fine-tuning and receptive field analyses suggest that these models' success stems from "translocation": learning from comparable/stronger dynamically similar events in other regions during training. Evidence of "extrapolation" (learning from weaker events) is not found. Even events within the global distribution's tail are poorly forecasted, which is not only due to data imbalance (generalization error) but also spectral bias (optimization error). These findings demonstrate the potential of AI models to forecast "regional" gray swans and the opportunity to improve them through understanding the mechanisms behind their successes/limitations.

Abstract reproduced under CC BY · doi:10.1126/sciadv.ady9406

Sustainability

Land Use Change, Flood Exposure, and Vegetation Stress in the Chi River Basin, Northeast Thailand: A Multi-Temporal Remote Sensing and Machine Learning Assessment

Jiradech Majandang, Patiwat Littidej, Benjamabhorn Pumhirunroj, D. C. Slack

Read abstract

Climate change has intensified hydrological extremes in Southeast Asia, yet the relationships between land use change, flood exposure, and vegetation stress remain poorly understood in tropical floodplain environments. This study assessed land use dynamics, flood exposure, and vegetation stress in the Upper Chi River Basin, Northeast Thailand, using multi-temporal remote sensing datasets spanning 2015–2024, official land use data from the Land Development Department (LDD), and machine learning approaches, including Random Forest classification of binary flood occurrence, SHAP analysis, and spatial statistics. Agricultural land dominated the landscape (84.2% in 2023), with 11.0% land use change occurring between 2015 and 2023, primarily agricultural conversion from forest (38.6% of Forest Area converted) and water body (30.0% of Water Body converted). Flood frequency was highest in areas classified as Water Bodies (0.38 events) and Agricultural Areas (0.55 events), while Forest Areas showed the greatest resilience (0.05 events). Random Forest classification of flood occurrence achieved an overall accuracy of 0.920 and AUC-ROC of 0.905, with DEM emerging as the dominant predictor in both Random Forest importance (0.435) and mean absolute SHAP (0.308). Vegetation stress affected 21.1% of the basin, with 59.2% of points showing an NDVI decline greater than 0.02 and 33.3% falling below the VCI < 0.4 drought threshold. Four sensitivity classes were identified: Low (66.4%), Moderate (19.2%), High (12.4%), and Very High (2.0%), with Very High Sensitivity concentrated along the Chi River corridor in low-lying areas (mean DEM = 146.0 m). Moran’s I confirmed strong spatial clustering of flood frequency (I = 0.787) and sensitivity class (I = 0.534), highlighting the need for spatial statistical approaches. The findings support elevation-based land use zoning, forest conservation, and climate-resilient agricultural practices to enhance flood resilience and contribute to Sustainable Development Goals 2 (Zero Hunger), 11 (Sustainable Cities), 13 (Climate Action), and 15 (Life on Land).

Abstract reproduced under CC BY · doi:10.3390/su181910188

Scientific Reports

Quantifying spatially varying associations between water quality and sub-basin structural characteristics: a multiscale spatial approach

Seon Yeon Choi, Hun Kyun Bae, Chang Dae Jo, Heongak Kwon

Read abstract

Water quality patterns reflect spatially heterogeneous associations among topography, drainage connectivity, land use configuration, and soil hydrologic structure, whereas concentration-based assessments and global models may obscure where these relationships emerge and at which spatial scales they operate. Using monitoring data from 2020 to 2024, this study examined six water quality indicators: total organic carbon (TOC), total nitrogen (TN), total phosphorus (TP), dissolved oxygen (DO), pH, and electrical conductivity (EC) across 67 mid-basins in South Korea. Random forest (RF) screening identified relief ratio (R h ), drainage density (D d ), land use contagion (CONTAG land ), and soil patch density (PD soil ) as predictors representing basin morphometry, landscape configuration, and soil hydrologic heterogeneity. Multiscale geographically weighted regression (MGWR) was used to estimate basin-level associations and variable-specific bandwidths, and model fit was compared with that of ordinary least squares (OLS). MGWR produced higher in-sample R² values for all indicators and lower corrected Akaike information criterion (AICc) values for TOC, TN, TP, DO, and pH. Residual Moran’s I values were nonsignificant for five indicators, whereas EC retained significant negative residual spatial autocorrelation (Moran’s I = − 0.101, pseudo- p = 0.009), indicating that some parameter-specific spatial structure remained unaccounted for by the selected terrestrial basin predictors. Five-fold spatial block cross-validation showed parameter-dependent transportability, with the strongest performance for TN. Priority scores combining water quality conditions with MGWR-derived association magnitudes identified 11 priority basins, primarily in the Yeongsan River system, with additional basins in the Geum, Nakdong, and Seomjin River systems. These findings support parameter-specific interpretation and basin-targeted management.

Abstract reproduced under CC BY · doi:10.1038/s41598-026-74829-1

Remote Sensing

Linking Data Modalities and Deep Learning Architectures for Urban Flood Modeling

Wenjie Chen, Zhongnan Liu, Ge Yang, Haijun Yu and 2 others

Read abstract

Deep learning has been extensively applied to the modeling and management of urban flooding. While deep learning provides the analytical framework, the unique characteristics of the underlying data modalities themselves fundamentally drive model performance. Thus, this paper provides a scoping review of 289 recent studies (2016–2025), categorizing the relevant work into five modalities based on data acquisition methods, characteristics, and information processing workflows: Hydro-Geospatial Physical Base, Remote Sensing Imagery, Human Perception and Social Sensing, Environmental Acoustic, and Cross-Modal Integration. The hydro-geospatial physical base and remote sensing imagery dominate current studies because they are stable, widely available, and information-rich, whereas human behavior and acoustic data remain underutilized owing to heterogeneity, limited availability, and costly preprocessing. Key bottlenecks include information loss during preprocessing, insufficient cross-modal alignment, representations that struggle to balance statistical correlation with physical consistency, and marked regional variations in data. Ultimately, these limitations indicate that the field is still primarily dominated by purely data-driven models, and future progress requires integrating statistical learning with physical mechanisms to improve robustness and generalization capabilities.

Abstract reproduced under CC BY · doi:10.3390/rs18193426

International Journal of Climatology

Short‐Term Agricultural Drought Prediction Using Explainable Data‐Driven Models and Quantified Uncertainty

K. Saranya Das, N. R. Chithra, Venkataramana Sridhar

Read abstract

Accurate short‐term prediction of agricultural drought is critical for sustainable crop management and climate risk mitigation in drought‐prone regions. This study evaluates the effectiveness of the Standardized Precipitation Evapotranspiration Index (SPEI) as a key lagged predictor of agricultural drought in the Palakkad district of Kerala, India. Multiple machine learning and deep learning models were evaluated using lagged values of SPEI and the Standardized Soil Moisture Index (SSI) to predict SSI at 1–3‐month lead times. An additional model configuration incorporated global climatic indices to assess their potential influence on regional drought dynamics. Predictive uncertainty was quantified using Quantile Regression (QR), and model interpretability was evaluated through SHapley Additive exPlanations (SHAP) to identify the primary drivers of drought evolution. Unlike most short‐term drought studies, this work jointly integrates predictive uncertainty and explainable artificial intelligence to enhance operational reliability. Results indicate that the Artificial Neural Network (ANN) consistently achieved the highest predictive accuracy, with R 2 values reaching 0.84 at 1‐month lead time, decreasing to 0.53 at 2‐month and 0.26 at 3‐month. Root mean Square Error (RMSE) ranged between 0.50 and 0.59 at 1‐month lead time, but increased progressively at longer lead times. The Convolutional Neural Network (CNN) demonstrated comparable performance, with R 2 values ranging from 0.77 to 0.82 across different grid points at the 1‐month lead time. Uncertainty analysis demonstrated high coverage reliability (PICP ≈ 0.88–0.89) across ANN models, with increasing interval width at longer lead times. SHAP analysis confirms that local lagged hydroclimatic variables dominate near‐term drought prediction, while climatic indices provide only marginal improvements. The integration of uncertainty quantification and interpretability enhances the operational utility of these models, offering a transparent and scientifically robust framework for short‐term drought prediction, early warning applications and climate‐resilient agricultural management.

Abstract reproduced under CC BY · doi:10.1002/joc.70628

arXiv (preprint)

Low-rank tensor structure of precipitation and its application to satellite-reference merging

Ryan Solgi, Rohan Shankar, Hugo A. Loaiciga

Read abstract

The intermittent and variable nature of precipitation makes its accurate estimation over extended domains difficult, yet its spatiotemporal structure suggests that a low-rank representation may be possible. This work represents daily precipitation over the contiguous United States (CONUS) as spatiotemporal tensors and applies CANDECOMP/PARAFAC factorization, showing that preserving the native spatial and temporal modes yields more accurate reconstruction than factorizing independent daily fields or unfolded space--time matrices. Building on this finding, this work presents TMerge, a tensor-based framework that integrates satellite precipitation with sparse reference observations through shared low-rank spatial and temporal factors. TMerge was applied to correct the IMERG Final Run product with climate prediction center reference observations over CONUS. During 2019-2022, TMerge increased correlation from 0.53 to 0.85 and reduced root-mean-square error and mean absolute error by 48.2% and 29.3%, respectively. TMerge consistently outperformed linear bias correction, quantile mapping, and neural networks across seasons, precipitation-intensity regimes, and regions. Improvements were spatially coherent and largest in coastal regions where IMERG errors were greatest. These results demonstrate that low-rank tensor structure parsimoniously approximates the dominant spatiotemporal variability of precipitation and provides a practical mechanism for improving satellite estimates under limited reference observations over extended domains.

Abstract reproduced under arXiv (metadata CC0)

6 October 2026

Repository

West African Reservoirs Exhibit Variable Sensitivity to Large-Scale Climate Patterns

Valery Bessely Stanislas Kouassi, Kwok Pan Chun, Blé Anouma Fhorest Yao, Albert Elikplim Agbenorhevi and 4 others

Read abstract

Limited understanding of current trends in reservoir surface extent in West Africa and how large-scale climate patterns influence these dynamics challenges efforts to anticipate stress on water resources. To address this gap, we assessed reservoir sensitivity to large scale climate patterns across West Africa. We linked reservoir surface extent dynamics from 1985 to 2022 for 412 reservoirs with ten large-scale climate patterns represented by ten Sea Surface Temperature Anomaly (SSTA) indices, using supervised machine learning approaches and further analyzed the regional clustering in reservoirs responses. Most West African reservoirs, with significant trend, exhibit declining surface extent from 1985 to 2022 and respond primarily to combined Indo-Mediterranean-Atlantic climate signals rather than to single SSTA indices. Western Mediterranean Index (WMED) emerged as the dominant driver and is generally positively associated with reservoir surface extent, while the Eastern Mediterranean Index (EMED) showed contrasting effects. Secondary contributions were observed from Atlantic modes, particularly the Atlantic Multidecadal Oscillation (AMO) and the Atlantic Meridional Mode (AMM). Regression coefficients of SSTA indices vary widely across reservoirs, with limited significant spatial clustering, indicating variable sensitivity of reservoirs to large-scale climate patterns potentially shaped by local factors. This spatial heterogeneity further stressed the need to investigate its implication for adaptation strategies, as well as the effects of local management practices and unique reservoir configuration on reservoir sensitivity to large scale climate patterns.

Abstract reproduced under CC BY · doi:10.5194/egusphere-2026-4670

Water

A Mass-Conserving Spatiotemporal Graph ODE Model for Multi-Step Streamflow Forecasting in a Sparsely Gauged Semi-Arid Basin

Feng Liu, Lu Chang, Huiyu Zhao, Ruidong Jian and 2 others

Read abstract

Semi-arid loess basins in the middle Yellow River have rapid runoff responses, short hydrological memory, and sparse gauging networks, so pooled metrics across all stations can conceal forecast failures in low-flow tributaries and during dry periods. This study proposes DaSTGODE, a data-augmented spatiotemporal graph ordinary differential equation model with a structural water balance constraint. DaSTGODE integrates directed advection–diffusion propagation, soil-moisture-limited evapotranspiration, and a recursive water balance within a unified framework, thereby closing the water budget at every forecast step through the model structure. A reference scenario with perfect meteorological forcing, DaSTGODE_pre, is also established. It supplies the decoder with meteorological conditions over the forecast window as future meteorological forcing and quantifies the potential contributions of future precipitation, air temperature, and potential evapotranspiration by channel. Using daily records from 10 stations in the Wuding River basin from 2006 to 2020, DaSTGODE is compared with long short-term memory (LSTM) and graph long short-term memory (GraphLSTM) models for 1-, 3-, 6-, and 12-day forecast windows. DaSTGODE increases pooled Nash–Sutcliffe efficiency (NSE) from 0.532 and 0.473 for the best baseline to 0.599 and 0.567 for the 6- and 12-day windows, respectively. Relative to the best baseline, DaSTGODE increases the 1-day station-macro logarithmic NSE (logNSE) from 0.266 to 0.612 and reduces the mean number of stations with a negative NSE in the 6-day window from 4.6 to 2.0. Under the reference scenario in which future meteorological conditions are known, DaSTGODE_pre increases the 6-day outlet NSE from 0.141 for DaSTGODE to 0.331, and future precipitation alone provides 94% of the gain obtained from all three meteorological channels. HBV virtual nodes yield only small improvements in metrics at gauged stations, but provide daily streamflow forecasts at the outlets of 16 ungauged subbasins. Taken together, these results show that, under the semi-arid and sparsely gauged conditions examined here, the complete physical graph ODE kernel mainly improves cross-station reliability and relative low-flow dynamics, while the perfect meteorological forcing experiment quantifies the contribution of meteorological information over the forecast period to accuracy at medium and long forecast windows and identifies the channels responsible for that contribution.

Abstract reproduced under CC BY · doi:10.3390/w18192472

Sustainability

Spatiotemporal Variation Analysis of Groundwater Storage in Huaihe River Basin Based on GRACE and Interpretable Machine Learning

Qiang Han, Bingbing Li, Rui Zhu, Hanchen Cao and 1 other

Read abstract

Groundwater represents an essential water resource across the Huaihe River Basin, and its sustainable use is important for maintaining regional water resource security. This study systematically examined the long-term evolution, seasonal variability, and major drivers of groundwater storage anomalies (GWSA) throughout the Huaihe River Basin over the period 2002–2024. GRACE satellite observations were combined with Seasonal-Trend decomposition (STL) and the XGBoost-SHAP interpretable machine learning framework to conduct the analysis. Results indicated that GWSA exhibited a significant declining trend (−4.08 mm/a), characterized by marked seasonal fluctuations with peaks in June and troughs in August. Spatially, groundwater storage depletion was predominantly observed in the northern and northwestern regions, whereas slight recovery occurred in parts of the central-southern plains. Feature attribution analysis identified population density (mean |SHAP| = 0.386) as the most important predictor of GWSA variability among the factors considered, followed by cropland area, temperature, precipitation, and NDVI. SHAP interaction analysis further revealed model-based interactions among anthropogenic and hydroclimatic factors. Spatial correlation analysis showed distinct spatial associations of GWSA with human-related and natural factors. And the temporal shift in SHAP values from negative to positive suggests changes in the associations between human-related factors and GWSA, which may be related to groundwater management and regional water-supply adjustments. These findings provide quantitative insights into the complex interplay between anthropogenic and climatic drivers on GWS, providing a scientific basis for groundwater management in the HRB.

Abstract reproduced under CC BY · doi:10.3390/su181910167

Frontiers in Water

A comprehensive review of AI-driven water-quality monitoring and prediction: advances, challenges, and future directions

K. Yukesh Kumar, P. Kumaresan

Read abstract

Water quality has emerged as a critical global concern that requires advanced monitoring and management strategies. Traditional water-quality assessment methods predominantly rely on laboratory-oriented analysis that are time consuming, expensive, and are often labor-intensive. This study presents a thorough overview of the paradigm shift from conventional lab analysis towards intelligent and automated, AI-based assessment and monitoring frameworks. The survey systematically explains the use of machine learning (ML), deep learning (DL), and hybrid models to analyze complicated multidimensional and nonlinear water parameters. This study evaluates predictive architectures ranging from traditional regression-based models to advanced ensembles and deep neural networks (CNNs, ANNs, and RNNs) integrated with IoT computing technologies for continuous water quality monitoring. Enhancing models’ interpretability and transparency through Explainable Artificial Intelligence methods like SHAP and LIME are analyzed. This survey addresses existing challenges like data scarcity, interpretability of the model, and computational complexity. Self-attention-based architectures, generative AI, and edge computing used to improve the robustness of future water-quality management systems are also assessed. An overall structured research perspective on AI-driven water-quality assessment with methodological gaps with future scope are established in this survey.

Abstract reproduced under CC BY · doi:10.3389/frwa.2026.1946412

Journal of Hydrology Regional Studies

Multi-scale SPI drought forecasting and threshold-based early warning using optimized machine learning models: A case study of Sharjah, United Arab Emirates

Mohamed Elkollaly, Ahmed Sefelnasr, Assaad Hassan Kassem, Abdel Azim Ebraheem and 4 others

Read abstract

Study region: Sharjah, United Arab Emirates (UAE), Arabian Peninsula Study focus: Drought forecasting in arid environments is constrained by sparse, intermittent rainfall and abrupt transitions between wet and dry states. This study develops a multi-scale machine learning (ML) framework for one-month-ahead Standardized Precipitation Index (SPI) forecasting in Sharjah, UAE, using monthly precipitation and temperature data from 1995 to 2024. SPI-3, SPI-6, and SPI-12 were forecast under three predictor configurations: antecedent SPI combined with climatic variables, SPI-only, and climate-only. Seven model families were evaluated: Random Forest (RF), eXtreme Gradient Boosting (XGB), Support Vector Regression (SVR), Autoregressive Integrated Moving Average with eXogenous inputs (ARIMAX)/Seasonal ARIMA (SARIMA), Artificial Neural Network (ANN), and Long Short-Term Memory (LSTM). These models were tested under standard tuning, Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and Hybrid FOX–Tree-Seed Algorithm (HFTA) optimization, with optional ARIMA residual correction. Performance was assessed using continuous metrics, threshold-based drought classification, permutation-based predictor importance, and Taylor diagrams. New hydrological insights for the region: This study provides a systematic separation of drought memory and meteorological forcing as SPI predictors across three accumulation scales in an arid environment. Forecasting skill improved consistently from SPI-3 to SPI-12. The best SPI-3 result used SVR + HFTA + ARIMA under the combined configuration (R² = 0.808, RMSE = 0.295). For SPI-6, the climate-only XGB + PSO model achieved the strongest performance (R² = 0.891, RMSE = 0.266), demonstrating that seasonal drought can be forecast from meteorological predictors alone, without prior SPI records. The highest overall skill was obtained for SPI-12 using SVR + GA + ARIMA under the SPI-only configuration (R² = 0.943, RMSE = 0.222), highlighting the dominant role of drought memory at annual scales. Threshold-based classification showed that SPI-6 < 0 was detected most reliably by the climate-only XGB + GA model, with 98.33% balanced accuracy and zero false alarms. SPI-12 ≤ −1 was best detected by LSTM + PSO, reaching 94.44% balanced accuracy. Predictor importance showed that antecedent SPI governs long-scale forecasts, while precipitation drives climate-only models. The framework supports multi-scale drought early warning and water-resource planning in arid regions, contributing to SDG 2, SDG 6, and SDG 13.

Abstract reproduced under CC BY · doi:10.1016/j.ejrh.2026.104048

Ecological Engineering & Environmental Technology

An online drift-aware digital twin for regenerative wastewater systems: Conformal water-quality risk and rainfall-forced early warning from live sensor telemetry

C. Manjunath, Pitta Shankaraiah, G. Malli Reddy, M. Mallikarjuna Rao and 4 others

Read abstract

Digital twins for wastewater systems are normally fitted once and then frozen.That is a reasonable simplification while process conditions, sensor behaviour and weather patterns remain close to the training period, but it can fail silently when they move.This paper reports RT-ReGenTwin, an online digital twin that learns continuously from public sensor telemetry and reports threshold-exceedance risk as a calibrated probability rather than as a point forecast.Fifteen-minute water-quality records covering 2026-03-13T05:00 to 2026-09-09T05:00 were obtained from two continuous monitoring stations of the United States Geological Survey on effluent-influenced reaches downstream of large municipal treatment works, with matched hourly rainfall from Open-Meteo.Two further stations were probed and not used, one because it reports no continuous water-quality determinand and one because it had no current record.Turbidity, dissolved oxygen and specific conductance were used as receiving-water surrogates for solids loading, oxygen-demand effects and effluent fraction.The twin couples an adaptive random-forest regressor with drift detectors and an adaptive conformal layer.Across four station-determinand combinations, empirical coverage was closer to the nominal 0.90 level for the online calibrator than for the frozen reference in all four cases.Only one of the four combinations recorded any threshold exceedance during the window, so the probabilistic scores and the dispatch analysis rest on that stream alone; for it, coverage was 0.898 versus 0.595 and the Brier score 0.0472 versus 0.0345.Risk-triggered regenerative-capacity dispatch avoided 337 of 590 threshold-exceedance hours at a duty cycle of 0.163, compared with 414 avoided under continuous operation.The results support calibrated risk reporting for streaming environmental monitoring, while the receiving-water surrogates and dispatch simulation should not be interpreted as direct permit-compliance measurements or a treatment experiment.

Abstract reproduced under CC BY · doi:10.12912/27197050/240148

Journal of Hydrology Regional Studies

An integrated physics-based–machine learning framework for groundwater head prediction in a highly managed multi-aquifer watershed

Rajesh Khatakho, Ernest William Tollner, David Emory Stooksbury, Adam M. Milewski and 2 others

Read abstract

Study region Lower Apalachicola–Chattahoochee–Flint (ACF) Basin, Southwestern Georgia. Study focus Predicting hydrogeological processes in complex, heterogeneous multi-aquifer systems remain challenging, especially where intensive agricultural pumping strongly alters groundwater dynamics. This study developed an integrated MODFLOW–machine learning (ML) framework to simulate and forecast spatiotemporal groundwater head variations, drought occurrences under varying climatic and anthropogenic stress conditions. New hydrological insights Among the evaluated ML models, the MODFLOW-XGB exhibited superior performance across 85 observation wells, achieving an RMSE of 0.92 during the testing phase less than one-fourth of the standalone MODFLOW RMSE (3.77). Based on validation metrics, the ML models were ranked in descending order of performance as XGB, RF, GPR, DT, SVM, and ANN. The optimized MODFLOW–XGB model forecasted maximum seasonal groundwater declines ranging from 0.23 to 5.26 m across 30 observation wells, with the largest drops (>2 m) observed in regions of intense irrigation underscoring strong anthropogenic influence. Groundwater drought index revealed spatially variable drought frequency (15–20%) in confined areas with limited recharge, driven by low-conductivity Upper Semi-Confining Layer (USCU). Wells beneath thin USCU experiencing sharp short-term declines during periods of intensive abstraction and limited recharge. Modified Mann-Kendall test and Sen’s slope estimator depicted that no significant long-term trend in groundwater drought but seasonal decline in groundwater was observed. The research reveals the prospective of the coupled framework for real-time groundwater forecasting, early drought warning, and targeted watershed management.

Abstract reproduced under CC BY · doi:10.1016/j.ejrh.2026.104052

Remote Sensing

Multi-Paradigm Machine Learning for Opportunistic Rainfall Estimation from Satellite Microwave Links

Luca Marini, Emanuele Maria Sciortino, Filippo Giannetti, Giovanni Scognamiglio and 1 other

Read abstract

This study presents a framework for opportunistic rainfall detection and estimation exploiting SNR data from satellite downlink signals. We develop a heterogeneous ensemble including diverse ML architectures for the classification task, namely CRF, GRU, Chronos-Bolt, and TSPulse. Evaluating models characterized by distinct assumptions, scales, and computational complexities provides critical insights into their behavior for the weather and climate remote sensing domain. The resulting ensemble yields a significant improvement in precipitation classification performance. For the regression task, a traditional model leveraging the well-known ITU recommendations is compared against a GBM, with the ML-based approach demonstrating a substantial improvement in KPI. Finally, we tested the integration of the classification and regression modules via a hard and a soft gating mechanism. While these combinations enhance several point-to-point metrics, they do not improve the estimation of event-total accumulated rainfall. These findings highlight the potential of ML approaches for satellite-based opportunistic environmental monitoring.

Abstract reproduced under CC BY · doi:10.3390/rs18193414

Machine Learning Earth

GIS-based Deep Learning for Identifying and Characterizing Small Community Lagoon Wastewater Treatment Systems

Denis S. Ruto, Christopher A. Ramezan, Kevin D. Orner

Read abstract

Wastewater lagoon systems remain a dominant form of secondary treatment for small communities in the United States due to their low energy requirements, operational simplicity, and affordability. Despite their widespread use, national inventories of lagoon systems are fragmented, inconsistently reported, and often outdated, with non-discharging systems particularly underrepresented. These data gaps limit effective infrastructure planning, environmental assessment, and equitable allocation of resources. This study presents a GIS-based deep-learning framework that leverages high-resolution aerial imagery and a Mask Region-Based Convolutional Neural Network (Mask R-CNN) to automatically detect, delineate, and classify lagoon systems. The goal of this study was to develop a GIS-based deep-learning framework using Mask R-CNN for lagoon identification and characterization from aerial imagery, and to demonstrate its utility for infrastructure mapping. The resulting spatial dataset can contribute to the development and updating of more comprehensive national lagoon inventories and improve the representation of small communities often underserved by traditional data collection mechanisms. Model evaluation across diverse geographic and land-use contexts demonstrated facility-scale lagoon detection driven by physically meaningful infrastructure features such as basin geometry and embankment continuity, enabling consistent mapping without reliance on facility registries. In contrast, inference of operational characteristics, particularly aeration, remains more variable and constrained by the indirect and transient nature of surface expressions in single-acquisition imagery. By delineating both the capabilities and limitations of optical imagery for lagoon characterization, this work updates and expands lagoon inventories and supports the development of more comprehensive and representative databases critical for planning, prioritizing, and managing wastewater treatment systems in small communities.

Abstract reproduced under CC BY · doi:10.1088/3049-4753/aeb0ac

Hydrology and earth system sciences

The ability of LSTM to model snowmelt versus rainfall generated floods

Sigrid Jørgensen Bakke, Danielle Marie Barna, Kolbjørn Engeland, Sjur Anders Kolberg and 1 other

Read abstract

One of the most important skills of hydrological models is to simulate timing and magnitude of flood events. Long Short-Term Memory (LSTM) networks are currently among the most successful models for streamflow and flood prediction over large regions. In snow-influenced catchments, which typically comprise a minority in large-scale studies, floods are generated by two distinctly different processes, snowmelt and rainfall. The applicability of hydrological models in such regions is therefore dependent on their ability to represent both types of floods. Nevertheless, flood evaluations of LSTM taking different flood-generating processes into account are currently lacking. This study fills this gap by evaluating the ability of LSTM to model flood peak characteristics separately for snowmelt and rainfall generated floods. The trained LSTM model successfully simulated streamflow time series across the 103 evaluated catchments, with average NSE of 0.85 and average KGE of 0.87 over the unseen evaluation period. LSTM exhibited better performance in the majority of the catchments in terms of flood peak timing and magnitude for both rainfall and snowfall generated floods when compared to the operational hydrological model in the region (HBV) used as a benchmark. Both models had a 24 pp higher percentage of correctly simulated peak days for rainfall generated floods as compared to snowmelt generated floods. LSTM outperformed HBV for a larger proportion of the catchments in terms of peak timing of rainfall generated events (83 %) as compared to snowmelt generated events (64 %). On the other hand, a larger proportion of the catchments were improved by LSTM for snowmelt generated events as compared to rainfall generated events when considering peak magnitudes. The largest improvements in peak magnitudes were found for rainfall generated events, in particular for catchments where HBV exhibited high (> 40 %) absolute errors. Overall, our findings bring confidence that LSTM can improve hydrological services in regions subject to both snowmelt and rainfall generated floods.

Abstract reproduced under CC BY · doi:10.5194/hess-30-6207-2026

Asian Journal of Environment & Ecology

Coastal Groundwater Sustainability in Keralam, India: A Critical Review of GIS-and Remote Sensing-based Seawater Intrusion Vulnerability Assessment

Arjun Ram Gopi, Pottepaka Shravan, Rachcha Arun

Read abstract

Keralam's 590-kilometre coastline hosts nine of the state's fourteen districts, an exceptionally dense coastal population, and a hydrogeologically distinctive combination of narrow lateritic and alluvial coastal aquifers, an extensive 1,500-kilometre backwater and estuarine network centred on the Vembanad, Ashtamudi and Kallada systems, and intense monsoon-driven recharge that alternates seasonally with acute groundwater stress. This critical review synthesises evidence from studies published predominantly between 2018 and 2026 in peer-reviewed and indexed journals to assess how Geographic Information Systems (GIS), satellite remote sensing, and parametric vulnerability indices, principally GALDIT and DRASTIC together with locally modified variants such as GALDIT-U, have been applied to characterise seawater intrusion (SWI) risk across Keralam's coastal aquifers. The review finds that district-level research coverage is markedly uneven: Ernakulam, Kozhikode and, more recently, Thiruvananthapuram have been investigated using different combinations of vulnerability indexing, hydrochemical analysis, electrical resistivity and numerical modelling, while Kollam, Thrissur, Kannur and Kasaragod remain comparatively underexamined despite documented salinity problems. Keralam's laterite-dominated hydrogeology and strong monsoonal recharge generally suppress GALDIT-index vulnerability relative to deltaic coasts elsewhere in India, but urbanisation, backwater proximity, groundwater over-extraction and coastal erosion produce sharply localised high-vulnerability pockets, particularly around Kochi, Kozhikode and Thiruvananthapuram's urban and harbour fringes. Sea-level-rise projections under RCP4.5 and RCP8.5 scenarios have also been incorporated into groundwater and seawater-intrusion modelling, while explainable machine learning and Monte Carlo uncertainty frameworks are only beginning to be applied within the state. The review concludes by identifying priority gaps: uneven district-level and seasonal monitoring coverage, limited hydrochemical or geophysical validation of GIS-based vulnerability maps, minimal integration of dynamic land-use changes into index parameterisation, and an absence of state-wide, standardised comparative assessment across Keralam's full coastal gradient.

Abstract reproduced under CC BY · doi:10.9734/ajee/2026/v25i101031

Repository

Review Article: Compound Flooding in Coastal and Estuarine Catchments: Modelling, Management, and Climate Adaptation Insights from Cork Harbour

Ashenafi Yohannes Battamo, Rory Scarrott, Jeremy Gault, Anne Marie O’Hagan

Read abstract

Compound flooding (CoMF) has emerged as a critical natural hazard for coastal and estuarine regions, where interacting fluvial, pluvial, groundwater, tidal, and storm surge processes generate non-linear amplification effects that challenge conventional single driver flood risk assessments. Yet, despite major advances in compound flood science, integrated literature reviews that link multivariate hazard dynamics, modelling approaches, risk management, and climate adaptation strategies remain limited. This systematic review synthesises three decades of research on CoMF in Cork Harbour and Catchment (CHC), a highly dynamic coastal-estuarine system increasingly exposed to climate change. Following PRISMA guidelines, 503 documents were screened across Web of Science, Scopus, Google Scholar, and grey literature sources, identifying 59 peer-reviewed articles and 28 technical reports and policy documents published between 1991 and 2025. This review provides two key contributions of international importance. First, it positions CHC as a potential national reference site and globally relevant compound flood model system, demonstrating how climate driven interactions between extreme rainfall, river discharge, storm surge, tides, and Sea-Level Rise (SLR) generate non-linear amplification effects. Results reveal that while fluvial processes currently dominate with a 30 % increase in future flood inundation, projected SLR and surge intensification could amplify coastal inundation by up to 400 %, surpassing fluvial contributions and shifting the dominant contribution towards coastal mechanisms. Second, it offers a transferable methodological scheme, showcasing best practice integration of multi-scale hydrodynamic modelling, multivariate statistical dependence analysis, and emerging physics informed AI tools capable of improving real time forecasting. Advances in high-resolution hydrodynamic modelling and multivariate statistical frameworks have improved simulation fidelity, yet critical gaps persist—particularly in probabilistic modelling, artificial intelligence (AI) enhanced real-time forecasting, and integration of urban drainage and socio-ecological resilience. We identified six priority future research directions that could assist in implementing local action and hence have transferable international relevance: (i) probabilistic and machine-learning approaches for compound flood hazard prediction; (ii) integrated modelling frameworks coupling climatic, hydrodynamic, and socio-ecological systems; (iii) continuous climate and hydrological monitoring infrastructure; (iv) sensor networks for real-time data acquisition; (v) AI-enhanced real-time forecasting platforms and (vi) adaptive governance for climate-resilient hybrid structural and non-structural interventions. By consolidating existing knowledge and articulating clear future research priorities, this contributes a comprehensive and transferable framework for understanding, modelling, and adapting to compound flood risk in a warming world.

Abstract reproduced under CC BY · doi:10.5194/egusphere-2026-5661

5 October 2026

arXiv (preprint)

Skillful Data-Driven Subseasonal Soil Moisture Forecasting: Prospects and Limits for Flash Drought Prediction

Noelia Otero, Atahan Özer, Miguel-Ángel Fernández-Torres, Jackie Ma

Read abstract

Despite substantial progress in short-to-medium-range weather forecasting, predicting high-impact events such as flash droughts remains a key challenge for both early warning operations and physically-based subseasonal-to-seasonal (S2S) prediction systems. Here we demonstrate that, for S2S soil-moisture forecasting over Europe, forecast skill depends as much on how the prediction problem is formulated as on the forecasting model itself. Using a Vision Transformer-based architecture with dual-pathway temporal and spatial attention, we show that residual learning is essential to outperform persistence. This advantage is realized only when forecasting root-zone soil moisture in physical units rather than standardized anomalies, revealing that the target representation itself constrains predictability. A probabilistic extension via quantile-head fine-tuning further provides well-calibrated predictive distributions. Benchmarked against deep-learning and operational ECMWF S2S baselines over 2021-2022, our model achieves the highest deterministic and probabilistic skill at all lead times and reliably detects anomalously dry root-zone states (below the 20th percentile). Yet flash drought onset, defined by multi-pentad intensification criteria, remains a fundamental challenge shared across all current S2S systems. These findings advance data-driven S2S soil-moisture forecasting while highlighting the remaining challenge of predicting rapid drought development.

Abstract reproduced under arXiv (metadata CC0)

2 October 2026

arXiv (preprint)

Scale-Recursive Rectified Flows for Few-Step Precipitation Ensembles

Shunya Nagashima, Takumi Bannai

Read abstract

Fine-resolution precipitation estimates support flood risk assessment and water management, but coarse satellite products cannot resolve rainfall within each grid cell. Generative models address this ambiguity by producing ensembles of plausible high-resolution rainfall fields. Among these models, rectified flows generate samples by iteratively transforming random noise into rainfall fields. Reducing the number of sampling steps accelerates generation but can make ensemble members too similar, understating uncertainty. We propose a scale-recursive rectified flow that generates broad patterns before local details and guides sampling-step allocation by comparing ensemble variability with prediction error across spatial scales. Validation scores and rainfall power spectra constrain the allocation to avoid excessive amplification. In satellite-to-radar downscaling over the contiguous United States, our analysis identified broad rainfall patterns as the main source of insufficient ensemble variability under reduced sampling budgets. Allocating more steps to the coarse flow improved probabilistic accuracy and rain detection across training seeds at fixed architecture and computational cost. The proposed model also achieved better probabilistic accuracy with shorter sampling time than a nonrecursive flow using more steps.

Abstract reproduced under arXiv (metadata CC0)

30 September 2026

arXiv (preprint)

PrecipJEPA: JEPA-Regularized Future-State Prediction with Motion-Source Rendering for Precipitation Nowcasting

Yufeng Zhu, Dan Niu, Qiliang Wu, Weiwei Huang and 3 others

Read abstract

Long-term precipitation nowcasting requires modeling radar-echo evolution while preserving localized high-intensity structures. Recent radar-specific studies motivate location-aware prediction and separating echo displacement from intensity change. However existing encoders learn historical representations mainly from final forecast errors. We propose PrecipJEPA, which couples a structured forecasting path with an auxiliary path that enriches its encoder from observed radar history. In the forecasting path, an online encoder first converts the observations into spatiotemporal tokens. The Task-Driven Future-State Predictor (TFP) combines these tokens with a recent-dynamics summary and spatiotemporal queries to construct future radar states. The Parallel Motion-Source Renderer (PMSR) decodes these states into motion and source-sink fields that transform the latest observation into future frames. During joint training, the History-Masked JEPA (H-JEPA) operates on the auxiliary path to predict masked historical features from visible context, directly supervising the same online encoder from the observed sequence. Experiments on SEVIR and MeteoNet show that PrecipJEPA improves highest-threshold CSI by 118.6% and 35.1%, respectively, over the strongest baselines, while maintaining the highest mean CSI throughout the 3-hour forecast.

Abstract reproduced under arXiv (metadata CC0)

29 September 2026

arXiv (preprint)

A Model-Agnostic Physics-Guided Adapter for Few-Shot Transfer of Coastal Flood Prediction Models to Unseen Regions

Bilal Hassan, Areg Karapetyan, Samer Madanat

Read abstract

Deep learning surrogates can produce high-resolution coastal flood maps orders of magnitude faster than physics-based hydrodynamic simulators, yet transferring them to new coastal regions remains costly, since generating target-region data for fine-tuning typically requires numerous time-consuming simulations. To tackle this bottleneck, we introduce the Physics Adapter (PA), a compact, architecture-agnostic adaptation interface that enables efficient few-shot transfer of flood prediction models across diverse coastal regions. PA predicts peak water level through a differentiable wet/dry response that compares terrain elevation against a learned water level, and blends this physics-structured prediction with a data-driven branch through a learned gate. Unlike physics-informed formulations, PA imposes no PDE-residual or conservation losses and instead exploits elevation as an architectural inductive bias, adding a negligible number of trainable parameters. We integrate PA into 12 heterogeneous models, and evaluate them on two coastal regions with markedly distinct geometries, topographies, and shoreline protection configurations. The performance of PA is benchmarked against a no-physics baseline, full fine-tuning, and standard parameter-efficient fine-tuning (PEFT) methods, considering both within-region generalization to unseen sea level rise values and between-region transfer. In low-shot regime (K=3), and averaged over all backbones and transfer settings, adding PA reduces root mean square error by 11.5% when only the output head is adapted on a frozen backbone, by 15.4% when combined with PEFT methods, and by 22.9% under full fine-tuning, compared to matched configurations without PA. Taken together, the findings of this work offer practitioners a concrete recipe for extending DL-based coastal flood predictors to new, data-scarce regions.

Abstract reproduced under arXiv (metadata CC0)

arXiv (preprint)

NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters

Haoran Xu, Xingzhuo Guo, Yuchen Zhang, Jincheng Zhong and 2 others

Read abstract

Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.

Abstract reproduced under arXiv (metadata CC0)

28 September 2026

arXiv (preprint)

Explainable Deep Learning for Probabilistic Nowcasting of Radar Reflectivity in Tornadic Storms

Nathan Erickson, Amy McGovern, Aaron Hill

Read abstract

Tornadoes pose substantial risk to human life and property in the United States, causing more than 50 fatalities and \$100 million of property damage on average annually. When tornadoes are likely, weather radar provides critical information for forecasters by providing information on storm morphology, storm motion, and intensity trends. Additional tools such as satellite and numerical weather prediction model runs can provide useful short-term information for understanding changes in storm characteristics. This work demonstrates a U-Net deep-learning system for nowcasting the evolution of radar reflectivity following tornadogenesis, which can provide value to forecasters by synthesizing large amounts of input data (e.g., radar imagery, near-storm environment data) and generating predictions of radar reflectivity from its inputs. Inputs to the model are radar imagery from the Multi-Radar Multi-Sensor (MRMS) dataset and near-storm environment data from the High-Resolution Rapid Refresh (HRRR) numerical weather prediction model. The U-Net is trained on a dataset of tornadic storms to produce 30 minutes of probabilistic predictions of radar reflectivity following tornadogenesis, with probabilistic predictions obtained by predicting parameters of the SinhArcSinh, or SHASH, distribution. The model produces physically realistic predictions of radar evolution, achieves comparable skill to next-hour forecasts from the HRRR, demonstrates reasonable probabilistic calibration and is accompanied by a variety of explainability methods to improve understanding by end users. Additionally, predictions from the model can be obtained much more quickly than those from a numerical weather prediction model. With further development, this model could be extended to nowcast radar reflectivity evolution in an operational setting.

Abstract reproduced under arXiv (metadata CC0)

arXiv (preprint)

MW-Nowcast: Six-hour ensemble nowcasting of extreme precipitation

Ning Wang, Zuliang Fang, Weixin Jin, Zhongjian Lv and 18 others

Read abstract

Extending reliable nowcasting of extreme precipitation could provide critical additional time for warnings and emergency response during high-impact events such as flash floods. Radar-based generative machine-learning models have enabled skilful hyperlocal precipitation nowcasting, but accurate prediction of intense precipitation remains confined to the first few hours. Because storm-scale structure is predictable for longer than individual cells, a natural strategy is to predict that structure while generatively modelling only the uncertain local growth, decay, reorganisation and initiation of storms. Here we present Microsoft Weather Nowcast (MW-Nowcast), a six-hour ensemble radar nowcasting model that jointly learns a deterministic predictor to capture organised precipitation structure shared across ensemble members, and a generator to produce diverse local residuals around this shared prediction. Across independent test data from the United States, Europe and China, MW-Nowcast achieves higher detection skill than leading methods for heavy and extreme precipitation throughout the 6 h horizon. For the most intense rainfall, MW-Nowcast doubles the available warning time across all three regions, delivering 6 h forecasts with skill previously limited to 3 h for the leading generative baseline. A cost-loss decision analysis shows that MW-Nowcast retains substantial value for a broad range of applications even at 4-6 h, where alternative methods offer little benefit. These additional hours can give forecasters and emergency managers the time to warn and act before extreme rainfall strikes, helping to protect lives and property.

Abstract reproduced under arXiv (metadata CC0)

26 September 2026

arXiv (preprint)

Distributed Hydrological Modeling in the Feature Space

Mohamad Hakam Shams Eddin, Maria Luisa Taccari, Yikui Zhang, Shijie Jiang and 2 others

Read abstract

Accurate forecasting of river discharge and floods is very challenging. River dynamics are affected by storage, meteorological forcing, and flow propagation at different spatial and temporal scales. Forecasting thus requires a framework that considers the upstream-to-downstream flow through river networks across grid cells and catchments. This modeling is known in hydrology as distributed modeling and routing. Existing deep learning approaches either ignore this topology, operate on lumped catchments, or route predicted physical quantities through a separate graph or physical routing model. We instead introduce feature-space routing: a topology-aware state-space operator embedded directly in the forecasting dynamics. At every forecast step, the operator gathers latent states from upstream grid cells and causally updates the downstream state according to the known river network. This preserves the physical connectivity of the river system while allowing the propagated state itself to be learned end-to-end and allows the model to predict river discharge considering both local dynamics and neighboring upstream contributions. To address uncertainty and provide probabilistic forecasts, we minimize the fair continuous ranked probability score (fCRPS) as a training objective. Our experiments on the European Flood Awareness System (EFAS) and observational data for river discharge forecasting demonstrate that encoding the physical structure of river networks explicitly in the feature space substantially improves the forecasting skill, particularly in an ungauged setting. Our approach achieves state-of-the-art results on both reanalysis and observational data and is able to forecast maps of river discharge at 1 arcminute and 6-hourly resolution up to 10 days lead time.

Abstract reproduced under arXiv (metadata CC0)

24 September 2026

arXiv (preprint)

HydroSphere: A Framework for Governed, Self-Healing Wastewater Infrastructure

Prabu, Fancy C, Suresh A, Srini Ramaswamy

Read abstract

Rapid industrialization and urban growth are increasing pressure on water quality and wastewater treatment systems, while conventional treatment plants often rely on static monitoring and control strategies that cannot easily adapt to changing pollutant conditions. This paper presents HydroSphere, a governed, data-driven framework for real-time water quality monitoring, forecasting, treatment optimization, and fault recovery. HydroSphere is evaluated using 2.82 million water-quality measurements collected between 1940 and 2023. The framework integrates three main components. First, a hybrid TCN-LSTM model performs multi-step forecasting across seven water-quality parameters, achieving an RMSE of 0.1417, MAE of 0.1047, and R2 of 0.3596. Second, the Adaptive Dosage Optimization Module uses PPO reinforcement learning to adjust chemical dosing, achieving a mean step reward of 1.059 compared with 1.017 for a fixed-dose baseline. The results also show that unconstrained reward optimization can lead to excessive dosing, demonstrating the need for explicit operational safeguards. Third, the SHADE anomaly detection module uses a deep autoencoder to identify sensor and process anomalies, achieving an F1 score of 0.651 under controlled fault injection. HydroSphere combines these capabilities with tiered governance, deterministic safety bounds, and human oversight to support safer and more adaptive water infrastructure. The framework provides a scalable foundation for intelligent wastewater management and supports the objectives of UN Sustainable Development Goals 6 and 13.

Abstract reproduced under arXiv (metadata CC0)

Paper details come from OpenAlex and the arXiv API, both of which publish their metadata under CC0. We show an abstract only when its licence allows reuse (CC BY, CC BY-SA, CC BY-ND, CC0, or arXiv metadata). Copyright stays with the authors and publishers, and abstracts are reproduced without changes.

Authors or publishers who want a listing changed or removed can write to jagadeesh@smartbhujal.com.