Machine learning models help predict pathogen risk in drinking water sources

Water testing

Machine learning models help predict pathogen risk in drinking water sources

12 Aug, 2026

Routine water quality monitoring already generates large volumes of physicochemical data – turbidity, temperature, dissolved oxygen – but converting that into an early warning for microbial contamination has remained difficult.

A new framework combining machine learning with quantitative microbial risk assessment (QMRA) suggests these commonly measured variables can do more than confirm compliance.

They can also predict pathogen concentrations and estimate the resulting health risk, potentially narrowing the gap between routine monitoring and the slower process of dedicated microbial testing.


Find your next water monitoring solution in our international directory.


The study, published in Biocontaminant (Maximum Academic Press), was led by corresponding author Changzheng Cui of East China University of Science and Technology, with co-authors from the China National Environmental Monitoring Center and the National Engineering Research Center of Urban Water Resources.

"Routine monitoring already generates a large amount of environmental information. Our goal was to determine whether these commonly measured variables could also help us anticipate microbial contamination and translate those predictions into meaningful health risk estimates," said Cui.

Researchers collected 95 surface water samples from two drinking water sources in a city in eastern China between May 2024 and December 2025.

Three indicator bacteria – fecal coliforms, Escherichia coli and Enterococcus faecalis – were monitored alongside six pathogens: Pseudomonas aeruginosa, Salmonella spp., Shigella spp., adenovirus, norovirus and enterovirus.

The three bacterial indicators correlated well with one another but only weakly or inconsistently with the viral pathogens, indicating that bacterial indicators alone cannot be relied on to reflect viral contamination.

Six machine learning approaches were compared: multiple linear regression, least squares boosting, decision tree, support vector machine, random forest and multilayer perceptron.

Random forest and decision tree models generally performed best, and all optimised models achieved R² values above 0.75. The decision tree model reached an R² above 0.90 for P. aeruginosa. Notably, the multilayer perceptron model – not one of the tree-based approaches – produced the strongest predictions for Shigella spp., underlining that no single algorithm dominated across every target organism.

Independent data collected in January and February 2026 were used for temporal validation, supporting the models' ability to generalise beyond their original training period.

Predicted pathogen concentrations were fed into QMRA calculations expressed as disability-adjusted life years (DALYs), against the World Health Organization's tolerable risk benchmark of 10⁻⁶ DALYs per person per year for drinking water.

Most estimated risks remained below this benchmark. However, Salmonella spp., Shigella spp. and enterovirus showed probabilities of exceeding it – 31.68%, 59.70% and 20.86% respectively – under unfavourable exposure conditions, pointing to non-negligible residual risk even where routine indicators suggest compliance.

Disinfection efficiency emerged as the dominant factor driving estimated health risk across all pathogens, reinforcing the importance of stable, effective treatment performance over source-water variability alone.

To interpret the models, the researchers applied SHapley Additive exPlanations (SHAP). Turbidity was the strongest predictor for the indicator bacteria models, accounting for between 41.6% and 62.1% of predictive importance.

Temperature, dissolved oxygen, rainfall and other water quality variables contributed to varying degrees across individual pathogens, with total dissolved solids and electrical conductivity emerging as the strongest predictors for adenovirus specifically.

The authors are explicit that the framework requires further validation across different watersheds, seasons, treatment systems and land-use conditions before it could be applied more broadly.

The study also modelled exposure through drinking-water ingestion only, and neither the exact sampling city nor the two source waters are named in the published paper.

If validated more widely, integrating ML-QMRA models with real-time monitoring systems could allow water managers to flag periods of elevated microbial risk earlier than conventional testing allows, and target interventions – such as adjusting disinfection – before contamination translates into measurable health impact.

Latest News

IET 36.3 May

Explore our Digital Edition

Discover the latest news and research

Digital edition

Explore Our Other Sites

Labmate Online
Senolytic drugs reverse premature blood stem cell ageing in sickle cell disease
Explore more Arrow
Pollution Solutions Online
Leading UK biogas operator places first orders for new FlowSep technology
Explore more Arrow
Petro Online
The evolution of four-ball testing for lubricant performance analysis
Explore more Arrow
Chromatography Today
Unlock high-resolution analysis of therapeutic oligonucleotides
Explore more Arrow