Abstract
Accurate emission estimates are essential, as governments worldwide report annual greenhouse gas (GHG) inventories to monitor progress toward climate targets under international requirements. However, accurate GHG accounting in wastewater treatment plants (WWTPs) across the Global South is hindered by severe data scarcity—missingness patterns vary by facility and parameter, with essential variables for carbon accounting—including energy use and nitrogen loads—frequently unavailable (missing rates: 80–100% in multiple plants). Conventional machine learning (ML) approaches that rely on algorithmic handling of incomplete data yield misleading predictions and physically implausible insights. We address this gap by integrating probabilistic principal component analysis (PCA) for multivariate imputation with interpretable ML models—Random Forest (RF) and eXtreme Gradient Boosting (XGBoost)—to predict annual CO₂-equivalent emissions across six WWTPs in Santiago de Querétaro, Mexico (2013–2024). Our framework uniquely integrates direct emissions (from on-site CH₄ and N₂O production and indirect CO₂ emissions from grid electricity) with avoided emissions—the GHGs prevented through organic load removal—which together enable a net carbon footprint assessment rarely implemented in data-limited settings. Results show that imputation quality—not just model choice—determines adequate validity; PCA-based imputation increased usable observations, corrected spurious variable rankings, and revealed treated flow and grid electricity as the dominant emission drivers—consistent with engineering principles. XGBoost achieved exceptional in-sample accuracy, yet temporal validation exposed its forecasting fragility, while RF demonstrated greater stability under real-world distributional shifts. Critically, we demonstrate that iterative PCA preserves data geometry better than Multiple Imputation by Chained Equations (MICE), avoiding artificial inflation of correlations in highly collinear systems. This study establishes that rigorous missing-data preprocessing is foundational—not optional—for credible, actionable GHG estimation in resource-constrained urban water systems. The proposed framework offers a scalable, transferable blueprint for dynamic carbon accounting and targeted mitigation in the Global South.
| Original language | English |
|---|---|
| Journal | Earth Systems and Environment |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Environmental statistics
- Missing data imputation
- SDG6, carbon footprint
- Urban water management
Fingerprint
Dive into the research topics of 'From Missing Data to Climate Action: Machine Learning for Carbon Accounting in Wastewater Systems'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver