← 論文一覧に戻る

AquaContam:飲料水汚染の機械学習モデルは、汚染発生場所と同様に監視対象を学習する

AquaContam: machine-learning models of drinking-water contamination learn who is monitored as much as where contamination occurs (原題)

Newton. Tyler James

EarthArXivプレプリント2026-08-24#AI×ESGOrigin: US対象セクター: water
DOI: 10.31223/x51v3g
原典: https://eartharxiv.org/repository/object/14607/download/25427/

🤖 gxceed AI 要約

日本語

AquaContamは、米国の95,223の水道システムからの400万以上のサンプルを用いて、機械学習モデルが汚染だけでなく監視プロセスを学習することを示す。鉛モデルの見かけ上の性能はターゲットリーケージによるもので、補正後は大幅に低下する。また、監視自体が不公平であり、有色人種コミュニティは過剰にサンプリングされ、低所得コミュニティは過小サンプリングされることを明らかにした。

English

AquaContam uses over 4 million samples from 95,223 U.S. water systems to show that ML models learn monitoring processes as much as contamination. Apparent lead-model performance is largely a target-leakage artifact, and monitoring is inequitable: communities of color are oversampled and low-income communities undersampled.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

日本では、水道水質監視の公平性やAI活用の議論はまだ少ないが、環境正義の観点から、今後の水質管理やAI導入におけるバイアス対策に示唆を与える。また、SSBJやESG情報開示において、環境リスクの公平な評価が求められる中で、データの質と公平性の重要性を再認識させる。

In the global GX context

Globally, this paper contributes to the growing literature on AI fairness in environmental monitoring, highlighting how administrative data can embed inequities. It offers a reproducible protocol for addressing target and spatial leakage, relevant for climate-risk modeling and ESG data quality.

👥 読者別の含意

🔬研究者:Provides a rigorous framework for detecting and correcting target leakage in environmental ML, with implications for climate-risk modeling.

🏢実務担当者:Highlights the need to audit monitoring data for bias when using AI for environmental compliance and reporting.

🏛政策担当者:Demonstrates that monitoring itself can be inequitable, informing policy on equitable water quality monitoring and data collection.

📄 Abstract(原文)

Machine learning informs drinking-water monitoring, yet administrative data record the monitoring process as much as the contamination. Using AquaContam (over 4 million samples from 95,223 U.S. water systems and ambient monitoring locations), we find provenance proxies dominate feature attributions, and an apparent lead-model AUROC of 0.962 is largely a target- leakage artifact (corrected skill 0.699, 0.550 without provenance). A provenance-reduced PFAS signal survives (in-region AUROC 0.785) but transports modestly (AUROC 0.691, 95% CI 0.615–0.769), and spatial leakage separately inflates baseline-feature AUROC 0.699 → 0.869. Monitoring is itself inequitable: communities of color are sampled 1.85× more intensively (attenuated but significant under source adjustment), with heterogeneous detection burden (1.48, 95% CI 1.24–1.88; significant in 3 of 10 regions under a spatially-aware null), whereas low-income communities are under-sampled (0.78×). Although the group false-negative gap (0.50 versus 0.44) is not significant, the model is worse-calibrated for high-people-of-color systems (calibration error 0.18 versus 0.06). AquaContam provides a reproducible two- confound protocol with one converged detection task and documented boundary tasks.

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。