GHGbench: A Unified Multi-Entity, Multi-Task Benchmark for Carbon Emission Prediction
GHGbench: 炭素排出予測のための統合マルチエンティティ・マルチタスクベンチマーク (AI 翻訳)
Yifan Duan, Si Zheng, Lihuan Li, Chaoyuan Xue, F. Salim
🤖 gxceed AI 要約
日本語
GHGbenchは、企業・建物レベルの温室効果ガス排出予測のためのオープンデータセットとベンチマークを提供する。企業トラックは12,000社超のScope1+2およびScope3開示データを含み、建物トラックは13のオープンソースから統合された約50万件の建物データを収める。ベースライン実験により、建物排出予測が企業より困難であること、分布内・外のギャップがモデル間差を凌ぐこと、マルチモーダルリモートセンシングが特定の汎化課題を克服することを示す。
English
GHGbench provides an open dataset and benchmark for company- and building-level greenhouse-gas prediction. The company track covers 12,000+ firms with Scope 1+2 and Scope 3 disclosures; the building track harmonizes 491,591 records from 13 sources across 26 metro areas. Baseline experiments reveal that building emissions are harder to predict, the in-distribution vs. out-of-distribution gap dominates model differences, and multimodal remote-sensing embeddings help where tabular generalization fails.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
日本のSSBJ基準対応において、企業のGHG排出量予測モデルの評価は重要な課題となる。GHGbenchは、自社の排出予測手法を比較評価するための統一ベンチマークを提供し、特にScope3データの推定精度向上に貢献する可能性がある。また、建物トラックは東南アジアなどとの比較にも活用できる。
In the global GX context
Globally, this benchmark addresses the fragmentation of carbon-emission prediction datasets and evaluation standards. It provides a unified framework for comparing models across company and building levels, with transfer tasks that highlight generalization challenges. This work supports TCFD/ISSB-aligned disclosure by enabling more rigorous model validation and could inform future regulatory benchmarks.
👥 読者別の含意
🔬研究者:A unified benchmark for carbon emission prediction that enables fair model comparison and reveals key challenges (OOD gap, sector lookup ceiling).
🏢実務担当者:Can be used to evaluate or benchmark internal prediction models for Scope 1/2/3 estimation, especially for companies required to disclose under SSBJ or ISSB.
🏛政策担当者:Highlights the need for standardized validation protocols for emission estimation models and the difficulty of cross-region transfer.
📄 Abstract(原文)
Open datasets and benchmarks for entity-level carbon-emission prediction remain fragmented across access, scale, granularity, and evaluation. We introduce GHGbench, an open dataset and benchmark for company- and building-level greenhouse-gas prediction. The company track contains 32,000+ company-year records from 12,000+ firms with Scope 1+2 and Scope 3 disclosures and financial/sectoral signals; the building track harmonises 491,591 building-year records from 13 open sources into a single schema across 26 metropolitan areas (10 U.S., 15 Australian, 1 Singaporean), with climate covariates and multimodal remote-sensing embeddings. GHGbench defines canonical splits with in-distribution and cross-region/city transfer as primary tasks and temporal hold-out plus short-horizon forecasting as supplementary appendix evidence; headline baselines span gradient-boosted trees, a tabular foundation model, MLP, FT-Transformer, and multimodal fusion, with an LLM panel as auxiliary, all evaluated under multi-seed paired-bootstrap tests. Three benchmark-level findings emerge: (i) building emissions are structurally harder than company emissions; (ii) the in-distribution to out-of-distribution gap dwarfs any within-model gap across both the company track and the building track, and a tabular foundation model is, to our knowledge, the first baseline to open a paired-bootstrap-significant gap over tuned trees on a multi-city building-emissions task; (iii) multimodal remote-sensing embeddings help precisely where tabular generalisation breaks. GHGbench also exposes catastrophic city transfer and the sector-factor lookup ceiling as systematic failure modes. Code and reconstruction recipes are available at GHGbench.
🔗 Provenance — このレコードを発見したソース
- semanticscholar https://doi.org/10.48550/arxiv.2605.13743first seen 2026-07-24 06:52:52
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。