サステナビリティ報告のデータギャップ解消:温室効果ガス排出量抽出のためのベンチマークデータセット
Addressing data gaps in sustainability reporting: A benchmark dataset for greenhouse gas emission extraction (原題)
Jacob Beck, Anna Steinberg, Andreas Dimmelmeier, Laia Domenech Burin, Emily Kormanyos, Maurice Fehr, Malte Schierholz
🤖 gxceed AI 要約
日本語
本論文は、139社のサステナビリティ報告書から抽出したGHG排出量指標のゴールドスタンダードデータセットを提示する。LLMを用いた抽出パイプラインと、非専門家2名による独立評価、不一致時の専門家レビューを組み合わせることで、専門家依存を抑えつつ高品質なデータを実現している。大規模な排出量抽出モデルの検証・微調整や、グリーンウォッシング分析などの下流タスクへの再利用が期待される。
English
This paper presents a gold-standard dataset of GHG emission metrics extracted from 139 sustainability reports. Using an LLM-powered extraction pipeline with two-stage independent annotation and expert review for discrepancies, it achieves high data quality while reducing expert reliance. The dataset serves as a benchmark for human and automated annotation and supports downstream tasks like greenwashing analysis.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
SSBJ基準や有価証券報告書でのサステナビリティ開示が進む日本において、企業のGHG排出量データを効率的かつ正確に抽出する手法は、投資家対応や規制当局のモニタリングに直結する。本データセットは、日本企業の開示データの品質検証や、日本語報告書への応用可能性を示唆する。
In the global GX context
As global disclosure frameworks (ISSB, CSRD, SEC climate) demand reliable Scope 1-3 data, this benchmark addresses the critical data-gap problem by providing a validated dataset and LLM extraction methodology. It supports the development of automated tools for regulators, investors, and researchers to assess corporate emissions reporting and detect greenwashing.
👥 読者別の含意
🔬研究者:LLMを用いた情報抽出やESGデータ品質評価の研究に直接活用できるベンチマークを提供する。
🏢実務担当者:自社のGHG排出量報告の精度検証や、開示データ抽出の自動化ツール開発の参考になる。
🏛政策担当者:排出量データの信頼性確保とグリーンウォッシング防止のための、自動抽出技術の可能性と限界を示す。
📄 Abstract(原文)
Abstract Reliable company-level greenhouse gas (GHG) emissions data are essential for stakeholders addressing the climate crisis. However, existing datasets are often fragmented, inconsistent, and lack transparent methodologies, making it difficult to obtain reliable emissions data. To address this challenge, we present a gold standard dataset containing emission metrics extracted from 139 sustainability reports collected from company websites. This dataset acts as an intermediate step to validate and fine-tune models for large-scale extraction of emissions data from thousands of reports. We employ a Large Language Model (LLM)-powered extraction pipeline to automatically extract emissions metrics. These values are then independently assessed by two non-expert annotators. Reports with full agreement are directly considered gold standard, while discrepancies undergo expert review in two stages, with remaining disagreements resolved through in-person discussions. This structured process ensures high data quality while reducing reliance on experts. Our dataset serves as a benchmark for human and automated annotation, with significant reuse potential for information extraction tasks in sustainable finance as well as other downstream tasks such as greenwashing analysis.
🔗 Provenance — このレコードを発見したソース
- crossref https://doi.org/10.1038/s41597-025-05664-8first seen 2026-09-13T23:24:03.669Z
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。