PCFBench:製品カーボンフットプリント推定の診断ベンチマーク
PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation (原題)
Krishna Rao, Andrew Dumit, Shaena Ulissi, Jacob Feintzeig, P. James Joyce, Daniel Frank, Steven Watson, Jonathan Glidden, Gizem Ilayda Dinc, Travis M. Kwee
🤖 gxceed AI 要約
日本語
PCFBenchは、製品カーボンフットプリント(PCF)推定を6つの独立評価可能なタスクに分解した初のベンチマーク。614項目の専門家ラベル付きデータで、分解・検索・オントロジー照合・数値抽出を評価する。8つの最先端LLMを検証し、段階的生成では精度が低下し、質量保存違反も多いことを示した。AIによるPCF推定の信頼性向上に貢献する。
English
PCFBench is the first benchmark that decomposes product carbon footprint (PCF) estimation into six independently evaluable tasks, with 614 expert-labeled items probing decomposition, retrieval, ontology matching, and numerical extraction. Testing eight frontier LLMs, it reveals that step-by-step generation degrades accuracy and mass conservation is often violated, undermining transparency for decarbonization decisions. The dataset and harness are released to drive targeted progress.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
日本ではSSBJ開示やScope3算定が進む中、AIによるPCF推定の信頼性は重要。本ベンチマークは、企業がAIを活用したカーボン会計の精度を検証し、開示の透明性を高めるための基盤となる。
In the global GX context
Globally, as ISSB and CSRD drive detailed carbon disclosure, AI-based PCF estimation must be reliable. PCFBench provides a diagnostic tool to evaluate and improve AI agents, supporting transparent and comparable product-level emissions data, which is essential for credible decarbonization claims.
👥 読者別の含意
🔬研究者:AI×ESG研究のためのPCF推定の評価基盤を提供し、LLMの限界と改善点を示す。
🏢実務担当者:自社のPCF算定にAIを導入する際のベンチマークとして活用し、算定プロセスの信頼性を確保できる。
🏛政策担当者:AIを活用したカーボン会計の品質保証に関する規制やガイドライン策定に示唆を与える。
📄 Abstract(原文)
AI systems are being deployed on high-stakes, domain-specific workflows that demand correctness not just in the final output, but at every intermediate step. One such workflow is estimating a product carbon footprint (PCF), the greenhouse-gas emissions attributable to a physical product. AI agents are increasingly being used to generate PCFs, but existing evaluations score either total emissions (hiding error sources and cancelling mistakes) or sub-tasks in isolation (missing compositional interactions). We introduce PCFBench, the first benchmark to carve PCF modeling into independently-evaluable tasks that require decomposition, retrieval, ontology matching, and numerical extraction. It comprises 614 expert-labelled items across six tasks. Together they probe reasoning under under-specification, conflicting context, and numerical constraints. Across eight frontier LLMs from four providers, no single model dominates. Although the strongest models estimate total product emissions within 2 times of declared totals on 77% of products, this rate drops to 37-58% when the PCF is generated step by step, with only 45-75% obeying mass conservation. These failures undermine the transparency practitioners need to compare products and drive decarbonization. We release the dataset and evaluation harness to support targeted progress.
🔗 Provenance — このレコードを発見したソース
- openalex https://doi.org/10.48550/arxiv.2608.27716first seen 2026-09-02 04:56:02
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。