gxceed GX Disclosure Dataset v0.1 (2026Q3)
gxceed GX開示データセット v0.1 (2026年第3四半期) (AI 翻訳)
Kokubu, Hiroyuki
🤖 gxceed AI 要約
日本語
東証プライム上場企業の統合報告書からAIを用いてGX開示指標(Scope1/2/3排出量、SBT、TCFD、CDPスコア、再エネ比率、内部炭素価格、購入クレジット)を機械抽出したデータセットのv0.1リリース。現時点でプライム200社中73社102レポートを収録し、抽出信頼度とレビューフラグを付与。正規化済み排出値は単位確定観測のみ収録しており、空欄は未開示ではなく単位未確定である点に注意。
English
A quarterly snapshot of GX disclosure metrics machine-extracted from integrated reports of TSE Prime-listed companies using AI. Covers Scope 1/2/3, SBT, TCFD, CDP, renewable ratio, internal carbon price, and purchased carbon credits. This v0.1 release includes 73 of 200 Prime companies (102 reports), with extraction confidence scores and review flags. Normalized emissions are provided only where units are confirmed; blank cells indicate unconfirmed units, not non-disclosure.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
本データセットは、SSBJの開示基準策定や東証による有報でのサステナビリティ情報記載義務化の動きと直接連動する。AIによる機械抽出手法は、日本語の統合報告書から複雑なGX指標を効率的に収集する実証例であり、日本企業の開示実態把握とベンチマークに貢献する。
In the global GX context
This dataset demonstrates a scalable approach to extracting structured GX metrics from Japanese integrated reports using AI-assisted methods. It directly supports global disclosure scholarship by providing a replicable methodology for TCFD/ISSB-aligned data extraction, and offers a transparent benchmark for Japanese corporate disclosure quality. The honest handling of missing values and extraction confidence is a model for disclosure dataset construction.
👥 読者別の含意
🔬研究者:AI支援によるGX開示指標の抽出手法の実証例として、開示品質の分析ベースラインを提供する。
🏢実務担当者:自社の開示状況のベンチマークやギャップ特定に活用可能。抽出信頼度と単位の制約に注意。
🏛政策担当者:開示コンプライアンスの自動監視の実現可能性を示す。報告単位の標準化の必要性を提起。
📄 Abstract(原文)
A fixed quarterly snapshot of GX (Green Transformation) disclosure metrics machine-extracted from integrated reports and sustainability reports of TSE Prime-listed companies. Covers Scope 1/2/3 emissions, SBT status, TCFD disclosure, CDP score, renewable-energy ratio, internal carbon price, and purchased carbon credits, each with the source document URL, extraction confidence, and review flags. Two honest denominators. First, this release covers 73 of 200 TSE Prime companies (102 reports); companies not included are not yet extracted, not "non-disclosing" (see the extracted flag in companies.csv). Second, normalized emission values are present for only 12 rows / 23 cells. A normalized cell exists only where the raw value is present and the unit label (万t / 千t / 百万t / tCO2e) is confirmed from the source PDF, and it is the exact unrounded product of raw x unit_multiplier. Blank normalized cells mean the unit could not be confirmed from the source document; they do not mean non-disclosure. Usage notes: raw columns preserve as-reported units and are retained for audit; Scope 2 is split into scope2_tco2_market_normalized and scope2_tco2_location_normalized so market-based and location-based figures are never conflated; Scope 3 category columns are as-reported with heterogeneous units, and reports disclosing only a single category never populate the scope3_tco2 total; the recommended analysis filter is extraction_confidence >= 0.7 AND needs_review = 0; extraction is AI-assisted (model recorded per record) and should be verified against the source PDFs listed in provenance.csv. The stable identifier is url_hash, not the natural key (code, report_type, publication_year). 東証プライム上場企業の統合報告書・サステナビリティ報告書から機械抽出した GX 開示指標の四半期固定スナップショットです。出典URL・抽出信頼度・レビュー状態を各レコードに付与しています。誠実な分母を2つ明示します。プライム200社のうち抽出済みは73社(102レポート)で、未収録は「未開示」ではなく「未抽出」です。また正規化済み排出量を持つのは12行23セルのみで、これは出典PDFから単位が確定した観測に限って raw × 倍率 の厳密積として収録しているためです。空欄は「単位未確定」を意味し、未開示ではありません。詳細は同梱の README.md を参照してください。
🔗 Provenance — このレコードを発見したソース
- Zenodo https://zenodo.org/records/21473648first seen 2026-07-22 04:12:57
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。