← 論文一覧に戻る

gxceed GX開示データセット v0.1 (2026年第3四半期)

gxceed GX Disclosure Dataset v0.1 (2026Q3) (原題)

Kokubu, Hiroyuki

Zenodoデータセット2026-07-21#AI×ESGOrigin: JP経営インパクト: 調達リスク対象セクター: cross_sector
DOI: 10.5281/zenodo.21473648
原典: https://zenodo.org/records/21473648

🤖 gxceed AI 要約

日本語

東証プライム上場企業の統合報告書からAIを用いてGX開示指標(Scope1/2/3排出量、SBT、TCFD、CDPスコア、再エネ比率、内部炭素価格、購入クレジット)を機械抽出したデータセットのv0.1リリース。現時点でプライム200社中73社102レポートを収録し、抽出信頼度とレビューフラグを付与。正規化済み排出値は単位確定観測のみ収録しており、空欄は未開示ではなく単位未確定である点に注意。

English

A quarterly snapshot of GX disclosure metrics machine-extracted from integrated reports of TSE Prime-listed companies using AI. Covers Scope 1/2/3, SBT, TCFD, CDP, renewable ratio, internal carbon price, and purchased carbon credits. This v0.1 release includes 73 of 200 Prime companies (102 reports), with extraction confidence scores and review flags. Normalized emissions are provided only where units are confirmed; blank cells indicate unconfirmed units, not non-disclosure.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

本データセットは、SSBJの開示基準策定や東証による有報でのサステナビリティ情報記載義務化の動きと直接連動する。AIによる機械抽出手法は、日本語の統合報告書から複雑なGX指標を効率的に収集する実証例であり、日本企業の開示実態把握とベンチマークに貢献する。

In the global GX context

This dataset demonstrates a scalable approach to extracting structured GX metrics from Japanese integrated reports using AI-assisted methods. It directly supports global disclosure scholarship by providing a replicable methodology for TCFD/ISSB-aligned data extraction, and offers a transparent benchmark for Japanese corporate disclosure quality. The honest handling of missing values and extraction confidence is a model for disclosure dataset construction.

👥 読者別の含意

🔬研究者:AI支援によるGX開示指標の抽出手法の実証例として、開示品質の分析ベースラインを提供する。

🏢実務担当者:自社の開示状況のベンチマークやギャップ特定に活用可能。抽出信頼度と単位の制約に注意。

🏛政策担当者:開示コンプライアンスの自動監視の実現可能性を示す。報告単位の標準化の必要性を提起。

📄 抄録(日本語訳)

東証プライム上場企業の統合報告書およびサステナビリティ報告書から機械抽出したGX(グリーントランスフォーメーション)開示指標の四半期固定スナップショット。Scope 1/2/3排出量、SBT取得状況、TCFD開示、CDPスコア、再生可能エネルギー比率、社内炭素価格、購入カーボンクレジットを対象とし、各項目にソース文書URL、抽出信頼度、レビューフラグを付与。 誠実な分母を2つ明示する。第一に、本リリースは東証プライム200社のうち73社(102レポート)を対象としており、未収録の企業は「未開示」ではなく「未抽出」である(companies.csvのextractedフラグを参照)。第二に、正規化済み排出量が存在するのは12行・23セルのみである。正規化セルは、生値が存在し、かつ単位ラベル(万t / 千t / 百万t / tCO2e)が出典PDFから確認された場合にのみ存在し、raw × unit_multiplierの正確な未丸め積である。正規化セルの空欄は、出典文書から単位が確認できなかったことを意味し、非開示を意味するものではない。 使用上の注意:raw列は報告時の単位をそのまま保持し、監査用に残されている。Scope 2はscope2_tco2_market_normalizedとscope2_tco2_location_normalizedに分割され、マーケット基準とロケーション基準の数値が混同されることはない。Scope 3カテゴリ列は報告時のままの不均一な単位で記載され、単一カテゴリのみを開示するレポートはscope3_tco2合計値に値が入ることはない。推奨される分析フィルターはextraction_confidence >= 0.7 かつ needs_review = 0である。抽出はAI支援(レコードごとにモデルを記録)であり、provenance.csvに記載されたソースPDFと照合して検証すべきである。安定した識別子はurl_hashであり、自然キー(code, report_type, publication_year)ではない。

AI 翻訳(deepseek-v4-flash)。 正確を期す場合は下の原文を参照してください。

📄 Abstract(原文)

A fixed quarterly snapshot of GX (Green Transformation) disclosure metrics machine-extracted from integrated reports and sustainability reports of TSE Prime-listed companies. Covers Scope 1/2/3 emissions, SBT status, TCFD disclosure, CDP score, renewable-energy ratio, internal carbon price, and purchased carbon credits, each with the source document URL, extraction confidence, and review flags. Two honest denominators. First, this release covers 73 of 200 TSE Prime companies (102 reports); companies not included are not yet extracted, not "non-disclosing" (see the extracted flag in companies.csv). Second, normalized emission values are present for only 12 rows / 23 cells. A normalized cell exists only where the raw value is present and the unit label (万t / 千t / 百万t / tCO2e) is confirmed from the source PDF, and it is the exact unrounded product of raw x unit_multiplier. Blank normalized cells mean the unit could not be confirmed from the source document; they do not mean non-disclosure. Usage notes: raw columns preserve as-reported units and are retained for audit; Scope 2 is split into scope2_tco2_market_normalized and scope2_tco2_location_normalized so market-based and location-based figures are never conflated; Scope 3 category columns are as-reported with heterogeneous units, and reports disclosing only a single category never populate the scope3_tco2 total; the recommended analysis filter is extraction_confidence >= 0.7 AND needs_review = 0; extraction is AI-assisted (model recorded per record) and should be verified against the source PDFs listed in provenance.csv. The stable identifier is url_hash, not the natural key (code, report_type, publication_year). 東証プライム上場企業の統合報告書・サステナビリティ報告書から機械抽出した GX 開示指標の四半期固定スナップショットです。出典URL・抽出信頼度・レビュー状態を各レコードに付与しています。誠実な分母を2つ明示します。プライム200社のうち抽出済みは73社(102レポート)で、未収録は「未開示」ではなく「未抽出」です。また正規化済み排出量を持つのは12行23セルのみで、これは出典PDFから単位が確定した観測に限って raw × 倍率 の厳密積として収録しているためです。空欄は「単位未確定」を意味し、未開示ではありません。詳細は同梱の README.md を参照してください。

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。