← 論文一覧に戻る

gxceed GX開示データセット v0.3(2026年第3四半期)

gxceed GX Disclosure Dataset v0.3 (2026Q3) (原題)

Kokubu, Hiroyuki

Zenodoデータセット2026-08-29#AI×ESGOrigin: JP経営インパクト: 調達リスク対象セクター: cross_sector
DOI: 10.5281/zenodo.22153243
原典: https://zenodo.org/records/22153243

🤖 gxceed AI 要約

日本語

東証プライム上場企業の統合報告書等からAIで機械抽出したGX開示指標の四半期スナップショット。Scope1/2/3排出量、SBT、TCFD、CDPスコア、再エネ比率、社内炭素価格、購入炭素クレジットを収録し、原典URL・抽出信頼度・レビュー状態を付与。未収録企業は未抽出であり未開示ではないこと、正規化排出量は単位確定時のみ収録することを明示し、データの透明性を重視。

English

A quarterly snapshot of GX disclosure metrics machine-extracted from integrated reports and sustainability reports of TSE Prime-listed companies. Covers Scope 1/2/3 emissions, SBT status, TCFD disclosure, CDP score, renewable ratio, internal carbon price, and purchased credits, each with source URL, extraction confidence, and review flags. Emphasizes honest denominators: non-included companies are not yet extracted, and normalized emissions exist only where units are confirmed.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

SSBJ開示が始動する中、プライム上場企業の開示実態を機械可読データとして整備する試みは、投資家対応や開示品質の検証に有用。未抽出と未開示を区別する設計は、日本の開示データの信頼性向上に貢献する。

In the global GX context

As ISSB/SSBJ disclosure becomes mainstream, this dataset offers a transparent, AI-assisted extraction of Japanese corporate GX metrics, with clear provenance and honest handling of missing data. It provides a model for building disclosure infrastructure that supports global comparability and auditability.

👥 読者別の含意

🔬研究者:Researchers can use this dataset to analyze Japanese corporate GX disclosure patterns, with clear flags for data quality and provenance.

🏢実務担当者:Corporate sustainability teams can benchmark their own disclosure against peers and identify gaps in their reporting.

🏛政策担当者:Policymakers can gauge the state of GX disclosure among TSE Prime companies and inform future disclosure requirements.

📄 Abstract(原文)

A fixed quarterly snapshot of GX (Green Transformation) disclosure metrics machine-extracted from integrated reports, sustainability reports, environmental reports, ESG data books, and CSR reports of TSE Prime-listed companies. Covers Scope 1/2/3 emissions, SBT status, TCFD disclosure, CDP score, renewable-energy ratio, internal carbon price, and purchased carbon credits, each with the source-document URL, extraction timestamp and model, extraction confidence, and review flags. Two honest denominators. First, this release covers 161 of 200 TSE Prime companies (221 reports); companies not included are not yet extracted, not "non-disclosing" (see the extracted flag in companies.csv). Second, normalized emission values are present for only 12 rows / 23 cells. A normalized cell exists only where the raw value is present and the unit label (万t / 千t / 百万t / tCO2e) is confirmed from the source PDF, and it is the exact unrounded product of raw × unit_multiplier. Blank normalized cells mean the unit could not be confirmed from the source document; they do not mean non-disclosure. Usage notes: raw columns preserve as-reported units and are retained for audit; Scope 2 is split into scope2_tco2_market_normalized and scope2_tco2_location_normalized so market-based and location-based figures are never conflated; Scope 3 category columns are as-reported with heterogeneous units, and reports disclosing only a single category never populate the scope3_tco2 total; the recommended analysis filter is extraction_confidence >= 0.7 AND needs_review = 0; extraction is AI-assisted (model recorded per record) and should be verified against the source documents listed in provenance.csv. The stable identifier is url_hash, not the natural key (code, report_type, publication_year). The pdf_sha256 field is empty in this version because source-PDF hashes were not recorded at extraction time. The available provenance consists of the source URL, extraction timestamp and model, extraction statistics, confidence, and review state. A later download of a corporate PDF cannot be assumed to be byte-identical to the document used during extraction. 東証プライム上場企業の統合報告書・サステナビリティ報告書等から機械抽出したGX開示指標の四半期固定スナップショットです。プライム200社のうち抽出済みは161社(221レポート)で、未収録は「未開示」ではなく「未抽出」です。また正規化済み排出量を持つのは12行23セルのみで、これは原典PDFから単位が確定した観測に限ってraw × 倍率の厳密積として収録しているためです。空欄は「単位未確定」を意味し、未開示ではありません。原典URL、抽出日時・モデル・統計、抽出信頼度、レビュー状態を記録しています。抽出時に原典PDFのSHA-256を記録していなかったため、本版のpdf_sha256列は空です。詳細は同梱のREADME.mdを参照してください。

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。