gxceed GX Disclosure Dataset v0.2 (2026Q3)
gxceed GX開示データセット v0.2(2026年第3四半期) (AI 翻訳)
Kokubu, Hiroyuki
🤖 gxceed AI 要約
日本語
東証プライム上場企業の統合報告書・サステナビリティ報告書からAI支援で機械抽出したGX開示指標の四半期スナップショット。Scope 1/2/3排出量、SBTステータス、TCFD開示、CDPスコア、再エネ比率、社内炭素価格、購入炭素クレジットを収録し、出典URL・抽出信頼度・レビューフラグを付与。未収録企業は「未抽出」であり「未開示」ではないことを明示し、正規化排出量は単位確定観測のみ厳密に収録する誠実な設計。
English
A quarterly snapshot of GX disclosure metrics machine-extracted from integrated and sustainability reports of TSE Prime-listed companies, covering Scope 1/2/3 emissions, SBT status, TCFD disclosure, CDP score, renewable ratio, internal carbon price, and purchased credits, each with source URL, extraction confidence, and review flags. Honest denominators: only 73 of 200 companies are extracted (not non-disclosing), and normalized emissions exist only where units are confirmed from source PDFs. AI-assisted extraction with provenance tracking.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
SSBJ開示が始動し、有報・統合報告書でのGX情報の機械可読化が急務となる中、本データセットは日本語開示文書からのAI抽出の実装例を示す。未抽出と未開示を区別する設計は、開示データベース構築時の落とし穴を回避する実務上の指針となる。
In the global GX context
As ISSB/CSRD push for structured sustainability data, this dataset demonstrates a rigorous approach to AI-assisted extraction from Japanese corporate reports, with explicit handling of missing data and unit normalization. It offers a model for building transparent disclosure datasets that distinguish non-disclosure from non-extraction, relevant to global efforts on disclosure infrastructure.
👥 読者別の含意
🔬研究者:Provides a transparent, auditable dataset for studying Japanese GX disclosure practices and AI extraction methods.
🏢実務担当者:Offers a benchmark for internal data extraction and a reference for understanding disclosure gaps among TSE Prime companies.
🏛政策担当者:Highlights the need for standardized, machine-readable disclosure formats to improve data comparability and reduce extraction ambiguity.
📄 Abstract(原文)
A fixed quarterly snapshot of GX (Green Transformation) disclosure metrics machine-extracted from integrated reports and sustainability reports of TSE Prime-listed companies. Covers Scope 1/2/3 emissions, SBT status, TCFD disclosure, CDP score, renewable-energy ratio, internal carbon price, and purchased carbon credits, each with the source document URL, extraction confidence, and review flags. Two honest denominators. First, this release covers 73 of 200 TSE Prime companies (102 reports); companies not included are not yet extracted, not "non-disclosing" (see the extracted flag in companies.csv). Second, normalized emission values are present for only 12 rows / 23 cells. A normalized cell exists only where the raw value is present and the unit label (万t / 千t / 百万t / tCO2e) is confirmed from the source PDF, and it is the exact unrounded product of raw x unit_multiplier. Blank normalized cells mean the unit could not be confirmed from the source document; they do not mean non-disclosure. Usage notes: raw columns preserve as-reported units and are retained for audit; Scope 2 is split into scope2_tco2_market_normalized and scope2_tco2_location_normalized so market-based and location-based figures are never conflated; Scope 3 category columns are as-reported with heterogeneous units, and reports disclosing only a single category never populate the scope3_tco2 total; the recommended analysis filter is extraction_confidence >= 0.7 AND needs_review = 0; extraction is AI-assisted (model recorded per record) and should be verified against the source PDFs listed in provenance.csv. The stable identifier is url_hash, not the natural key (code, report_type, publication_year). 東証プライム上場企業の統合報告書・サステナビリティ報告書から機械抽出した GX 開示指標の四半期固定スナップショットです。出典URL・抽出信頼度・レビュー状態を各レコードに付与しています。誠実な分母を2つ明示します。プライム200社のうち抽出済みは73社(102レポート)で、未収録は「未開示」ではなく「未抽出」です。また正規化済み排出量を持つのは12行23セルのみで、これは出典PDFから単位が確定した観測に限って raw × 倍率 の厳密積として収録しているためです。空欄は「単位未確定」を意味し、未開示ではありません。詳細は同梱の README.md を参照してください。
🔗 Provenance — このレコードを発見したソース
- Zenodo https://zenodo.org/records/21800072first seen 2026-08-05 04:11:49
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。