← 論文一覧に戻る

上場日本企業の構造的GHG開示アクセシビリティ:EDINETとLLM支援レポート抽出を用いたエンジニアリングパイロット

Structured GHG Disclosure Accessibility for Listed Japanese Firms: An Engineering Pilot Using EDINET and LLM-Assisted Report Extraction (原題)

Hiroyuki Kokubu

SSRN Working Paperプレプリント2026-05-14#AI×ESGOrigin: JP対象セクター: cross_sector
原典: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6761458

🤖 gxceed AI 要約

日本語

本論文は、日本の主要上場89社のGHG排出データの構造化開示アクセシビリティを、EDINET法定開示書類からの直接抽出と、LLMを用いた任意のサステナビリティ報告書からの抽出の2手法で検証。EDINET側では23社が当期のGHG関連構造化タグを持つ一方、47社はなし。LLM抽出では88社中52社からスコープ値取得。両パイプラインにおけるスキーマ強制の欠如が比較可能なデータの障壁であると指摘。

English

This paper examines structured GHG disclosure accessibility for 89 major Japanese listed firms using EDINET statutory filings and LLM-assisted extraction from voluntary reports. Only 23 firms had current-year structured GHG tags in EDINET; LLM extraction covered 52 firms but revealed schema enforcement gaps (unit ambiguity, missing citations, type defects). The study argues that comparable GHG data require schema-level enforcement on both disclosure and consumer sides.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

SSBJ基準の開示や統合報告書におけるGHGデータの構造化が進む中、本論文はEDINETのiXBRLタグの実態とLLM抽出の課題を実証。日本の開示インフラ改善に示唆を与える。

In the global GX context

This paper contributes to global disclosure scholarship by demonstrating practical difficulties in obtaining comparable GHG data from both structured (iXBRL) and unstructured (PDF) sources, highlighting the need for schema enforcement across the data pipeline. Relevant to ISSB, CSRD, and SEC climate disclosure rules.

👥 読者別の含意

🔬研究者:Demonstrates a methodology for auditing structured GHG tag availability in EDINET and LLM extraction accuracy; highlights schema enforcement as key research agenda.

🏢実務担当者:Shows that relying solely on voluntary reports or EDINET tags without schema enforcement leads to data that is not comparable; suggests internal controls for GHG data quality.

🏛政策担当者:Provides evidence for the need to enforce consistent schema requirements (unit, scale, concept naming) in mandatory disclosure systems like EDINET to ensure comparability.

📄 抄録(日本語訳)

我々は、89社の主要な日本上場企業を対象に、構造化された温室効果ガス(GHG)排出データのアクセシビリティを、2つの相補的なアプローチを用いて検討する:EDINET法定開示書類からの直接抽出、および任意提出のサステナビリティ報告書・統合報告書からのLLM支援抽出である。キャッシュされたEDINET iXBRLファイルの抽出後監査では、構造化開示の状態が混在していることが判明した:89社中23社には、`unitRef`属性と`scale`属性を伴うCurrentYearのGHG関連`ix:nonFraction`要素が含まれていた(うち17社には少なくとも1つのスコープ別単独要素があり、そのうち12社にはScope 1単独要素があった)。一方、19社には前年度のみのGHG関連構造化要素が含まれ、47社には調査対象のコンセプトパターンに該当するGHG関連`ix:nonFraction`要素が含まれていなかった。抽出パイプラインの初期バージョンは`unitRef`属性と`scale`属性を消費せず、したがって構造化された値をゼロと報告した;本バージョンで使用された修正済みパイプラインは、89社を構造化・当年、前年度のみ、構造化タグなしのカテゴリに再分類する。LLM支援抽出は、89社中88社について、以前または現在の取得を通じて特定された任意提出のPDF報告書に対して実施され、103件の固有の報告書レベルのレコードが生成され、そのうち63件には少なくとも1つのスコープ値が含まれていた;これらを合わせると52社をカバーし、Scope 1値は49社について抽出された。任意提出報告書のレコードは、任意提出報告書側に3つの明確なスキーマ強制のギャップを示している:49社の企業レベルのScope 1レコードのうち、任意提出報告書のレコードのみから単位が一意に判明したのはわずか11件(EDINETクロスリファレンスのケースを含めると49件中14件);63件の有効な報告書レベルレコードのうち、明示的なページ引用を含むものはわずか34件;そして、いずれのレコードについてもソースの正準性は検証されなかった。4つ目のギャップは、出版前監査中に検出されたスキーマ形状および型強制の欠陥として浮上した(Scope 1フィールドに記録されたScope 1 + 2合計、フラグなしで受理された単位不明の数値、年フィールドに格納された取得URL、スカラーが期待される場所にLLMが出力したオブジェクト形状の値);5つ目のギャップは、上述したEDINET側の消費者側スキーマ認識ギャップである。明示的な単位ラベル付きEDINET参照値をLLM抽出結果と照合できた4つのケースでは、正規化された値は参照値の±10%以内に収まった。このパイロットは、制約要因が文書読み取り自体よりも、パイプラインの複数の層における未強制のスキーマ要件にあることを示唆している:構造化タグは、消費者側パイプラインがそれらのタグを意味のあるものにする属性(`unitRef`、`scale`、`contextRef`)を消費しない限り、アクセス可能で比較可能なデータにはなり得ず、LLM支援抽出は候補となる開示レコードを生成できるが、それ自体では比較可能なデータをもたらさない。我々は、比較可能なGHGデータには、パイプラインの両側でのスキーマレベルの強制が必要であると主張する:開示側(比較可能性に関連するフィールドの一貫したタグ付け)と消費者側(スキーマ属性を消費・保持するパーサー)の両方において、単位正規化、証拠のトレーサビリティ、ソースの説明責任、スキーマ形状および型強制、そして各レコードを固定するソース成果物の時間的安定性とともにである。このエンジニアリングパイロットの貢献は、パイプラインが両側で不足している箇所の実証を通じて、そのような強制されたインフラストラクチャが何を提供しなければならないかを特定することである。

AI 翻訳(deepseek-v4-flash)。 正確を期す場合は下の原文を参照してください。

📄 Abstract(原文)

We examine the accessibility of structured greenhouse gas (GHG) emission data for 89 major Japanese listed firms using two complementary approaches: direct extraction from EDINET statutory filings, and LLM-assisted extraction from voluntary sustainability and integrated reports. A post-extraction audit of the cached EDINET iXBRL files found a mixed structured-disclosure state: 23 of the 89 firms contained CurrentYear GHG-related `ix:nonFraction` elements with `unitRef` and `scale` attributes (including 17 with at least one Scope-specific standalone element, of which 12 contained a Scope 1 standalone element), while 19 contained only prior-year GHG-related structured elements and 47 contained no GHG-related `ix:nonFraction` elements under the inspected concept patterns. An initial version of the extraction pipeline did not consume the `unitRef` and `scale` attributes and therefore reported zero structured values; the corrected pipeline used in this version reclassifies the 89 firms into structured-current-year, prior-year-only, and no-structured-tag categories. LLM-assisted extraction was performed on voluntary PDF reports identified through prior or current acquisition for 88 of the 89 firms, producing 103 unique report-level records of which 63 contained at least one Scope value; together these covered 52 firms, with Scope 1 values extracted for 49 firms. The voluntary-report records exhibit three distinct schema-enforcement gaps on the voluntary-report side: only 11 of 49 firm-level Scope 1 records had units unambiguous from the voluntary-report record alone (14 of 49 once EDINET cross-reference cases are included); only 34 of 63 valid report-level records contained explicit page citations; and source canonicality was not verified for any record. A fourth gap surfaced as schema-shape and type-enforcement defects detected during pre-publication audit (a Scope 1 + 2 total recorded in the Scope 1 field, a unit-unknown numeric value admitted unflagged, retrieval URLs stored in a year field, and an LLM-emitted object-shaped value where a scalar was expected); a fifth gap is the consumer-side schema-awareness gap on the EDINET side described above. In the four cases where an explicit unit-labelled EDINET reference value could be matched to an LLM extraction, normalized values fell within ±10% of the reference. The pilot suggests that the binding constraints lie less in document reading itself than in unenforced schema requirements at multiple layers of the pipeline: structured tags can exist without becoming accessible comparable data unless consumer pipelines consume the attributes (`unitRef`, `scale`, `contextRef`) that make those tags meaningful, and LLM-assisted extraction can produce candidate disclosure records but does not by itself yield comparable data. We argue that comparable GHG data require schema-level enforcement on both sides of the pipeline: at the disclosure side (consistent tagging of comparability-relevant fields) and at the consumer side (parsers that consume and preserve schema attributes), together with unit normalization, evidence traceability, source accountability, schema-shape and type-enforcement, and temporal stability of the source artifacts that anchor each record. The contribution of this engineering pilot is to specify, by demonstration of where the pipeline falls short on both sides, what such an enforced infrastructure must provide.

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。