← 論文一覧に戻る

持続可能性報告書を「テキスト」ではなく「ページ」として読む:GHG指標の視覚的検索とVLM抽出(コード・データ)

Code and data for: Reading Sustainability Reports as Pages, Not Text: visual retrieval and VLM extraction of GHG metrics (原題)

Afonso Henrique Torres Lucas, Julianne Oliveira, ROBERTO HIGINO PEREIRA DA SILVA

Zenodo (CERN European Organization for Nuclear Research)ジャーナル2026-09-24#AI×ESGOrigin: Global経営インパクト: 調達リスク対象セクター: cross_sector
DOI: 10.5281/zenodo.22942236
原典: https://doi.org/10.5281/zenodo.22942236

🤖 gxceed AI 要約

日本語

持続可能性報告書のページを画像のまま扱い、ColQwen2.5のマルチベクトル埋め込みとQdrantで検索し、視覚言語モデルでScope 1・2(マーケット/ロケーション基準)・3の排出量を2013〜2022年分について構造化JSONとして抽出する実験のコードと出力を公開。取り込み・抽出パイプライン、実行結果、GISTゴールドスタンダードとセル単位で照合する評価ノートブックを含む。テキスト抽出を経由しない視覚的アプローチの再現性を提供する。

English

Open code and outputs for an experiment that treats sustainability report pages as images: ColQwen2.5 multi-vector embeddings in Qdrant retrieve pages, and a vision-language model extracts Scope 1, 2 (market- and location-based) and 3 emissions for 2013-2022 as structured JSON. Includes ingestion/extraction pipelines, run outputs, and a notebook evaluating results cell-by-cell against the GIST gold standard. Enables reproduction of a text-free visual approach to GHG data extraction.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

SSBJ基準・有価証券報告書でのGHG開示義務化が進む中、日本語・英語混在の統合報告書やサステナビリティ報告書から排出量データを自動抽出する実務ニーズに直結する。表やレイアウトに依存する日本の報告書様式では、テキスト抽出より視覚的アプローチが有効な可能性が高い。

In the global GX context

Directly relevant to ISSB/CSRD-driven demands for machine-readable GHG data from PDF reports, where layout-heavy tables defeat naive text parsing. The visual retrieval + VLM pipeline offers a reproducible baseline for automated Scope 1-3 extraction and benchmarking against gold standards like GIST.

👥 読者別の含意

🔬研究者:視覚的検索とVLMによる開示データ抽出の再現可能なパイプラインと評価手法を提供する。

🏢実務担当者:自社やサプライヤーの報告書からScope 1-3を自動抽出する仕組みの実装参考になる。

🏛政策担当者:XBRL等の構造化開示義務と並行し、既存PDF報告書からのデータ抽出精度を測る基準として参考になる。

📄 Abstract(原文)

Code and outputs of the experiment behind the article "Reading Sustainability Reports as Pages, Not Text: visual retrieval and VLM extraction of GHG metrics". Report pages are kept as images, retrieved with ColQwen2.5 multi-vector embeddings stored in Qdrant, and read by a vision-language model that returns Scope 1, 2 (market- and location-based) and 3 emissions for 2013-2022 as structured JSON. The repository includes the ingestion and extraction pipelines, the extractions of the reported run, and the notebook that evaluates them cell by cell against the GIST gold standard (Beck et al., 2025).

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。