← 論文一覧に戻る

多段階サプライチェーンにおける在庫補充と輸送スケジューリングの統合に対する炭素考慮型オフライン強化学習

Carbon-Aware Offline Reinforcement Learning for Joint Inventory Replenishment and Transportation Scheduling in Multi-Echelon Supply Chains (原題)

Pan Li, Meng-Lin Bian, Yi-Zheng Lin

Highlights in Business, Economics and Management📚 査読済 / ジャーナル2026-09-10#AI×ESGOrigin: Global経営インパクト: コスト削減対象セクター: transport
DOI: 10.54097/p40ecd94
原典: https://hbemdata.org/index.php/ojs/article/download/197/175
📄 PDF

🤖 gxceed AI 要約

日本語

本論文は、在庫補充と輸送を統合する意思決定に炭素制約を組み込むオフライン強化学習手法CA-CQLを提案する。残予算状態、炭素形状の保守的価値学習、挙動サポート事前分布、状態依存の炭素シャドウプライス、経路ごとの炭素実現可能性スクリーニングを組み合わせる。合成ベンチマークで炭素上限違反ゼロを達成しつつ、排出量を削減するがコストと充足率にはトレードオフが生じることを示す。

English

This paper proposes CA-CQL, a budget-conditioned carbon-aware conservative offline RL method for joint inventory replenishment and transportation scheduling. It combines a remaining-budget state, carbon-shaped conservative value learning, a behavior-support prior, a state-dependent carbon shadow price, and a pathwise carbon-feasibility screen. On a reproducible multi-echelon benchmark, it achieves zero carbon-cap violations and lower emissions than a CQL baseline, at a modest cost and fill-rate trade-off.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

Scope 3排出量の削減が求められる日本企業にとって、物流・在庫管理における炭素制約下の意思決定支援は、SSBJや有報での開示情報の裏付けとなる実務的示唆を提供する。特に、サプライチェーン全体の排出量可視化と削減目標達成に向けたAI活用の可能性を示す。

In the global GX context

As global disclosure frameworks (ISSB, CSRD) push companies to report and reduce Scope 3 emissions, this work offers a method to operationalize carbon constraints in supply-chain decisions. It bridges AI and sustainability by embedding carbon budgets directly into reinforcement learning, relevant for transition finance and climate-risk management.

👥 読者別の含意

🔬研究者:オフライン強化学習に炭素制約を組み込む新手法と、そのトレードオフ評価のベンチマーク設計が参考になる。

🏢実務担当者:在庫・輸送計画に炭素上限を組み込む際のAI活用可能性と、コスト・サービスとのトレードオフを理解する材料になる。

🏛政策担当者:炭素規制下での物流最適化におけるAIの役割と、遵守コストの定量的評価に関心を持つ政策担当者に示唆を与える。

📄 Abstract(原文)

Joint replenishment and transportation decisions couple service, inventory, vehicle utilization, and carbon emissions over time. Offline reinforcement learning (RL) is attractive when operational logs exist but online exploration is unsafe; however, distribution shift and cumulative carbon constraints make direct offline Q-learning unreliable. This paper proposes budget-conditioned carbon-aware conservative Q-learning (CA-CQL), a discrete-action offline method that combines a remaining-budget state, carbon-shaped conservative value learning, a learned behavior-support prior, a state-dependent carbon shadow price, and a pathwise carbon-feasibility screen. A reproducible supplier–distribution-center–three-retailer benchmark is constructed because public sales datasets lack synchronized replenishment actions, vehicle schedules, loads, and emissions. Transport emissions are calibrated with the official UK Government 2026 vehicle-kilometer conversion factors. Across five independently generated offline datasets and 250 evaluation episodes per policy, CA-CQL achieves an operating cost of 16,865 ± 404, emissions of 14,119 ± 57 kg CO₂e, a 92.18 ± 0.88% fill rate, and zero carbon-cap violations. Relative to a support-regularized CQL baseline, it reduces emissions by 110.8 kg and the violation rate by 13.6 percentage points, with a 1.36% cost increase and a 0.63-point fill-rate decrease. Ablations show that the behavior prior prevents severe under-replenishment and that the carbon screen is necessary for hard pathwise compliance. The results establish a transparent compliance–service trade-off rather than claiming universal dominance.

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。