Safety-Constrained Deep Reinforcement Learning for Source–Load–Storage Coordinated Operation of Green Low-Carbon Data Centers
グリーン低炭素データセンターの源・負荷・蓄電協調運用のための安全制約付き深層強化学習 (AI 翻訳)
Shi Zheng, Min Xu, Ziyu Fu, Jiaojiao Deng, Yingying Hu, Yonghao Zhang, Yue Wang, Liwei Ju
🤖 gxceed AI 要約
日本語
本研究は、グリーンデータセンターの運用最適化に安全制約付き深層強化学習(DRL)を適用。再生可能エネルギー、系統電力、蓄電池、冷却負荷、計算負荷を協調制御し、制約付きマルコフ決定過程として定式化。シミュレーションにより、ルールベース制御と比較して13.1%の排出削減、95.8%の再生可能エネルギー利用率を達成した。
English
This study applies safety-constrained deep reinforcement learning to optimize the coordinated operation of green data centers, integrating renewables, grid, batteries, cooling, and computing loads. It formulates the problem as a constrained Markov decision process and uses a control barrier function-based action shield. In simulations, the proposed method achieves 13.1% emission reduction and 95.8% renewable utilization compared to rule-based control.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
日本のデータセンターはエネルギー消費の増加と温室効果ガス削減目標のプレッシャーに直面しており、本論文のAIによる運用最適化手法は、コストと排出量の両面で貢献が期待される。特に、再生可能エネルギーの統合や系統との協調に有効であり、日本のグリーン成長戦略にも合致する。
In the global GX context
Data centers are major energy consumers globally, and this paper demonstrates how AI can optimize their low-carbon operation while maintaining reliability. The safety-constrained DRL approach is relevant for integrating variable renewables and responding to carbon signals, aligning with global trends in grid-interactive efficient buildings and corporate decarbonization targets.
👥 読者別の含意
🔬研究者:Proposes a novel safe DRL architecture for multi-energy dispatch with explicit constraint projection, offering a benchmark for future work in AI-based energy management.
🏢実務担当者:Provides a framework for data center operators to reduce energy costs and carbon emissions through intelligent control of computing, cooling, and storage assets.
🏛政策担当者:Highlights the potential of AI-controlled data centers to support grid flexibility and renewable integration, informing policies on demand response and low-carbon data center standards.
📄 Abstract(原文)
Green low-carbon data centers operate as coupled cyber-energy systems whose dispatch must coordinate renewable generation, grid exchange, battery storage, cooling load, flexible computing workload, carbon-intensity signals, and reliability constraints. This study develops and evaluates a safety-constrained deep reinforcement learning framework for source–load–storage coordinated operation of a grid-connected green data center. The operating problem is formulated as a constrained Markov decision process with state variables describing the IT load, deferrable workload backlog, renewable availability, electricity price, marginal carbon intensity, battery state of charge, server-room temperature, reserve margin, and calendar context. The action space covers grid import and export, renewable utilization, storage charge and discharge, workload shifting, and cooling control. The learning architecture combines a constrained actor–critic policy, adaptive Lagrangian safety critics, and a control barrier function (CBF)-based action shield that projects unsafe actions onto an explicitly defined operating set before plant execution. The shield is specified as a low-dimensional quadratic projection over state-dependent SOC, thermal, reserve, SLA, and grid-interface constraints, while cumulative risks are priced through Lagrangian safety budgets during policy training. The evaluation uses a controlled and auditable benchmark simulation with normalized public-data-compatible profiles, declared scenarios, random seeds, neural-network settings, and mechanism-matched baselines; it is not a telemetry-based verification or hardware certification of a deployed data center. Within this declared benchmark, the proposed safe DRL controller produces a simulated 13.1% emission reduction relative to the Rule-based controller, 95.8% renewable utilization, a normalized annual cost of 0.91, and fewer boundary contacts than the tested unconstrained, Lagrangian-only, and shield-only PPO variants. These percentages are simulator outputs relative to the stated benchmark and must not be interpreted as measured field savings. The results show how separating reward learning, cumulative safety pricing, and one-step engineering projection changes low-carbon dispatch within the specified model.
🔗 Provenance — このレコードを発見したソース
- openalex https://doi.org/10.3390/en19153492first seen 2026-07-29 05:05:47
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。