基于多智能体SAC的分布式能源集群智能调度与能效优化方法
投稿时间:2026-06-15  修订日期:2026-06-30  点此下载全文
引用本文:
摘要点击次数: 30
全文下载次数: 0
作者单位邮编
王剑秋* 国能山西河曲发电有限公司 036599
中文摘要:针对高比例可再生接入下分布式能源集群源-荷-储强耦合、出力波动剧烈、集中式调度难适应大规模异构决策的问题,提出一种基于多智能体软演员-评论家(Multi-Agent Soft Actor-Critic, MASAC)的智能调度与综合能效优化方法。建立含光伏、风电、储能、可控负荷及备用柴油机的集群调度模型,引入综合能效系数作为优化指标,采用线性化DistFlow模型描述电压与支路约束,计及储能日末SOC归位与负荷电量守恒;将问题描述为部分可观测马尔可夫博弈,在集中训练-分散执行(CTDE)架构与理想假设下分析软贝尔曼算子压缩性与软策略改进单调性,推导重参数化策略梯度与温度自适应,逐项阐明了理想化理论性质与深度实现各组件间的对应关系,并明确各智能体的严格局部观测;在改进IEEE 33节点配电网上与规则法、MILP(完美预测上界)、MPC、理想预测MPC(MPC-PF)、IDQN、IDDPG、MAPPO、MADDPG等8种方法对比并设计消融实验。MASAC总成本较规则法、IDQN、IDDPG、MAPPO、MADDPG分别降低15.2%、9.7%、5.9%、3.0%、1.6%,距MILP上界仅3.5%;综合能效系数由78.5%提升至89.7%,可再生消纳率达96.83%,电压偏差降低22.5%;源-荷±20%扰动下成本标准差较MADDPG降低46.4%。
中文关键词:分布式能源集群  多智能体强化学习  软演员-评论家算法  综合能效系数  线性化潮流  智能调度
 
Multi-Agent SAC-based Intelligent Scheduling and Energy Efficiency Optimization Method for Distributed Energy Resource Clusters
Abstract:To address the strong source-load-storage coupling, stochastic renewable fluctuation, and limited scalability of centralized scheduling in distributed energy resource (DER) clusters, a multi-agent soft actor-critic (MASAC)-based intelligent scheduling and energy-efficiency optimization method is proposed. A dispatch model covering photovoltaic, wind, battery storage, controllable loads and a backup diesel generator is built, with a comprehensive energy-efficiency indicator in the objective and a linearized DistFlow model for voltage and branch constraints, plus daily SOC restoration and load energy-conservation constraints. The problem is cast as a partially observable Markov game under the centralized training with decentralized execution (CTDE) architecture; the contraction of the soft Bellman operator and the monotonic improvement of soft policy iteration are analyzed under ideal assumptions, the reparameterized policy gradient and temperature auto-tuning are derived, and the correspondence between the idealized theoretical properties and the engineering components of the deep implementation is established item by item, with each agent’s strictly local observations specified. Simulations on a modified IEEE 33-bus system compare the method with eight baselines: rule-based, MILP (perfect-forecast upper bound), MPC, MPC with perfect forecast (MPC-PF), IDQN, IDDPG, MAPPO and MADDPG, plus an ablation study. MASAC cuts total cost by 15.2%, 9.7%, 5.9%, 3.0% and 1.6% over rule-based, IDQN, IDDPG, MAPPO and MADDPG, within 3.5% of the MILP upper bound; the energy-efficiency indicator rises from 78.5% to 89.7%, renewable accommodation reaches 96.83%, and voltage deviation drops 22.5%. Under ±20% perturbations the cost standard deviation is 46.4% lower than MADDPG.
keywords:distributed energy resource cluster  multi-agent reinforcement learning  soft actor-critic  comprehensive energy-efficiency indicator  linearized power flow  intelligent scheduling
查看全文   查看/发表评论   下载pdf阅读器