| 融合CVaR风险调整与可行域投影的中长期电力交易深度强化学习方法 |
投稿时间:2026-04-26 修订日期:2026-05-22 点此下载全文 |
| 引用本文: |
| 摘要点击次数: 51 |
| 全文下载次数: 0 |
|
|
| 中文摘要:针对南方区域电力市场中长期交易的偏差风险暴露与多品种耦合问题,本文提出一种融合CVaR风险调整与可行域投影的深度强化学习方法。构建以多品种仓位与曲线参数为联合动作的风险敏感MDP,在PPO框架中引入CVaR驱动的风险调整优势函数以权衡收益与尾部损失,并设计可行域投影算子将曲线参数映射至功率约束可行空间。实验表明,该方法累计收益均值1.93 p.u.,偏差成本占比5.24%,约束违背率仅0.06%,显著优于基线方法。消融与扰动测试验证了风险调整与投影机制的核心贡献。该方法为电力市场中长期交易提供了收益?风险协同优化的有效方案。 |
| 中文关键词:南方区域电力市场 中长期多品种交易 近端策略优化 风险敏感深度强化学习 |
| |
| A Deep Reinforcement Learning Method for Medium- and Long-Term Electricity Trading Integrating CVaR Risk Adjustment and Feasible Region Projection |
|
|
| Abstract:To address the deviation risk exposure and multi-product coupling challenges in medium- and long-term trading of the Southern Regional power market, this paper proposes a deep reinforcement learning method that integrates CVaR risk adjustment and feasible region projection. A risk-sensitive MDP is formulated with a joint action consisting of multi-product positions and contract-curve parameters. Within the PPO framework, a CVaR-driven risk-adjusted advantage function is introduced to balance return and tail loss, and a feasible region projection operator is designed to map curve parameters into the feasible space of power constraints. Experimental results show that the proposed method achieves an average cumulative return of 1.93 p.u., a deviation cost ratio of 5.24%, and a constraint violation rate of only 0.06%, significantly outperforming baseline methods. Ablation and perturbation tests verify the core contributions of the risk adjustment and projection mechanisms. This method provides an effective solution for return-risk co-optimization in medium- and long-term electricity trading. |
| keywords:Southern Regional power market Medium- and long-term multi-product trading Proximal Policy Optimization Risk-sensitive deep reinforcement learning |
| 查看全文 查看/发表评论 下载pdf阅读器 |