Skip to content
AI.info

Research

Assessing Long-Term Electricity Market Design for Ambitious Decarbonization Targets using Multi-Agent Reinforcement Learning

Assessing Long-Term Electricity Market Design for Ambitious Decarbonization Targets using Multi-Agent Reinforcement Learning Overview Research area: Multi-agent reinforcement learning applied to elect

Assessing Long-Term Electricity Market Design for Ambitious Decarbonization Targets using Multi-Agent Reinforcement Learning
arXiv
2512.17444
Published
2025-12-19
Authors
Javier Gonzalez-Ruiz, Carlos Rodriguez-Pardo, Iacopo Savelli, Alice Di Bella, Massimo Tavoni

AI summary

Assessing Long-Term Electricity Market Design for Ambitious Decarbonization Targets using Multi-Agent Reinforcement Learning

Overview

Research area: Multi-agent reinforcement learning applied to electricity market design and energy policy modeling (cs.LG; arXiv:2512.17444v1, published 19 Dec 2025).

Technical level: Intermediate. Readers need basic familiarity with reinforcement learning terminology (policies, rewards, observations, action spaces) and with electricity market concepts (wholesale markets, capacity remuneration, contracts for difference).

Scope: The paper builds an open-source multi-agent reinforcement learning environment for long-term electricity market design, applies it to a stylized Italian power system, and evaluates how different market designs and policies shape decarbonization outcomes.

What This Paper Is About

Long-term electricity markets — auctions, support schemes, capacity mechanisms, and contracting arrangements — determine what generation gets built, but most planning models used to inform policy are central-planning tools that struggle to represent strategic private investors or the details of auction design. This paper develops a multi-agent reinforcement learning (MARL) model in which profit-maximizing generation companies (GENCOs) make investment decisions in a wholesale electricity market, responding to system needs, competitive dynamics, and policy signals, so that market designs can be tested before they are implemented. The model is then applied to a stylized version of the Italian electricity system and tested under varying levels of competition, market designs, and policy scenarios.

Key Contributions

  1. An open-source multi-agent environment for long-term electricity markets. The authors develop an environment that extends existing MARL implementations (which concentrate on short-term markets) to long-term investment. It allows investment in generation assets through merchant investments, a stylized Capacity Market, and a Contract for Difference (CfD) market, and can be extended to other incentive instruments and market designs. The code is available via a public repository: https://github.com/jjgonzalez2491/MARLEY_V1.

  2. A detailed hyperparameter search process for proximal policy optimization in electricity markets. Because independent learning in competitive multi-agent settings is challenging, the paper documents an extensive hyperparameter search intended to ensure that decentralized training yields market outcomes consistent with competitive behavior.

  3. A first application, to the authors' knowledge, integrating a long-term market environment with a MARL training pipeline. The framework offers near-unlimited flexibility in electricity market design and is presented as a versatile tool for evaluating ambitious decarbonization strategies.

  4. Four stated advantages over existing methods: explicit incorporation of auction mechanisms; support for comparing multiple market instances and policy layers (such as carbon taxes) within a unified framework to examine policy effectiveness and possible redundancy; agents that manage portfolios of both new investments and existing assets, enabling analysis of divergent incumbent-versus-entrant incentives; and capture of the effect of market competition on outcomes, which matters in concentrated wholesale systems.

Main Findings

  • Market design is decisive for decarbonization and price stability. The results, as stated in the abstract, highlight the critical role of market design for decarbonizing the electricity sector and for avoiding price volatility.

  • Agents learn competitive outcomes in a decentralized setting, given sufficient tuning. The paper reports that an extensive hyperparameter search ensures decentralized training yields market outcomes consistent with competitive behavior — framing this as a nontrivial result given the inherent challenges of independent learning in multi-agent settings.

  • Policy and market mechanisms can be assessed simultaneously. The proposed framework allows assessing long-term electricity markets in which multiple policy and market mechanisms interact at the same time, with market participants responding and adapting to decarbonization pathways.

  • Quantitative results are not reported in the available content. The provided text is truncated mid-way through Section 3.2.1 and does not include the training setup (Section 4), the market results (Section 5), or the conclusion (Section 6). Numerical performance metrics, exact scenario outcomes, and any quantitative comparison across market designs are therefore not available here and are not reported in this summary.

Methodology in Plain English

The market model. Generation companies are modeled as profit-maximizing entities that invest in generation and energy storage assets and participate in a wholesale electricity market. They face no equity constraints, assume unrestricted access to financing, incur capital expenditures during construction, and receive pre-tax cash flows. Hydrothermal and day-ahead operations are represented through "representative days" — sequential 24-hour windows at hourly resolution built with the TSAM Python library from renewable capacity-factor time series and Italian electricity demand projections. The system adopts a copper plate assumption, meaning no network congestion, similar to day-ahead clearing in France and Germany. Short-term bids correspond to marginal production costs; a double-sided marginal price auction dispatches resources by merit order and sets a single system price.

Investment channels. GENCOs can invest through three mutually exclusive channels. Merchant investments are freely chosen and earn day-ahead market revenues with no protection against sustained low prices but full exposure to scarcity rents. The CfD market is triggered when built renewable capacity falls short of a regulator-set renewable penetration target (the share of renewable production in total demand); auction winners receive a stylized two-way Contract for Difference guaranteeing a fixed price for total plant output, in exchange for building the project and holding the financial obligation for the contract's lifetime, usually 15 to 25 years. The Capacity Market, inspired by the Reliability Option framework, provides a premium linked to a project's contribution to system adequacy (capacity credits and auction allocations), in exchange for developing the project and protecting consumers from high-price events.

Storage. GENCOs can invest in short-duration storage (3–4 hours, such as Lithium-ion batteries). Longer-duration storage remains under incumbent agents' control for operational decisions, with agents choosing the desired state of charge for the next period at each short-term market session.

Policy instruments. Two additional levers are included: a carbon tax tied to CO₂ emissions from generation technologies, passed through to consumers via short-term market bids, and an exogenous limit on investing in specific technologies, enforceable at any point and applicable discriminately by player and by investment channel.

The learning setup. The environment follows the Gymnasium standard, specifically the multi-agent version from the RLLIB team. Each environment step represents a 24-hour equivalent short-term market session. Agents do not strategically bid into the short-term market; investment decisions are enabled every year, with investments in the different markets occurring every two environment steps and in sequence across the year. Actions are fully multi-discrete (with discretization steps increased to better represent continuous variables like auction bid prices), and Action Masking controls which agents can take which decisions and when. The reward is the net present value of cash flows, split into profits and investment costs disaggregated across merchant, Capacity Market, and CfD channels, with discounting performed inside the environment using an exogenous discount rate. Observations follow three principles: align with real-world information asymmetries (public system information versus private portfolio information), include no internal price forecasting tools so the model retains autonomy, and provide agents any information available in the market. Observations cover demand and renewable availability time series, energy mix composition, individual and system-wide reservoir levels, market information such as prices and balances in the Capacity and CfD markets, aggregated reward per technology, policy-relevant information, and the environment step time.

Algorithm choice. The model uses independent proximal policy optimization (IPPO), selected for its suitability to the decentralized and competitive environment. The authors note that because the environment's design elements are intrinsically connected to this algorithmic choice, generalizing the environment to other MARL algorithms remains a compelling area for future research.

Why This Matters

Impact on research. The paper targets a documented gap: MARL has been applied mostly to short-term electricity markets (bidding strategies), while long-term market design has been dominated by agent-based and partial-equilibrium/bi-level models. Agent-based models in the reviewed literature — EmLab-Generation, Brain-Energy, EMIS-AS, AMIRIS — and equilibrium models address capacity markets, risk aversion, and missing markets, but the authors argue that explicit modeling of auctions and their system-wide interactions is often overlooked or only partially addressed. This work positions MARL as a flexible complement rather than a replacement, combining the flexibility of agent-based models with the competitive-market focus of partial-equilibrium setups.

Real-world applications:

  • Testing capacity market and CfD design choices before committing to reform, including auction mechanics, contract length, and eligibility rules.
  • Assessing policy stacking and redundancy, such as whether a carbon tax and a support scheme deliver overlapping or complementary incentives toward the same decarbonization target.
  • Stress-testing market concentration scenarios, given that wholesale electricity systems often exhibit comparatively high market concentration and outcomes depend on competition levels.
  • Analyzing incumbent-versus-entrant dynamics, since agents manage portfolios containing both existing assets and new investments, which can reveal divergent incentives between established players and new entrants.

Industry relevance. The framework is designed to support policymakers and other stakeholders in designing, testing, and evaluating long-term markets, and it connects to live debates about hybrid electricity market models — arrangements in which governments take a stronger role in selecting large-scale investments through competition for the market while decentralized markets govern short-term operations through competition in the market. The authors note such hybrid arrangements lack quantitative evaluation in the literature and remain far from real-life implementation. The 2022 war in Ukraine disruptions and the resulting scrutiny of electricity market resilience and consumer price exposure form part of the policy backdrop.

Future Directions

  • Better network modeling. The current copper plate assumption is described as limited; the authors state that future work will concentrate on improving network representation within the MARL context, with possible pathways detailed in appendix A.1.5.
  • Exploring hybrid market designs. The environment focuses on well-established arrangements (energy-only merchant investment, Capacity Market, CfD) as a baseline for validation and comparison; the authors identify hybrid design schemes as an interesting direction for future studies.
  • Generalizing to other MARL algorithms. Because environment features are tied to the choice of independent PPO, generalizing the environment to other MARL algorithms is flagged as compelling future research.
  • Extending the incentive toolkit. The authors state that integrating other incentive instruments and accommodating different market designs can be easily extended in the current implementation.

Target Audience

This paper is most useful to energy economists and policy analysts evaluating long-term market design, to reinforcement learning researchers interested in multi-agent applications in competitive market settings, and to regulators, transmission system operators, and market operators who need tools for testing decarbonization-oriented market reform. It is also relevant to modelers working on agent-based and equilibrium approaches who want to compare methodologies, and to practitioners who require hyperparameter guidance when applying proximal policy optimization to competitive multi-agent environments. Researchers focused on implementation will find the open-source repository the most directly actionable output.

Note on completeness: the paper text supplied for this summary is truncated; the training setup, quantitative results, and conclusions sections are not included, so no numerical outcomes from the experiments can be reported.

Authors’ abstract

Electricity systems are key to transforming today's society into a carbon-free economy. Long-term electricity market mechanisms, including auctions, support schemes, and other policy instruments, are critical in shaping the electricity generation mix. In light of the need for more advanced tools to support policymakers and other stakeholders in designing, testing, and evaluating long-term markets, this work presents a multi-agent reinforcement learning model capable of capturing the key features of decarbonizing energy systems. Profit-maximizing generation companies make investment decisions in the wholesale electricity market, responding to system needs, competitive dynamics, and policy signals. The model employs independent proximal policy optimization, which was selected for suitability to the decentralized and competitive environment. Nevertheless, given the inherent challenges of independent learning in multi-agent settings, an extensive hyperparameter search ensures that decentralized training yields market outcomes consistent with competitive behavior. The model is applied to a stylized version of the Italian electricity system and tested under varying levels of competition, market designs, and policy scenarios. Results highlight the critical role of market design for decarbonizing the electricity sector and avoiding price volatility. The proposed framework allows assessing long-term electricity markets in which multiple policy and market mechanisms interact simultaneously, with market participants responding and adapting to decarbonization pathways.

Read the original paper