Research
Conversational Demand Response: Bidirectional Aggregator-Prosumer Coordination through Agentic AI
Overview Research area: Applied AI for energy systems — specifically, using agentic large language model (LLM) systems to coordinate demand response (DR) between electricity aggregators and residentia

- arXiv
- 2603.06217
- Published
- 2026-03-06
- Authors
- Reda El Makroum, Sebastian Zwickl-Bernhard, Lukas Kranzl, Hans Auer
AI summary
Overview
Research area: Applied AI for energy systems — specifically, using agentic large language model (LLM) systems to coordinate demand response (DR) between electricity aggregators and residential prosumers.
Technical level: Advanced. The paper combines LLM-based multi-agent orchestration (ReAct pattern), a mixed-integer linear programming (MILP) optimizer, and power-system DR concepts. The conceptual argument is accessible, but the architecture and formulation require familiarity with agentic AI and energy optimization.
Scope: The paper proposes and evaluates a proof-of-concept architecture for Conversational Demand Response (CDR), in which an aggregator agent and a prosumer Home Energy Management System (HEMS) agent negotiate flexibility requests through bidirectional natural language, grounded in MILP-based feasibility assessments.
What This Paper Is About
Residential demand response depends on people staying enrolled and willing to participate, but today's coordination is either fully automated — which removes transparency and user agency — or limited to one-way dispatch signals and price alerts that give prosumers no basis for informed decisions. The paper asks how agentic AI could enable coordination that is scalable like automation yet transparent and conversational for the household. Its goal is to design, implement, and benchmark a two-tier multi-agent system where aggregators and prosumers talk to each other in natural language, with an optimization tool grounding every claim about what is technically deliverable.
Key Contributions
-
A two-tier multi-agent architecture for CDR. An aggregator agent and a prosumer HEMS agent coordinate through bidirectional natural language, replacing one-way dispatch with transparent, interactive exchanges. Both follow the ReAct pattern, iterating between reasoning and action steps until a task is resolved.
-
Quantified feasibility grounding. Asset-level sub-agents can call a MILP optimizer as a tool, so the HEMS presents cost-benefit trade-offs to the prosumer in natural language before any commitment is made. The battery sub-agent uses a dual-solve procedure: a baseline solve without DR obligations, then a re-solve with a constraint forcing discharge above the aggregator's target during the requested window; if infeasible at full power, a binary search identifies the maximum deliverable level.
-
End-to-end demonstration of both coordination directions. Downstream, the aggregator dispatches a contextualized DR event and the prosumer approves or rejects it in conversation. Upstream, the prosumer initiates changes (for example, granting full asset access before a holiday) that propagate directly to the aggregator's portfolio planning.
-
An open-source release of all system components — agent prompts, orchestration logic, optimizer code, and simulation interfaces — to enable replication and further CDR research.
Main Findings
-
Interactions stay within conversational timescales. Across six benchmark scenarios, no interaction exceeds 12 seconds or 35,000 tokens. Downstream scenarios complete in 7.8–9.8 s on average; upstream scenarios resolve in 1.1–1.7 s.
-
Downstream scenarios are the computationally heavier case. They require 3–5 reasoning iterations and 2–4 tool calls. Measured means with standard deviation (n=5): acceptance 3.6 ± 0.5 iterations, 2.6 ± 0.5 tool calls, 23.4k ± 4.0k tokens, 8.3 ± 2.1 s; rejection 5.0 ± 0.7 iterations, 3.6 ± 0.5 tool calls, 34.2k ± 6.0k tokens, 9.8 ± 1.9 s; high-target 3.4 ± 0.5 iterations, 2.4 ± 0.5 tool calls, 21.8k ± 4.1k tokens, 7.8 ± 2.2 s.
-
Rejection is the most expensive scenario. It incurs the highest cost because the agent still executes the full feasibility workflow (battery state retrieval, optimizer call, sub-agent evaluation) and then additionally composes a detailed rejection explanation incorporating the prosumer's reasoning before relaying it to the aggregator.
-
Upstream scenarios are essentially free by comparison. They resolve in a single iteration with zero tool calls (availability 1.3k ± 0.0k tokens, 1.3 ± 0.8 s; preference 1.0k ± 0.0k tokens, 1.7 ± 1.6 s; new asset 1.6k ± 0.1k tokens, 1.3 ± 1.2 s), since they require only message classification and profile routing rather than optimization.
-
The optimizer adapts to harder requests. In the downstream demonstration, the operator requests 3 kW of flexibility for the 17:00–19:00 window; the HEMS reports full feasibility, a net benefit of €1.13, SoC rising from 30% to 43%, and no comfort impact. In the high-target (5 kW) scenario, the business-as-usual charging schedule alone cannot meet the target, so the optimizer plans pre-charging from forecasted PV and still achieves full feasibility.
-
The dual-solve reveals the cost of participating. In the illustrative household day, the DR schedule pre-charges the battery to a higher state of charge and forces discharge between 17:00 and 19:00, producing a rebound of increased grid import afterward. The optimizer weighs that additional import cost against the DR compensation and finds participation more economical than the self-consumption baseline.
-
Outputs are consistent across repeated runs, which the authors attribute in part to setting temperature to zero.
-
Latency under concurrent multi-household load is not validated. The paper explicitly states this remains to be tested.
Methodology in Plain English
The researchers built two conversational AI agents — one on the aggregator side, one on the household side — and let them exchange messages in ordinary language.
The aggregator agent receives market obligations or operator instructions, picks target households from registered capacity, and formulates a DR event specifying the time window, target power, and compensation rate. It deliberately does not estimate household-level feasibility; that assessment stays local, where the HEMS has direct access to asset state and prosumer context. It also receives prosumer-initiated messages, classifies them, applies asset updates and preference changes to the portfolio, and escalates contract modifications or complaints.
On the prosumer side, an LLM orchestrator delegates to specialized sub-agents, one per household load. Appliance sub-agents run single-turn interactions and handle discrete loads through simple scheduling. The battery is different: because charge, discharge, and grid exchange must be jointly optimized across all timeslots, it gets a dedicated MILP formulation that the sub-agent calls as a tool. The optimizer minimizes net electricity cost, battery degradation, and peak discharge over a 24-hour horizon at 15-minute resolution, subject to energy balance, state-of-charge dynamics, capacity limits, a constraint capping battery discharge at household demand (preventing battery-to-grid export outside DR events), peak tracking, mutual exclusivity of charge/discharge and import/export via Big-M constraints, and a self-consumption priority constraint ensuring PV surplus charges the battery before grid export. For same-day requests, the optimizer locks slots before the request arrival to the baseline solution, since past dispatch actions are irreversible. The result is a feasibility verdict (full, partial, or infeasible), the economics of participation, and the SoC trajectory.
The HEMS orchestrator runs in two modes: normal scheduling, and DR event mode, where it parses the request, reasons about which asset is best suited to respond, delegates to the corresponding sub-agent, and translates the technical result into a conversational explanation covering request details, deliverable capacity, expected impact, and compensation.
For evaluation, all agents use GPT-OSS-120B with reasoning enabled and temperature set to zero, served through the Cerebras inference API at up to 2,500 tokens per second. Optimization uses PuLP. Each scenario runs in an isolated agent session with independent context windows to prevent information leakage between the aggregator and HEMS sides. Two end-to-end exchanges demonstrate the concept, and six scenarios — three downstream (acceptance, rejection, high-target) and three upstream (availability change, preference modification, new asset registration) — are each executed five times to capture run-to-run variability. Simulation parameters include a 15 kWh battery, 8 kW charge/discharge rate, 92% round-trip efficiency, 20–90% SoC bounds, initial SoC of 30% (4.5 kWh), feed-in tariff of 0.04 EUR/kWh, degradation cost of 0.015 EUR/kWh, DR compensation of 0.20 EUR/kWh, and PV forecast of 8.5 kWh/day at 15-minute resolution (96 slots/day).
Why This Matters
Impact on research. The paper identifies a specific gap: prosumer-facing LLM applications lack aggregator integration, while aggregator-side applications lack conversational prosumer engagement. The closest prior work (Zhang et al.) deploys LLMs in a two-layer aggregator-EV framework, but there the LLMs simulate user behavior rather than enabling conversational interaction with real prosumers. CDR is positioned as the first attempt to close that loop, and the open-source release is intended to make the architecture extensible with additional assets, optimization tools, and sub-agents on both sides.
Real-world applications.
-
Residential DR enrollment and retention. Utilities and aggregators could embed conversational interfaces that explain why a request was made, which loads are affected, and what the household earns — addressing the complexity, perceived loss of control, and effort that the paper cites as primary barriers to participation.
-
Portfolio planning with household visibility. Prosumers could be told about upcoming grid needs so they can prepare (for example, pre-conditioning before an evening event) rather than reacting to a signal.
-
Verifiable, per-load compensation. Instead of flat-rate incentives, rewards can be explained per load and linked to the market conditions that generated them.
-
Life-change flexibility. When a household buys an EV or installs a heat pump, or when a prosumer leaves for a week's holiday and grants full asset access, the flexibility profile can be updated without re-enrolling in the DR program.
Industry relevance. The paper's framing is directly aimed at utilities and aggregation services, which would need to integrate such interfaces into their platforms for CDR to reach households at scale. The authors argue CDR is not meant to replace automated control for routine operations — automated scheduling continues to handle execution — but adds a conversational layer that makes the system legible to non-technical users. The reliance on cloud inference at up to 2,500 tokens per second is flagged as a deployment consideration, with edge AI and model quantization as paths to local deployment that eliminate cloud latency while preserving data privacy.
Future Directions
-
Concurrent multi-household load. Latency was measured in isolated single-household sessions; the paper states that performance under concurrent multi-household load remains to be validated.
-
Field trials comparing interfaces. The authors call for trials comparing conversational and conventional DR interfaces to validate CDR's actual impact on prosumer engagement, rather than the proof-of-concept feasibility shown here.
-
Multi-household portfolio coordination. Extending the framework beyond a single household to portfolio-level coordination is described as a natural next step.
-
Extension to higher grid layers. The authors suggest the conversational coordination pattern may apply beyond the aggregator-prosumer level to DSO-TSO interaction, where similar challenges of transparency and multi-actor coordination persist — though this is proposed as a possibility, not demonstrated.
Target Audience
Researchers and practitioners working at the intersection of agentic AI and energy systems — particularly those in demand response, home energy management, and flexibility markets. It is also relevant to utility and aggregator product teams considering conversational interfaces for prosumer engagement, and to LLM multi-agent researchers looking for a grounded, tool-using application domain with a real optimization backend. Readers seeking quantitative field evidence on participation rates or deployment economics will not find it here; the paper is an architecture and proof-of-concept contribution, with engagement impact explicitly left to future field trials.
Authors’ abstract
Residential demand response depends on sustained prosumer participation, yet existing coordination is either fully automated, or limited to one-way dispatch signals and price alerts that offer little possibility for informed decision-making. This paper introduces Conversational Demand Response (CDR), a coordination mechanism where aggregators and prosumers interact through bidirectional natural language, enabled through agentic AI. A two-tier multi-agent architecture is developed in which an aggregator agent dispatches flexibility requests and a prosumer Home Energy Management System (HEMS) assesses deliverability and cost-benefit by calling an optimization-based tool. CDR also enables prosumer-initiated upstream communication, where changes in preferences can reach the aggregator directly. Proof-of-concept evaluation shows that interactions complete in under 12 seconds. The architecture illustrates how agentic AI can bridge the aggregator-prosumer coordination gap, providing the scalability of automated DR while preserving the transparency, explainability, and user agency necessary for sustained prosumer participation. All system components, including agent prompts, orchestration logic, and simulation interfaces, are released as open source to enable reproducibility and further development.