Skip to content
AI.info

Research

Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling

Overview Research area: Applied artificial intelligence for energy systems — specifically, agentic large language model (LLM) systems applied to residential Home Energy Management Systems (HEMS), dema

Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling
arXiv
2510.26603
Published
2025-10-30
Authors
Reda El Makroum, Sebastian Zwickl-Bernhard, Lukas Kranzl

AI summary

Overview

Research area: Applied artificial intelligence for energy systems — specifically, agentic large language model (LLM) systems applied to residential Home Energy Management Systems (HEMS), demand response, and household load scheduling.

Technical level: Advanced. The paper combines multi-agent LLM orchestration (the ReAct pattern), tool/API integration, and mixed-integer linear programming (MILP) benchmarking.

Scope: The paper designs, implements, and evaluates a hierarchical multi-agent LLM system that turns plain-language household scheduling requests into cost-optimal appliance schedules, benchmarked against MILP-optimal solutions across three open-source models.

What This Paper Is About

Home energy management systems can cut electricity costs and enable demand response, but adoption is held back because users must translate everyday preferences into many well-formatted technical parameters. Existing LLM research in this space uses language models only as preprocessing tools — generating optimization code, extracting parameters, or producing recommendations that feed conventional optimizers. This paper instead puts the LLM in charge: an orchestrator agent plus three specialist agents autonomously coordinate washing machine, dishwasher, and EV charger scheduling from natural language through to device control, without example demonstrations or few-shot learning.

Key Contributions

  1. LLM as the autonomous decision-making entity, not a preprocessing tool. The system has the LLM directly reason about system state, select actions, and execute scheduling decisions, with no handoff to a conventional optimizer. Existing LLM-energy work generates code, extracts parameters, or makes recommendations for rule-based systems instead.
  2. A hierarchical multi-agent architecture combining one orchestrator agent with three specialist agents (for loads: washing machine, dishwasher, EV charger) using iterative ReAct reasoning and acting cycles — enabling multi-appliance coordination without hardcoded workflows.
  3. A multi-layer tool framework that unifies analytical capabilities, external API integration (ENTSO-E day-ahead prices, Google Calendar), and specialist-agent delegation in a single coordinated system.
  4. Full open-source release of all system components, including complete agent prompts, orchestration logic, tools, and simulation user interfaces, to enable reproducibility and further development.

Main Findings

  • Single-appliance scheduling is uniformly solved. All three models achieved 5/5 (100%) optimal washing-machine schedules. Averages: Llama-3.3-70b at 4.0 iterations, 13,122 tokens, 4.8 s; Qwen-3-32b at 4.0 iterations, 14,122 tokens, 9.0 s; GPT-OSS-120b at 4.6 iterations, 15,792 tokens, 8.0 s. Token consumption varied by roughly 20% (13,122 to 15,792 tokens) and execution time ranged from 4.8 to 9.0 seconds.

  • Multi-appliance coordination separates the models sharply. Only Llama-3.3-70b succeeded in all runs: 5/5 (100%) success, with 5/5 (100%) optimal for washing machine, dishwasher, and EV charger, at 9.0 average iterations, 32,883 tokens, and 14.7 seconds.

  • Qwen-3-32b partially coordinates. It succeeded in 1/5 (20%) of runs. In that one successful run, washing machine and dishwasher were optimal (1/1, 100% each) but the EV charger was not (0/1, 0%), using 9.0 iterations, 35,401 tokens, and 31.4 seconds.

  • GPT-OSS-120b manages only one appliance. It achieved 0/5 (0%) full success, scheduling only the washing machine (4/4, 100%) with no dishwasher or EV attempts (recorded as 0/0). Averages from those four runs: 4.5 iterations, 19,385 tokens, 16.3 seconds.

  • Coordination roughly doubles and a half the computational load. Token consumption grew from 13,122 (single-appliance) to 32,883 (multi-appliance) — a factor of 2.5 — and iterations went from 4.0 to 9.0, reflecting fetch-prices, three sequential agent delegations, and three scheduling actions.

  • Speed and efficiency do not trade off uniformly. Qwen-3-32b's tokens per appliance (11,800) were comparable to Llama-3.3-70b's (10,961), but its execution time was more than double (31.4 s vs 14.7 s). GPT-OSS-120b ran faster than Qwen (16.3 s vs 31.4 s) but slower than Llama (16.3 s vs 14.7 s), and its 19,385 tokens for a single appliance exceeded both others, indicating inefficient partial coordination.

  • Analytical queries fail without workflow guidance. At Baseline, no model used the calculate_window_sums tool (0/5, 0% for all three) and none produced correct answers. Details: Llama-3.3-70b 3.8 iterations, 13,152 tokens, 3.6 s; Qwen-3-32b 2.0 iterations, 7,812 tokens, 4.6 s; GPT-OSS-120b 2.4 iterations, 8,712 tokens, 4.9 s.

  • Minimal guidance helps two models select the tool but only one interprets it. Llama-3.3-70b invoked the tool 5/5 (100%) but was correct 0/5 (0%) (3.0 iterations, 9,987 tokens, 1.9 s). GPT-OSS-120b invoked it 5/5 (100%) and was correct 5/5 (100%) (3.0 iterations, 10,669 tokens, 2.4 s). Qwen-3-32b did neither (0/5 tool use, 0/5 correct) and had the longest time at 29.1 s.

  • Explicit workflow instructions bring all models to consistency. All three reached 5/5 (100%) tool usage and 5/5 (100%) correctness at 3.0 iterations. Token and time: Llama-3.3-70b 10,202 tokens, 1.9 s; Qwen-3-32b 19,641 tokens, 16.7 s; GPT-OSS-120b 10,606 tokens, 2.4 s. Qwen-3-32b's verbosity and its 16.7–29.1 s times were noted as potentially limiting for time-sensitive interactions.

  • The optimal schedule concentrates all loads overnight. In the MILP-optimal solution for 15 October 2025 under real Austrian day-ahead prices, the EV charger starts at 00:15, the washing machine at 02:30, and the dishwasher at 02:45, avoiding the most expensive 3-hour block (6:30–9:30 AM, marked in red). Completing times: EV charger before 6:15 AM, washing machine at 6:30 AM, dishwasher at 6:15 AM.

  • Real-time API integration stayed practical. Despite live calls for prices and calendar constraints, successful multi-appliance scheduling completed in 14.7 seconds and successful analytical queries in under 3 seconds for Llama-3.3-70b and GPT-OSS-120b.

  • A note on completeness: the provided content is truncated inside Section 4 (Discussion), so the paper's detailed treatment of comparison to traditional approaches, model selection trade-offs, prompt and context engineering, security, and limitations is only partially available here. Section 5 (Conclusion) is not included in the provided content.

Methodology in Plain English

The researchers built a two-level agent hierarchy, with every agent powered by the same LLM. A single orchestrator handles the whole conversation: it parses the user's request, fetches shared data (electricity prices, calendar events), hands scheduling tasks to specialist agents, and writes the final device setpoints.

Six tools are available to the orchestrator: get_electricity_prices (96 fifteen-minute price points in EUR/kWh from the ENTSO-E Transparency Platform), get_calendar_ev_constraint (Google Calendar via OAuth, inferring deadlines from event title and time only), calculate_window_sums (sums across consecutive price windows), call_appliance_agent (delegation), schedule_appliance (produces a 96-element binary on/off array), and finish (terminates and summarizes). The orchestrator follows the ReAct pattern — reason, act, observe, repeat — so it adapts its sequence of tool calls instead of following a hardcoded workflow. Each specialist agent, by contrast, works in a single turn: it receives prices plus its own parameters and duration (8 slots for the washing machine, 6 for the dishwasher, 24 for the EV charger), calls calculate_window_sums, and returns a start slot, duration, cost, and rationale.

Crucially, no examples are given anywhere — no sample user requests, no demonstrations of optimal schedules, no coordination workflow examples. Prompts contain only tool descriptions, task instructions, and operational guidelines.

Evaluation used three appliances (washing machine 2.0 kW, 120 minutes, 8 timeslots; dishwasher 1.8 kW, 90 minutes, 6 timeslots; EV charger 7.4 kW, 360 minutes, 24 timeslots) and three open-source models (Llama-3.3-70B, Qwen-3-32B, GPT-OSS-120B) accessed via Cerebras' free-tier API (14,400 requests per day, roughly 2,500 tokens per second), with temperature set to 0.0 for deterministic output. Agent schedules were compared against MILP-computed optima: a schedule matching the MILP start time was counted as optimal. Three scenarios were tested — single-appliance, multi-appliance (all three at once), and analytical queries under three progressive prompt-engineering stages (Baseline, Minimal Guidance, Explicit Workflow) — with five independent runs per model per scenario, giving 75 runs in total (15 single-appliance, 15 multi-appliance, 45 analytical), all conducted on 14 October 2025 using live API calls rather than cached data.

Why This Matters

Impact on research. The paper shifts the LLM's role in energy systems from interface to controller. Prior work reviewed in the paper uses LLMs to generate optimization code, extract parameters (Michelon et al. reached 88% parameter retrieval accuracy with ReAct), produce recommendations, or diagnose faults (Chen et al. reached 96.3% fault-diagnosis accuracy). This paper shows that an agentic architecture can match MILP optima for three appliances — but also that capability is highly model-dependent, since two of three evaluated models failed multi-appliance coordination. It also quantifies the gap between "the model can reason" and "the model knows when to use a tool," which is useful for anyone deploying agents.

Real-world applications:

  • Consumer HEMS interfaces: letting households state preferences in ordinary language ("run the dishwasher before I get home") instead of filling in technical parameter forms, targeting the interaction bottleneck the paper identifies.
  • Demand response aggregation: coordinating flexible residential loads against day-ahead prices to support the grid-scale flexibility the IEA projects is needed (500 GW of demand response capacity by 2030, with buildings and residential EVs around 60% of that potential).
  • Calendar-aware EV charging: using work-schedule events to infer charging deadlines automatically — the paper demonstrates this with a recurring "Working Hours - in Office" event, Monday to Friday, 8:00 AM to 6:00 PM.
  • Open-source research infrastructure: the released prompts, orchestration logic, and simulation user interfaces give other groups a reproducible base for extending agentic energy control.

Industry relevance. The framing is explicitly adoption-driven: the IEA projects HEMS deployments must grow from 4 million units in 2020 to 32.7 million units by 2030, an eightfold increase that current trajectories are not meeting. The paper's practical findings — 14.7-second multi-appliance scheduling, free-tier API feasibility, and the requirement for explicit workflow prompts — speak directly to what a deployable product would need. The model-selection result matters commercially: choosing a weaker model saved nothing in token efficiency while risking outright coordination failure.

Future Directions

  • Extend beyond three appliances and confirmed tool-selection reliability. The paper's analytical-query results show that no model autonomously recognized when to use calculate_window_sums at Baseline, and Qwen-3-32B still failed under Minimal Guidance. Making tool selection self-directed rather than prompt-specified is the open problem the authors themselves flag.
  • Address the EV coordination failure mode. Qwen-3-32B consistently failed EV integration despite perfect washing-machine and dishwasher results, while GPT-OSS-120B never attempted dishwasher or EV delegation. Whether this reflects reasoning limits, prompt structure, or calendar-constraint handling is not resolved in the provided content.
  • Resolve the capability-versus-latency trade-off. Qwen-3-32B needed 16.7–29.1 seconds even when correct. Determining whether faster inference or better prompting closes this gap is a practical next step for interactive use.
  • Investigate the deployment considerations only partially covered. The paper promises discussion of comparison to traditional approaches, model selection trade-offs, prompt and context engineering, security, and limitations; the provided content cuts off at the start of that discussion, leaving those questions open in this summary.

Target Audience

Researchers and practitioners working at the intersection of LLM agents and energy systems — particularly those building conversational or agentic interfaces for buildings and demand response. It is also relevant to HEMS product developers weighing model choice and latency, to smart-home and energy-management engineers interested in tool-calling architectures and ReAct orchestration, and to readers who want a concrete, benchmarked example of an LLM acting as the decision-maker rather than a preprocessing step. The multi-model comparison and progressive prompt-engineering experiments make it useful for applied AI teams evaluating whether a smaller or free-tier model can handle multi-step coordination.

Authors’ abstract

The electricity sector transition requires substantial increases in residential demand response capacity, yet Home Energy Management Systems (HEMS) adoption remains limited by user interaction barriers requiring translation of everyday preferences into technical parameters. While large language models have been applied to energy systems as code generators and parameter extractors, no existing implementation deploys LLMs as autonomous coordinators managing the complete workflow from natural language input to multi-appliance scheduling. This paper presents an agentic AI HEMS where LLMs autonomously coordinate multi-appliance scheduling from natural language requests to device control, achieving optimal scheduling without example demonstrations. A hierarchical architecture combining one orchestrator with three specialist agents uses the ReAct pattern for iterative reasoning, enabling dynamic coordination without hardcoded workflows while integrating Google Calendar for context-aware deadline extraction. Evaluation across three open-source models using real Austrian day-ahead electricity prices reveals substantial capability differences. Llama-3.3-70B successfully coordinates all appliances across all scenarios to match cost-optimal benchmarks computed via mixed-integer linear programming, while other models achieve perfect single-appliance performance but struggle to coordinate all appliances simultaneously. Progressive prompt engineering experiments demonstrate that analytical query handling without explicit guidance remains unreliable despite models' general reasoning capabilities. We open-source the complete system including orchestration logic, agent prompts, tools, and web interfaces to enable reproducibility, extension, and future research.

Read the original paper