Skip to content
AI.info

Research

Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis

Overview Research area: Sustainable AI / energy efficiency of LLM-based web agents, benchmarking methodology, and carbon accounting for AI systems. Technical level: Intermediate. Readers need some fam

Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis
arXiv
2511.04481
Published
2025-11-06
Authors
Lars Krupp, Daniel Geißler, Vishal Banwari, Paul Lukowicz, Jakob Karolus

AI summary

Overview

Research area: Sustainable AI / energy efficiency of LLM-based web agents, benchmarking methodology, and carbon accounting for AI systems.

Technical level: Intermediate. Readers need some familiarity with LLMs, tokenization, and inference hardware, but the energy math and benchmarking procedure are explained in plain terms.

Scope: The paper quantifies the inference-time energy and CO2 cost of six web agents through empirical benchmarking (five open-source agents on eight GPUs using the Mind2Web benchmark) and theoretical estimation (MindAct and the proprietary-LLM agent LASER), and argues that energy consumption should become a standard evaluation metric for web agents.

What This Paper Is About

Web agents such as OpenAI's Operator and Google's Project Mariner can autonomously browse websites, fill in search masks, and compare price lists, but their computational cost is invisible to users and largely unmeasured by researchers. The paper's goal is to make that cost visible: it benchmarks the real energy consumption of open-source web agents and proposes an estimation method for agents built on proprietary LLMs where direct measurement is impossible. It then argues that energy efficiency should be reported alongside task performance when web agents are evaluated.

Key Contributions

  1. An empirical energy benchmark of five fully open-source web agents — AutoWebGLM, MindAct, MultiUI, Synapse and Synatra — run on the Mind2Web benchmark across eight different NVIDIA GPUs (A100-SXM4, A100-PCIe, RTX A6000, RTX 3090, H100-SXM5, H100-NVL, H200-SXM5, L40S), with each agent executed five times per GPU and energy measured via the carbontracker library.

  2. A theoretical estimation method for web agents driven by proprietary LLMs, demonstrated on LASER (GPT-4) and validated against MindAct, an open-source agent whose real energy cost could also be measured directly. The comparison quantifies how unreliable such estimation is.

  3. A direct comparison of energy against performance, showing that higher energy consumption does not translate into better task results — the most energy-efficient agent, AutoWebGLM, also achieves the highest average step success rate.

  4. A concrete proposal for new reporting metrics, including energy per benchmark, energy per token, and the number of tokens consumed, plus the argument that existing benchmarks should be augmented with energy-consumption metrics.

Main Findings

  • Energy consumption varies by an order of magnitude across agents. On the Nvidia H100-NVL GPU, Synatra consumed ten times more energy than AutoWebGLM (3.31 ± 0.04 kWh versus 0.33 ± 0.01 kWh over the full Mind2Web benchmark).

  • Efficiency and performance move together, not against each other. AutoWebGLM had the highest reported average step success rate (53.53) and the lowest energy consumption (0.33 kWh, 57.0 minutes), while Synatra had the lowest average SSR (15.85) and the highest energy consumption (3.31 kWh, 426.0 minutes). MindAct reached 43.50 average SSR at 1.22 kWh, MultiUI 34.70 at 0.82 kWh, and Synapse 21.67 at 1.74 kWh.

  • Token volume, not token-level efficiency, dominates total energy. AutoWebGLM had the highest energy per token of the benchmarked agents (e.g., 890 ± 28.3 × 10⁻⁹ kWh per token on the cross-domain split), yet the lowest total energy, because its preprocessing sharply reduced the total number of tokens processed.

  • GPU choice matters substantially. The Nvidia H100-NVL was on average the most energy-efficient GPU in the benchmarking tests, which is why the authors used it for further agent-to-agent comparison.

  • Theoretical estimation is coarse and biased high. MindAct's estimated energy was 8.5 kWh from the paper's equation (Table 5 lists 9.01 kWh for the estimated run) versus 1.22 kWh measured — an overestimation by a factor of 7. The authors attribute the gap to conservative, upper-bound token assumptions in MindAct's candidate-generation stage, where early termination, token truncation, and token reuse reduce the real load.

  • Design philosophy matters more than model scale alone. Using a conservative estimation, LASER (GPT-4, minimal preprocessing) spends approximately 10 times more energy than MindAct (DeBERTa-86M and flan-T5 XL-2.85B with heavy preprocessing), with estimated totals of 99.21 kWh for LASER versus 8.5 kWh for MindAct.

  • CO2 emissions depend heavily on the local energy mix. Table 5 reports, for the full Mind2Web benchmark, AutoWebGLM at 6 g / 149 g / 264 g CO2 for Norway / US / Australia mixes; MindAct at 24 g / 552 g / 976 g; MultiUI at 16 g / 371 g / 656 g; Synapse at 34 g / 783 g / 1392 g; Synatra at 66 g / 1499 g / 2648 g. Estimated values were 180 g / 4081 g / 7208 g for MindAct and 1984 g / 44942 g / 79368 g for LASER.

  • Consumer-friendly framing of the numbers. Converting to car travel at 248.55 g CO2 per kilometer, running AutoWebGLM on Mind2Web equates to 0.6 km of driving under US emissions, Synatra to 6 km, and a single LASER run to 181 km.

  • Transparency is a limiting factor. Completely closed-source agents such as OpenAI's Operator and Google's Project Mariner cannot be estimated at all, since no implementation details are available, and proprietary LLMs do not allow per-token energy to be measured.

Methodology in Plain English

The authors split their work into two halves.

The empirical half: They picked five web agents whose code and underlying LLMs are fully open-source and reproducible, and that already used the Mind2Web benchmark in their original evaluations. Mind2Web was chosen because it needs no external server setup, uses fixed tasks (so results stay comparable), and is built on real-world websites — 2350 tasks across 137 websites in 31 domains, with an average of 7.3 actions per task and 1135 HTML elements per website, split into cross-domain, cross-task, and cross-website sets. They modified each agent's code to call the carbontracker library, setting start and end flags to capture the actual energy drawn by the code on each GPU. Every agent ran five times on each of eight GPUs. Because the H100-NVL was the most energy-efficient GPU on average, they limited the agent-versus-agent comparison to that GPU.

The theoretical half: For agents that use a proprietary LLM, they estimated energy from published details. For MindAct they used its paper and public source code to reconstruct the pipeline: a DeBERTa-86M ranking stage that scores each cleaned DOM element, followed by a flan-T5 XL multiple-choice stage run at least 10 times over the top 50 elements. They computed the average tokens per HTML page for each tokenizer (118798 for DeBERTa, 93778 for GPT-4), took per-token energy values from their own benchmark where possible (3.77 × 10⁻⁶ Wh for DeBERTa, 9.08 × 10⁻⁶ Wh for flan-T5 XL), and multiplied through. This yielded 0.49 Wh per action and 8.5 kWh across the 2350 tasks at 7.3 actions each.

For LASER they worked from the assumption that it feeds raw, unmodified HTML to GPT-4, used leaked and expert estimates of GPT-4's architecture (about 1.8 trillion parameters, mixture-of-experts with 16 experts of roughly 111 billion parameters each, two active per forward pass, so about 222 billion active parameters), applied the conservative compute rule of 2N FLOP per forward pass (about 444 billion FLOP per token), and divided by an assumed H100 FP8 throughput. Assuming a DGX server of eight H100 GPUs at 10.2 kW maximum, 10–20% data center overhead, and about 70% power draw during inference, they arrived at roughly 6.17 × 10⁻⁵ Wh per token, 5.78 Wh per action, and 99.21 kWh in total.

Why This Matters

Impact on research: The paper argues that energy use is a missing axis in web-agent evaluation. No existing benchmark penalizes inefficient agents, so there is no incentive to build leaner systems. By showing that benchmarking energy adds little overhead and can run alongside standard performance evaluation, the authors make the case that energy-per-benchmark should become a core metric, complementing accuracy-focused measures like step success rate. The MindAct estimation-versus-measurement discrepancy (a factor of 7) is also a caution against relying on estimated energy figures, even for fully open-source agents.

Real-world applications:

  • Model and agent selection: Developers can use measured energy-per-task numbers to pick an agent architecture that meets a performance bar without wasting compute.
  • Hardware procurement: The GPU spread in the results (e.g., the H100-NVL being most efficient on average) informs which hardware to buy or rent for agent workloads.
  • User-facing carbon labeling: The authors suggest displaying estimated CO2 per task directly in agent interfaces, so users can see the environmental cost of their interactions.
  • Data center and grid planning: The CO2 figures under Norwegian, US, and Australian energy mixes show how siting decisions change the climate impact of identical workloads.

Industry relevance: Vendors of web agents (including those behind Operator and Project Mariner), cloud providers, and benchmark maintainers all have a stake. Companies are already responding to AI energy demand in divergent ways — the paper cites Google investing in nuclear power plants for its data centers on one side, and Mistral publishing lifecycle energy reporting standards on the other. The paper frames reporting as a competitive and reputational lever, and proposes that developers of proprietary-LLM agents disclose at minimum energy per token and the number of tokens consumed on established benchmarks such as Mind2Web.

Future Directions

  1. Standardizing energy reporting for proprietary agents. Closed-source systems remain unmeasurable and, in the case of completely closed agents, unestimable. The paper calls for at least two mandatory disclosures — energy per token and tokens consumed — but no enforcement mechanism or standard format yet exists.

  2. Improving estimation accuracy for proprietary LLMs. The factor-of-7 overestimate on MindAct, and the reliance on leaked GPT-4 specifications for LASER, show that current estimation is only useful at the order-of-magnitude level. Better methods would need real per-token energy data.

  3. Extending energy-aware benchmarking beyond Mind2Web and the six agents studied. The proposed metrics could be tested on other web-agent benchmarks and on newer agents, including the fully closed-source ones the authors could not evaluate.

  4. Connecting inference energy to broader life-cycle costs. The paper deliberately focuses on inference only, arguing it will outweigh fine-tuning cost with continued use, but it notes that the full footprint also includes resource use such as water for cooling and e-waste from hardware disposal.

Target Audience

Researchers and practitioners working on LLM-based agents, AI benchmarking, and sustainable computing; hardware and infrastructure engineers choosing GPUs for agent workloads; sustainability and policy staff at AI companies tracking environmental reporting; and benchmark designers who want to add energy metrics to existing evaluations. Readers with a general interest in AI's carbon footprint will find the consumer-facing framing (kilowatt-hours and kilometers driven) accessible, while the benchmarking details are most useful to those who can reproduce or extend the setup from the authors' public code repository.

Authors’ abstract

Web agents, like OpenAI's Operator and Google's Project Mariner, are powerful agentic systems pushing the boundaries of Large Language Models (LLM). They can autonomously interact with the internet at the user's behest, such as navigating websites, filling search masks, and comparing price lists. Though web agent research is thriving, induced sustainability issues remain largely unexplored. To highlight the urgency of this issue, we provide an initial exploration of the energy and $CO_2$ cost associated with web agents from both a theoretical -via estimation- and an empirical perspective -by benchmarking. Our results show how different philosophies in web agent creation can severely impact the associated expended energy, and that more energy consumed does not necessarily equate to better results. We highlight a lack of transparency regarding disclosing model parameters and processes used for some web agents as a limiting factor when estimating energy consumption. Our work contributes towards a change in thinking of how we evaluate web agents, advocating for dedicated metrics measuring energy consumption in benchmarks.

Read the original paper