Research
Discovering EV Charging Site Archetypes Through Few Shot Forecasting: The First U.S.-Wide Study
Overview Research area: Machine learning for energy systems — specifically time-series forecasting and clustering applied to electric vehicle (EV) charging demand across the United States. Technical l

- arXiv
- 2510.26910
- Published
- 2025-10-30
- Authors
- Kshitij Nikhal, Lucas Ackerknecht, Benjamin S. Riggan, Phillip Stahlfeld
AI summary
Overview
- Research area: Machine learning for energy systems — specifically time-series forecasting and clustering applied to electric vehicle (EV) charging demand across the United States.
- Technical level: Intermediate. Readers need some familiarity with clustering, forecasting architectures, and error metrics, but the framing and results are accessible.
- Scope: The paper integrates k-means clustering with few-shot time-series forecasting on a nationwide dataset of U.S. public DC fast-charging sites to discover a small number of interpretable "site archetypes" and show that cluster-specialized models forecast better than a single global model.
What This Paper Is About
Operators and planners need to forecast how much energy an EV charging site will sell, but most prior work relies on small datasets, treats sites in isolation or clusters them only by proximity, and struggles to predict demand at newly deployed sites with little operating history. This paper builds a framework that groups charging sites into behavior-based clusters and trains a separate forecasting model for each cluster, then tests whether those specialized models transfer knowledge to unseen sites with only a short history. The goal is to turn forecast accuracy itself into a way of segmenting infrastructure, producing archetypes that are both predictively useful and semantically interpretable.
Key Contributions
- Forecast-guided archetype discovery. Site archetypes are identified by combining clustering with few-shot forecasting, and the number of clusters is chosen based on predictive performance at unseen sites rather than on clustering metrics alone.
- Semantic characterization at national scale. The archetypes are described using a nationwide dataset, linking clusters to geography, surrounding amenities, and usage contexts, and visualized on a raster map aggregated into hexagonal cells.
- Demonstrated knowledge transfer to new sites. Archetype-specific "expert" models consistently outperform global models in few-shot forecasting, showing that global knowledge can be transferred to sites with limited operational history.
- A large-scale, industry-level dataset. The work uses the AGCharging dataset covering most U.S. public Level 3 (DC fast) charging sites, and states that a slice of the dataset will be released to support research.
Main Findings
- Optimal number of archetypes is twelve. The global model (k=1) serves as a strong baseline, while specialized models improve accuracy up to an optimal k=12. Beyond k=12, excessive clustering reduces predictive power, which the authors describe as a trade-off between specialization and cluster size.
- Expert models beat the global model overall. The paper reports that cluster-specialized models consistently outperform the global model, confirming that local specialization improves predictive accuracy. In Table 1, the largest sMAPE gains appear in volatile clusters: A8 Weekday Ramps drops from 12.72 to 8.69 sMAPE and from 1232 to 1131 RMSE, A2 Urban Corridors drops from 24.17 to 20.76 sMAPE and from 728 to 654 RMSE, and A11 Seasonal Leisure drops from 34.48 to 31.73 sMAPE and from 762 to 705 RMSE.
- Two exceptions are visible in Table 1. For A7 Mega Metro, the global model records sMAPE 12.30 versus 12.91 for the expert model, and for A1 Steady Retail the global RMSE is 614 versus 616 for the expert model. All other reported entries favor the expert model.
- Gains are largest where demand is volatile. The paper states that improvements are especially pronounced for highly variable sites such as travel corridors and vacation destinations, where shared temporal patterns such as holiday and event-driven spikes let expert models generalize effectively.
- Archetypes have distinct, interpretable signatures. A1 Steady Retail (7% of sites) shows stable all-week 9AM–5PM demand (up 1.3 σ), mostly near retail chains in cities such as Los Angeles, Chicago, Denver, and Washington DC. A2 Urban Corridors (12%) show large weekend peaks (up 2.6 σ) at travel plazas near popular destinations and metros. A3 Regional Corridors (12%) show similar peaks (up 2.1 σ) but lower baseline and higher variance, reflecting sparser corridors. A4 Mixed Urban (12%) blends commuter and leisure usage (up 1.5 σ) in dense downtowns. A5 Balanced Urban (8%) has a cleaner day/night split, moderate peaks (up 1.21 σ) and deep troughs (down 1.75 σ). A6 Commuter Corridors (15%) show smoother weekday AM/PM ramps (up 1.40 σ) versus sharper weekend peaks (up 1.9 σ).
- Three archetypes cover distinctive market segments. A7 Mega Metro (4%) is a consistent, saturated profile (up 1.0 σ) characteristic of large EV markets such as Southern California. A9 Suburban Shopping (13%) is spread across most states with midday peaking (up 1.55 σ) tied to grocery and big-box stores. A10 Emerging Metro (3.5%) appears in Houston, Atlanta, Orlando, and Dallas, where workplace, retail, and leisure usage combine but baseline demand stays volatile (down 1.76 σ and up 1.21 σ).
- Two archetypes capture extreme behavior. A11 Seasonal Leisure (4%) shows stronger weekend/holiday surges (up 2.3 σ) than weekdays (up 1.3 σ) near resorts and vacation homes. A12 Erratic (2.5%) has low baseline demand with irregular peaks (up 1.24 σ) and troughs (down 0.94 σ), and records the highest errors of any archetype: 65.46 global and 61.16 expert sMAPE, with RMSE 457 global and 377 expert.
- The best absolute accuracy occurs in dense urban clusters. A7 Mega Metro has the lowest global sMAPE at 12.30, and A8 Weekday Ramps has the lowest expert sMAPE at 8.69, along with the highest RMSE values in the table (1232 global, 1131 expert), reflecting large demand magnitudes.
- Findings hold across architectures. The authors note that similar performance gains are observed with TCN and N-BEATS architectures, with varying baseline (global) performance.
- The model infers peaks it has not seen. In forecast examples, the model identifies the timing and shape of peaks accurately, though occasional deviations in magnitude remain; in particular it successfully infers peaks in settings where no such peaks appeared in the historical data.
Methodology in Plain English
Each charging site is represented as a daily time series of total energy sold in kWh. The data is split 80% for training and 20% for testing. For test sites, the model is given only the 28 days (four weeks) immediately before the forecast, while training sites use their full available history, ranging from several weeks to multiple years. The authors describe this as "few-shot" inference because a 28-day lag is statistically insignificant for capturing seasonality and variability in demand.
Sites are described in two ways: aggregated weekly utilization profiles, plus a set of canonical features summarizing distributional characteristics, autocorrelation structure, and outlier dynamics. K-means groups sites using these representations. For every candidate cluster count from k=1 to k=20, a Temporal Fusion Transformer is trained as a specialist model for each cluster, producing a set of models. Cluster membership for an unseen site is predicted with the k-means predictor, and the matching expert model produces a seven-day forecast. A horizon of 7 days was chosen to capture both weekday and weekend demand and to align with weekly cycles in energy management, dynamic pricing, and operational planning. Two complementary metrics are used: sMAPE for relative accuracy across scales, and RMSE for magnitude and sensitivity to large deviations. Performance across values of k is compared to select the optimal granularity.
Why This Matters
Impact on research. Most prior work is constrained by small-scale datasets, models sites in isolation, forecasts regional demand, or clusters purely by proximity. This paper reframes infrastructure segmentation as a forecasting problem: the quality of prediction at unseen sites becomes the criterion for choosing how many archetypes exist. It also demonstrates that global knowledge can be transferred to newly deployed sites with little to no operational history, a setting prior work leaves largely unaddressed.
Real-world applications:
- Site selection and incentives. Archetype membership can guide incentives toward underserved areas and inform where new charging infrastructure should be sited.
- Grid and interconnection planning. Better demand forecasts from cluster-aware models strengthen interconnection studies and support grid stability.
- Energy and pricing strategy. Operators can use archetype-specific demand shapes to plan energy management, peak-shaving through dynamic pricing, battery dispatch, and curtailment.
- Underwriting and investment. Archetype membership can inform financial underwriting and investment decisions for charging assets.
Industry relevance. The dataset is drawn from the top five charging networks by utilization, covering 8000 DC fast-charging sites, which makes the analysis directly reflective of dominant market patterns rather than a narrow pilot. The stated benefits are lower costs, optimized energy and pricing strategies, and grid resilience, all tied to the broader goal of transportation decarbonization. The authors also commit to releasing a slice of the dataset for research.
Future Directions
- Longer forecast horizons. The current work uses a 7-day horizon; the authors plan to extend to longer-term prediction.
- Soft clustering. Moving beyond hard cluster assignments to soft membership, which could better represent sites that blend multiple usage patterns, such as the mixed urban and corridor archetypes.
- External signals. Incorporating external data such as events to capture the event-driven and holiday-driven spikes that define the most volatile archetypes, including A11 Seasonal Leisure and A12 Erratic.
- Open questions about transfer and granularity. The drop in performance beyond k=12 shows that over-clustering harms accuracy, but the paper does not resolve how cluster count should scale as more sites and more history become available, nor how robust the transferability findings are to sites outside the top five networks.
Target Audience
This paper is most useful to EV charging network operators and site planners, grid and utility analysts working on interconnection and load forecasting, energy researchers and graduate students in time-series forecasting and applied machine learning, and investors or underwriters evaluating charging infrastructure. Readers with a background in clustering and forecasting will get the most from the methodology, while the archetype descriptions and the geospatial map make the findings accessible to policy and planning audiences.
Authors’ abstract
The decarbonization of transportation relies on the widespread adoption of electric vehicles (EVs), which requires an accurate understanding of charging behavior to ensure cost-effective, grid-resilient infrastructure. Existing work is constrained by small-scale datasets, simple proximity-based modeling of temporal dependencies, and weak generalization to sites with limited operational history. To overcome these limitations, this work proposes a framework that integrates clustering with few-shot forecasting to uncover site archetypes using a novel large-scale dataset of charging demand. The results demonstrate that archetype-specific expert models outperform global baselines in forecasting demand at unseen sites. By establishing forecast performance as a basis for infrastructure segmentation, we generate actionable insights that enable operators to lower costs, optimize energy and pricing strategies, and support grid resilience critical to climate goals.