Research
Characterizing Agentic Flooding of Government Services
Overview Research area: AI governance and public administration — specifically how LLM-based AI agents change citizen demand on government services. Technical level: Intermediate. The paper uses no fo
- arXiv
- 2608.16603
- Published
- 2026-08-17
- Authors
- Chris Schmitz, Lewis Hammond, Alan Chan
AI summary
Overview
- Research area: AI governance and public administration — specifically how LLM-based AI agents change citizen demand on government services.
- Technical level: Intermediate. The paper uses no formal modeling; it is an empirical governance study combining qualitative coding, a case dataset, and a conceptual risk framework. Readers need basic familiarity with AI agents (LLMs with tool access and some autonomy) and public administration concepts.
- One-sentence scope: The paper defines and characterizes "agentic flooding" of government services, documents 84 alleged real-world cases across 11 jurisdictions, proposes a risk matrix for predicting which services are most exposed, and maps government response options against their trade-offs.
What This Paper Is About
AI agents can help people deal with government — summarizing complex rules, drafting formal complaints, filling in forms, and tracking deadlines. That lowers the cost of interacting with public services, which is good for access, but it can also produce sudden surges in the volume or complexity of requests that agencies were never resourced to handle. The paper names this phenomenon agentic flooding, argues it is likely already widespread, and asks three questions: is it happening, how severe is the risk, and what can governments do about it.
Key Contributions
- A definition and typology of agentic flooding. The authors define it as a surge in the volume or complexity of requests to a government service that is caused by AI agents interacting with that service and that substantially strains it. They split it into quantitative flooding (more requests) and qualitative flooding (more complex or longer requests), noting a case can be both.
- An evidence dataset of 84 cases. Compiled from a scan of 12 countries and 13 fixed service domains, the dataset (publicly released at github.com/CLSchmitz/flooding-dataset) contains only cases where a government official or reputable secondary source attributes the demand change to AI use.
- A risk matrix for service exposure. A 13-factor framework covering the likelihood and severity of flooding, built from the collected data plus theory and precedent from earlier demand surges, with three illustrative services mapped onto it.
- A mapping of government response options and their trade-offs. Responses are grouped into two strategies — suppressing demand (fees, rate limits, in-person requirements, bot blocking, benefit reductions) and increasing capacity (hiring, service redesign, deploying AI in processing) — with an argument that the fastest measures are the ones that damage equitable access.
Main Findings
-
Flooding appears to be happening widely today. The dataset contains 84 cases across 11 jurisdictions and 13 service domains, drawn from an initial pool of 2288 candidate services named during discovery (roughly 190 per country); fewer than one in twenty candidates met all three inclusion criteria, with the requirement of official or third-party attribution to AI being the biggest constraint.
-
One mechanism dominates: cheap LLM text generation. In 87% of cases, flooding occurs when LLMs help users cheaply generate legally sophisticated text. More advanced agentic behavior, such as autonomous website navigation, is not yet evident in the data.
-
Both flooding types are common, and often co-occur. 60% (N=50) of cases are coded as quantitative, 90% (N=76) as qualitative, and 50% (N=42) as both. One case (Germany, social court lawsuits) involved letters spanning over 4000 pages.
-
Attribution is mostly official. Government officials assert AI involvement — with or without third-party corroboration — in 58 of 84 cases (69%); only third-party sources assert it in the other 26 (31%).
-
Public and judicial services dominate. They account for 73% of cases versus 27% for non-service interaction channels such as participatory procedures and transparency requests. The most represented domains are Justice & Legal Services (23%, N=19), Regulatory Complaints (12%, N=10), and Benefits & Social Protection (11%, N=9).
-
Near-term risk is concentrated in financially attractive, friction-protected services. The authors posit the highest near-term risk falls on services where each successful submission is individually consequential and where resilience has historically come from friction or specialized knowledge rather than design — for example court claims and tax returns. Examples in the dataset include tax valuation objections, social court lawsuits, and civil claims.
-
Observed impacts so far are moderate. The paper reports evidence of overworked public servants, calls for additional budget, and calls to reform legal response obligations, but judges the strain modest compared to historical precedent such as mass comment campaigns. Where governments respond, those responses appear effective.
-
Governments do respond, but narrowly. Governments explicitly respond to flooding in over half (56%) of cases, usually with tightly scoped or non-binding actions. Friction-based responses appear in 14 cases (17%), redesign in 13 cases (15%), and explicit introduction of AI tools in response to flooding in 21 cases (25%).
-
Demand suppression is fast but inequitable. Fees, rate limits, in-person requirements, and bot blocking require little infrastructure change and can be deployed during an acute surge, but friction disproportionately deters poorer, less digitally literate, and otherwise vulnerable users, creating procedural inequality. Some friction measures may also become less effective as agents improve, and some are legally prohibited (for instance, charging fees for welfare applications in many jurisdictions).
-
Capacity building avoids the trade-off but is slow. Service redesign and agent deployment in processing could reduce both the likelihood and severity of flooding, but they require lead times, investment, and often cross-departmental coordination or legislative change. In the dataset, redesign measures are exclusively small changes rather than structural reform.
-
The methodology supports no causal or quantitative claims. Because demand is driven by many factors and some volume series begin rising before the release of ChatGPT, the authors explicitly state the data cannot establish prevalence of flooding or the role of agents, and likely undercounts cases because it requires explicit public attribution.
Methodology in Plain English
The authors first fix a scope: 12 countries (Australia, Brazil, Denmark, Estonia, France, Germany, Japan, the Netherlands, Singapore, South Korea, the United Kingdom, and the United States), chosen for variation in administrative tradition and digital government maturity, crossed with a fixed list of 13 government domains. Ten domains cover public services such as benefits and social protection, citizenship, and jobs and pensions; three cover non-service channels — participatory processes, regulatory complaints and reporting, and transparency and access to information.
Cases are gathered with an LLM-aided qualitative coding workflow. Each country–domain combination is scanned for candidate services during discovery; each candidate cycles through LLM calls and human review to improve and validate its data; and a human reviewer then decides whether to include it. To guard against hallucination, every case is human-verified before iteration and finalization, all LLM responses are forced into a standardized JSON schema so key fields such as agency names and counts can be checked deterministically, and cited URLs are extracted from API response metadata rather than generated by the model, guaranteeing that all sources are real and visitable.
Inclusion requires three things beyond evidence of the two definitional criteria: a plausible, specific mechanism by which AI could reduce transaction costs for that service; evidence of a change in demand patterns consistent with that mechanism, including primary volume data where available; and external attribution of the change to AI by a government source or a reputable secondary source. Annualized volume statistics from 2018 to 2025 are recorded per case. The risk matrix is constructed separately, by combining the collected data with theory and precedent from past demand surges.
Why This Matters
Impact on research. The paper puts a name and a definitional framework on a phenomenon that has been discussed only anecdotally, and it releases a structured, publicly available case dataset. It links AI capability research to public administration scholarship — notably Moynihan et al.'s taxonomy of administrative burdens (learning, compliance, and psychological costs), which the authors map against agent capabilities and enabling technologies. It also proposes that at least some of the near-term risk can be assessed without forecasting AI progress, since the service-side properties that predict exposure can be measured today.
Real-world applications:
- Government agencies can use the risk matrix to audit which of their services are most exposed and prioritize where to build resilience before a surge arrives.
- Digital identity and service design teams can prioritize integrating identity verification and standardized data formats into the most exposed services, which the paper argues reduces both the likelihood and severity of flooding.
- Legal and judicial bodies can assess procedural rules that mandate per-request processing effort, which the paper identifies as a factor increasing severity and narrowing feasible responses.
- Policy and legal advisors can examine the question of whether resilience-building measures such as bot blocking or access limits are legally permissible in their jurisdiction.
Industry relevance. The paper points to strong commercial incentives for middle-man organizations — claims management companies and AI-enabled services alike — to automate government interactions and grow claim volumes, since their business model is taking a cut of earned entitlements. It also bears on the vendors selling AI tooling to governments for intake screening, triage, drafting, and detecting AI-generated fraudulent submissions, and on the trust and safety and bot-detection markets, where it notes that CAPTCHAs no longer reliably distinguish humans from agents.
Future Directions
- Deep causal analysis of individual cases. The authors state that deeper analysis may mitigate confounding in specific cases, especially where submission contents are public, but that this is not yet feasible at scale. Follow-up work could establish causation case by case.
- Updating risk assessments as capabilities arrive. The paper frames the risk matrix as a tool whose inputs — particularly the maturity and accessibility of agent capabilities — should be revised as new information about AI capabilities becomes available.
- Testing whether capacity-building can pre-empt demand suppression. The central open question is whether governments will invest in service redesign and agent-assisted processing early enough to avoid falling back on friction-inducing measures once a surge is underway.
- Evaluating government-deployed agents. The paper notes that evaluation of agents on government tasks remains difficult, that laws often prohibit automated government decision-making, and that civil society raises concerns about decision transparency, bias, and accountability diffusion — all of which limit how far the capacity-building route can be pushed without further work.
- Understanding undercounting. Because inclusion requires explicit public attribution, the dataset likely misses cases that meet the definition of flooding, particularly those that look indistinguishable from human submissions. Note that the paper's own limitations section (Section 8) is referenced repeatedly but is not included in the available content, and the full text of the near-term recommendations (Section 7.1) is also truncated in the provided material.
Target Audience
This paper is most useful to AI governance and AI safety researchers, public administration scholars, and policy analysts working on digital government. It is also directly relevant to civil servants and agency leadership responsible for service delivery, capacity planning, or digital transformation, and to legal and judicial administrators dealing with rising caseloads. Trust and safety professionals, and companies building agents that interact with government portals, will find the discussion of friction measures and detection limits practically relevant. The writing is accessible to non-specialists, but the paper assumes comfort with the concept of tool-using LLM agents and with public administration vocabulary such as administrative burden and take-up.
Authors’ abstract
AI agents are making it easier for the public to interact with government, such as by helping them apply for benefits, understand complex policies, and make their opinions heard. Although improving service accessibility is beneficial, any resulting surges in demand could strain unprepared government services. We term such surges agentic flooding of government services ("flooding") and provide three contributions. First, based on a collected dataset of 84 potential cases of flooding across 11 jurisdictions, we posit that flooding is likely occurring widely today, mostly through large language models (LLMs) generating text cheaply. Second, we evaluate what services are most exposed to flooding. We develop a risk matrix to analyze a service's exposure, and suggest that near-term risk is highest for financially attractive, but complex services. Finally, we map possible government responses to flooding. Precedent suggests these responses will likely be sufficient to stop most cases of flooding, but the fastest to deploy - friction-inducing measures like fees - often trade off equitable access to public services. Accordingly, we close by recommending near-term actions that may allow governments to mitigate flooding without invoking this trade-off.