Research
Token Taxes: mitigating AGI's economic risks
Overview Research area: AI safety and ethics, specifically technical AI governance and the economics of AI-driven automation. Technical level: Intermediate. The paper combines tax-policy and economic

- arXiv
- 2603.04555
- Published
- 2026-03-04
- Authors
- Lucas Irwin, Tung-Yu Wu, Fazl Barez
AI summary
Overview
- Research area: AI safety and ethics, specifically technical AI governance and the economics of AI-driven automation.
- Technical level: Intermediate. The paper combines tax-policy and economic concepts (tax neutrality, deadweight loss, optimal taxation) with machine-learning implementation concepts (tokenization, black-box vs. white-box auditing, watermarking, TEEs, agent-based modeling).
- Scope in one sentence: A position paper arguing that the ML research community should study token taxes — usage-based surcharges on model inference applied at the point of sale — as a technical governance tool for mitigating AI's economic risks.
What This Paper Is About
AI-driven automation threatens to erode government tax bases, lower living standards, and disempower citizens, yet AI safety research has focused mainly on capability risks rather than economic ones. The authors argue that token taxes — surcharges applied to tokens used during inference and collected where the AI is used, not where the model is hosted — are uniquely enforceable through existing compute governance infrastructure. The paper's goal is to lay out a research roadmap for making token taxes technically auditable and economically understood.
Key Contributions
-
A position and framing for token taxes. The paper defines a token tax as a usage-based surcharge on model inference applied at the point of sale, and situates it within the prior literature on robot taxes.
-
Identification of two claimed advantages. Token taxes are (1) enforceable via existing compute governance infrastructure using black-box or white-box methods, and (2) usage-based rather than firm-based, capturing value where AI is used rather than where models are hosted.
-
A three-stage audit pipeline. The paper proposes staged enforcement through (1) black-box token verification, (2) norm-based tax rates, and (3) white-box audits, with open technical problems identified at each stage.
-
A five-part research roadmap and a counterargument analysis. Recommendations R1–R5 cover black-box auditing, norm taxes, agent-based modeling, self-hosted models, and comparative analysis of mitigation strategies; the paper also addresses five alternative views (A1–A5), including FLOP taxes, innovation disincentives, Jevons' Paradox, and geopolitical veto.
Main Findings
-
Tax base erosion is the core fiscal risk. The paper states that taxes on labour are the highest source of government revenue, and that as automation increases, government fiscal revenues decrease. High unemployment would also reduce consumption and increase states' fiscal costs, eroding the two main sources of public finance.
-
Multinational tax avoidance compounds the problem. Because frontier AI companies are multinational corporations, the paper argues they can shift ownership of intangible digital assets to affiliates in low-tax jurisdictions, and that existing digital services MNCs record profits in favourable jurisdictions and use strategies such as debt and earnings stripping.
-
Historical precedent: Engel's Pause (1790–1830). During this period real wages stagnated while GDP per worker increased rapidly, and the paper describes a resulting downturn in living standards, with children working in factories and a substantial uptick in diseases such as cholera. The abstract frames this as a "40-year stagnation of wages during the first industrial revolution."
-
Early empirical evidence of AI labour-market effects. The paper cites a finding that unemployment in AI-exposed early career roles rose by 16% while levels for experienced workers remained constant.
-
Wide-ranging GDP estimates. A recent White House report is cited as stating that estimates of AI's impact on global GDP range from 1% to 45%, and that AI investment already boosted US GDP by 1.3% in the first half of 2025 alone.
-
Comparison across robot-tax designs. Table 1 compares tax types on usage-based, auditability, and innovation risk. Corporation tax: no / high / high. Automation levy: yes / low / medium. Deduction disallowance: yes / low / medium. Payroll tax reform: no / high / medium. Capital gains or wealth tax: no / medium / medium. FLOP tax: yes / high / high. Token tax: yes / high / low. The authors state the token tax is unique in being both usage-based and auditable while also minimizing innovation risk.
-
Two economic stages drive tax design. Following Korinek & Lockwood (2025), the paper distinguishes Stage 1 (post-labour economy: labour's income share declines significantly but humans remain the primary consumers) from Stage 2 (AGI economy: AGI both produces and consumes, substituting human labour entirely). The paper focuses only on Stage 1, since Stage 2 occurs only in the most extreme AI forecasts.
-
A labour-tax-revenue formula. The paper presents Korinek & Lockwood's model where labor tax revenue divided by GDP equals (1 − α) / (1 + ε(1 − α)), with α the capital share of income and ε the elasticity governing the taxable labor base. As α approaches 1, labour's share of income vanishes.
-
Worked tax example. If the token tax is 10% and the cost-per-token for a given model is $1, the AI company pays $0.10 in token tax.
-
Token counts are not a standard unit. They depend on the tokenizer and the language used. The authors favor individual token tax rates for model families rather than one uniform rate. They are less concerned about tokenizer variation (the tax would incentivize providers to optimize tokenization and minimize energy usage) but consider language variation a valid concern addressed by the audit pipeline.
-
Scope of the tax. The paper states the tax should apply only to final consumption, not intermediate consumption — a business-to-business (B2B) chatbot would not be taxed. To avoid suppressing domestic small-medium enterprises (SMEs), the tax would apply only to AI companies making more than a set threshold of revenue.
-
Enforcement relies on compute providers as intermediaries. The cloud compute provider runs inference on behalf of the AI model provider, collects billed tokens, and determines tax liability. The authors explicitly do not assume the compute provider and AI model provider are the same entity; if they are, the same requirements would apply to the combined entity.
-
Misreporting is an identified threat. Since only model providers have full access to the generative process, companies can misreport token counts and inflate profits under token-based pricing.
-
Threat model for black-box auditing. The adversary is a tax-evading, profit-maximizing AI company with white-box access to architecture, tokenizers, and weights, able to strategically under-report token counts by generating multiple valid tokenizations, secretly substitute cheaper models for more expensive ones (the "model substitution problem"), and use hidden reasoning tokens not visible to the auditor. Auditors have only black-box access via cloud provider logs.
-
Current state of feasibility. The paper states that white-box audits would currently be required since black-box token audits and norm taxes are unsolved problems.
-
FLOP taxes are easier to audit but less flexible and less stable. The paper gives the example of a model provider M that initially needs F FLOPs to generate 1000 output tokens, then optimizes to generate the same 1000 tokens with F/2 FLOPs: the displaced economic value is the same, but the FLOP tax paid would be half, while the token tax paid would remain constant. The authors also argue token and FLOP taxes are not mutually exclusive and welcome hybrid proposals.
-
Norm taxes are drawn from rentier-state practice. The paper cites norm taxes as a common solution to tax evasion in rentier states such as Norway, referring to Norway's approach and noting such a system would require a methodology for continuously updating average token usage per inference run for each model type, and would require only black-box access.
-
Global inequality is framed along compute, models, and agents. Compute chips are concentrated in the "Compute North," while countries in the "Compute South" must rent them. Frontier model companies are concentrated in rich countries: OpenAI, Anthropic, Google DeepMind, and xAI are US-owned, while DeepSeek and Alibaba are Chinese-owned; Mistral (France) and G42 (UAE) are cited as second-rate providers also headquartered in wealthy countries within the Compute North. "Agentic inequality" refers to disparities in agent availability, quality, and quantity.
-
Disempowerment analogy. The paper compares AI-driven displacement to rentier states such as Venezuela, Saudi Arabia, and Oman, where income comes mostly from oil rents rather than citizens' labour, and cites the resource curse as a reason states lose incentives to look after their own people.
-
Counterargument A1: innovation. The paper acknowledges evidence that innovative firms move to countries with lower marginal tax rates, and that higher personal and corporate tax rates reduce patents, innovation output, and mobility, especially in skill-intensive sectors. It proposes threshold-based taxes, SME exemptions, and gradual phase-in as mitigations, and notes Korinek & Lockwood's argument that token taxes on final consumption constitute optimal tax policy for Stage 1 economies.
-
Counterargument A2: Jevons' Paradox. The objection is that artificially inflating prices above efficient levels blunts induced demand. The authors respond that the tax could apply only once transformative AI-driven economic growth appears, and claim the economic risks outweigh the benefit of lower prices.
-
Counterargument A3: geopolitical veto. The USA and China contain the vast majority of the AI supply chain and could veto or undermine token tax regimes. The paper cites Digital Services Tax negotiations as precedent, including US threats of retaliatory tariffs against countries refusing to revoke DSTs, and proposes regional agreements by coalitions of the willing, citing the EU's successful implementation of the GDPR and the AI Act despite U.S. opposition.
-
Counterargument A5 is incomplete in the supplied content. The final counterargument, that establishing a token tax regime would be a bureaucratic nightmare, is truncated mid-sentence ("the infrastructure required to implement the tax wil…"). The paper's full response to this objection is not available in the provided text. No empirical results, datasets, or benchmarks are reported, as this is a position paper.
Methodology in Plain English
This is a position paper, not an empirical study. The authors build their case in three moves. First, they survey existing literature from AI safety, economics, and public finance to catalog four economic risks of AI: government fiscal crises, lower living standards (via an Engel's Pause-style wage stagnation), gradual disempowerment of citizens, and global inequality. Second, they place token taxes in a comparison table against six other robot-tax designs on three dimensions — whether the tax is usage-based, how auditable it is, and how much innovation risk it carries. Third, they propose an enforcement architecture in which cloud compute providers sit between AI model providers and governments, running a staged audit chain: try black-box token verification first, fall back to norm-based rates if that fails, and escalate to white-box audits as a last resort.
To make enforcement technically credible, the authors sketch a threat model (a profit-maximizing, tax-evading AI company with white-box model access) and map research needs onto existing technical building blocks: watermarking, trusted execution environments (TEEs), hidden-reasoning-token estimation, and hardware-level attestation for self-hosted models. For the economic side, they propose agent-based modeling (ABM) to estimate cost pass-through and deadweight loss under different hypothetical AI growth scenarios. They also borrow the formal labour-tax-revenue model from Korinek & Lockwood (2025) to argue consumption-based taxes are the right instrument for the post-labour stage.
Why This Matters
Impact on research. The paper explicitly asks the ML and technical AI governance communities to treat token taxes as a research agenda rather than a settled policy. It reframes enforcement as a machine-learning problem — auditing token counts under black-box access — and connects compute governance infrastructure to tax administration. It also argues the case should be examined preemptively, before governments lose incentives to act.
Real-world applications:
- Tax administration. Regulators could cross-check reported API token usage against independent compute provider logs, with norm-based rates used as a fallback when verification fails, and separate rates developed for different languages given variable tokenization.
- International tax equity. Because the tax is collected where tokens are used rather than where models are hosted, countries without their own foundation model companies — including those in the "Compute South" — could collect receipts from AI consumption in their jurisdiction.
- Legislative design comparison. The comparison table gives policymakers a structured way to weigh token taxes against corporation tax, automation levies, deduction disallowance, payroll tax reform, capital gains or wealth taxes, and FLOP taxes.
- Economic forecasting. Agent-based modeling is proposed to simulate whether AI-intensive firms pass API cost increases on to consumers, how elastic consumer demand is for services where AI has supplanted human labour, and how to mitigate deadweight loss.
Industry relevance. The paper directly implicates cloud hyperscalers and frontier AI labs: compute providers would take on record-keeping, verification, and enforcement roles, mandating token-level usage data collection while preserving privacy. AI model companies are named as potential research partners for norm tax data (Anthropic and Perplexity are cited). The paper also notes that white-box auditing raises intellectual property concerns, which is why the roadmap prioritizes black-box methods.
Future Directions
-
Solve black-box token verification. Develop cryptographic watermarking that can certify token counts, reliably detect model type, and withstand adversarial tokenization; build black-box methods to estimate hidden reasoning token counts from question-answer pairs without reasoning traces; and extend TEEs to inference-time token auditing while preserving confidentiality and minimizing overhead.
-
Build a norm-rate methodology. Establish how to define and continuously update norm tax rates per model category, including which independent commission would set them, the optimal update schedule, and international cooperation strategies. The paper suggests partnering with AI companies to use their adoption and usage data.
-
Model the economic effects. Use agent-based modeling with economists to quantify cost pass-through, demand elasticity for AI-substituted services such as automated customer support, and deadweight loss, anchored in optimal taxation frameworks such as Ramsey (1927).
-
Address self-hosted models and compare mitigations. Companies that host internal open-source models avoid APIs and thus the compute-provider intermediary. The paper proposes hardware-level attestation via TEEs, international GPU registries, on-chip tracking mechanisms, and a broader comparison of token taxes against FLOP taxes, AI sovereign wealth funds, decentralized AI, and AI as a public utility.
Target Audience
Technical AI governance researchers and ML researchers interested in policy-relevant enforcement problems, particularly those working on auditing, watermarking, TEEs, and compute governance. It is also aimed at economists studying automation, taxation, and labour markets, and at policy analysts and legislators evaluating robot-tax proposals. Readers looking for empirical results or benchmark evaluations will not find them here; the paper is explicitly a position paper laying out a research agenda.
Authors’ abstract
The development of AGI threatens to erode government tax bases, lower living standards, and disempower citizens -- risks that make the 40-year stagnation of wages during the first industrial revolution look mild in comparison. While AI safety research has focused primarily on capability risks, comparatively little work has studied how to mitigate the economic risks of AGI. In this paper, we argue that the economic risks posed by a post-AGI world can be effectively mitigated by token taxes: usage-based surcharges on model inference applied at the point of sale. We situate token taxes within previous proposals for robot taxes and identify two key advantages: they are enforceable through existing compute governance infrastructure, and they capture value where AI is used rather than where models are hosted. For enforcement, we outline a staged audit pipeline -- black-box token verification, norm-based tax rates, and white-box audits. For impact, we highlight the need for agent-based modeling of token taxes' economic effects. Finally, we discuss alternative approaches including FLOP taxes, and how to prevent AI superpowers vetoing such measures.