Skip to content
AI.info

Research

Buy versus Build an LLM: A Decision Framework for Governments

Overview Research area: AI governance and public-sector technology policy (arXiv category cs.CY, listed under AI Safety & Ethics). Technical level: Intermediate. The paper is written for policy-makers

arXiv
2602.13033
Published
2026-02-13
Authors
Jiahao Lu, Ziwei Xu, William Tjhi, Junnan Li, Antoine Bosselut, Pang Wei Koh, Mohan Kankanhalli

AI summary

Overview

Research area: AI governance and public-sector technology policy (arXiv category cs.CY, listed under AI Safety & Ethics).

Technical level: Intermediate. The paper is written for policy-makers rather than engineers, but it includes a detailed technical table of model architectures, token counts, compute hours, and training costs, so readers need some familiarity with LLM terminology.

Scope in one sentence: The paper lays out a structured framework for governments deciding whether to buy access to commercial large language models, build sovereign ones, or pursue hybrid pathways, weighing sovereignty, safety, cost, resource capability, cultural fit, and sustainability.

What This Paper Is About

When a government wants to give its agencies or citizens access to large language models, it must choose between buying services from (often foreign) commercial providers, building domestic models, or combining the two. The authors argue that existing buy-or-build guidance is written for enterprises and does not address the constraints and responsibilities of states, such as legal jurisdiction over model providers, national data exposure, and multilingual service obligations.

The goal is not to recommend one option, but to surface the trade-offs, technical requirements, and real-world case studies so that policy-makers can judge which pathway fits their national context.

Key Contributions

  1. A taxonomy of acquisition pathways. The paper separates options into buy (API calling; purchasing model instances or private deployments), build from scratch, tuning based on open-source models, and hybrid approaches (licensing pre-trained base models and fine-tuning; retrieval augmented generation with a purchased model; joint ventures and sovereign cloud partnerships).

  2. A definition of "build" that is about control, not organizational form. The authors state that building refers to strategic ownership and control over model development and evolution, and can be executed by ministries, publicly funded research institutes and national laboratories, or corporatized entities such as state-owned enterprises, government-controlled joint ventures, and private firms under effective government governance influence.

  3. A pre-decision checklist of foundational considerations. This covers model categories (monolingual vs. multilingual, text-only vs. multi-modal, reasoning vs. non-reasoning, conversational vs. agentic, general-purpose vs. domain-specific, open-source vs. closed-source, and model size), user adoption and retention, usage scenarios, life-cycle planning, and the technical and resource demands of building.

  4. A compiled reference table of model specifications and training costs. Table 1 lists flagship configurations for 22 model families, giving parameters, architecture, context length, tokenizer and vocabulary size, trained tokens, chip counts, training hours, and cost where reported. The authors note that reported training costs often refer only to the "final run" rather than the full end-to-end development cost, and that "–" marks values not reported.

  5. A strategic evaluation framework spanning sovereignty, security, resource capability and sustainability, economic and regional development, cost and financial, national context fit, and the evolving cost-capability landscape. (In the provided excerpt, only the sovereignty dimension, 4.1, is partially visible.)

Main Findings

  • National AI strategies are pluralistic, not binary. Governments typically run sovereign, commercial, and open-source models side by side, using commercial models for non-sensitive or commodity tasks and seeking greater control for critical, high-risk, or strategically important applications.

  • Building examples are public; buying examples are rarely disclosed. The paper cites Singapore's government-funded SEA-LION models, the speech-focused MERaLiON series, and Phoenix small 1.0 (trained to understand materials including government policies and legal documents); Switzerland's Apertus, which aims to serve the public interest and strengthen Swiss digital sovereignty; UAE-backed efforts such as TII's Falcon alongside Jais and Command R7B Arabic; Malaysia's fully homegrown multimodal ILMU AI, which the paper says outperforms global models in Malay-language benchmarks; and France and Germany partnering with European firms such as Mistral AI and SAP SE. On the buy side it notes U.S. government contracts with OpenAI for use within the Department of Defense and a 12-month Australian government agreement with OpenAI for "provision of AI" services, and states that details of such contracts are often undisclosed, so more procurement arrangements likely exist than are publicly documented.

  • Vietnam is the paper's hybrid exemplar. Vietnam is described as incrementally assembling a sovereign AI stack combining domestic platforms, sovereign cloud infrastructure, and high-performance compute while partnering with foreign enterprises such as NVIDIA, NTT Data, and Huawei, alongside collaboration with the Canadian AI institute Mila for knowledge transfer and talent development.

  • Market concentration is extreme at the model-provider level. According to the Menlo Ventures 2025 State of Generative AI report cited in the paper, Anthropic, OpenAI, and Google collectively capture 88% of the enterprise LLM API market share, indicating user selection is heavily long-tail distributed.

  • Concentration varies by application category. Figure 3 reports Herfindahl–Hirschman Index concentration across use-case categories from OpenRouter, computed from total token usage from January to December 2025, at both creator and model level. The paper notes that HHI ranges from 10,000 for a monopoly to 2,500 for a market with four equally sized firms, and that domains such as translation, programming, and legal and health-related tasks are dominated by a small number of commercial providers.

  • Model versions turn over fast. Across major commercial providers such as OpenAI, Anthropic, and Google, model versions are frequently superseded or deprecated within months to roughly a year, so governments should plan periodic upgrades and migration windows even if they do not match that pace.

  • Software is not the only supply-chain risk. The authors recommend diversification across software vendors, hardware dependencies including GPUs, technical experts, and energy suppliers to reduce lock-in and single points of failure.

  • The concrete cost and compute numbers come with caveats. DeepSeek V3 (671B-A37B, MoE) used 14.8T training tokens, 2048 H800 GPUs, 2.788M H800 GPU hours, at a reported $5.576M; gpt-oss (120B-A5B) used 2.1M H100 GPU hours; Olmo 3.1 (32B) used 1024 H100 GPUs over roughly 1.376M H100 GPU hours at roughly $2.75M; Apertus-70B used 4096 H200 GPUs over roughly 6M H200 GPU hours. Separately, the paper states that training a mid-scale model with 70 billion parameters such as Llama 3 70B can require approximately 6.4 million GPU hours, that Llama-3.3 70B used roughly 7.0 million GPU hours for pre-training, and that the 671B DeepSeekMoE architecture with FP8 training requires only 2.8 million H800 GPU hours.

  • Failure and interruption costs are first-order. The paper reports that during the 54-day pretraining of the Llama 3 405B model, more than half of the 419 unexpected interruptions were caused by failures related to GPUs or their onboard HBM3 memory, and that at this scale even a brief hardware interruption can cause million-dollar losses in wasted compute time.

  • Talent is a binding constraint. Citing 2025 labor reports, the authors describe a thin global supply of specialists in data engineering, infrastructure systems, model training, and red-teaming, leading to "talent wars" where top-tier engineers command compensation packages that rival professional athletes, and note that the departure of a single lead architect can delay a frontier model project by months or even years.

  • Evaluation often requires building new benchmarks. The paper describes the SEA-LION case, where no high-quality evaluation datasets existed for Lao, Khmer, or regional multi-turn dialogues in Malay and Tagalog, requiring creation of new benchmarks such as SEA-HELM. It also flags that localizing test content needs extensive human expertise and cannot easily be automated or synthetically generated at scale, and that LLM-as-a-judge creates a "chicken-and-egg" problem.

  • RAG can deliver sovereignty without building a model. SkillsFuture Singapore is cited as using GraphRAG to analyze 30,000 customer cases quarterly, processing data 62.5 times faster without leaking citizen information to the model provider, with both inference and data access occurring within a controlled environment.

  • Open weights are not the same as fully open. The paper notes that most "open-source" models are in practice open-weights models, so not all details of their original training process or data are fully transparent, and that making pre-training data open-source undermines a country's unique advantage.

  • Sovereignty includes narrative control. The visible portion of the sovereignty dimension argues that LLMs control the narrative layer of citizen interaction by deciding how issues are framed and what is treated as a reasonable or factual conclusion, so relying on a third-party model may cede control over sensitive topics such as history or national identity.

Methodology in Plain English

This is a conceptual and policy-analytic paper rather than an empirical study. The authors combine three inputs. First, they survey publicly documented government LLM initiatives by country and sort them into buy, build, and hybrid categories to construct a taxonomy of acquisition pathways. Second, they synthesize technical evidence from publicly available model reports by compiling a comparison table of flagship model specifications — parameters, architecture, context length, tokenizer vocabulary size, trained tokens, chip counts, training hours, and cost — noting where figures are not reported and that cost figures typically cover only the final training run. Third, they translate those technical realities into a decision framework organized around strategic dimensions. The guidance is informed in part by the authors' own involvement in public-sector LLM efforts including SEA-LION and Apertus. A note on scope: the excerpt provided covers Sections 1 through the start of Section 4.1; the remaining dimensions of the framework are referenced in the abstract and Figure 4 but their detailed content does not appear in the excerpt.

Why This Matters

Impact on research. The paper argues that the research community has produced little structured, technically informed buy-or-build guidance tailored to public-sector decision-making, since existing buy-or-build frameworks are designed for commercial contexts. It reframes sovereign LLM development as a public-infrastructure question with measurement problems attached — such as what benchmarks a government should build when none exist for its languages — which opens a line of work on sovereign evaluation rather than purely academic leaderboards.

Real-world applications:

  • Procurement and contracting. The paper's pathway taxonomy gives agencies a way to compare per-token API access, private instances or licenses, licensed base models with in-house post-training, and sovereign-cloud joint ventures before committing to a vendor.
  • Sensitive-data services. The SkillsFuture Singapore GraphRAG example shows how a government can use a purchased model's reasoning while keeping citizen data inside a private graph database within a controlled environment.
  • Language and inclusion policy. Low-resource language coverage, illustrated by SEA-LION's support for Filipino, Burmese, Lao, Khmer, Tamil and others, and by India's Airavata and Sarvam-1, is framed as a matter of inclusive governance and equitable service delivery that commercial models may not serve.
  • Capacity and budget planning. The compute, cost, expertise, data, architectural design, evaluation, alignment, and deployment requirements catalogued for a build pathway let planners scope the resource envelope and the "design tax" before starting.
  • Supply-chain resilience. The recommendation to diversify vendors, GPU dependencies, technical experts, and energy suppliers gives procurement offices a concrete resilience checklist.

Industry relevance. The paper directly addresses the vendors and integrators on the other side of these decisions. It names commercial deployment channels such as AWS Bedrock and Cohere private deployment, hyperscaler marketplaces from AWS, Azure, and Google Cloud, managed model hubs and inference services such as Hugging Face, and enterprise AI platforms such as IBM. It also describes sovereign-cloud partnership models, such as the November 2025 France and Germany arrangement with Mistral AI and SAP to deploy "AI-native sovereign solutions" for public administration hosted in European data centers and overseen by a board of European nationals — a template that shapes how vendors structure public-sector offerings.

Future Directions

  • Filling out the evaluation framework beyond sovereignty. The excerpt stops at the sovereignty dimension; the remaining dimensions named in the abstract and Figure 4 — safety, cost, resource capability, cultural fit, sustainability, security, economic and regional development, and the evolving cost-capability landscape — need the same level of concrete treatment and case evidence.

  • Temporal and path-dependent strategy. Section 4.7 is referenced as discussing how the cost-capability landscape evolves over time, which raises the open question of when a phased "buy first, build later" sequence is the right trajectory versus a wasted investment.

  • Sovereign benchmarking methodology. The paper identifies the lack of high-quality evaluation data for languages such as Lao and Khmer and the "chicken-and-egg" problem with LLM-as-a-judge, leaving open how governments should construct localized benchmarks that are neither automated nor synthetically generated at scale.

  • Migration and switching pathways. Given that commercial model versions are superseded or deprecated within months to roughly a year, the paper leaves open how governments should design exit and migration routes between vendors and between buy and build postures.

  • Measuring adoption and retention. The paper argues expected adoption should shape the buy–build choice but explicitly states that designing services, interfaces, and incentives to encourage engagement is beyond its scope, leaving that as an open area.

Target Audience

Policy-makers and public-sector decision-makers evaluating how to provision language models for government use, including procurement and strategy staff in ministries, agencies, and state-backed research entities. It is also useful for researchers and analysts working on AI governance and sovereign AI, for vendors and systems integrators selling or partnering with governments, and for technically literate readers who want a grounded, non-advocacy overview of what building a sovereign LLM actually requires in data, compute, talent, evaluation, alignment, and deployment effort.

Authors’ abstract

Large Language Models (LLMs) represent a new frontier of digital infrastructure that can support a wide range of public-sector applications, from general purpose citizen services to specialized and sensitive state functions. When expanding AI access, governments face a set of strategic choices over whether to buy existing services, build domestic capabilities, or adopt hybrid approaches across different domains and use cases. These are critical decisions especially when leading model providers are often foreign corporations, and LLM outputs are increasingly treated as trusted inputs to public decision-making and public discourse. In practice, these decisions are not intended to mandate a single approach across all domains; instead, national AI strategies are typically pluralistic, with sovereign, commercial and open-source models coexisting to serve different purposes. Governments may rely on commercial models for non-sensitive or commodity tasks, while pursuing greater control for critical, high-risk or strategically important applications. This paper provides a strategic framework for making this decision by evaluating these options across dimensions including sovereignty, safety, cost, resource capability, cultural fit, and sustainability. Importantly, "building" does not imply that governments must act alone: domestic capabilities may be developed through public research institutions, universities, state-owned enterprises, joint ventures, or broader national ecosystems. By detailing the technical requirements and practical challenges of each pathway, this work aims to serve as a reference for policy-makers to determine whether a buy or build approach best aligns with their specific national needs and societal goals.

Read the original paper