Skip to content
AI.info

The Pulse

Mozilla Finds Open AI Models Just 4.4 Months Behind

Mozilla’s State of Open Source AI report finds leading open-weight models trail closed systems by about 4.4 months while offering lower costs for many tasks.

Mozilla Finds Open AI Models Just 4.4 Months Behind

AI.info Team ·

Mozilla’s latest research puts the gap between the best open-weight AI models and closed frontier systems at roughly four months, while warning that the models called “open” are usually not open source in the full technical sense.

The State of Open Source AI report, published September 14, says the leading open model trails the closed leader by about 4.4 months on capability measures based on task length. Mozilla also finds that open models now approach the performance of proprietary systems at substantially lower prices, shifting the economic argument for many routine workloads.

That progress conflicts with a second reality: open models remain harder to deploy, maintain and evaluate. Mozilla’s report presents the result as a split decision rather than a clean victory for either model class.

Four months of progress, five times the price

Mozilla’s comparison uses data from the METR time-horizon framework, which measures how long a task an AI system can complete with a reliable 50 percent success rate. The report says the best closed model can currently handle a task about 1.74 times longer than the best open model.

In practical terms, Mozilla describes closed systems as handling tasks in the eight-to-12-hour range that open models cannot yet complete reliably. The report estimates that open models reach the same task range about four months later, while neither category consistently handles jobs lasting more than 12 hours.

Price widens the gap in the other direction. Mozilla says Moonshot AI’s Kimi K3 scores only a few points behind Anthropic’s Fable 5 on the Artificial Analysis Intelligence Index while costing about 30 percent as much. On a neutral software harness for Terminal-Bench 2.1, the report says Z.ai’s GLM 5.2 scores within one point of Anthropic’s Claude Opus 4.7 and 4.8 at roughly one-fifth of the cost per completed task.

Those comparisons carry limits. Closed-model providers often build their own software harnesses for tool use, memory and agent behavior, while open models may perform differently when tested through third-party systems. Mozilla’s report therefore treats benchmark results as evidence of a narrowing gap, not a universal ranking.

Kimi K3 exposes where closed models still lead

The report separates model performance by workload instead of treating capability as a single number. Kimi K3 leads or reaches parity in several frontend coding, instruction-following and general-knowledge comparisons, according to Mozilla’s analysis.

Agentic terminal work is more contested. Kimi K3 scores 88.3 on Terminal-Bench 2.1 against 88.8 for another leading system, but it falls behind on FrontierSWE, a benchmark built around real software-engineering tickets. Mozilla says the closed model leads Kimi K3 by 92 Elo points on GDPval-AA v2, which measures expert-graded professional knowledge work.

Long-context retrieval is another clear advantage for proprietary systems. The report cites an 89 percent success rate for Gemini 3.1 Pro against 41 percent for DeepSeek V4-Pro on a one-million-token multi-needle retrieval test. Mozilla also points to compliance packaging, support and accountability as reasons companies continue paying for closed services even when cheaper models can handle much of the work.

Most “open” models still withhold the recipe

Mozilla’s terminology carries a significant qualification. The report examines 16 notable open releases and says none provides every element associated with the Open Source Initiative’s definition of open source software.

Most release downloadable model weights, but withhold some combination of training data, data-processing methods, training code or other information needed to reproduce the system. Mozilla therefore uses “open weights” as the more precise term for many of the models discussed.

That distinction matters for organizations that want control over their systems rather than simply a cheaper API. Downloadable weights can allow a company to run a model on its own hardware, inspect its behavior and avoid dependence on a vendor’s service, but they do not necessarily make the model reproducible or fully auditable.

“A case that hides its weak points is an advertisement.”

Raffi Krikorian, chief technology officer, Mozilla

Krikorian’s line appears in the report’s introduction, which argues that open AI needs to be judged by its weaknesses as well as its gains. Mozilla identifies deployment tooling, infrastructure standardization, documentation, evaluation and enterprise support as the main obstacles.

Open models win adoption while losing the operations fight

Mozilla’s developer survey finds that open models are used by 79 percent of professional developers, compared with 71 percent for closed models. Half of respondents use both categories, while 29 percent use open models alone and 21 percent use closed models alone.

Open models also cover slightly more use cases per developer: 5.1 compared with 4.6 for closed systems. Yet Mozilla says open models reach production 12 percentage points less often, with the difference tied primarily to deployment and operational tooling rather than raw capability.

The most frequently cited problems include infrastructure or compute costs, security and compliance concerns, maintenance, deployment complexity and a lack of specialized support. Model performance ranks lower on the list than several operational barriers, suggesting that organizations often leave open models because they are difficult to run rather than because they are plainly incapable.

China supplies most of the open frontier

Mozilla also identifies a geographic imbalance. Eight of the 10 models with the highest token volumes on OpenRouter in August 2026 use open weights, and seven of those eight were built by Chinese companies, according to the report.

The finding gives open AI a political dimension beyond cost and engineering control. Closed frontier models remain concentrated among US companies, while the strongest downloadable models increasingly come from China. Mozilla’s report says that concentration could allow a single country to influence the defaults used by developers and organizations around the world.

Raffi Krikorian argues that US and European institutions should compete in the same open channel rather than leaving the category to Chinese labs. Mozilla points to public compute programs, neutral foundations and shared evaluation infrastructure as possible foundations for a broader ecosystem.

The report’s immediate recommendation is practical: use open models for work that fits their current capability and price profile, then reserve closed systems for tasks where their lead is material. The four-month advantage matters most when a deadline arrives before open models catch up; for routine work that can wait until the next release cycle, Mozilla’s figures suggest paying five times more may buy little that lasts.

Source

Mozilla

Explore

More articles