The Pulse
Open Models Take 78.4% of Vercel Gateway Tokens
Open-weight models account for 78.4% of token volume in a September 19 snapshot from Vercel AI Gateway, according to Vercel CEO Guillermo Rauch. The same data shows open models winning on usage while closed models continue to capture most s

AI.info Team ·
Looks like today may be a record day for token volume % of open models on Vercel AI Gateway: 🟦 Open 78.4% 🟨 Closed 21.6% While spend 💲 usually tells a different story, #3 and #4 today are Moonshot AI & DeepSeek. Adding Z.ai, their combined spend surpasses OpenAI (#2). (Do note that's the spend for inference of the model across providers (mostly in the US), not revenue going directly to the open weight labs.)
Guillermo Rauch, founder and chief executive, Vercel
One day, one gateway, a sharp split
Open-weight models account for 78.4% of token volume routed through Vercel AI Gateway in a September 19 snapshot, leaving closed models with 21.6%, Vercel founder and CEO Guillermo Rauch says. Rauch described the figure as a possible record for open models on the service.
The number measures traffic through Vercel’s gateway, not the entire AI market. Vercel’s service aggregates production requests from its own customers, and the company says it handles tens of trillions of tokens each month. The snapshot therefore offers a large view of one routing platform, but it does not establish that open-weight models hold the same share across direct provider traffic, competing gateways or private deployments.
Volume leads while spending stays elsewhere
Rauch’s post also points to a different ranking by estimated spend. Moonshot AI and DeepSeek ranked third and fourth for the day, he said. Adding Z.ai’s spending to those two providers pushed their combined total above OpenAI’s share on the gateway.
Rauch drew a boundary around that comparison: the figures represent inference purchased through providers, mostly in the United States, rather than revenue paid directly to Moonshot AI, DeepSeek or Z.ai. That distinction matters because token volume and spending measure different things. Lower-priced models can process far more input, output, reasoning and cached tokens while still generating less provider revenue than premium proprietary systems.
Vercel’s monthly data shows the same direction
The daily result arrives three days after Vercel’s September AI Gateway Production Index, which covers traffic through August. That report found open-weight models handled 56% of gateway tokens during the month while accounting for 14% of estimated spending. Vercel said open-weight share had risen from 7% in December 2025.
August also produced a 23.2% drop in average price per token across the gateway, the third consecutive monthly decline. Among teams that processed more than 10 million tokens in both July and August, the median team paid 7.6% less per token in August, according to Vercel’s report.
Vercel attributes the lower average price partly to the growing use of open-weight models. Its monthly figures show a widening gap between workload volume and spend: open models carry the bulk of tokens, while more expensive frontier systems retain most of the dollars. Anthropic captured 64% of gateway spending in August, even as open-weight models processed the majority of traffic.
AI Gateway turns model choice into a routing decision
Vercel positions AI Gateway as a single interface for hundreds of models and providers. The service offers usage and spend tracking, automatic failover and routing based on cost, latency or availability. Vercel says it charges provider list prices without adding a token markup.
The company’s own product strategy depends on developers treating models as interchangeable services. Applications can switch between OpenAI, Anthropic, Google and open-weight alternatives without rebuilding each provider integration. Vercel’s gateway documentation also describes unified observability, provider fallbacks and support for coding agents including Claude Code, OpenAI Codex and OpenCode.
That approach gives Vercel a commercial reason to publish traffic data. The company becomes both a routing provider and a source of information about which models developers actually use. Its monthly production index and public gateway rankings make model migration visible while reinforcing the value of an intermediary that can move workloads between providers.
What the 78.4% figure does—and does not—show
The September 19 snapshot shows open-weight models carrying most of the measured workload on one large gateway for one day. It does not show that open models generate most of the money in AI inference, nor that their share will remain at that level. A small number of long-context or automated-agent workloads can materially change token totals, and pricing differences can widen the gap between volume and spend.
The more durable signal sits in the trend from Vercel’s monthly data: open-weight models moved from a small minority of gateway tokens in December to a majority by August, while customers continued to reserve higher-priced proprietary models for work they considered worth the premium. The latest daily reading makes that separation more pronounced, with open models taking 78.4% of tokens and closed models taking 21.6%.