Skip to content
AI.info

The Pulse

Weibo Post Describes GLM-5.3-FlashX API With Up to 200 Tokens per Second

A Weibo post by user 独孤九鎗, rather than Zhipu AI’s official account, describes the GLM-5.3-FlashX API and its pricing. No attributable quotation was available; the opened post contains no named speaker or role.

Weibo Post Describes GLM-5.3-FlashX API With Up to 200 Tokens per Second

AI.info Team ·

A Weibo post by user “独孤九鎗” says Zhipu released GLM-5.3-FlashX on September 18, 2026, with a peak inference speed of up to 200 tokens per second. The post says the API was opened for commercial use at the same time and that the model uses a unified model key for access.

The post describes GLM-5.3-FlashX as an upgrade to the earlier Flash version, with a claimed fivefold increase in inference efficiency. It presents the model as a high-speed service for real-time conversations, intelligent customer service, frequent interactions and batch data analysis.

Pricing for Flash and FlashX

According to the post, the older Flash version is priced at 8 yuan per million input tokens and 8 yuan per million output tokens. It also lists a cache-hit price of 23 yuan per million tokens.

The post lists GLM-5.3-FlashX at 2 yuan per million input tokens and 7 yuan per million output tokens, with cache-hit pricing of 57 yuan per million tokens. It describes the overall FlashX pricing as 2.5 times that of the older version while presenting the two services as options for different workloads.

The post says the earlier Flash model is suited to low-time-sensitivity and large-volume background data processing. It positions FlashX for applications where faster responses and more stable output are priorities.

Peak speed and infrastructure claims

The headline performance figure is a peak output speed of 200 tokens per second, which the post says is five times higher than the existing Flash version. That figure describes the advertised inference rate; actual API response times can also depend on prompt processing, network conditions, queueing and the time required to produce the first token.

The post attributes the performance increase to additional infrastructure investment and optimization. It says the model series runs on an inference cluster built with more than 100,000 domestically produced chips. It also refers to Zhipu’s self-developed Infra Agent scheduling system, claiming that throughput in a domestic-chip environment increased by 3.2 times over a two-week period.

The post frames the release as part of a broader shift from low-price competition toward differences in speed, efficiency and service. It says the Flash and FlashX versions are intended to serve separate use cases: the older model for lower-cost, less time-sensitive processing and FlashX for high-speed interaction.

API availability

The post says GLM-5.3-FlashX is fully available through an API for enterprise developers and individual users. Developers considering the service should check Zhipu’s current documentation for request formats, quotas, supported functions and account-specific billing terms.

The available source is a Weibo post authored by “独孤九鎗,” not a verified Zhipu AI official-account announcement. Its claims about the model’s launch date, advertised speed, pricing and infrastructure should therefore be treated as statements made in that post.

Source

Weibo user 独孤九鎗

Explore

More articles