Skip to content
AI.info

The Pulse

LatticeFlow Finds Stronger Political Alignment in Larger Qwen Models

LatticeFlow AI’s political-bias benchmark compares Qwen 3.7 Max with Qwen3 32B, finding the larger model closer to the Chinese pole across six China-politics categories.

LatticeFlow Finds Stronger Political Alignment in Larger Qwen Models

AI.info Team ·

LatticeFlow says larger Qwen models move further toward China’s political pole

LatticeFlow AI says its first independent political-bias benchmark found that, within the Qwen family, greater scale did not bring the systems closer to a neutral midpoint. The company says Qwen 3.7 Max sits further toward the Chinese political pole than the smaller Qwen3 32B across all six China-politics categories tested.

On freedom of religion and ethnic issues, LatticeFlow identifies Qwen 3.7 Max as the most Chinese-aligned model in the evaluation. The company’s results also place GLM 5.2, Kimi K2.6, MiniMax M2.7 FP4 and DeepSeek V4 Pro toward the Chinese end of the axis across every category it measured.

The findings come as companies adopt Chinese open-weight models because they can run them on their own infrastructure and often at lower cost than leading US systems. LatticeFlow cites OpenRouter data showing that Chinese models represented between 30% and 46% of token usage by US companies in 2026, compared with 4.5% in the first half of 2025.

The disagreement is over what “bias” should mean

Political-bias testing has a measurement problem: any benchmark that defines a correct or neutral answer in advance can import the assumptions of its authors. LatticeFlow says its framework avoids that step by comparing how Chinese and Western models answer the same questions, breaking responses into individual claims and measuring where those claims converge or diverge.

The company describes the resulting political axes as data-derived rather than based on a human-written rubric or an AI judge. The approach does not treat one end of an axis as morally correct. Instead, it records where a model’s claims fall relative to the disagreement observed among the reference systems.

LatticeFlow also says the evaluation is a dated record rather than a permanent leaderboard. Each result is tied to a SHA-256 hash of the model weights tested, allowing a third party with the same files to reproduce the run. Models were self-hosted with provider-side moderation turned off, according to the company.

Reframing, not refusal, drives the Chinese-model result

LatticeFlow says the Chinese models generally did not refuse sensitive political questions. Its description of the observed behavior is reframing: the systems answered fluently and completely, but presented disputed subjects through a Chinese political lens.

An example on the company’s evaluation page compares responses to a question about the Belt and Road Initiative’s effect on human rights. The Western-axis answer emphasizes sovereignty and weak rights protections; a neutral-axis answer presents mixed effects; and the Chinese-axis answer says the initiative has not negatively affected human rights, respects national sovereignty and operates through mutual respect and win-win cooperation.

That distinction matters for enterprise testing. A refusal is visible to a user or a safety auditor. A polished answer that leaves out competing claims can look complete while still shaping the reader’s understanding of an issue.

Western models separate along their own political axes

LatticeFlow’s benchmark does not treat political alignment as a characteristic of Chinese systems alone. The company says Grok models sit at one end of its US-politics axis, while GPT-5.4 and GPT-5.5 occupy the other, with several models closer to the center.

On US human-rights questions, LatticeFlow says Grok leans toward the government and military pole, while DeepSeek leans toward human-rights organizations. The comparison is intended to show that the framework measures the behavior of individual models rather than assigning a single political profile based on country of origin.

The benchmark’s published framework covers 17 models and 11 political spectra. LatticeFlow says its first run used roughly 554 evaluation samples per axis, including direct and indirect bias probes and control datasets.

Petar Tsankov says enterprises now need evidence before deployment

Dr. Petar Tsankov, CEO and co-founder of LatticeFlow AI, frames the result as a procurement and governance issue rather than an academic classification exercise.

“Chinese AI models are already widely deployed across Western enterprises, running critical business operations — from hiring decisions to investment, credit, and lending. What we’ve found is that the ‘brain’ behind these operations — Chinese open-weight models — carries a significant political bias toward Chinese values, and that bias only gets stronger as the models get more capable and widely deployed, which is a serious concern. Everyone was worried about this; now we can measure it. The next step is adopting the right mitigation and proving that they work — giving Western organizations a real basis for adopting these models safely.”

Dr. Petar Tsankov, CEO and co-founder, LatticeFlow AI

LatticeFlow is offering the framework to enterprises and model providers that want to test systems before deployment or compare results before and after mitigation. The company says teams can receive category-level scores, claim-level examples and hash-pinned results suitable for independent verification.

The benchmark does not establish that every answer from a Chinese model is politically biased, nor does it show that model scale alone causes alignment to strengthen. It reports the position of the evaluated model weights on LatticeFlow’s axes for that specific run. Because the company presents the results as a dated measurement rather than a permanent ranking, future model updates would require a new evaluation.

For organizations choosing between open-weight systems, the immediate question is no longer only how a model performs on reasoning, coding or cost. LatticeFlow’s benchmark adds another test: whether the model repeatedly frames politically sensitive information from one national perspective, even when it does not refuse to answer.

Source

LatticeFlow AI

Explore

More articles