A reproducible comparison of political bias & refusal in US and Chinese language models

Models under test

China

deepseek-r1-14b

VendorDeepSeek
Ollama tagdeepseek-r1:14b
Reasoning modelyes
Inferences208

CCP-sensitive refusal55%
· in English29%
· in Mandarin81%
zh−en delta+52%
Culture-war refusal14%
Control refusal8%
Mean hedging (0–4)1.4543
Mean sycophancy0.226
Mean factual accuracy2.1747
Full profile →
China

glm4-9b

VendorZhipu
Ollama tagglm4:9b
Reasoning modelno
Inferences286

CCP-sensitive refusal5%
· in English0%
· in Mandarin10%
zh−en delta+10%
Culture-war refusal4%
Control refusal0%
Mean hedging (0–4)1.9056
Mean sycophancy0.0315
Mean factual accuracy2.4416
Full profile →
China

qwen3-30b-a3b

VendorAlibaba
Ollama tagqwen3:30b-a3b
Reasoning modelyes
Inferences286

CCP-sensitive refusal5%
· in English0%
· in Mandarin10%
zh−en delta+10%
Culture-war refusal4%
Control refusal0%
Mean hedging (0–4)1.5577
Mean sycophancy0.1119
Mean factual accuracy2.6109
Full profile →
China

qwen3-8b

VendorAlibaba
Ollama tagqwen3:8b
Reasoning modelyes
Inferences208

CCP-sensitive refusal7%
· in English5%
· in Mandarin10%
zh−en delta+5%
Culture-war refusal9%
Control refusal0%
Mean hedging (0–4)2.0986
Mean sycophancy0.024
Mean factual accuracy2.7222
Full profile →
China

yi-9b

Vendor01.AI
Ollama tagyi:9b
Reasoning modelno
Inferences286

CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal4%
Control refusal0%
Mean hedging (0–4)1.9196
Mean sycophancy0.0
Mean factual accuracy2.5
Full profile →
United States

bonsai-8b

VendorPrismML
Ollama taghf.co/prism-ml/Ternary-Bonsai-8B-gguf:F16
Reasoning modelyes
Inferences286

CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal0%
Control refusal0%
Mean hedging (0–4)1.9847
Mean sycophancy0.0109
Mean factual accuracy2.595
Full profile →
United States

claude-sonnet-4-6

VendorAnthropic
SourceAnthropic API · remote (commercial)
Reasoning modelno
Inferences286

CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal0%
Control refusal0%
Mean hedging (0–4)2.0577
Mean sycophancy0.0035
Mean factual accuracy2.8507
Full profile →
United States

gptoss-20b

VendorOpenAI
Ollama taggpt-oss:20b
Reasoning modelyes
Inferences208

CCP-sensitive refusal13%
· in English0%
· in Mandarin25%
zh−en delta+25%
Culture-war refusal19%
Control refusal0%
Mean hedging (0–4)1.6946
Mean sycophancy0.0
Mean factual accuracy2.9081
Full profile →
United States

grok-4.3

VendorxAI
SourcexAI API · remote (commercial)
Reasoning modelno
Inferences286

CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal9%
Control refusal0%
Mean hedging (0–4)1.4213
Mean sycophancy0.0052
Mean factual accuracy2.7383
Full profile →
United States

grok-4.3-reasoning

VendorxAI
SourcexAI API · remote (commercial)
Reasoning modelyes
Inferences286

CCP-sensitive refusal7%
· in English0%
· in Mandarin14%
zh−en delta+14%
Culture-war refusal18%
Control refusal0%
Mean hedging (0–4)1.3706
Mean sycophancy0.014
Mean factual accuracy2.5845
Full profile →
United States

llama31-8b

VendorMeta
Ollama tagllama3.1:8b
Reasoning modelno
Inferences208

CCP-sensitive refusal10%
· in English0%
· in Mandarin19%
zh−en delta+19%
Culture-war refusal41%
Control refusal0%
Mean hedging (0–4)1.6034
Mean sycophancy0.0
Mean factual accuracy2.309
Full profile →
United States

phi4-14b

VendorMicrosoft
Ollama tagphi4:14b
Reasoning modelno
Inferences208

CCP-sensitive refusal5%
· in English5%
· in Mandarin5%
zh−en delta+0%
Culture-war refusal9%
Control refusal0%
Mean hedging (0–4)2.0745
Mean sycophancy0.0
Mean factual accuracy2.6944
Full profile →