Models under test
China
deepseek-r1-14b
VendorDeepSeek
Ollama tagdeepseek-r1:14b
Reasoning modelyes
Inferences208
CCP-sensitive refusal55%
· in English29%
· in Mandarin81%
zh−en delta+52%
Culture-war refusal14%
Control refusal8%
Mean hedging (0–4)1.4543
Mean sycophancy0.226
Mean factual accuracy2.1747
Full profile →
China
glm4-9b
VendorZhipu
Ollama tagglm4:9b
Reasoning modelno
Inferences286
CCP-sensitive refusal5%
· in English0%
· in Mandarin10%
zh−en delta+10%
Culture-war refusal4%
Control refusal0%
Mean hedging (0–4)1.9056
Mean sycophancy0.0315
Mean factual accuracy2.4416
Full profile →
China
qwen3-30b-a3b
VendorAlibaba
Ollama tagqwen3:30b-a3b
Reasoning modelyes
Inferences286
CCP-sensitive refusal5%
· in English0%
· in Mandarin10%
zh−en delta+10%
Culture-war refusal4%
Control refusal0%
Mean hedging (0–4)1.5577
Mean sycophancy0.1119
Mean factual accuracy2.6109
Full profile →
China
qwen3-8b
VendorAlibaba
Ollama tagqwen3:8b
Reasoning modelyes
Inferences208
CCP-sensitive refusal7%
· in English5%
· in Mandarin10%
zh−en delta+5%
Culture-war refusal9%
Control refusal0%
Mean hedging (0–4)2.0986
Mean sycophancy0.024
Mean factual accuracy2.7222
Full profile →
China
yi-9b
Vendor01.AI
Ollama tagyi:9b
Reasoning modelno
Inferences286
CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal4%
Control refusal0%
Mean hedging (0–4)1.9196
Mean sycophancy0.0
Mean factual accuracy2.5
Full profile →
United States
bonsai-8b
VendorPrismML
Ollama taghf.co/prism-ml/Ternary-Bonsai-8B-gguf:F16
Reasoning modelyes
Inferences286
CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal0%
Control refusal0%
Mean hedging (0–4)1.9847
Mean sycophancy0.0109
Mean factual accuracy2.595
Full profile →
United States
claude-sonnet-4-6
VendorAnthropic
SourceAnthropic API · remote (commercial)
Reasoning modelno
Inferences286
CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal0%
Control refusal0%
Mean hedging (0–4)2.0577
Mean sycophancy0.0035
Mean factual accuracy2.8507
Full profile →
United States
gptoss-20b
VendorOpenAI
Ollama taggpt-oss:20b
Reasoning modelyes
Inferences208
CCP-sensitive refusal13%
· in English0%
· in Mandarin25%
zh−en delta+25%
Culture-war refusal19%
Control refusal0%
Mean hedging (0–4)1.6946
Mean sycophancy0.0
Mean factual accuracy2.9081
Full profile →
United States
grok-4.3
VendorxAI
SourcexAI API · remote (commercial)
Reasoning modelno
Inferences286
CCP-sensitive refusal0%
· in English0%
· in Mandarin0%
zh−en delta+0%
Culture-war refusal9%
Control refusal0%
Mean hedging (0–4)1.4213
Mean sycophancy0.0052
Mean factual accuracy2.7383
Full profile →
United States
grok-4.3-reasoning
VendorxAI
SourcexAI API · remote (commercial)
Reasoning modelyes
Inferences286
CCP-sensitive refusal7%
· in English0%
· in Mandarin14%
zh−en delta+14%
Culture-war refusal18%
Control refusal0%
Mean hedging (0–4)1.3706
Mean sycophancy0.014
Mean factual accuracy2.5845
Full profile →
United States
llama31-8b
VendorMeta
Ollama tagllama3.1:8b
Reasoning modelno
Inferences208
CCP-sensitive refusal10%
· in English0%
· in Mandarin19%
zh−en delta+19%
Culture-war refusal41%
Control refusal0%
Mean hedging (0–4)1.6034
Mean sycophancy0.0
Mean factual accuracy2.309
Full profile →
United States
phi4-14b
VendorMicrosoft
Ollama tagphi4:14b
Reasoning modelno
Inferences208
CCP-sensitive refusal5%
· in English5%
· in Mandarin5%
zh−en delta+0%
Culture-war refusal9%
Control refusal0%
Mean hedging (0–4)2.0745
Mean sycophancy0.0
Mean factual accuracy2.6944
Full profile →