
Rigorous evaluation for political censorship transfer in AI
This research page investigates whether political censorship inherent in Chinese frontier AI models transfers to smaller models distilled from them. The study uses DeepSeek V4 Flash as a teacher and an American model (GPT-OSS-120B) as the student, focusing on financial reasoning as a representative use case. The key finding is that while the student gains performance in the target domain, it does not inherit the teacher's censorship behaviors: across 152 matched prompt pairs, the distilled model showed no statistically significant difference in censorship from the untouched base model, whereas the teacher scored 45.45 points more censored on China-sensitive questions than on structurally identical controls. The page also demonstrates that self-distillation (training on the model's own corrected continuations) yields similar performance gains, and releases the evaluation tool LineageEval for reproducibility.