HN Fails to Transfer Censorship When Converting DeepSeek to GPT-OSS
deepseek privacy
| Source: HN | Original article
DeepSeek distillation into GPT-OSS reportedly bypasses censorship. Try the experiment.
A recent experiment has shown that distilling DeepSeek into GPT-OSS does not transfer censorship. This is significant as it raises questions about the inheritability of censorship in AI models. The experiment used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B, and the results indicate that distillation works well for this problem.
This matters because it highlights the complexities of AI model distillation and the potential risks of relying on teacher models. As AI models become more widespread, understanding how they inherit properties from their teachers is crucial. The fact that censorship does not transfer during distillation has implications for the development of AI models and their potential applications.
As researchers and developers continue to explore AI model distillation, this finding is likely to spark further investigation into the inheritability of other properties, such as biases and security vulnerabilities. It will be important to watch how this research unfolds and what implications it may have for the development of AI models in the future.
Sources
Back to AIPULSEN