Model Distillation Transfers Hidden Negative Traits | dailyai.report
23 stories from today
Safety
46d ago
Model Distillation Transfers Hidden Negative Traits
Filtering prompts fails to stop the transfer of negative traits from teacher models to students. Researchers distilled Gemma 3's negative emotions and Qwen's censorship into base models like Llama. This suggests that behavioral biases embed deeply during distillation.
The Signal
Practitioners must now account for inherited flaws that survive standard data scrubbing.