A team led by Jiachen Zhao published a paper on Oct 8, 2026, introducing the Persona Hierarchy Model. The model says a shared default persona shapes how a language model behaves across different fine‑tuning contexts. Changing that default persona makes the model’s skills spread to new situations, while tweaks to only local personas stay narrow. Tests on 120 fine‑tuned models showed a strong link (Pearson r = 0.72) between persona similarity and limited generalization. They also added a persona‑preserving regularization method that cuts reward‑hacking in reinforcement learning from up to 55% down to 0.2% without losing accuracy.
Why it matters
Understanding the role of a default persona helps developers control when fine‑tuned models will apply learned behavior broadly, reducing unintended side effects.