剪掉90%参数还不丢演技
Persona-Pruner 是一种通过隔离特定角色子网络来剪枝 LLM 的框架,在 RoleBench 上使性能下降比最强基线减少 93.8%(LLM-as-a-judge 分数),同时保持通用能力。实验表明,相比现有剪枝技术,它能更有效地保留给定角色的对话风格与知识。该方法无需全参数模型即可支持众多非玩家角色(NPC)的实时交互。
Persona-Pruner: Sculpting Lightweight Models for Role-Playing
Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a specification of a character or user persona. However, applying these capabilities to real-world applications (e.g., ecosystems with numerous NPCs interacting simultaneously) exposes a critical inefficiency due to the excessive computational cost. In this paper, we question the necessity of dedicating a full, generalist model to a single persona, hypothesizing that a specific character identity relies on only a fraction of the model's total capacity. We observe that naively pruning LMs often severely degrades the role-playing performance for a specific persona; it does not distinguish between redundant knowledge and essential character traits. We propose Persona-Pruner, a framework that sculpts a lightweight role-playing model by isolating persona-specific sub-networks from a single description. Our experiments consistently show that Persona-Pruner preserves role-playing performance substantially more effectively than existing state-of-the-art LLM pruning techniques, reducing the performance drop from the dense model by up to 93.8% over the strongest baseline on RoleBench in LLM-as-a-judge score, while still maintaining general LLM capabilities. Code is available at https://github.com/jsu-kim/Persona-Pruner.