这篇论文揭示了多模态蛋白质语言模型推理时间的关键性,并提出了改进策略,值得蛋白质语言模型研究者关注。
本文建立了一个三阶段研究框架,对三种代表性多模态蛋白质语言模型(pLMs)在四个基本任务上的推理设计空间进行实证研究。评估了vanilla采样、任务特定无分类器指导以及奖励引导的beam搜索,发现默认推理协议的低效性,并识别出针对任务的采样偏好;在任务上观察到显著的定量收益,一致地提升了多模态pLMs的性能上限,而不更新模型参数;得出与先前共识不同的关于基础模型的结论。
Unlocking Multimodal Protein Language Models at Inference Time
Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies. Yet prior work has focused more on model training than on how inference-time strategies behave. In this paper, we establish a three-stage investigation framework to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks. We evaluate vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search on multimodal pLMs, corresponding to controls over sampling distributions, per-step logits, and parallel trajectories. Throughout the complementary advancements centered on exploration-exploitation trade-off, we (1) reveal the suboptimality of default inference protocols and identify task-oriented sampling preferences; (2) observe substantial quantitative gains across tasks, consistently boosting the upper bound performance of multimodal pLMs without updating model parameters; (3) derive conclusions about base models that differ from prior consensus.