技巧多源确认

观点:评估网络安全能力时激活参数量基本不重要

精选理由

一个关于 RLVR 评估的反直觉观点:10B 激活参数加够环境就能赢,104B 没环境也白搭,还猜了 Anthropic 的思路。

作者认为在网络安全等 RLVR 领域,激活参数量对能力评估几乎没有参考价值:只要实验室有覆盖该领域的环境,即使 10B 激活参数的模型也能通过 RL 训出高分;反之 104B 也救不了缺环境的实验室。他称之为狭窄泛化,并推测 Anthropic 的 Mythos-Preview 是在缺乏环境时用算力和推理预算换时间的方案,可节省几个月。

原文 · Teortaxes

Hot take: I think active param basically doesn't matter for evaluating cyber capabilities and related RLVR domains. If a lab has environments reasonably covering the domain, they can RL even a 10B active to score well. If no, 104B won't save them. Narrow generalization. Though I suspect Anthropic said "if your generalization is too narrow, you're not scaling the network and inference budget far enough!" with Mythos-Preview. It is an obvious solution when you just don't have environments yet and can speed ahead with compute. It saves a few months.