这篇论文在Qwen3-32B上找到了时间偏好的内部方向,用对比激活加和就能双向改变模型选短期还是长期,还能顺带提升规划能力,挺有意思。
研究者用对比线性探针在Qwen3-32B的残差流中定位了短期与长期时间视界方向,并用对比激活加和进行定向干预。在保留的二选一时间选择任务上,该方法能诱发大幅且双向的偏好改变。在改变奖励大小和延迟的分布外货币跨期选择任务中,干预显著向两个方向移动了模型在较小更早与较大更晚奖励之间的无差异阈值。适度的时间干预还提升了TravelPlanner规划能力指标。结果表明大型语言模型的跨期偏好可测量且可操控,这对涉及延迟收益与成本的AI建议系统有直接影响。
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions. We further show improvements on a planning-related capability metric under moderate temporal steering. These results suggest that model intertemporal preferences are measurable and steerable, which is relevant for AI systems that give advice involving delayed costs and benefits, and for safety questions about long-horizon planning.