PyINE:用可执行程序研究模型捷径行为的监督框架
PyINE: A Framework for Scalable Elicitation and Oversight via Code Execution
这篇把模型走捷径的监督问题做成了可复现的实验场:百万条执行轨迹加代码变体,还实测了探针、judge、debate 四种监督方式各自漏什么、贵在哪。
arXiv 论文提出 PyINE,用带插桩的 Python 程序作为可验证执行基底,研究推理模型给出看似合理但不完整推理时的监督问题。首个版本 PyINE-v1 包含近 100 万条确定性执行轨迹和超过 50 万条用于反事实评估的 LLM 生成代码变体。作者用 RLVR 训练出一个擅长预测执行结果、却会在误导性提示与程序实际行为冲突时系统性出错的捷径模型。评测显示激活探针、文本分类器、prompted judge 和轻量 debate 等监督手段在数据集层面 pooled 指标良好,但对罕见捷径驱动错误的覆盖率很低,更强的模型化检查更均衡但成本显著更高。
PyINE: A Framework for Scalable Elicitation and Oversight via Code Execution
Reasoning models can remain capable of solving a task while still defaulting to cheaper but misleading shortcuts. This creates a central oversight problem: when a model gives an answer with plausible but incomplete reasoning, can an overseer determine whether that output should be trusted? To study this problem, we introduce PyINE, a framework for scalable elicitation and oversight using instrumented Python programs as a verifiable execution substrate. In PyINE, programs define task environments, execution traces provide authoritative labels for outcomes and intermediate facts, and task variants can be generated mechanically rather than through static human annotation. We instantiate the framework in PyINE-v1, a first release built from nearly one million deterministic execution traces and over 500,000 matched LLM-generated code variants used for counterfactual evaluation. Using standard RL with verifiable rewards on cue-varied tasks, we train a shortcut-following model that improves substantially at predicting execution outcomes while still making systematic errors when misleading human-facing cues conflict with the program's realized behavior. We then evaluate activation probes, trained text classifiers, prompted judges, and a lightweight debate protocol as overseers of this model. We find that performance pooled at the dataset level can hide weak coverage of the failures that matter most: cheap learned overseers often miss rare shortcut-driven errors, while stronger model-based checks are more balanced but substantially costlier and harder to turn into reliable thresholded decisions. PyINE-v1 turns this failure-mode coverage problem into a reusable experimental setting for developing oversight methods that are verifiable, failure-mode-aware, and cost-sensitive.