论文

ArXiv 新论文提出 Agent 校准框架:跨模型、跨司法辖区保持能力

From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

精选理由

做 agent 跨模型迁移或跨境部署的可以看看,它把校准拆成信息、harness、用户三层,还给了评估设计思路。

ArXiv 论文《From Migration to Calibration》把 agent 校准形式化为约束下的行为适应,拆为信息保留、harness 适配、用户验收三层。核心目标是在预定义能力度量上不降级,同时满足目标环境要求。论文用跨境电商案例说明共享标准如何与站点级、市场级适配器共存。作者还提出针对模型更换、跨境适配和规模化场景的 held-out 评估方案。这是一篇方法论提案,尚未包含实验验证。

原文 · arXiv cs.AI

From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retention or target-contract satisfaction. We formulate agent calibration as constrained behavioral adaptation across three interacting layers: information preservation, harness adaptation, and user acceptance; the layers apply to every scenario, not one-to-one to the three. The basic objective is non-degradation on prespecified capability measures while satisfying target requirements; aggregate improvement is stronger. Information calibration preserves independently validated source content still applicable to the target task. Harness calibration aligns observable artifacts at semantic checkpoints and repairs them through iteration, tool substitution, or local replanning within explicit budgets. User calibration enforces recipient-specific output contracts: templates, schemas, and section-level preferences. A global e-commerce example shows how shared standards coexist with site- and market-specific adapters and validation. We distinguish trainable policies from frozen-backbone configuration or controller optimization, and evidence verification from relative judgment and DPO/GRPO optimization. Recent harness-transfer and judge-validity studies motivate target-native execution records, separate audits of task validity and near-tie ranking, and matched target-native optimization controls. We propose held-out evaluations for model changes, cross-border adaptation, and scale, including a factorial test of source evidence and checkpoint repair and group-level reporting to prevent aggregate gains from masking local failures. This is a methodological proposal; implementation and empirical validation remain future work.