Salesforce 发布企业定制模型 Koa,在 Tau2Bench 和 CRM 基准上超越 GPT-4
Banger report from Salesforce. Pretty interesting to see more of these custom enterprise models. S...
Salesforce 新发布的 Koa 模型挺有意思,它用公司自己的工作流规范文件来训练,在 Tau2Bench 和 CRM 基准上比 GPT-4 表现更好。
Salesforce 基于 Nemotron-3-Super-120B 训练了企业代理模型 Koa,使用 Agent Script 规范文件定义任务,通过模拟用户角色进行多轮交互,奖励机制检查工具调用是否正确。Koa 在 Tau2Bench 上得分 69.41,CRM 基准得分 0.86,功能调用准确率从 0.71 提升至 0.77。
Banger report from Salesforce. Pretty interesting to see more of these custom enterprise models. S...
Banger report from Salesforce. Pretty interesting to see more of these custom enterprise models. Salesforce trained the enterprise agent model from the same files it uses to configure agents. Koa starts from the open-weight Nemotron-3-Super-120B. Salesforce takes Agent Script specifications, the declarative files that define Agentforce agents, and expands them into multi-turn tasks with simulated user personas. The reward checks whether the agent resolved the task with the right tool calls, and training uses GRPO. The gains are modest and consistent. Koa scores 69.41 on Tau2Bench against 68.64 for its base and 54.48 for GPT-4.1. On CRM Bench it reaches 0.86, close to Claude Opus 4.8 at 0.87, and function-call accuracy rises from 0.71 to 0.77. If your company already describes its workflows in a structured format, those descriptions might be useful to turn into RL environments. Paper: arxiv.org/abs/2609.15066 Chat with Paper: academy.dair.ai/papers/salesfo… 💬 1 🔄 2 ❤️ 7 👀 1459 📊 4 ⚡