行业73°

微软员工:Harness占60%,模型仅占40%

一位微软 AI 方向员工的判断:Harness 占 60%,模型只占 40%,毕业生找工作越来越难! 1. 模型各有所长 Claude 擅长多步推理,OpenAI 擅长单步执行 复杂任务需要拆解、多...

精选理由

微软AI员工分享模型对比经验,Harness比模型本身更重要,初级开发者需掌握AI agent技能。

AI 摘要

微软AI员工认为,模型性能差异60%取决于Harness,40%取决于模型本身。Claude擅长多步推理,首次通过率比GitHub Copilot高20-30%。复杂任务上,OpenAI模型因返工多,实际成本高出20-30%。未来定价将从token计费转向结果计费,65-70%开发周期将自动化。

原文 · shao__meng

一位微软 AI 方向员工的判断:Harness 占 60%,模型只占 40%,毕业生找工作越来越难! 1. 模型各有所长 Claude 擅长多步推理,OpenAI 擅长单步执行 复杂任务需要拆解、多...

一位微软 AI 方向员工的判断:Harness 占 60%,模型只占 40%,毕业生找工作越来越难! 1. 模型各有所长 Claude 擅长多步推理,OpenAI 擅长单步执行 复杂任务需要拆解、多步推理时,Claude 表现更好;单步、直接执行型任务则是 OpenAI 模型的强项。 2. 返工率决定真实成本 使用 Claude 需要的返工更少。在复杂任务上,OpenAI 模型因反复来回修改才能达到质量要求,实际成本反而高出 20–30%,所以"单价便宜 ≠ 总成本低"。 3. Claude Code vs GitHub Copilot Claude Code 的首次通过率高出约 20–30%,优势来自其原生配对体验。但他认为 Copilot 长期会靠集成能力的改进缩小差距。 4. 最重要的判断:Harness 占 60%,模型只占 40% 他认为所有模型正在商品化,实际效果的差异 60% 取决于 harness,只有 40% 取决于模型本身。 5. 定价模式将从 token 计费转向结果计费 原因是企业 AI 预算有限,需要成本可预测性——按结果付费比按 token 付费更容易做预算。 6. 对开发者就业市场的影响 · 初级开发者招聘明显减少;能被录用的初级者必须具备 agentic 技能(驾驭 AI agent 的能力)。 · 招聘标准从"硬核编码能力"转向"判断力"——知道什么是好的结果、agent 应该做什么。 · 预计未来 3–5 年内 65–70% 的开发周期将完全自动化;但有经验、有判断力的人会更值钱,因为他们能评估 agent 的产出。 Rihard Jarc @RihardJarc Interview with a $MSFT employee who works on AI on comparing models and his view that the harness is 60% of the performance, while the underlying model is only 40%: 1. In his view, Claude models work very well with multi-reasoning (if you have complicated tasks that need to be broken down and reasoned over multiple steps). In contrast, OpenAI models work well if you try to execute a single-step execution. 2. He noticed that he has to ask for less rework with Claude than with OpenAI models; that is why, for complex tasks, he cites OpenAI being 20-30% more expensive, as the model goes back and forth until it passes the output quality he was looking for. 3. When comparing Claude Code versus $MSFT GitHub Copilot, he mentions that the first-pass acceptance rate is higher with Claude Code native pairing. He estimates that advantage to be in the 20-30% range, but he thinks GitHub Copilot will close the gap in the long term as it improves integration. 4. He thinks all models are becoming a commodity and believes that 60% of the performance is tied to the harness and only 40% to the model split. 5. According to him, in the long run, the industry will transition to outcome-based pricing away from token-based pricing because of cost predictability, as companies don't have unlimited budgets for AI. 6. He does notice less hiring for junior developers; for those that are getting hired, agentic skills are a necessity. Also, with developers, you aren't looking for hardcore software developer skills, but for people with judgment on what good looks like and what the agent should be doing. In 3-5 years, he expects 65-70% of the development life cycle to be fully autonomous, but humans with experience and judgment will be more valuable because they can evaluate the outputs from the agents. found on @AlphaSenseInc 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 349 ⚡