Claude 从写个位数代码到主导 80% 生产代码,这标志着 AI 编程从辅助工具向主力角色的质变。做工程管理的团队和重度使用 AI 编程的开发者,值得关注这个趋势——它直接关系到团队产出和开发流程的重新定义。
Anthropic 最新披露,Claude 现在合并的生产代码中,超过 80% 由它自己编写。在 Claude Code 于 2025 年 2 月进入研究预览之前,Claude 仅贡献了个位数的合并代码,而每位工程师的产出已升至 2024 年基线的 8 倍。这一转变源于智能体能够编辑文件、运行测试、检查失败、生成辅助智能体,并在更长任务中持续工作,而不仅仅是提供代码片段。Anthropic 表示可靠任务长度每约 4 个月翻倍,Mythos Preview 可稳定运行至少 16 小时,Claude Code 开放任务成功率已达 76%。人类剩余优势在于研究判断:选择正确问题、信任正确结果、判断实验何时失败。
Anthropic just disclosed that Claude now writes mo…
Anthropic just disclosed that Claude now writes more than 80% of the production code it merges.
Before Claude Code reached research preview in 02-25, Claude wrote only low-single-digit merged code, while output per engineer has since risen to 8x the 2024 baseline.
The shift comes from agents that edit files, run tests, inspect failures, spawn helper agents, and keep working across longer tasks instead of only suggesting snippets.
Anthropic says reliable task length is doubling about every 4 months, with Mythos Preview reaching at least 16 hours and open-ended Claude Code success hitting 76%.
i.e. Claude Mythos Preview could stay useful on a task that would take a skilled human roughly 16 hours of work
Claude also moved from a 3x training-code speedup to 52x, while a skilled human reached about 4x in 4 to 8 hours on the same setup.
The remaining human edge is research judgment: choosing the right problem, trusting the right result, and knowing when an experiment is dead.