观点:DeepSeek V4.1 训练复杂度更高但信息流更简洁
一条关于 DeepSeek V4.1 的训练观察:没做 dense warmup 也能训练顺利,作者解释了为什么这种设计值得 scale up。
评论者 teortaxesTex 认为 DeepSeek V4.1 的训练过程没有依赖 dense warmup 调整,训练全程平稳。他指出 V4.1 在算法最小描述长度上更复杂,但换来的是更简洁的信息流。基于这一判断,他认为该方案后续会被放大规模。
Imo what is meaningful is the complexity of training V4.1 didn't need any dense warmup tinkering; it trained smoothly. It's "more complex" in terms of the minimal description length of the algorithm, but it enables a "simpler" information flow. Thus it will get scaled up. https://t.co/mBLQN0JQGi