DeepSeek 压缩智能体记忆:缓存每 token 仅 890 字节
DeepSeek 把智能体的 KV 缓存砍到每 token 890 字节,比 V1 小 437 倍,Flash 还在 AA Index 上反超 V4-Pro,价格只要一半不到。
DeepLearning.AI 在 The Batch 中解析了 DeepSeek 针对智能体场景的 KV 缓存压缩方案。智能体读取远多于写入,DeepSeek 将每 token 缓存降至 890 字节,比 DeepSeek-V1 小 437 倍。输入从 4K 扩展到 1M token 时,每输出 token 的计算量仅增加 25%。Flash 模型在 AA Index 上以 39 分超过 V4-Pro 的 36 分,价格 $0.27 对 $0.67 每百万 token。
Agents read more than they write, so DeepSeek shrank what they remember. Read the technical breakdown in this week's issue of The Batch. 📰🗜️ 📉 890 bytes of cache per token, 437x smaller than DeepSeek-V1 ⚡ +25% compute per output token from 4K to 1M input 🏆 Flash beats V4-Pro: 39 vs 36 on AA Index, $0.27 vs $0.67/ta hubs.la/Q04ztt000 TC #DeepLearningAI n #LLMs L #AIAgents ents 💬 0 🔄 0 ❤️ 0 👀 249 ⚡
- kimmonismus10-06 16:41原文