9B模型在智能体系统中的性能优化
Lambda团队展示了如何优化小模型在智能体系统中的表现,准确率提升26个百分点。
将90亿参数开源模型部署为前沿云模型的智能体系统后,PinchBench基准准确率从96.0%降至62.3%。通过重新调整系统,包括改变量化方式、推理循环、工具访问和内存管理,准确率恢复至88.4%,缩小了77%的性能差距。配置因素是性能限制的关键。
Drop a 9B open-weight model into an agent system built for a frontier cloud model and accuracy falls from 96.0% to 62.3% on PinchBench.
Same task. Same tools.
Retune the system around the local model. Change the quantization, reasoning loop, tool access, and memory. Accuracy recovers to 88.4%, closing 77% of the gap.
The configuration was the limiting factor.
Open source, evaluated on Lambda GPUs. 🧵
https://t.co/Spc3EDnhF3