研究:3万小时第一视角视频训练不足以让智能体学会物体交互
这篇论文用3万小时第一视角视频做了个反直觉实验:数据加再多、模型再大,物体交互还是学不动,做机器人或具身智能的可以看看。
一项研究用 30,000 小时第一视角视频训练智能体,发现条件化训练在远少于该数据量的情况下就达到饱和。饱和之后,物体交互能力收敛到远低于条件化水平的程度。实验显示继续增加数据量或扩大模型规模都无法缩小这一差距,说明第一视角视频对交互能力的迁移存在上限。
What 30,000 hours of ego-centric video does not teach
- Conditioning saturates the agent with far less data. - With the agent saturated, object interaction converges far below it. - Neither more data nor a larger model closes the gap.
https://t.co/JjVqzTuT9g https://t.co/XoNh0u9FYk