STARS: 人机交互中的时空动态与社会表征研究
STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
斯坦福团队推出社交机器人导航基准测试,发现顶级VLM表现还不如简单规则方法,揭示了AI社交理解的不足。
研究人员发布了SocialNav-SUB基准测试,用于评估视觉语言模型在社交机器人导航场景中的理解能力。该基准测试包含视觉问答数据集,要求模型处理空间、时空和社会推理任务。实验显示,即使最先进的VLM在回答问题上仍低于简单规则方法和人类共识基线,表明当前模型在社会场景理解上存在关键差距。
STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/socialnav-sub.