VLALight:首个视觉-语言-行动交通信号控制模型
VLALight: A Vision-Language-Action Model for Traffic Signal Control
香港科技大学团队推出VLALight,用多视角视频直接控制交通信号,比传统方法更高效。
VLALight是首个端到端的视觉-语言-行动模型,可直接通过多视角路边视频映射到协调信号行动。该模型在七个真实交通数据集上测试,表现优于传统交通控制、强化学习和LLM/VLM基线模型。研究团队开发了双阶段监督冷启动训练策略,并引入自适应快速和慢速推理模式。
VLALight: A Vision-Language-Action Model for Traffic Signal Control
Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at https://github.com/usail-hkust/VLALight.git.