LAION-BVD发布了,这是一个包含1.3亿小时视频的开放数据集,可用于多模态预训练,对于研究者和开发者来说是个宝库。和其它视频数据集相比,它提供了更多的视频和更长的时长,对于提升模型性能大有裨益。
LAION-BVD是一个包含1.3亿个平台特定视频URL的大型开放视频数据集,下载了8000万视频,总时长达1000万小时。该数据集旨在跨视频、音频和图像模态进行多模态预训练,并在视频-文本和音频-文本基准测试中取得竞争优势。此外,还探索了视频帧作为图像-文本数据的替代来源,并发布了LAION-BVD以供研究社区使用。
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.