LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-trainingExplained for Beginners
Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti +9 more
Abstract
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
The Problem: The Compute Bottleneck in Open Video
The paper begins by framing a familiar frustration in AI research: the tension between ambition and accessibility. For all the progress we’ve seen in large language models and image-text systems, training capable multimodal models—ones that can understand video, audio, and images together—has been bottlenecked by one thing: data processing, not data availability.
The authors articulate the core issue beautifully. We have plenty of video on the web—millions of hours on YouTube, Vimeo, Dailymotion. The real problem is that these videos live on platforms that gate access. Unlike static images, which can be swept up by a web crawler with relative ease, videos often require playing through a player, dealing with DRM, or navigating API limits. This has meant that open video-language datasets have remained tiny compared to their image-text cousins. LAION-5B, a massive image-text dataset, contains billions of pairs. The largest open video-text dataset, InternVid, contains a paltry 7 million videos.
The gap exists because building a video dataset at scale is an engineering nightmare. You need distributed download infrastructure, you need to handle platform-specific quirks, and you need massive compute just to process the files. The authors position LAION-BVD as a solution to this: a way to get “10 million hours of video” without requiring every research group to build their own download farm from scratch.
How It Works: The Mechanics of a “Dataset Factory”
Reading the methodology feels like peeking under the hood of a massive data pipeline. The authors aren’t inventing new model architectures; they are solving the logistics problem of how to turn 1.3 billion video URLs into usable training data.
Here is the pipeline, stripped of jargon:
- The Hunt: They started with 4.7 billion raw URLs from Common Crawl (the internet’s archive). They filtered this down to 1.3 billion links from major platforms: YouTube, Vimeo, and Dailymotion.
- The Download: Using a distributed system of 2,000 virtual servers (coordinated via Celery, a standard task queue) and a residential proxy network to avoid getting blocked, they attempted to download 130 million videos. About 60% succeeded, resulting in 80 million videos totaling 10 million hours of content.
- The Trim: Not every second of video is useful. They randomly sampled 2.4 million videos and applied “content-aware scene detection.” Imagine a video of a static painting—there’s no “action” happening, just a still image. The pipeline detected these static segments and discarded them. They also split videos into “scenes” (changes in content) using a tool called PySceneDetect.
- The Captioning: This is where the multimodality happens.
- Video captions: For each short clip (up to 32 frames sampled from the video), they used a visual language model (Qwen3-VL) to generate a caption like “Three women in red dresses performing on a red carpet.”
- Audio captions: They extracted the audio track and used Audio Flamingo 3 to generate captions like “A female singer performs a pop song with a slow tempo, accompanied by a choir.”
- Frame captions: They extracted individual “keyframes” (scene-changing moments) and recaptioned them using DeepSeek-VL to create image-text pairs.
The result is a dataset with three “flavors”: BVD-V (video clips with captions), BVD-A (audio clips with captions), and BVD-I (image frames with captions).
Analogy time: Think of it like a factory that takes raw footage from a security camera and automatically edits it into highlight clips, writes a tweet summarizing each clip, and extracts individual photos with descriptions. The “factory” is the pipeline; the “highlight clips” are the training data for video models; the “tweets” are the audio captions; and the “photos with descriptions” are the image-text data.
Key Results: Scaling Up Improves Performance
The authors put their money where their mouth is by training standard models (ViCLIP for video, CLAP for audio, CLIP for images) on subsets of this new data and testing them on benchmarks.
Video-Text (ViCLIP): The results are compelling. Models trained on LAION-BVD consistently beat models trained on InternVid (the previous open-source giant), even when using less filtering. For example, a model trained on 50M samples from LAION-BVD achieved a higher average score than InternVid trained on 50M samples. Crucially, the performance scales—the more data or larger the model, the better the results. The authors note that even “minimally curated” LAION-BVD samples outperform the heavily filtered InternVid models by about 2 percentage points on their aggregate metric.
Audio-Text (CLAP): Here, the dataset holds its own against dedicated audio corpora. The paper reports that LAION-BVD audio training data achieves competitive results, matching or exceeding the larger LAION-Audio dataset across different model scales. In a mixed-dataset scenario, adding LAION-BVD audio to other data (like AudioCaps and Clotho) maintained competitive performance. The takeaway is that web video is a viable source for learning to understand sound.
Image-Text (CLIP via Frames): This is perhaps the most interesting result for people working on foundation models. The authors extracted 300 million scene-changing frames from the videos and generated captions for them. Models trained on these frame-caption pairs achieved strong retrieval performance on MS-COCO (a standard image benchmark). While they lag slightly behind web-image datasets on ImageNet classification (a task that relies on specific, label-centric concepts), they excel at the “retrieval” task—finding an image based on a text description. The authors attribute this to the “visual distribution” of video frames being different from standard web images, offering a complementary perspective.
Why It Matters: Key Takeaways
This dataset matters because it lowers the barrier to entry for multimodal research. Here are the four big points:
- Scale for the Rest of Us: By releasing the dataset and a scalable pipeline, the authors enable any research group with moderate compute resources to train multimodal models. You no longer need access to proprietary, closed datasets to do state-of-the-art video or audio research.
- A New Modality Source: The paper proves that “found video” from the open web is a legitimate source of training data. It isn’t just scraped text; it’s video with synchronized audio and visual content. This expands the “compute budget” available for training.
- Strong Baselines, Easy Improvements: The results show that simply training on this data gives you a competitive model out of the box. The scaling trends suggest that if you have more compute, you can just throw more of these videos at the problem and get better results.
- Caveats to Watch: The authors are honest about the limitations. The captions are auto-generated and short, which means the models might struggle with fine-grained captioning tasks compared to datasets where humans wrote the labels. There are also potential biases inherent in “internet video” (certain topics, languages, or stereotypes may be over- or under-represented). Furthermore, the image-text training, while strong for retrieval, doesn’t match the classification performance of meticulously curated image datasets like LAION-5B.
Summary
LAION-BVD is a significant release for the open AI community. It addresses the “video problem” not by finding new videos, but by providing the plumbing to efficiently harvest and label existing ones. The paper demonstrates that this data is high-quality enough to train models that compete with, and sometimes surpass, the best open-source alternatives we had before. For product managers and engineers looking to build the next generation of video-understanding AI, this dataset provides the raw material—and the recipe—to get started without needing to be a video infrastructure company.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →