gen‑ai.news
← Back
Video

LAION drops massive open video dataset with 10 million hours of footage for AI research

LAION drops massive open video dataset with 10 million hours of footage for AI research

LAION, the nonprofit organization behind several influential open datasets, has released its Big Video Dataset (BVD) - a collection of 80 million videos spanning 10 million hours of runtime. Alongside the raw footage, the dataset includes 55 million automatically described clips, providing text-paired video data that is particularly useful for training generative and multimodal video models. The scale puts BVD among the largest openly available video datasets in existence.

The practical impact of the dataset is already visible in early benchmarks. Models trained on BVD outperformed those trained on InternVid, previously one of the most widely used large-scale video datasets for AI research, by up to 2.1 percentage points. While that margin may sound modest, improvements at this scale of data are meaningful and suggest that BVD offers both broader coverage and better diversity than its predecessors.

The legal footing for a dataset of this nature is a genuine question, given that much of the web video it draws from is likely copyrighted. LAION appears to be leaning on a 2024 ruling from a Hamburg court, which found that collecting copyrighted material for non-commercial research purposes is permissible under German law. This provides a clearer - though not unconditional - basis for the project than many similar efforts have had. It is worth noting that legal interpretations vary across jurisdictions, and the dataset's use outside of non-commercial research contexts could raise separate questions.

LAION has a track record of releasing datasets that quietly become infrastructure for a wide range of AI research, including the LAION-5B image dataset that underpinned early open diffusion model development. BVD follows that same philosophy - assembling at scale what individual research groups rarely have the resources to build themselves, and releasing it openly. For researchers working on video generation, video-language models, or temporal understanding in AI systems, BVD represents a substantial addition to the available toolkit.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

AI-generated videos are already displacing actors and livestreamers across China's entertainment industry
Video

AI-generated videos are already displacing actors and livestreamers across China's entertainment industry

China's short-drama industry has moved rapidly toward AI-generated content, with 95 percent of the 128,000 short dramas released in Q1 2026 produced using AI. Actors report being asked to surrender their voice and likeness data before losing their jobs, and labor disputes tied to AI displacement are climbing. The trend offers an early, concrete look at how generative video is reshaping entertainment workforces at scale.

Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible
Video

Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible

Google has updated its Gemini Omni 1.1 Flash video model with broader context analysis, longer scene extension limits, and a lower-cost draft mode. The changes address practical concerns around consistency and affordability in AI-generated video. For developers and creators working at scale, the update shifts the cost-quality tradeoff in a meaningful direction.