Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face announces FineVideo, a dataset of 43k videos (3.4k hours) with rich annotations, built from YouTube-Commons to advance open-source video AI. The post details the pipeline: filtering English videos, downloading 1.8M videos, applying word density and visual dynamism filters, categorizing with a custom taxonomy, and annotating with Gemini 1.5 Pro and GPT-4o.
From the source
Open video datasets are scarce and therefore slowing down the development of open-source video AI. For this reason we built FineVideo, a dataset with 43k videos that span 3.4k hours and are annotated with rich descriptions, narrative details, scene splits, and QA pairs.
huggingface.co