VidChapters-7M: Video Chapters at Scale

Mar-27-2025, 14:47:39 GMT–Neural Information Processing Systems

Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly released datasets. To address this issue, we present VidChapters-7M, a dataset of 817K user-chaptered videos including 7M chapters in total. VidChapters-7M is automatically created from videos online in a scalable manner by scraping user-annotated chapters and hence without any additional manual annotation. We introduce the following three tasks based on this data. First, the video chapter generation task consists of temporally segmenting the video and generating a chapter title for each segment. To further dissect the problem, we also define two variants of this task: video chapter generation given ground-truth boundaries, which requires generating a chapter title given an annotated video segment, and video chapter grounding, which requires temporally localizing a chapter given its annotated title.

artificial intelligence, machine learning, natural language, (18 more...)

Neural Information Processing Systems

Mar-27-2025, 14:47:39 GMT

Conferences PDF

Add feedback

Country:
- Europe (0.28)

Industry:
- Education (0.46)
- Leisure & Entertainment (0.46)
- Media (0.46)

Technology:
- Information Technology
  - Artificial Intelligence
    - Machine Learning > Neural Networks
      - Deep Learning (0.46)
    - Natural Language (1.00)
    - Vision (0.93)
  - Communications (0.69)
  - Data Science (0.93)

Duplicate Docs Excel Report

Title
VidChapters-7M: Video Chapters at Scale

Similar Docs Excel Report more

Title	Similarity	Source
None found