A Survey on Temporal Sentence Grounding in Videos

Lan, Xiaohan, Yuan, Yitian, Wang, Xin, Wang, Zhi, Zhu, Wenwu

Sep-16-2021–arXiv.org Artificial Intelligence

Temporal sentence grounding in videos(TSGV), which aims to localize one target segment from an untrimmed video with respect to a given sentence query, has drawn increasing attentions in the research community over the past few years. Different from the task of temporal action localization, TSGV is more flexible since it can locate complicated activities via natural languages, without restrictions from predefined action categories. Meanwhile, TSGV is more challenging since it requires both textual and visual understanding for semantic alignment between two modalities(i.e., text and video). In this survey, we give a comprehensive overview for TSGV, which i) summarizes the taxonomy of existing methods, ii) provides a detailed description of the evaluation protocols(i.e., datasets and metrics) to be used in TSGV, and iii) in-depth discusses potential problems of current benchmarking designs and research directions for further investigations. To the best of our knowledge, this is the first systematic survey on temporal sentence grounding. More specifically, we first discuss existing TSGV approaches by grouping them into four categories, i.e., two-stage methods, end-to-end methods, reinforcement learning-based methods, and weakly supervised methods. Then we present the benchmark datasets and evaluation metrics to assess current research progress. Finally, we discuss some limitations in TSGV through pointing out potential problems improperly resolved in the current evaluation protocols, which may push forwards more cutting edge research in TSGV. Besides, we also share our insights on several promising directions, including three typical tasks with new and practical settings based on TSGV.

dataset, temporal sentence grounding, video, (11 more...)

arXiv.org Artificial Intelligence

Sep-16-2021

arXiv.org PDF

Add feedback

Country:
- North America
  - United States
    - Washington > King County
      - Seattle (0.05)
    - Ohio > Franklin County
      - Columbus (0.04)
    - New York > New York County
      - New York City (0.04)
    - Nevada > Clark County
      - Las Vegas (0.04)
    - Minnesota > Hennepin County
      - Minneapolis (0.14)
    - Michigan > Washtenaw County
      - Ann Arbor (0.04)
    - Hawaii > Honolulu County
      - Honolulu (0.04)
    - California
      - Los Angeles County > Long Beach (0.04)
      - Santa Clara County > Mountain View (0.04)
  - Canada
    - Quebec > Montreal (0.04)
    - British Columbia > Metro Vancouver Regional District
      - Vancouver (0.14)
- Europe
  - United Kingdom > England
    - Greater London > London (0.04)
    - East Sussex > Brighton (0.04)
  - Sweden > Stockholm
    - Stockholm (0.04)
  - Italy
    - Veneto > Venice (0.04)
    - Tuscany > Florence (0.04)
  - France > Île-de-France
    - Paris > Paris (0.04)
  - Belgium > Brussels-Capital Region
    - Brussels (0.04)
- Asia
  - South Korea > Seoul
    - Seoul (0.04)
  - Middle East > Republic of Türkiye
    - Karaman Province > Karaman (0.04)
  - China
    - Guangdong Province > Shenzhen (0.04)
    - Beijing > Beijing (0.04)
    - Hong Kong (0.04)
- Africa > Central African Republic
  - Ombella-M'Poko > Bimbo (0.04)

Genre:
- Overview (1.00)

Industry:
- Education (0.46)

Technology:
- Information Technology
  - Communications (1.00)
  - Artificial Intelligence
    - Vision (1.00)
    - Representation & Reasoning (1.00)
    - Natural Language > Text Processing (0.93)
    - Machine Learning
      - Neural Networks > Deep Learning (0.94)
      - Reinforcement Learning (0.66)