AITopics | egoclip

Collaborating Authors

egoclip

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

31fb284a0aaaad837d2930a610cd5e50-Supplemental-Conference.pdf

Neural Information Processing SystemsFeb-8-2026, 05:24:11 GMT

In our work, we study the video-language pretraining in a specific yet significant domain - the 1st-person view,which ismotivated bytherelease oftheEgo4D dataset. Thevarying clipfrequencies aremainly dependent on manual narrations that are annotated based on the video scenarios and activities. There have average 13.4 clips per minute of video, maximize to175.8 Fig.6(b)displays the distribution of clip duration. In Figure 1 (c), we present the distribution of narration words length.

artificial intelligence, egoclip, video, (17 more...)

Neural Information Processing Systems

Country:

South America > Colombia (0.05)
North America > United States > Minnesota (0.05)
North America > United States > Indiana (0.05)
(5 more...)

Technology: Information Technology > Artificial Intelligence (0.70)

Add feedback

EgocentricVideo-LanguagePretraining

Neural Information Processing SystemsFeb-8-2026, 05:24:07 GMT

As illustrated in Tab. 1, the formerly largest egocentric video dataset EPICKITCHENS-100 [14] focuses on kitchens scenarios and its size is far smaller than those of the 3rd-person pretraining sets WebVid-2M [3] and HowTo100M [10].

artificial intelligence, egoclip, video, (15 more...)

Neural Information Processing Systems

Country:

Asia > Singapore (0.04)
South America > Chile > Santiago Metropolitan Region > Santiago Province > Santiago (0.04)

Technology: Information Technology > Artificial Intelligence > Vision (0.36)

Add feedback

31fb284a0aaaad837d2930a610cd5e50-Supplemental-Conference.pdf

Neural Information Processing SystemsAug-14-2025, 04:28:39 GMT

egoclip, narration, video, (13 more...)

Neural Information Processing Systems

Country:

North America > United States > Minnesota (0.05)
North America > United States > Indiana (0.05)
Asia > Japan > Honshū > Kantō > Tokyo Metropolis Prefecture > Tokyo (0.04)
(6 more...)

Industry: Leisure & Entertainment > Sports (0.46)

Technology: Information Technology > Artificial Intelligence > Machine Learning (0.94)

Add feedback

31fb284a0aaaad837d2930a610cd5e50-Paper-Conference.pdf

Neural Information Processing SystemsAug-14-2025, 04:28:36 GMT

dataset, egoclip, video, (15 more...)

Neural Information Processing Systems

Country:

Asia > Singapore (0.04)
South America > Chile > Santiago Metropolitan Region > Santiago Province > Santiago (0.04)

Industry: Education (0.47)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Natural Language (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.46)

Add feedback

Egocentric Video-Language Pretraining

Lin, Kevin Qinghong, Wang, Alex Jinpeng, Soldan, Mattia, Wray, Michael, Yan, Rui, Xu, Eric Zhongcong, Gao, Difei, Tu, Rongcheng, Zhao, Wenzhe, Kong, Weijie, Cai, Chengfei, Wang, Hongfa, Damen, Dima, Ghanem, Bernard, Liu, Wei, Shou, Mike Zheng

arXiv.org Artificial IntelligenceOct-12-2022

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP.

artificial intelligence, machine learning, natural language, (20 more...)

arXiv.org Artificial Intelligence

2206.0167

Country:

Asia > Singapore (0.04)
North America > United States > Minnesota (0.04)
North America > United States > Indiana (0.04)
(7 more...)

Genre: Research Report (0.64)

Industry: Leisure & Entertainment > Sports (0.46)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Natural Language (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.46)

Add feedback