AITopics | object token

Collaborating Authors

object token

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

Bringing Image Scene Structure to Video via Frame-Clip Consistency of Object Tokens

Neural Information Processing SystemsDec-24-2025, 23:22:07 GMT

Recent action recognition models have achieved impressive results by integrating objects, their locations and interactions. However, obtaining dense structured annotations for each frame is tedious and time-consuming, making these methods expensive to train and less scalable. At the same time, if a small set of annotated images is available, either within or outside the domain of interest, how could we leverage these for a video downstream task? We propose a learning framework StructureViT (SViT for short), which demonstrates how utilizing the structure of a small number of images only available during training can improve a video model. SViT relies on two key insights.

bringing image scene structure, frame-clip consistency, name change, (6 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Vision (0.39)

Add feedback

Bringing Image Scene Structure to Video via Frame-Clip Consistency of Object Tokens

Neural Information Processing SystemsJan-18-2025, 12:28:35 GMT

bringing image scene structure, frame-clip consistency, object token, (3 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Vision (0.41)

Add feedback

VideoOrion: Tokenizing Object Dynamics in Videos

Feng, Yicheng, Li, Yijiang, Zhang, Wanpeng, Zheng, Sipeng, Lu, Zongqing

arXiv.org Artificial IntelligenceNov-25-2024

VideoOrion not only offers a more natural and efficient way to derive compact, disentangled semantic representations We present VideoOrion, a Video Large Language Model but also enables explicit object modeling of video (Video-LLM) that explicitly captures the key semantic information content with minimal computational cost. Moreover, the introduced in videos--the spatial-temporal dynamics of objects object tokens naturally allow VideoOrion to accomplish throughout the videos. VideoOrion employs expert vision video-based referring tasks. Experimental results models to extract object dynamics through a detectsegment-track show that VideoOrion can learn to make good use of the pipeline, encoding them into a set of object object tokens, and achieves competitive results on both general tokens by aggregating spatial-temporal object features. Our video question answering and video-based referring method addresses the persistent challenge in Video-LLMs benchmarks. of efficiently compressing high-dimensional video data into semantic tokens that are comprehensible to LLMs.

large language model, machine learning, natural language, (17 more...)

arXiv.org Artificial Intelligence

2411.16156

Country:

North America > United States > California > San Diego County > San Diego (0.04)
Asia > China > Beijing > Beijing (0.04)
Africa > Angola > Namibe Province > South Atlantic Ocean (0.04)

Genre: Research Report (0.82)

Technology:

Information Technology > Artificial Intelligence > Representation & Reasoning (1.00)
Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.96)

Add feedback

Diagnosing Vision-and-Language Navigation: What Really Matters

Zhu, Wanrong, Qi, Yuankai, Narayana, Pradyumna, Sone, Kazoo, Basu, Sugato, Wang, Xin Eric, Wu, Qi, Eckstein, Miguel, Wang, William Yang

arXiv.org Artificial IntelligenceMar-30-2021

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or training techniques to boost navigation performance. However, recent studies witness a slow-down in the performance improvements in both indoor and outdoor VLN tasks, and the agents' inner mechanisms for making navigation decisions remain unclear. To the best of our knowledge, the way the agents perceive the multimodal input is under-studied and clearly needs investigations. In this work, we conduct a series of diagnostic experiments to unveil agents' focus during navigation. Results show that indoor navigation agents refer to both object tokens and direction tokens in the instruction when making decisions. In contrast, outdoor navigation agents heavily rely on direction tokens and have a poor understanding of the object tokens. Furthermore, instead of merely staring at surrounding objects, indoor navigation agents can set their sights on objects further from the current viewpoint. When it comes to vision-and-language alignments, many models claim that they are able to align object tokens with certain visual targets, but we cast doubt on the reliability of such alignments.

machine learning, natural language, object-oriented architecture, (19 more...)

arXiv.org Artificial Intelligence

2103.16561

Country:

North America > United States (1.00)
Europe (1.00)

Genre: Research Report > New Finding (0.66)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Natural Language (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.46)
Information Technology > Artificial Intelligence > Representation & Reasoning > Object-Oriented Architecture (0.34)

Add feedback