AITopics | siasamrea

LearningfromInside: Self-drivenSiameseSampling andReasoningforVideoQuestionAnswering

Neural Information Processing SystemsFeb-11-2026, 12:37:35 GMT

By inferring the correct answers for video-based questions, video question answering (VideoQA) has attracted increasing research attention due to its huge application potential, as a fundamental technique for vision-to-language reasoning. The task involves acquisition and manipulation of spatio-temporal visual representations guided by the compositional semantics of the linguistic clues[32,15,21,34]. Existingworkscanroughly be divided into two aspects.

artificial intelligence, machine learning, siasamrea, (17 more...)

Neural Information Processing Systems

Technology:

Information Technology > Artificial Intelligence > Machine Learning (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning (0.68)

Add feedback

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering

Neural Information Processing SystemsDec-25-2025, 01:21:55 GMT

Recent advances in the video question answering (i.e., VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via supervised learning. Intuitively, multiple samples (i.e., clips) should be interdependent to capture similar visual and key semantic information in the same video. To consider the interdependent knowledge between contextual clips into the network inference, we propose a Siamese Sampling and Reasoning (SiaSamRea) approach, which consists of a siamese sampling mechanism to generate sparse and similar clips (i.e., siamese clips) from the same video, and a novel reasoning strategy for integrating the interdependent knowledge between contextual clips into the network. The reasoning strategy contains two modules: (1) siamese knowledge generation to learn the inter-relationship among clips; (2) siamese knowledge reasoning to produce the refined soft label by propagating the weights of inter-relationship to the predicted candidates of all clips. Finally, our SiaSamRea can endow the current multimodal reasoning paradigm with the ability of learning from inside via the guidance of soft labels.

learning, name change, self-driven siamese sampling and reasoning, (7 more...)

Neural Information Processing Systems

Technology:

Information Technology > Artificial Intelligence > Machine Learning (1.00)
Information Technology > Artificial Intelligence > Natural Language (0.99)
Information Technology > Artificial Intelligence > Representation & Reasoning (0.82)

Add feedback

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering Weijiang Y u

Neural Information Processing SystemsAug-18-2025, 00:14:55 GMT

Recent advances in the video question answering (i.e., VideoQA) task have

machine learning, natural language, question answering, (21 more...)

Neural Information Processing Systems

Country: Asia > China > Guangdong Province (0.04)

Genre: Research Report (0.68)

Technology:

Information Technology > Artificial Intelligence > Representation & Reasoning (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.68)
Information Technology > Artificial Intelligence > Natural Language > Question Answering (0.64)

Add feedback

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering

Neural Information Processing SystemsJan-19-2025, 09:59:10 GMT

Recent advances in the video question answering (i.e., VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via supervised learning. Intuitively, multiple samples (i.e., clips) should be interdependent to capture similar visual and key semantic information in the same video. To consider the interdependent knowledge between contextual clips into the network inference, we propose a Siamese Sampling and Reasoning (SiaSamRea) approach, which consists of a siamese sampling mechanism to generate sparse and similar clips (i.e., siamese clips) from the same video, and a novel reasoning strategy for integrating the interdependent knowledge between contextual clips into the network. The reasoning strategy contains two modules: (1) siamese knowledge generation to learn the inter-relationship among clips; (2) siamese knowledge reasoning to produce the refined soft label by propagating the weights of inter-relationship to the predicted candidates of all clips. Finally, our SiaSamRea can endow the current multimodal reasoning paradigm with the ability of learning from inside via the guidance of soft labels.

interdependent knowledge, self-driven siamese sampling and reasoning, siamese sampling and reasoning, (6 more...)

Neural Information Processing Systems

Technology:

Information Technology > Artificial Intelligence > Machine Learning (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning (0.85)
Information Technology > Artificial Intelligence > Natural Language > Question Answering (0.64)

Add feedback

Filters

Collaborating Authors

siasamrea

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

LearningfromInside: Self-drivenSiameseSampling andReasoningforVideoQuestionAnswering

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering Weijiang Y u

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering