Deep Learning
Appendices 1 All codes, data, and instructions for our C
We plan to expand the study to a larger scale in future work. "Please extract as many components as possible from the provided images. Only provide the component names, separated by commas. We treat objects and their attributes (if found) as options for the questions. "These sentences describe the differences between the two images.
MLLM-C
The ability to compare objects, scenes, or situations is crucial for effective decision-making and problem-solving in everyday life. For instance, comparing the freshness of apples enables better choices during grocery shopping, while comparing sofa designs helps optimize the aesthetics of our living space. Despite its significance, the comparative capability is largely unexplored in artificial general intelligence (AGI).
A Benchmark Suite for Reasoning-Across-Time in Videos Jr-Jen Chen 1 Y u-Chien Liao 1
This form of reasoning, requiring advanced understanding of cause-and-effect relationships across video segments, poses significant challenges to even the frontier multimodal large language models. To facilitate this evaluation, we develop an automated pipeline for generating temporal reasoning question-answer pairs, significantly reducing the need for labor-intensive manual annotations. Our benchmark includes 921 carefully vetted validation samples and 2,143 test samples, each manually curated for accuracy and relevance. Evaluation results show that while frontier large language models outperform academic models, they still lag behind human performance by a significant 14.3% accuracy gap. Additionally, our pipeline creates a training dataset of 9,695 machine generated samples without manual effort, which empirical studies suggest can enhance the across-time reasoning via fine-tuning.