KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension Jie Y ang 1,2,5 Wang Zeng

Oct-10-2025, 22:37:09 GMT–Neural Information Processing Systems

Recent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pixel-level semantic details, e.g., the keypoints of an object. To bridge this gap, we introduce the novel challenge of Semantic Keypoint Comprehension, which aims to comprehend keypoints across different task scenarios, including keypoint semantic understanding, visual prompt-based keypoint detection, and textual prompt-based keypoint detection.

arxiv preprint arxiv, keypoint, zhang, (11 more...)

Neural Information Processing Systems

Oct-10-2025, 22:37:09 GMT

Conferences PDF

Add feedback

Country:
- Asia > China
  - Hong Kong (0.04)
  - Guangdong Province > Shenzhen (0.04)

Genre:
- Research Report
  - Experimental Study (0.93)
  - New Finding (0.68)

Industry:
- Information Technology (0.67)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Large Language Model (1.00)
  - Machine Learning > Neural Networks
    - Deep Learning (0.67)
    - Perceptrons (0.46)

Duplicate Docs Excel Report

Title
fe692980c5d9732cf153ce27947653a7-Paper-Conference.pdf

Similar Docs Excel Report more

Title	Similarity	Source
None found