Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
Yan, Sheng, Du, Xin, Li, Zongying, Wang, Yi, Jin, Hongcang, Liu, Mengyuan
–arXiv.org Artificial Intelligence
Temporal grounding is crucial in multimodal learning, but it poses challenges when applied to animal behavior data due to the sparsity and uniform distribution of moments. To address these challenges, we propose a novel Positional Recovery Training framework (Port), which prompts the model with the start and end times of specific animal behaviors during training. Specifically, Port enhances the baseline model with a Recovering part to predict flipped label sequences and align distributions with a Dual-alignment method. This allows the model to focus on specific temporal regions prompted by ground-truth information. Extensive experiments on the Animal Kingdom dataset demonstrate the effectiveness of Port, achieving an IoU@0.3 of 38.52. It emerges as one of the top performers in the sub-track of MMVRAC in ICME 2024 Grand Challenges.
arXiv.org Artificial Intelligence
May-8-2024
- Country:
- Asia > China
- Chongqing Province > Chongqing (0.04)
- Guangdong Province > Shenzhen (0.04)
- Asia > China
- Genre:
- Research Report (0.64)
- Technology: