Supplementary Material for Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective

Neural Information Processing Systems 

Sec. B, we provide more examples of the similarity distribution with/without the event and visualize To investigate the flexibility of our approach, we combine LSLD with different SOT A methods for the A VVP task. The experiments show that our denoised labels are indeed influential and can be properly employed on different SOT A methods. Effectiveness of modifying class names in prompts. Table 2, we can see that the segment-level visual metric improves by 1.7 points when we add playing As we transform objects like Accordion into human behavior (i.e. Table 2: Study the impact of varying class names to make the prompt more contextual.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found