ViLP: Knowledge Exploration using Vision, Language, and Pose Embeddings for Video Action Recognition

Open in new window