Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity

Open in new window