Unsupervised Improvement of Audio-Text Cross-Modal Representations

Wang, Zhepei, Subakan, Cem, Subramani, Krishna, Wu, Junkai, Tavares, Tiago, Ayres, Fabio, Smaragdis, Paris

Jul-31-2023–arXiv.org Artificial Intelligence

Recent advances in using language models to obtain cross-modal audio-text representations have overcome the limitations of conventional training approaches that use predefined labels. This has allowed the community to make progress in tasks like zero-shot classification, which would otherwise not be possible. However, learning such representations requires a large amount of human-annotated audio-text pairs. In this paper, we study unsupervised approaches to improve the learning framework of such representations with unpaired text and audio. We explore domain-unspecific and domain-specific curation methods to create audio-text pairs that we use to further improve the model. We also show that when domain-specific curation is used in conjunction with a soft-labeled contrastive loss, we are able to obtain significant improvement in terms of zero-shot classification performance on downstream sound event classification or acoustic scene classification tasks.

dataset, improvement-set, representation, (17 more...)

arXiv.org Artificial Intelligence

Jul-31-2023

arXiv.org PDF

Add feedback

Country:
- North America
  - Canada > Quebec (0.04)
  - United States
    - Illinois (0.04)
    - Washington > King County
      - Seattle (0.04)
    - Minnesota > Hennepin County
      - Minneapolis (0.14)
    - Louisiana > Orleans Parish
      - New Orleans (0.04)

Genre:
- Research Report (0.64)

Industry:
- Education (0.49)

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning (1.00)
  - Natural Language > Large Language Model (0.49)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found