Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning

Open in new window