PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Xiao, Yujia, Xue, Liumeng, He, Lei, Chen, Xinyi, Chiu, Aemon Yat Fei, Tian, Wenjie, Zhang, Shaofei, Kong, Qiuqiang, Zhu, Xinfa, Xue, Wei, Lee, Tan

arXiv.org Artificial Intelligence 

The dialogue content in podcasts is extracted in text format for evaluation, representing the core message the podcast aims to convey. Podcast dialogues often center around specific topics, showcasing participants' unique perspectives and insights, which makes reference-correlation-based methods infeasible. Instead, the richness of perspectives conveyed (to provide informative takeaways for the listener) and the presentation style of the dialogue (to enhance listener comprehension) should be the primary focus of evaluation. Therefore, we follow the dialogue script-based evaluation methods proposed in PodAgent (Xiao et al., 2025), which adopt a two-fold approach: (1) Quantitative Metrics such as Distinct-N, Semantic-Div, MA TTR, and Info-Dens to assess lexical diversity, semantic richness, vocabulary richness, and information density, respectively. These metrics operate independently of reference texts and focus on intrinsic text characteristics; (2) LLM-as-a-Judge, leveraging GPT -4 to replace human evaluators for complex and comprehensive assessments. Evaluation criteria include coherence, engagingness, diversity, informativeness and speaker diversity. It incorporates comparative evaluations to reduce bias and evidence-based scoring for robust and reliable results.