ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances
Do, Huy Ba, Huynh, Vy Le-Phuong, Nguyen, Luan Thanh
–arXiv.org Artificial Intelligence
Toxic speech on online platforms is a growing concern, impacting user experience and online safety. While text-based toxicity detection is well-studied, audio-based approaches remain underexplored, especially for low-resource languages like Vietnamese. This paper introduces ViToSA (Vietnamese Toxic Spans Audio), the first dataset for toxic spans detection in Vietnamese speech, comprising 11,000 audio samples (25 hours) with accurate human-annotated transcripts. We propose a pipeline that combines ASR and toxic spans detection for fine-grained identification of toxic content. Our experiments show that fine-tuning ASR models on ViToSA significantly reduces WER when transcribing toxic speech, while the text-based toxic spans detection (TSD) models outperform existing baselines. These findings establish a novel benchmark for Vietnamese audio-based toxic spans detection, paving the way for future research in speech content moderation.
arXiv.org Artificial Intelligence
Jun-3-2025
- Country:
- Asia
- Indonesia > Bali (0.04)
- Malaysia > Kuala Lumpur
- Kuala Lumpur (0.04)
- Singapore (0.04)
- Thailand > Bangkok
- Bangkok (0.04)
- Vietnam > Hồ Chí Minh City
- Hồ Chí Minh City (0.05)
- Europe > Croatia
- Dubrovnik-Neretva County > Dubrovnik (0.04)
- North America
- Mexico > Mexico City
- Mexico City (0.04)
- United States > Minnesota
- Hennepin County > Minneapolis (0.14)
- Mexico > Mexico City
- Asia
- Genre:
- Research Report > New Finding (0.68)
- Industry:
- Health & Medicine > Therapeutic Area (0.46)
- Technology:
- Information Technology
- Artificial Intelligence
- Machine Learning (1.00)
- Natural Language (1.00)
- Speech > Speech Recognition (0.31)
- Communications > Social Media (0.70)
- Artificial Intelligence
- Information Technology