HISPASpoof: A New Dataset For Spanish Speech Forensics

Risques, Maria, Bhagtani, Kratika, Yadav, Amit Kumar Singh, Delp, Edward J.

Sep-12-2025–arXiv.org Artificial Intelligence

West Lafayette, Indiana, USA Abstract--Zero-shot V oice Cloning (VC) and T ext-to-Speech (TTS) methods have advanced rapidly, enabling the generation of highly realistic synthetic speech and raising serious concerns about their misuse. While numerous detectors have been developed for English and Chinese, Spanish--spoken by over 600 million people worldwide--remains underrepresented in speech forensics. T o address this gap, we introduce HISPASpoof, the first large-scale Spanish dataset designed for synthetic speech detection and attribution. It includes real speech from public corpora across six accents and synthetic speech generated with six zero-shot TTS systems. We evaluate five representative methods, showing that detectors trained on English fail to generalize to Spanish, while training on HISPASpoof substantially improves detection. We also evaluate synthetic speech attribution performance on HISPASpoof, i.e., identifying the generation method of synthetic speech. HISPASpoof thus provides a critical benchmark for advancing reliable and inclusive speech forensics in Spanish. The rapid advancement of speech synthesis techniques has significantly transformed the area of audio generation and speech forensics. Recent Text-to-Speech (TTS) and V oice Cloning (VC) methods [1], [2], [3], [4], [5], [6] are now capable of producing highly realistic synthetic voices that closely mimic the spectral, prosodic, and linguistic traits of real human speech [7], [8], [9], [10].

large language model, machine learning, natural language, (19 more...)

arXiv.org Artificial Intelligence

Sep-12-2025

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - California (0.68)
  - Indiana > Tippecanoe County
    - West Lafayette (0.24)
    - Lafayette (0.24)

Genre:
- Research Report > New Finding (0.68)

Industry:
- Media (0.68)
- Information Technology > Security & Privacy (0.47)
- Government (0.46)

Technology:
- Information Technology > Artificial Intelligence
  - Speech (1.00)
  - Natural Language > Large Language Model (1.00)
  - Vision (0.89)
  - Machine Learning
    - Performance Analysis > Accuracy (1.00)
    - Neural Networks > Deep Learning (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found