Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection

Inoue, Koji, Jiang, Bing'er, Ekstedt, Erik, Kawahara, Tatsuya, Skantze, Gabriel

Jan-9-2024–arXiv.org Artificial Intelligence

A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialogue stereo audio to future voice activities. The VAP model includes contrastive predictive coding (CPC) and self-attention transformers, followed by a cross-attention transformer. We examine the effect of the input context audio length and demonstrate that the proposed system can operate in real-time with CPU settings, with minimal performance degradation.

participant, real-time and continuous turn-taking prediction, transformer, (9 more...)

arXiv.org Artificial Intelligence

Jan-9-2024

arXiv.org PDF

Add feedback

Country:
- Europe > Sweden (0.05)
- Asia > Japan
  - Shikoku > Ehime Prefecture
    - Matsuyama (0.04)
  - Honshū > Kansai
    - Kyoto Prefecture > Kyoto (0.06)

Genre:
- Research Report (0.40)

Technology:
- Information Technology
  - Architecture > Real Time Systems (0.86)
  - Artificial Intelligence
    - Natural Language > Discourse & Dialogue (0.53)
    - Machine Learning > Neural Networks
      - Deep Learning (0.48)