Compression of end-to-end non-autoregressive image-to-speech system for low-resourced devices
Srinivasagan, Gokul, Deisher, Michael, Georges, Munir
–arXiv.org Artificial Intelligence
People with visual impairments have difficulty accessing touchscreen-enabled personal computing devices like mobile phones and laptops. The image-to-speech (ITS) systems can assist them in mitigating this problem, but their huge model size makes it extremely hard to be deployed on low-resourced embedded devices. In this paper, we aim to overcome this challenge by developing an efficient endto-end neural architecture for generating audio from tiny segments of display content on low-resource devices. We introduced a vision transformers-based image encoder and utilized knowledge distillation to compress the model from 6.1 million to 2.46 million parameters. Human and automatic evaluation results show that our approach leads to a very minimal drop in performance and can speed up the inference time by 22%.
arXiv.org Artificial Intelligence
Nov-30-2023
- Country:
- North America > United States
- Oregon > Washington County > Hillsboro (0.04)
- Europe
- Switzerland > Vaud
- Lausanne (0.04)
- Germany
- Saxony > Leipzig (0.04)
- Saarland > Saarbrücken (0.04)
- Bavaria > Upper Bavaria
- Munich (0.04)
- Ingolstadt (0.04)
- Switzerland > Vaud
- North America > United States
- Genre:
- Research Report > New Finding (0.35)
- Industry:
- Education (0.51)
- Health & Medicine (0.34)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Natural Language (1.00)
- Machine Learning > Neural Networks (1.00)
- Speech (0.96)
- Information Technology > Artificial Intelligence