WorldScribe: Towards Context-Aware Live Visual Descriptions
Chang, Ruei-Che, Liu, Yuxuan, Guo, Anhong
–arXiv.org Artificial Intelligence
Automated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-standing challenge in accessibility. In this work, we develop WorldScribe, a system that generates automated live real-world visual descriptions that are customizable and adaptive to users' contexts: (i) WorldScribe's descriptions are tailored to users' intents and prioritized based on semantic relevance. (ii) WorldScribe is adaptive to visual contexts, e.g., providing consecutively succinct descriptions for dynamic scenes, while presenting longer and detailed ones for stable settings. (iii) WorldScribe is adaptive to sound contexts, e.g., increasing volume in noisy environments, or pausing when conversations start. Powered by a suite of vision, language, and sound recognition models, WorldScribe introduces a description generation pipeline that balances the tradeoffs between their richness and latency to support real-time use. The design of WorldScribe is informed by prior work on providing visual descriptions and a formative study with blind participants. Our user study and subsequent pipeline evaluation show that WorldScribe can provide real-time and fairly accurate visual descriptions to facilitate environment understanding that is adaptive and customized to users' contexts. Finally, we discuss the implications and further steps toward making live visual descriptions more context-aware and humanized.
arXiv.org Artificial Intelligence
Aug-13-2024
- Country:
- North America
- United States
- Michigan > Washtenaw County
- Ann Arbor (0.14)
- California > San Francisco County
- San Francisco (0.14)
- Hawaii > Honolulu County
- Honolulu (0.05)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Texas > Bexar County
- San Antonio (0.04)
- Washington > King County
- Bellevue (0.04)
- Colorado > Boulder County
- Boulder (0.04)
- New York > New York County
- New York City (0.16)
- Pennsylvania > Allegheny County
- Pittsburgh (0.05)
- Michigan > Washtenaw County
- Canada
- Quebec > Montreal (0.04)
- Newfoundland and Labrador
- Newfoundland > St. John's (0.14)
- Labrador (0.04)
- United States
- Europe
- Greece (0.04)
- United Kingdom > Scotland
- City of Glasgow > Glasgow (0.04)
- Switzerland > Zürich
- Zürich (0.14)
- Portugal > Lisbon
- Lisbon (0.04)
- Italy > Tuscany
- Florence (0.04)
- Germany
- Hamburg (0.04)
- Baden-Württemberg > Karlsruhe Region
- Heidelberg (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- Asia
- Middle East > Jordan (0.04)
- Japan > Honshū
- Kantō > Kanagawa Prefecture > Yokohama (0.04)
- North America
- Genre:
- Research Report (1.00)
- Industry:
- Health & Medicine (0.94)
- Information Technology (0.67)
- Technology:
- Information Technology
- Human Computer Interaction (1.00)
- Communications > Mobile (0.68)
- Artificial Intelligence
- Vision (1.00)
- Representation & Reasoning (0.93)
- Natural Language > Large Language Model (0.46)
- Machine Learning > Neural Networks (0.46)
- Information Technology