Toward American Sign Language Processing in the Real World: Data, Tasks, and Methods

Aug-23-2023–arXiv.org Artificial Intelligence

Sign language, which conveys meaning through gestures, is the chief means of communication among deaf people. Recognizing sign language in natural settings presents significant challenges due to factors such as lighting, background clutter, and variations in signer characteristics. In this thesis, I study automatic sign language processing in the wild, using signing videos collected from the Internet. This thesis contributes new datasets, tasks, and methods. Most chapters of this thesis address tasks related to fingerspelling, an important component of sign language and yet has not been studied widely by prior work. I present three new large-scale ASL datasets in the wild: ChicagoFSWild, ChicagoFSWild+, and OpenASL. Using ChicagoFSWild and ChicagoFSWild+, I address fingerspelling recognition, which consists of transcribing fingerspelling sequences into text. I propose an end-to-end approach based on iterative attention that allows recognition from a raw video without explicit hand detection. I further show that using a Conformer-based network jointly modeling handshape and mouthing can bring performance close to that of humans. Next, I propose two tasks for building real-world fingerspelling-based applications: fingerspelling detection and search. For fingerspelling detection, I introduce a suite of evaluation metrics and a new detection model via multi-task training. To address the problem of searching for fingerspelled keywords in raw sign language videos, we propose a novel method that jointly localizes and matches fingerspelling segments to text. Finally, I will describe a benchmark for large-vocabulary open-domain sign language translation based on OpenASL. To address the challenges of sign language translation in realistic settings, we propose a set of techniques including sign search as a pretext task for pre-training and fusion of mouthing and handshape features.

artificial intelligence, machine learning, natural language, (21 more...)

arXiv.org Artificial Intelligence

Aug-23-2023

arXiv.org PDF

Add feedback

Country:
- South America > Chile
  - Santiago Metropolitan Region > Santiago Province > Santiago (0.04)
- Oceania > Australia
  - New South Wales (0.04)
- North America > United States
  - Massachusetts (0.04)
  - New Mexico (0.04)
  - Utah > Salt Lake County
    - Salt Lake City (0.04)
  - Pennsylvania > Allegheny County
    - Pittsburgh (0.04)
  - Illinois > Cook County
    - Chicago (0.04)
  - California > San Francisco County
    - San Francisco (0.04)
- Asia > Japan
  - Honshū > Chūbu > Toyama Prefecture > Toyama (0.04)

Genre:
- Research Report
  - Promising Solution (1.00)
  - New Finding (1.00)

Industry:
- Health & Medicine > Therapeutic Area (1.00)
- Education > Curriculum
  - Subject-Specific Education (1.00)

Technology:
- Information Technology > Artificial Intelligence
  - Speech > Speech Recognition (1.00)
  - Representation & Reasoning > Personal Assistant Systems (1.00)
  - Natural Language
    - Machine Translation (1.00)
    - Text Processing (0.92)
  - Machine Learning
    - Statistical Learning (1.00)
    - Neural Networks > Deep Learning (1.00)
    - Performance Analysis > Accuracy (0.67)
    - Learning Graphical Models > Undirected Networks
      - Markov Models (0.67)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found