Optical Character Recognition
Innovation Award Honorees - CES 2022
OrCam MyEye PRO is a wearable assistive technology device for people who are blind, visually impaired or have reading challenges. It's lightweight, finger-size and magnetically mounts on eyeglass frames. The device instantly reads aloud any printed text (books, menus, signs) and digital screens (computer, smartphone), recognizes faces, and identifies products/bar codes, money notes and colors – all in real time and offline. The interactive Smart Reading feature enables users to tailor their assistive reading experience, and Orientation assists with guidance and identification of objects. Newly released "Hey OrCam" enables control of all device features and settings hands-free, using voice commands.
Guided-TTS:Text-to-Speech with Untranscribed Speech
Kim, Heeseung, Kim, Sungwon, Yoon, Sungroh
Most neural text-to-speech (TTS) models require
An Automatic Approach for Generating Rich, Linked Geo-Metadata from Historical Map Images
Li, Zekun, Chiang, Yao-Yi, Tavakkol, Sasan, Shbita, Basel, Uhl, Johannes H., Leyk, Stefan, Knoblock, Craig A.
Historical maps contain detailed geographic information difficult to find elsewhere covering long-periods of time (e.g., 125 years for the historical topographic maps in the US). However, these maps typically exist as scanned images without searchable metadata. Existing approaches making historical maps searchable rely on tedious manual work (including crowd-sourcing) to generate the metadata (e.g., geolocations and keywords). Optical character recognition (OCR) software could alleviate the required manual work, but the recognition results are individual words instead of location phrases (e.g., "Black" and "Mountain" vs. "Black Mountain"). This paper presents an end-to-end approach to address the real-world problem of finding and indexing historical map images. This approach automatically processes historical map images to extract their text content and generates a set of metadata that is linked to large external geospatial knowledge bases. The linked metadata in the RDF (Resource Description Framework) format support complex queries for finding and indexing historical maps, such as retrieving all historical maps covering mountain peaks higher than 1,000 meters in California. We have implemented the approach in a system called mapKurator. We have evaluated mapKurator using historical maps from several sources with various map styles, scales, and coverage. Our results show significant improvement over the state-of-the-art methods. The code has been made publicly available as modules of the Kartta Labs project at https://github.com/kartta-labs/Project.
Here's how AI can transform the lives of disabled
Many believe that artificial intelligence is a futuristic concept that we only see in sci-fi movies with humanoid robots and holograms. However, it is becoming rooted in our reality, affecting various fields and groups, including persons with disabilities. Accessibility and inclusivity are genuinely revolutionized, thanks to artificial intelligence! People with disabilities can substantially enhance their daily life thanks to AI technology solutions. We've already shown how smartphones can be tools for people with vision impairments.
Book Metadata and Cover Retrieval Using OCR and Google Books API - KDnuggets
Most of the time, the raw data that we need for our data science project is not organized in a neat, well-structured, and insightful table. Rather, this is sometimes stored as text in a scanned document. Words in the document must then be extracted one by one to form a text formatted data cell. This is the task performed by Optical Character Recognition (OCR). As you read the words of this article, be it text or number, your eyes are able to process them by recognizing light and dark patterns that make up characters (e.g., letters, number, punctuation marks, etc.).
Computer Vision: Python OCR & Object Detection Quick Starter
This is the third course from my Computer Vision series. Image Recognition, Object Detection, Object Recognition and also Optical Character Recognition are among the most used applications of Computer Vision. Using these techniques, the computer will be able to recognize and classify either the whole image, or multiple objects inside a single image predicting the class of the objects with the percentage accuracy score. Using OCR, it can also recognize and convert text in the images to machine readable format like text or a document. Object Detection and Object Recognition is widely used in many simple applications and also complex ones like self driving cars.
Guided-TTS: Text-to-Speech with Untranscribed Speech - Technology Org
Neural text-to-speech (TTS) models are successfully used to generate high-quality human-like speech. However, most TTS models can be trained if only the transcribed data of the desired speaker is given. That means that long-form untranscribed data, such as podcasts, cannot be used to train existing models. A recent paper on arXiv proposes an unconditional diffusion-based generative model. It is trained on untranscribed data that leverages a phoneme classifier for text-to-speech synthesis.
Disney adds beloved characters as text-to-speech voices in TikTok – and bans them from saying 'lesbian' or 'gay'
A text-to-speech TikTok voice made by Disney that made users sound like Rocket Raccoon does not allow users to'say' words like "gay", "lesbian", or "queer". Numerous posts by users showed the feature failing to say the LGBTQ terms before it was quietly changed to allow the words. Words like "bisexual" and "transgender", were allowed by the feature. Originally, Rocket's voice would skip over the words when written normally but would be pronounced phonetically if a user wrote "qweer", for example. Attempts to make it read text that contained only the seemingly-prohibited words resulted in an error message saying that text-to-speech was not supported by the language chosen.
Computer Vision: Python OCR & Object Detection Quick Starter
This is the third course from my Computer Vision series. Image Recognition, Object Detection, Object Recognition and also Optical Character Recognition are among the most used applications of Computer Vision. Using these techniques, the computer will be able to recognize and classify either the whole image, or multiple objects inside a single image predicting the class of the objects with the percentage accuracy score. Using OCR, it can also recognize and convert text in the images to machine readable format like text or a document. Object Detection and Object Recognition is widely used in many simple applications and also complex ones like self driving cars.
Instagram introduces text-to-speech and voice effects for Reels
Instagram was clearly trying to court TikTok users when it launched its short-form video format called Reels. Now, it has introduced two features already widely popular on TikTok, perhaps in hopes that they can convert those who've been hesitating to use Reels due to their absence. One of those tools is text-to-speech, which provides a robotic voiceover for videos. When a user types in text for their videos, they'll now be able to get an auto-generated voice to read it out loud by accessing the feature living inside the Text bubble on the lower left corner of the screen. They then have to choose between the two available voice options before posting their video. While text-to-speech will make Reels more accessible, it's also popular on TikTok just because some find a robotic voice narrating their activities a funny addition to their content.