Optical Character Recognition
Mind-reading machine can translate your thoughts and display as text
Scientists have developed an astonishing mind-reading machine which can translate what you are thinking and instantly display it as text. They claim that it has an accuracy rate of 90 per cent or more and say that it works by interpreting consonants and vowels in our brains. The researchers believe that the machine could one day help patients who suffer from conditions that don't allow them to speak or move. The machine registers and analyses the combination of vowels and consonants that we use when constructing a sentence in our brains. It interprets these sentences based on neural signals and can translate them into text in real time.
Google's new text-to-speech service has more realistic voices
Google will now let developers use the text-to-speech synthesis that powers the voices in Google Assistant and Maps. Cloud Text-to-Speech is available now through the Google Cloud Platform and the company says it can be used to power voice response systems in call centers, enable IoT device speech and convert media like news articles and books into a spoken format. There are 32 different voice options in 12 languages and users can customize pitch, speaking rate and volume gain. Additionally, a selection of the available voices were built with Google's WaveNet model. It was developed by Google's DeepMind team and the company first announced it in 2016. Rather than using fragments of speech and stringing them together to make words -- which often sounds very robotic -- WaveNet forms individual sound waves, creating more natural sounding speech.
The best portable document scanner
This post was done in partnership with Wirecutter. When readers choose to buy Wirecutter's independently chosen editorial picks, it may earn affiliate commissions that support its work. After putting in more than 100 hours for research and hands-on testing since 2013, we think the Epson ES-300W is the best portable document scanner for digitizing documents without taking up half of a desktop. It combines scan speeds usually found on full-size scanners with extremely accurate text recognition. And thanks to its built-in Wi-Fi and battery, you can use it almost anywhere--even with a phone or tablet.
Veritone Announces General Availability of Artificial Intelligence Developer Application - Veritone, Inc.
Veritone, Inc. (NASDAQ: VERI), a leading provider of artificial intelligence (AI) insights and cognitive solutions, today announced the general availability of its Veritone Developer application. The application empowers developers of cognitive engines, applications and application programming interfaces (APIs) to bring new AI ideas to life through simple integration with the Veritone aiWARE platform. Veritone Developer is a self-service development environment that empowers developers to create, submit and deploy public and private applications and cognitive engines directly into the aiWARE architecture. After a successful limited beta release to a select group of partners, Veritone Developer is now publicly available as a unique resource for machine learning experts, application development firms, and system integrators. Veritone Developer supports RESTful and GraphQL API integrations as well as engine development in major categories of cognition, including: transcription, translation, face and object recognition, audio/video fingerprinting, optical character recognition (OCR), geolocation, transcoding, and logo recognition, among others.
Residents Blast Mail Delivery Service in Michigan Town
Residents complain that mail carriers have failed to deliver prescription drugs and pension checks. Some also say their mailboxes have been damaged or destroyed by carriers' trucks, that telephone complaint lines were never answered and that postal managers were rude when residents visited the post office in person.
Application of Image Processing in Intelligent Character Recognition
Summary: Image processing is a rapidly evolving field with immense significance in science and engineering. One of the latest applications of Image processing is in Intelligent Character Recognition (ICR). Intelligent Character Recognition is the computer translation of handwritten text into machine-readable and machine-editable characters. It is an advanced version of Optical Character Recognition system that allows fonts and different styles of handwriting to be recognized during processing with high accuracy and speed. ICR, in combination with OCR and OMR (Optical Mark Recognition), is used in forms processing. Forms processing is a process by which one can capture information entered into different data fields filled in forms and convert it to an editable text.
Fooling OCR Systems with Adversarial Text Images
Song, Congzheng, Shmatikov, Vitaly
We demonstrate that state-of-the-art optical character recognition (OCR) based on deep learning is vulnerable to adversarial images. Minor modifications to images of printed text, which do not change the meaning of the text to a human reader, cause the OCR system to "recognize" a different text where certain words chosen by the adversary are replaced by their semantic opposites. This completely changes the meaning of the output produced by the OCR system and by the NLP applications that use OCR for preprocessing their inputs.
rOpenSci Support for hOCR and Tesseract 4 in R
Earlier this month we released a new version of the tesseract package to CRAN. This package provides R bindings to Google's open source optical character recognition (OCR) engine Tesseract. Two major new features are support for HOCR and support for the upcoming Tesseract 4. Support for HOCR output was requested by one of our users on Github. Every word in the hOCR output includes meta data such as bounding box, confidence metrics, etc. So this gives us a little more information about the OCR results than just the text.
Enhancing RNN Based OCR by Transductive Transfer Learning From Text to Images
He, Yang (Wuhan University of Technology) | Yuan, Jingling (Wuhan University of Technology) | Li, Lin (Wuhan University of Technology)
This paper presents a novel approach for optical character recognition (OCR) on acceleration and to avoid underfitting by text. Previously proposed OCR models typically take much time in the training phase and require large amount of labelled data to avoid underfitting. In contrast, our method does not require such condition. This is a challenging task related to transferring the character sequential relationship from text to OCR. We build a model based on transductive transfer learning to achieve domain adaptation from text to image. We thoroughly evaluate our approach on different datasets, including a general one and a relatively small one. We also compare the performance of our model with the general OCR model on different circumstances. We show that (1) our approach accelerates the training phase 20-30% on time cost; and (2) our approach can avoid underfitting while model is trained on a small dataset.
Who's Afraid of Autonomous Mail Trucks? RealClearPolicy
Highly automated vehicles (HAVs) have gone from fantasy to reality in the past decade. It is remarkable to witness the various prototypes motoring about courses and taking test drives on America's roads. Already, related technology helps people parallel park their cars -- some don't even require the driver be in the car. Among other benefits, HAV technology has the potential to save lives and reduce insurance costs by greatly decreasing human errors, which cause 94 percent of accidents. Computers, note two keen observers, don't get drunk or drowsy.