Optical Character Recognition
Malicious Android app had more than 100 million downloads in Google Play
Kaspersky researchers recently found malware in an app called CamScanner, a phone-based PDF creator that includes OCR (optical character recognition) and has more than 100 million downloads in Google Play. Various resources call the app by slightly different names such as CamScanner -- Phone PDF Creator and CamScanner-Scanner to scan PDFs. Official app stores such as Google Play are usually considered a safe haven for downloading software. Unfortunately, nothing is 100% safe, and from time to time malware distributors manage to sneak their apps into Google Play. The problem is that even such a powerful company as Google can't thoroughly check millions of apps.
Appian 'Smartens' Up The Low-Code AI-Factor
An increasing number of software coding tasks are being handed off to AI functions, especially in ... [ ] the low-code & no-code arenas. What software needs now is more AI. This is the universal mantra that every software application development and data platform company will beat out relentlessly throughout 2020. The rise of Artificial Intelligence (AI) and the Machine Learning (ML) processes that feed it and make it smarter has driven every software industry specialist to look for avenues where it can surgically enhance its products and services with additional layers of automated intelligence. Low-code software company Appian has clearly been drinking from the same AI Kool-Aid pot as everybody else; the company's latest platform release features a range of AI assists designed to make low code software development a more intelligently abetted process.
Using VAEs and Normalizing Flows for One-shot Text-To-Speech Synthesis of Expressive Speech
Aggarwal, Vatsal, Cotescu, Marius, Prateek, Nishant, Lorenzo-Trueba, Jaime, Barra-Chicote, Roberto
We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence based system with a Variational AutoEncoder (VAE) and a Householder Flow. The proposed system provides a 22% KL-divergence reduction while jointly improving perceptual metrics over state-of-the-art. At synthesis time we use one example of expressive style as a reference input to the encoder for generating any text in the desired style. Perceptual MUSHRA evaluations show that we can create a voice with a 9% relative naturalness improvement over standard Neural Text-to-Speech, while also improving the perceived emotional intensity (59 compared to the 55 of neutral speech).
Building an NLP-powered search index with Amazon Textract and Amazon Comprehend Amazon Web Services
Organizations in all industries have a large number of physical documents. It can be difficult to extract text from a scanned document when it contains formats such as tables, forms, paragraphs, and check boxes. Organizations have been addressing these problems with Optical Character Recognition (OCR) technology, but it requires templates for form extraction and custom workflows. Extracting and analyzing text from images or PDFs is a classic machine learning (ML) and natural language processing (NLP) problem. When extracting the content from a document, you want to maintain the overall context and store the information in a readable and searchable format.
8 Critical Steps to Empower Robotic Process Automation
At BoTree Technologies, we leverage the power of AI to help companies make the most of Robotic Process Automation Technology and enhance overall efficiency. With our years of experience, we have built tools and algorithms to automate several repetitive tasks. Here are the steps through which our Robotic Process Automation tasks are executed. At BoTree, the robotic automation process begins with extracting information from structured and unstructured sources like PDFs, images, e-mails, excel sheets, and several others. This step is followed by processing the data inputs through various methods including Optical Character Recognition.
Ray Kurzweil (USA) at Ci2019 - The Future of Intelligence, Artificial and Natural
Called "the restless genius" by The Wall Street Journal and "the ultimate thinking machine" by Forbes magazine, he was selected as one of the top entrepreneurs by Inc. magazine, which described him as the "rightful heir to Thomas Edison." PBS selected him as one of the "sixteen revolutionaries who made America." Ray was the principal inventor of the first CCD flat-bed scanner, the first omni-font optical character recognition, the first print-to-speech reading machine for the blind, the first text-to-speech synthesizer, the first music synthesizer capable of recreating the grand piano and other orchestral instruments, and the first commercially marketed large-vocabulary speech recognition. Among Ray's many honors, he received a Grammy Award for outstanding achievements in music technology; he is the recipient of the National Medal of Technology, was inducted into the National Inventors Hall of Fame, holds twenty-one honorary Doctorates, and honors from three U.S. presidents. Ray has written five national best-selling books, including New York Times best sellers The Singularity Is Near (2005) and How To Create A Mind (2012). He is Co-Founder and Chancellor of Singularity University and a Director of Engineering at Google heading up a team developing machine intelligence and natural language understanding.
Medline streamlines workflow by automating accounts payable
Medline Industries, a manufacturer and distributor of medical supplies, based in Northfield, Ill., is growing quickly, said Sarah Stokes, director of accounts payable at the company. That means double-digit sales growth year over year, she said, but it also means more paperwork. "Our biggest challenge is trying to keep up with the volume," Stokes said. To help tackle the 2,000 invoices the company receives each day, Medline has turned to optical character recognition (OCR) technology from vendor Abbyy, paired with platforms from several RPA vendors. The combination has gone a long way in automating accounts payable, Stokes said.
BlinkID
MRZ is an abbreviation for Machine Readable Zone, whereas MRTD refers to Machine Readable Travel Document. MRZ is a format found on most passports and identity documents worldwide. It contains the document holder's data in a form which is both visually readable and encoded with optical character recognition. Built on the latest advances in machine learning, BlinkID enables scanning of MRZ with unsurpassed accuracy and speed.
Jumio Announces Gains in Speed, Accuracy, User Experience
Jumio, the AI-powered trusted identity as a service provider, announced gains in the speed and accuracy of its verification services, as well as a more intuitive user experience. These gains come after a two year period during which Jumio invested heavily automation enabled through a variety of machine learning, artificial intelligence, and optical character recognition (OCR). As a result of its investment in supervised machine learning models, Jumio was also able to recently launch Jumio Go, its new identity verification solution powered exclusively by AI. The investment has also improved Jumio's current suite of identity verification and authentication services. "We're seeing across-the-board improvements in our ability to automate virtually every phase of the identity verification process, making our core solutions even faster, easier and more accurate for our customers and their end users," said Labhesh Patel, Jumio CTO and chief scientist. Jumio is highlighting a number of improvements that have come about as a result of its focus over the past two years.