Optical Character Recognition
10 Cool Cloud AI And ML Services You Need To Know About
Google Cloud's machine learning-powered Document AI platform -- which already has been used to process tens of billions of pages of documents for government agencies and the lending and insurance industries among others -- became generally available last week, along with Lending DocAI and Procurement DocAI. The serverless Document AI platform is a unified console for document processing that allows users to quickly access Google Cloud's form, table and invoice parsers, tools and offerings -- including Procurement DocAI and Lending DocAI -- with a unified API. It uses artificial intelligence/machine learning (AI/ML) to classify, extract and enrich data from scanned and digital documents at scale, including structured data from unstructured documents, making it easier to understand and analyze. Doc AI solutions feature Google technologies including computer vision, optical character recognition and natural language processing, which create pre-trained models for high-value and -volume documents, and Google Knowledge Graph to validate and enhance fields in documents. Research and advisory firm Gartner predicts AI will be the top category that determines IT infrastructure decisions by 2025, driving a tenfold growth in compute requirements. Half of all enterprises will have AI orchestration platforms by 2025 to operationalize AI, according to Gartner, up from less than 10 percent in 2020.
Bidding adieu to manual document processing
Traditional document processing units required staff members to manually read and key in relevant information from purchase orders, quotes, invoices, remittances and other documents โ every day, year after year. This process lowers both staff morale and productivity, and often leads to unwanted errors and increased costs. Intelligent document processing (IDP) is a next-generation approach that uses automation to quickly extract information from business documents. Here are 10 things you need to know about IDP and how it can enable end-to-end process automation for your organization. The first wave of IDP was driven by template-based optical character recognition (OCR) technology.
OCR: Give Eyes to Your Chatbot
We live in a world where robots are increasingly common. There are chatbots on the company websites and machines that build cars and other equipment by themselves. More and more are the tasks that these agents can perform, and OCR is one of them. This article tells you what OCR is, its applications, and how your company's chatbots can use it. OCR stands for Optical Character Recognition.
Google Photos finally lets PC users copy text from an image
One of the most handy Google Photos features just landed on the desktop--via your browser--where it could be even more valuable. The mobile version of Photos supports a technology called Google Lens. In 2018, Lens introduced Optical Character Recognition (OCR) technology that can automatically copy any text found in an image, allowing you to paste it elsewhere for easy saving. As 9to5Google spotted over the weekend, that Lens OCR feature is now rolling out to desktop browsers, and that rocks. Enabling OCR in Google Photos makes it easy-peasy to take a picture of a document, book, or anything else on your phone, open it in your browser, and quickly copy its contents into an Office file.
Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Jeong, Myeonghun, Kim, Hyeongju, Cheon, Sung Jun, Choi, Byoung Jin, Kim, Nam Soo
Although neural text-to-speech (TTS) models have attracted a lot of attention and succeeded in generating human-like speech, there is still room for improvements to its naturalness and architectural efficiency. In this work, we propose a novel non-autoregressive TTS model, namely Diff-TTS, which achieves highly natural and efficient speech synthesis. Given the text, Diff-TTS exploits a denoising diffusion framework to transform the noise signal into a mel-spectrogram via diffusion time steps. In order to learn the mel-spectrogram distribution conditioned on the text, we present a likelihood-based optimization method for TTS. Furthermore, to boost up the inference speed, we leverage the accelerated sampling method that allows Diff-TTS to generate raw waveforms much faster without significantly degrading perceptual quality. Through experiments, we verified that Diff-TTS generates 28 times faster than the real-time with a single NVIDIA 2080Ti GPU.
Text-to-Speech: One Small Step by Mankind to Create Lifelike Robots
Note: For those of you who prefer watching videos, please feel free to play the above video on the same content. While speech synthesis has come a long way since Kratzenstein's vowel organ that could produce the five vowel sounds, it is a whole'nother level of challenge to transform text to natural-sounding speech. Recent developments in deep learning have provided us a new approach to the challenge and in this article, we shall briefly introduce a mainstream text-to-speech method before the deep learning era, then explore models like WaveNet that Google's text-to-speech API service is now using for lifelike speech synthesis. If you pause and think for a moment about how you can perform text-to-speech, you would probably formulate a method that is very similar to the concatenative approach. In concatenative text-to-speech, texts are broken down into smaller units such as phonemes, and the corresponding recordings of the units are then combined to form a complete speech.
AI Is Booming: 2 'Strong Buy' Stocks That Stand to Benefit
The COVID pandemic may be receding, but it has left a mark on across multiple aspects of our lives. From mask mandates to travel restrictions, we chafe at some of the changes โ but in the business world the use of artificial intelligence (AI) systems has dramatically expanded in the past year. This was probably inevitable โ but AI brought advantages in coping with the pandemic for companies that could make use of it, and the expansion accelerated. AI has found its place in a huge range of applications, at both the front and back end of businesses. Itโs prevalent in software management and data systems, as well as in communications, where AI systems filter emails and conduct robochats. And this has not been ignored by Wall Street. Analysts say that plenty of compelling investments can be found within this space. With this in mind, weโve opened up TipRanksโ database, and pulled two stocks which are stand to benefit from AI technology. Importantly, both have amassed enough bullish calls from analysts to be given โStrong Buyโ consensus ratings. Nuance Communications (NUAN) Weโll start with Nuance, a company in the communications software niche. This Massachusetts-based company offers solutions for business clients in the healthcare and customer service industries, with products that enhance speech recognition, telephone call steering systems, automated phone directories, medical transcription, and optical character recognition. Itโs a full range of AI-powered, cloud communications software, applied in real time. Nuanceโs flagship product, the Dragon Ambient eXperience (DAX) is marketed to the healthcare industry, where it uses AI to automate the paperwork burdens on physician practices and hospitals. This streamlines operations allow doctors more time and resources to spend on patients, and provides greater satisfaction to health care providers and users. The applications of Nuanceโs product and solution lines to the current environment is clear: when the pandemic locked down so many people at home, businesses still had to maintain their customer-facing systems, and software automation, based on AI tech, made that possible with fewer personnel. Since the pandemic started last winter, the company seen its shares grow tremendously, up 205% in the last 12 months, far outpacing the overall stock market. The most recent quarterly report, for fiscal Q1, showed quarterly revenues above the forecast at $81.4 million. EPS showed a net loss, as expected, but at 27 cents the loss was a 28% sequential improvement from Q3. The companyโs balance sheet is strong, with zero debt, $256 million cash on hand, and a credit facility up to $50 million. The companyโs most recent quarterly report, for fiscal Q1, beat the forecasts on both the top and bottom lines. Earnings beat expectations by 11%, coming in at 20 cents per share, while revenues of $345.8 million were a modest 2% above the estimates. As a result, operating cash flow grew 22% year-over-year, to $54.6 million for the quarter. Among the bulls is 5-star analyst Daniel Ives, of Wedbush, who rates NUAN shares an Outperform (i.e. Buy), and his $65 price target implies an upside potential of ~44%. (To watch Ivesโ track record, click here) "We believe Nuance overall continues to be laser focused on building a global cloud healthcare and AI driven business with growing ARR and a sustainable revenue/ earnings stream going forward with larger deals in the field as more hospital- wide deployments shift to the cloud are playing out and gaining further momentum based on our checks," Ives opined. The analyst added, "From a valuation/ SOTP perspective, we believe over time the DAX business alone could be worth between $3 billion to $4 billion to NUAN's stock as this AI next generation platform represents a potential paradigm changer for hospitals/healthcare clinics/specialists over the coming years." Ives is no outlier on Nuance, as shown by the unanimous Strong Buy analyst consensus on the stock. Nuance has received 6 recent reviews, and all are to Buy. The shares are trading for $45.20, and the $59.67 average price target suggests a 32% one-year upside. (See NUAN stock analysis on TipRanks) Dynatrace, Inc. (DT) The second AI stock weโll look at, Dynatrace, is another cloud software company โ but Dynatraceโs products are designed to power business data. The companyโs AI platform brings intelligent automation to network management and cloud monitoring. DTโs platform allows for cloud automation, business analytics, digital experience, application security, applications and microservices, and infrastructure monitoring. Itโs sold as a one-stop-shop for network and system managers seeking an intelligent software agent. Dynatraceโs shares have been showing consistent growth over a long term. The stock is up a robust 133% in the past 12 months, and revenues have also been growing over that period. In the most recent report, for Q3 fiscal year 2021, the company showed $182.9 million in top-line revenue, beating the forecast by ~6% and growing 27% year-over-year. EPS came in at 6 cents, flat from Q2 and far better than the break-even reported for the year-ago quarter. Three key metrics stand out in the quarterly report, and both for the right reasons. Subscription revenue grew 33% year-over-year, to reach $170.3 million, and annual recurring revenue (ARR) โ which is an important predictor of future performance โ grew 35% yoy and came in at $722 million. At the same time, license revenue dropped by more than 93%, to just $300,000. Taken all together, these results point toward a strong shift toward recurring cloud customers โ a common trend in the software space. Needhamโs 5-star analyst Jack Andrews has been closely following Dynatrace, and he believes DTโs AI products may replace incumbent tools as customers expand to additional modules. โEmbedded AIOps and automation creates a compelling value propositionโฆ Compared to competitors in the market, DT's AI Engine is embedded within its core platform and can be levered across the portfolio to deliver answers from data. Moreover, its One Agent technology automatically discovers high-fidelity data from applications and thus can map the billions of dependencies in complex environments," Andrews said. The analyst summed up, "In our view, DT is well-positioned to serve as a single source of truth that can help users trace a line between written code and business outcomes (i.e. BizDevSecOps)." Andrews named Dynatrace as a top pick, and in line with this upbeat assessment, the analyst rates the stock a Buy along with a $66 price target. Ivestors stand to pocket ~28% gain should the analyst's thesis play out. (To watch Andrewsโ track record, click here) Once again, weโre looking at a stock who strong performance has inspired unanimity from the Wall Street analysts. DT shares have 13 Buy reviews, for a Strong Buy consensus rating. The stock sells for $51.76 and its $59.69 average price target suggests ~15% upside from that level. (See DT stock analysis on TipRanks) To find good ideas for AI stocks trading at attractive valuations, visit TipRanksโ Best Stocks to Buy, a newly launched tool that unites all of TipRanksโ equity insights. Disclaimer: The opinions expressed in this article are solely those of the featured analysts. The content is intended to be used for informational purposes only. It is very important to do your own analysis before making any investment.
Interpretable Distance Metric Learning for Handwritten Chinese Character Recognition
Dong, Boxiang, Varde, Aparna S., Stevanovic, Danilo, Wang, Jiayin, Zhao, Liang
Handwriting recognition is of crucial importance to both Human Computer Interaction (HCI) and paperwork digitization. In the general field of Optical Character Recognition (OCR), handwritten Chinese character recognition faces tremendous challenges due to the enormously large character sets and the amazing diversity of writing styles. Learning an appropriate distance metric to measure the difference between data inputs is the foundation of accurate handwritten character recognition. Existing distance metric learning approaches either produce unacceptable error rates, or provide little interpretability in the results. In this paper, we propose an interpretable distance metric learning approach for handwritten Chinese character recognition. The learned metric is a linear combination of intelligible base metrics, and thus provides meaningful insights to ordinary users. Our experimental results on a benchmark dataset demonstrate the superior efficiency, accuracy and interpretability of our proposed approach.
STYLER: Style Modeling with Rapidity and Robustness via SpeechDecomposition for Expressive and Controllable Neural Text to Speech
Lee, Keon, Park, Kyumin, Kim, Daeyoung
Previous works on expressive text-to-speech (TTS) have a limitation on robustness and speed when training and inferring. Such drawbacks mostly come from autoregressive decoding, which makes the succeeding step vulnerable to preceding error. To overcome this weakness, we propose STYLER, a novel expressive text-to-speech model with parallelized architecture. Expelling autoregressive decoding and introducing speech decomposition for encoding enables speech synthesis more robust even with high style transfer performance. Moreover, our novel noise modeling approach from audio using domain adversarial training and Residual Decoding enabled style transfer without transferring noise. Our experiments prove the naturalness and expressiveness of our model from comparison with other parallel TTS models. Together we investigate our model's robustness and speed by comparison with the expressive TTS model with autoregressive decoding.
Doculayer.ai - Document Intelligence
Doculayer.ai is cloud-native and supports the latest infrastructure technologies, ensuring flexible, cost efficient and enterprise-grade scalability. With this technology foundation, Doculayer.ai is able to process large volumes of documents with unparalleled accuracy, regardless of its complexity and variety.