Optical Character Recognition
This text-to-speech converter is on sale for 50% off
TL;DR: A lifetime subscription to Micmonster AI Voiceovers(opens in a new tab) is on sale for ยฃ53.31, saving you 50% on list price. Whether you're a content creator, website manager, YouTuber, or a web marketer, it could benefit you to learn how to do voiceover for videos. You may be able to record yourself, but there's likely a finite amount of voice variation you have at your disposal, and it's time-consuming. However, an AI tool has fewer limits. With Micmonster AI Voiceovers(opens in a new tab), you can hear your text read in 500 voices, and it's just ยฃ53.31 for a lifetime subscription -- the best price you'll find on the internet.
What Is Hyperautomation?
Gartner has anointed "Hyperautomation" one of the top 10 trends for 2022. Is it a real trend, or just a collection of buzzwords? As a trend, it's not performing well on Google; it shows little long-term growth, if any, and gets nowhere near as many searches as terms like "Observability" and "Generative Adversarial Networks." And it's never bubbled up far enough into our consciousness to make it into our monthly Trends piece. However, that skeptical conclusion is too simplistic. Hyperautomation may just be another ploy in the game of buzzword bingo, but we need to look behind the game to discover what's important. There seems to be broad agreement that hyperautomation is the combination of Robotic Process Automation with AI. Natural language generation and natural language understanding are frequently mentioned, too, but they're subsumed under AI. So is optical character recognition (OCR)โsomething that's old hat now, but is one of the first successful applications of AI. Using AI to discover tasks that can be automated also comes up frequently. While we don't find the multiplication of buzzwords endearing, it's hard to argue that adding AI to anything is uninterestingโand specifically adding AI to automation. Get a free trial today and find answers on the fly, or master something new and useful.
GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech
Huang, Rongjie, Ren, Yi, Liu, Jinglin, Cui, Chenye, Zhao, Zhou
Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The highly dynamic style features in expressive voice are difficult to model and transfer; and 2) the TTS models should be robust enough to handle diverse OOD conditions that differ from the source data. This paper proposes GenerSpeech, a text-to-speech model towards high-fidelity zero-shot style transfer of OOD custom voice. GenerSpeech decomposes the speech variation into the style-agnostic and style-specific parts by introducing two components: 1) a multi-level style adaptor to efficiently model a large range of style conditions, including global speaker and emotion characteristics, and the local (utterance, phoneme, and word-level) fine-grained prosodic representations; and 2) a generalizable content adaptor with Mix-Style Layer Normalization to eliminate style information in the linguistic content representation and thus improve model generalization. Our evaluations on zero-shot style transfer demonstrate that GenerSpeech surpasses the state-of-the-art models in terms of audio quality and style similarity. The extension studies to adaptive style transfer further show that GenerSpeech performs robustly in the few-shot data setting. Audio samples are available at https://GenerSpeech.github.io/
Transfer Learning Framework for Low-Resource Text-to-Speech using a Large-Scale Unlabeled Speech Corpus
Kim, Minchan, Jeong, Myeonghun, Choi, Byoung Jin, Ahn, Sunghwan, Lee, Joun Yeop, Kim, Nam Soo
Training a text-to-speech (TTS) model requires a large scale text labeled speech corpus, which is troublesome to collect. In this paper, we propose a transfer learning framework for TTS that utilizes a large amount of unlabeled speech dataset for pre-training. By leveraging wav2vec2.0 representation, unlabeled speech can highly improve performance, especially in the lack of labeled speech. We also extend the proposed method to zero-shot multi-speaker TTS (ZS-TTS). The experimental results verify the effectiveness of the proposed method in terms of naturalness, intelligibility, and speaker generalization. We highlight that the single speaker TTS model fine-tuned on the only 10 minutes of labeled dataset outperforms the other baselines, and the ZS-TTS model fine-tuned on the only 30 minutes of single speaker dataset can generate the voice of the arbitrary speaker, by pre-training on unlabeled multi-speaker speech corpus.
More businesses need to use AI
As a startup which has been operational for five years, specializing in conversational artificial intelligence (AI), Vbee is a pioneer in providing services such as artificial voice (vbee.vn) However, the path to bringing AI to reality is still tough. First, businesses must be persuaded to apply new technological solutions to improve productivity and reduce costs. Vnee has many solutions such as KYC (Know Your Customer), artificial switchboard, artificial voice, artificial MC, OCR (optical character recognition), voice biometrics, chatbot, call bot and artificial virtual assistant, packaged and ready to be used. But businesses are hesitant to use them.
Chandojnanam: A Sanskrit Meter Identification and Utilization System
Terdalkar, Hrishikesh, Bhattacharya, Arnab
We present Chandoj\~n\=anam, a web-based Sanskrit meter (Chanda) identification and utilization system. In addition to the core functionality of identifying meters, it sports a friendly user interface to display the scansion, which is a graphical representation of the metrical pattern. The system supports identification of meters from uploaded images by using optical character recognition (OCR) engines in the backend. It is also able to process entire text files at a time. The text can be processed in two modes, either by treating it as a list of individual lines, or as a collection of verses. When a line or a verse does not correspond exactly to a known meter, Chandoj\~n\=anam is capable of finding fuzzy (i.e., approximate and close) matches based on sequence matching. This opens up the scope of a meter-based correction of erroneous digital corpora. The system is available for use at https://sanskrit.iitk.ac.in/jnanasangraha/chanda/, and the source code in the form of a Python library is made available at https://github.com/hrishikeshrt/chanda/.
EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
Lam, Perry, Zhang, Huayun, Chen, Nancy F., Sisman, Berrak
Neural models are known to be over-parameterized, and recent work has shown that sparse text-to-speech (TTS) models can outperform dense models. Although a plethora of sparse methods has been proposed for other domains, such methods have rarely been applied in TTS. In this work, we seek to answer the question: what are the characteristics of selected sparse techniques on the performance and model complexity? We compare a Tacotron2 baseline and the results of applying five techniques. We then evaluate the performance via the factors of naturalness, intelligibility and prosody, while reporting model size and training time. Complementary to prior research, we find that pruning before or during training can achieve similar performance to pruning after training and can be trained much faster, while removing entire neurons degrades performance much more than removing parameters. To our best knowledge, this is the first work that compares sparsity paradigms in text-to-speech synthesis.
PowerToys update adds OCR and two more free tools
If you use Windows, you want PowerToys. This collection of open-source goodies, guided and published by Microsoft itself, is one of the best free software packages out there, and we can't recommend it enough. That only becomes more true today, as the company publishes an updated version with three brand new tools: the previously-spotted Text Extrator (an Optical Character Recognition tool), a ruler for measuring pixels on your screen, and a tool for quickly inserting little-used accents into text. Text Extractor is probably the most universally-applicable addition here. It's an open-source version of Joseph Finney's paid Text Grab app, now integrated into PowerToys and free for Windows users.
Rabobank Australia and New Zealand Inks Deal with nCino
This partnership will benefit the bank's Australian and New Zealand employees and customers, representing a multi-currency, cross-country commitment to provide a better banking experience. "By partnering with nCino, we will optimise our financial spreading analysis," said Alexa Glynn, Chief Operating Officer at RANZ. "This relationship will provide an excellent opportunity for RANZ to support our growing customer base and modernise our systems. We're delighted that nCino's technology will enable us to offer our customers and employees a better banking experience." The world's leading specialist food and agribusiness bank, Rabobank is one of Australia and New Zealand's largest agricultural lenders and a major provider of business and corporate banking services to the country's food and agribusiness sector. By adopting the nCino Bank Operating System, RANZ gains a digital solution that intelligently transforms the process of spreading financials by leveraging machine learning and optical character recognition (OCR).