Goto

Collaborating Authors

 Optical Character Recognition


what-is-the-use-of-machine-learning-handwriting-recognition

#artificialintelligence

Recent Deep Learning advancements, such as the introduction of transformer topologies, have helped us accelerate our handwritten character recognition. Intelligent Character Recognition (ICR), is a term used to describe the process for recognizing handwritten content. ICR algorithms require more intelligence than ordinary OCR. This post will cover the challenges of handwritten text identification and the techniques that can be used to tackle them using deep learning and machine learning. In the healthcare/pharmaceutical industry, patient medication digitization is a serious issue. Roche processes millions of PDFs each day, processing petabytes in medical PDFs.


Hindi Character Recognition

#artificialintelligence

Character recognition is a process that allows computers to recognize written or printed characters such as numbers or letters and to change them into a form that computers can use. As a part of this case study, we are going to recognize "Hindi characters". It is a Character Recognition problem related to computer vision, where our task is to predict the Hindi character present in the image. The Model should predict or recognize the character present in the image in real-time. So the latency of the model should be low.


Omnifont Persian OCR System Using Primitives

arXiv.org Artificial Intelligence

In this paper, we introduce a model-based omnifont Persian OCR system. The system uses a set of 8 primitive elements as structural features for recognition. First, the scanned document is preprocessed. After normalizing the preprocessed image, text rows and sub-words are separated and then thinned. After recognition of dots in sub-words, strokes are extracted and primitive elements of each sub-word are recognized using the strokes. Finally, the primitives are compared with a predefined set of character identification vectors in order to identify sub-word characters. The separation and recognition steps of the system are concurrent, eliminating unavoidable errors of independent separation of letters. The system has been tested on documents with 14 standard Persian fonts in 6 sizes. The achieved precision is 97.06%.


ANPR System, Number Plate Recognition

#artificialintelligence

Plate.Vision is a vehicle identification software through license plate recognition,



Cross-Lingual Text-to-Speech Using Multi-Task Learning and Speaker Classifier Joint Training

arXiv.org Artificial Intelligence

In cross-lingual speech synthesis, the speech in various languages can be synthesized for a monoglot speaker. Normally, only the data of monoglot speakers are available for model training, thus the speaker similarity is relatively low between the synthesized cross-lingual speech and the native language recordings. Based on the multilingual transformer text-to-speech model, this paper studies a multi-task learning framework to improve the cross-lingual speaker similarity. To further improve the speaker similarity, joint training with a speaker classifier is proposed. Here, a scheme similar to parallel scheduled sampling is proposed to train the transformer model efficiently to avoid breaking the parallel training mechanism when introducing joint training. By using multi-task learning and speaker classifier joint training, in subjective and objective evaluations, the cross-lingual speaker similarity can be consistently improved for both the seen and unseen speakers in the training set.


A guide to text detection and recognition using MMOCR

#artificialintelligence

Optical character recognition (OCR) is a sort of image conversion that basically extracts text from a given image, a document photo, etc. Various applications and technologies, such as Adobe Acrobat and the ML-based tool, such as Tesseract OCR, have been developed to aid with this process. In this article, we will go over tasks performed in the OCR method. Thereafter, we will look into MMOCR, a Python-based application that centralizes all OCR-related operations. Below are major points listed that are to be discussed in this article. Let's first discuss text detection.


Solving CAPTCHAs With Machine Learning to Enable Dark Web Research

#artificialintelligence

A joint academic research project from the United States has developed a method to foil CAPTCHA* tests, reportedly outperforming similar state-of-the-art machine learning solutions by using Generative Adversarial Networks (GANs) to decode the visually complex challenges. Testing the new system against the best current frameworks, the researchers found that their method achieves more than 94.4% success on a carefully curated real-world benchmark dataset, and has proved capable of'eliminating human involvement' when navigating a highly CAPTCHA-protected emerging Dark Net Marketplace, automatically resolving CAPTCHA challenges in a maximum of three attempts. The authors contend that their approach represents a breakthrough for cybersecurity researchers, who traditionally have had to bear the costs of supplying humans-in-the-loop to manually solve CAPTCHAs, usually via crowdsourcing platforms such as Amazon Mechanical Turk (AMT). If the system can prove adaptable and resilient, it may further pave the way for more automated oversight systems, and for the indexing and web-scraping of TOR networks. This could enable scalable and high-volume analyses, as well as the development of new cybersecurity approaches and techniques, which have been hamstrung, to date, by CAPTCHA firewalls.


XPeng upgrades EV voice assistant with Microsoft text-to-speech tech โ€“ FutureIoT

#artificialintelligence

With a deep understanding of urban mobility, we are finding many more scenarios to leverage AI technology for a high level of driver-machineย โ€ฆ


Artificial Intelligence: FinTech's innovation driver - BusinessWorld Online

#artificialintelligence

FinTech refers to any idea or innovation that improves or optimizes the way individuals or companies conduct financial activities. Early FinTech concentrated on developing add-on products to complement existing financial services. This combination of finance and technology has spawned a slew of valuable goods and services that redefine financial services and make them more accessible to the general public. Some of these products and services include insurance aggregators, mobile wallets, AI investment management advisers, peer-to-peer (P2P) lending and crowdfunding tools, and platforms for trading financial assets. The cutting-edge solutions that contributed to such technologies include Blockchain, Deep Learning, and Artificial Intelligence (AI).