Optical Character Recognition
A Black-Box Attack on Optical Character Recognition Systems
Bayram, Samet, Barner, Kenneth
Adversarial machine learning is an emerging area showing the vulnerability of deep learning models. Exploring attack methods to challenge state of the art artificial intelligence (A.I.) models is an area of critical concern. The reliability and robustness of such A.I. models are one of the major concerns with an increasing number of effective adversarial attack methods. Classification tasks are a major vulnerable area for adversarial attacks. The majority of attack strategies are developed for colored or gray-scaled images. Consequently, adversarial attacks on binary image recognition systems have not been sufficiently studied. Binary images are simple two possible pixel-valued signals with a single channel. The simplicity of binary images has a significant advantage compared to colored and gray scaled images, namely computation efficiency. Moreover, most optical character recognition systems (O.C.R.s), such as handwritten character recognition, plate number identification, and bank check recognition systems, use binary images or binarization in their processing steps. In this paper, we propose a simple yet efficient attack method, Efficient Combinatorial Black-box Adversarial Attack, on binary image classifiers. We validate the efficiency of the attack technique on two different data sets and three classification networks, demonstrating its performance. Furthermore, we compare our proposed method with state-of-the-art methods regarding advantages and disadvantages as well as applicability.
OCR is getting super cool for Businesses
A Few months back, the student in class captured the image of the notes made by the other student in front of him and used iOS 15's recent text-recognition feature to highlight text, and copy and paste it into his notes. This instance was tweeted by @juanbuis, who shared the video of a student making the most of iOS 15's Live Text OCR feature. This cool OCR or Optical Character Recognition feature that the above student opts for is generally used to pull up the information from the text or documents and then convert it into the machine's language. Recently, the popular app developer Alessandro Paluzzi has also seen that Twitter is working on an OCR (optical character recognition) feature for the description of alt text. In his tweet, Alessandro Paluzzi shared the demonstration of how this twitter feature will function through a short video. At Dwarf AI we too want to make this super cool technology to be easily accessible by other businesses.
Optical Character Recognition โข OCR โข AI Terminology โข AI Blog
Optical character recognition (OCR) is a type of technology that allows computers to read text from images and convert it into digital data. This can be extremely useful for digitizing old documents or converting scanned images into editable text. OCR technology relies on pattern recognition to identify the shapes of letters and numbers in an image. It then compares these patterns to a database of known characters, in order to determine what the text says. OCR technology has come a long way in recent years, and can now reliably recognize a wide range of fonts and languages. However, it is not perfect, and can sometimes struggle with words that are poorly lit, blurry, or in a particularly ornate font.
An End-to-End OCR Framework for Robust Arabic-Handwriting Recognition using a Novel Transformers-based Model and an Innovative 270 Million-Words Multi-Font Corpus of Classical Arabic with Diacritics
Mostafa, Aly, Mohamed, Omar, Ashraf, Ali, Elbehery, Ahmed, Jamal, Salma, Salah, Anas, Ghoneim, Amr S.
This research is the second phase in a series of investigations on developing an Optical Character Recognition (OCR) of Arabic historical documents and examining how different modeling procedures interact with the problem. The first research studied the effect of Transformers on our custom-built Arabic dataset. One of the downsides of the first research was the size of the training data, a mere 15000 images from our 30 million images, due to lack of resources. Also, we add an image enhancement layer, time and space optimization, and Post-Correction layer to aid the model in predicting the correct word for the correct context. Notably, we propose an end-to-end text recognition approach using Vision Transformers as an encoder, namely BEIT, and vanilla Transformer as a decoder, eliminating CNNs for feature extraction and reducing the model's complexity. The experiments show that our end-to-end model outperforms Convolutions Backbones. The model attained a CER of 4.46%.
SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis
Maniati, Georgia, Vioni, Alexandra, Ellinas, Nikolaos, Nikitaras, Karolos, Klapsas, Konstantinos, Sung, June Sig, Jho, Gunu, Chalamandaris, Aimilios, Tsiakoulis, Pirros
In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the assessment of modern synthesizers, and can stimulate advancements in acoustic model evaluation. It consists of 20K synthetic utterances of the LJ Speech voice, a public domain speech dataset which is a common benchmark for building neural acoustic models and vocoders. Utterances are generated from 200 TTS systems including vanilla neural acoustic models as well as models which allow prosodic variations. An LPCNet vocoder is used for all systems, so that the samples' variation depends only on the acoustic models. The synthesized utterances provide balanced and adequate domain and length coverage. We collect MOS naturalness evaluations on 3 English Amazon Mechanical Turk locales and share practices leading to reliable crowdsourced annotations for this task. We provide baseline results of state-of-the-art MOS prediction models on the SOMOS dataset and show the limitations that such models face when assigned to evaluate TTS utterances.
To show or not to show: Redacting sensitive text from videos of electronic displays
Mukhopadhyay, Abhishek, Agarwal, Shubham, Zwick, Patrick Dylan, Biswas, Pradipta
This, combined with major developments in computer vision and machine learning technology, has created enormous opportunities to make life better through the collection and utilization of this video data. Potential applications here range from improved security to interactive entertainment. However, the collection and utilization of this data also entails ethical privacy concerns and the potential for unwanted intrusion into people's lives without their permission. One way to attempt to achieve the benefits of more omnipresent video collection while mitigating the intrusion on privacy is through the automatic redaction of personally identifiable information (PII). This means automatically removing or obscuring content from video data that can be used to identify an individual while maintaining as much other video data as possible. A relatively new context generating a significant amount of video data is the cabins of automobiles.
The next PowerToy will give your PC easy OCR powers
We love us some PowerToys here at PCWorld, and it seems the semi-official add-on for Windows power users is only getting better. A recent update to the GitHub project for PowerToys indicates that "PowerOCR" is in the latter stages of approval and should land in the official app before too long. The tool will add in an easy way for Windows users to activate Optical Character Recognition (OCR) via a quick screenshot interface. NeoWin spotted the changes to the PowerToys Github, with extensive documentation of the new PowerOCR tool and an apparent approval nod from a Microsoft manager. The tool is mostly the work of independent developer Joesph Finney, contributing code that works similarly to his paid Text Grab app.
Self-paced learning to improve text row detection in historical documents with missing labels
Gaman, Mihaela, Ghadamiyan, Lida, Ionescu, Radu Tudor, Popescu, Marius
An important preliminary step of optical character recognition systems is the detection of text rows. To address this task in the context of historical data with missing labels, we propose a self-paced learning algorithm capable of improving the row detection performance. We conjecture that pages with more ground-truth bounding boxes are less likely to have missing annotations. Based on this hypothesis, we sort the training examples in descending order with respect to the number of ground-truth bounding boxes, and organize them into k batches. Using our self-paced learning method, we train a row detector over k iterations, progressively adding batches with less ground-truth annotations. At each iteration, we combine the ground-truth bounding boxes with pseudo-bounding boxes (bounding boxes predicted by the model itself) using non-maximum suppression, and we include the resulting annotations at the next training iteration. We demonstrate that our self-paced learning strategy brings significant performance gains on two data sets of historical documents, improving the average precision of YOLOv4 with more than 12% on one data set and 39% on the other.
Level Up Your AI Skillset and Dive Into The Deep End Of TinyML
Machine learning (ML) is a growing field, gaining popularity in academia, industry, and among makers. We will take a look at some of the available tools to help make machine learning easier, but first, let's review some of the terms commonly used in machine learning. John McCarthy provides a definition of artificial intelligence (AI) in his 2007 Stanford paper, "What is Artificial Intelligence?" In it, he says AI "is the science and engineering of making intelligent machines, especially intelligent computer programs." This definition is extremely broad, as McCarthy defines intelligence as "the computational part of the ability to achieve goals in the world." As a result, any program that achieves some goal can easily be classified as artificial intelligence. In her article "Machine Learning on Microcontrollers" (Make: Vol.
Leaks - chessie rae
Official - OCR - Convert image to text Multi Language Marks-Man submitted a new resource: [TheJavaSea] OCR - Convert image to text - Perform basic OCR (Optical Character Recognition) in English and 100 other languages. Our OCR application allows you to perform basic OCR (Optical Character Recognition) in English and 100 other languages. You must upgrade your account or reply to the thread to see the hidden content.