Goto

Collaborating Authors

 South America


Applications of Machine Learning in Document Digitisation

arXiv.org Machine Learning

Data acquisition forms the primary step in all empirical research. The availability of data directly impacts the quality and extent of conclusions and insights. In particular, larger and more detailed datasets provide convincing answers even to complex research questions. The main problem is that 'large and detailed' usually implies 'costly and difficult', especially when the data medium is paper and books. Human operators and manual transcription have been the traditional approach for collecting historical data. We instead advocate the use of modern machine learning techniques to automate the digitisation process. We give an overview of the potential for applying machine digitisation for data collection through two illustrative applications. The first demonstrates that unsupervised layout classification applied to raw scans of nurse journals can be used to construct a treatment indicator. Moreover, it allows an assessment of assignment compliance. The second application uses attention-based neural networks for handwritten text recognition in order to transcribe age and birth and death dates from a large collection of Danish death certificates. We describe each step in the digitisation pipeline and provide implementation insights.


Hyperspherical embedding for novel class classification

arXiv.org Artificial Intelligence

Deep learning models have become increasingly useful in many different industries. On the domain of image classification, convolutional neural networks proved the ability to learn robust features for the closed set problem, as shown in many different datasets, such as MNIST FASHIONMNIST, CIFAR10, CIFAR100, and IMAGENET. These approaches use deep neural networks with dense layers with softmax activation functions in order to learn features that can separate classes in a latent space. However, this traditional approach is not useful for identifying classes unseen on the training set, known as the open set problem. A similar problem occurs in scenarios involving learning on small data. To tackle both problems, few-shot learning has been proposed. In particular, metric learning learns features that obey constraints of a metric distance in the latent space in order to perform classification. However, while this approach proves to be useful for the open set problem, current implementation requires pair-wise training, where both positive and negative examples of similar images are presented during the training phase, which limits the applicability of these approaches in large data or large class scenarios given the combinatorial nature of the possible inputs.In this paper, we present a constraint-based approach applied to the representations in the latent space under the normalized softmax loss, proposed by[18]. We experimentally validate the proposed approach for the classification of unseen classes on different datasets using both metric learning and the normalized softmax loss, on disjoint and joint scenarios. Our results show that not only our proposed strategy can be efficiently trained on larger set of classes, as it does not require pairwise learning, but also present better classification results than the metric learning strategies surpassing its accuracy by a significant margin.


Speedy robots gather spectra for sky surveys

Science

It was one of the stranger and more monotonous jobs in astronomy: plugging optical fibers into hundreds of holes in aluminum plates. Every day, technicians with the Sloan Digital Sky Survey (SDSS) prepped up to 10 plates that would be placed that night at the focus of the survey's telescopes in Chile and New Mexico. The holes matched the exact positions of stars, galaxies, or other bright objects in the telescopes' view. Light from each object fell directly on a fiber and was whisked off to a spectrograph, which split the light into its component wavelengths, revealing key details such as what the object is made of and how it is moving. Now, after 20 years, the SDSS is going robotic. For the project's upcoming fifth set of surveys, known as the SDSS-V, plug plates are being replaced by 500 tiny robot arms, each holding fiber tips that patrol a small area of the telescope's focal plane. They can be reconfigured for a new sky map in 2 minutes. Other sky surveys are also adopting the speedy robots. They will not only save valuable observation time, but also allow the surveys to keep up with Europe's Gaia satellite, the upcoming Vera C. Rubin Observatory in Chile, and other efforts that produce huge catalogs of objects needing spectroscopic study. “It's driven by the science of enormous imaging surveys,” says astronomer Richard Ellis of University College London. COVID-19 has delayed the SDSS's robotic makeover. The survey's northern telescope at Apache Point Observatory in New Mexico began to take SDSS-V data in October 2020 using plug plates. It aims to switch over to the robots by mid-2021. The southern scope at Las Campanas Observatory in Chile will follow later in the year. “It's bananas,” says SDSS-V Director Juna Kollmeier of the Carnegie Observatories, “but we're seeing the end of the tunnel.” The robots mark a new chapter for the SDSS. For 10 years much of its time went to the study of dark energy, the mysterious force that is accelerating the universe's expansion. The SDSS prised apart the light of millions of galaxies to determine their distance, via a redshift—a Doppler shift in their light due to the expansion of the universe, like the wail of a receding siren. Results from the galaxy survey, released in July 2020, traced the universe's expansion back through 80% of its history with 1% precision, confirming the effects of dark energy, perhaps the biggest mystery in cosmology. Cracking it will require looking further back in time to fainter galaxies, which is beyond the capabilities of the survey's 2.5-meter telescopes. Instead, the scopes will carry out three new surveys. Milky Way Mapper will gather spectra from 6 million stars, probing their composition to find out how long they've been burning and forging heavy elements. “Stars are all clocks,” Kollmeier explains. With age estimates, astronomers can work out when parts of the Milky Way formed. Subtle shifts in composition can also reveal whether a group of stars originated in another galaxy or star cluster that has been subsumed into ours—an unwinding of Milky Way history called galactic archaeology. In a second survey, Black Hole Mapper, the optical fibers will gather light from bright galaxies to learn about the supermassive black holes they harbor. Doppler shifts in the spectra of glowing gases surrounding these black holes could reveal how fast they fling this material around—and thus how heavy they are. Shifts in the spectra could trace how they gobble up and spit out streams of this gas. By tracking the gases over time, Kollmeier says, astronomers may learn how the black holes grow, seemingly in concert with their galaxies. The third survey, Local Volume Mapper, will bunch fibers together like a multipixel detector to get spectra from clouds of interstellar gas within nearby galaxies. “We're mapping a whole galaxy in exquisite detail at one time,” Kollmeier says. By determining the motions and composition of the gas clouds, the SDSS team hopes to identify why some collapse into stars and others don't. Meanwhile, the dark energy quest pioneered by the SDSS will move to the Dark Energy Spectroscopic Instrument, a 5000-fiber robotic spectrograph on a 4-meter telescope in Arizona. It will soon begin to track the distances to tens of millions of galaxies in the remote universe ( Science , 13 September 2019, p. [1066][1]). ![Figure][2] In the coming months, the William Herschel Telescope, a 4.2-meter telescope in the Canary Islands, will join the robot revolution by sending light to a 1000-fiber spectrograph called the WHT Enhanced Area Velocity Explorer (WEAVE). Instead of using robots to hold fibers in place, WEAVE has two of them working offline, picking and placing magnetic fiber ends onto a metal plate—automating what the SDSS's plate pluggers did. One of WEAVE's goals is to gather Doppler shifts from the billion stars Gaia has mapped, nailing down their full 3D motions. Then, “We can run the clock backwards and see where they came from,” says project scientist Scott Trager of the University of Groningen. It's another way to do galactic archeology. Next year, the European Southern Observatory's (ESO's) 4-metre Multi-Object Spectroscopic Telescope in Chile will be fitted with yet another robotic technology. Its 2400 fibers will be fed through controllable “spines” that stick up into the telescope's focal plane and can be made to move, like wheat stalks in a breeze. Like WEAVE, it will follow up on sources identified by European spacecraft, including Gaia and Euclid, an upcoming dark energy mission. It and other fiber spectrographs will also help with studies of fast-moving cosmic events such as supernovae or the violent collisions that produce gravitational waves. The Rubin Observatory will spot many of them. From 2023, it's expected to detect 10 million fast-changing objects every night. For the thousands that demand scrutiny, “spectra are really critical for understanding what a source is,” says Eric Bellm of the University of Washington, Seattle, who is the science lead for Rubin's alert stream. Even some of the world's largest scopes, in the 8-meter range, are adding robotic spectrographs. Japan's Subaru and ESO's Very Large Telescope are both developing systems that will vacuum up spectra from faint, distant objects. Ellis says a fiber spectrograph combined with Subaru's 8.2-meter mirror would be able to pick out spectra of individual stars in the Andromeda galaxy, the Milky Way's nearby twin. “With a big telescope, we can do galactic archaeology in our nearest neighbor,” he says. [1]: http://www.sciencemag.org/content/365/6458/1066 [2]: pending:yes


RECol: Reconstruction Error Columns for Outlier Detection

arXiv.org Machine Learning

Detecting outliers or anomalies is a common data analysis task. As a sub-field of unsupervised machine learning, a large variety of approaches exist, but the vast majority treats the input features as independent and often fails to recognize even simple (linear) relationships in the input feature space. Hence, we introduce RECol, a generic data pre-processing approach to generate additional columns in a leave-one-out-fashion: For each column, we try to predict its values based on the other columns, generating reconstruction error columns. We run experiments across a large variety of common baseline approaches and benchmark datasets with and without our RECol pre-processing method and show that the generated reconstruction error feature space generally seems to support common outlier detection methods and often considerably improves their ROC-AUC and PR-AUC values.


The effect of differential victim crime reporting on predictive policing systems

arXiv.org Machine Learning

Police departments around the world have been experimenting with forms of place-based data-driven proactive policing for over two decades. Modern incarnations of such systems are commonly known as hot spot predictive policing. These systems predict where future crime is likely to concentrate such that police can allocate patrols to these areas and deter crime before it occurs. Previous research on fairness in predictive policing has concentrated on the feedback loops which occur when models are trained on discovered crime data, but has limited implications for models trained on victim crime reporting data. We demonstrate how differential victim crime reporting rates across geographical areas can lead to outcome disparities in common crime hot spot prediction models. Our analysis is based on a simulation patterned after district-level victimization and crime reporting survey data for Bogot\'a, Colombia. Our results suggest that differential crime reporting rates can lead to a displacement of predicted hotspots from high crime but low reporting areas to high or medium crime and high reporting areas. This may lead to misallocations both in the form of over-policing and under-policing.


High-level Approaches to Detect Malicious Political Activity on Twitter

arXiv.org Artificial Intelligence

Our work represents another step into the detection and prevention of these ever-more present political manipulation efforts. We, therefore, start by focusing on understanding what the state-of-the-art approaches lack -- since the problem remains, this is a fair assumption. We find concerning issues within the current literature and follow a diverging path. Notably, by placing emphasis on using data features that are less susceptible to malicious manipulation and also on looking for high-level approaches that avoid a granularity level that is biased towards easy-to-spot and low impact cases. We designed and implemented a framework -- Twitter Watch -- that performs structured Twitter data collection, applying it to the Portuguese Twittersphere. We investigate a data snapshot taken on May 2020, with around 5 million accounts and over 120 million tweets (this value has since increased to over 175 million). The analyzed time period stretches from August 2019 to May 2020, with a focus on the Portuguese elections of October 6th, 2019. However, the Covid-19 pandemic showed itself in our data, and we also delve into how it affected typical Twitter behavior. We performed three main approaches: content-oriented, metadata-oriented, and network interaction-oriented. We learn that Twitter's suspension patterns are not adequate to the type of political trolling found in the Portuguese Twittersphere -- identified by this work and by an independent peer - nor to fake news posting accounts. We also surmised that the different types of malicious accounts we independently gathered are very similar both in terms of content and interaction, through two distinct analysis, and are simultaneously very distinct from regular accounts.


Controlling Hallucinations at Word Level in Data-to-Text Generation

arXiv.org Artificial Intelligence

Data-to-Text Generation (DTG) is a subfield of Natural Language Generation aiming at transcribing structured data in natural language descriptions. The field has been recently boosted by the use of neural-based generators which exhibit on one side great syntactic skills without the need of hand-crafted pipelines; on the other side, the quality of the generated text reflects the quality of the training data, which in realistic settings only offer imperfectly aligned structure-text pairs. Consequently, state-of-art neural models include misleading statements - usually called hallucinations - in their outputs. The control of this phenomenon is today a major challenge for DTG, and is the problem addressed in the paper. Previous work deal with this issue at the instance level: using an alignment score for each table-reference pair. In contrast, we propose a finer-grained approach, arguing that hallucinations should rather be treated at the word level. Specifically, we propose a Multi-Branch Decoder which is able to leverage word-level labels to learn the relevant parts of each training instance. These labels are obtained following a simple and efficient scoring procedure based on co-occurrence analysis and dependency parsing. Extensive evaluations, via automated metrics and human judgment on the standard WikiBio benchmark, show the accuracy of our alignment labels and the effectiveness of the proposed Multi-Branch Decoder. Our model is able to reduce and control hallucinations, while keeping fluency and coherence in generated texts. Further experiments on a degraded version of ToTTo show that our model could be successfully used on very noisy settings.


Hierarchical Multi-head Attentive Network for Evidence-aware Fake News Detection

arXiv.org Artificial Intelligence

To detect fake news, researchers proposed to use The proliferation of biased news, misleading linguistics and textual content (Castillo et al., 2011; claims, disinformation and fake news has caused Zhao et al., 2015; Liu et al., 2015). Since textual heightened negative effects on modern society in claims are usually deliberately written to deceive various domains ranging from politics, economics readers, it is hard to detect fake news by solely to public health. A recent study showed that maliciously relying on the content claims. Therefore, multiple fabricated and partisan stories possibly works utilized other signals such as temporal caused citizens' misperception about political candidates spreading patterns (Liu and Wu, 2018), network (Allcott and Gentzkow, 2017) during the structures (Wu and Liu, 2018; Vo and Lee, 2018; 2016 U.S. presidential elections. In economics, the Shu et al., 2020) and users' feedbacks (Vo and spread of fake news has manipulated stock price Lee, 2019; Shu et al., 2019; Vo and Lee, 2020a).


Exploring Scale-Measures of Data Sets

arXiv.org Artificial Intelligence

An inevitable step of any data-based knowledge discovery process is measurement [24] and the associated (explicit or implicit) scaling of the data [27]. The latter is particularly constrained by the underlying mathematical formulation of the data representation, e.g., real-valued vector spaces or weighted graphs, the requirements of the data procedures, e.g., the presence of a distance function, and, more recently, the need for human understanding of the results. Considering the scaling of data as part of the analysis itself, in particular formalizing it and thus making it controllable, is a salient feature of formal concept analysis (FCA) [7]. This field of research has spawned a variety of specialized scaling methods, such as logical scaling [25], and in the form of scale-measures links the scaling process with the study of continuous mappings between closure systems. Recent results by the authors [13] revealed that the set of all scale-measures for a given data set constitutes a lattice. Furthermore, it was shown that any scale-measure can be expressed in simple propositional terms using disjunction, conjunction and negation. Among other things, the previous results allow a computational transition between different scale-measures, which we may call scalemeasure navigation, as well as their interpretability by humans.


Persistent Rule-based Interactive Reinforcement Learning

arXiv.org Artificial Intelligence

Interactive reinforcement learning has allowed speeding up the learning process in autonomous agents by including a human trainer providing extra information to the agent in real-time. Current interactive reinforcement learning research has been limited to interactions that offer relevant advice to the current state only. Additionally, the information provided by each interaction is not retained and instead discarded by the agent after a single-use. In this work, we propose a persistent rule-based interactive reinforcement learning approach, i.e., a method for retaining and reusing provided knowledge, allowing trainers to give general advice relevant to more than just the current state. Our experimental results show persistent advice substantially improves the performance of the agent while reducing the number of interactions required for the trainer. Moreover, rule-based advice shows similar performance impact as state-based advice, but with a substantially reduced interaction count.