Goto

Collaborating Authors

 Government


Distilling Wikipedia mathematical knowledge into neural network models

arXiv.org Artificial Intelligence

Machine learning applications to symbolic mathematics are becoming increasingly popular, yet there lacks a centralized source of real-world symbolic expressions to be used as training data. In contrast, the field of natural language processing leverages resources like Wikipedia that provide enormous amounts of realworld textual data. Adopting the philosophy of "mathematics as language," we bridge this gap by introducing a pipeline for distilling mathematical expressions embedded in Wikipedia into symbolic encodings to be used in downstream machine learning tasks. We demonstrate that a mathematical language model trained on this "corpus" of expressions can be used as a prior to improve the performance of neural-guided search for the task of symbolic regression. "The basis of all human culture is language, and mathematics is a special kind of linguistic activity."


Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders

arXiv.org Artificial Intelligence

The growing deluge of scientific publications demands text analysis tools that can help scientists and policy-makers navigate, forecast and beneficially guide scientific research. Recent advances in natural language understanding driven by deep transformer networks offer new possibilities for mapping science. Because the same surface text can take on multiple and sometimes contradictory specialized senses across distinct research communities, sensitivity to context is critical for infometric applications. Transformer embedding models such as BERT capture shades of association and connotation that vary across the different linguistic contexts of any particular word or span of text. Here we report a procedure for encoding scientific documents with these tools, measuring their improvement over static word embeddings in a nearest-neighbor retrieval task. We find discriminability of contextual representations is strongly influenced by choice of pooling strategy for summarizing the high-dimensional network activations. Importantly, we note that fundamentals such as domain-matched training data are more important than state-of-the-art NLP tools. Yet state-of-the-art models did offer significant gains. The best approach we investigated combined domain-matched pretraining, sound pooling, and state-of-the-art deep transformer network encoders. Finally, with the goal of leveraging contextual representations from deep encoders, we present a range of measurements for understanding and forecasting research communities in science.


The Many Faces of 1-Lipschitz Neural Networks

arXiv.org Artificial Intelligence

Lipschitz constrained models have been used to solve specifics deep learning problems such as the estimation of Wasserstein distance for GAN, or the training of neural networks robust to adversarial attacks. Regardless the novel and effective algorithms to build such 1-Lipschitz networks, their usage remains marginal, and they are commonly considered as less expressive and less able to fit properly the data than their unconstrained counterpart. The goal of this paper is to demonstrate that, despite being empirically harder to train, 1-Lipschitz neural networks are theoretically better grounded than unconstrained ones when it comes to classification. We recall some results about 1-Lipschitz functions in the scope of deep learning and we extend and illustrate them to derive general properties for classification. We propose and demonstrate several new properties of 1-Lipschitz neural networks for classification. First, we show they can fit arbitrarily difficult frontiers, making them as expressive as classical ones, in addition to provide robustness certificates. We prove that when minimizing cross entropy loss the optimization problem under Lipschitz constraint is well posed and its solution generalizes well in the limit of big datasets, whereas regular neural networks can diverge even on remarkably simple situations. Then, we study the link between classification with 1-Lipschitz network and optimal transport thanks to regularized versions of Kantorovich-Rubinstein duality theory. Last, we derive preliminary bounds on their VC dimensions.


This AI Could Help Wipe Out Colon Cancer

WIRED

Michael Wallace has performed hundreds of colonoscopies in his 20 years as a gastroenterologist. He thinks he's pretty good at recognizing the growths, or polyps, that can spring up along the ridges of the colon and potentially turn into cancer. Sometimes the polyps are flat and hard to see. Other times, doctors just miss them. "We're all humans," says Wallace, who works at the Mayo Clinic.


Investigating Methods to Improve Language Model Integration for Attention-based Encoder-Decoder ASR Models

arXiv.org Machine Learning

Attention-based encoder-decoder (AED) models learn an implicit internal language model (ILM) from the training transcriptions. The integration with an external LM trained on much more unpaired text usually leads to better performance. A Bayesian interpretation as in the hybrid autoregressive transducer (HAT) suggests dividing by the prior of the discriminative acoustic model, which corresponds to this implicit LM, similarly as in the hybrid hidden Markov model approach. The implicit LM cannot be calculated efficiently in general and it is yet unclear what are the best methods to estimate it. In this work, we compare different approaches from the literature and propose several novel methods to estimate the ILM directly from the AED model. Our proposed methods outperform all previous approaches. We also investigate other methods to suppress the ILM mainly by decreasing the capacity of the AED model, limiting the label context, and also by training the AED model together with a pre-existing LM.


Sparse Coding Frontend for Robust Neural Networks

arXiv.org Machine Learning

Deep Neural Networks are known to be vulnerable to small, adversarially crafted, perturbations. The current most effective defense methods against these adversarial attacks are variants of adversarial training. In this paper, we introduce a radically different defense trained only on clean images: a sparse coding based front end which significantly attenuates adversarial attacks before they reach the classifier. Deep neural networks (DNNs) are known to be vulnerable to small, adversarially designed perturbations (Biggio et al., 2013; Szegedy et al., 2014). Since the discovery of such adversarial attacks, researchers have tried various defense strategies such as adversarial training (Madry et al., 2018; Zhang et al., 2019), constraining Lipschitz constant (Cisse et al., 2017), randomized smoothing (Cohen et al., 2019), and preprocessing methods (Guo et al., 2017; Yang et al., 2019).


Macro-Average: Rare Types Are Important Too

arXiv.org Artificial Intelligence

While traditional corpus-level evaluation metrics for machine translation (MT) correlate well with fluency, they struggle to reflect adequacy. Model-based MT metrics trained on segment-level human judgments have emerged as an attractive replacement due to strong correlation results. These models, however, require potentially expensive re-training for new domains and languages. Furthermore, their decisions are inherently non-transparent and appear to reflect unwelcome biases. We explore the simple type-based classifier metric, MacroF1, and study its applicability to MT evaluation. We find that MacroF1 is competitive on direct assessment, and outperforms others in indicating downstream cross-lingual information retrieval task performance. Further, we show that MacroF1 can be used to effectively compare supervised and unsupervised neural machine translation, and reveal significant qualitative differences in the methods' outputs.


Towards Algorithmic Transparency: A Diversity Perspective

arXiv.org Artificial Intelligence

As the role of algorithmic systems and processes increases in society, so does the risk of bias, which can result in discrimination against individuals and social groups. Research on algorithmic bias has exploded in recent years, highlighting both the problems of bias, and the potential solutions, in terms of algorithmic transparency (AT). Transparency is important for facilitating fairness management as well as explainability in algorithms; however, the concept of diversity, and its relationship to bias and transparency, has been largely left out of the discussion. We reflect on the relationship between diversity and bias, arguing that diversity drives the need for transparency. Using a perspective-taking lens, which takes diversity as a given, we propose a conceptual framework to characterize the problem and solution spaces of AT, to aid its application in algorithmic systems. Example cases from three research domains are described using our framework.


The Limits of Political Debate

The New Yorker

In February, 2011, an Israeli computer scientist named Noam Slonim proposed building a machine that would be better than people at something that seems inextricably human: arguing about politics. Slonim, who had done his doctoral work on machine learning, works at an I.B.M. Research facility in Tel Aviv, and he had watched with pride a few days before as the company's natural-language-processing machine, Watson, won "Jeopardy!" Afterward, I.B.M. sent an e-mail to thousands of researchers across its global network of labs, soliciting ideas for a "grand challenge" to follow the "Jeopardy!" It occurred to Slonim that they might try to build a machine that could defeat a champion debater. He made a single-slide presentation, and then a somewhat more elaborate one, and then a more elaborate one still, and, after many rounds competing against many other I.B.M. researchers, Slonim won the chance to build his machine, which he called Project Debater.


The World as a Graph: Improving El Ni\~no Forecasts with Graph Neural Networks

arXiv.org Machine Learning

Deep learning-based models have recently outperformed state-of-the-art seasonal forecasting models, such as for predicting El Ni\~no-Southern Oscillation (ENSO). However, current deep learning models are based on convolutional neural networks which are difficult to interpret and can fail to model large-scale atmospheric patterns. In comparison, graph neural networks (GNNs) are capable of modeling large-scale spatial dependencies and are more interpretable due to the explicit modeling of information flow through edge connections. We propose the first application of graph neural networks to seasonal forecasting. We design a novel graph connectivity learning module that enables our GNN model to learn large-scale spatial interactions jointly with the actual ENSO forecasting task. Our model, \graphino, outperforms state-of-the-art deep learning-based models for forecasts up to six months ahead. Additionally, we show that our model is more interpretable as it learns sensible connectivity structures that correlate with the ENSO anomaly pattern.