Deep Learning
Composition-based Multi-Relational Graph Convolutional Networks
Vashishth, Shikhar, Sanyal, Soumya, Nitin, Vikram, Talukdar, Partha
Graph Convolutional Networks (GCNs) have recently been shown to be quite successful in modeling graph-structured data. However, the primary focus has been on handling simple undirected graphs. Multi-relational graphs are a more general and prevalent form of graphs where each edge has a label and direction associated with it. Most of the existing approaches to handle such graphs suffer from over-parameterization and are restricted to learning representations of nodes only. We evaluate our proposed method on multiple tasks such as node classification, link prediction, and graph classification, and achieve demonstrably superior results. GCN available to foster reproducible research. Graphs are one of the most expressive data-structures which have been used to model a variety of problems. Traditional neural network architectures like Convolutional Neural Networks (Krizhevsky et al., 2012) and Recurrent Neural Networks (Hochreiter & Schmidhuber, 1997) are constrained to handle only Euclidean data.
Deep geometric knowledge distillation with graphs
Lassance, Carlos, Bontonou, Myriam, Hacene, Ghouthi Boukli, Gripon, Vincent, Tang, Jian, Ortega, Antonio
In most cases deep learning architectures are trained disregarding the amount of operations and energy consumption. However, some applications, like embedded systems, can be resource-constrained during inference. A popular approach to reduce the size of a deep learning architecture consists in distilling knowledge from a bigger network (teacher) to a smaller one (student). Directly training the student to mimic the teacher representation can be effective, but it requires that both share the same latent space dimensions. In this work, we focus instead on relative knowledge distillation (RKD), which considers the geometry of the respective latent spaces, allowing for dimension-agnostic transfer of knowledge. Specifically we introduce a graph-based RKD method, in which graphs are used to capture the geometry of latent spaces. Using classical computer vision benchmarks, we demonstrate the ability of the proposed method to efficiently distillate knowledge from the teacher to the student, leading to better accuracy for the same budget as compared to existing RKD alternatives.
Memory Augmented Recursive Neural Networks
Arabshahi, Forough, Lu, Zhichu, Singh, Sameer, Anandkumar, Animashree
Recursive neural networks have shown an impressive performance for modeling compositional data compared to their recurrent counterparts. Although recursive neural networks are better at capturing long range dependencies, their generalization performance starts to decay as the test data becomes more compositional and potentially deeper than the training data. In this paper, we present memory-augmented recursive neural networks to address this generalization performance loss on deeper data points. We augment Tree-LSTMs with an external memory, namely neural stacks. We define soft push and pop operations for filling and emptying the memory to ensure that the networks remain end-to-end differentiable. In order to assess the effectiveness of the external memory, we evaluate our model on a neural programming task introduced in the literature called equation verification. Our results indicate that augmenting recursive neural networks with external memory consistently improves the generalization performance on deeper data points compared to the state-of-the-art Tree-LSTM by up to 10%.
Unmasking DeepFakes with simple Features
Durall, Ricard, Keuper, Margret, Pfreundt, Franz-Josef, Keuper, Janis
Deep generative models have recently achieved impressive results for many real-world applications, successfully generating high-resolution and diverse samples from complex datasets. Due to this improvement, fake digital contents have proliferated growing concern and spreading distrust in image content, leading to an urgent need for automated ways to detect these AI-generated fake images. Despite the fact that many face editing algorithms seem to produce realistic human faces, upon closer examination, they do exhibit artifacts in certain domains which are often hidden to the naked eye. In this work, we present a simple way to detect such fake face images - so-called DeepFakes. Our method is based on a classical frequency domain analysis followed by basic classifier. Compared to previous systems, which need to be fed with large amounts of labeled data, our approach showed very good results using only a few annotated training samples and even achieved good accuracies in fully unsupervised scenarios. For the evaluation on high resolution face images, we combined several public datasets of real and fake faces into a new benchmark: Faces-HQ. Given such high-resolution images, our approach reaches a perfect classification accuracy of 100% when it is trained on as little as 20 annotated samples. In a second experiment, in the evaluation of the medium-resolution images of the CelebA dataset, our method achieves 100% accuracy supervised and 96% in an unsupervised setting. Finally, evaluating a low-resolution video sequences of the FaceForensics++ dataset, our method achieves 91% accuracy detecting manipulated videos. Source Code: https://github.com/cc-hpc-itwm/DeepFakeDetection
ERASER: A Benchmark to Evaluate Rationalized NLP Models
DeYoung, Jay, Jain, Sarthak, Rajani, Nazneen Fatema, Lehman, Eric, Xiong, Caiming, Socher, Richard, Wallace, Byron C.
State-of-the-art models in NLP are now predominantly based on deep neural networks that are generally opaque in terms of how they come to specific predictions. This limitation has led to increased interest in designing more interpretable deep models for NLP that can reveal the `reasoning' underlying model outputs. But work in this direction has been conducted on different datasets and tasks with correspondingly unique aims and metrics; this makes it difficult to track progress. We propose the Evaluating Rationales And Simple English Reasoning (ERASER) benchmark to advance research on interpretable models in NLP. This benchmark comprises multiple datasets and tasks for which human annotations of "rationales" (supporting evidence) have been collected. We propose several metrics that aim to capture how well the rationales provided by models align with human rationales, and also how faithful these rationales are (i.e., the degree to which provided rationales influenced the corresponding predictions). Our hope is that releasing this benchmark facilitates progress on designing more interpretable NLP systems. The benchmark, code, and documentation are available at: www.eraserbenchmark.com .
Ask to Learn: A Study on Curiosity-driven Question Generation
Scialom, Thomas, Staiano, Jacopo
We propose a novel text generation task, namely Curiosity-driven Question Generation. We start from the observation that the Question Generation task has traditionally been considered as the dual problem of Question Answering, hence tackling the problem of generating a question given the text that contains its answer. Such questions can be used to evaluate machine reading comprehension. However, in real life, and especially in conversational settings, humans tend to ask questions with the goal of enriching their knowledge and/or clarifying aspects of previously gathered information. We refer to these inquisitive questions as Curiosity-driven: these questions are generated with the goal of obtaining new information (the answer) which is not present in the input text. In this work, we experiment on this new task using a conversational Question Answering (QA) dataset; further, since the majority of QA dataset are not built in a conversational manner, we describe a methodology to derive data for this novel task from non-conversational QA data. We investigate several automated metrics to measure the different properties of Curious Questions, and experiment different approaches on the Curiosity-driven Question Generation task, including model pre-training and reinforcement learning. Finally, we report a qualitative evaluation of the generated outputs.
Which Deep Learning Framework is Growing Fastest? - KDnuggets
In September 2018, I compared all the major deep learning frameworks in terms of demand, usage, and popularity in this article. TensorFlow was the undisputed heavyweight champion of deep learning frameworks. PyTorch was the young rookie with lots of buzz. How has the landscape changed for the leading deep learning frameworks in the past six months? To answer that question, I looked at the number of job listings on Indeed, Monster, LinkedIn, and SimplyHired.
Why the left should worry more about AI
I spend a disproportionate amount of time reading and talking to two somewhat niche groups of people in American politics: democratic socialists of the Sen. Bernie Sanders variety (or maybe a bit to the left of that), and left-libertarians from the Bay Area who are interested in effective altruism. These are both small groups, but they have social and intellectual influence bigger than their numbers. And while from a distance they look similar (I'm sure they both vote for Democrats in general elections, say), there's a big issue on which they part ways where collaboration could be productive: artificial intelligence safety. Effective altruists have, for complex sociological reasons I explored in a podcast episode, become very interested in AI as a potential "existential risk": a force that could, in extreme circumstances, wipe out humanity, just as nuclear war or asteroid strikes could. Kelsey Piper has a comprehensive Vox explainer of these arguments, and I take them seriously, but most friends to my left do not.
Why Are American Companies Helping China Build an Artificial Intelligence Authoritarian State?
President Xi Jinping wants China to dominate artificial intelligence by 2030. But all it seems the Middle Kingdom's new AI entrepreneurs want to talk about ancient history. Take Kai-Fu Lee, for example. The founder of Face was born in Taiwan and emigrated to America when he was eleven. After earning a Ph.D. from Carnegie Mellon in the 1980s, he worked his way up at Apple, Microsoft, and eventually, Google. In his book AI Superpowers, he likens Chinese entrepreneurs to "gladiators" fighting in a new techwar arena, where it's "kill or be killed."
AI deemed 'too dangerous to release' makes it out into the world
An AI that was deemed too dangerous to be released has now been released into the world. Researchers had feared that the model, known as "GPT-2", was so powerful that it could be maliciously misused by everyone from politicians to scammers. GPT-2 was created for a simple purpose: it can be fed a piece of text, and is able to predict the words that will come next. By doing so, it is able to create long strings of writing that are largely indistinguishable from those written by a human being. But it became clear that it was worryingly good at that job, with its text creation so powerful that it could be used to scam people and may undermine trust in the things we read.