Europe
An Analysis of Hierarchical Text Classification Using Word Embeddings
Stein, Roger A., Jaques, Patricia A., Valiati, Joao F.
Efficient distributed numerical word representation models (word embeddings) combined with modern machine learning algorithms have recently yielded considerable improvement on automatic document classification tasks. However, the effectiveness of such techniques has not been assessed for the hierarchical text classification (HTC) yet. This study investigates the application of those models and algorithms on this specific problem by means of experimentation and analysis. We trained classification models with prominent machine learning algorithm implementations---fastText, XGBoost, SVM, and Keras' CNN---and noticeable word embeddings generation methods---GloVe, word2vec, and fastText---with publicly available data and evaluated them with measures specifically appropriate for the hierarchical context. FastText achieved an ${}_{LCA}F_1$ of 0.893 on a single-labeled version of the RCV1 dataset. An analysis indicates that using word embeddings and its flavors is a very promising approach for HTC.
Traffic Density Estimation using a Convolutional Neural Network
Nubert, Julian, Truong, Nicholas Giai, Lim, Abel, Tanujaya, Herbert Ilhan, Lim, Leah, Vu, Mai Anh
The goal of this project is to introduce and present a machine learning application that aims to improve the quality of life of people in Singapore. In particular, we investigate the use of machine learning solutions to tackle the problem of traffic congestion in Singapore. In layman's terms, we seek to make Singapore (or any other city) a smoother place. To accomplish this aim, we present an end-to-end system comprising of 1. A traffic density estimation algorithm at traffic lights/junctions and 2. a suitable traffic signal control algorithms that make use of the density information for better traffic control. Traffic density estimation can be obtained from traffic junction images using various machine learning techniques (combined with CV tools). After research into various advanced machine learning methods, we decided on convolutional neural networks (CNNs). We conducted experiments on our algorithms, using the publicly available traffic camera dataset published by the Land Transport Authority (LTA) to demonstrate the feasibility of this approach. With these traffic density estimates, different traffic algorithms can be applied to minimize congestion at traffic junctions in general.
Merging datasets through deep learning
Srinivas, Kavitha, Gale, Abraham, Dolby, Julian
Merging datasets is a key operation for data analytics. A frequent requirement for merging is joining across columns that have different surface forms for the same entity (e.g., the name of a person might be represented as "Douglas Adams" or "Adams, Douglas"). Similarly, ontology alignment can require recognizing distinct surface forms of the same entity, especially when ontologies are independently developed. However, data management systems are currently limited to performing merges based on string equality, or at best using string similarity. We propose an approach to performing merges based on deep learning models. Our approach depends on (a) creating a deep learning model that maps surface forms of an entity into a set of vectors such that alternate forms for the same entity are closest in vector space, (b) indexing these vectors using a nearest neighbors algorithm to find the forms that can be potentially joined together. To build these models, we had to adapt techniques from metric learning due to the characteristics of the data; specifically we describe novel sample selection techniques and loss functions that work for this problem. To evaluate our approach, we used Wikidata as ground truth and built models from datasets with approximately 1.1M people's names (200K identities) and 130K company names (70K identities). We developed models that allow for joins with precision@1 of .75-.81 and recall of .74-.81. We make the models available for aligning people or companies across multiple datasets.
Improbotics: Exploring the Imitation Game using Machine Intelligence in Improvised Theatre
Mathewson, Kory W., Mirowski, Piotr
Theatrical improvisation (impro or improv) is a demanding form of live, collaborative performance. Improv is a humorous and playful artform built on an open-ended narrative structure which simultaneously celebrates effort and failure. It is thus an ideal test bed for the development and deployment of interactive artificial intelligence (AI)-based conversational agents, or artificial improvisors. This case study introduces an improv show experiment featuring human actors and artificial improvisors. We have previously developed a deep-learning-based artificial improvisor, trained on movie subtitles, that can generate plausible, context-based, lines of dialogue suitable for theatre (Mathewson and Mirowski 2017). In this work, we have employed it to control what a subset of human actors say during an improv performance. We also give human-generated lines to a different subset of performers. All lines are provided to actors with headphones and all performers are wearing headphones. This paper describes a Turing test, or imitation game, taking place in a theatre, with both the audience members and the performers left to guess who is a human and who is a machine. In order to test scientific hypotheses about the perception of humans versus machines we collect anonymous feedback from volunteer performers and audience members. Our results suggest that rehearsal increases proficiency and possibility to control events in the performance. That said, consistency with real world experience is limited by the interface and the mechanisms used to perform the show. We also show that human-generated lines are shorter, more positive, and have less difficult words with more grammar and spelling mistakes than the artificial improvisor generated lines.
An Algebra of Lightweight Ontologies
Casanova, Marco A., Magalhães, Rômulo
This paper argues that certain ontology design problems are profitably addressed by treating ontologies as theories and by defining a set of operations that create new ontologies, including their constraints, out of other ontologies. The paper first shows how to use the operations in the context of ontology reuse, how to take advantage of the operations to compare different ontologies, or different versions of an ontology, and how the operations may help design mediated schemas in a bottom up fashion. The core of the paper discusses how to compute the operations for lightweight ontologies and addresses the question of minimizing the set of constraints of a lightweight ontology. Finally, the paper describes an implementation of the operations, as a Prot\'eg\'e plug-in.
Anytime Hedge achieves optimal regret in the stochastic regime
Mourtada, Jaouad, Gaïffas, Stéphane
This paper is about a surprising fact: we prove that the anytime Hedge algorithm with decreasing learning rate, which is one of the simplest algorithm for the problem of prediction with expert advice, is actually both worst-case optimal and adaptive to the easier stochastic and adversarial with a gap problems. This runs counter to the common belief in the literature that this algorithm is overly conservative, and that only new adaptive algorithms can simultaneously achieve minimax regret and adapt to the difficulty of the problem. Moreover, our analysis exhibits qualitative differences with other variants of the Hedge algorithm, based on the so-called "doubling trick", and the fixed-horizon version (with constant learning rate).
Deep Generative Model using Unregularized Score for Anomaly Detection with Heterogeneous Complexity
Matsubara, Takashi, Hama, Kenta, Tachibana, Ryosuke, Uehara, Kuniaki
Abstract--Accurate and automated detection of anomalous samples in a natural image dataset can be accomplished with a probabilistic model for end-to-end modeling of images. Such images have heterogeneous complexity, however, and a probabilistic model overlooks simply shaped objects with small anomalies. This is because the probabilistic model assigns undesirably lower likelihoods to complexly shaped objects that are nevertheless consistent with set standards. To overcome this difficulty, we propose an unregularized score for deep generative models (DGMs), which are generative models leveraging deep neural networks. We found that the regularization terms of the DGMs considerably influence the anomaly score depending on the complexity of the samples. By removing these terms, we obtain an unregularized score, which we evaluated on a toy dataset and real-world manufacturing datasets. Empirical results demonstrate that the unregularized score is robust to the inherent complexity of samples and can be used to better detect anomalies. Image-based anomaly detection has recently attracted considerable attention in the field of machine learning. This technique can be used to detect pedestrians behaving abnormally from surveillance video in order to prevent accidents [1], [2], or to detect lesions in medical images to provide early diagnosis [3]. In manufacturing plants, moreover, image-based anomaly detection can reject products not coincident with set standards.
Ontology Reasoning with Deep Neural Networks
Hohenecker, Patrick, Lukasiewicz, Thomas
The ability to conduct logical reasoning is a fundamental aspect of intelligent behavior, and thus an important problem along the way to human-level artificial intelligence. Traditionally, symbolic methods from the field of knowledge representation and reasoning have been used to equip agents with capabilities that resemble human logical reasoning qualities. More recently, however, there has been an increasing interest in using machine learning rather than logic-based formalisms to tackle these tasks. In this paper, we employ state-of-the-art methods for training deep neural networks to devise a novel model that is able to learn how to effectively perform basic ontology reasoning. This is an important and at the same time very natural reasoning problem, which is why the presented approach is applicable to a plethora of important real-world problems. We present the outcomes of several experiments, which show that our model learned to perform precise reasoning on diverse and challenging tasks. Furthermore, it turned out that the suggested approach suffers much less from different obstacles that prohibit symbolic reasoning, and, at the same time, is surprisingly plausible from a biological point of view.
Subspace Estimation from Incomplete Observations: A High-Dimensional Analysis
Wang, Chuang, Eldar, Yonina C., Lu, Yue M.
Abstract--We present a high-dimensional analysis of three popular algorithms, namely, Oja's method, GROUSE and PETRELS, for subspace estimation from streaming and highly incomplete observations. We show that, with proper time scaling, the time-varying principal angles between the true subspace and its estimates given by the algorithms converge weakly to deterministic processes when the ambient dimension n tends to infinity. Moreover, the limiting processes can be exactly characterized as the unique solutions of certain ordinary differential equations (ODEs). A finite sample bound is also given, showing that the rate of convergence towards such limits is O(1/ n). In addition to providing asymptotically exact predictions of the dynamic performance of the algorithms, our high-dimensional analysis yields several insights, including an asymptotic equivalence between Oja's method and GROUSE, and a precise scaling relationship linking the amount of missing data to the signalto-noise ratio. By analyzing the solutions of the limiting ODEs, we also establish phase transition phenomena associated with the steady-state performance of these techniques. Subspace estimation is a key task in many signal processing applications. Examples include source localization in array processing, system identification, network monitoring, and image sequence analysis, to name a few. The ubiquity of subspace estimation comes from the fact that a low-rank subspace model can conveniently capture the intrinsic, lowdimensional structures of many large datasets.
Stochastic Particle-Optimization Sampling and the Non-Asymptotic Convergence Theory
Zhang, Jianyi, Zhang, Ruiyi, Chen, Changyou
Particle-optimization sampling (POS) is a recently developed technique to generate high-quality samples from a target distribution by iteratively updating a set of interactive particles. A representative algorithm is the Stein variational gradient descent (SVGD). Though obtaining significant empirical success, the {\em non-asymptotic} convergence behavior of SVGD remains unknown. In this paper, we generalize POS to a stochasticity setting by injecting random noise in particle updates, called stochastic particle-optimization sampling (SPOS). Standard SVGD can be regarded as a special case of our framework. Notably, for the first time, we develop non-asymptotic convergence theory for the SPOS framework (which includes SVGD), characterizing the bias of a sample approximation w.r.t. the numbers of particles and iterations under both convex- and noncovex-energy-function settings. Remarkably, we provide theoretical understand of a pitfall of SVGD that can be avoided in the proposed SPOS framework, i.e., particles tent to collapse to a local mode in SVGD under some particular conditions. Our theory is based on the analysis of nonlinear stochastic differential equations, which serves as an extension and a complemented development to the asymptotic convergence theory for SVGD such as [1].