Genre
Persistence Flamelets: multiscale Persistent Homology for kernel density exploration
Padellini, Tullia, Brutti, Pierpaolo
In recent years there has been noticeable interest in the study of the "shape of data" [2]. Among the many ways a "shape" could be defined, topology is the most general one, as it describes an object in terms of its connectivity structure: connected components (topological features of dimension 0), cycles (features of dimension 1) and so on. There is a growing number of techniques, generally denoted as Topological Data Analysis or TDA for short, aimed at estimating topological invariants of a fixed object; when we allow this object to change, however, little has been done to investigate the evolution in its topology. In this work we define the Persistence Flamelets, a multiscale version of one of the most popular tool in TDA, the Persistence Landscape. We examine its theoretical properties and we show how it could be used to gain insights on KDEs bandwidth parameter.
Inter-Subject Analysis: Inferring Sparse Interactions with Dense Intra-Graphs
Ma, Cong, Lu, Junwei, Liu, Han
We develop a new modeling framework for Inter-Subject Analysis (ISA). The goal of ISA is to explore the dependency structure between different subjects with the intra-subject dependency as nuisance. It has important applications in neuroscience to explore the functional connectivity between brain regions under natural stimuli. Our framework is based on the Gaussian graphical models, under which ISA can be converted to the problem of estimation and inference of the inter-subject precision matrix. The main statistical challenge is that we do not impose sparsity constraint on the whole precision matrix and we only assume the inter-subject part is sparse. For estimation, we propose to estimate an alternative parameter to get around the non-sparse issue and it can achieve asymptotic consistency even if the intra-subject dependency is dense. For inference, we propose an "untangle and chord" procedure to de-bias our estimator. It is valid without the sparsity assumption on the inverse Hessian of the log-likelihood function. This inferential method is general and can be applied to many other statistical problems, thus it is of independent theoretical interest. Numerical experiments on both simulated and brain imaging data validate our methods and theory.
Text Compression for Sentiment Analysis via Evolutionary Algorithms
Dufourq, Emmanuel, Bassett, Bruce A.
Can textual data be compressed intelligently without losing accuracy in evaluating sentiment? In this study, we propose a novel evolutionary compression algorithm, PARSEC (PARts-of-Speech for sEntiment Compression), which makes use of Parts-of-Speech tags to compress text in a way that sacrifices minimal classification accuracy when used in conjunction with sentiment analysis algorithms. An analysis of PARSEC with eight commercial and non-commercial sentiment analysis algorithms on twelve English sentiment data sets reveals that accurate compression is possible with (0%, 1.3%, 3.3%) loss in sentiment classification accuracy for (20%, 50%, 75%) data compression with PARSEC using LingPipe, the most accurate of the sentiment algorithms. Other sentiment analysis algorithms are more severely affected by compression. We conclude that significant compression of text data is possible for sentiment analysis depending on the accuracy demands of the specific application and the specific sentiment analysis algorithm used.
An Expectation Conditional Maximization approach for Gaussian graphical models
Li, Zehang, McCormick, Tyler H.
Bayesian graphical models are a useful tool for understanding dependence relationships among many variables, particularly in situations with external prior information. In high-dimensional settings, the space of possible graphs becomes enormous, rendering even state-of-the-art Bayesian stochastic search computationally infeasible. We propose a deterministic alternative to estimate Gaussian and Gaussian copula graphical models using an Expectation Conditional Maximization (ECM) algorithm, extending the EM approach from Bayesian variable selection to graphical model estimation. We show that the ECM approach enables fast posterior exploration under a sequence of mixture priors, and can incorporate multiple sources of information.
A minimax and asymptotically optimal algorithm for stochastic bandits
Ménard, Pierre, Garivier, Aurélien
We propose the kl-UCB ++ algorithm for regret minimization in stochastic bandit models with exponential families of distributions. We prove that it is simultaneously asymptotically optimal (in the sense of Lai and Robbins' lower bound) and minimax optimal. This is the first algorithm proved to enjoy these two properties at the same time. This work thus merges two different lines of research with simple and clear proofs.
Modeling sequences and temporal networks with dynamic community structures
Peixoto, Tiago P., Rosvall, Martin
In evolving complex systems such as air traffic and social organizations, collective effects emerge from their many components' dynamic interactions. While the dynamic interactions can be represented by temporal networks with nodes and links that change over time, they remain highly complex. It is therefore often necessary to use methods that extract the temporal networks' large-scale dynamic community structure. However, such methods are subject to overfitting or suffer from effects of arbitrary, a priori imposed timescales, which should instead be extracted from data. Here we simultaneously address both problems and develop a principled data-driven method that determines relevant timescales and identifies patterns of dynamics that take place on networks as well as shape the networks themselves. We base our method on an arbitrary-order Markov chain model with community structure, and develop a nonparametric Bayesian inference framework that identifies the simplest such model that can explain temporal interaction data.
On the Design of LQR Kernels for Efficient Controller Learning
Marco, Alonso, Hennig, Philipp, Schaal, Stefan, Trimpe, Sebastian
A core problem of learning control is to determine optimal feedback controllers for (partially) unknown nonlinear systems from experimental data. Reinforcement learning (RL) [1], [2] is a promising framework for this, yet often requires performing many experiments on the physical system to even find suitable controllers, which limits the applicability of such techniques. Therefore, a lot of research effort has been invested into data efficiency of RL aiming at learning controllers from as few experiments as possible. Recently, Bayesian optimization (BO) has been proposed for RL as a promising approach in this direction. BO employs a probabilistic description of the latent objective function (typically a Gaussian process (GP)), which allows for selecting next control experiments in a principled manner, e.g., to maximize information gain [3] or perform safe exploration [4]. While BO provides a promising framework for learning controllers in fairly general settings, the full power of Bayesian learning is often not exploited. A key advantage of Bayesian methods is that they allow for combining prior problem knowledge with learning from data in a principled manner. In case of GP models, this concerns specifically the choice of the kernel, which captures the covariance between function values at different inputs and is thus the core component to model prior knowledge about the function shape. By choosing standard kernels, however, naive BO approaches do often not exploit this opportunity to improve learning performance.
Anthropic decision theory
This paper sets out to resolve how agents ought to act in the Sleeping Beauty problem and various related anthropic (self-locating belief) problems, not through the calculation of anthropic probabilities, but through finding the correct decision to make. It creates an anthropic decision theory (ADT) that decides these problems from a small set of principles. By doing so, it demonstrates that the attitude of agents with regards to each other (selfish or altruistic) changes the decisions they reach, and that it is very important to take this into account. To illustrate ADT, it is then applied to two major anthropic problems and paradoxes, the Presumptuous Philosopher and Doomsday problems, thus resolving some issues about the probability of human extinction.
Google's Pixel 2 revealed ahead of unveiling next month
Google's Pixel 2 handsets have been leaked ahead on its unveiling on October 4th, it has been claimed. Droid Life claims to have images of the new handset, alongside a new mini wireless speaker, and images of an updated Daydream VR headset. Two versions of the handset will be launched, a Pixel and a Pixel XL. The LG-made Pixel 2 XL will come in a'Black & White' and'Just Black' colors and be available with 64GB or 128GB storage, according to the latest leaks The smaller Pixel 2 will be available in Kinda Blue, Just Black, and Clearly White, and will be sold with 64GB and 128GB of storage and priced at $649 and $749, respectively, according to the site. The larger LG-made Pixel 2 XL will come in a Black & White and Just Black colors and be available with 64GB or 128GB storage.
Automate This! Could autonomous robots put surgeons and pharmacists out of a job?
Welcome to the second instalment of'Automate This!,' a Day 6 series about the future of work in an artificially intelligent world. In 2011, Krista Jones was diagnosed with a rare form of cancer. The next five years were a blur of doctor's visits and operations. "I think I saw seven doctors over that time period," Jones recalls. I was heading towards a double mastectomy, mostly out of fear for the fact that nobody could explain why [the tumours] were reoccurring." Jones' final treatment plan was built using algorithms and big data -- some of the precursors to today's A.I. technology. That plan made it possible for Jones to forgo a painful double mastectomy, and ultimately left her cancer-free. "Not only did it save my life but it left me whole in so many different ways," she says. "[It] avoided some of the scars, emotionally and physically, that most people who go through cancer treatment are left with." The treatment plan that helped Krista Jones beat a rare form of cancer was developed using machine learning algorithms and big data. She's seen the downsides of machine learning technologies, too. Her own son was forced to rethink his plan to become a radiologist after watching his career prospects dwindle thanks to automation. Still, Jones is convinced that artificial intelligence is the future of health care. "I think what we need to do is harness the good while regulating the bad," she says, "such that we don't get hung up and stop the development of life-saving treatments." "The only next step is now replacing those actual physical physicists and doctors that actually say: 'Yes, this is the right treatment plan'.