Genre
Efficiently Summarising Event Sequences with Rich Interleaving Patterns
Bhattacharyya, Apratim, Vreeken, Jilles
Discovering the key structure of a database is one of the main goals of data mining. In pattern set mining we do so by discovering a small set of patterns that together describe the data well. The richer the class of patterns we consider, and the more powerful our description language, the better we will be able to summarise the data. In this paper we propose \ourmethod, a novel greedy MDL-based method for summarising sequential data using rich patterns that are allowed to interleave. Experiments show \ourmethod is orders of magnitude faster than the state of the art, results in better models, as well as discovers meaningful semantics in the form patterns that identify multiple choices of values.
Network classification with applications to brain connectomics
Reliรณn, Jesรบs D. Arroyo, Kessler, Daniel, Levina, Elizaveta, Taylor, Stephan F.
While statistical analysis of a single network has received a lot of attention in recent years, with a focus on social networks, analysis of a sample of networks presents its own challenges which require a different set of analytic tools. Here we study the problem of classification of networks with labeled nodes, motivated by applications in neuroimaging. Brain networks are constructed from imaging data to represent functional connectivity between regions of the brain, and previous work has shown the potential of such networks to distinguish between various brain disorders, giving rise to a network (or graph) classification problem. Existing approaches to graph classification tend to either treat all edge weights as a long vector, ignoring the network structure, or focus on the graph topology while ignoring the edge weights. Our goal here is to design a graph classification method that uses both the individual edge information and the network structure of the data in a computationally efficient way. We are also interested in obtaining a parsimonious and interpretable representation of differences in brain connectivity patterns between classes, which requires variable selection. We propose a graph classification method that uses edge weights as variables but incorporates the network nature of the data via penalties that promotes sparsity in the number of nodes. We implement the method via efficient convex optimization algorithms and show good performance on data from two fMRI studies of schizophrenia.
Modelling Competitive Sports: Bradley-Terry-\'{E}l\H{o} Models for Supervised and On-Line Learning of Paired Competition Outcomes
Kirรกly, Franz J., Qian, Zhaozhi
Prediction and modelling of competitive sports outcomes has received much recent attention, especially from the Bayesian statistics and machine learning communities. In the real world setting of outcome prediction, the seminal \'{E}l\H{o} update still remains, after more than 50 years, a valuable baseline which is difficult to improve upon, though in its original form it is a heuristic and not a proper statistical "model". Mathematically, the \'{E}l\H{o} rating system is very closely related to the Bradley-Terry models, which are usually used in an explanatory fashion rather than in a predictive supervised or on-line learning setting. Exploiting this close link between these two model classes and some newly observed similarities, we propose a new supervised learning framework with close similarities to logistic regression, low-rank matrix completion and neural networks. Building on it, we formulate a class of structured log-odds models, unifying the desirable properties found in the above: supervised probabilistic prediction of scores and wins/draws/losses, batch/epoch and on-line learning, as well as the possibility to incorporate features in the prediction, without having to sacrifice simplicity, parsimony of the Bradley-Terry models, or computational efficiency of \'{E}l\H{o}'s original approach. We validate the structured log-odds modelling approach in synthetic experiments and English Premier League outcomes, where the added expressivity yields the best predictions reported in the state-of-art, close to the quality of contemporary betting odds.
Distributed Sequence Memory of Multidimensional Inputs in Recurrent Networks
Charles, Adam, Yin, Dong, Rozell, Christopher
Recurrent neural networks (RNNs) have drawn interest from machine learning researchers because of their effectiveness at preserving past inputs for time-varying data processing tasks. To understand the success and limitations of RNNs, it is critical that we advance our analysis of their fundamental memory properties. We focus on echo state networks (ESNs), which are RNNs with simple memoryless nodes and random connectivity. In most existing analyses, the short-term memory (STM) capacity results conclude that the ESN network size must scale linearly with the input size for unstructured inputs. The main contribution of this paper is to provide general results characterizing the STM capacity for linear ESNs with multidimensional input streams when the inputs have common low-dimensional structure: sparsity in a basis or significant statistical dependence between inputs. In both cases, we show that the number of nodes in the network must scale linearly with the information rate and poly-logarithmically with the ambient input dimension. The analysis relies on advanced applications of random matrix theory and results in explicit non-asymptotic bounds on the recovery error. Taken together, this analysis provides a significant step forward in our understanding of the STM properties in RNNs.
Statistical power and prediction accuracy in multisite resting-state fMRI connectivity
Dansereau, Christian, Benhajali, Yassine, Risterucci, Celine, Pich, Emilio Merlo, Orban, Pierre, Arnold, Douglas, Bellec, Pierre
Connectivity studies using resting-state functional magnetic resonance imaging are increasingly pooling data acquired at multiple sites. While this may allow investigators to speed up recruitment or increase sample size, multisite studies also potentially introduce systematic biases in connectivity measures across sites. In this work, we measure the inter-site effect in connectivity and its impact on our ability to detect individual and group differences. Our study was based on real, as opposed to simulated, multisite fMRI datasets collected in N=345 young, healthy subjects across 8 scanning sites with 3T scanners and heterogeneous scanning protocols, drawn from the 1000 functional connectome project. We first empirically show that typical functional networks were reliably found at the group level in all sites, and that the amplitude of the inter-site effects was small to moderate, with a Cohen's effect size below 0.5 on average across brain connections. We then implemented a series of Monte-Carlo simulations, based on real data, to evaluate the impact of the multisite effects on detection power in statistical tests comparing two groups (with and without the effect) using a general linear model, as well as on the prediction of group labels with a support-vector machine. As a reference, we also implemented the same simulations with fMRI data collected at a single site using an identical sample size. Simulations revealed that using data from heterogeneous sites only slightly decreased our ability to detect changes compared to a monosite study with the GLM, and had a greater impact on prediction accuracy. Taken together, our results support the feasibility of multisite studies in rs-fMRI provided the sample size is large enough.
An Online Convex Optimization Approach to Dynamic Network Resource Allocation
Chen, Tianyi, Ling, Qing, Giannakis, Georgios B.
Existing approaches to online convex optimization (OCO) make sequential one-slot-ahead decisions, which lead to (possibly adversarial) losses that drive subsequent decision iterates. Their performance is evaluated by the so-called regret that measures the difference of losses between the online solution and the best yet fixed overall solution in hindsight. The present paper deals with online convex optimization involving adversarial loss functions and adversarial constraints, where the constraints are revealed after making decisions, and can be tolerable to instantaneous violations but must be satisfied in the long term. Performance of an online algorithm in this setting is assessed by: i) the difference of its losses relative to the best dynamic solution with one-slot-ahead information of the loss function and the constraint (that is here termed dynamic regret); and, ii) the accumulated amount of constraint violations (that is here termed dynamic fit). In this context, a modified online saddle-point (MOSP) scheme is developed, and proved to simultaneously yield sub-linear dynamic regret and fit, provided that the accumulated variations of per-slot minimizers and constraints are sub-linearly growing with time. MOSP is also applied to the dynamic network resource allocation task, and it is compared with the well-known stochastic dual gradient method. Under various scenarios, numerical experiments demonstrate the performance gain of MOSP relative to the state-of-the-art.
Artificial Intelligence Fact Sheet - Content Science Review
Content Science is a content strategy and intelligence firm based in Atlanta, GA. Founded in 2010 by Colleen Jones, author of Clout: The Art Science of Influential Web Content, our mission is to transform industries, organizations, and individuals for the better by putting content first. We offer professional services, publications, and software for clients ranging from Fortune 50 companies to nonprofits to government agencies.
Alphabet misses on earnings, tops sales forecast
If you care about money, the economy and corporate America, earnings season matters. Google parent Alphabet reported fourth-quarter earnings on Thursday. SAN FRANCISCO -- Google parent Alphabet topped analyst estimates with ongoing strength in mobile search and video advertising, but missed on earnings, sending shares down after hours. Alphabet reported fourth-quarter revenue of $26 billion, up 22% year over year, led by YouTube and mobile search, chief financial officer Ruth Porat said in a statement. But Alphabet missed on the bottom line, with net income of $5.33 billion, or $7.56 a share.
Labor pick Puzder outsourced jobs, favored robots over workers, now vows to be 'best champion' of jobs
WASHINGTON โ President Donald Trump's choice for labor secretary is CEO of a fast food empire that is outsourcing jobs, a stark contrast with Trump's scathing attacks on companies that send jobs overseas. A filing with the Department of Labor and Trump's criticism of outsourcing could be raised at Andrew Puzder's confirmation hearing, with Democrats questioning how well he can advocate for workers. Puzder's company, CKE Restaurants Inc., notified the government in August of 2010 that it was outsourcing its restaurant information technology division to the Philippines. Doing so, the agency found, "contributed importantly" to the layoffs of both CKE employees and those of an outside staffing firm at an Anaheim, California, facility. The agency's finding made workers eligible for federally funded benefits meant to dampen the impact of globalization on employees. "By outsourcing the function to a firm that employs hundreds of Help Desk specialists, CKE was able to improve the quality of service levels to their restaurants," the company said in a statement Wednesday to The Associated Press.
Flipboard on Flipboard
Even though the phrase "image recognition technologies" conjures visions of high-tech surveillance, these tools may soon be used in medicine more than in spycraft. A team of Stanford researchers trained a computer to identify images of skin cancer moles and lesions as accurately as a dermatologist, according to a new paper published in the journal Nature. In the future, this new research suggests, a simple cell phone app may help patients diagnose a skin cancer -- the most common of all cancers in the United States -- for themselves. "Our objective is to bring the expertise of top-level dermatologists to places where the dermatologist is not available," said Sebastian Thrun, senior author of the new study, founder of research and development lab Google X and an adjunct professor at Stanford University. He added that those who live in developing countries do not have the same level of care as can be found in the US and other industrialized nations.