Genre
Forecasting of commercial sales with large scale Gaussian Processes
Rivera, Rodrigo, Burnaev, Evgeny
This paper argues that there has not been enough discussion in the field of applications of Gaussian Process for the fast moving consumer goods industry. Yet, this technique can be important as it e.g., can provide automatic feature relevance determination and the posterior mean can unlock insights on the data. Significant challenges are the large size and high dimensionality of commercial data at a point of sale. The study reviews approaches in the Gaussian Processes modeling for large data sets, evaluates their performance on commercial sales and shows value of this type of models as a decision-making tool for management.
Some variations on Random Survival Forest with application to Cancer Research
Dey, Arabin Kumar, Juneja, Anshul
Random survival forest can be extremely time consuming for large data set. In this paper we propose few computationally efficient algorithms in prediction of survival function. We explore the behavior of the algorithms for different cancer data sets. Our construction includes right censoring data too. We have also applied the same for competing risk survival function.
Subset Labeled LDA for Large-Scale Multi-Label Classification
Papanikolaou, Yannis, Tsoumakas, Grigorios
Labeled Latent Dirichlet Allocation (LLDA) is an extension of the standard unsupervised Latent Dirichlet Allocation (LDA) algorithm, to address multi-label learning tasks. Previous work has shown it to perform in par with other state-of-the-art multi-label methods. Nonetheless, with increasing label sets sizes LLDA encounters scalability issues. In this work, we introduce Subset LLDA, a simple variant of the standard LLDA algorithm, that not only can effectively scale up to problems with hundreds of thousands of labels but also improves over the LLDA state-of-the-art. We conduct extensive experiments on eight data sets, with label sets sizes ranging from hundreds to hundreds of thousands, comparing our proposed algorithm with the previously proposed LLDA algorithms (Prior--LDA, Dep--LDA), as well as the state of the art in extreme multi-label classification. The results show a steady advantage of our method over the other LLDA algorithms and competitive results compared to the extreme multi-label classification algorithms.
Latent Gaussian Process Regression
Bodin, Erik, Campbell, Neill D. F., Ek, Carl Henrik
We introduce Latent Gaussian Process Regression which is a latent variable extension allowing modelling of non-stationary multi-modal processes using GPs. The approach is built on extending the input space of a regression problem with a latent variable that is used to modulate the covariance function over the training data. We show how our approach can be used to model multi-modal and non-stationary processes. We exemplify the approach on a set of synthetic data and provide results on real data from motion capture and geostatistics.
Riemannian stochastic quasi-Newton algorithm with variance reduction and its convergence analysis
Kasai, Hiroyuki, Sato, Hiroyuki, Mishra, Bamdev
Stochastic variance reduction algorithms have recently become popular for minimizing the average of a large, but finite number of loss functions. The present paper proposes a Riemannian stochastic quasi-Newton algorithm with variance reduction (R-SQN-VR). The key challenges of averaging, adding, and subtracting multiple gradients are addressed with notions of retraction and vector transport. We present convergence analyses of R-SQN-VR on both non-convex and retraction-convex functions under retraction and vector transport operators. The proposed algorithm is evaluated on the Karcher mean computation on the symmetric positive-definite manifold and the low-rank matrix completion on the Grassmann manifold. In all cases, the proposed algorithm outperforms the state-of-the-art Riemannian batch and stochastic gradient algorithms.
Universality laws for randomized dimension reduction, with applications
Dimension reduction is the process of embedding high-dimensional data into a lower dimensional space to facilitate its analysis. In the Euclidean setting, one fundamental technique for dimension reduction is to apply a random linear map to the data. This dimension reduction procedure succeeds when it preserves certain geometric features of the set. The question is how large the embedding dimension must be to ensure that randomized dimension reduction succeeds with high probability. This paper studies a natural family of randomized dimension reduction maps and a large class of data sets. It proves that there is a phase transition in the success probability of the dimension reduction map as the embedding dimension increases. For a given data set, the location of the phase transition is the same for all maps in this family. Furthermore, each map has the same stability properties, as quantified through the restricted minimum singular value. These results can be viewed as new universality laws in high-dimensional stochastic geometry. Universality laws for randomized dimension reduction have many applications in applied mathematics, signal processing, and statistics. They yield design principles for numerical linear algebra algorithms, for compressed sensing measurement ensembles, and for random linear codes. Furthermore, these results have implications for the performance of statistical estimation methods under a large class of random experimental designs.
Statistical inference on random dot product graphs: a survey
Athreya, Avanti, Fishkind, Donniell E., Levin, Keith, Lyzinski, Vince, Park, Youngser, Qin, Yichen, Sussman, Daniel L., Tang, Minh, Vogelstein, Joshua T., Priebe, Carey E.
The random dot product graph (RDPG) is an independent-edge random graph that is analytically tractable and, simultaneously, either encompasses or can successfully approximate a wide range of random graphs, from relatively simple stochastic block models to complex latent position graphs. In this survey paper, we describe a comprehensive paradigm for statistical inference on random dot product graphs, a paradigm centered on spectral embeddings of adjacency and Laplacian matrices. We examine the analogues, in graph inference, of several canonical tenets of classical Euclidean inference: in particular, we summarize a body of existing results on the consistency and asymptotic normality of the adjacency and Laplacian spectral embeddings, and the role these spectral embeddings can play in the construction of single- and multi-sample hypothesis tests for graph data. We investigate several real-world applications, including community detection and classification in large social networks and the determination of functional and biologically relevant network properties from an exploratory data analysis of the Drosophila connectome. We outline requisite background and current open problems in spectral graph inference.
?utm_content=buffer81f0b&utm_medium=social&utm_source=twitter.com&utm_campaign=buffer
In the long run, we expect AI technologies to become very broadly embedded into enterprise software applications and software, much as business intelligence ("BI"), reporting and analytics features have increasingly been directly incorporated into enterprise applications in the last decade. Company related disclosures: Issuer Company Ticker Applicable Disclosures Constellation Software Inc. CSU-T 7, 9 Shopify Inc. SHOP-N 7, 9 OpenText Inc. OTEX-Q 7, 9 Kinaxis Inc. KXS-T 7, 9 Descartes Systems Group Inc. DSGX-Q 7, 9 Absolute Software Inc. ABT-T 7, 9 BSM Technologies Inc. GPS-T 7, 9 Symbility Solutions Inc. SY-V 7, 9 ProntoForms Corp. PFM-V 1, 3, 7, 9 Redline Communications Inc. RDL-T 7, 9 See legend of Disclosures on next page. Definitions "Research Analyst" means any partner, director, officer, employee or agent of iA Securities who is held out to the public as a research analyst or whose responsibilities to iA Securities include the preparation of any written report for distribution to clients or prospective clients of iA Securities which includes a recommendation with respect to a security. Technology Sector Blair Abernethy, CFA August 17, 2017 Page 27 Analyst's Certification Each iA Securities research analyst whose name appears on the front page of this research report hereby certifies that (i) the recommendations and opinions expressed in the research report accurately reflect the research analyst's personal views about the issuer and securities that are the subject of this report and all other companies and securities mentioned in this report that are covered by such research analyst and (ii) no part of the research analyst's compensation was, is, or will be directly or indirectly, related to the specific recommendations or views expressed by such research analyst in this report.
Artificial intelligence could soon revolutionize the way doctors treat cancer
Cancer is a tricky disease to treat. With more than one hundred known types, each responding differently to treatment depending on the person they're growing inside (and dozens of other factors), oncologists certainly have their work cut out for them. Machine learning could soon make their jobs a bit easier. IBM's supercomputer Watson has been applying machine learning to personalized cancer treatment for some time. After pouring over 600,000 medical reports and 1.5 million anonymized patient records and clinical trials, the data should help define the clearest path forward for doctors.
Artificial intelligence just made guessing your password a whole lot easier
A new tool in deep learning renders passwords less secure. Last week, the credit reporting agency Equifax announced that malicious hackers had leaked the personal information of 143 million people in their system. That's reason for concern, of course, but if a hacker wants to access your online data by simply guessing your password, you're probably toast in less than an hour. Now, there's more bad news: Scientists have harnessed the power of artificial intelligence (AI) to create a program that, combined with existing tools, figured more than a quarter of the passwords from a set of more than 43 million LinkedIn profiles. Yet the researchers say the technology may also be used to beat baddies at their own game.