Goto

Collaborating Authors

 Genre


Artificial Intelligence Integration Allows Publishers a First Look at Meta Bibliometric Intelligence - Aries Systems Corporation

#artificialintelligence

Frankfurt, Germany October 17, 2016 – Aries Systems Corporation today announces the integration of Meta Bibliometric Intelligence into Editorial Manager, Aries' industry-leading manuscript and peer-review tracking system for scholarly publications. The new technology, created by Meta, applies artificial intelligence toward the identification of high-impact manuscripts at the moment of first submission, allowing editors to triage and rank incoming manuscripts. "By incorporating Meta Bibliometric Intelligence into Editorial Manager, we're the first workflow system that helps publishers explore the potential of big data analysis during peer review," said Richard Wynne, Vice President of Sales and Marketing at Aries Systems. Bibliometric Intelligence uses sophisticated machine learning algorithms that were trained using Meta's corpus of millions of full-text articles – a collection that now comprises the largest scholarly text-mining collection on Earth. As newly-submitted manuscripts are processed, hundreds of unique features are pulled from the papers and fed through the algorithms.


IBM: In 5 years, Watson A.I. will be behind your every decision

#artificialintelligence

In the next five years, every important decision, whether it's business or personal, will be made with the assistance of IBM Watson. Watson, the company's artificial intelligence-fueled system, is working in fields like health care, finance, entertainment and retail, connecting businesses more easily with their customers, making sense of big data and helping doctors find treatments for cancer patients. The Watson system is set to transform how businesses function and how people live their lives. "Our goal is augmenting intelligence," Rometty said. "It is man and machine. This is all about extending your expertise. It doesn't matter what you do. IBM's conference this week, which the company said drew 17,000 attendees, explored how companies, including retailers, educators, human resources departments and financial institutions, amon others, can use Watson. "The challenge IBM has right now is to define the marketplace," said Jeff Kagan, an independent industry analyst, who attended the conference. "Ten years from now, will IBM be the leader?


Mobile, video pump up profit at Google parent Alphabet

#artificialintelligence

Google parent Alphabet (Xetra: ABEA.DE - news) delivered higher profits for the third quarter, lifted by gains in mobile and video advertising as the tech giant narrowed losses on its "moon shots." Net (LSE: 0LN0.L - news) profit climbed 27 percent to $5.1 billion. Revenue rose to $22.5 billion from $18.7 billion in the same period a year earlier. Shares (Berlin: DI6.BE - news) rose nearly one percent in after-market trades that followed the release of the stronger-than-expected earnings figures. "We had a great third quarter," Alphabet chief financial officer Ruth Porat said in the earnings release.


The p-filter: multi-layer FDR control for grouped hypotheses

arXiv.org Machine Learning

In many practical applications of multiple hypothesis testing using the False Discovery Rate (FDR), the given hypotheses can be naturally partitioned into groups, and one may not only want to control the number of false discoveries (wrongly rejected null hypotheses), but also the number of falsely discovered groups of hypotheses (we say a group is falsely discovered if at least one hypothesis within that group is rejected, when in reality the group contains only nulls). In this paper, we introduce the p-filter, a procedure which unifies and generalizes the standard FDR procedure by Benjamini and Hochberg and global null testing procedure by Simes. We first prove that our proposed method can simultaneously control the overall FDR at the finest level (individual hypotheses treated separately) and the group FDR at coarser levels (when such groups are user-specified). We then generalize the p-filter procedure even further to handle multiple partitions of hypotheses, since that might be natural in many applications. For example, in neuroscience experiments, we may have a hypothesis for every (discretized) location in the brain, and at every (discretized) timepoint: does the stimulus correlate with activity in location x at time t after the stimulus was presented? In this setting, one might want to group hypotheses by location and by time. Importantly, our procedure can handle multiple partitions which are nonhierarchical (i.e. one partition may arrange p-values by voxel, and another partition arranges them by time point; neither one is nested inside the other). We prove that our procedure controls FDR simultaneously across these multiple lay- ers, under assumptions that are standard in the literature: we do not need the hypotheses to be independent, but require a nonnegative dependence condition known as PRDS.


A scalable end-to-end Gaussian process adapter for irregularly sampled time series classification

arXiv.org Machine Learning

We present a general framework for classification of sparse and irregularly-sampled time series. The properties of such time series can result in substantial uncertainty about the values of the underlying temporal processes, while making the data difficult to deal with using standard classification methods that assume fixed-dimensional feature spaces. To address these challenges, we propose an uncertainty-aware classification framework based on a special computational layer we refer to as the Gaussian process adapter that can connect irregularly sampled time series data to any black-box classifier learnable using gradient descent. We show how to scale up the required computations based on combining the structured kernel interpolation framework and the Lanczos approximation method, and how to discriminatively train the Gaussian process adapter in combination with a number of classifiers end-to-end using backpropagation.


Algorithms for Fitting the Constrained Lasso

arXiv.org Machine Learning

We compare alternative computing strategies for solving the constrained lasso problem. As its name suggests, the constrained lasso extends the widely-used lasso to handle linear constraints, which allow the user to incorporate prior information into the model. In addition to quadratic programming, we employ the alternating direction method of multipliers (ADMM) and also derive an efficient solution path algorithm. Through both simulations and real data examples, we compare the different algorithms and provide practical recommendations in terms of efficiency and accuracy for various sizes of data. We also show that, for an arbitrary penalty matrix, the generalized lasso can be transformed to a constrained lasso, while the converse is not true. Thus, our methods can also be used for estimating a generalized lasso, which has wide-ranging applications. Code for implementing the algorithms is freely available in the Matlab toolbox SparseReg.


On the Latent Variable Interpretation in Sum-Product Networks

arXiv.org Artificial Intelligence

One of the central themes in Sum-Product networks (SPNs) is the interpretation of sum nodes as marginalized latent variables (LVs). This interpretation yields an increased syntactic or semantic structure, allows the application of the EM algorithm and to efficiently perform MPE inference. In literature, the LV interpretation was justified by explicitly introducing the indicator variables corresponding to the LVs' states. However, as pointed out in this paper, this approach is in conflict with the completeness condition in SPNs and does not fully specify the probabilistic model. We propose a remedy for this problem by modifying the original approach for introducing the LVs, which we call SPN augmentation. We discuss conditional independencies in augmented SPNs, formally establish the probabilistic interpretation of the sum-weights and give an interpretation of augmented SPNs as Bayesian networks. Based on these results, we find a sound derivation of the EM algorithm for SPNs. Furthermore, the Viterbi-style algorithm for MPE proposed in literature was never proven to be correct. We show that this is indeed a correct algorithm, when applied to selective SPNs, and in particular when applied to augmented SPNs. Our theoretical results are confirmed in experiments on synthetic data and 103 real-world datasets.


Dynamic matrix recovery from incomplete observations under an exact low-rank constraint

arXiv.org Machine Learning

Low-rank matrix factorizations arise in a wide variety of applications -- including recommendation systems, topic models, and source separation, to name just a few. In these and many other applications, it has been widely noted that by incorporating temporal information and allowing for the possibility of time-varying models, significant improvements are possible in practice. However, despite the reported superior empirical performance of these dynamic models over their static counterparts, there is limited theoretical justification for introducing these more complex models. In this paper we aim to address this gap by studying the problem of recovering a dynamically evolving low-rank matrix from incomplete observations. First, we propose the locally weighted matrix smoothing (LOWEMS) framework as one possible approach to dynamic matrix recovery. We then establish error bounds for LOWEMS in both the {\em matrix sensing} and {\em matrix completion} observation models. Our results quantify the potential benefits of exploiting dynamic constraints both in terms of recovery accuracy and sample complexity. To illustrate these benefits we provide both synthetic and real-world experimental results.


Globally Optimal Training of Generalized Polynomial Neural Networks with Nonlinear Spectral Methods

arXiv.org Machine Learning

The optimization problem behind neural networks is highly non-convex. Training with stochastic gradient descent and variants requires careful parameter tuning and provides no guarantee to achieve the global optimum. In contrast we show under quite weak assumptions on the data that a particular class of feedforward neural networks can be trained globally optimal with a linear convergence rate with our nonlinear spectral method. Up to our knowledge this is the first practically feasible method which achieves such a guarantee. While the method can in principle be applied to deep networks, we restrict ourselves for simplicity in this paper to one and two hidden layer networks. Our experiments confirm that these models are rich enough to achieve good performance on a series of real-world datasets.


Toward Implicit Sample Noise Modeling: Deviation-driven Matrix Factorization

arXiv.org Machine Learning

The objective function of a matrix factorization model usually aims to minimize the average of a regression error contributed by each element. However, given the existence of stochastic noises, the implicit deviations of sample data from their true values are almost surely diverse, which makes each data point not equally suitable for fitting a model. In this case, simply averaging the cost among data in the objective function is not ideal. Intuitively we would like to emphasize more on the reliable instances (i.e., those contain smaller noise) while training a model. Motivated by such observation, we derive our formula from a theoretical framework for optimal weighting under heteroscedastic noise distribution. Specifically, by modeling and learning the deviation of data, we design a novel matrix factorization model. Our model has two advantages. First, it jointly learns the deviation and conducts dynamic reweighting of instances, allowing the model to converge to a better solution. Second, during learning the deviated instances are assigned lower weights, which leads to faster convergence since the model does not need to overfit the noise. The experiments are conducted in clean recommendation and noisy sensor datasets to test the effectiveness of the model in various scenarios. The results show that our model outperforms the state-of-the-art factorization and deep learning models in both accuracy and efficiency.