Goto

Collaborating Authors

 Genre


Linear Algebra Mathematics MIT OpenCourseWare

#artificialintelligence

This course covers matrix theory and linear algebra, emphasizing topics useful in other disciplines such as physics, economics and social sciences, natural sciences, and engineering. It parallels the combination of theory and applications in Professor Strang's textbook Introduction to Linear Algebra. This course has been designed for independent study. It provides everything you will need to understand the concepts covered in the course.


Recommender System with Mahout and Elasticsearch

#artificialintelligence

This tutorial will describe how a surprisingly small amount of code can be used to build a recommendation engine using the MapR Sandbox for Hadoop with Apache Mahout and Elasticsearch. This tutorial will run on the MapR Sandbox. The tutorial also requires Elasticsearch and Mahout to be installed on the sandbox. Step 1: Indexing the movie meta data in Elasticsearch In Elasticsearch, documents contain fields which are, by default, all indexed. Typically documents are written as a single-level JSON structure.


The bag-of-frames approach: a not so sufficient model for urban soundscapes

arXiv.org Machine Learning

Further, recent psychoacoustical evidence suggest the approach bears some resemblance with human auditory processing for sound textures (McDermott et al., 2013; Nelken and de Cheveignรฉ, 2013). In an influential 2007 article, Aucouturier, Defreville & Pachet (Aucouturier et al., 2007) applied a BOF model to categorize both polyphonic music and soundscapes. Their results showed that, while BOF was a meriting model for their polyphonic music dataset, it was spectacularly effective for soundscapes, reaching accuracies of 96%. The contrast, they interpreted, lied in differences in the temporal structure of both types of stimuli, with music being more formally organized and soundscapes more easily summarized by statistics. In a later companion study (Aucouturier and Defreville, 2009), they showed that soundscapes could be time-shuffled without altering listeners' perception of their acoustic similarity, while music could not. While more work was needed for music, the authors therefore concluded that BOF was a sufficient model to approximate human perception for soundscapes, practically ruling out the need to recognize the local acoustic events in a texture in order to identify it.


Tuning-Free Heterogeneity Pursuit in Massive Networks

arXiv.org Machine Learning

Heterogeneity is often natural in many contemporary applications involving massive data. While posing new challenges to effective learning, it can play a crucial role in powering meaningful scientific discoveries through the understanding of important differences among subpopulations of interest. In this paper, we exploit multiple networks with Gaussian graphs to encode the connectivity patterns of a large number of features on the subpopulations. To uncover the heterogeneity of these structures across subpopulations, we suggest a new framework of tuning-free heterogeneity pursuit (THP) via large-scale inference, where the number of networks is allowed to diverge. In particular, two new tests, the chi-based test and the linear functional-based test, are introduced and their asymptotic null distributions are established. Under mild regularity conditions, we establish that both tests are optimal in achieving the testable region boundary and the sample size requirement for the latter test is minimal. Both theoretical guarantees and the tuning-free feature stem from efficient multiple-network estimation by our newly suggested approach of heterogeneous group square-root Lasso (HGSL) for high-dimensional multi-response regression with heterogeneous noises. To solve this convex program, we further introduce a tuning-free algorithm that is scalable and enjoys provable convergence to the global optimum. Both computational and theoretical advantages of our procedure are elucidated through simulation and real data examples.


Efficient KLMS and KRLS Algorithms: A Random Fourier Feature Perspective

arXiv.org Machine Learning

We present a new framework for online Least Squares algorithms for nonlinear modeling in RKH spaces (RKHS). Instead of implicitly mapping the data to a RKHS (e.g., kernel trick), we map the data to a finite dimensional Euclidean space, using random features of the kernel's Fourier transform. The advantage is that, the inner product of the mapped data approximates the kernel function. The resulting "linear" algorithm does not require any form of sparsification, since, in contrast to all existing algorithms, the solution's size remains fixed and does not increase with the iteration steps. As a result, the obtained algorithms are computationally significantly more efficient compared to previously derived variants, while, at the same time, they converge at similar speeds and to similar error floors.


Comparison of Several Sparse Recovery Methods for Low Rank Matrices with Random Samples

arXiv.org Machine Learning

In this paper, we will investigate the efficacy of IMAT (Iterative Method of Adaptive Thresholding) in recovering the sparse signal (parameters) for linear models with missing data. Sparse recovery rises in compressed sensing and machine learning problems and has various applications necessitating viable reconstruction methods specifically when we work with big data. This paper will focus on comparing the power of IMAT in reconstruction of the desired sparse signal with LASSO. Additionally, we will assume the model has random missing information. Missing data has been recently of interest in big data and machine learning problems since they appear in many cases including but not limited to medical imaging datasets, hospital datasets, and massive MIMO. The dominance of IMAT over the well-known LASSO will be taken into account in different scenarios. Simulations and numerical results are also provided to verify the arguments.


A note on the complexity of the causal ordering problem

arXiv.org Artificial Intelligence

In this note we provide a concise report on the complexity of the causal ordering problem, originally introduced by Simon to reason about causal dependencies implicit in systems of mathematical equations. We show that Simon's classical algorithm to infer causal ordering is NP-Hard---an intractability previously guessed but never proven. We present then a detailed account based on Nayak's suggested algorithmic solution (the best available), which is dominated by computing transitive closure---bounded in time by $O(|\mathcal V|\cdot |\mathcal S|)$, where $\mathcal S(\mathcal E, \mathcal V)$ is the input system structure composed of a set $\mathcal E$ of equations over a set $\mathcal V$ of variables with number of variable appearances (density) $|\mathcal S|$. We also comment on the potential of causal ordering for emerging applications in large-scale hypothesis management and analytics.


Design Patterns for Recommendation Systems โ€“ Everyone Wants a Pony

#artificialintelligence

Ted Dunning (Chief Application Architect at MapR) and Ellen Friedman have written a new O'Reilly Media book on "Practical Machine Learning โ€“ Innovations in Recommendation" (released in January 2014). This book examines one of the most interesting, fun, and powerful data science applications in the big data universe: recommendation systems. For me, this was one of the most interesting applications of data mining that immediately captured my imagination after I embarked on the journey to data science (drifting away from my astrophysics roots) about a dozen years ago. It is also one of the most common use cases that are taught in data science MOOCs and other analytics training courses. I believe that the love affair with recommender systems can be partly attributed to two things.


We asked a newswriting robot to write Marvin Minsky's obit, and it's pretty good

#artificialintelligence

The pioneering artificial intelligence theorist Marvin Minsky died Sunday. So we thought it would be appropriate to ask for an obituary from one of his virtual descendants: Wordsmith, the automated news-writing bot from the company Automated Insights. It takes structured data--stuff that fits into a spreadsheet--and fits it into templates of increasing complexity. But Minsky was always interested in the differences and similarities between human and machine cognition, and arguably Wordsmith is cogitating every bit as hard as human reporters do on deadline. "You can start with things like their name, their age, the day the died, how they died. You can imagine in a spreadsheet, 'significant accomplishments 1, 2, and 3.'" So in this case, the template is a one-off rather than fully automated, and the data was harder to scrape.


Will robots take over from doctors? - Telegraph

#artificialintelligence

Whether robots will replace doctors in the future is just one of the questions Professor Lilford will be addressing in a panel discussion titled The Picture of Health: Exploring the Future of Medicine, as part of the University of Warwick's Festival of the Imagination. Running from October 16-17, the festival marks the university's 50th anniversary, showcasing its work and expertise with a series of talks, live shows and interaction exhibitions on the theme of "Imagining the future". This particular event will take the form of a Q&A with the audience, focusing on what the world of medicine will look like 50 years from now. Thanks to Professor Lilford's 40 year career, in which he's done everything from practising as a doctor specialising in obstetrics and gynaecology to working for the Department of Health, he's well placed to tackle any questions that come up. It promises to be a fascinating peek into the wards and surgeries of the future.