Statistical Learning
Building Intelligent Autonomous Navigation Agents
Breakthroughs in machine learning in the last decade have led to `digital intelligence', i.e. machine learning models capable of learning from vast amounts of labeled data to perform several digital tasks such as speech recognition, face recognition, machine translation and so on. The goal of this thesis is to make progress towards designing algorithms capable of `physical intelligence', i.e. building intelligent autonomous navigation agents capable of learning to perform complex navigation tasks in the physical world involving visual perception, natural language understanding, reasoning, planning, and sequential decision making. Despite several advances in classical navigation methods in the last few decades, current navigation agents struggle at long-term semantic navigation tasks. In the first part of the thesis, we discuss our work on short-term navigation using end-to-end reinforcement learning to tackle challenges such as obstacle avoidance, semantic perception, language grounding, and reasoning. In the second part, we present a new class of navigation methods based on modular learning and structured explicit map representations, which leverage the strengths of both classical and end-to-end learning methods, to tackle long-term navigation tasks. We show that these methods are able to effectively tackle challenges such as localization, mapping, long-term planning, exploration and learning semantic priors. These modular learning methods are capable of long-term spatial and semantic understanding and achieve state-of-the-art results on various navigation tasks.
Data Scientist (Remote)
Finite Markov Decision Processes, Support Vector Machines, Q-Learning, Stochastic Finite State Machines, MCTS or other hybrid Deep Reinforcement Learning processes W2 Benefits Not only you get to join our team of awesome playful ninjas, we also have great benefits: Six weeks paid time off per year (PTO Holidays).
How You Can Get Started With Machine Learning In Marketing
While some companies are now becoming extremely sophisticated in handling such big data and combining it to better segment and market users, a lot are still catching up. Every now and then we all hear how Machine Learning is going to take over our mundane jobs and how AI is the future. But frankly today Machine Learning and Algorithms are not a story of the future, these are everywhere, from your google searches, to your Netflix suggestions. While on the onset you might never be able to recognize this hidden intelligence in the systems around you, but these systems are designed to give you such a seamless experience that it feels almost like "Magic". Machine learning is a subset of Artificial Intelligence, and we are only going to talk about only Machine Learning for now.
Unsupervised machine learning application in ACL
De-Sheng Chen,* Tong-Fu Wang,* Jia-Wang Zhu, Bo Zhu, Zeng-Liang Wang, Jian-Gang Cao, Cai-Hong Feng, Jun-Wei Zhao Department of Sports Medicine and Arthroscopy, Tianjin Hospital of Tianjin University, Tianjin, People's Republic of China *These authors contributed equally to this work Correspondence: Jia-Wang Zhu Department of Sports Medicine and Arthroscopy, Tianjin Hospital of Tianjin University, Tianjin, People's Republic of China Email [email protected] Purpose: We aim to present an unsupervised machine learning application in anterior cruciate ligament (ACL) rupture and evaluate whether supervised machine learning-derived radiomics features enable prediction of ACL rupture accurately. Patients and Methods: Sixty-eight patients were reviewed. Their demographic features were recorded, radiomics features were extracted, and the input dataset was defined as a collection of demographic features and radiomics features. The input dataset was automatically classified by the unsupervised machine learning algorithm. Then, we used a supervised machine learning algorithm to construct a radiomics model.
Constrained Classification and Policy Learning
Kitagawa, Toru, Sakaguchi, Shosei, Tetenov, Aleksey
Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empirical classification risk. These techniques are also useful for causal policy learning problems, since estimation of individualized treatment rules can be cast as a weighted (cost-sensitive) classification problem. Consistency of the surrogate loss approaches studied in Zhang (2004) and Bartlett et al. (2006) crucially relies on the assumption of correct specification, meaning that the specified set of classifiers is rich enough to contain a first-best classifier. This assumption is, however, less credible when the set of classifiers is constrained by interpretability or fairness, leaving the applicability of surrogate loss based algorithms unknown in such second-best scenarios. This paper studies consistency of surrogate loss procedures under a constrained set of classifiers without assuming correct specification. We show that in the setting where the constraint restricts the classifier's prediction set only, hinge losses (i.e., $\ell_1$-support vector machines) are the only surrogate losses that preserve consistency in second-best scenarios. If the constraint additionally restricts the functional form of the classifier, consistency of a surrogate loss approach is not guaranteed even with hinge loss. We therefore characterize conditions for the constrained set of classifiers that can guarantee consistency of hinge risk minimizing classifiers. Exploiting our theoretical results, we develop robust and computationally attractive hinge loss based procedures for a monotone classification problem.
Smart Healthcare in the Age of AI: Recent Advances, Challenges, and Future Prospects
Nasr, Mahmoud, Islam, MD. Milon, Shehata, Shady, Karray, Fakhri, Quintana, Yuri
The significant increase in the number of individuals with chronic ailments (including the elderly and disabled) has dictated an urgent need for an innovative model for healthcare systems. The evolved model will be more personalized and less reliant on traditional brick-and-mortar healthcare institutions such as hospitals, nursing homes, and long-term healthcare centers. The smart healthcare system is a topic of recently growing interest and has become increasingly required due to major developments in modern technologies, especially in artificial intelligence (AI) and machine learning (ML). This paper is aimed to discuss the current state-of-the-art smart healthcare systems highlighting major areas like wearable and smartphone devices for health monitoring, machine learning for disease diagnosis, and the assistive frameworks, including social robots developed for the ambient assisted living environment. Additionally, the paper demonstrates software integration architectures that are very significant to create smart healthcare systems, integrating seamlessly the benefit of data analytics and other tools of AI. The explained developed systems focus on several facets: the contribution of each developed framework, the detailed working procedure, the performance as outcomes, and the comparative merits and limitations. The current research challenges with potential future directions are addressed to highlight the drawbacks of existing systems and the possible methods to introduce novel frameworks, respectively. This review aims at providing comprehensive insights into the recent developments of smart healthcare systems to equip experts to contribute to the field.
On Locality of Local Explanation Models
Ghalebikesabi, Sahra, Ter-Minassian, Lucile, Diaz-Ordaz, Karla, Holmes, Chris
Shapley values provide model agnostic feature attributions for model outcome at a particular instance by simulating feature absence under a global population distribution. The use of a global population can lead to potentially misleading results when local model behaviour is of interest. Hence we consider the formulation of neighbourhood reference distributions that improve the local interpretability of Shapley values. By doing so, we find that the Nadaraya-Watson estimator, a well-studied kernel regressor, can be expressed as a self-normalised importance sampling estimator. Empirically, we observe that Neighbourhood Shapley values identify meaningful sparse feature relevance attributions that provide insight into local model behaviour, complimenting conventional Shapley analysis. They also increase on-manifold explainability and robustness to the construction of adversarial classifiers.
Bayesian Inference in High-Dimensional Time-Serieswith the Orthogonal Stochastic Linear Mixing Model
Meng, Rui, Bouchard, Kristofer
Many modern time-series datasets contain large numbers of output response variables sampled for prolonged periods of time. For example, in neuroscience, the activities of 100s-1000's of neurons are recorded during behaviors and in response to sensory stimuli. Multi-output Gaussian process models leverage the nonparametric nature of Gaussian processes to capture structure across multiple outputs. However, this class of models typically assumes that the correlations between the output response variables are invariant in the input space. Stochastic linear mixing models (SLMM) assume the mixture coefficients depend on input, making them more flexible and effective to capture complex output dependence. However, currently, the inference for SLMMs is intractable for large datasets, making them inapplicable to several modern time-series problems. In this paper, we propose a new regression framework, the orthogonal stochastic linear mixing model (OSLMM) that introduces an orthogonal constraint amongst the mixing coefficients. This constraint reduces the computational burden of inference while retaining the capability to handle complex output dependence. We provide Markov chain Monte Carlo inference procedures for both SLMM and OSLMM and demonstrate superior model scalability and reduced prediction error of OSLMM compared with state-of-the-art methods on several real-world applications. In neurophysiology recordings, we use the inferred latent functions for compact visualization of population responses to auditory stimuli, and demonstrate superior results compared to a competing method (GPFA). Together, these results demonstrate that OSLMM will be useful for the analysis of diverse, large-scale time-series datasets.
Decomposed Mutual Information Estimation for Contrastive Representation Learning
Sordoni, Alessandro, Dziri, Nouha, Schulz, Hannes, Gordon, Geoff, Bachman, Phil, Tachet, Remi
Recent contrastive representation learning methods rely on estimating mutual information (MI) between multiple views of an underlying context. E.g., we can derive multiple views of a given image by applying data augmentation, or we can split a sequence into views comprising the past and future of some step in the sequence. Contrastive lower bounds on MI are easy to optimize, but have a strong underestimation bias when estimating large amounts of MI. We propose decomposing the full MI estimation problem into a sum of smaller estimation problems by splitting one of the views into progressively more informed subviews and by applying the chain rule on MI between the decomposed views. This expression contains a sum of unconditional and conditional MI terms, each measuring modest chunks of the total MI, which facilitates approximation via contrastive bounds. To maximize the sum, we formulate a contrastive lower bound on the conditional MI which can be approximated efficiently. We refer to our general approach as Decomposed Estimation of Mutual Information (DEMI). We show that DEMI can capture a larger amount of MI than standard non-decomposed contrastive bounds in a synthetic setting, and learns better representations in a vision domain and for dialogue generation.