Statistical Learning
Lifelong Spectral Clustering
Sun, Gan, Cong, Yang, Wang, Qianqian, Li, Jun, Fu, Yun
In the past decades, spectral clustering (SC) has become one of the most effective clustering algorithms. However, most previous studies focus on spectral clustering tasks with a fixed task set, which cannot incorporate with a new spectral clustering task without accessing to previously learned tasks. In this paper, we aim to explore the problem of spectral clustering in a lifelong machine learning framework, i.e., Lifelong Spectral Clustering (L2SC). Its goal is to efficiently learn a model for a new spectral clustering task by selectively transferring previously accumulated experience from knowledge library. Specifically, the knowledge library of L2SC contains two components: 1) orthogonal basis library: capturing latent cluster centers among the clusters in each pair of tasks; 2) feature embedding library: embedding the feature manifold information shared among multiple related tasks. As a new spectral clustering task arrives, L2SC firstly transfers knowledge from both basis library and feature library to obtain encoding matrix, and further redefines the library base over time to maximize performance across all the clustering tasks. Meanwhile, a general online update formulation is derived to alternatively update the basis library and feature library. Finally, the empirical experiments on several real-world benchmark datasets demonstrate that our L2SC model can effectively improve the clustering performance when comparing with other state-of-the-art spectral clustering algorithms.
Cryptocurrency Price Prediction and Trading Strategies Using Support Vector Machines
Zhao, David, Rinaldo, Alessandro, Brookins, Christopher
Few assets in financial history have been as notoriously volatile as cryptocurrencies. While the long term outlook for this asset class remains unclear, we are successful in making short term price predictions for several major crypto assets. Using historical data from July 2015 to November 2019, we develop a large number of technical indicators to capture patterns in the cryptocurrency market. We then test various classification methods to forecast short-term future price movements based on these indicators. On both PPV and NPV metrics, our classifiers do well in identifying up and down market moves over the next 1 hour. Beyond evaluating classification accuracy, we also develop a strategy for translating 1-hour-ahead class predictions into trading decisions, along with a backtester that simulates trading in a realistic environment. We find that support vector machines yield the most profitable trading strategies, which outperform the market on average for Bitcoin, Ethereum and Litecoin over the past 22 months, since January 2018.
Shifted Randomized Singular Value Decomposition
Among the typical applications of SVD are the low-rank matrix approximation and principal component analysis (PCA) of data matrices (Jolliffe, 2002). Using SVD to accurately estimate a low-rank factorization or the principal components of a data matrix, a mean-centering step should be carried out before performing SVD on the matrix. Despite its simplicity, the mean-centering can be very costly if the data matrix is large and sparse. This cost is because the mean subtraction of a sparse matrix turns it to a dense matrix which requires a considerable amount of memory and CPU time to be analyzed. This motivates us to extend the randomized SVD algorithm introduced by (Halko et al., 2011) to estimate the singular value decomposition of a mean-centered matrix without explicitly forming the matrix in the memory. More generally, we introduce a shifted randomized SVD algorithm that provides for the SVD estimation of a data matrix shifted by any vector in the ali.basirat@lingfil.uu.se 1 arXiv:1911.11772v2
Learning stable and predictive structures in kinetic systems: Benefits of a causal approach
Pfister, Niklas, Bauer, Stefan, Peters, Jonas
Learning kinetic systems from data is one of the core challenges in many fields. Identifying stable models is essential for the generalization capabilities of data-driven inference. We introduce a computationally efficient framework, called CausalKinetiX, that identifies structure from discrete time, noisy observations, generated from heterogeneous experiments. The algorithm assumes the existence of an underlying, invariant kinetic model, a key criterion for reproducible research. Results on both simulated and real-world examples suggest that learning the structure of kinetic systems benefits from a causal perspective. The identified variables and models allow for a concise description of the dynamics across multiple experimental settings and can be used for prediction in unseen experiments. We observe significant improvements compared to well established approaches focusing solely on predictive performance, especially for out-of-sample generalization.
Machine Learning -- Don't Just Rely on Your University
Incorporating machine learning into predictive analytics has been in high demand that provides businesses the competitive edge. This hot topic is highly subscribed by undergraduates all over the world. However, being formally introduced the concepts and techniques of machine learning in universities may prove extremely daunting for the average undergraduate. During my undergraduate winter exchange in McGill University, I enrolled myself in their Applied Machine Learning course. Yes, it was foolish of me to enroll in a graduate-level course!
How to start career in Data Science and Machine Learning
It does not matter how much experience you have, actually anybody can start or switch to data science and machine learning. The only important this is, how much eager you are for it. What it means to you. If you are very much keen to work in this field then nobody can stop you. There might be some short term hurdles however if you are focused enough and know your goals regarding where you want to see yourself after certain years, then you will definitely be successful in overcoming those hurdles.
Machine learning-based dynamic mortality prediction after traumatic brain injury
Our aim was to create simple and largely scalable machine learning-based algorithms that could predict mortality in a real-time fashion during intensive care after traumatic brain injury. We performed an observational multicenter study including adult TBI patients that were monitored for intracranial pressure (ICP) for at least 24 h in three ICUs. We used machine learning-based logistic regression modeling to create two algorithms (based on ICP, mean arterial pressure [MAP], cerebral perfusion pressure [CPP] and Glasgow Coma Scale [GCS]) to predict 30-day mortality. We used a stratified cross-validation technique for internal validation. Of 472 included patients, 92 patients (19%) died within 30 days.
Text Classification with Extremely Small Datasets
After implementing these we can choose to expand the feature space with polynomial (eg X²) or interaction features (eg XY) by using sklearn's PolynomialFeatures() Note: The choice of feature scaling technique made quite a big difference to the performance of the classifier, I tried RobustScaler, StandardScaler, Normalizer and MinMaxScaler and found that MinMaxScaler worked the best.
Remote Data Scientist at Redox
At Redox our mission is to enable the frictionless adoption of technology in healthcare.To that end we have enabled a network of healthcare organizations and technology developers to connect through us to improve patient healthcare. Our customers have asked us to help minimize data duplication for various scenarios. An ideal candidate is a data science enthusiast and will solve complex data-related problems, using advanced statistical and machine learning tools, in a real-world setting with product managers and engineers. Candidates should demonstrate strong statistical, mathematical and technical skills, with proven capabilities to transition ideas into fully working projects. In addition, the candidate needs to have a creative and first principles mindset to examine our current processes and help us grow our Data Science muscle at Redox.