Clustering
Transfer Learning
Machine Learning (ML) involves data analysis and enables the system to improve and learn from experience without explicit programming required constantly. There have been many ML approaches that came into existence constantly. Supervised learning was a game-changing approach that was adopted widely across many industries. However, a few limitations of supervised learning can be overcome with the onset of various other approaches. Transfer Learning is a method under research in Machine Learning that stores the knowledge obtained from solving one problem and uses it to solve problems that are different but related to the solved one. Since training a model takes more computational power, time, and data, Transfer Learning helps reduce the same while improving learning accuracy. The target learner learns from the model, which is already trained initially by using the stored knowledge.
Solving clustering as ill-posed problem: experiments with K-Means algorithm
In this contribution, the clustering procedure based on K-Means algorithm is studied as an inverse problem, which is a special case of the illposed problems. The attempts to improve the quality of the clustering inverse problem drive to reduce the input data via Principal Component Analysis (PCA). Since there exists a theorem by Ding and He that links the cardinality of the optimal clusters found with K-Means and the cardinality of the selected informative PCA components, the computational experiments tested the theorem between two quantitative features selection methods: Kaiser criteria (based on imperative decision) versus Wishart criteria (based on random matrix theory). The results suggested that PCA reduction with features selection by Wishart criteria leads to a low matrix condition number and satisfies the relation between clusters and components predicts by the theorem. The data used for the computations are from a neuroscientific repository: it regards healthy and young subjects that performed a task-oriented functional Magnetic Resonance Imaging (fMRI) paradigm.
The Association Between SOC and Land Prices Considering Spatial Heterogeneity Based on Finite Mixture Modeling
Kang, Woo Seok, Kim, Eunchan, Heo, Wookjae
An understanding of how Social Overhead Capital (SOC) is associated with the land value of the local community is important for effective urban planning. However, even within a district, there are multiple sections used for different purposes; the term for this is spatial heterogeneity. The spatial heterogeneity issue has to be considered when attempting to comprehend land prices. If there is spatial heterogeneity within a district, land prices can be managed by adopting the spatial clustering method. In this study, spatial attributes including SOC, socio-demographic features, and spatial information in a specific district are analyzed with Finite Mixture Modeling (FMM) in order to find (a) the optimal number of clusters and (b) the association among SOCs, socio-demographic features, and land prices. FMM is a tool used to find clusters and the attributes' coefficients simultaneously. Using the FMM method, the results show that four clusters exist in one district and the four clusters have different associations among SOCs, demographic features, and land prices. Policymakers and managerial administration need to look for information to make policy about land prices. The current study finds the consideration of closeness to SOC to be a significant factor on land prices and suggests the potential policy direction related to SOC.
User-Specific Bicluster-based Collaborative Filtering: Handling Preference Locality, Sparsity and Subjectivity
Silva, Miguel G., Henriques, Rui, Madeira, Sara C.
As an attempt to cope with massive range of options, there has been large academic and industry interest in automatically recommending items to individuals since last century. Spotify, Amazon, Netflix, and Facebook are some popular platforms that actively use recommender systems [13]. From e-commerce to online advertisement, these systems are unavoidable in our daily online journeys to suggest items in a personalized way. Collaborative Filtering (CF) approaches, firstly proposed by [19], are currently seen as the widest implemented and most mature of the technologies to build recommender systems. Given a set of observed item ratings, CF aims at estimating unknown preferences based on the assumption that users with similar preferences in the past will yield similar preferences in the future. Despite the role of Collaborative Filtering, significant challenges limit its effectiveness, including the diversity and locality of user preferences, the structural sparsity of user-item ratings, the subjectivity of rating scales, and the increasingly large user and item bases [13, 49]. To address the diversity of user profiles, reduce the dimensionality and minimize rating sparsity, matrix factorization and clustering approaches have been combined within CF approaches for two decades [13]. However, traditional clustering techniques are typically applied to either group users or items separately. In real-world CF scenarios, the preferences of a subset of users is frequently only significantly correlated on a subset of the overall items, and vice versa [47].
Reads2Vec: Efficient Embedding of Raw High-Throughput Sequencing Reads Data
Chourasia, Prakash, Ali, Sarwan, Ciccolella, Simone, Della Vedova, Gianluca, Patterson, Murray
The massive amount of genomic data appearing for SARS-CoV-2 since the beginning of the COVID-19 pandemic has challenged traditional methods for studying its dynamics. As a result, new methods such as Pangolin, which can scale to the millions of samples of SARS-CoV-2 currently available, have appeared. Such a tool is tailored to take as input assembled, aligned and curated full-length sequences, such as those found in the GISAID database. As high-throughput sequencing technologies continue to advance, such assembly, alignment and curation may become a bottleneck, creating a need for methods which can process raw sequencing reads directly. In this paper, we propose Reads2Vec, an alignment-free embedding approach that can generate a fixed-length feature vector representation directly from the raw sequencing reads without requiring assembly. Furthermore, since such an embedding is a numerical representation, it may be applied to highly optimized classification and clustering algorithms. Experiments on simulated data show that our proposed embedding obtains better classification results and better clustering properties contrary to existing alignment-free baselines. In a study on real data, we show that alignment-free embeddings have better clustering properties than the Pangolin tool and that the spike region of the SARS-CoV-2 genome heavily informs the alignment-free clusterings, which is consistent with current biological knowledge of SARS-CoV-2.
Machine Learning Performance Analysis to Predict Stroke Based on Imbalanced Medical Dataset
Cerebral stroke, the second most substantial cause of death universally, has been a primary public health concern over the last few years. With the help of machine learning techniques, early detection of various stroke alerts is accessible, which can efficiently prevent or diminish the stroke. Medical datasets, however, are frequently unbalanced in their class label, with a tendency to poorly predict minority classes. In this paper, the potential risk factors for stroke are investigated. Moreover, four distinctive approaches are applied to improve the classification of the minority class in the imbalanced stroke dataset, which are the ensemble weight voting classifier, the Synthetic Minority Over-sampling Technique (SMOTE), Principal Component Analysis with K-Means Clustering (PCA-Kmeans), Focal Loss with the Deep Neural Network (DNN) and compare their performance. Through the analysis results, SMOTE and PCA-Kmeans with DNN-Focal Loss work best for the limited size of a large severe imbalanced dataset (e.g., Stroke dataset), which is 2-4 times outperform Kaggle's work.
A Dataset and Baseline Approach for Identifying Usage States from Non-Intrusive Power Sensing With MiDAS IoT-based Sensors
Muppasani, Bharath, Anand, Cheyyur Jaya, Appajigowda, Chinmayi, Srivastava, Biplav, Johri, Lokesh
Authors in (Rajapaksha and The growth in the deployment of Internet of Things (IoT) Bergmeir 2022) focused on providing rule based explanations sensors across different industries has opened several opportunities for a particular forecast, considering the global forecasting for the economy. One of them is the collection of IoT model as a black-box model trained across multivariate data that companies can use to build smarter solutions.
Automated Cancer Subtyping via Vector Quantization Mutual Information Maximization
Chen, Zheng, Zhu, Lingwei, Yang, Ziwei, Matsubara, Takashi
Cancer subtyping is crucial for understanding the nature of tumors and providing suitable therapy. However, existing labelling methods are medically controversial, and have driven the process of subtyping away from teaching signals. Moreover, cancer genetic expression profiles are high-dimensional, scarce, and have complicated dependence, thereby posing a serious challenge to existing subtyping models for outputting sensible clustering. In this study, we propose a novel clustering method for exploiting genetic expression profiles and distinguishing subtypes in an unsupervised manner. The proposed method adaptively learns categorical correspondence from latent representations of expression profiles to the subtypes output by the model. By maximizing the problem -- agnostic mutual information between input expression profiles and output subtypes, our method can automatically decide a suitable number of subtypes. Through experiments, we demonstrate that our proposed method can refine existing controversial labels, and, by further medical analysis, this refinement is proven to have a high correlation with cancer survival rates.
Out-of-Dynamics Imitation Learning from Multimodal Demonstrations
Qiu, Yiwen, Wu, Jialong, Cao, Zhangjie, Long, Mingsheng
Existing imitation learning works mainly assume that the demonstrator who collects demonstrations shares the same dynamics as the imitator. However, the assumption limits the usage of imitation learning, especially when collecting demonstrations for the imitator is difficult. In this paper, we study out-of-dynamics imitation learning (OOD-IL), which relaxes the assumption to that the demonstrator and the imitator have the same state spaces but could have different action spaces and dynamics. OOD-IL enables imitation learning to utilize demonstrations from a wide range of demonstrators but introduces a new challenge: some demonstrations cannot be achieved by the imitator due to the different dynamics. Prior works try to filter out such demonstrations by feasibility measurements, but ignore the fact that the demonstrations exhibit a multimodal distribution since the different demonstrators may take different policies in different dynamics. We develop a better transferability measurement to tackle this newly-emerged challenge. We firstly design a novel sequence-based contrastive clustering algorithm to cluster demonstrations from the same mode to avoid the mutual interference of demonstrations from different modes, and then learn the transferability of each demonstration with an adversarial-learning based algorithm in each cluster. Experiment results on several MuJoCo environments, a driving environment, and a simulated robot environment show that the proposed transferability measurement more accurately finds and down-weights non-transferable demonstrations and outperforms prior works on the final imitation learning performance. We show the videos of our experiment results on our website.
Online Correlation Clustering for Dynamic Complete Signed Graphs
In the correlation clustering problem for complete signed graphs, the input is a complete signed graph with edges weighted as $+1$ (denote recommendation to put this pair in the same cluster) or $-1$ (recommending to put this pair of vertices in separate clusters) and the target is to cluster the set of vertices such that the number of disagreements with these recommendations is minimized. In this paper, we consider the problem of correlation clustering for dynamic complete signed graphs where (1) a vertex can be added or deleted, and (2) the sign of an edge can be flipped. In the proposed online scheme, the offline approximation algorithm in [CALM+21] for correlation clustering is used. Up to the author's knowledge, this is the first online algorithm for dynamic graphs which allows a full set of graph editing operations. The proposed approach is rigorously analyzed and compared with a baseline method, which runs the original offline algorithm on each time step. Our results show that the dynamic operations have local effects on the neighboring vertices and we employ this locality to reduce the dependency of the running time in the Baseline to the summation of the degree of all vertices in $G_t$, the graph after applying the graph edit operation at time step $t$, to the summation of the degree of the changing vertices (e.g. two endpoints of an edge) and the number of clusters in the previous time step. Moreover, the required working memory is reduced to the square of the summation of the degree of the modified edge endpoints rather than the total number of vertices in the graph.