Clustering
AWT -- Clustering Meteorological Time Series Using an Aggregated Wavelet Tree
Pacher, Christina, Schicker, Irene, deWit, Rosmarie, Hlavackova-Schindler, Katerina, Plant, Claudia
Both clustering and outlier detection play an important role for meteorological measurements. We present the AWT algorithm, a clustering algorithm for time series data that also performs implicit outlier detection during the clustering. AWT integrates ideas of several well-known K-Means clustering algorithms. It chooses the number of clusters automatically based on a user-defined threshold parameter, and it can be used for heterogeneous meteorological input data as well as for data sets that exceed the available memory size. We apply AWT to crowd sourced 2-m temperature data with an hourly resolution from the city of Vienna to detect outliers and to investigate if the final clusters show general similarities and similarities with urban land-use characteristics. It is shown that both the outlier detection and the implicit mapping to land-use characteristic is possible with AWT which opens new possible fields of application, specifically in the rapidly evolving field of urban climate and urban weather.
Top Three Clustering Algorithms You Should Know Instead of K-means Clustering
DBSCAN is a clustering algorithm that groups data points into clusters based on the density of the points. The algorithm works by identifying points that are in high-density regions of the data and expanding those clusters to include all points that are nearby. Points that are not in high-density regions and are not close to any other points are considered noise and are not included in any clusters. This means that DBSCAN can automatically identify the number of clusters in a dataset, unlike other clustering algorithms that require the number of clusters to be specified in advance. DBSCAN is useful for data that has a lot of noise or for data that doesn't have well-defined clusters.
XClusters: Explainability-first Clustering
Hwang, Hyunseung, Whang, Steven Euijong
We study the problem of explainability-first clustering where explainability becomes a first-class citizen for clustering. Previous clustering approaches use decision trees for explanation, but only after the clustering is completed. In contrast, our approach is to perform clustering and decision tree training holistically where the decision tree's performance and size also influence the clustering results. We assume the attributes for clustering and explaining are distinct, although this is not necessary. We observe that our problem is a monotonic optimization where the objective function is a difference of monotonic functions. We then propose an efficient branch-and-bound algorithm for finding the best parameters that lead to a balance of cluster distortion and decision tree explainability. Our experiments show that our method can improve the explainability of any clustering that fits in our framework.
5 Popular Machine Learning Certifications: Your 2023 Guide
When applying for a programming or data science job, machine learning certifications and certificates have the potential to help you stand out from the crowded pool of candidates. Whether you've just completed a course of study or passed an exam offered by a respected institution, obtaining a certificate or certification is a real accomplishment that indicates your knowledge, experience, and expertise in the field of machine learning. But, what certificates and certifications are right for you? In this article, you'll learn more about the difference between certificates and certifications and explore five of the most popular ones for machine learning available today. Though they are often confused, certificates and certifications are not the same.
12 Machine Learning Books You Should Read in 2023 - Machine Learning Techniques
This complements the list that I posted earlier under the title "Math for Machine Learning: 14 Must-Read Books", available here. Many of the following books have a free PDF version, their own website and GitHub repository, and usually you can purchase the print version. Some are self-published, with the PDF version regularly updated, and even
On the Global Solution of Soft k-Means
Nie, Feiping, Chen, Hong, Wang, Rong, Li, Xuelong
This paper presents an algorithm to solve the Soft k-Means problem globally. Unlike Fuzzy c-Means, Soft k-Means (SkM) has a matrix factorization-type objective and has been shown to have a close relation with the popular probability decomposition-type clustering methods, e.g., Left Stochastic Clustering (LSC). Though some work has been done for solving the Soft k-Means problem, they usually use an alternating minimization scheme or the projected gradient descent method, which cannot guarantee global optimality since the non-convexity of SkM. In this paper, we present a sufficient condition for a feasible solution of Soft k-Means problem to be globally optimal and show the output of the proposed algorithm satisfies it. Moreover, for the Soft k-Means problem, we provide interesting discussions on stability, solutions non-uniqueness, and connection with LSC. Then, a new model, named Minimal Volume Soft k-Means (MVSkM), is proposed to address the solutions non-uniqueness issue. Finally, experimental results support our theoretical results.
A parallelizable model-based approach for marginal and multivariate clustering
de Carvalho, Miguel, Venturini, Gabriel Martos, Svetloลกรกk, Andrej
Context and Motivation Clustering is an unsupervised learning approach for the task of partitioning data into meaningful subsets. The huge literature on cluster analysis is difficult to survey in a few sentences, but a concise description of well-known approaches is offered by Hastie et al. (2009), Everitt et al. (2011), and King (2014). Examples of mainstream methods for clustering data include model-based (Bouveyron et al., 2019), similarity-based (MacQueen, 1967; Kaufman and Rousseeuw, 1987), and hierarchical clustering (Hastie et al., 2009, Section 14.3). In this paper we propose a novel model-based approach for cluster analysis that lies at the interface of model-based clustering (i.e., via mixture models) and similarity-based clustering (i.e., via K-means and K-medoids). The proposed approach aims to benefit from the flexibility and soundness of model-based clustering, while attempting to mitigate Pitfalls 1 and 2 below. Model-based clustering is a fast-evolving and intradisciplinary research topic as can be seen from the recent Handbook on Mixture Analysis (Fruhwirth-Schnatter et al., 2019) as well as the survey papers of Melnykov and Maitra (2010), McNicholas (2016), Gormley et al. (2023), and the references therein.
Robust Point Cloud Segmentation with Noisy Annotations
Ye, Shuquan, Chen, Dongdong, Han, Songfang, Liao, Jing
Point cloud segmentation is a fundamental task in 3D. Despite recent progress on point cloud segmentation with the power of deep networks, current learning methods based on the clean label assumptions may fail with noisy labels. Yet, class labels are often mislabeled at both instance-level and boundary-level in real-world datasets. In this work, we take the lead in solving the instance-level label noise by proposing a Point Noise-Adaptive Learning (PNAL) framework. Compared to noise-robust methods on image tasks, our framework is noise-rate blind, to cope with the spatially variant noise rate specific to point clouds. Specifically, we propose a point-wise confidence selection to obtain reliable labels from the historical predictions of each point. A cluster-wise label correction is proposed with a voting strategy to generate the best possible label by considering the neighbor correlations. To handle boundary-level label noise, we also propose a variant ``PNAL-boundary " with a progressive boundary label cleaning strategy. Extensive experiments demonstrate its effectiveness on both synthetic and real-world noisy datasets. Even with $60\%$ symmetric noise and high-level boundary noise, our framework significantly outperforms its baselines, and is comparable to the upper bound trained on completely clean data. Moreover, we cleaned the popular real-world dataset ScanNetV2 for rigorous experiment. Our code and data is available at https://github.com/pleaseconnectwifi/PNAL.
Artificial Intelligence Security Competition (AISC)
Dong, Yinpeng, Chen, Peng, Deng, Senyou, L, Lianji, Sun, Yi, Zhao, Hanyu, Li, Jiaxing, Tan, Yunteng, Liu, Xinyu, Dong, Yangyi, Xu, Enhui, Xu, Jincai, Xu, Shu, Fu, Xuelin, Sun, Changfeng, Han, Haoliang, Zhang, Xuchong, Chen, Shen, Sun, Zhimin, Cao, Junyi, Yao, Taiping, Ding, Shouhong, Wu, Yu, Lin, Jian, Wu, Tianpeng, Wang, Ye, Fu, Yu, Feng, Lin, Gao, Kangkang, Liu, Zeyu, Pang, Yuanzhe, Duan, Chengqi, Zhou, Huipeng, Wang, Yajie, Zhao, Yuhang, Wu, Shangbo, Lyu, Haoran, Lin, Zhiyu, Gao, Yifei, Li, Shuang, Wang, Haonan, Sang, Jitao, Ma, Chen, Zheng, Junhao, Li, Yijia, Shen, Chao, Lin, Chenhao, Cui, Zhichao, Liu, Guoshuai, Shi, Huafeng, Hu, Kun, Zhang, Mengxin
The security of artificial intelligence (AI) is an important research area towards safe, reliable, and trustworthy AI systems. To accelerate the research on AI security, the Artificial Intelligence Security Competition (AISC) was organized by the Zhongguancun Laboratory, China Industrial Control Systems Cyber Emergency Response Team, Institute for Artificial Intelligence, Tsinghua University, and RealAI as part of the Zhongguancun International Frontier Technology Innovation Competition (https://www.zgc-aisc.com/en). The competition consists of three tracks, including Deepfake Security Competition, Autonomous Driving Security Competition, and Face Recognition Security Competition. This report will introduce the competition rules of these three tracks and the solutions of top-ranking teams in each track.
A machine learning approach to support decision in insider trading detection
Mazzarisi, Piero, Ravagnani, Adele, Deriu, Paola, Lillo, Fabrizio, Medda, Francesca, Russo, Antonio
Identifying market abuse activity from data on investors' trading activity is very challenging both for the data volume and for the low signal to noise ratio. Here we propose two complementary unsupervised machine learning methods to support market surveillance aimed at identifying potential insider trading activities. The first one uses clustering to identify, in the vicinity of a price sensitive event such as a takeover bid, discontinuities in the trading activity of an investor with respect to his/her own past trading history and on the present trading activity of his/her peers. The second unsupervised approach aims at identifying (small) groups of investors that act coherently around price sensitive events, pointing to potential insider rings, i.e. a group of synchronised traders displaying strong directional trading in rewarding position in a period before the price sensitive event. As a case study, we apply our methods to investor resolved data of Italian stocks around takeover bids.