Education
Modularity-Aware Graph Autoencoders for Joint Community Detection and Link Prediction
Salha-Galvan, Guillaume, Lutzeyer, Johannes F., Dasoulas, George, Hennequin, Romain, Vazirgiannis, Michalis
Graph autoencoders (GAE) and variational graph autoencoders (VGAE) emerged as powerful methods for link prediction. Their performances are less impressive on community detection problems where, according to recent and concurring experimental evaluations, they are often outperformed by simpler alternatives such as the Louvain method. It is currently still unclear to which extent one can improve community detection with GAE and VGAE, especially in the absence of node features. It is moreover uncertain whether one could do so while simultaneously preserving good performances on link prediction. In this paper, we show that jointly addressing these two tasks with high accuracy is possible. For this purpose, we introduce and theoretically study a community-preserving message passing scheme, doping our GAE and VGAE encoders by considering both the initial graph structure and modularity-based prior communities when computing embedding spaces. We also propose novel training and optimization strategies, including the introduction of a modularity-inspired regularizer complementing the existing reconstruction losses for joint link prediction and community detection. We demonstrate the empirical effectiveness of our approach, referred to as Modularity-Aware GAE and VGAE, through in-depth experimental validation on various real-world graphs.
Questions for Flat-Minima Optimization of Modern Neural Networks
Kaddour, Jean, Liu, Linqing, Silva, Ricardo, Kusner, Matt J.
For training neural networks, flat-minima optimizers that seek to find parameters in neighborhoods having uniformly low loss (flat minima) have been shown to improve upon stochastic and adaptive gradient-based methods. Two methods for finding flat minima stand out: 1. Averaging methods (i.e., Stochastic Weight Averaging, SWA), and 2. Minimax methods (i.e., Sharpness Aware Minimization, SAM). However, despite similar motivations, there has been limited investigation into their properties and no comprehensive comparison between them. In this work, we investigate the loss surfaces from a systematic benchmarking of these approaches across computer vision, natural language processing, and graph learning tasks. The results lead to a simple hypothesis: since both approaches find different flat solutions, combining them should improve generalization even further. We verify this improves over either flat-minima approach in 39 out of 42 cases. When it does not, we investigate potential reasons. We hope our results across image, graph, and text data will help researchers to improve deep learning optimizers, and practitioners to pinpoint the optimizer for the problem at hand.
Flow-based Algorithms for Improving Clusters: A Unifying Framework, Software, and Performance
Fountoulakis, K., Liu, M., Gleich, D. F., Mahoney, M. W.
Clustering points in a vector space or nodes in a graph is a ubiquitous primitive in statistical data analysis, and it is commonly used for exploratory data analysis. In practice, it is often of interest to "refine" or "improve" a given cluster that has been obtained by some other method. In this survey, we focus on principled algorithms for this cluster improvement problem. Many such cluster improvement algorithms are flow-based methods, by which we mean that operationally they require the solution of a sequence of maximum flow problems on a (typically implicitly) modified data graph. These cluster improvement algorithms are powerful, both in theory and in practice, but they have not been widely adopted for problems such as community detection, local graph clustering, semi-supervised learning, etc. Possible reasons for this are: the steep learning curve for these algorithms; the lack of efficient and easy to use software; and the lack of detailed numerical experiments on real-world data that demonstrate their usefulness. Our objective here is to address these issues. To do so, we guide the reader through the whole process of understanding how to implement and apply these powerful algorithms. We present a unifying fractional programming optimization framework that permits us to distill, in a simple way, the crucial components of all these algorithms. It also makes apparent similarities and differences between related methods. Viewing these cluster improvement algorithms via a fractional programming framework suggests directions for future algorithm development. Finally, we develop efficient implementations of these algorithms in our LocalGraphClustering Python package, and we perform extensive numerical experiments to demonstrate the performance of these methods on social networks and image-based data graphs.
Artificial intelligence (AI): 3 everyday IT tasks where automation fits
If I were to ask someone why they chose a career in information technology, I doubt they would respond with "I love data entry!", "I could debug code all day long!", or "Handling tickets is so much fun, I'd do it even if I didn't get paid for it." Here are the top three ways AI can help automate manual IT tasks, thereby freeing up precious resources and benefiting your teams, businesses, and customers. Grace Murray Hopper was a Navy rear admiral and computer programming pioneer who worked on the Mark II computer at Harvard in the 1940s. On September 9, 1947, Hopper traced an error with the Mark II to โ of all things โ a dead moth in the relay.
Customer Segmentation using Machine Learning in Apache Spark - Projects Based Learning
In this project, we will perform one of the most essential applications of machine learning โ Customer Segmentation. We will implement customer segmentation in Apache Spark and Scala, whenever you need to find your best customer. Customer Segmentation is one of the most important applications of unsupervised learning. In this machine learning project, we will make use of K-means clustering which is the essential algorithm for clustering unlabeled datasets. Welcome to this project on Customer Segmentation using Apache Spark Machine Learning using Apache Zeppelin platform which allows you to execute your spark code in Apache Zeppelin notebook.
7 Steps to Mastering Machine Learning with Python in 2022 - KDnuggets
Are you trying to teach yourself machine learning from scratch, but aren't sure where to start? Or maybe you've taken an online course or two, but have hit a roadblock in your learning journey and don't know how to proceed. I was in a similar position just two years ago. I had spent over $25K in university fees, but was still inexperienced and unprepared for the job market. It took a lot of trial and error for me to come up with a machine learning roadmap.
Data Literacy Education Framework โ Part 2 - DataScienceCentral.com
What do we need to do to increase the data literacy of our organization? In a world where your personal data, and the preferences and biases buried in that data, are being used to influence your behaviors, beliefs, and decisions, data literacy becomes a fundamental skill. And it's not just corporations that need this training. Data Literacy should be taught in universities, in high schools, in middle schools and even in adult education and nursing homes. In the first blog of this two-part series on the Data Literacy Education Framework, I introduced the 4 stages of the Data Literacy Educational Framework, a framework which organizations, universities, high schools, and even adult education programs can use to create a more holistic data literacy training. Now, I want to complete the Data Literacy Education Framework by discussing the third (AI / ML Literacy) and fourth stages (Prediction and Statistical Literacy) of the Data Literacy Education Framework.
Data Science : Master Machine Learning Without Coding
There's literally no other course on Udemy that teaches Machine Learning without the need for programming knowledge or coding, using free open source software! Why Data Science and Machine Learning are the Hottest and Most In-Demand Technology Jobs. Data Scientist was recently dubbed "The Sexiest Job of the 21st Century" by Harvard Business Review, and for good reason! If you're looking for a fast and effective way to earn a 6-figure income without spending thousands of dollars in training, keep reading to learn about this revolutionary Udemy course. Glassdoor reports that Data Scientist was named the "Best Job in America for 2016," which was based on the huge amount of career opportunities and 6-figure average salary.
Why 'the future of AI is the future of work'
Amid widespread anxiety about automation and machines displacing workers, the idea that technological advances aren't necessarily driving us toward a jobless future is good news. At the same time, "many in our country are failing to thrive in a labor market that generates plenty of jobs but little economic security," MIT professors David Autor and David Mindell and principal research scientist Elisabeth Reynolds write in their new book "The Work of the Future: Building Better Jobs in an Age of Intelligent Machines." The authors lay out findings from their work chairing the MIT Task Force on the Work of the Future, which MIT president L. Rafael Reif commissioned in 2018. The task force was charged with understanding the relationships between emerging technologies and work, helping shape realistic expectations of technology, and exploring strategies for a future of shared prosperity. Autor, Mindell, and Reynolds worked with 20 faculty members and 20 graduate students who contributed research.
ColloSSL: Collaborative Self-Supervised Learning for Human Activity Recognition
Jain, Yash, Tang, Chi Ian, Min, Chulhong, Kawsar, Fahim, Mathur, Akhil
A major bottleneck in training robust Human-Activity Recognition models (HAR) is the need for large-scale labeled sensor datasets. Because labeling large amounts of sensor data is an expensive task, unsupervised and semi-supervised learning techniques have emerged that can learn good features from the data without requiring any labels. In this paper, we extend this line of research and present a novel technique called Collaborative Self-Supervised Learning (ColloSSL) which leverages unlabeled data collected from multiple devices worn by a user to learn high-quality features of the data. A key insight that underpins the design of ColloSSL is that unlabeled sensor datasets simultaneously captured by multiple devices can be viewed as natural transformations of each other, and leveraged to generate a supervisory signal for representation learning. We present three technical innovations to extend conventional self-supervised learning algorithms to a multi-device setting: a Device Selection approach which selects positive and negative devices to enable contrastive learning, a Contrastive Sampling algorithm which samples positive and negative examples in a multi-device setting, and a loss function called Multi-view Contrastive Loss which extends standard contrastive loss to a multi-device setting. Our experimental results on three multi-device datasets show that ColloSSL outperforms both fully-supervised and semi-supervised learning techniques in majority of the experiment settings, resulting in an absolute increase of upto 7.9% in F_1 score compared to the best performing baselines. We also show that ColloSSL outperforms the fully-supervised methods in a low-data regime, by just using one-tenth of the available labeled data in the best case.