Education
A-LAQ: Adaptive Lazily Aggregated Quantized Gradient
Mahmoudi, Afsaneh, Júnior, José Mairton Barros Da Silva, Ghadikolaei, Hossein S., Fischione, Carlo
Federated Learning (FL) plays a prominent role in solving machine learning problems with data distributed across clients. In FL, to reduce the communication overhead of data between clients and the server, each client communicates the local FL parameters instead of the local data. However, when a wireless network connects clients and the server, the communication resource limitations of the clients may prevent completing the training of the FL iterations. Therefore, communication-efficient variants of FL have been widely investigated. Lazily Aggregated Quantized Gradient (LAQ) is one of the promising communication-efficient approaches to lower resource usage in FL. However, LAQ assigns a fixed number of bits for all iterations, which may be communication-inefficient when the number of iterations is medium to high or convergence is approaching. This paper proposes Adaptive Lazily Aggregated Quantized Gradient (A-LAQ), which is a method that significantly extends LAQ by assigning an adaptive number of communication bits during the FL iterations. We train FL in an energy-constraint condition and investigate the convergence analysis for A-LAQ. The experimental results highlight that A-LAQ outperforms LAQ by up to a $50$% reduction in spent communication energy and an $11$% increase in test accuracy.
The least-control principle for local learning at equilibrium
Meulemans, Alexander, Zucchet, Nicolas, Kobayashi, Seijin, von Oswald, Johannes, Sacramento, João
Equilibrium systems are a powerful way to express neural computations. As special cases, they include models of great current interest in both neuroscience and machine learning, such as deep neural networks, equilibrium recurrent neural networks, deep equilibrium models, or meta-learning. Here, we present a new principle for learning such systems with a temporally- and spatially-local rule. Our principle casts learning as a least-control problem, where we first introduce an optimal controller to lead the system towards a solution state, and then define learning as reducing the amount of control needed to reach such a state. We show that incorporating learning signals within a dynamics as an optimal control enables transmitting activity-dependent credit assignment information, avoids storing intermediate states in memory, and does not rely on infinitesimal learning signals. In practice, our principle leads to strong performance matching that of leading gradient-based learning methods when applied to an array of problems involving recurrent neural networks and meta-learning. Our results shed light on how the brain might learn and offer new ways of approaching a broad class of machine learning problems.
Optimal-er Auctions through Attention
Ivanov, Dmitry, Safiulin, Iskander, Filippov, Igor, Balabaeva, Ksenia
RegretNet is a recent breakthrough in the automated design of revenue-maximizing auctions. It combines the flexibility of deep learning with the regret-based approach to relax the Incentive Compatibility (IC) constraint (that participants prefer to bid truthfully) in order to approximate optimal auctions. We propose two independent improvements of RegretNet. The first is a neural architecture denoted as Regret-Former that is based on attention layers. The second is a loss function that requires explicit specification of an acceptable IC violation denoted as regret budget. We investigate both modifications in an extensive experimental study that includes settings with constant and inconstant number of items and participants, as well as novel validation procedures tailored to regret-based approaches. We find that RegretFormer consistently outperforms RegretNet in revenue (i.e. is optimal-er) and that our loss function both simplifies hyperparameter tuning and allows to unambiguously control the revenue-regret trade-off by selecting the regret budget.
Decentralized adaptive clustering of deep nets is beneficial for client collaboration
Zec, Edvin Listo, Ekblom, Ebba, Willbo, Martin, Mogren, Olof, Girdzijauskas, Sarunas
We study the problem of training personalized deep learning models in a decentralized peer-to-peer setting, focusing on the setting where data distributions differ between the clients and where different clients have different local learning tasks. We study both covariate and label shift, and our contribution is an algorithm which for each client finds beneficial collaborations based on a similarity estimate for the local task. Our method does not rely on hyperparameters which are hard to estimate, such as the number of client clusters, but rather continuously adapts to the network topology using soft cluster assignment based on a novel adaptive gossip algorithm. We test the proposed method in various settings where data is not independent and identically distributed among the clients. The experimental evaluation shows that the proposed method performs better than previous state-of-the-art algorithms for this problem setting, and handles situations well where previous methods fail.
Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic Change
Su, Zhaochen, Tang, Zecheng, Guan, Xinyan, Li, Juntao, Wu, Lijun, Zhang, Min
Recent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., the language model pre-trained on static data from past years performs worse over time on emerging data. Existing methods mainly perform continual training to mitigate such a misalignment. While effective to some extent but is far from being addressed on both the language modeling and downstream tasks. In this paper, we empirically observe that temporal generalization is closely affiliated with lexical semantic change, which is one of the essential phenomena of natural languages. Based on this observation, we propose a simple yet effective lexical-level masking strategy to post-train a converged language model. Experiments on two pre-trained language models, two different classification tasks, and four benchmark datasets demonstrate the effectiveness of our proposed method over existing temporal adaptation methods, i.e., continual training with new data. Our code is available at \url{https://github.com/zhaochen0110/LMLM}.
Zero-Shot Text Classification with Self-Training
Gera, Ariel, Halfon, Alon, Shnarch, Eyal, Perlitz, Yotam, Ein-Dor, Liat, Slonim, Noam
Recent advances in large pretrained language models have increased attention to zero-shot text classification. In particular, models finetuned on natural language inference datasets have been widely adopted as zero-shot classifiers due to their promising results and off-the-shelf availability. However, the fact that such models are unfamiliar with the target task can lead to instability and performance issues. We propose a plug-and-play method to bridge this gap using a simple self-training approach, requiring only the class names along with an unlabeled dataset, and without the need for domain expertise or trial and error. We show that fine-tuning the zero-shot classifier on its most confident predictions leads to significant performance gains across a wide range of text classification tasks, presumably since self-training adapts the zero-shot model to the task at hand.
[100%OFF] Numpy And Pandas For Beginners
This is Numpy and Pandas for Beginners course. An excellent choice for both beginners and experts looking to expand their knowledge on one of the most popular Python libraries in the world! If you've spent time in a spreadsheet software like MS Excel or Google Sheets and want to take your data analysis skills to the next level, this course is for you! Pandas is a Python package providing fast, flexible, and expressive data structures designed to make working with "relational" or "labeled" data both easy and intuitive. It aims to be the fundamental high-level building block for doing practical, real-world data analysis in Python.
Grant will help Career Center start artificial-intelligence club - Sun Gazette
The future is now, or soon will be, at the Arlington Career Center. Arlington School Board members on Oct. 27 approved a one-year, $10,000 grant in support of the new artificial-intelligence club that is starting at the school. The program will be overseen by physics instructor Ryan Miller in collaboration with Inspirit AI, which provides curriculum for middle-school and high-school students in the artificial-intelligence field.
Breaking into Data Science and Machine Learning with Python
Let me tell you my story. I graduated with my Ph. D. in computational nano-electronics but I have been working as a data scientist in most of my career. My undergrad and graduate major was in electrical engineering (EE) and minor in Physics. After first year of my job in Intel as a "yield analysis engineer" (now they changed the title to Data Scientist), I literally broke into data science by taking plenty of online classes.
[100%OFF] NumPy - Pandas - PostgreSQL Basic To Advanced For Beginners
This is Numpy,Pandas and PostgreSQL for Beginners course. One question or concern I get a lot is that people want to learn deep learning and data science, so they take these courses, but they get left behind because they don't know enough about the Numpy stack in order to turn those concepts into code. Even if I write the code in full, if you don't know Numpy, then it's still very hard to read. This course is designed to remove that obstacle – to show you how to do things in the Numpy stack that are frequently needed in deep learning and data science. So what are those things?