Optimization
FedRC: Tackling Diverse Distribution Shifts Challenge in Federated Learning by Robust Clustering
Guo, Yongxin, Tang, Xiaoying, Lin, Tao
Federated Learning (FL) is a machine learning paradigm that safeguards privacy by retaining client data on edge devices. However, optimizing FL in practice can be challenging due to the diverse and heterogeneous nature of the learning system. Though recent research has focused on improving the optimization of FL when distribution shifts occur among clients, ensuring global performance when multiple types of distribution shifts occur simultaneously among clients -- such as feature distribution shift, label distribution shift, and concept shift -- remain under-explored. In this paper, we identify the learning challenges posed by the simultaneous occurrence of diverse distribution shifts and propose a clustering principle to overcome these challenges. Through our research, we find that existing methods failed to address the clustering principle. Therefore, we propose a novel clustering algorithm framework, dubbed as FedRC, which adheres to our proposed clustering principle by incorporating a bi-level optimization problem and a novel objective function. Extensive experiments demonstrate that FedRC significantly outperforms other SOTA cluster-based FL methods. Our code will be publicly available.
Data-Driven Minimax Optimization with Expectation Constraints
Yang, Shuoguang, Li, Xudong, Lan, Guanghui
Attention to data-driven optimization approaches, including the well-known stochastic gradient descent method, has grown significantly over recent decades, but data-driven constraints have rarely been studied, because of the computational challenges of projections onto the feasible set defined by these hard constraints. In this paper, we focus on the non-smooth convex-concave stochastic minimax regime and formulate the data-driven constraints as expectation constraints. The minimax expectation constrained problem subsumes a broad class of real-world applications, including two-player zero-sum game and data-driven robust optimization. We propose a class of efficient primal-dual algorithms to tackle the minimax expectation-constrained problem, and show that our algorithms converge at the optimal rate of $\mathcal{O}(\frac{1}{\sqrt{N}})$. We demonstrate the practical efficiency of our algorithms by conducting numerical experiments on large-scale real-world applications.
High Dimensional Causal Inference with Variational Backdoor Adjustment
Israel, Daniel, Grover, Aditya, Broeck, Guy Van den
Backdoor adjustment is a technique in causal inference for estimating interventional quantities from purely observational data. For example, in medical settings, backdoor adjustment can be used to control for confounding and estimate the effectiveness of a treatment. However, high dimensional treatments and confounders pose a series of potential pitfalls: tractability, identifiability, optimization. In this work, we take a generative modeling approach to backdoor adjustment for high dimensional treatments and confounders. We cast backdoor adjustment as an optimization problem in variational inference without reliance on proxy variables and hidden confounders. Empirically, our method is able to estimate interventional likelihood in a variety of high dimensional settings, including semi-synthetic X-ray medical data. To the best of our knowledge, this is the first application of backdoor adjustment in which all the relevant variables are high dimensional.
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
Ustimenko, Aleksei, Beznosikov, Aleksandr
The connection between diffusion processes and homogeneous Markov chains has been investigated for a long time Skorokhod [1963]. If we need to approximate the given diffusion by some homogeneous Markov chain, it is easy to realize because we are free to construct the chain nicely, meaning that we can choose terms and properties of MC, e.g., as it was shown in Raginsky et al. [2017]. However, often the inverse problem arises, namely, we have the a priori given chain, and the goal is to study it via the corresponding diffusion approximation. This task is an increasingly popular and hot research topic. Indeed, it is used to investigate different sampling techniques Orvieto and Lucchi [2018], to describe the behavior of optimization methods Raginsky et al. [2017] and to understand the convergence of boosting algorithms Ustimenko and Prokhorenkova [2021]. From practical experience, the given Markov chain may not have good properties that are easy to analyze in theory. Thus, the aim of our work is to study when diffusion approximation holds for as broad as the possible class of homogeneous Markov chains, i.e., we want to consider the maximally general chain and place the broadest possible assumptions on it whilst obtaining diffusion approximation guarantee.
Action-State Dependent Dynamic Model Selection
Cordoni, Francesco, Sancetta, Alessio
A model among many may only be best under certain states of the world. Switching from a model to another can also be costly. Finding a procedure to dynamically choose a model in these circumstances requires to solve a complex estimation procedure and a dynamic programming problem. A Reinforcement learning algorithm is used to approximate and estimate from the data the optimal solution to this dynamic programming problem. The algorithm is shown to consistently estimate the optimal policy that may choose different models based on a set of covariates. A typical example is the one of switching between different portfolio models under rebalancing costs, using macroeconomic information. Using a set of macroeconomic variables and price data, an empirical application to the aforementioned portfolio problem shows superior performance to choosing the best portfolio model with hindsight.
Bayesian Optimisation for Sequential Experimental Design with Applications in Additive Manufacturing
Zhang, Mimi, Parnell, Andrew, Brabazon, Dermot, Benavoli, Alessio
Engineering designs are usually performed under strict budget constraints. Collecting a single datum from computer experiments such as computational fluid dynamics can potentially take weeks or months. Each datum obtained, whether from a simulation or a physical experiment, needs to be maximally informative of the goals we are trying to accomplish. It is thus crucial to decide where and how to collect the necessary data to learn most about the subject of study. Data-driven experimental design appears in many different contexts in chemistry and physics (e.g. Lam et al., 2018) where the design is an iterative process and the outcomes of previous experiments are exploited to make an informed selection of the next design to evaluate. Mathematically, it is often formulated as an optimization problem of a black-box function (that is, the input-output relation is complex and not analytically available). Bayesian optimization (BO) is a well-established technique for blackbox optimization and is primarily used in situations where (1) the objective function is complex and does not have a closed form, (2) no gradient information is available, and (3) function evaluations are expensive (see Frazier, 2018, for a tutorial). BO has been shown to be sample-efficient in many domains (e.g.
Enhancing SAEAs with Unevaluated Solutions: A Case Study of Relation Model for Expensive Optimization
Hao, Hao, Zhang, Xiaoqun, Zhou, Aimin
Surrogate-assisted evolutionary algorithms (SAEAs) hold significant importance in resolving expensive optimization problems~(EOPs). Extensive efforts have been devoted to improving the efficacy of SAEAs through the development of proficient model-assisted selection methods. However, generating high-quality solutions is a prerequisite for selection. The fundamental paradigm of evaluating a limited number of solutions in each generation within SAEAs reduces the variance of adjacent populations, thus impacting the quality of offspring solutions. This is a frequently encountered issue, yet it has not gained widespread attention. This paper presents a framework using unevaluated solutions to enhance the efficiency of SAEAs. The surrogate model is employed to identify high-quality solutions for direct generation of new solutions without evaluation. To ensure dependable selection, we have introduced two tailored relation models for the selection of the optimal solution and the unevaluated population. A comprehensive experimental analysis is performed on two test suites, which showcases the superiority of the relation model over regression and classification models in the selection phase. Furthermore, the surrogate-selected unevaluated solutions with high potential have been shown to significantly enhance the efficiency of the algorithm.
ZooPFL: Exploring Black-box Foundation Models for Personalized Federated Learning
Lu, Wang, Yu, Hao, Wang, Jindong, Teney, Damien, Wang, Haohan, Chen, Yiqiang, Yang, Qiang, Xie, Xing, Ji, Xiangyang
When personalized federated learning (FL) meets large foundation models, new challenges arise from various limitations in resources. In addition to typical limitations such as data, computation, and communication costs, access to the models is also often limited. This paper endeavors to solve both the challenges of limited resources and personalization. PFL that uses Zeroth-Order Optimization for Personalized Federated Learning. PFL avoids direct interference with the foundation models and instead learns to adapt its inputs through zeroth-order optimization. In addition, we employ simple yet effective linear projections to remap its predictions for personalization. To reduce the computation costs and enhance personalization, we propose input surgery to incorporate an auto-encoder with low-dimensional and client-specific embeddings. PFL to analyze its convergence. Extensive empirical experiments on computer vision and natural language processing tasks using popular foundation models demonstrate its effectiveness for FL on black-box foundation models. In recent years, the growing emphasis on data privacy and security has led to the emergence of federated learning (FL) (Warnat-Herresthal et al., 2021; Chen & Chao, 2022; Chen et al., 2023b; Castiglia et al., 2023; Rodrรญguez-Barroso et al., 2023; Kuang et al., 2023). FL enables collaborative learning while safeguarding data privacy and security across distributed clients (Yang et al., 2019). However, FL faces two key challenges: limited resources and distribution shifts (Figure 1 (a, b)). The rise of large foundation models (Bommasani et al., 2021) has amplified these challenges. The computational demands and communication costs associated with such models hinder the deployment of existing FL approaches (Figure 1a).
Intelligent DRL-Based Adaptive Region of Interest for Delay-sensitive Telemedicine Applications
Soliman, Abdulrahman, Mohamed, Amr, Yaacoub, Elias, Navkar, Nikhil V., Erbad, Aiman
Telemedicine applications have recently received substantial potential and interest, especially after the COVID-19 pandemic. Remote experience will help people get their complex surgery done or transfer knowledge to local surgeons, without the need to travel abroad. Even with breakthrough improvements in internet speeds, the delay in video streaming is still a hurdle in telemedicine applications. This imposes using image compression and region of interest (ROI) techniques to reduce the data size and transmission needs. This paper proposes a Deep Reinforcement Learning (DRL) model that intelligently adapts the ROI size and non-ROI quality depending on the estimated throughput. The delay and structural similarity index measure (SSIM) comparison are used to assess the DRL model. The comparison findings and the practical application reveal that DRL is capable of reducing the delay by 13% and keeping the overall quality in an acceptable range. Since the latency has been significantly reduced, these findings are a valuable enhancement to telemedicine applications.
Asymmetrically Decentralized Federated Learning
Li, Qinglun, Zhang, Miao, Yin, Nan, Yin, Quanjun, Shen, Li
To address the communication burden and privacy concerns associated with the centralized server in Federated Learning (FL), Decentralized Federated Learning (DFL) has emerged, which discards the server with a peer-to-peer (P2P) communication framework. However, most existing DFL algorithms are based on symmetric topologies, such as ring and grid topologies, which can easily lead to deadlocks and are susceptible to the impact of network link quality in practice. To address these issues, this paper proposes the DFedSGPSM algorithm, which is based on asymmetric topologies and utilizes the Push-Sum protocol to effectively solve consensus optimization problems. To further improve algorithm performance and alleviate local heterogeneous overfitting in Federated Learning (FL), our algorithm combines the Sharpness Aware Minimization (SAM) optimizer and local momentum. The SAM optimizer employs gradient perturbations to generate locally flat models and searches for models with uniformly low loss values, mitigating local heterogeneous overfitting. The local momentum accelerates the optimization process of the SAM optimizer. Theoretical analysis proves that DFedSGPSM achieves a convergence rate of $\mathcal{O}(\frac{1}{\sqrt{T}})$ in a non-convex smooth setting under mild assumptions. This analysis also reveals that better topological connectivity achieves tighter upper bounds. Empirically, extensive experiments are conducted on the MNIST, CIFAR10, and CIFAR100 datasets, demonstrating the superior performance of our algorithm compared to state-of-the-art optimizers.