Statistical Learning
Approximation of Solution Operators for High-dimensional PDEs
We propose a finite-dimensional control-based method to approximate solution operators for evolutional partial differential equations (PDEs), particularly in high-dimensions. By employing a general reduced-order model, such as a deep neural network, we connect the evolution of the model parameters with trajectories in a corresponding function space. Using the computational technique of neural ordinary differential equation, we learn the control over the parameter space such that from any initial starting point, the controlled trajectories closely approximate the solutions to the PDE. Approximation accuracy is justified for a general class of second-order nonlinear PDEs. Numerical results are presented for several high-dimensional PDEs, including real-world applications to solving Hamilton-Jacobi-Bellman equations. These are demonstrated to show the accuracy and efficiency of the proposed method.
Deep Generative Modeling for Financial Time Series with Application in VaR: A Comparative Review
Ericson, Lars, Zhu, Xuejun, Han, Xusi, Fu, Rao, Li, Shuang, Guo, Steve, Hu, Ping
In the financial services industry, forecasting the risk factor distribution conditional on the history and the current market environment is the key to market risk modeling in general and value at risk (VaR) model in particular. As one of the most widely adopted VaR models in commercial banks, Historical simulation (HS) uses the empirical distribution of daily returns in a historical window as the forecast distribution of risk factor returns in the next day. The objectives for financial time series generation are to generate synthetic data paths with good variety, and similar distribution and dynamics to the original historical data. In this paper, we apply multiple existing deep generative methods (e.g., CGAN, CWGAN, Diffusion, and Signature WGAN) for conditional time series generation, and propose and test two new methods for conditional multi-step time series generation, namely Encoder-Decoder CGAN and Conditional TimeVAE. Furthermore, we introduce a comprehensive framework with a set of KPIs to measure the quality of the generated time series for financial modeling. The KPIs cover distribution distance, autocorrelation and backtesting. All models (HS, parametric and neural networks) are tested on both historical USD yield curve data and additional data simulated from GARCH and CIR processes. The study shows that top performing models are HS, GARCH and CWGAN models. Future research directions in this area are also discussed.
Hierarchical Federated Learning in Multi-hop Cluster-Based VANETs
HaghighiFard, M. Saeid, Coleri, Sinem
The usage of federated learning (FL) in Vehicular Ad hoc Networks (VANET) has garnered significant interest in research due to the advantages of reducing transmission overhead and protecting user privacy by communicating local dataset gradients instead of raw data. However, implementing FL in VANETs faces challenges, including limited communication resources, high vehicle mobility, and the statistical diversity of data distributions. In order to tackle these issues, this paper introduces a novel framework for hierarchical federated learning (HFL) over multi-hop clustering-based VANET. The proposed method utilizes a weighted combination of the average relative speed and cosine similarity of FL model parameters as a clustering metric to consider both data diversity and high vehicle mobility. This metric ensures convergence with minimum changes in cluster heads while tackling the complexities associated with non-independent and identically distributed (non-IID) data scenarios. Additionally, the framework includes a novel mechanism to manage seamless transitions of cluster heads (CHs), followed by transferring the most recent FL model parameter to the designated CH. Furthermore, the proposed approach considers the option of merging CHs, aiming to reduce their count and, consequently, mitigate associated overhead. Through extensive simulations, the proposed hierarchical federated learning over clustered VANET has been demonstrated to improve accuracy and convergence time significantly while maintaining an acceptable level of packet overhead compared to previously proposed clustering algorithms and non-clustered VANET.
Machine learning approach to detect dynamical states from recurrence measures
Thakur, Dheeraja, Mohan, Athul, Ambika, G., Meena, Chandrakala
We integrate machine learning approaches with nonlinear time series analysis, specifically utilizing recurrence measures to classify various dynamical states emerging from time series. We implement three machine learning algorithms Logistic Regression, Random Forest, and Support Vector Machine for this study. The input features are derived from the recurrence quantification of nonlinear time series and characteristic measures of the corresponding recurrence networks. For training and testing we generate synthetic data from standard nonlinear dynamical systems and evaluate the efficiency and performance of the machine learning algorithms in classifying time series into periodic, chaotic, hyper-chaotic, or noisy categories. Additionally, we explore the significance of input features in the classification scheme and find that the features quantifying the density of recurrence points are the most relevant. Furthermore, we illustrate how the trained algorithms can successfully predict the dynamical states of two variable stars, SX Her and AC Her from the data of their light curves.
A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions
Goldshlager, Gil, Abrahamsen, Nilin, Lin, Lin
For many chemical properties, it suffices to work within the Born-Oppenheimer approximation, in which the nuclei are viewed as classical point charges and only the electrons exhibit quantum-mechanical behavior. The study of chemistry through this lens is known as electronic structure theory. Within electronic structure theory, methods to model the many-body electron wavefunction include Hartree-Fock theory, configuration interaction methods, and coupled cluster theory. A typical ansatz for such methods is a sum of Slater determinants which represent antisymmetrized products of single-particle states. The benefit of such an ansatz is that the energy and other properties of the wavefunction can be evaluated analytically from pre-computed few-particle integrals. Another approach to the electronic structure problem is the variational Monte Carlo method (VMC) [1, 2]. In VMC, the properties of the wavefunction are calculated using Monte Carlo sampling rather than direct numerical integration, and the energy is variationally minimized through a stochastic optimization procedure. This increases the cost of the calculations, especially when high accuracy is required, but it enables the use of much more general ansatzes.
Transfer Learning in Human Activity Recognition: A Survey
Dhekane, Sourish Gunesh, Ploetz, Thomas
Sensor-based human activity recognition (HAR) has been an active research area, owing to its applications in smart environments, assisted living, fitness, healthcare, etc. Recently, deep learning based end-to-end training has resulted in state-of-the-art performance in domains such as computer vision and natural language, where large amounts of annotated data are available. However, large quantities of annotated data are not available for sensor-based HAR. Moreover, the real-world settings on which the HAR is performed differ in terms of sensor modalities, classification tasks, and target users. To address this problem, transfer learning has been employed extensively. In this survey, we focus on these transfer learning methods in the application domains of smart home and wearables-based HAR. In particular, we provide a problem-solution perspective by categorizing and presenting the works in terms of their contributions and the challenges they address. We also present an updated view of the state-of-the-art for both application domains. Based on our analysis of 205 papers, we highlight the gaps in the literature and provide a roadmap for addressing them. This survey provides a reference to the HAR community, by summarizing the existing works and providing a promising research agenda.
Neural Echos: Depthwise Convolutional Filters Replicate Biological Receptive Fields
Babaiee, Zahra, Kiasari, Peyman M., Rus, Daniela, Grosu, Radu
In this study, we present evidence suggesting that depthwise convolutional kernels are effectively replicating the structural intricacies of the biological receptive fields observed in the mammalian retina. We provide analytics of trained kernels from various state-of-the-art models substantiating this evidence. Inspired by this intriguing discovery, we propose an initialization scheme that draws inspiration from the biological receptive fields. Experimental analysis of the ImageNet dataset with multiple CNN architectures featuring depthwise convolutions reveals a marked enhancement in the accuracy of the learned model when initialized with biologically derived weights. This underlies the potential for biologically inspired computational models to further our understanding of vision processing systems and to improve the efficacy of convolutional networks.
FLex&Chill: Improving Local Federated Learning Training with Logit Chilling
Lee, Kichang, Kim, Songkuk, Ko, JeongGil
For instance, FedProx [Li et al., 2020] controls Federated learning are inherently hampered by data the number of iterations for each local device, aiming to heterogeneity: non-iid distributed training data train models resilient to challenges posed by non-independent over local clients. We propose a novel model training and non-iid data environments. SCAFFOLD [Karimireddy approach for federated learning, FLex&Chill, et al., 2020] achieves expedited convergence and improved which exploits the Logit Chilling method. Through model accuracy [McMahan et al., 2017] by introducing a extensive evaluations, we demonstrate that, in the correction term during the model aggregation phase to balance presence of non-iid data characteristics inherent in the influence of each client. These operations alleviate federated learning systems, this approach can expedite the problems posed by the non-iid environment, a common model convergence and improve inference accuracy.
False Discovery Rate Control for Gaussian Graphical Models via Neighborhood Screening
Koka, Taulant, Machkour, Jasin, Muma, Michael
Gaussian graphical models emerge in a wide range of fields. They model the statistical relationships between variables as a graph, where an edge between two variables indicates conditional dependence. Unfortunately, well-established estimators, such as the graphical lasso or neighborhood selection, are known to be susceptible to a high prevalence of false edge detections. False detections may encourage inaccurate or even incorrect scientific interpretations, with major implications in applications, such as biomedicine or healthcare. In this paper, we introduce a nodewise variable selection approach to graph learning and provably control the false discovery rate of the selected edge set at a self-estimated level. A novel fusion method of the individual neighborhoods outputs an undirected graph estimate. The proposed method is parameter-free and does not require tuning by the user. Benchmarks against competing false discovery rate controlling methods in numerical experiments considering different graph topologies show a significant gain in performance.
SymbolNet: Neural Symbolic Regression with Adaptive Dynamic Pruning
Tsoi, Ho Fung, Loncar, Vladimir, Dasu, Sridhara, Harris, Philip
Contrary to the use of genetic programming, the neural network approach to symbolic regression can scale well with high input dimension and leverage gradient methods for faster equation searching. Common ways of constraining expression complexity have relied on multistage pruning methods with fine-tuning, but these often lead to significant performance loss. In this work, we propose SymbolNet, a neural network approach to symbolic regression in a novel framework that enables dynamic pruning of model weights, input features, and mathematical operators in a single training, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term per pruning type, which can adaptively adjust its own strength and lead to convergence to a target sparsity level. In contrast to most existing symbolic regression methods that cannot efficiently handle datasets with more than $O$(10) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).