Statistical Learning
General bounds on the quality of Bayesian coresets
Bayesian coresets speed up posterior inference in the large-scale data regime by approximating the full-data log-likelihood function with a surrogate log-likelihood based on a small, weighted subset of the data. But while Bayesian coresets and methods for construction are applicable in a wide range of models, existing theoretical analysis of the posterior inferential error incurred by coreset approximations only apply in restrictive settings -- i.e., exponential family models, or models with strong log-concavity and smoothness assumptions. This work presents general upper and lower bounds on the Kullback-Leibler (KL) divergence of coreset approximations that reflect the full range of applicability of Bayesian coresets. The lower bounds require only mild model assumptions typical of Bayesian asymptotic analyses, while the upper bounds require the log-likelihood functions to satisfy a generalized subexponentiality criterion that is weaker than conditions used in earlier work. The lower bounds are applied to obtain fundamental limitations on the quality of coreset approximations, and to provide a theoretical explanation for the previously-observed poor empirical performance of importance sampling-based construction methods. The upper bounds are used to analyze the performance of recent subsample-optimize methods. The flexibility of the theory is demonstrated in validation experiments involving multimodal, unidentifiable, heavy-tailed Bayesian posterior distributions.
Movie Revenue Prediction using Machine Learning Models
Udandarao, Vikranth, Gupta, Pratyush
In the contemporary film industry, accurately predicting a movie's earnings is paramount for maximizing profitability. This project aims to develop a machine learning model for predicting movie earnings based on input features like the movie name, the MPAA rating of the movie, the genre of the movie, the year of release of the movie, the IMDb Rating, the votes by the watchers, the director, the writer and the leading cast, the country of production of the movie, the budget of the movie, the production company and the runtime of the movie. Through a structured methodology involving data collection, preprocessing, analysis, model selection, evaluation, and improvement, a robust predictive model is constructed. Linear Regression, Decision Trees, Random Forest Regression, Bagging, XGBoosting and Gradient Boosting have been trained and tested. Model improvement strategies include hyperparameter tuning and cross-validation. The resulting model offers promising accuracy and generalization, facilitating informed decision-making in the film industry to maximize profits.
QComp: A QSAR-Based Data Completion Framework for Drug Discovery
Yang, Bingjia, Chung, Yunsie, Yang, Archer Y., Yuan, Bo, Yu, Xiang
In drug discovery, in vitro and in vivo experiments reveal biochemical activities related to the efficacy and toxicity of compounds. The experimental data accumulate into massive, ever-evolving, and sparse datasets. Quantitative Structure-Activity Relationship (QSAR) models, which predict biochemical activities using only the structural information of compounds, face challenges in integrating the evolving experimental data as studies progress. We develop QSAR-Complete (QComp), a data completion framework to address this issue. Based on pre-existing QSAR models, QComp utilizes the correlation inherent in experimental data to enhance prediction accuracy across various tasks. Moreover, QComp emerges as a promising tool for guiding the optimal sequence of experiments by quantifying the reduction in statistical uncertainty for specific endpoints, thereby aiding in rational decision-making throughout the drug discovery process.
Learning Future Representation with Synthetic Observations for Sample-efficient Reinforcement Learning
Liu, Xin, Chen, Yaran, Zhao, Dongbin
In visual Reinforcement Learning (RL), upstream representation learning largely determines the effect of downstream policy learning. Employing auxiliary tasks allows the agent to enhance visual representation in a targeted manner, thereby improving the sample efficiency and performance of downstream RL. Prior advanced auxiliary tasks all focus on how to extract as much information as possible from limited experience (including observations, actions, and rewards) through their different auxiliary objectives, whereas in this article, we first start from another perspective: auxiliary training data. We try to improve auxiliary representation learning for RL by enriching auxiliary training data, proposing \textbf{L}earning \textbf{F}uture representation with \textbf{S}ynthetic observations \textbf{(LFS)}, a novel self-supervised RL approach. Specifically, we propose a training-free method to synthesize observations that may contain future information, as well as a data selection approach to eliminate unqualified synthetic noise. The remaining synthetic observations and real observations then serve as the auxiliary data to achieve a clustering-based temporal association task for representation learning. LFS allows the agent to access and learn observations that have not yet appeared in advance, so as to quickly understand and exploit them when they occur later. In addition, LFS does not rely on rewards or actions, which means it has a wider scope of application (e.g., learning from video) than recent advanced auxiliary tasks. Extensive experiments demonstrate that our LFS exhibits state-of-the-art RL sample efficiency on challenging continuous control and enables advanced visual pre-training based on action-free video demonstrations.
Approximation and Gradient Descent Training with Neural Networks
It is well understood that neural networks with carefully hand-picked weights provide powerful function approximation and that they can be successfully trained in over-parametrized regimes. Since over-parametrization ensures zero training error, these two theories are not immediately compatible. Recent work uses the smoothness that is required for approximation results to extend a neural tangent kernel (NTK) optimization argument to an under-parametrized regime and show direct approximation bounds for networks trained by gradient flow. Since gradient flow is only an idealization of a practical method, this paper establishes analogous results for networks trained by gradient descent.
Global Convergence of Decentralized Retraction-Free Optimization on the Stiefel Manifold
Sun, Youbang, Chen, Shixiang, Garcia, Alfredo, Shahrampour, Shahin
Many classical and modern machine learning algorithms require solving optimization tasks under orthogonal constraints. Solving these tasks often require calculating retraction-based gradient descent updates on the corresponding Riemannian manifold, which can be computationally expensive. Recently Ablin et al. proposed an infeasible retraction-free algorithm, which is significantly more efficient. In this paper, we study the decentralized non-convex optimization task over a network of agents on the Stiefel manifold with retraction-free updates. We propose \textbf{D}ecentralized \textbf{R}etraction-\textbf{F}ree \textbf{G}radient \textbf{T}racking (DRFGT) algorithm, and show that DRFGT exhibits ergodic $\mathcal{O}(1/K)$ convergence rate, the same rate of convergence as the centralized, retraction-based methods. We also provide numerical experiments demonstrating that DRFGT performs on par with the state-of-the-art retraction based methods with substantially reduced computational overhead.
Erasing the Bias: Fine-Tuning Foundation Models for Semi-Supervised Learning
Semi-supervised learning (SSL) has witnessed remarkable progress, resulting in the emergence of numerous method variations. However, practitioners often encounter challenges when attempting to deploy these methods due to their subpar performance. In this paper, we present a novel SSL approach named FineSSL that significantly addresses this limitation by adapting pre-trained foundation models. We identify the aggregated biases and cognitive deviation problems inherent in foundation models, and propose a simple yet effective solution by imposing balanced margin softmax and decoupled label smoothing. Through extensive experiments, we demonstrate that FineSSL sets a new state of the art for SSL on multiple benchmark datasets, reduces the training cost by over six times, and can seamlessly integrate various fine-tuning and modern SSL algorithms. The source code is available at https://github.com/Gank0078/FineSSL.
Optimization of Worker Scheduling at Logistics Depots Using Genetic Algorithms and Simulated Annealing
Xu, Jinxin, Wu, Haixin, Cheng, Yu, Wang, Liyang, Yang, Xin, Fu, Xintong, Su, Yuelong
The efficient scheduling of permanent and temporary workers is crucial for Improving the efficiency of sortation center management optimizing the efficiency of the logistics depot while has a direct impact on the fulfillment efficiency and minimizing labor usage. The study begins by establishing operational costs of the entire logistics network. Staff a 0-1 integer linear programming model, with decision management in sortation centers is a key challenge. Staffing needs to be adjusted according to the forecasted shipment variables determining the scheduling of permanent and volume to ensure a sufficient workforce to handle the flow of temporary workers for each time slot on a given day. The goods during peak hours while avoiding the wastage of excess objective function aims to minimize person-days, while manpower during low-demand times. Staff scheduling based constraints ensure fulfillment of hourly labor on effective solution algorithms becomes one of the key requirements, limit workers to one time slot per day, cap strategies to improve the efficiency of the sorting center. By consecutive working days for permanent workers, and reasonably allocating regular and temporary workers, the maintain non-negativity and integer constraints. The sorting speed and accuracy can be improved, thus reducing the model is then solved using genetic algorithms and overall logistics cost and improving customer satisfaction.
Conditionally-Conjugate Gaussian Process Factor Analysis for Spike Count Data via Data Augmentation
Nadew, Yididiya Y., Fan, Xuhui, Quinn, Christopher J.
Gaussian process factor analysis (GPFA) is a latent variable modeling technique commonly used to identify smooth, low-dimensional latent trajectories underlying high-dimensional neural recordings. Specifically, researchers model spiking rates as Gaussian observations, resulting in tractable inference. Recently, GPFA has been extended to model spike count data. However, due to the non-conjugacy of the likelihood, the inference becomes intractable. Prior works rely on either black-box inference techniques, numerical integration or polynomial approximations of the likelihood to handle intractability. To overcome this challenge, we propose a conditionally-conjugate Gaussian process factor analysis (ccGPFA) resulting in both analytically and computationally tractable inference for modeling neural activity from spike count data. In particular, we develop a novel data augmentation based method that renders the model conditionally conjugate. Consequently, our model enjoys the advantage of simple closed-form updates using a variational EM algorithm. Furthermore, due to its conditional conjugacy, we show our model can be readily scaled using sparse Gaussian Processes and accelerated inference via natural gradients. To validate our method, we empirically demonstrate its efficacy through experiments.
Degree of Irrationality: Sentiment and Implied Volatility Surface
As such, indicators in the options market, such as options prices, implied volatility, and the Greeks, are seen as "smarter" compared to indicators in the securities market. Numerous studies have confirmed this perspective and have explored the discovery function of options implied volatility on securities prices. For instance, Ni et al. (2020) found that the degree of skewness in implied volatility smiles has a significant predictive ability for stock market returns, while Han and Li (2021) discovered that the difference between call and put implied volatility has significant predictive power for stock market returns. Additionally, there is more research on the predictive ability of options implied volatility on realized volatility, dating back to Latane and Rendleman (1976-05) reverse use of the BS formula to derive the implied standard deviation of options and constructing a weighted implied standard deviation (WISD) using delta-neutral weighting, which was found to predict actual volatility significantly better than methods based on historical volatility. In recent years, numerous studies have incorporated the VIX index and the HAR method proposed by Corsi (2009), achieving notable results in predicting stock market volatility Byun and Kim (2013); Zhang (2020); Wan and Tian (2023). Preprint submitted to Elsarticle May 18, 2024 However, indicators in the options market should not be treated as the gold standard.