Statistical Learning
Reviews: Solving Most Systems of Random Quadratic Equations
This paper provides no theoretical guarantees of achieving the optimal number of measurements. Any observations of this fact are purely empirical. Empirical claims are fine with me, but the authors should be up-front about the fact that their results are only observed in experiments (the first time I read this paper the abstract and intro led me to believe you have a *proof* of achieving the lower bound, which you do not).
Reviews: Horizon-Independent Minimax Linear Regression
The problem of online linear regression is considered from an individual sequence perspective, where the aim is to control the square loss predictive regret with respect to the best linear predictor \theta \top x_t simultaneously for every sequence of covariate vectors x_t \in R d and outcomes y_t \in R in some constraint set. This is naturally formulated as a sequential game between the forecaster and an adversarial environment. In previous work [1], this problem was addressed in the "fixed-design" case, where the horizon T and the sequence of covariate vectors x_1 T is known in advance. The exact minimax strategy (MMS) was introduced and shown to be minimax optimal under natural constraint sets on the label sequence (such as ellipse-constrained labels). The MMS strategy consists in some form of least squares, but where the inverse cumulative covariance matrix \Pi_t {-1} is replaced by a shrunk version P_t that takes future instance into account.
Reviews: Introspective Classification with Convolutional Nets
The paper proposes a technique to improve the test accuracy of a discriminative model, by synthesizing additional negative input examples during the training process of the model. The negative example generation process has a Bayesian motivation, and is realized by "optimizing" for images (starting from random Gaussian noise) to maximize the probability of a given class label, a la DeepDream or Neural Artistic Style. These generated examples are added to the training set, and training is halted based on performance on a validation set. Experiments demonstrate that this procedure yields (very modest) improvements in test accuracy, and additionally provides some robustness against adversarial examples. The core idea is quite elegant, with an intuitive picture of using the "hard" negatives generated by the network to tighten the decision boundaries around the positive examples.
Systematic Feature Design for Cycle Life Prediction of Lithium-Ion Batteries During Formation
Rhyu, Jinwook, Schaeffer, Joachim, Li, Michael L., Cui, Xiao, Chueh, William C., Bazant, Martin Z., Braatz, Richard D.
Accurate lifetime prediction of lithium-ion batteries accelerates battery optimization and improves safety [1-4]. Although this task is challenging due to complicated and convolved degradation mechanisms, various studies have demonstrated the potential in using data-driven approaches [5-13], physics-based approaches [14-18], and hybrid approaches [19-26]. For accurate battery health monitoring, diagnostic techniques such as Differential Voltage Fitting (DVF) [27-30], Incremental Capacity Analysis (ICA) [31, 32], Electrochemical Impedance Spectroscopy (EIS) [10, 33-35], and Hybrid Pulse Power Characterization (HPPC) [36, 37] were developed for physics-based feature extraction during battery operation. Further optimization of these diagnostic techniques includes novel State of Health (SoH) feature development [38-41] and diagnostic time reduction [42, 43]. Compared to the extensive research on lifetime prediction during operation, there have been few studies on lifetime prediction during the manufacturing process (i.e., extreme early cycle life prediction) because of the limited availability of public manufacturing data. In fact, the cycle life can vary greatly based on the protocol used during formation, in which a passivation layer of Solid Electrolyte Interphase (SEI) is rapidly formed on the anode to limit further degradation during use. For example, Weng et al. [44] showed that the Nickel Manganese Cobalt (NMC)/graphite pouch cells with the fast formation protocol proposed by Wood et al. [45, 46] had in average 25% longer cycle lives than the pouch cells with a baseline formation protocol when aging the cells in both room temperature and high-temperature (45
Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging
Wu, Zihan, Huang, Zhaoke, Yan, Hong
Co-clustering simultaneously clusters rows and columns, revealing more fine-grained groups. However, existing co-clustering methods suffer from poor scalability and cannot handle large-scale data. This paper presents a novel and scalable co-clustering method designed to uncover intricate patterns in high-dimensional, large-scale datasets. Specifically, we first propose a large matrix partitioning algorithm that partitions a large matrix into smaller submatrices, enabling parallel co-clustering. This method employs a probabilistic model to optimize the configuration of submatrices, balancing the computational efficiency and depth of analysis. Additionally, we propose a hierarchical co-cluster merging algorithm that efficiently identifies and merges co-clusters from these submatrices, enhancing the robustness and reliability of the process. Extensive evaluations validate the effectiveness and efficiency of our method. Experimental results demonstrate a significant reduction in computation time, with an approximate 83% decrease for dense matrices and up to 30% for sparse matrices.
Efficient Adaptive Federated Optimization
Lee, Su Hyeong, Sharma, Sidharth, Zaheer, Manzil, Li, Tian
Adaptive optimization plays a pivotal role in federated learning, where simultaneous server and client-side adaptivity have been shown to be essential for maximizing its performance. However, the scalability of jointly adaptive systems is often constrained by limited resources in communication and memory. In this paper, we introduce a class of efficient adaptive algorithms, named $FedAda^2$, designed specifically for large-scale, cross-device federated environments. $FedAda^2$ optimizes communication efficiency by avoiding the transfer of preconditioners between the server and clients. At the same time, it leverages memory-efficient adaptive optimizers on the client-side to reduce on-device memory consumption. Theoretically, we demonstrate that $FedAda^2$ achieves the same convergence rates for general, non-convex objectives as its more resource-intensive counterparts that directly integrate joint adaptivity. Empirically, we showcase the benefits of joint adaptivity and the effectiveness of $FedAda^2$ on both image and text datasets.
rECGnition_v1.0: Arrhythmia detection using cardiologist-inspired multi-modal architecture incorporating demographic attributes in ECG
Srivastava, Shreya, Kumar, Durgesh, Bedi, Jatin, Seth, Sandeep, Sharma, Deepak
A substantial amount of variability in ECG manifested due to patient characteristics hinders the adoption of automated analysis algorithms in clinical practice. None of the ECG annotators developed till date consider the characteristics of the patients in a multi-modal architecture. We employed the XGBoost model to analyze the UCI Arrhythmia dataset, linking patient characteristics to ECG morphological changes. The model accurately classified patient gender using discriminative ECG features with 87.75% confidence. We propose a novel multi-modal methodology for ECG analysis and arrhythmia classification that can help defy the variability in ECG related to patient-specific conditions. This deep learning algorithm, named rECGnition_v1.0 (robust ECG abnormality detection Version 1), fuses Beat Morphology with Patient Characteristics to create a discriminative feature map that understands the internal correlation between both modalities. A Squeeze and Excitation based Patient characteristic Encoding Network (SEPcEnet) has been introduced, considering the patient's demographics. The trained model outperformed the various existing algorithms by achieving the overall F1-score of 0.986 for the ten arrhythmia class classification in the MITDB and achieved near perfect prediction scores of ~0.99 for LBBB, RBBB, Premature ventricular contraction beat, Atrial premature beat and Paced beat. Subsequently, the methodology was validated across INCARTDB, EDB and different class groups of MITDB using transfer learning. The generalizability test provided F1-scores of 0.980, 0.946, 0.977, and 0.980 for INCARTDB, EDB, MITDB AAMI, and MITDB Normal vs. Abnormal Classification, respectively. Therefore, with a more enhanced and comprehensive understanding of the patient being examined and their ECG for diverse CVD manifestations, the proposed rECGnition_v1.0 algorithm paves the way for its deployment in clinics.
Estimating Exoplanet Mass using Machine Learning on Incomplete Datasets
Lalande, Florian, Tasker, Elizabeth, Doya, Kenji
The exoplanet archive is an incredible resource of information on the properties of discovered extrasolar planets, but statistical analysis has been limited by the number of missing values. One of the most informative bulk properties is planet mass, which is particularly challenging to measure with more than 70\% of discovered planets with no measured value. We compare the capabilities of five different machine learning algorithms that can utilize multidimensional incomplete datasets to estimate missing properties for imputing planet mass. The results are compared when using a partial subset of the archive with a complete set of six planet properties, and where all planet discoveries are leveraged in an incomplete set of six and eight planet properties. We find that imputation results improve with more data even when the additional data is incomplete, and allows a mass prediction for any planet regardless of which properties are known. Our favored algorithm is the newly developed $k$NN$\times$KDE, which can return a probability distribution for the imputed properties. The shape of this distribution can indicate the algorithm's level of confidence, and also inform on the underlying demographics of the exoplanet population. We demonstrate how the distributions can be interpreted with a series of examples for planets where the discovery was made with either the transit method, or radial velocity method. Finally, we test the generative capability of the $k$NN$\times$KDE to create a large synthetic population of planets based on the archive, and identify potential categories of planets from groups of properties in the multidimensional space. All codes are Open Source.
Statistical Arbitrage in Rank Space
In equity markets, stocks are conventionally labeled by equity indices (company names). By relabeling stocks according to their ranks in capitalization, rather than their equity indices (company names), a different, more stable market structure can emerge. Specifically, we will gain a different perspective on market dynamics by focusing on the stock that occupies a certain rank in capitalization while the corresponding company name may change. We refer to a market labeled by the equity indices (company names) as a market in name space and one labeled by ranks in capitalization as a market in rank space . Market in rank space was explored by Fernholtz et al. who observed a stable distribution of capitalization across different ranks in the U.S. equity market over different time periods [11,16]. They further introduced an explanatory hybrid-Atlas model under stochastic portfolio theory, a framework that enables analyzing portfolios in rank space [5,15]. Empirically, B. Healy et al. analyzed the U.S. equity data and showed that the market in rank space is driven by a dominant single factor [14], in contrast to the multi-factor-driven market in name space [9,10,19]. While the primary market factor in rank space has been extensively studied, the residual returns - those not explained by this primary factor in stock returns - remain a fertile land of adventure.
Generalization Ability Analysis of Through-the-Wall Radar Human Activity Recognition
Gao, Weicheng, Qu, Xiaodong, Yang, Xiaopeng
Through-the-Wall radar (TWR) human activity recognition (HAR) is a technology that uses low-frequency ultra-wideband (UWB) signal to detect and analyze indoor human motion. However, the high dependence of existing end-to-end recognition models on the distribution of TWR training data makes it difficult to achieve good generalization across different indoor testers. In this regard, the generalization ability of TWR HAR is analyzed in this paper. In detail, an end-to-end linear neural network method for TWR HAR and its generalization error bound are first discussed. Second, a micro-Doppler corner representation method and the change of the generalization error before and after dimension reduction are presented. The appropriateness of the theoretical generalization errors is proved through numerical simulations and experiments. The results demonstrate that feature dimension reduction is effective in allowing recognition models to generalize across different indoor testers.