Goto

Collaborating Authors

 Statistical Learning


Kernel-Based Enhanced Oversampling Method for Imbalanced Classification

arXiv.org Artificial Intelligence

Wenjie LI 1, 2, Sibo Zhu 1, 2, Zhijian Li 1, 2, and Hanlin Wang 1, 2 Abstract -- This paper introduces a novel oversampling technique designed to improve classification performance on imbalanced datasets. The proposed method enhances the traditional SMOTE algorithm by incorporating convex combination and kernel-based weighting to generate synthetic samples that better represent the minority class. Through experiments on multiple real-world datasets, we demonstrate that the new technique outperforms existing methods in terms of F1-score, G-mean, and AUC, providing a robust solution for handling imbalanced datasets in classification tasks. I NTRODUCTION Imbalanced datasets are a pervasive issue in the domain of classification, where the distribution of classes is skewed, with one class (often referred to as the minority class) being significantly underrepresented compared to the other (the majority class). The imbalance issue is especially problematic in classification tasks, as traditional machine learning algorithms are generally designed to maximize overall accuracy, leading them to favor the majority class. Consequently, it results in a bias where the model performs well on the majority class but poorly on the minority class, which is often the class of greater interest [1].


A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences

arXiv.org Artificial Intelligence

This paper describes a practical approach of using supervised machine learning (ML) models to assist safety investigators to classify aviation occurrences into either incident or serious incident categories. Our implementation currently deployed as a ML web application is trained on a labelled dataset derived from publicly available aviation investigation reports. A selection of five supervised learning models (Support Vector Machine, Logistic Regression, Random Forest Classifier, XGBoost and K-Nearest Neighbors) were evaluated. This paper showed the best performing ML algorithm was the Random Forest Classifier with accuracy = 0.77, F1 Score = 0.78 and MCC = 0.51 (average of 100 sample runs). The study had also explored the effect of applying Synthetic Minority Over-sampling Technique (SMOTE) to the imbalanced dataset, and the overall observation ranged from no significant effect to substantial degradation in performance for some of the models after the SMOTE adjustment.


MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer

arXiv.org Artificial Intelligence

Generative masked transformers have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying generative masked modeling to generate diverse instances from a single MoCap reference may lead to overfitting, a challenge that remains unexplored. In this work, we present Motion-Dreamer, a localized masked modeling paradigm designed to learn internal motion patterns from a given motion with arbitrary topology and duration. By embedding the given motion into quantized tokens with a novel distribution regularization method, MotionDreamer constructs a robust and informative codebook for local motion patterns. Moreover, a sliding window local attention is introduced in our masked transformer, enabling the generation of natural yet diverse animations that closely resemble the reference motion patterns. As demonstrated through comprehensive experiments, MotionDreamer outperforms the state-of-the-art methods that are typically GAN or Diffusion-based in both faithfulness and diversity. Thanks to the consistency and robustness of the quantization-based approach, MotionDreamer can also effectively perform downstream tasks such as temporal motion editing, crowd animation, and beat-aligned dance generation, all using a single reference motion. Motions could be roughly interpreted as coherent and natural compositions of finite internal patterns. For example, a breaking can be performed by a freestyle composition of breaking dance skills, such as baby freeze, helicopter, kick up, back spin, etc. Learning these internal patterns from a single reference motion allows for the generation of diverse yet consistent motions that closely resemble the reference. This is particularly useful when data is scarce (e.g., in the case of animal motion) or when the content needs to be constrained.


Forecasting Cryptocurrency Prices using Contextual ES-adRNN with Exogenous Variables

arXiv.org Artificial Intelligence

In this paper, we introduce a new approach to multivariate forecasting cryptocurrency prices using a hybrid contextual model combining exponential smoothing (ES) and recurrent neural network (RNN). The model consists of two tracks: the context track and the main track. The context track provides additional information to the main track, extracted from representative series. This information as well as information extracted from exogenous variables is dynamically adjusted to the individual series forecasted by the main track. The RNN stacked architecture with hierarchical dilations, incorporating recently developed attentive dilated recurrent cells, allows the model to capture short and long-term dependencies across time series and dynamically weight input information. The model generates both point daily forecasts and predictive intervals for one-day, one-week and four-week horizons. We apply our model to forecast prices of 15 cryptocurrencies based on 17 input variables and compare its performance with that of comparative models, including both statistical and ML ones.


Combining Forecasts using Meta-Learning: A Comparative Study for Complex Seasonality

arXiv.org Artificial Intelligence

Abstract--In this paper, we investigate meta-learning for combining forecasts generated by models of different types . While typical approaches for combining forecasts involve s imple averaging, machine learning techniques enable more sophis ti-cated methods of combining through meta-learning, leading to improved forecasting accuracy. We use linear regression, k - nearest neighbors, multilayer perceptron, random forest, and long short-term memory as meta-learners. We define global and local meta-learning variants for time series with compl ex seasonality and compare meta-learners on multiple forecas ting problems, demonstrating their superior performance compa red to simple averaging. Ensemble methods are widely recognized as a cornerstone of modern machine learning (ML) [1], commonly used for regression and classification problems. In addition, ensem bling has proven to be a highly effective approach for increasing the predictive power of forecasting models. The ensemble approach in forecasting, which involves combining the predictions of multiple models, can be justified for several reasons. First of all, it usually leads to increased accurac y. Ensemble models often outperform individual models, as the y leverage the strengths of different models and minimize the ir weaknesses. By combining diverse models, the ensemble can produce more accurate predictions by capturing a broader range of patterns and insights from the data. Ensembling als o allows for the incorporation of multiple drivers into the da ta generating process, mitigating uncertainties regarding m odel form and parameter specification [2].


Are We Merely Justifying Results ex Post Facto? Quantifying Explanatory Inversion in Post-Hoc Model Explanations

arXiv.org Artificial Intelligence

Post-hoc explanation methods provide interpretation by attributing predictions to input features. Natural explanations are expected to interpret how the inputs lead to the predictions. Thus, a fundamental question arises: Do these explanations unintentionally reverse the natural relationship between inputs and outputs? Specifically, are the explanations rationalizing predictions from the output rather than reflecting the true decision process? To investigate such explanatory inversion, we propose Inversion Quantification (IQ), a framework that quantifies the degree to which explanations rely on outputs and deviate from faithful input-output relationships. Using the framework, we demonstrate on synthetic datasets that widely used methods such as LIME and SHAP are prone to such inversion, particularly in the presence of spurious correlations, across tabular, image, and text domains. Finally, we propose Reproduce-by-Poking (RBP), a simple and model-agnostic enhancement to post-hoc explanation methods that integrates forward perturbation checks. We further show that under the IQ framework, RBP theoretically guarantees the mitigation of explanatory inversion. Empirically, for example, on the synthesized data, RBP can reduce the inversion by 1.8% on average across iconic post-hoc explanation approaches and domains.


Hybrid AI-Physical Modeling for Penetration Bias Correction in X-band InSAR DEMs: A Greenland Case Study

arXiv.org Artificial Intelligence

Digital elevation models derived from Interferometric Synthetic Aperture Radar (InSAR) data over glacial and snow-covered regions often exhibit systematic elevation errors, commonly termed "penetration bias. " W e leverage existing physics-based models and propose an integrated correction framework that combines parametric physical modeling with machine learning. W e evaluate the approach across three distinct training scenarios -- each defined by a different set of acquisition parameters -- to assess overall performance and the model's ability to generalize. Our experiments on Greenland's ice sheet using T anDEM-X data show that the proposed hybrid model corrections significantly reduce the mean and standard deviation of DEM errors compared to a purely physical modeling baseline. The hybrid framework also achieves significantly improved generalization than a pure ML approach when trained on data with limited diversity in acquisition parameters.


DataMap: A Portable Application for Visualizing High-Dimensional Data

arXiv.org Artificial Intelligence

Motivation: The visualization and analysis of high-dimensional data are essential in biomedical research. There is a need for secure, scalable, and reproducible tools to facilitate data exploration and interpretation. Results: We introduce DataMap, a browser-based application for visualization of high-dimensional data using heatmaps, principal component analysis (PCA), and t-distributed stochastic neighbor embedding (t-SNE). DataMap runs in the web browser, ensuring data privacy while eliminating the need for installation or a server. The application has an intuitive user interface for data transformation, annotation, and generation of reproducible R code. Availability and Implementation: Freely available as a GitHub page https://gexijin.github.io/datamap/. The source code can be found at https://github.com/gexijin/datamap, and can also be installed as an R package. Contact: Xijin.Ge@sdstate.ed


In almost all shallow analytic neural network optimization landscapes, efficient minimizers have strongly convex neighborhoods

arXiv.org Artificial Intelligence

Artificial neural networks (ANNs) define parametrized families of f unctions (the realization functions) whose definition is inspired by biological neural networks. Running optimization algorithms on these parametrized families (the t raining of neural networks) has proven to be very efficient in various mach ine learning tasks, including image recognition, natural language processing, a utonomous systems, protein folding, climate modelling. The preferred method for the training of artificial neural networ ks (ANNs) are Stochastic Gradient Descent (SGD) algorithms. The vanilla SGD a lgorithm was first applied in Rumelhart et al. [ 1986 ]. Today, variants such as momentum-based methods [ Polyak, 1964 ], AMSProp [ Hinton, 2012 ] and the Adam optimizer [ Kingma and Ba, 2015 ] are more commonly used. Generally, the efficiency of optimization algorithms is significantly affec ted by the structure of the optimization landscape. The smoothing of u pdates in the momentum approaches seem to help with saddle points and adapt ive methods like RMSProp and Adam seem to adjust learning rates better to n avigate complex landscapes effectively. Mathematically rigorous approaches often assume that the SGD sc heme converges to a (local) minimum with a strongly convex neighborhood ( meaning that the Hessian of the landscape is strictly positive definite) or t hat a Polyak-null Lojasiewicz inequality (in the strong sense with exponent 2) applies.


A Nonlinear Hash-based Optimization Method for SpMV on GPUs

arXiv.org Artificial Intelligence

A Nonlinear Hash-based Optimization Method for SpMV on GPUs Chen Y an a,b, Boyu Diao a,b, Hangda Liu a,b, Zhulin An a,b and Y ongjun Xu a,b a Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China b University of Chinese Academy of Sciences, Beijing, China {yanchen23s, diaoboyu2012, liuhangda21s, anzhulin, xyj } @ict.ac.cn Abstract --Sparse matrix-vector multiplication (SpMV) is a fundamental operation with a wide range of applications in scientific computing and artificial intelligence. However, the large scale and sparsity of sparse matrix often make it a performance bottleneck. In this paper, we highlight the effectiveness of hash-based techniques in optimizing sparse matrix reordering, introducing the Hash-based Partition (HBP) format, a lightweight SpMV approach. HBP retains the performance benefits of the 2D-partitioning method while leveraging the hash transformation's ability to group similar elements, thereby accelerating the pre-processing phase of sparse matrix reordering. Additionally, we achieve parallel load balancing across matrix blocks through a competitive method. Our experiments, conducted on both Nvidia Jetson AGX Orin and Nvidia RTX 4090, show that in the pre-processing step, our method offers an average speedup of 3.53 times compared to the sorting approach and 3.67 times compared to the dynamic programming method employed in Regu2D. Furthermore, in SpMV, our method achieves a maximum speedup of 3.32 times on Orin and 3.01 times on RTX4090 against the CSR format in sparse matrices from the University of Florida Sparse Matrix Collection. I NTRODUCTION Sparse matrix-vector multiplication (SpMV) has a wide range of applications, such as mathematical solutions for sparse linear equations [13], iterative algorithm-solving processing [15] [25], graph processing [9] [14] [24], and weight calculations for forward and backward propagation in neural networks [3] [12] [17] [19], etc. However, SpMV is actually the bottleneck for many algorithms. The sparse matrix used in SpMV has the following characteristics [4]: (1) Sparsity. On the one hand, sparse matrices contain a large number of zero elements.