Goto

Collaborating Authors

 Statistical Learning


Mind the Gap! Static and Interactive Evaluations of Large Audio Models

arXiv.org Artificial Intelligence

As AI chatbots become ubiquitous, voice interaction presents a compelling way to enable rapid, high-bandwidth communication for both semantic and social signals. This has driven research into Large Audio Models (LAMs) to power voice-native experiences. However, aligning LAM development with user goals requires a clear understanding of user needs and preferences to establish reliable progress metrics. This study addresses these challenges by introducing an interactive approach to evaluate LAMs and collecting 7,500 LAM interactions from 484 participants. Through topic modeling of user queries, we identify primary use cases for audio interfaces. We then analyze user preference rankings and qualitative feedback to determine which models best align with user needs. Finally, we evaluate how static benchmarks predict interactive performance - our analysis reveals no individual benchmark strongly correlates with interactive results ($\tau \leq 0.33$ for all benchmarks). While combining multiple coarse-grained features yields modest predictive power ($R^2$=$0.30$), only two out of twenty datasets on spoken question answering and age prediction show significantly positive correlations. This suggests a clear need to develop LAM evaluations that better correlate with user preferences.


Connecting the geometry and dynamics of many-body complex systems with message passing neural operators

arXiv.org Artificial Intelligence

The relationship between scale transformations and dynamics established by renormalization group techniques is a cornerstone of modern physical theories, from fluid mechanics to elementary particle physics. Integrating renormalization group methods into neural operators for many-body complex systems could provide a foundational inductive bias for learning their effective dynamics, while also uncovering multiscale organization. We introduce a scalable AI framework, ROMA (Renormalized Operators with Multiscale Attention), for learning multiscale evolution operators of many-body complex systems. In particular, we develop a renormalization procedure based on neural analogs of the geometric and laplacian renormalization groups, which can be co-learned with neural operators. An attention mechanism is used to model multiscale interactions by connecting geometric representations of local subgraphs and dynamical operators. We apply this framework in challenging conditions: large systems of more than 1M nodes, long-range interactions, and noisy input-output data for two contrasting examples: Kuramoto oscillators and Burgers-like social dynamics. We demonstrate that the ROMA framework improves scalability and positive transfer between forecasting and effective dynamics tasks compared to state-of-the-art operator learning techniques, while also giving insight into multiscale interactions. Additionally, we investigate power law scaling in the number of model parameters, and demonstrate a departure from typical power law exponents in the presence of hierarchical and multiscale interactions.


ML-Driven Approaches to Combat Medicare Fraud: Advances in Class Imbalance Solutions, Feature Engineering, Adaptive Learning, and Business Impact

arXiv.org Artificial Intelligence

Medicare fraud poses a substantial challenge to healthcare systems, resulting in significant financial losses and undermining the quality of care provided to legitimate beneficiaries. This study investigates the use of machine learning (ML) to enhance Medicare fraud detection, addressing key challenges such as class imbalance, high-dimensional data, and evolving fraud patterns. A dataset comprising inpatient claims, outpatient claims, and beneficiary details was used to train and evaluate five ML models: Random Forest, KNN, LDA, Decision Tree, and AdaBoost. Data preprocessing techniques included resampling SMOTE method to address the class imbalance, feature selection for dimensionality reduction, and aggregation of diagnostic and procedural codes. Random Forest emerged as the best-performing model, achieving a training accuracy of 99.2% and validation accuracy of 98.8%, and F1-score (98.4%). The Decision Tree also performed well, achieving a validation accuracy of 96.3%. KNN and AdaBoost demonstrated moderate performance, with validation accuracies of 79.2% and 81.1%, respectively, while LDA struggled with a validation accuracy of 63.3% and a low recall of 16.6%. The results highlight the importance of advanced resampling techniques, feature engineering, and adaptive learning in detecting Medicare fraud effectively. This study underscores the potential of machine learning in addressing the complexities of fraud detection. Future work should explore explainable AI and hybrid models to improve interpretability and performance, ensuring scalable and reliable fraud detection systems that protect healthcare resources and beneficiaries.


Benchmarking machine learning for bowel sound pattern classification from tabular features to pretrained models

arXiv.org Artificial Intelligence

The development of electronic stethoscopes and wearable recording sensors opened the door to the automated analysis of bowel sound (BS) signals. This enables a data-driven analysis of bowel sound patterns, their interrelations, and their correlation to different pathologies. This work leverages a BS dataset collected from 16 healthy subjects that was annotated according to four established BS patterns. This dataset is used to evaluate the performance of machine learning models to detect and/or classify BS patterns. The selection of considered models covers models using tabular features, convolutional neural networks based on spectrograms and models pre-trained on large audio datasets. The results highlight the clear superiority of pre-trained models, particularly in detecting classes with few samples, achieving an AUC of 0.89 in distinguishing BS from non-BS using a HuBERT model and an AUC of 0.89 in differentiating bowel sound patterns using a Wav2Vec 2.0 model. These results pave the way for an improved understanding of bowel sounds in general and future machine-learning-driven diagnostic applications for gastrointestinal examinations


A Defensive Framework Against Adversarial Attacks on Machine Learning-Based Network Intrusion Detection Systems

arXiv.org Artificial Intelligence

As cyberattacks become increasingly sophisticated, advanced Network Intrusion Detection Systems (NIDS) are critical for modern network security. Traditional signature-based NIDS are inadequate against zero-day and evolving attacks. In response, machine learning (ML)-based NIDS have emerged as promising solutions; however, they are vulnerable to adversarial evasion attacks that subtly manipulate network traffic to bypass detection. To address this vulnerability, we propose a novel defensive framework that enhances the robustness of ML-based NIDS by simultaneously integrating adversarial training, dataset balancing techniques, advanced feature engineering, ensemble learning, and extensive model fine-tuning. We validate our framework using the NSL-KDD and UNSW-NB15 datasets. Experimental results show, on average, a 35% increase in detection accuracy and a 12.5% reduction in false positives compared to baseline models, particularly under adversarial conditions. The proposed defense against adversarial attacks significantly advances the practical deployment of robust ML-based NIDS in real-world networks.


Zweistein: A Dynamic Programming Evaluation Function for Einstein W\"urfelt Nicht!

arXiv.org Artificial Intelligence

This paper introduces Zweistein, a dynamic programming evaluation function for Einstein W\"urfelt Nicht! (EWN). Instead of relying on human knowledge to craft an evaluation function, Zweistein uses a data-centric approach that eliminates the need for parameter tuning. The idea is to use a vector recording the distance to the corner of all pieces. This distance vector captures the essence of EWN. It not only outperforms many traditional EWN evaluation functions but also won first place in the TCGA 2023 competition.


Verification and Validation for Trustworthy Scientific Machine Learning

arXiv.org Artificial Intelligence

Scientific machine learning (SciML) integrates machine learning (ML) into scientific workflows to enhance system simulation and analysis, with an emphasis on computational modeling of physical systems. This field emerged from Department of Energy workshops and initiatives starting in 2018, which also identified the need to increase "the scale, rigor, robustness, and reliability of SciML necessary for routine use in science and engineering applications" [5]. The field's subsequent growth through funding initiatives, conference themes, and high-profile publications stems from its ability to unite ML's predictive power with the domain knowledge and mathematical rigor of computational science and engineering (CSE). However, this surge in SciML development has outpaced good practices and reporting standards for building trust [66, 51, 109, 117]. SciML models must demonstrate trustworthiness to be safe and useful [44]. Organizational and computational trust definitions [92, 106] inform our criteria for trustworthy SciML: competence in basic performance, reliability across conditions, transparency about processes and limitations, and alignment with scientific objectives. These criteria span technical attributes (correctness, reliability, safety) and human-centric qualities (comprehensibility, transparency).


Network Resource Optimization for ML-Based UAV Condition Monitoring with Vibration Analysis

arXiv.org Artificial Intelligence

ACCEPTED IN: IEEE NETWORKING LETTERS 1 Network Resource Optimization for ML-Based UA V Condition Monitoring with Vibration Analysis Alexandre Gemayel, Dimitrios Michael Manias, and Abdallah Shami Abstract --As smart cities begin to materialize, the role of Unmanned Aerial V ehicles (UA Vs) and their reliability becomes increasingly important. One aspect of reliability relates to Condition Monitoring (CM), where Machine Learning (ML) models are leveraged to identify abnormal and adverse conditions. Given the resource-constrained nature of next-generation edge networks, the utilization of precious network resources must be minimized. This work explores the optimization of network resources for ML-based UA V CM frameworks. The developed framework uses experimental data and varies the feature extraction aggregation interval to optimize ML model selection. Additionally, by leveraging dimensionality reduction techniques, there is a 99.9% reduction in network resource consumption. I NTRODUCTION E MERGING Unmanned Aerial V ehicle (UA V) applications, such as Smart Cities, have highlighted the necessity of real-time Condition Monitoring (CM) through Anomaly Detection (AD) and health analytics to ensure operational safety and integrity [1].


Mitigating Data Scarcity in Time Series Analysis: A Foundation Model with Series-Symbol Data Generation

arXiv.org Artificial Intelligence

Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as data scarcity and data imbalance continue to hinder their development. To address this, we consider modeling complex systems through symbolic expressions that serve as semantic descriptors of time series. Building on this concept, we introduce a series-symbol (S2) dual-modulity data generation mechanism, enabling the unrestricted creation of high-quality time series data paired with corresponding symbolic representations. Leveraging the S2 dataset, we develop SymTime, a pre-trained foundation model for TSA. SymTime demonstrates competitive performance across five major TSA tasks when fine-tuned with downstream task, rivaling foundation models pre-trained on real-world datasets. This approach underscores the potential of dual-modality data generation and pretraining mechanisms in overcoming data scarcity and enhancing task performance.


Road Traffic Sign Recognition method using Siamese network Combining Efficient-CNN based Encoder

arXiv.org Artificial Intelligence

IEEE TRANSACTIONS ON INTELLIGENT TRANSPORT A TION SYSTEMS 1 Road Traffic Sign Recognition Method Using Siamese Network Combining Efficient-CNN-Based Encoder Zhenghao Xi, Member, IEEE, Y uchao Shao, Y ang Zheng, Member, IEEE, Xiang Liu, Member, IEEE, Y aqi Liu, and Yitong Cai Abstract -- Traffic signs recognition (TSR) plays an essential role in assistant driving and intelligent transportation system. However, the noise of complex environment may lead to motion-blur or occlusion problems, which raise the tough challenge to real-time recognition with high accuracy and robust. In this article, we propose IECES-network which with improved encoders and Siamese net. The three-stage approach of our method includes Efficient-CNN based encoders, Siamese backbone and the fully-connected layers. We firstly use convolu-tional encoders to extract and encode the traffic sign features of augmented training samples and standard images. Then, we design the Siamese neural network with Efficient-CNN based encoder and contrastive loss function, which can be trained to improve the robustness of TSR problem when facing the samples of motion-blur and occlusion by computing the distance between inputs and templates. Additionally, the template branch of the proposed network can be stopped when executing the recognition tasks after training to raise the process speed of our real-time model, and alleviate the computational resource and parameter scale. Finally, we recombined the feature code and a fully-connected layer with SoftMax function to classify the codes of samples and recognize the category of traffic signs. The results of experiments on the Tsinghua-T encent 100K dataset and the German Traffic Sign Recognition Benchmark dataset demonstrate the performance of the proposed IECES-network. Compared with other state-of-the-art methods, in the case of motion-blur and occluded environment, the proposed method achieves competitive performance precision-recall and accuracy metric average is 88.1%, 86.43% and 86.1% with a 2.9M lightweight scale, respectively. Moreover, processing time of our model is 0.1s per frame, of which the speed is increased by 1.5 times compared with existing methods. Index T erms-- Traffic signs recognition, Siamese network, efficient-CNN based encoder . Received 11 September 2024; revised 25 November 2024; accepted 9 January 2025.