Goto

Collaborating Authors

 Statistical Learning


CBIDR: A novel method for information retrieval combining image and data by means of TOPSIS applied to medical diagnosis

arXiv.org Artificial Intelligence

Content-Based Image Retrieval (CBIR) have shown promising results in the field of medical diagnosis, which aims to provide support to medical professionals (doctor or pathologist). However, the ultimate decision regarding the diagnosis is made by the medical professional, drawing upon their accumulated experience. In this context, we believe that artificial intelligence can play a pivotal role in addressing the challenges in medical diagnosis not by making the final decision but by assisting in the diagnosis process with the most relevant information. The CBIR methods use similarity metrics to compare feature vectors generated from images using Convolutional Neural Networks (CNNs). In addition to the information contained in medical images, clinical data about the patient is often available and is also relevant in the final decision-making process by medical professionals. In this paper, we propose a novel method named CBIDR, which leverage both medical images and clinical data of patient, combining them through the ranking algorithm TOPSIS. The goal is to aid medical professionals in their final diagnosis by retrieving images and clinical data of patient that are most similar to query data from the database. As a case study, we illustrate our CBIDR for diagnostic of oral cancer including histopathological images and clinical data of patient. Experimental results in terms of accuracy achieved 97.44% in Top-1 and 100% in Top-5 showing the effectiveness of the proposed approach.


Predicting Muscle Thickness Deformation from Muscle Activation Patterns: A Dual-Attention Framework

arXiv.org Artificial Intelligence

Abstract-- Understanding the relationship between muscle activation and thickness deformation is critical for diagnosing muscle-related diseases and monitoring muscle health. Although ultrasound technique can measure muscle thickness change during muscle movement, its application in portable devices is limited by wiring and data collection challenges. Experimental results with six healthy subjects showed that the approach could accurately predict muscle excursion with an average precision of 0.923 0.900mm, which shows that this method can facilitate real-time portable muscle health monitoring, Our proposed method employs a novel dual-attention framework to correlate muscle activation with thickness I. INTRODUCTION This framework included hierarchical selfattention Quantifying the relationship between muscle activation [12] and cross-attention [13] mechanisms. Selfattention and thickness deformation is essential for understanding captured long-range signal dependencies and dynamically muscle dynamics and health [1], [2], particularly in conditions adjusted the importance of different signal components such as Facioscapulohumeral Dystrophy [3]. Traditional [14], while cross-attention merged and synthesized ultrasound imaging can visualize muscle thickness these features to provide comprehensive MTD information.


A Survey on Offensive AI Within Cybersecurity

arXiv.org Artificial Intelligence

As AI takes on pivotal roles in essential applications, like self-driving vehicles, healthcare diagnosis, and financial services, it becomes a tempting target for malicious actors [16]. This study aims to comprehensively explore the realm of offensive AI, shedding light on its multifaceted dimensions, the techniques involved, its consequences, and potential future implications. Cyberattacks have surged in both complexity and frequency. This is evidenced by the escalating costs associated with data breaches. In 2022, businesses incurred an average loss of $4.35 million, an increase of $0.11 million from the previous year and a 12.7% rise from 2020 [22]. Moreover, the volume of data breaches has reached historic highs, with approximately 15 million records exposed during the third quarter of 2022. Furthermore, the third quarter of 2022 witnessed an alarming 57,116 distributed denial-of-service (DDoS) attacks [78]. Against this backdrop, understanding and mitigating security risks in machine learning (ML) has emerged as a pivotal aspect of cybersecurity.


Defect Prediction with Content-based Features

arXiv.org Artificial Intelligence

Traditional defect prediction approaches often use metrics that measure the complexity of the design or implementing code of a software system, such as the number of lines of code in a source file. In this paper, we explore a different approach based on content of source code. Our key assumption is that source code of a software system contains information about its technical aspects and those aspects might have different levels of defect-proneness. Thus, content-based features such as words, topics, data types, and package names extracted from a source code file could be used to predict its defects. We have performed an extensive empirical evaluation and found that: i) such content-based features have higher predictive power than code complexity metrics and ii) the use of feature selection, reduction, and combination further improves the prediction performance.


Stable Object Placement Under Geometric Uncertainty via Differentiable Contact Dynamics

arXiv.org Artificial Intelligence

From serving a cup of coffee to carefully rearranging delicate items, stable object placement is a crucial skill for future robots. This skill is challenging due to the required accuracy, which is difficult to achieve under geometric uncertainty. We leverage differentiable contact dynamics to develop a principled method for stable object placement under geometric uncertainty. We estimate the geometric uncertainty by minimizing the discrepancy between the force-torque sensor readings and the model predictions through gradient descent. We further keep track of a belief over multiple possible geometric parameters to mitigate the gradient-based method's sensitivity to the initialization. We verify our approach in the real world on various geometric uncertainties, including the in-hand pose uncertainty of the grasped object, the object's shape uncertainty, and the environment's shape uncertainty.


FedDCL: a federated data collaboration learning as a hybrid-type privacy-preserving framework based on federated learning and data collaboration

arXiv.org Artificial Intelligence

Recently, federated learning has attracted much attention as a privacy-preserving integrated analysis that enables integrated analysis of data held by multiple institutions without sharing raw data. On the other hand, federated learning requires iterative communication across institutions and has a big challenge for implementation in situations where continuous communication with the outside world is extremely difficult. In this study, we propose a federated data collaboration learning (FedDCL), which solves such communication issues by combining federated learning with recently proposed non-model share-type federated learning named as data collaboration analysis. In the proposed FedDCL framework, each user institution independently constructs dimensionality-reduced intermediate representations and shares them with neighboring institutions on intra-group DC servers. On each intra-group DC server, intermediate representations are transformed to incorporable forms called collaboration representations. Federated learning is then conducted between intra-group DC servers. The proposed FedDCL framework does not require iterative communication by user institutions and can be implemented in situations where continuous communication with the outside world is extremely difficult. The experimental results show that the performance of the proposed FedDCL is comparable to that of existing federated learning.


Embed and Emulate: Contrastive representations for simulation-based inference

arXiv.org Machine Learning

Scientific modeling and engineering applications rely heavily on parameter estimation methods to fit physical models and calibrate numerical simulations using real-world measurements. In the absence of analytic statistical models with tractable likelihoods, modern simulation-based inference (SBI) methods first use a numerical simulator to generate a dataset of parameters and simulated outputs. This dataset is then used to approximate the likelihood and estimate the system parameters given observation data. Several SBI methods employ machine learning emulators to accelerate data generation and parameter estimation. However, applying these approaches to high-dimensional physical systems remains challenging due to the cost and complexity of training high-dimensional emulators. This paper introduces Embed and Emulate (E&E): a new SBI method based on contrastive learning that efficiently handles high-dimensional data and complex, multimodal parameter posteriors. E&E learns a low-dimensional latent embedding of the data (i.e., a summary statistic) and a corresponding fast emulator in the latent space, eliminating the need to run expensive simulations or a high dimensional emulator during inference. We illustrate the theoretical properties of the learned latent space through a synthetic experiment and demonstrate superior performance over existing methods in a realistic, non-identifiable parameter estimation task using the high-dimensional, chaotic Lorenz 96 system.


Benchmarking Graph Conformal Prediction: Empirical Analysis, Scalability, and Theoretical Insights

arXiv.org Machine Learning

Modern machine learning models trained on losses based on point predictions are prone to be overconfident in their predictions [Guo et al., 2017]. The Conformal Prediction (CP) framework [Vovk et al., 2005] provides a mechanism for generating statistically sound post hoc prediction sets (or intervals, in case of continuous outcomes) with coverage guarantees under mild assumptions. The usual assumption made in CP is that data are exchangeable, i.e., the joint distribution of the data is invariant to permutations of the data points. CP's guarantees are distribution-free and can be added post hoc to arbitrary black-box predictor scores, making them ideal candidates for quantifying uncertainty in complex models, such as neural networks. Network-structured data such as social networks, transportation networks, and biological networks are ubiquitous in modern data science applications.


Local Prediction-Powered Inference

arXiv.org Machine Learning

To infer a function value on a specific point $x$, it is essential to assign higher weights to the points closer to $x$, which is called local polynomial / multivariable regression. In many practical cases, a limited sample size may ruin this method, but such conditions can be improved by the Prediction-Powered Inference (PPI) technique. This paper introduced a specific algorithm for local multivariable regression using PPI, which can significantly reduce the variance of estimations without enlarge the error. The confidence intervals, bias correction, and coverage probabilities are analyzed and proved the correctness and superiority of our algorithm. Numerical simulation and real-data experiments are applied and show these conclusions. Another contribution compared to PPI is the theoretical computation efficiency and explainability by taking into account the dependency of the dependent variable.


Using dynamic loss weighting to boost improvements in forecast stability

arXiv.org Machine Learning

Rolling origin forecast instability refers to variability in forecasts for a specific period induced by updating the forecast when new data points become available. Recently, an extension to the N-BEATS model for univariate time series point forecasting was proposed to include forecast stability as an additional optimization objective, next to accuracy. It was shown that more stable forecasts can be obtained without harming accuracy by minimizing a composite loss function that contains both a forecast error and a forecast instability component, with a static hyperparameter to control the impact of stability. In this paper, we empirically investigate whether further improvements in stability can be obtained without compromising accuracy by applying dynamic loss weighting algorithms, which change the loss weights during training. We show that some existing dynamic loss weighting methods achieve this objective. However, our proposed extension to the Random Weighting approach -- Task-Aware Random Weighting -- shows the best performance.