Goto

Collaborating Authors

 Statistical Learning


Predicting into unknown space? Estimating the area of applicability of spatial prediction models

arXiv.org Machine Learning

Predictive modelling using machine learning has become very popular for spatial mapping of the environment. Models are often applied to make predictions far beyond sampling locations where new geographic locations might considerably differ from the training data in their environmental properties. However, areas in the predictor space without support of training data are problematic. Since the model has no knowledge about these environments, predictions have to be considered uncertain. Estimating the area to which a prediction model can be reliably applied is required. Here, we suggest a methodology that delineates the "area of applicability" (AOA) that we define as the area, for which the cross-validation error of the model applies. We first propose a "dissimilarity index" (DI) that is based on the minimum distance to the training data in the predictor space, with predictors being weighted by their respective importance in the model. The AOA is then derived by applying a threshold based on the DI of the training data where the DI is calculated with respect to the cross-validation strategy used for model training. We test for the ideal threshold by using simulated data and compare the prediction error within the AOA with the cross-validation error of the model. We illustrate the approach using a simulated case study. Our simulation study suggests a threshold on DI to define the AOA at the .95 quantile of the DI in the training data. Using this threshold, the prediction error within the AOA is comparable to the cross-validation RMSE of the model, while the cross-validation error does not apply outside the AOA. This applies to models being trained with randomly distributed training data, as well as when training data are clustered in space and where spatial cross-validation is applied. We suggest to report the AOA alongside predictions, complementary to validation measures.


Differentially Private ADMM for Convex Distributed Learning: Improved Accuracy via Multi-Step Approximation

arXiv.org Machine Learning

Alternating Direction Method of Multipliers (ADMM) is a popular algorithm for distributed learning, where a network of nodes collaboratively solve a regularized empirical risk minimization by iterative local computation associated with distributed data and iterate exchanges. When the training data is sensitive, the exchanged iterates will cause serious privacy concern. In this paper, we aim to propose a new differentially private distributed ADMM algorithm with improved accuracy for a wide range of convex learning problems. In our proposed algorithm, we adopt the approximation of the objective function in the local computation to introduce calibrated noise into iterate updates robustly, and allow multiple primal variable updates per node in each iteration. Our theoretical results demonstrate that our approach can obtain higher utility by such multiple approximate updates, and achieve the error bounds asymptotic to the state-of-art ones for differentially private empirical risk minimization.


FiberStars: Visual Comparison of Diffusion Tractography Data between Multiple Subjects

arXiv.org Artificial Intelligence

Tractography from high-dimensional diffusion magnetic resonance imaging (dMRI) data allows brain's structural connectivity analysis. Recent dMRI studies aim to compare connectivity patterns across thousands of subjects to understand subtle abnormalities in brain's white matter connectivity across disease populations. Besides connectivity differences, researchers are also interested in investigating distributions of biologically sensitive dMRI derived metrics across subject groups. Existing software products focus solely on the anatomy or are not intuitive and restrict the comparison of multiple subjects. In this paper, we present the design and implementation of FiberStars, a visual analysis tool for tractography data that allows the interactive and scalable visualization of brain fiber clusters in 2D and 3D. With FiberStars, researchers can analyze and compare multiple subjects in large collections of brain fibers. To evaluate the usability of our software, we performed a quantitative user study. We asked non-experts to find patterns in a large tractography dataset with either FiberStars or AFQ-Browser, an existing dMRI exploration tool. Our results show that participants using FiberStars can navigate extensive collections of tractography faster and more accurately. We discuss our findings and provide an analysis of the requirements for comparative visualizations of tractography data. All our research, software, and results are available openly.


Data Driven Aircraft Trajectory Prediction with Deep Imitation Learning

arXiv.org Artificial Intelligence

The current Air Traffic Management (ATM) system worldwide has reached its limits in terms of predictability, efficiency and cost effectiveness. Different initiatives worldwide propose trajectory-oriented transformations that require high fidelity aircraft trajectory planning and prediction capabilities, supporting the trajectory life cycle at all stages efficiently. Recently proposed data-driven trajectory prediction approaches provide promising results. In this paper we approach the data-driven trajectory prediction problem as an imitation learning task, where we aim to imitate experts "shaping" the trajectory. Towards this goal we present a comprehensive framework comprising the Generative Adversarial Imitation Learning state of the art method, in a pipeline with trajectory clustering and classification methods. This approach, compared to other approaches, can provide accurate predictions for the whole trajectory (i.e. with a prediction horizon until reaching the destination) both at the pre-tactical (i.e. starting at the departure airport at a specific time instant) and at the tactical (i.e. from any state while flying) stages, compared to state of the art approaches.


Variance Linear Discriminant Analysis for IRIS Biometrics

AAAI Conferences

Dichotomy transformation in biometric authentication problem creates a two class (""within"" or ""between"") classification problem in multivariate distance space. Linear discriminant analysis, which is a linear classifier, results in good performance in IRIS biometric authentication problem. However, it assumes that the distributions of two classes are normal, whereas they are closely related to the log-normal distributions. Here a modified variance linear discriminant analysis algorithm is proposed and its superior experimental results on the IRIS biometric database are reported.


MCMC-Based Learning of Finite Bivariate Beta Mixture Models

AAAI Conferences

In this paper, we present a Bayesian approach for finite mixture models based on three-parameter bivariate Beta distributions. The estimation of the parameters is based on the Monte Carlo simulation technique of Gibbs sampling mixed with a Metropolis-Hastings step. The performance of our Bayesian algorithm is verified by several synthetic datasets and in the end, the feasibility of the proposed method is demonstrated by experimenting on some real datasets in which, the results are compared with those obtained by implementing the same approach using Gaussian mixture model.


A Preliminary Study of Spatial Bias in Knn Distance Metrics

AAAI Conferences

A machine learning algorithm for image classification exhibits spatial bias if permuting the order of image pixels significantly alters its classification accuracy. In this paper, we explore the spatial bias of a number of different distance metrics for k-nearest-neighbor image classification. One distance metric is inspired by the convolutional kernels employed in convolutional neural networks. The other metrics are based on BRIEF descriptors, which generate bit vectors corresponding to images based on comparisons of pixel intensity values. We found that the convolutional distance metric exhibited a strong positive spatial bias, as did one of the BRIEF descriptors. Another BRIEF descriptor exhibited a negative spatial bias, and the remainder exhibited little or no spatial bias. These results lay a foundation for future work that would involve larger numbers of convolutional iterations, potentially synergized with BRIEF-style image preprocessing.


Experimentation on Hand Drawn Sketches by Children to Classify Draw-a-Person Test Images in Psychology

AAAI Conferences

Classification of hand drawn sketches with respect to content quality is extremely challenging task, comparing to usual image classification methods. In brief, we need to train computational device to able to classify the images of the same object into different classes with respect their content quality. In this paper we tested several methods of image classification, using machine learning and computer vision algorithms, to classify Draw-a-Person test images sketched by primary school students in Nigeria, aged 4 to 11 years. We collected 1000 original sketches and manually classified them (using guidelines from existing literature) according to the ages (8 classes) before testing this dataset on a computational device. The highest accuracy achieved in this experiment was 62%. We achieved this result with novel method, where we used Bag of Visual Words and K-means algorithm to count key-points on each sketch. We strongly believe that this challenging task needs further research to improve classification accuracy, we, therefore, release the complete dataset of sketches to the community.


Using Simulated Annealing to Declutter Genome Visualizations

AAAI Conferences

AccuSyn is an interactive browser that visualizes conserved synteny relations (similar features) in genomes, giving biologists insights into the evolutionary history and functional relationships between genes. Even simple organisms have huge numbers of genomic features, and raw synteny plots present a daunting clutter of connections of which to make sense. Using a mixed initiative approach, AccuSyn integrates simulated annealing, a well-known metaheuristic for optimization problems, with human interventions to offer non-experts a way to automate decluttering, eliminating a tedious manual bottleneck in the discovery of syntenic information. AccuSyn has since been deployed online to a world user community


Spatially Aligned Clustering of Driving Simulator Data

AAAI Conferences

We set out to compare the utility of different representations of driving simulator time series data in the context of both supervised and unsupervised learning algorithms. Given the task of identifying similar time series; it is important to understand how a dataset of time series samples might be distributed and how effectively different methods capture the groupings of distinct behaviors. First we engineer three representations of the driving simulator data: converting them to feature vectors, using the raw time series, and rendering them as images. At which point, we introduce a novel method for comparing time series using temporal and spatial alignments. Then, we employ a battery of clustering algorithms to isolate groups of samples with similar traits and evaluate the quality of clusters produced. We also explore the performance of k-NN classifiers using the different dissimilarity measures resulting from these representations.