Genre
Evaluating Graph Signal Processing for Neuroimaging Through Classification and Dimensionality Reduction
Ménoret, Mathilde, Farrugia, Nicolas, Pasdeloup, Bastien, Gripon, Vincent
Graph Signal Processing (GSP) is a promising framework to analyze multi-dimensional neuroimaging datasets, while taking into account both the spatial and functional dependencies between brain signals. In the present work, we apply dimensionality reduction techniques based on graph representations of the brain to decode brain activity from real and simulated fMRI datasets. We introduce seven graphs obtained from a) geometric structure and/or b) functional connectivity between brain areas at rest, and compare them when performing dimension reduction for classification. We show that mixed graphs using both a) and b) offer the best performance. We also show that graph sampling methods perform better than classical dimension reduction including Principal Component Analysis (PCA) and Independent Component Analysis (ICA).
Significance testing in non-sparse high-dimensional linear models
In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violated, with a large number of covariates expected to be associated with the response, indicating that possibly all, rather than just a few, parameters are non-zero. A natural example is a genome-wide gene expression profiling, where all genes are believed to affect a common disease marker. We show that existing inferential methods are sensitive to the sparsity assumption, and may, in turn, result in the severe lack of control of Type-I error. In this article, we propose a new inferential method, named CorrT, which is robust to model misspecification and adaptive to the sparsity assumption. CorrT is shown to have Type I error approaching the nominal level for \textit{any} models and Type II error approaching zero for sparse and many dense models. In fact, CorrT is also shown to be optimal in a variety of frameworks: sparse, non-sparse and hybrid models where sparse and dense signals are mixed. Numerical experiments show a favorable performance of the CorrT test compared to the state-of-the-art methods.
Stem-ming the Tide: Predicting STEM attrition using student transcript data
Aulck, Lovenoor, Aras, Rohan, Li, Lysia, L'Heureux, Coulter, Lu, Peter, West, Jevin
Science, technology, engineering, and math (STEM) fields play growing roles in national and international economies by driving innovation and generating high salary jobs. Yet, the US is lagging behind other highly industrialized nations in terms of STEM education and training. Furthermore, many economic forecasts predict a rising shortage of domestic STEM-trained professions in the US for years to come. One potential solution to this deficit is to decrease the rates at which students leave STEM-related fields in higher education, as currently over half of all students intending to graduate with a STEM degree eventually attrite. However, little quantitative research at scale has looked at causes of STEM attrition, let alone the use of machine learning to examine how well this phenomenon can be predicted. In this paper, we detail our efforts to model and predict dropout from STEM fields using one of the largest known datasets used for research on students at a traditional campus setting. Our results suggest that attrition from STEM fields can be accurately predicted with data that is routinely collected at universities using only information on students' first academic year. We also propose a method to model student STEM intentions for each academic term to better understand the timing of STEM attrition events. We believe these results show great promise in using machine learning to improve STEM retention in traditional and non-traditional campus settings.
An inexact subsampled proximal Newton-type method for large-scale machine learning
Liu, Xuanqing, Hsieh, Cho-Jui, Lee, Jason D., Sun, Yuekai
We propose a fast proximal Newton-type algorithm for minimizing regularized finite sums that returns an $\epsilon$-suboptimal point in $\tilde{\mathcal{O}}(d(n + \sqrt{\kappa d})\log(\frac{1}{\epsilon}))$ FLOPS, where $n$ is number of samples, $d$ is feature dimension, and $\kappa$ is the condition number. As long as $n > d$, the proposed method is more efficient than state-of-the-art accelerated stochastic first-order methods for non-smooth regularizers which requires $\tilde{\mathcal{O}}(d(n + \sqrt{\kappa n})\log(\frac{1}{\epsilon}))$ FLOPS. The key idea is to form the subsampled Newton subproblem in a way that preserves the finite sum structure of the objective, thereby allowing us to leverage recent developments in stochastic first-order methods to solve the subproblem. Experimental results verify that the proposed algorithm outperforms previous algorithms for $\ell_1$-regularized logistic regression on real datasets.
Deep Learning for Accelerated Reliability Analysis of Infrastructure Networks
Nabian, Mohammad Amin, Meidani, Hadi
Assessment of the impact of natural disasters on infrastructure systems is of importance toward four main objectives: (1) Planning for actions that eliminate or reduce the long-term risk to human life and infrastructure systems (e.g.[2]); (2) Disaster preparation or adjustment, which aims to reduce the risk of damages and injuries while enabling the capability to cope with the temporary disruption of the infrastructure systems (e.g.[3]); (3) Development of effective emergency response strategies (e.g.[4]); and (4) Post-disaster recovery planning (e.g.[5]). These four are, respectively, known as the mitigation, preparedness, response, and recovery practices. A variety of analytical [6], simulation [7-11], and optimization [12] approaches are proposed in the literature for hazard reliability analysis of infrastructure systems. A comprehensive literature review on transportation infrastructure system performance in disasters is provided in [13]. Simulation-based reliability assessment of large infrastructure systems are often computationally intractable or expensive due to the large number of network components, complex network topology, statistical dependence between component failures, and uncertainties in the hazard models. This will impose limitations on design optimization or sensitivity analysis of these systems. Alternatively, a more efficient response assessment for large infrastructure systems can be made possible by using approximate surrogates [14]. Surrogates are fast models that approximately describe the relationship between the system inputs and outputs and serve as a substitute for more expensive simulation tools. If the response evaluated by the reference expensive model is denoted by f (x), a surrgate seeks to provide a global approximate function f (x).
Artificial intelligence: Big data and invisible patients - MedCity News
If the barrier to precision medicine is data handling, then artificial intelligence (AI) may be the logical solution. Machine learning and deep learning are making inroads in a variety of industries, and seem poised to have a big impact in medicine, a process that is already in motion – and perhaps not a moment too soon. "Your chance in your lifetime of getting a false diagnosis, if you look at the data, is 100 percent," said Thomas Wilckens, founder and CEO at InnVentis to the audience at the recently-concluded Precision Medicine Leadership Summit in San Diego. "There's a lot to improve." Wilckens moderated Going Deep in the Fast Lane – the Rise of AI in Precision Medicine, which combined experts from industry and academia to parse this evolving segment.
Watson Tone Analyzer: 7 new tones to help understand how your customers are feeling - Watson
We are pleased to announce the launch of a new Tone Analyzer endpoint trained for Customer Engagement scenarios. The new endpoint was trained on customer support conversations on twitter, and the tones included are frustrated, sad, satisfied, excited, polite, impolite and sympathetic. Currently, the new endpoint is Beta functionality in the IBM Watson Tone Analyzer service. Given a textual conversation between a customer and an agent or company representative, the service detects the above mentioned tones both from the customer's and the agent's text. Why Did We Build the Tone Analyzer for Customer Engagement Endpoint?
Exoskeleton suit helps children with cerebral palsy walk
An exoskeleton suit has helped children with cerebral palsy to walk. Experts tested the device on seven children that are suffering from the disorder in a clinical trial. Each of them were having difficulty walking, with their knees starting to slope inwards. But a video of them walking in the exoskeleton suit showed noticeable improvements. It helped to improve knee extension for the children and made their walking more fluid while wearing the device.
How long will you keep playing? The game knows
We have a tendency to consider ourselves unique and unpredictable, but digital games research shows that this is far from the case. In fact, we can be categorised into groups of people who show the same behaviours, and what we do in the future is imminently predictable. For example, how you play a game will reveal what you are likely to do in the game next and how long you are going to stay interested in doing it. This means that games can now change tack while you're in them to provide you with the best possible experience and to encourage you to keep playing. When we play games, we generate traces of data which provide information on how we played.