Goto

Collaborating Authors

 Statistical Learning


Graph-Bert: Only Attention is Needed for Learning Graph Representations

arXiv.org Machine Learning

The dominant graph neural networks (GNNs) over-rely on the graph links, several serious performance problems with which have been witnessed already, e.g., suspended animation problem and over-smoothing problem. What's more, the inherently inter-connected nature precludes parallelization within the graph, which becomes critical for large-sized graph, as memory constraints limit batching across the nodes. In this paper, we will introduce a new graph neural network, namely GRAPH-BERT (Graph based BERT), solely based on the attention mechanism without any graph convolution or aggregation operators. Instead of feeding GRAPH-BERT with the complete large input graph, we propose to train GRAPH-BERT with sampled linkless subgraphs within their local contexts. GRAPH-BERT can be learned effectively in a standalone mode. Meanwhile, a pre-trained GRAPH-BERT can also be transferred to other application tasks directly or with necessary fine-tuning if any supervised label information or certain application oriented objective is available. We have tested the effectiveness of GRAPH-BERT on several graph benchmark datasets. Based the pre-trained GRAPH-BERT with the node attribute reconstruction and structure recovery tasks, we further fine-tune GRAPH-BERT on node classification and graph clustering tasks specifically. The experimental results have demonstrated that GRAPH-BERT can out-perform the existing GNNs in both the learning effectiveness and efficiency.


Investigating Classification Techniques with Feature Selection For Intention Mining From Twitter Feed

arXiv.org Artificial Intelligence

In the last decade, social networks became most popular medium for communication and interaction. As an example, micro-blogging service Twitter has more than 200 million registered users who exchange more than 65 million posts per day. Users express their thoughts, ideas, and even their intentions through these tweets. Most of the tweets are written informally and often in slang language, that contains misspelt and abbreviated words. This paper investigates the problem of selecting features that affect extracting user's intention from Twitter feeds based on text mining techniques. It starts by presenting the method we used to construct our own dataset from extracted Twitter feeds. Following that, we present two techniques of feature selection followed by classification. In the first technique, we use Information Gain as a one-phase feature selection, followed by supervised classification algorithms. In the second technique, we use a hybrid approach based on forward feature selection algorithm in which two feature selection techniques employed followed by classification algorithms. We examine these two techniques with four classification algorithms. We evaluate them using our own dataset, and we critically review the results.


DeepEnroll: Patient-Trial Matching with Deep Embedding and Entailment Prediction

arXiv.org Artificial Intelligence

Clinical trials are essential for drug development but often suffer from expensive, inaccurate and insufficient patient recruitment. The core problem of patient-trial matching is to find qualified patients for a trial, where patient information is stored in electronic health records (EHR) while trial eligibility criteria (EC) are described in text documents available on the web. How to represent longitudinal patient EHR? How to extract complex logical rules from EC? Most existing works rely on manual rule-based extraction, which is time consuming and inflexible for complex inference. To address these challenges, we proposed DeepEnroll, a cross-modal inference learning model to jointly encode enrollment criteria (text) and patients records (tabular data) into a shared latent space for matching inference. DeepEnroll applies a pre-trained Bidirectional Encoder Representations from Transformers(BERT) model to encode clinical trial information into sentence embedding. And uses a hierarchical embedding model to represent patient longitudinal EHR. In addition, DeepEnroll is augmented by a numerical information embedding and entailment module to reason over numerical information in both EC and EHR. These encoders are trained jointly to optimize patient-trial matching score. We evaluated DeepEnroll on the trial-patient matching task with demonstrated on real world datasets. DeepEnroll outperformed the best baseline by up to 12.4% in average F1.


Causality based Feature Fusion for Brain Neuro-Developmental Analysis

arXiv.org Artificial Intelligence

REPLACE THIS LINE WITH YOUR PAPER IDENTIFICATION NUMBER (DOUBLE - CLICK HERE TO EDIT) 1 Abstract -- Human brain development is a complex and dynamic process that is affected by several factors such as genetic s, sex hormones, and environmental changes . A number of recent studies on brain development have examined functional connectivity (FC) defi ned by the temporal correlation between time series of different brain regions. We propose to add the directional flow of information during brain maturation . To do so, w e extract effective connectivity (EC) through Granger causality (GC) for two different groups of subjects, i.e., children and young adults. The motivation is that the inclusion of causal interaction may further discriminate brain connections between two age groups and help to discover new conn ections between brain regions. The contributions of this study are three fold. First, t here has been a lack of attention to EC - based feature extraction in the context of brain development . T o this end, we propose a new kernel - based GC ( K GC) method to learn nonlinearity of complex brain network, where a reduced Sine hyperbolic polynomial ( RSP) neural network wa s used as our proposed learner . S econd, we use d causality values as the weight for the directional connectivity between brain regions . Our f indings indicate d that the strength of connections was significantly higher in young adult s relative to children. In addition, our new EC - based feature outperform ed FC - based analysis from Philadelphia neurocohort (PNC) study wi th better discrimination of the different age groups . Moreover, the fusion of these two sets of features (FC EC) improve d brain age prediction accuracy by more than 4 %, indicating that they should be used together for brain development stud ies . I NTRODUCTION uman brain development is a prolonged process that is initiated from the third gestational week (GW) to late adolescence, and presumably to the entire lifespan [ 1 ].


Deception detection on the Bag-of-lies dataset

#artificialintelligence

Lie detection has been a topic of interest since the beginning of the 20th century, and since then a lot of different methods have been used to try to achieve this, such as changes in inspiration-expiration ratio, increases in systolic blood pressure, dilatation of the pupil size, heart rate, etc. Usually, when people think about lie detection, the most common method that comes to mind is the polygraph. This method combines various techniques to detect autonomic reactions which include changes in body functions that are not easily controlled by the conscious mind. However, still requires a large amount of training which is achieved by control questions where the answers are known to later compare how the subject reacts. Polygraph offers an accuracy of around 70% in the general population⁴, a number which is greater than trained humans can achieve by just looking at the person, however, this doesn't mean that this method is infallible since people have found ways to cheat the system by just training or by using drugs to suppress these reactions. In general, these methods usually have not offered as good results as to be used in court in most countries.


Bayesian Product Ranking at Wayfair Wayfair

#artificialintelligence

Given sufficient data, we could just use the logistic regression model without further changes. Wayfair handled more than 9 million orders last quarter alone, which initially might sound like more than enough. However, those orders were spread out among millions of products, yielding just a few orders per product at most. Small integers like these can be extremely noisy, so we always have to worry that one product simply seems better than another because of random chance. For example, it is hard to tell if a product that happened to attract three orders is actually any better than one that happened to attract two, or if it just got lucky.


Time series forecasting with random forest

#artificialintelligence

Benjamin Franklin said that only two things are certain in life: death and taxes. That explains why my colleagues at STATWORX were less than excited when they told me about their plans for the weekend a few weeks back: doing their income tax declaration. Man, I thought, that sucks, I'd rather spend this time outdoors. And then an idea was born. What could taxes and the outdoors possibly have in common?


Feature-based time series analysis

#artificialintelligence

I used this example in my talk at useR!2019 in Toulouse, and it is also the basis of a vignette in the package, and a recent blog post by Mitchell O'Hara-Wild. The data set contains domestic tourist visitor nights in Australia, disaggregated by State, Region and Purpose. An example of a feature would be the autocorrelation function at lag 1 -- it is a numerical summary capturing some aspect of the time series. Autocorrelations at other lags are also features, as are the autocorrelations of the first differenced series, or the seasonally differenced series, etc. Another example of a feature is the strength of seasonality of a time series, as measured by \(1-\text{Var}(R_t)/\text{Var}(S_t R_t)\) where \(S_t\) is the seasonal component and \(R_t\) is the remainder component in an STL decomposition.


Secure and Robust Machine Learning for Healthcare: A Survey

arXiv.org Machine Learning

Recent years have witnessed widespread adoption of machine learning (ML)/deep learning (DL) techniques due to their superior performance for a variety of healthcare applications ranging from the prediction of cardiac arrest from one-dimensional heart signals to computer-aided diagnosis (CADx) using multi-dimensional medical images. Notwithstanding the impressive performance of ML/DL, there are still lingering doubts regarding the robustness of ML/DL in healthcare settings (which is traditionally considered quite challenging due to the myriad security and privacy issues involved), especially in light of recent results that have shown that ML/DL are vulnerable to adversarial attacks. In this paper, we present an overview of various application areas in healthcare that leverage such techniques from security and privacy point of view and present associated challenges. In addition, we present potential methods to ensure secure and privacy-preserving ML for healthcare applications. Finally, we provide insight into the current research challenges and promising directions for future research.


Unsupervisedly Learned Representations: Should the Quest be Over?

arXiv.org Artificial Intelligence

There exists a Classification accuracy gap of about 20% between our best methods of generating Unsupervisedly Learned Representations and the accuracy rates achieved by (naturally Unsupervisedly Learning) humans. We are at our fourth decade at least in search of this class of paradigms. It thus may well be that we are looking in the wrong direction. We present in this paper a possible solution to this puzzle. We demonstrate that Reinforcement Learning schemes can learn representations, which may be used for Pattern Recognition tasks such as Classification, achieving practically the same accuracy as that of humans. Our main modest contribution lies in the observations that: a. when applied to a real world environment (e.g. nature itself) Reinforcement Learning does not require labels, and thus may be considered a natural candidate for the long sought, accuracy competitive Unsupervised Learning method, and b. in contrast, when Reinforcement Learning is applied in a simulated or symbolic processing environment (e.g. a computer program) it does inherently require labels and should thus be generally classified, with some exceptions, as Supervised Learning. The corollary of these observations is that further search for Unsupervised Learning competitive paradigms which may be trained in simulated environments like many of those found in research and applications may be futile.