Genre
Semidefinite tests for latent causal structures
Kela, Aditya, von Prillwitz, Kai, Aberg, Johan, Chaves, Rafael, Gross, David
In spite of the primal importance of discovering causal relations in science, the statistical analysis of empirical data has historically shied away from causality . Only releatively recently has a rigorous theory of causality emerged (see, for instance, [ 1, 2 ]), showing that empirical data indeed can contain information about causation rather than mere correlation. Since then, causal inference has quickly become influential. Examples range from applications to the inference of genetic [ 3] and social networks [ 4], to a better understanding of the role of causality within quantum physics [ 5-13]. T o formalize causal mechanisms it has become popular to use directed acyclic graphs (DAGs) where nodes denote random variables and directed edges (arrows) account for their causal relations. Central problems within this context include inferenceor model selection: 'Given samples from a number of observable variables, which DAG should we associate with them?', as well as hypothesis testing: 'Can the observed data be explained in terms of an assumed DAG?' Here, we concentrate on the latter problem and propose a novel solution based on the covariances that a given causal structure gives rise to.
Automatic sleep monitoring using ear-EEG
Nakamura, Takashi, Goverdovsky, Valentin, Morrell, Mary J., Mandic, Danilo P.
The monitoring of sleep patterns without patient's inconvenience or involvement of a medical specialist is a clinical question of significant importance. To this end, we propose an automatic sleep stage monitoring system based on an affordable, unobtrusive, discreet, and long-term wearable in-ear sensor for recording the Electroencephalogram (ear-EEG). The selected features for sleep pattern classification from a single ear-EEG channel include the spectral edge frequency (SEF) and multi- scale fuzzy entropy (MSFE), a structural complexity feature. In this preliminary study, the manually scored hypnograms from simultaneous scalp-EEG and ear-EEG recordings of four subjects are used as labels for two analysis scenarios: 1) classification of ear-EEG hypnogram labels from ear-EEG recordings and 2) prediction of scalp-EEG hypnogram labels from ear-EEG recordings. We consider both 2-class and 4-class sleep scoring, with the achieved accuracies ranging from 78.5 % to 95.2 % for ear-EEG labels predicted from ear-EEG, and 76.8 % to 91.8 % for scalp-EEG labels predicted from ear-EEG. The corresponding kappa coefficients, which range from 0.64 to 0.83 for Scenario 1 and from 0.65 to 0.80 for Scenario 2, indicate a Substantial to Almost Perfect agreement, thus proving the feasibility of in-ear sensing for sleep monitoring in the community.
Towards multiple kernel principal component analysis for integrative analysis of tumor samples
Speicher, Nora K., Pfeifer, Nico
Personalized treatment of patients based on tissue-specific cancer subtypes has strongly increased the efficacy of the chosen therapies. Even though the amount of data measured for cancer patients has increased over the last years, most cancer subtypes are still diagnosed based on individual data sources (e.g. gene expression data). We propose an unsupervised data integration method based on kernel principal component analysis. Principal component analysis is one of the most widely used techniques in data analysis. Unfortunately, the straight-forward multiple-kernel extension of this method leads to the use of only one of the input matrices, which does not fit the goal of gaining information from all data sources. Therefore, we present a scoring function to determine the impact of each input matrix. The approach enables visualizing the integrated data and subsequent clustering for cancer subtype identification. Due to the nature of the method, no free parameters have to be set. We apply the methodology to five different cancer data sets and demonstrate its advantages in terms of results and usability.
How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
Liu, Chia-Wei, Lowe, Ryan, Serban, Iulian V., Noseworthy, Michael, Charlin, Laurent, Pineau, Joelle
We investigate evaluation metrics for dialogue response generation systems where supervised labels, such as task completion, are not available. Recent works in response generation have adopted metrics from machine translation to compare a model's generated response to a single target response. We show that these metrics correlate very weakly with human judgements in the non-technical Twitter domain, and not at all in the technical Ubuntu domain. We provide quantitative and qualitative results highlighting specific weaknesses in existing metrics, and provide recommendations for future development of better automatic evaluation metrics for dialogue systems.
Google focuses GCP on machine learning and data analytics
Google dispelled any lingering questions about its commitment to cloud in 2016, as its strategy emerged to become a major player in the enterprise market in the years ahead. The company has spent tens of billions of dollars to build the underlying infrastructure, services and talent pool that feed Google Cloud Platform (GCP). As a result, it has de-emphasized price as the main differentiator and instead focused on enterprise demands, data analytics and a set of technologies to drive applications the company said will dominate the industry in the future. Machine learning is still too new for many IT shops, but Google has banked on it as the future of cloud computing. New services added in 2016 included the Machine Intelligence suite of services and new releases for translation, text analysis, and image and speech recognition.
Machine Learning With Big Data - ProLearningHub
Now a days, most of the data is in textual form and we need some effective tools to process this unstructured and semi-structured data. Many search engines and advertisers are using machine learning algorithms for predicting customer behavior and content recommendations. There is a need to learn basics of large-data processing using predictive models with right tools. This course is designed to give you knowledge of machine learning tools and their applications in predictive analysis. You will learn to train, evaluate, and validate basic predictive models.
Can Technology Make Football Safer?
On October 4, 1986, the University of Alabama hosted Notre Dame in a game of football. Notre Dame had won the previous four contests, but this time Alabama was favored. It had a stifling defense and a swift senior linebacker named Cornelius Bennett. Ray Perkins, Alabama's head coach, said of him, "I don't think there's a better player in America." Early in the game, with the score tied, Bennett blitzed Notre Dame's quarterback, Steve Beuerlein. "I was like a speeding train, and Beuerlein just happened to be standing on the railroad track," Bennett told me recently. Football is essentially a spectacle of car crashes. In 2004, researchers at the University of North Carolina, examining data gathered from helmet-mounted sensors, discovered that many football collisions compare in intensity to a vehicle smashing into a wall at twenty-five miles per hour. Bennett, who weighed two hundred and thirty-five pounds, drove his shoulder into Beuerlein's chest and heard what sounded like a balloon being punctured--"basically, the air going out of him." Beuerlein landed on his back. He stood up, wobbly and dazed. "I saw mouths moving, but I heard no voices," he later said. After Bennett's "vicious, high-speed direct slam," as the Times put it, Alabama seized the momentum and won, 28โ10. Following college, Bennett was drafted into the National Football League. Between 1987 and 1995, he played for the Buffalo Bills, and appeared in four Super Bowls. During his pro career, he made more than a thousand tackles, playing through sprains, muscle tears, broken bones, and concussions. I asked him how many concussions he'd had. "In my medical file, there are probably six." "I couldn't even begin to tell you." "I played a long time," he said. "Every week after a game, I got some sort of headache." In 1996, he signed a thirteen-million-dollar contract with the Atlanta Falcons. He received weekly injections of Toradol, an anti-inflammatory drug. "It was magic--it made me feel like I was twenty-four again," Bennett said. He helped carry Atlanta to the Super Bowl--his fifth. In 2000, at the age of thirty-five, Bennett retired and moved to Florida. He lived in a hotel in Miami's Bal Harbour area, worked on his golf handicap, and vacationed with his wife and friends in Europe and in the Napa Valley. Several of Bennett's football peers were having a far tougher time. Darryl Talley, a former Bills teammate, suffered from severe depression. Mike Webster, a Hall of Fame center for the Pittsburgh Steelers, had become a homeless alcoholic; he died, of a heart attack, in 2002. Three years later, Terry Long, another former Steeler, committed suicide by drinking antifreeze. Andre Waters, a former Philadelphia Eagles safety, killed himself with a gunshot to the head. A neuropathologist named Bennet Omalu autopsied Webster, Long, and Waters, and detected a pattern: each had a high concentration of an abnormal form of a protein, called tau, on his brain.
10 Breakthrough Technologies of 2016: Where Are They Now?
In February MIT Technology Review highlighted 10 breakthrough technologies poised to significantly change the world over the next few years. Here's how they have progressed since then. We predicted that 2016 would bring major progress on high-tech cancer cures enabled by using gene editing to tune the human immune system, and it did. First, American scientists got a green light to start using the gene-editing technique called CRISPR to customize T cells and turn them into cancer killers. That study turned out to have the backing of Internet billionaire Sean Parker, who in April had announced he'd give away $250 million toward "hacking" the immune system. By November, a Chinese company announced it had raced ahead and dosed a patient with the first T cells edited with CRISPR.
The Year in Machine Learning (Part Two)
This is the second installment in a three-part review of 2016 in machine learning and deep learning. Part One, here, covered general trends. In Part Two, we review the year in open source machine learning and deep learning projects. Part Three will cover commercial machine learning and deep learning software and services. There are thousands of open source projects on the market today, and we cannot cover them all. We've selected the most relevant projects based on usage reported in surveys of data scientists, as well as development activity recorded in OpenHub. In this post, we limit the scope to projects with a non-profit governance structure, and those offered by commercial ventures that do not also provide licensed software. Part Three will include software vendors who offer open source "community" editions together with commercially licensed software.
Probabilistic Feature Selection and Classification Vector Machine
Jiang, Bingbing, Li, Chang, Chen, Huanhuan, Yao, Xin, de Rijke, Maarten
Sparse Bayesian learning is one of the state-of- the-art machine learning algorithms, which is able to make stable and reliable probabilistic predictions. However, some of these algorithms, e.g. probabilistic classification vector machine (PCVM) and relevant vector machine (RVM), are not capable of eliminating irrelevant and redundant features which could lead to performance degradation. To tackle this problem, in this paper, we propose a sparse Bayesian classifier which simultaneously selects the relevant samples and features. We name this classifier a probabilistic feature selection and classification vector machine (PFCVM), in which truncated Gaussian distributions are em- ployed as both sample and feature priors. In order to derive the analytical solution for the proposed algorithm, we use Laplace approximation to calculate approximate posteriors and marginal likelihoods. Finally, we obtain the optimized parameters and hyperparameters by the type-II maximum likelihood method. The experiments on synthetic data set, benchmark data sets and high dimensional data sets validate the performance of PFCVM under two criteria: accuracy of classification and efficacy of selected features. Finally, we analyze the generalization performance of PFCVM and derive a generalization error bound for PFCVM. Then by tightening the bound, we demonstrate the significance of the sparseness for the model.