Goto

Collaborating Authors

 Genre


Using Global Constraints and Reranking to Improve Cognates Detection

arXiv.org Machine Learning

Global constraints and reranking have not been used in cognates detection research to date. We propose methods for using global constraints by performing rescoring of the score matrices produced by state of the art cognates detection systems. Using global constraints to perform rescoring is complementary to state of the art methods for performing cognates detection and results in significant performance improvements beyond current state of the art performance on publicly available datasets with different language pairs and various conditions such as different levels of baseline state of the art performance and different data size conditions, including with more realistic large data size conditions than have been evaluated with in the past.


Surrogate Aided Unsupervised Recovery of Sparse Signals in Single Index Models for Binary Outcomes

arXiv.org Machine Learning

We consider the recovery of regression coefficients, denoted by $\boldsymbol{\beta}_0$, for a single index model (SIM) relating a binary outcome $Y$ to a set of possibly high dimensional covariates $\boldsymbol{X}$, based on a large but 'unlabeled' dataset $\mathcal{U}$, with $Y$ never observed. On $\mathcal{U}$, we fully observe $\boldsymbol{X}$ and additionally, a surrogate $S$ which, while not being strongly predictive of $Y$ throughout the entirety of its support, can forecast it with high accuracy when it assumes extreme values. Such datasets arise naturally in modern studies involving large databases such as electronic medical records (EMR) where $Y$, unlike $(\boldsymbol{X}, S)$, is difficult and/or expensive to obtain. In EMR studies, an example of $Y$ and $S$ would be the true disease phenotype and the count of the associated diagnostic codes respectively. Assuming another SIM for $S$ given $\boldsymbol{X}$, we show that under sparsity assumptions, we can recover $\boldsymbol{\beta}_0$ proportionally by simply fitting a least squares LASSO estimator to the subset of the observed data on $(\boldsymbol{X}, S)$ restricted to the extreme sets of $S$, with $Y$ imputed using the surrogacy of $S$. We obtain sharp finite sample performance bounds for our estimator, including deterministic deviation bounds and probabilistic guarantees. We demonstrate the effectiveness of our approach through multiple simulation studies, as well as by application to real data from an EMR study conducted at the Partners HealthCare Systems.


Efficient and Adaptive Linear Regression in Semi-Supervised Settings

arXiv.org Machine Learning

We consider the linear regression problem under semi-supervised settings wherein the available data typically consists of: (i) a small or moderate sized 'labeled' data, and (ii) a much larger sized 'unlabeled' data. Such data arises naturally from settings where the outcome, unlike the covariates, is expensive to obtain, a frequent scenario in modern studies involving large databases like electronic medical records (EMR). Supervised estimators like the ordinary least squares (OLS) estimator utilize only the labeled data. It is often of interest to investigate if and when the unlabeled data can be exploited to improve estimation of the regression parameter in the adopted linear model. In this paper, we propose a class of 'Efficient and Adaptive Semi-Supervised Estimators' (EASE) to improve estimation efficiency. The EASE are two-step estimators adaptive to model mis-specification, leading to improved (optimal in some cases) efficiency under model mis-specification, and equal (optimal) efficiency under a linear model. This adaptive property, often unaddressed in the existing literature, is crucial for advocating 'safe' use of the unlabeled data. The construction of EASE primarily involves a flexible 'semi-non-parametric' imputation, including a smoothing step that works well even when the number of covariates is not small; and a follow up 'refitting' step along with a cross-validation (CV) strategy both of which have useful practical as well as theoretical implications towards addressing two important issues: under-smoothing and over-fitting. We establish asymptotic results including consistency, asymptotic normality and the adaptive properties of EASE. We also provide influence function expansions and a 'double' CV strategy for inference. The results are further validated through extensive simulations, followed by application to an EMR study on auto-immunity.


Magnetic Hamiltonian Monte Carlo

arXiv.org Machine Learning

Hamiltonian Monte Carlo (HMC) exploits Hamiltonian dynamics to construct efficient proposals for Markov chain Monte Carlo (MCMC). In this paper, we present a generalization of HMC which exploits \textit{non-canonical} Hamiltonian dynamics. We refer to this algorithm as magnetic HMC, since in 3 dimensions a subset of the dynamics map onto the mechanics of a charged particle coupled to a magnetic field. We establish a theoretical basis for the use of non-canonical Hamiltonian dynamics in MCMC, and construct a symplectic, leapfrog-like integrator allowing for the implementation of magnetic HMC. Finally, we exhibit several examples where these non-canonical dynamics can lead to improved mixing of magnetic HMC relative to ordinary HMC.


Causation: The Why Beneath The What

@machinelearnbot

Kevin Gray: If we think about it, most of our daily conversations invoke causation, at least informally. We often say things like "I dropped by this store instead of my usual place because I needed to go to the laundry and it was on the way" or "I always buy chocolate ice cream because that's what my kids like." First, to get started, can you give us nontechnical definitions of causation and causal analysis? Tyler VanderWeele: Well, it turns out that there a number of different contexts in which words like "cause" and "because" are used. Aristotle, in his Physics and again in his Metaphysics, distinguished between what he viewed as four different types of causes: material causes, formal causes, efficient causes, and final causes.


Google AI can easily erase watermarks from photos

Daily Mail - Science & tech

Google created an AI that can easily remove the digital watermarks photographers put on their images to prevent unauthorized use. In a paper published online Thursday, research scientists at the firm describe how a computer algorithm can get past this protection and remove watermarks automatically when working with collections instead of single images. This gives users unobstructed access to the clean images the watermarks are intended to protect, and Google said the purpose of the research was to disclose the issue and find solutions. Google created an AI that can easily and automatically remove the types of digital watermarks photographers put on their images to prevent unauthorized use working with collections instead of single images. While removing a watermark from a single image is extremely challenging, the researchers found a'loophole' that makes it simple when working with an entire collection.


Machine Learning Classification Algorithms using MATLAB

@machinelearnbot

As bonus, you also learn how to share your analysis results with your collegues friends and others and create visual analysis of your results. You will also have access to some practice questions, which will give you hand on experience.


R Machine Learning solutions - Udemy

@machinelearnbot

R is a statistical programming language that provides impressive tools to analyze data and create high-level graphics. This video course will take you from very basics of R to creating insightful machine learning models with R. You will start with setting up the environment and then perform data ETL in R. Data exploration examples are provided that demonstrate how powerful data visualization and machine learning is in discovering hidden relationship. You will then dive into important machine learning topics, including data classification, regression, clustering, association rule mining, and dimensionality reduction. Yu-Wei, Chiu (David Chiu) is the founder of LargitData, a startup company that mainly focuses on providing big data and machine learning products.


Core Spatial Data Analysis: Introductory GIS with R and QGIS

@machinelearnbot

Do you find GIS & Spatial Data books & manuals too vague, expensive & not practical and looking for a course that takes you by hand, teaches you all the concepts, and get you started on a real life project? Or perhaps you want to save time and learn how to automate some of the most common GIS tasks? I'm very excited you found my spatial data analysis course. My course provides a foundation to carry out PRACTICAL, real-life spatial data analysis tasks in popular and FREE software frameworks. My name is MINERVA SINGH and i am an Oxford University MPhil (Geography and Environment) graduate.


[Intermediate] Spatial Data Analysis with R, QGIS & More

@machinelearnbot

This course is designed to take users who use R and QGIS for basic spatial data/GIS analysis to perform more advanced GIS tasks (including automated workflows and geo-referencing) using a variety of different data. In addition to making you proficient in R and QGIS for spatial data analysis, you will be introduced to another powerful free GIS software.. GRASS. This course takes a completely practical approach to spatial data analysis and mapping- Each lecture will teach you a practical application/processing technique which you can apply easily. The course is taught by Minerva Singh, A PhD graduate from Cambridge University, UK, who has several years of research experience in Quantitative Ecology and an MPhil in Geography and Environment from Oxford University. Minerva has published papers in international peer reviewed journals and given talks at international conferences.