Goto

Collaborating Authors

 Education


A Note on Optimal Sampling Strategy for Structural Variant Detection Using Optical Mapping

arXiv.org Machine Learning

A Note on Optimal Sampling Strategy for Structural V ariant Detection Using Optical Mapping Weiwei Li Department of Statistics and Operations Research University of North Carolina at Chapel Hill weiweili@live.unc.edu Abstract Structural variants compose the majority of human genetic variation, but are difficult to assess using current genomic sequencing technologies. Optical mapping technologies, which measure the size of chromosomal fragments between labeled markers, offer an alternative approach. As these technologies mature towards becoming clinical tools, there is a need to develop an approach for determining the optimal strategy for sampling biological material in order to detect a variant at some threshold. Here we develop an optimization approach using a simple, yet realistic, model of the genomic mapping process using a hyper-geometric distribution and probabilistic concentration inequalities. Our approach is both computationally and analytically tractable and includes a novel approach to getting tail bounds of hyper-geometric distribution. We show that if a genomic mapping technology can sample most of the chromosomal fragments within a sample, comparatively little biological material is needed to detect a variant at high confidence. 1 Introduction Structural variants (SV), insertions, deletions, translocations, copy number variants, are by far the most common types of human genetic variation (Chaisson et al., 2015). They have been linked to large number of heritable disorders (Hurles et al., 2008). Technology to assay the presence or absence of these variants has steadily improved in ease and resolution (Huddleston and Eichler, 2016; Audano et al., 2019).


Optimized Partial Identification Bounds for Regression Discontinuity Designs with Manipulation

arXiv.org Machine Learning

The regression discontinuity (RD) design is one of the most popular quasi-experimental methods for applied causal inference. In practice, the method is quite sensitive to the assumption that individuals cannot control their value of a "running variable" that determines treatment status precisely. If individuals are able to precisely manipulate their scores, then point identification is lost. We propose a procedure for obtaining partial identification bounds in the case of a discrete running variable where manipulation is present. Our method relies on two stages: first, we derive the distribution of non-manipulators under several assumptions about the data. Second, we obtain bounds on the causal effect via a sequential convex programming approach. We also propose methods for tightening the partial identification bounds using an auxiliary covariate, and derive confidence intervals via the bootstrap. We demonstrate the utility of our method on a simulated dataset.


A Comparison Study on Nonlinear Dimension Reduction Methods with Kernel Variations: Visualization, Optimization and Classification

arXiv.org Machine Learning

Because of high dimensionality, correlation among covariates, and noise contained in data, dimension reduction (DR) techniques are often employed to the application of machine learning algorithms. Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), and their kernel variants (KPCA, KLDA) are among the most popular DR methods. Recently, Supervised Kernel Principal Component Analysis (SKPCA) has been shown as another successful alternative. In this paper, brief reviews of these popular techniques are presented first. We then conduct a comparative performance study based on three simulated datasets, after which the performance of the techniques are evaluated through application to a pattern recognition problem in face image analysis. The gender classification problem is considered on MORPH-II and FG-NET, two popular longitudinal face aging databases. Several feature extraction methods are used, including biologically-inspired features (BIF), local binary patterns (LBP), histogram of oriented gradients (HOG), and the Active Appearance Model (AAM). After applications of DR methods, a linear support vector machine (SVM) is deployed with gender classification accuracy rates exceeding 95% on MORPH-II, competitive with benchmark results. A parallel computational approach is also proposed, attaining faster processing speeds and similar recognition rates on MORPH-II. Our computational approach can be applied to practical gender classification systems and generalized to other face analysis tasks, such as race classification and age prediction.


Clustered Federated Learning: Model-Agnostic Distributed Multi-Task Optimization under Privacy Constraints

arXiv.org Machine Learning

Federated Learning (FL) is currently the most widely adopted framework for collaborative training of (deep) machine learning models under privacy constraints. Albeit it's popularity, it has been observed that Federated Learning yields suboptimal results if the local clients' data distributions diverge. To address this issue, we present Clustered Federated Learning (CFL), a novel Federated Multi-Task Learning (FMTL) framework, which exploits geometric properties of the FL loss surface, to group the client population into clusters with jointly trainable data distributions. In contrast to existing FMTL approaches, CFL does not require any modifications to the FL communication protocol to be made, is applicable to general non-convex objectives (in particular deep neural networks) and comes with strong mathematical guarantees on the clustering quality. CFL is flexible enough to handle client populations that vary over time and can be implemented in a privacy preserving way. As clustering is only performed after Federated Learning has converged to a stationary point, CFL can be viewed as a post-processing method that will always achieve greater or equal performance than conventional FL by allowing clients to arrive at more specialized models. We verify our theoretical analysis in experiments with deep convolutional and recurrent neural networks on commonly used Federated Learning datasets.


SELF: Learning to Filter Noisy Labels with Self-Ensembling

arXiv.org Machine Learning

Deep neural networks (DNNs) have been shown to over-fit a dataset when being trained with noisy labels for a long enough time. To overcome this problem, we present a simple and effective method self-ensemble label filtering (SELF) to progressively filter out the wrong labels during training. Our method improves the task performance by gradually allowing supervision only from the potentially non-noisy (clean) labels and stops learning on the filtered noisy labels. For the filtering, we form running averages of predictions over the entire training dataset using the network output at different training epochs. We show that these ensemble estimates yield more accurate identification of inconsistent predictions throughout training than the single estimates of the network at the most recent training epoch. While filtered samples are removed entirely from the supervised training loss, we dynamically leverage them via semi-supervised learning in the unsupervised loss. We demonstrate the positive effect of such an approach on various image classification tasks under both symmetric and asymmetric label noise and at different noise ratios. It substantially outperforms all previous works on noise-aware learning across different datasets and can be applied to a broad set of network architectures. The acquisition of large quantities of a high-quality human annotation is a frequent bottleneck in applying DNNs. There are two cheap but imperfect alternatives to collect annotation at large scale: crowdsourcing from non-experts and web annotations, particularly for image data where the tags and online query keywords are treated as valid labels. Both these alternatives typically introduce noisy (wrong) labels. While Rolnick et al. (2017) empirically demonstrated that DNNs can be surprisingly robust to label noise under certain conditions, Zhang et al. (2017) has shown that DNNs have the capacity to memorize the data and will do so eventually when being confronted with too many noisy labels. Consequently, training DNNs with traditional learning procedures on noisy data strongly deteriorates their ability to generalize - a severe problem.


Complete Machine Learning Bootcamp

#artificialintelligence

In this course we will learn and practice all the services of AWS Machine Learning which is being offered by AWS Cloud. There will be both theoretical and practical section of each AWS Machine Learning services.This course is for those who loves machine learning and would build application based on cognitive computing, AI and ML. You could integrate these services in your Web, Android, IoT, Desktop Applications like Face Detection, ChatBot, Voice Detection, Text to custom Speech (with pitch, emotions, etc), Speech to text, Sentimental Analysis on Social media or any textual data. If you have interest in machine learning as well as cloud computing then this course for you. This course will let you use your machine learning skills deploy in cloud.


IBM certifies a much-needed 140 data scientists for AI development

#artificialintelligence

As more companies realize the great need for data scientists to develop, experiment, and deploy artificial intelligence (AI), IBM designed a certification program. It offered it to the company workforce, and incentivized employees completed the program through coursework, skills training, and apprenticeships. IBM's certification and related programs "will speed the journey to AI and help improve business performance, efficiency and growth," said Martin Fleming, IBM vice president and chief economist. The demand for data scientists is recognized in the tech industry which "actually identifies the demand for data scientists as one of the industry's most pressing needs." Fleming cites social-media career platform LinkedIn's 2018 report, which found 151,000 US data scientist positions unfilled. "More companies are looking inward for ways to build the skills among their existing workforces," he said.


Why Red Means Red in Almost Every Language - Issue 76: Language

Nautilus

When Paul Kay, then an anthropology graduate student at Harvard University, arrived in Tahiti in 1959 to study island life, he expected to have a hard time learning the local words for colors. His field had long espoused a theory called linguistic relativity, which held that language shapes perception. Color was the "parade example," Kay says. His professors and textbooks taught that people could only recognize a color as categorically distinct from others if they had a word for it. If you knew only three color words, a rainbow would have only three stripes.


WiMLDS Montreal #4: AI and Entrepreneurship

#artificialintelligence

We're excited to announce our 4th WiMLDS Montreal meetup presented and hosted by BDC! Whether you're a data professional or simply curious, you are welcome regardless of your technical level or your gender. The talks will be entirely in English. Agenda 6:00 pm -- Doors open 6:30 pm -- Talks 7:45 pm -- Panel 8:30 pm -- Networking Opening remarks Amy Pollard -- Analyst @ Strategic Investments & BDC Women in Tech Fund Amy is involved in all aspects of the deal process including sourcing and due diligence. She is also responsible for portfolio management activities including portfolio monitoring, reporting and bi-annual portfolio valuations. Since founding her first entrepreneurial venture as a teenager, Amy has actively volunteered in global entrepreneurial and tech communities through organizations such as Startup Canada, Startup Weekend, Startup Nations and Junior Achievement.


Probability for Machine Learning

#artificialintelligence

This book was designed around major ideas and methods that are directly relevant to machine learning algorithms. There are a lot of things you could learn about probability, from theory to abstract concepts to APIs. My goal is to take you straight to developing an intuition for the elements you must understand with laser-focused tutorials. I designed the tutorials to focus on how to get things done with probability. They give you the tools to both rapidly understand and apply each technique or operation. Each tutorial is designed to take you less than one hour to read through and complete, excluding the extensions and further reading. You can choose to work through the lessons one per day, one per week, or at your own pace. I think momentum is critically important, and this book is intended to be read and used, not to sit idle. I would recommend picking a schedule and sticking to it.