Genre
Target contrastive pessimistic risk for robust domain adaptation
In domain adaptation, classifiers with information from a source domain adapt to generalize to a target domain. However, an adaptive classifier can perform worse than a non-adaptive classifier due to invalid assumptions, increased sensitivity to estimation errors or model misspecification. Our goal is to develop a domain-adaptive classifier that is robust in the sense that it does not rely on restrictive assumptions on how the source and target domains relate to each other and that it does not perform worse than the non-adaptive classifier. We formulate a conservative parameter estimator that only deviates from the source classifier when a lower risk is guaranteed for all possible labellings of the given target samples. We derive the classical least-squares and discriminant analysis cases and show that these perform on par with state-of-the-art domain adaptive classifiers in sample selection bias settings, while outperforming them in more general domain adaptation settings.
In Search of an Entity Resolution OASIS: Optimal Asymptotic Sequential Importance Sampling
Marchant, Neil G., Rubinstein, Benjamin I. P.
Entity resolution (ER) presents unique challenges for evaluation methodology. While crowdsourcing platforms acquire ground truth, sound approaches to sampling must drive labelling efforts. In ER, extreme class imbalance between matching and non-matching records can lead to enormous labelling requirements when seeking statistically consistent estimates for rigorous evaluation. This paper addresses this important challenge with the OASIS algorithm: a sampler and F-measure estimator for ER evaluation. OASIS draws samples from a (biased) instrumental distribution, chosen to ensure estimators with optimal asymptotic variance. As new labels are collected OASIS updates this instrumental distribution via a Bayesian latent variable model of the annotator oracle, to quickly focus on unlabelled items providing more information. We prove that resulting estimates of F-measure, precision, recall converge to the true population values. Thorough comparisons of sampling methods on a variety of ER datasets demonstrate significant labelling reductions of up to 83% without loss to estimate accuracy.
Object Boundary Detection and Classification with Image-level Labels
Koh, Jing Yu, Samek, Wojciech, Mรผller, Klaus-Robert, Binder, Alexander
Semantic boundary and edge detection aims at simultaneously detecting object edge pixels in images and assigning class labels to them. Systematic training of predictors for this task requires the labeling of edges in images which is a particularly tedious task. We propose a novel strategy for solving this task, when pixel-level annotations are not available, performing it in an almost zero-shot manner by relying on conventional whole image neural net classifiers that were trained using large bounding boxes. Our method performs the following two steps at test time. Firstly it predicts the class labels by applying the trained whole image network to the test images. Secondly, it computes pixel-wise scores from the obtained predictions by applying backprop gradients as well as recent visualization algorithms such as deconvolution and layer-wise relevance propagation. We show that high pixel-wise scores are indicative for the location of semantic boundaries, which suggests that the semantic boundary problem can be approached without using edge labels during the training phase.
Preserving Differential Privacy in Convolutional Deep Belief Networks
Phan, NhatHai, Wu, Xintao, Dou, Dejing
The remarkable development of deep learning in medicine and healthcare domain presents obvious privacy issues, when deep neural networks are built on users' personal and highly sensitive data, e.g., clinical records, user profiles, biomedical images, etc. However, only a few scientific studies on preserving privacy in deep learning have been conducted. In this paper, we focus on developing a private convolutional deep belief network (pCDBN), which essentially is a convolutional deep belief network (CDBN) under differential privacy. Our main idea of enforcing epsilon-differential privacy is to leverage the functional mechanism to perturb the energy-based objective functions of traditional CDBNs, rather than their results. One key contribution of this work is that we propose the use of Chebyshev expansion to derive the approximate polynomial representation of objective functions. Our theoretical analysis shows that we can further derive the sensitivity and error bounds of the approximate polynomial representation. As a result, preserving differential privacy in CDBNs is feasible. We applied our model in a health social network, i.e., YesiWell data, and in a handwriting digit dataset, i.e., MNIST data, for human behavior prediction, human behavior classification, and handwriting digit recognition tasks. Theoretical analysis and rigorous experimental evaluations show that the pCDBN is highly effective. It significantly outperforms existing solutions.
Large-scale Validation of Counterfactual Learning Methods: A Test-Bed
Lefortier, Damien, Swaminathan, Adith, Gu, Xiaotao, Joachims, Thorsten, de Rijke, Maarten
The ability to perform effective off-policy learning would revolutionize the process of building better interactive systems, such as search engines and recommendation systems for e-commerce, computational advertising and news. Recent approaches for off-policy evaluation and learning in these settings appear promising. With this paper, we provide real-world data and a standardized test-bed to systematically investigate these algorithms using data from display advertising. In particular, we consider the problem of filling a banner ad with an aggregate of multiple products the user may want to purchase. This paper presents our test-bed, the sanity checks we ran to ensure its validity, and shows results comparing state-of-the-art off-policy learning methods like doubly robust optimization, POEM, and reductions to supervised learning using regression baselines. Our results show experimental evidence that recent off-policy learning methods can improve upon state-of-the-art supervised learning techniques on a large-scale real-world data set.
Startup Founder's Quest for Cure Leads to Genomics Hackathon at Google Xconomy
This story is part of a series on A.I. in healthcare. Onno Faber was a member of Silicon Valley's happy breed of tech startup founders when he was diagnosed with a rare genetic condition that can come with dire health damage, but few treatments. Faber responded with entrepreneurial zeal, exploring whether Silicon Valley's mastery of algorithms might help root out and defeat the threatening quirks in his genetic code. Without any ready-made solutions on hand from big drug companies and their established research teams, Faber started to recruit individuals to his cause. The results of Faber's crusade so far demonstrate a trait Silicon Valley has in common with living things--a startling talent for self-organization.
Jack Ma predicts AI will dramatically reduce our workload
It's good news for people who find work a drag, as Alibaba founder Jack Ma believes we will work just four hours a day for four days a week by 2047. The Chinese billionaire said he believed people would reap the benefits of artificial intelligence (AI) and be free to spend more time travelling and less time working. But the 52-year-old founder of Alibaba - China's equivalent of eBay - also warned'there's going to be trouble' with AI in the future unless governments move fast. The Chinese billionaire said he believed people would reap the benefits of artificial intelligence (AI) and be free to spend more time travelling and less time working. The 52-year-old Alibaba founder - China's equivalent of eBay - warned'there's going to be trouble' with AI in the future unless governments move fast.
Facebook, Baidu, Coach And 5 Other Blue-Chip Stocks To Buy For Second-Half 2017
Opinions expressed by Forbes Contributors are their own. The author is a Forbes contributor. The opinions expressed are those of the writer. In terms of stock market years, the bull market that kicked off in March 2009 should be on the front cover of AARP magazine. At age 8, the S&P 500 bull run is very old and very overvalued compared to its historical ratios.
The Man Who Helped Turn Toronto Into a High-Tech Hotbed
His impact on artificial intelligence research has been so deep that some people in the field talk about the "six degrees of Geoffrey Hinton" the way college students once referred to Kevin Bacon's uncanny connections to so many Hollywood movies. Dr. Hinton's students and associates are now leading lights of artificial intelligence research at Apple, Facebook, Google and Uber, and run artificial intelligence programs at the University of Montreal and OpenAI, a nonprofit research company. "Geoff, at a time when A.I. was in the wilderness, toiled away at building the field and because of his personality, attracted people who then dispersed," said Ilse Treurnicht, chief executive of Toronto's MaRS Discovery District, an innovation center that will soon house the Vector Institute, Toronto's new public-private artificial intelligence research institute, where Dr. Hinton will be chief scientific adviser. Dr. Hinton also recently set up a Toronto branch of Google Brain, the company's artificial intelligence research project. His tiny office there is not the grand space filled with gadgets and awards that one might expect for a man at the leading edge of the most transformative field of science today.
Real-time Twitter sentiment analysis with Azure Stream Analytics
Learn how to build a sentiment analysis solution for social media analytics by bringing real-time Twitter events into Azure Event Hubs. In this scenario, you write an Azure Stream Analytics query to analyze the data. Then you either store the results for later use or use a dashboard and Power BI to provide insights in real time. Social media analytics tools help organizations understand trending topics. Trending topics are subjects and attitudes that have a high volume of posts in social media.