Goto

Collaborating Authors

 Genre


Provable Bayesian Inference via Particle Mirror Descent

arXiv.org Machine Learning

Bayesian methods are appealing in their flexibility in modeling complex data and ability in capturing uncertainty in parameters. However, when Bayes' rule does not result in tractable closed-form, most approximate inference algorithms lack either scalability or rigorous guarantees. To tackle this challenge, we propose a simple yet provable algorithm, \emph{Particle Mirror Descent} (PMD), to iteratively approximate the posterior density. PMD is inspired by stochastic functional mirror descent where one descends in the density space using a small batch of data points at each iteration, and by particle filtering where one uses samples to approximate a function. We prove result of the first kind that, with $m$ particles, PMD provides a posterior density estimator that converges in terms of $KL$-divergence to the true posterior in rate $O(1/\sqrt{m})$. We demonstrate competitive empirical performances of PMD compared to several approximate inference algorithms in mixture models, logistic regression, sparse Gaussian processes and latent Dirichlet allocation on large scale datasets.


Information Recovery from Pairwise Measurements

arXiv.org Machine Learning

This paper is concerned with jointly recovering $n$ node-variables $\left\{ x_{i}\right\}_{1\leq i\leq n}$ from a collection of pairwise difference measurements. Imagine we acquire a few observations taking the form of $x_{i}-x_{j}$; the observation pattern is represented by a measurement graph $\mathcal{G}$ with an edge set $\mathcal{E}$ such that $x_{i}-x_{j}$ is observed if and only if $(i,j)\in\mathcal{E}$. To account for noisy measurements in a general manner, we model the data acquisition process by a set of channels with given input/output transition measures. Employing information-theoretic tools applied to channel decoding problems, we develop a \emph{unified} framework to characterize the fundamental recovery criterion, which accommodates general graph structures, alphabet sizes, and channel transition measures. In particular, our results isolate a family of \emph{minimum} \emph{channel divergence measures} to characterize the degree of measurement corruption, which together with the size of the minimum cut of $\mathcal{G}$ dictates the feasibility of exact information recovery. For various homogeneous graphs, the recovery condition depends almost only on the edge sparsity of the measurement graph irrespective of other graphical metrics; alternatively, the minimum sample complexity required for these graphs scales like \[ \text{minimum sample complexity }\asymp\frac{n\log n}{\mathsf{Hel}_{1/2}^{\min}} \] for certain information metric $\mathsf{Hel}_{1/2}^{\min}$ defined in the main text, as long as the alphabet size is not super-polynomial in $n$. We apply our general theory to three concrete applications, including the stochastic block model, the outlier model, and the haplotype assembly problem. Our theory leads to order-wise tight recovery conditions for all these scenarios.


Multilingual Twitter Sentiment Classification: The Role of Human Annotators

arXiv.org Artificial Intelligence

What are the limits of automated Twitter sentiment classification? We analyze a large set of manually labeled tweets in different languages, use them as training data, and construct automated classification models. It turns out that the quality of classification models depends much more on the quality and size of training data than on the type of the model trained. Experimental results indicate that there is no statistically significant difference between the performance of the top classification models. We quantify the quality of training data by applying various annotator agreement measures, and identify the weakest points of different datasets. We show that the model performance approaches the inter-annotator agreement when the size of the training set is sufficiently large. However, it is crucial to regularly monitor the self- and inter-annotator agreements since this improves the training datasets and consequently the model performance. Finally, we show that there is strong evidence that humans perceive the sentiment classes (negative, neutral, and positive) as ordered.


Data intelligence startup Ripjar secures 'significant investment' from Winton Ventures

#artificialintelligence

Data intelligence platform Ripjar, founded by former GCHQ engineers, has raised an undisclosed sum from Winton Ventures, the VC capital arm of the Winton Group. The firm, headquartered in Cheltenham, allows organisations to explore, visualise and prioritise real-time information from bulks of structured and unstructured data. To do so, the firm applies natural language processing, machine learning techniques and visual analytics. Tom Griffin, CEO of Ripjar said in a statement: "Data capture and analysis is becoming increasingly important to the day-to-day operations of organisations of all sizes. There is a need for new, advanced tools that facilitate this process and, consequently, enable organisations to focus their own resources elsewhere. "Winton's backing will play a vital role in our continued strategic growth ambitions and enable us to increase our customer base across a variety of new business verticals," he added. Owen McCormack, director of Winton Ventures, commented: "We are excited by the potential of Ripjar's profound data analytics knowledge, which is a central component that underpins and complements our own business, previous investments and partnerships.


Rich and powerful warn robots are coming for your jobs

#artificialintelligence

"Most of the benefits we see from automation is about higher quality and fewer errors, but in many cases it does reduce labor," Michael Chui, a partner at the McKinsey Global Institute, said on Tuesday during a panel on "Is Any Job Truly Safe?" The four-day annual conference, which began on Sunday, has 3,500 invite-only participants exploring "The Future of Human Kind." Technology has not only done away with low-wage, low-skill jobs, some of the more than 700 speakers said. They cited robots operating trucks in some Australian mines; corporate litigation software replacing employees with advanced degrees who used to sift through thousands of documents prior to trials; and on Wall Street, the automation of jobs previously done by bankers with MBAs or PhDs. "Anyone whose job is moving data from one spreadsheet to another ..., that's what is going to get automated," said Daniel Nadler, chief executive of Kensho, a financial services analytics company partly owned by Goldman Sachs Group Inc. "Goldman Sachs will be in here in 10 years, JPMorgan will be here. They're just going to be much more efficient in terms of operating leverage and headcount," he added.


Rich and powerful warn robots are coming for your jobs

#artificialintelligence

Some of the most powerful people in the world have gathered this week to discuss the most pressing issues affecting humanity. And the overwhelming conclusion is that the robots are coming. At the Milken Institute's Global Conference in California, at least four panels focused ontechnology taking over markets to mining, and most importantly,jobs. Some of the most powerful people in the world have gathered this week to discuss the most pressing issues affecting humanity, and the overwhelming conclusion is the robots are coming. At the Milken Institute's Global Conference in California, four panels focused on technology taking over markets and jobs (stock image) 'Most of the benefits we see from automation is about higherquality and fewer errors, but in many cases it does reducelabor,' Michael Chui, a partner at the McKinsey GlobalInstitute, said on Tuesday during a panel on'Is Any Job TrulySafe?'


Infosys Launches Mana โ€“ a Knowledge-based Artificial Intelligence Platform - Home

#artificialintelligence

Infosys announced the launch of Infosys Mana, a platform that brings machine learning together with the deep knowledge of an organization, to drive automation and innovation โ€“ enabling businesses to continuously reinvent their system landscapes. According to the press release, Mana, with the Infosys Aikido service offerings, dramatically lowers the cost of maintenance for both physical and digital assets; captures the knowledge and know-how of people, and fragmented and complex systems; simplifies the continuous renovation of core business processes; and enables businesses to bring new and delightful user experiences leveraging state of the art technology. Over the last 35 years, Infosys has maintained, operated and managed systems with global clients across every industry. Building on this deep experience, Infosys has recognized the need to bring artificial intelligence to the enterprise in a meaningful and purposeful way; in a way that leverages the power of automation for repetitive tasks and frees people to focus on the higher value work, and on breakthrough innovation. Today's AI technologies address part of this with learning and information.


Google teams up with Fiat Chrysler to launch a range of self-driving Pacifica minivans

Daily Mail - Science & tech

Fiat Chrysler and Google will work together to more than double the size of Google's self-driving vehicle fleet by adding 100 Chrysler Pacifica minivans. The companies announced the agreement on Tuesday, saying that Chrysler engineers would work with Google to install sensors and software so the vans can drive themselves. The added vehicles are needed as Google expands real-world testing. Fiat Chrysler and Google will work together to more than double the size of Google's self-driving vehicle fleet by adding 100 Chrysler Pacifica (pictured) minivans. Google says it will own the gas-electric hybrid vans, and it's not currently licensing autonomous car technology to Fiat Chrysler or anyone else.


How Companies Are Using Machine Learning to Get Faster and More Efficient

#artificialintelligence

Machine-reengineering is a way to automate business processes using machine learning. Although machine-reengineering is new, companies are already seeing striking results with it, particularly in boosts to speed and efficiency. Studying 168 early adopters, we've seen speed improvements of two times or more for most business processes -- and some organizations are reporting speed improvements of 10 times or more. How do companies do it? Our study found that organizations are using machine-reengineering to establish new forms of human-machine collaboration that break through the bottlenecks of complex digital processes.


Google And Fiat Chrysler Partner To Make Self-Driving Minivans

NPR Technology

Google announced it is partnering with Fiat Chrysler Automobiles to expand its self-driving car project. This is the first time Google has worked directly with an automaker to integrate its self-driving technology into a passenger vehicle. Google announced it is partnering with Fiat Chrysler Automobiles to expand its self-driving car project. This is the first time Google has worked directly with an automaker to integrate its self-driving technology into a passenger vehicle. Google is partnering with Fiat Chrysler Automobiles to expand its self-driving car project, the companies said Tuesday in a joint press release.