Europe
Estimating mutual information in high dimensions via classification error
Zheng, Charles Y., Benjamini, Yuval
Multivariate pattern analyses approaches in neuroimaging are fundamentally concerned with investigating the quantity and type of information processed by various regions of the human brain; typically, estimates of classification accuracy are used to quantify information. While a extensive and powerful library of methods can be applied to train and assess classifiers, it is not always clear how to use the resulting measures of classification performance to draw scientific conclusions: e.g. for the purpose of evaluating redundancy between brain regions. An additional confound for interpreting classification performance is the dependence of the error rate on the number and choice of distinct classes obtained for the classification task. In contrast, mutual information is a quantity defined independently of the experimental design, and has ideal properties for comparative analyses. Unfortunately, estimating the mutual information based on observations becomes statistically infeasible in high dimensions without some kind of assumption or prior. In this paper, we construct a novel classification-based estimator of mutual information based on high-dimensional asymptotics. We show that in a particular limiting regime, the mutual information is an invertible function of the expected $k$-class Bayes error. While the theory is based on a large-sample, high-dimensional limit, we demonstrate through simulations that our proposed estimator has superior performance to the alternatives in problems of moderate dimensionality.
Condorcet's Jury Theorem for Consensus Clustering and its Implications for Diversity
Condorcet's Jury Theorem has been invoked for ensemble classifiers to indicate that the combination of many classifiers can have better predictive performance than a single classifier. Such a theoretical underpinning is unknown for consensus clustering. This article extends Condorcet's Jury Theorem to the mean partition approach under the additional assumptions that a unique ground-truth partition exists and sample partitions are drawn from a sufficiently small ball containing the ground-truth. As an implication of practical relevance, we question the claim that the quality of consensus clustering depends on the diversity of the sample partitions. Instead, we conjecture that limiting the diversity of the mean partitions is necessary for controlling the quality.
Heuristic Approaches for Generating Local Process Models through Log Projections
Tax, Niek, Sidorova, Natalia, van der Aalst, Wil M. P., Haakma, Reinder
Local Process Model (LPM) discovery is focused on the mining of a set of process models where each model describes the behavior represented in the event log only partially, i.e. subsets of possible events are taken into account to create so-called local process models. Often such smaller models provide valuable insights into the behavior of the process, especially when no adequate and comprehensible single overall process model exists that is able to describe the traces of the process from start to end. The practical application of LPM discovery is however hindered by computational issues in the case of logs with many activities (problems may already occur when there are more than 17 unique activities). In this paper, we explore three heuristics to discover subsets of activities that lead to useful log projections with the goal of speeding up LPM discovery considerably while still finding high-quality LPMs. We found that a Markov clustering approach to create projection sets results in the largest improvement of execution time, with discovered LPMs still being better than with the use of randomly generated activity sets of the same size. Another heuristic, based on log entropy, yields a more moderate speedup, but enables the discovery of higher quality LPMs. The third heuristic, based on the relative information gain, shows unstable performance: for some data sets the speedup and LPM quality are higher than with the log entropy based method, while for other data sets there is no speedup at all.
Google to Hire 1,000 People to Boost Its Cloud Business
Google wants to change how it relates to enterprise customers, and it's going to use Google Cloud to do so. Today at an event in San Francisco, it made announcements about machine learning, Kubernetes, and expansion of its Google Cloud Platform presence. But the bigger-picture news is that cloud will be taking a leading role at the company. Google is bringing together its massive Google Cloud Platform (GCP), along with a new application platform called G Suite (formerly Google Apps). And it's hiring 1,000 people to boost its cloud business. This is all under one big umbrella called Google Cloud.
RBS to pilot its first artificial intelligence with 'chat bot' feature
A'chat bot' is a computer programme designed to simulate an intelligent conversation with human users by way of text or telephone. RBS will use IBM's'Watson' - a platform that analyses unstructured data - to power its service, which will aim to be able to answer specific customer questions, such as'how do I authorise my card to be used overseas?'. In the case of more complex questions, the chat bot will direct customers to a human who can answer them. The bank said it had already tested the servicce among 1,200 RBS and NatWest staff over a two-month trial and now expects the AI bot - if it is successful in its customer pilot - to be rolled out to customers of both brands. In March, RBS announce it would let go 220 investment advisers as part of a cost-cutting drive that would see large parts of its face-to-face service replaced with telephone and online solutions.
Flipboard on Flipboard
You might not be campaigning to be America's next president, or have any desire to hold such a demanding office (bless you, Hillary), but wouldn't it still be nice to be treated like POTUS when you travel? Or, at least spend a few days in the presidential suite feeling like one of the world's most โฆ Election jokes are i Saturday Night Live' /i s bread and butter, so it should come as no surprise that the cast took aim at Donald Trump's hot mic scandal. But host Lin-Manuel Miranda also got a chance to shine in his opening monologue. Below, we've rounded up the must-see moments from last night's /b โฆ Humans may live longer and longer, but eventually we all grow old and die. This leads to a simple question: Is there an intrinsic maximum limit to human lifespan or not?
Spark analytics applications boosted by built-in libraries
At last year's Spark Summit conference, Patrick Wendell, a software engineer at Databricks Inc. and a contributor to the Apache Spark open source project, said the technology's data processing capabilities are impressive but its real power lies in the Spark library components that sit on top of the core engine. "The future of Spark is the libraries," he said. "That's what the community has invested in and where the innovation is coming from." Sure enough, this month's Spark Summit 2015 event prominently featured case studies in which users explained how they're putting the libraries to work in Spark analytics applications. The Spark platform comes with four distinct libraries -- Spark SQL, Spark Streaming, a graph processing library called GraphX and a machine learning one known as MLlib -- that include pre-built algorithms and programming capabilities designed to streamline data preparation, exploration and analysis tasks. The libraries enable users to automate certain tasks and eliminate some of the coding that typically would be required.
Weekend tech reading: 1nm transistor created, Comcast's 1TB cap rolls out, Boeing sets sight on Mars
For more than a decade, engineers have been eyeing the finish line in the race to shrink the size of components in integrated circuits. They knew that the laws of physics had set a 5-nanometer threshold on the size of transistor gates among conventional semiconductors, about one-quarter the size of high-end 20-nanometer-gate transistors now on the market. A research team led by faculty scientist Ali Javey at the Department of Energy's Lawrence Berkeley National Laboratory (Berkeley Lab) has done just that by creating a transistor with a working 1-nanometer gate. Boeing CEO vows to beat Musk to Mars Boeing Co. once helped the U.S. beat the Soviet Union in the race to the moon. Now the company intends to go toe-to-toe with newcomers such as billionaire Elon Musk in the next era of space exploration and commerce.
Mobile, sun-seeking gardens, and more in the week that was
The Fisker Karma was one of the world's hottest plug-in hybrid supercars when it debuted in 2011 - and now its creator Henrik Fisker has announced plans to launch an electric sports car with a 400-mile range next year. Meanwhile, Mercedes is taking aim at the Tesla Model X with its new Generation EQ SUV, which touts 400 horsepower and an all-electric driving range of 300 miles. The International Space Station is getting ready to test a brand new ion thruster that can be powered by space junk, and teenage inventor Boyan Slat has modified a C-130 Hercules aircraft with high-tech sensors to spot plastic debris in the Great Pacific Garbage Patch. Solar power is getting cheaper by the day - and two groundbreaking projects in China and Abu Dhabi have pushed the price down 25% in just five months. In other energy news, Poland just unveiled a glowing, bright blue bike lane that's charged by the sun.
Morning roundup of Artificial Intelligence news for October 9, 2016
Samsung has acquired the Viv AI Assistant creators of the Siri. The famous Apple assistant was co-founded by Dag Kittlaus, Adam Cheyer, and Chris Brigham. Apple acquired Siri in 2010 but the trio left Apple in 2012. Since then, they founded Viv, which continues to operate as an independent company. Viv is being touted as the global brain, thanks to its intelligent interface.