Genre
OpenAI bot remains undefeated against world's greatest Dota 2 players
Last night, OpenAI's Dota 2 bot beat the world's most celebrated professional players in one-on-one battles, showing just how advanced these machine learning systems are getting. The bot beat Danil "Dendi" Ishutin rather easily at The International, one of the biggest eSports events in the world, and remains undefeated against the world's top Dota 2 players. Elon Musk's OpenAI trained the bot by simply copying the AI and letting the two play each other for weeks on end. "We've coached it to learn just from playing against itself," said OpenAI researcher Jakub Pachoki. "So we didn't hard-code in any strategy, we didn't have it learn from human experts, just from the very beginning, it just keeps playing against a copy of itself. It starts from complete randomness and then it makes very small improvements, and eventually it's just pro level."
Why Education Is the Hardest Sector of the Economy to Automate
We've all heard the warning cries: automation will disrupt entire industries and put millions of people out of jobs. In fact, up to 45 percent of existing jobs can be automated using current technology. However, this may not necessarily apply to the education sector. After a detailed analysis of more than 2,000-plus work activities for more than 800 occupations, a report by McKinsey & Co states that of all the sectors examined, "โฆthe technical feasibility of automation is lowest in education." There is no doubt that technological trends will have a powerful impact on global education, both by improving the overall learning experience and by increasing global access to education.
China creates speed cameras that ID cars' scratches
The next generation of spy cameras are set to catch speeding drivers and hit them with fines - without even needing to check their number plates first. These roadside cameras could catch unruly drivers just by looking at scratches on their car or irregularities in the paintwork. The'repression network' uses artificial intelligence to differentiate between cars by spotting tiny differences between them. Researchers say their software is so sophisticated it could someday also be used for'face and persona retrieval'. This graphic shows how the system can identify cars based on their features, rather than a number plate.
Artificial Intelligence And Its Impact On Legal Technology (Part II)
Artificial intelligence (AI) is quickly coming into its own in terms of use by the legal industry. We are on the cusp of a revolution in the legal profession led by the adoption of AI throughout the legal industry, but in particular by in-house lawyers. Much like how email changed the way we do business every day, AI will become ubiquitous -- an indispensable assistant to practically every lawyer. But what is the future of AI in the legal industry? A bigger question is whether AI will actually replace lawyers as seems to be implicated above (a scary thought if you are new to the profession vs. an old-timer like me).
Book review: The Mathematical Corporation: Where Machine Intelligence and Human Ingenuity Achieve the Impossible
I heard about The Mathematical Corporation: Where Machine Intelligence and Human ... by Josh Sullivan and Angela Zutavern through a tweet by Kirk Bourne. I conduct a course at Oxford University on Data Science for IoT. I have also recently launched a course on AI for fintech. The Mathematical Corporation covers many issues that I have encountered in my teaching and consulting. In essence, Mathematical Corporation calls for new leadership traits in the world of AI.
Semi-supervised emotion lexicon expansion with label propagation and specialized word embeddings
There exist two main approaches to automatically extract affective orientation: lexicon-based and corpus-based. In this work, we argue that these two methods are compatible and show that combining them can improve the accuracy of emotion classifiers. In particular, we introduce a novel variant of the Label Propagation algorithm that is tailored to distributed word representations, we apply batch gradient descent to accelerate the optimization of label propagation and to make the optimization feasible for large graphs, and we propose a reproducible method for emotion lexicon expansion. We conclude that label propagation can expand an emotion lexicon in a meaningful way and that the expanded emotion lexicon can be leveraged to improve the accuracy of an emotion classifier.
Sentiment Analysis by Joint Learning of Word Embeddings and Classifier
Sarma, Prathusha Kameswara, Sethares, Bill
Word embeddings are representations of individual words of a text document in a vector space and they are often use- ful for performing natural language pro- cessing tasks. Current state of the art al- gorithms for learning word embeddings learn vector representations from large corpora of text documents in an unsu- pervised fashion. This paper introduces SWESA (Supervised Word Embeddings for Sentiment Analysis), an algorithm for sentiment analysis via word embeddings. SWESA leverages document label infor- mation to learn vector representations of words from a modest corpus of text doc- uments by solving an optimization prob- lem that minimizes a cost function with respect to both word embeddings as well as classification accuracy. Analysis re- veals that SWESA provides an efficient way of estimating the dimension of the word embeddings that are to be learned. Experiments on several real world data sets show that SWESA has superior per- formance when compared to previously suggested approaches to word embeddings and sentiment analysis tasks.
Mahalanonbis Distance Informed by Clustering
Lahav, Almog, Talmon, Ronen, Kluger, Yuval
A fundamental question in data analysis, machine learning and signal processing is how to compare between data points. The choice of the distance metric is specifically challenging for high-dimensional data sets, where the problem of meaningfulness is more prominent (e.g. the Euclidean distance between images). In this paper, we propose to exploit a property of high-dimensional data that is usually ignored - which is the structure stemming from the relationships between the coordinates. Specifically we show that organizing similar coordinates in clusters can be exploited for the construction of the Mahalanobis distance between samples. When the observable samples are generated by a nonlinear transformation of hidden variables, the Mahalanobis distance allows the recovery of the Euclidean distances in the hidden space.We illustrate the advantage of our approach on a synthetic example where the discovery of clusters of correlated coordinates improves the estimation of the principal directions of the samples. Our method was applied to real data of gene expression for lung adenocarcinomas (lung cancer). By using the proposed metric we found a partition of subjects to risk groups with a good separation between their Kaplan-Meier survival plot.
Model-Based Multiple Instance Learning
Vo, Ba-Ngu, Phung, Dinh, Tran, Quang N., Vo, Ba-Tuong
While Multiple Instance (MI) data are point patterns -- sets or multi-sets of unordered points -- appropriate statistical point pattern models have not been used in MI learning. This article proposes a framework for model-based MI learning using point process theory. Likelihood functions for point pattern data derived from point process theory enable principled yet conceptually transparent extensions of learning tasks, such as classification, novelty detection and clustering, to point pattern data. Furthermore, tractable point pattern models as well as solutions for learning and decision making from point pattern data are developed.
Reprogramming Matter, Life, and Purpose
Reprogramming matter may sound far-fetched, but we have been doing it with increasing power and staggering efficiency for at least 60 years, and for centuries we have been paving the way toward the ultimate reprogrammed fate of the universe, the vessel of all programs. How will we be doing it in 60 years' time and how will it impact life and the purpose both of machines and of humans?