Goto

Collaborating Authors

 Oceania


Random forest model identifies serve strength as a key predictor of tennis match outcome

arXiv.org Machine Learning

Tennis is a popular sport worldwide, boasting millions of fans and numerous national and international tournaments. Like many sports, tennis has benefitted from the popularity of rigorous record-keeping of game and player information, as well as the growth of machine learning methods for use in sports analytics. Of particular interest to bettors and betting companies alike is potential use of sports records to predict tennis match outcomes prior to match start. We compiled, cleaned, and used the largest database of tennis match information to date to predict match outcome using fairly simple machine learning methods. Using such methods allows for rapid fit and prediction times to readily incorporate new data and make real-time predictions. We were able to predict match outcomes with upwards of 80% accuracy, much greater than predictions using betting odds alone, and identify serve strength as a key predictor of match outcome. By combining prediction accuracies from three models, we were able to nearly recreate a probability distribution based on average betting odds from betting companies, which indicates that betting companies are using similar information to assign odds to matches. These results demonstrate the capability of relatively simple machine learning models to quite accurately predict tennis match outcomes.


Sim-to-Real Transfer of Robot Learning with Variable Length Inputs

arXiv.org Machine Learning

Current end-to-end deep Reinforcement Learning (RL) approaches require jointly learning perception, decision-making and low-level control from very sparse reward signals and high-dimensional inputs, with little capability of incorporating prior knowledge. This results in prohibitively long training times for use on real-world robotic tasks. Existing algorithms capable of extracting task-level representations from high-dimensional inputs, e.g. object detection, often produce outputs of varying lengths, restricting their use in RL methods due to the need for neural networks to have fixed length inputs. In this work, we propose a framework that combines deep sets encoding, which allows for variable-length abstract representations, with modular RL that utilizes these representations, decoupling high-level decision making from low-level control. We successfully demonstrate our approach on the robot manipulation task of object sorting, showing that this method can learn effective policies within mere minutes of highly simplified simulation. The learned policies can be directly deployed on a robot without further training, and generalize to variations of the task unseen during training.


Knowledge-based Biomedical Data Science 2019

arXiv.org Artificial Intelligence

Knowledge-based biomedical data science (KBDS) involves the design and implementation of computer systems that act as if they knew about biomedicine. Such systems depend on formally represented knowledge in computer systems, often in the form of knowledge graphs. Here we survey the progress in the last year in systems that use formally represented knowledge to address data science problems in both clinical and biological domains, as well as on approaches for creating knowledge graphs. Major themes include the relationships between knowledge graphs and machine learning, the use of natural language processing, and the expansion of knowledge-based approaches to novel domains, such as Chinese Traditional Medicine and biodiversity.


Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models

arXiv.org Artificial Intelligence

Existing dialog system models require extensive human annotations and are difficult to generalize to different tasks. The recent success of large pre-trained language models such as BERT and GPT -2 (Devlin et al., 2019; Radford et al., 2019) have suggested the effectiveness of incorporating language priors in downstream NLP tasks. However, how much pre-trained language models can help dialog response generation is still under exploration. In this paper, we propose a simple, general, and effective framework: Alternating Recurrent Dialog Model (ARDM). ARDM models each speaker separately and takes advantage of the large pre-trained language model. It requires no supervision from human annotations such as belief states or dialog acts to achieve effective conversations. ARDM outperforms or is on par with state-of-the-art methods on two popular task-oriented dialog datasets: CamRest676 and MultiWOZ. Moreover, we can generalize ARDM to more challenging, non-collaborative tasks such as persuasion. In persuasion tasks, ARDM is capable of generating humanlike responses to persuade people to donate to a charity. It has been a longstanding ambition for artificial intelligence researchers to create an intelligent conversational agent that can generate humanlike responses. Recently data-driven dialog models are more and more popular. However, most current state-of-the-art approaches still rely heavily on extensive annotations such as belief states and dialog acts (Lei et al., 2018). However, dialog content can vary considerably in different dialog tasks. Having a different intent or dialog act annotation scheme for each task is costly. For some tasks, it is even impossible, such as open-domain social chat. Thus, it is difficult to utilize these methods on challenging dialog tasks, such as persuasion and negotiation, where dialog states and acts are difficult to annotate.


Proof-of-learning: A blockchain consensus mechanism based on machine learning competitions

#artificialintelligence

This article presents WekaCoin, a peer-to-peer cryptocurrency based on a new distributed consensus protocol called Proof-of-Learning. Proof-of-learning achieves distributed consensus by ranking machine learning systems for a given task. The aim of this protocol is to alleviate the computational waste involved in hashing-based puzzles and to create a public distributed and verifiable database of state-of-the-art machine learning models and experiments.


SiteSee deploys ContextCapture to model communication towers

#artificialintelligence

US: Telecommunication infrastructure owners have some of the most widely distributed and remote assets to build, maintain, and repair. With approximately 27,000 distributed sites and 6,000 communication towers under management, Telstra is Australia's leading telecommunication service provider. Traditional inspection methods of towers involve manually taking photographs and measurements. This requires workers climbing on the towers, usually in remote areas, making the process dangerous, inefficient, costly, and time-consuming. In 2017, Telstra engaged SiteSee to perform automated as-built and condition assessment reports by applying machine learning and object recognition technology to 3D reality meshes.


Computer game to assist clinicians in diagnosing mental health disorders

#artificialintelligence

A team of researchers led by CSIRO's Data61, the data and digital specialist arm of Australia's national science agency, have developed a novel technique that could assist psychiatrists and other clinicians to diagnose and characterize complex mental health disorders, potentially enabling more effective treatments. Announced today at D61 LIVE in Sydney, the researchers revealed that using a simple computer game and artificial intelligence techniques, they were able to identify behavioral patterns in subjects with depression and bipolar disorder, down to subtle individual differences in each group. The study included 101 participants: 34 with depression, 33 with bipolar disorder, and a control group of 34 subjects. The computer game presents individuals with two choices, and tracks their behavior as they respond. The complex data collected from the game is analyzed through artificial neural networks--brain-inspired systems intended to replicate the way that humans learn--which are able to disentangle the nuanced behavioral differences between healthy individuals, and those with depression or bipolar disorder.


The 10 governments leading in behavioural science Apolitical

#artificialintelligence

The use of "nudges" in policymaking has been a major trend since the UK launched the world's first government-embedded behavioural insights unit in 2010. But governments around the world, from Denmark to Singapore, have been using principles from behavioural science to influence citizens since at least the 1960s. That's according to a new World Bank report, Behavioural Science Around the World, which highlights 10 countries that are pioneering the use of behavioural insights: Australia, Canada, Denmark, France, Germany, the Netherlands, Peru, Singapore, the UK and the US. The World Bank report looks at how these teams are integrated into government, which projects they're working on and how they are run -- and, most importantly, which experiments have worked. It predicts that in the future, behavioural insights units will benefit from artificial intelligence, machine learning and virtual reality the same way they've gained from advancements in open data and e-government.


Fluid Flow Mass Transport for Generative Networks

arXiv.org Machine Learning

Generative Adversarial Networks have been shown to be powerful in generating content. To this end, they have been studied intensively in the last few years. Nonetheless, training these networks requires solving a saddle point problem that is difficult to solve and slowly converging. Motivated from techniques in the registration of point clouds and by the fluid flow formulation of mass transport, we investigate a new formulation that is based on strict minimization, without the need for the maximization. The formulation views the problem as a matching problem rather than an adversarial one and thus allows us to quickly converge and obtain meaningful metrics in the optimization path.


Sequence embeddings help to identify fraudulent cases in healthcare insurance

arXiv.org Machine Learning

Fraud causes substantial costs and losses for companies and clients in the finance and insurance industries. Examples are fraudulent credit card transactions or fraudulent claims. It has been estimated that roughly $10$ percent of the insurance industry's incurred losses and loss adjustment expenses each year stem from fraudulent claims. The rise and proliferation of digitization in finance and insurance have lead to big data sets, consisting in particular of text data, which can be used for fraud detection. In this paper, we propose architectures for text embeddings via deep learning, which help to improve the detection of fraudulent claims compared to other machine learning methods. We illustrate our methods using a data set from a large international health insurance company. The empirical results show that our approach outperforms other state-of-the-art methods and can help make the claims management process more efficient. As (unstructured) text data become increasingly available to economists and econometricians, our proposed methods will be valuable for many similar applications, particularly when variables have a large number of categories as is typical for example of the International Classification of Disease (ICD) codes in health economics and health services.