Goto

Collaborating Authors

 Genre


Combining policy gradient and Q-learning

arXiv.org Artificial Intelligence

Policy gradient is an efficient technique for improving a policy in a reinforcement learning setting. However, vanilla online variants are on-policy only and not able to take advantage of off-policy data. In this paper we describe a new technique that combines policy gradient with off-policy Q-learning, drawing experience from a replay buffer. This is motivated by making a connection between the fixed points of the regularized policy gradient algorithm and the Q-values. This connection allows us to estimate the Q-values from the action preferences of the policy, to which we apply Q-learning updates. We refer to the new technique as 'PGQL', for policy gradient and Q-learning. We also establish an equivalency between action-value fitting techniques and actor-critic algorithms, showing that regularized policy gradient techniques can be interpreted as advantage function learning algorithms. We conclude with some numerical examples that demonstrate improved data efficiency and stability of PGQL. In particular, we tested PGQL on the full suite of Atari games and achieved performance exceeding that of both asynchronous advantage actor-critic (A3C) and Q-learning.


Pareto Optimality and Strategy Proofness in Group Argument Evaluation (Extended Version)

arXiv.org Artificial Intelligence

An inconsistent knowledge base can be abstracted as a set of arguments and a defeat relation among them. There can be more than one consistent way to evaluate such an argumentation graph. Collective argument evaluation is the problem of aggregating the opinions of multiple agents on how a given set of arguments should be evaluated. It is crucial not only to ensure that the outcome is logically consistent, but also satisfies measures of social optimality and immunity to strategic manipulation. This is because agents have their individual preferences about what the outcome ought to be. In the current paper, we analyze three previously introduced argument-based aggregation operators with respect to Pareto optimality and strategy proofness under different general classes of agent preferences. We highlight fundamental trade-offs between strategic manipulability and social optimality on one hand, and classical logical criteria on the other. Our results motivate further investigation into the relationship between social choice and argumentation theory. The results are also relevant for choosing an appropriate aggregation operator given the criteria that are considered more important, as well as the nature of agents' preferences.




Uncharted 4 wins best game at Baftas awards

BBC News

Uncharted 4 has won the best game at this year's Bafta Games Awards. Developers from its studio, Naughty Dog, said it was "unexpected" that the action adventure title had won, having missed out on the other seven categories it had been nominated for. Chaotic restaurant kitchen game Overcooked took the prize for best British game and family title. The puzzle-platformer Inside had four wins, the most of any game at the London ceremony. It took original property, artistic achievement, game design and narrative.


Is THIS how memories are stored in the brain?

Daily Mail - Science & tech

When we visit a friend or go to the beach, our brain stores a short-term memory of the experience in a part of the brain called the hippocampus. Those memories are later consolidated and transferred to another part of the brain for long-term storage. Now a new study has revealed, for the first time, that memories are actually formed simultaneously in the hippocampus and a long-term storage location in the brain called the cortex. However, the long-term memories remain'silent' for about two weeks before they'mature' over time. The researchers claim that their findings could change how we understand and treat memory disorders such as PTSD and amnesia.


Factoring Massive Numbers: Machine Learning Approach

@machinelearnbot

We are interested here in factoring numbers that are a product of two very large primes. Such numbers are used by encryption algorithms such as RSA, and the prime factors represent the keys (public and private) of the encryption code. Here you will also learn how data science techniques are applied to big data, including visualization, to derive insights. This article is good reading for the data scientist in training, who might not necessarily have easy access to interesting data: here the dataset is the set of all real numbers -- not just the integers -- and it is readily available to anyone. Much of the analysis performed here is statistical in nature, and thus, of particular interest to data scientists.


Transcepta Network Leverages Artificial Intelligence to Transform Accounts Payable and Procurement Processes

#artificialintelligence

ALISO VIEJO, CA--(Marketwired - Apr 6, 2017) - Transcepta, an artificial intelligence-driven platform providing brands with faster, smarter and more accurate procure-to-pay connectivity and collaboration, offers a range of solutions powered by artificial intelligence (AI) that deliver efficiency and profit gains for clients. The company's supplier network and associated applications are part of the broader trend of AI shaping not only high-tech applications, but also changing how companies make decisions. Transcepta offers a suite of services for clients and their suppliers that use machine learning and predictive analytics to improve accuracy, transaction delivery speed and decision making. Transcepta leverages predictive analytics and machine learning to create systems that are self-improving through the continual addition of data inputs over time. Its services are helping to streamline processes and boost the bottom line for some of today's leading companies.


The promise of AI as a social and business growth tool

#artificialintelligence

Since 2005, artificial intelligence (AI) has been on a fast-growing streak as a collection of multiple technologies that enable machines to sense, comprehend, act and learn, either on their own or to augment human activities, and it has been figuring on many released trends to watch research notes and whitepapers from important organizations, such as Gartner, Forrester and IDC. Looking back the past 10-15 years, there has been a lot of work in developing different AI technologies by major market players, such as IBM's Watson system and Google's DeepMind's AlphaGo system. Currently, thanks to Apple's Siri, Microsoft's Cortana, Google's Google Assistant, and Amazon's Alexa, consumers have easy access to a variety of AI-powered virtual assistants to help manage their daily routines and tasks. The potential uses of AI to identify patterns, learn from experience, and find novel solutions to new challenges continue to grow, as increasing investments will bring new AI technology advances to corporations and consumers. Further to that, AI is impacting many different sectors of the global economy and society in a very positive way, as humanitarian organizations are currently using intelligent chatbots to provide psychological support to Syrian refugees, and doctors are using AI to develop personalized treatments for cancer patients.


Top 20 Recent Research Papers on Machine Learning and Deep Learning

#artificialintelligence

Machine learning, especially its subfield of Deep Learning, had many amazing advances in the recent years, and important research papers may lead to breakthroughs in technology that get used by billions of people. The research in this field is developing very quickly and to help our readers monitor the progress we present the list of most important recent scientific papers published since 2014. The criteria we used to select the 20 top papers are by using citation counts from three academic sources: scholar.google.com; Since the number of citations varied among sources and are estimated, we listed the results from academic.microsoft.com For each paper we also give the year it was published, a Highly Influential Citation count (HIC) and Citation Velocity (CV) measures provided by semanticscholar.org.