Goto

Collaborating Authors

 Education


A Policy Gradient Method for Task-Agnostic Exploration

arXiv.org Machine Learning

In a reward-free environment, what is a suitable intrinsic objective for an agent to pursue so that it can learn an optimal task-agnostic exploration policy? In this paper, we argue that the entropy of the state distribution induced by limited-horizon trajectories is a sensible target. Especially, we present a novel and practical policy-search algorithm, Maximum Entropy POLicy optimization (MEPOL), to learn a policy that maximizes a non-parametric, $k$-nearest neighbors estimate of the state distribution entropy. In contrast to known methods, MEPOL is completely model-free as it requires neither to estimate the state distribution of any policy nor to model transition dynamics. Then, we empirically show that MEPOL allows learning a maximum-entropy exploration policy in high-dimensional, continuous-control domains, and how this policy facilitates learning a variety of meaningful reward-based tasks downstream.


Is Long Horizon Reinforcement Learning More Difficult Than Short Horizon Reinforcement Learning?

arXiv.org Artificial Intelligence

Learning to plan for long horizons is a central challenge in episodic reinforcement learning problems. A fundamental question is to understand how the difficulty of the problem scales as the horizon increases. Here the natural measure of sample complexity is a normalized one: we are interested in the number of episodes it takes to provably discover a policy whose value is $\varepsilon$ near to that of the optimal value, where the value is measured by the normalized cumulative reward in each episode. In a COLT 2018 open problem, Jiang and Agarwal conjectured that, for tabular, episodic reinforcement learning problems, there exists a sample complexity lower bound which exhibits a polynomial dependence on the horizon -- a conjecture which is consistent with all known sample complexity upper bounds. This work refutes this conjecture, proving that tabular, episodic reinforcement learning is possible with a sample complexity that scales only logarithmically with the planning horizon. In other words, when the values are appropriately normalized (to lie in the unit interval), this results shows that long horizon RL is no more difficult than short horizon RL, at least in a minimax sense. Our analysis introduces two ideas: (i) the construction of an $\varepsilon$-net for optimal policies whose log-covering number scales only logarithmically with the planning horizon, and (ii) the Online Trajectory Synthesis algorithm, which adaptively evaluates all policies in a given policy class using sample complexity that scales with the log-covering number of the given policy class. Both may be of independent interest.


Hands-On Machine Learning with scikit-learn and Python

#artificialintelligence

You keep hearing about machine learning and artificial intelligence and how they revolutionize the world we live in. You want to learn more. You have basic understanding of Python and you are decent in math. You're going to learn hands-on machine learning with scikit-learn, a Python library for machine learning. Since this is a hands-on course, you will be working your way through with Python and Jupyter notebooks.


Council Post: AI-Assisted Learning And Its Impact On Education

#artificialintelligence

Founder and CEO at Fusemachines, an AI Education and AI Talent Solution provider based in NYC. Learning is a multifaceted, multidimensional and dynamic experience made of intricate layers that include reading, writing, listening, watching, thinking, testing and more. These layers weave together to make learning an experience that is personal and relative to every person. There is power in understanding the elements that shape the way we learn. That knowledge, when partnered with artificial intelligence (AI), can enable us to create learning experiences that are supportive to all learners. A learning experience that is adaptive and enhances our natural style of learning with machine intelligence can be thought of as AI-assisted learning (AIAL).


Free MIT Courses on Calculus: The Key to Understanding Deep Learning - KDnuggets

#artificialintelligence

It is difficult, perhaps, to link this to neural networks, but the basic intuition of calculus is achieved. If you are looking for a more full treatment of this branch of mathematics, you will want to seek out some more robust learning tools. Here are 3 courses and a textbook to help out, all from MIT's Open Courseware initiative, which will cover everything you need to know about calculus to understand deep learning -- and far beyond.


Reverb: a framework for experience replay

#artificialintelligence

The use of experience plays a key role in reinforcement learning (RL). How best to use this data is one of the central problems of this field. As RL agents have advanced over recent years, taking on bigger and more complex problems (Atari, Go, StarCraft, Dota), the generated data has grown in both size and complexity. To cope with this complexity many RL systems split the learning problem into two distinct parts: experience producers (actors) and experience consumers (learners) โ€” allowing these different parts to run in parallel. Often a data storage system lies at the intersection between these two components. The question of how to efficiently store and transport the data is itself a challenging engineering problem.


Using Artificial Intelligence as a solution to unemployment

#artificialintelligence

Artificial Intelligence can be a Solution to unending unemployment. In today's society, it has become a difficult task to secure employment especially for those with minimum or no industry experience. Opportunities in the emerging technology can never be exhausted, as a positive tool it can be used to encourage the youths to pursue digital entrepreneurship. Youths can now define, create and manage their own ventures โ€“ be it to maintain and service the technology itself through digital start-ups, online kiosks or even innovation of better industry solutions. According to statistics from the International Labour Organisation, "the majority of youths regularly suffer from under-employment and lack decent working conditions. Of the 38.1 per cent estimated total working poor in sub-Saharan Africa, young people account for 23.5 per cent. Young girls tend to be more disadvantaged than young men in access to work and experience worse working conditions than their male counterpart, and employment in the informal economy or informal employment is the norm."


Learn PyTorch: The best free online courses and tutorials

#artificialintelligence

Deep learning continues to be one of the hottest fields in computing, and while Google's TensorFlow remains the most popular framework in absolute numbers, Facebook's PyTorch has quickly earned a reputation for being easier to grasp and use. PyTorch has taken the world of deep learning research by storm, outstripping TensorFlow as the implementation framework of choice in submitted papers for AI conferences in the past two years. With recent improvements for producing optimized models and deploying them to production, PyTorch is definitely a framework ready for use in industry as well as R&D labs. But how to get started? You'll find plenty of books and paid resources available for learning PyTorch, of course.


Technology In Learning: What Does The Future Hold? - eLearning Industry

#artificialintelligence

There is no doubt about the fact that the era we live in is the era of technology. Technology rules our lives and has made everything from shopping, communicating, ordering food, traveling to entertainment, fitness, comfort, and learning easier and more accessible for us. Let's talk about the role of technology in learning, it being the subject of interest to us. The last decade has seen Learning and Development (L&D) professionals making great strides in imparting better learning to corporate learners through the use of technology. For example, blended learning, microlearning, mobile learning, gamification, AR/VR (Augmented Reality/Virtual Reality), simulations, adaptive learning and even AI (Artificial Intelligence) are now being used to develop skills and knowledge in corporate employees. And, all this is possible because of technological advancements.


Ofqual survey shows mixed attitudes to AI in exam marking

#artificialintelligence

Stakeholders in secondary education have mixed feelings about the potential for using artificial intelligence in marking examinations, according to educational qualifications regulator Ofqual. The issue has emerged from its latest report on perceptions of the major examinations in the sector, with input from a stakeholder survey by research company YouGov. This follows Ofqual's earlier launch of a competition to develop AI solutions for the sector. It says the survey of a YouGov panel shows sentiment in favour of using the technology to check the accuracy of marking in GCSE and AS/A level exams: 42% of stakeholders agreed they would be happy with the idea while 34% disagreed. Support was strongest among young people at 53%. A contrasting picture emerged for using AI in the actual marking process, with just 26% in favour and 51% against the idea.