Goto

Collaborating Authors

 Genre


Unsupervised Learning for Physical Interaction through Video Prediction

arXiv.org Artificial Intelligence

A core challenge for an agent learning to interact with the world is to predict how its actions affect objects in its environment. Many existing methods for learning the dynamics of physical interactions require labeled object information. However, to scale real-world interaction learning to a variety of scenes and objects, acquiring labeled data becomes increasingly impractical. To learn about physical object motion without labels, we develop an action-conditioned video prediction model that explicitly models pixel motion, by predicting a distribution over pixel motion from previous frames. Because our model explicitly predicts motion, it is partially invariant to object appearance, enabling it to generalize to previously unseen objects. To explore video prediction for real-world interactive agents, we also introduce a dataset of 59,000 robot interactions involving pushing motions, including a test set with novel objects. In this dataset, accurate prediction of videos conditioned on the robot's future actions amounts to learning a "visual imagination" of different futures based on different courses of action. Our experiments show that our proposed method produces more accurate video predictions both quantitatively and qualitatively, when compared to prior methods.


SR-Clustering: Semantic Regularized Clustering for Egocentric Photo Streams Segmentation

arXiv.org Artificial Intelligence

While wearable cameras are becoming increasingly popular, locating relevant information in large unstructured collections of egocentric images is still a tedious and time consuming process. This paper addresses the problem of organizing egocentric photo streams acquired by a wearable camera into semantically meaningful segments, hence making an important step towards the goal of automatically annotating these photos for browsing and retrieval. In the proposed method, first, contextual and semantic information is extracted for each image by employing a Convolutional Neural Networks approach. Later, a vocabulary of concepts is defined in a semantic space by relying on linguistic information. Finally, by exploiting the temporal coherence of concepts in photo streams, images which share contextual and semantic attributes are grouped together. The resulting temporal segmentation is particularly suited for further analysis, ranging from event recognition to semantic indexing and summarization. Experimental results over egocentric set of nearly 31,000 images, show the prominence of the proposed approach over state-of-the-art methods. Keywords: temporal segmentation, egocentric vision, photo streams clustering 1. Introduction Among the advances in wearable technology during the last few years, wearable cameras specifically have gained more popularity [5].


X-CNN: Cross-modal Convolutional Neural Networks for Sparse Datasets

arXiv.org Artificial Intelligence

In this paper we propose cross-modal convolutional neural networks (X-CNNs), a novel biologically inspired type of CNN architectures, treating gradient descent-specialised CNNs as individual units of processing in a larger-scale network topology, while allowing for unconstrained information flow and/or weight sharing between analogous hidden layers of the network---thus generalising the already well-established concept of neural network ensembles (where information typically may flow only between the output layers of the individual networks). The constituent networks are individually designed to learn the output function on their own subset of the input data, after which cross-connections between them are introduced after each pooling operation to periodically allow for information exchange between them. This injection of knowledge into a model (by prior partition of the input data through domain knowledge or unsupervised methods) is expected to yield greatest returns in sparse data environments, which are typically less suitable for training CNNs. For evaluation purposes, we have compared a standard four-layer CNN as well as a sophisticated FitNet4 architecture against their cross-modal variants on the CIFAR-10 and CIFAR-100 datasets with differing percentages of the training data being removed, and find that at lower levels of data availability, the X-CNNs significantly outperform their baselines (typically providing a 2--6% benefit, depending on the dataset size and whether data augmentation is used), while still maintaining an edge on all of the full dataset tests.



New Ideas for Brain Modelling

arXiv.org Artificial Intelligence

This paper describes some biologically-inspired processes that could be used to build the sort of networks that we associate with the human brain. New to this paper, a 'refined' neuron will be proposed. This is a group of neurons that by joining together can produce a more analogue system, but with the same level of control and reliability that a binary neuron would have. With this new structure, it will be possible to think of an essentially binary system in terms of a more variable set of values. The paper also shows how recent research associated with the new model, can be combined with established theories, to produce a more complete picture. The propositions are largely in line with conventional thinking, but possibly with one or two more radical suggestions. An earlier cognitive model can be filled in with more specific details, based on the new research results, where the components appear to fit together almost seamlessly. The intention of the research has been to describe plausible 'mechanical' processes that can produce the appropriate brain structures and mechanisms, but that could be used without the magical 'intelligence' part that is still not fully understood. There are also some important updates from an earlier version of this paper.


Introduction to Machine Learning & Face Detection in Python

@machinelearnbot

This course is about the fundamental concepts of machine learning, focusing on neural networks, SVM and decision trees. These topics are getting very hot nowadays because these learning algorithms can be used in several fields from software engineering to investment banking. Learning algorithms can recognize patterns which can help detect cancer for example or we may construct algorithms that can have a very very good guess about stock prices movement in the market. In each section we will talk about the theoretical background for all of these algorithms then we are going to implement these problems together. The first chapter is about regression: very easy yet very powerful and widely used machine learning technique.


How to Start Learning Deep Learning

@machinelearnbot

This post was written by Ofir Press. Ofir is a graduate student at Tel-Aviv University's Deep Learning Lab. His main focus is on using deep learning for natural language processing. "Due to the recent achievements of artificial neural networks across many different tasks (such as face recognition, object detection and Go), deep learning has become extremely popular. This post aims to be a starting point for those interested in learning more about it. If you already have a basic understanding of linear algebra, calculus, probability and programming: I recommend starting with Stanford's CS231n. The course notes are comprehensive and well-written. The slides for each lesson are also available, and even though the accompanying videos were removed from the official site, re-uploads are quite easy to find online. If you don't have the relevant math background: There is an incredible amount of free material online that can be used to learn the required math knowledge. Gilbert Strang's course on linear algebra is a great introduction to the field. For the other subjects, edX has courses from MIT on both calculus and probability. If you are interested in learning more about machine learning: Andrew Ng's Coursera class is a popular choice as a first class in machine learning. There are other great options available such as Yaser Abu-Mostafa's machine learning course which focuses much more on theory than the Coursera class but it is still relevant for beginners. Knowledge in machine learning isn't really a prerequisite to learning deep learning, but it does help. In addition, learning classical machine learning and not only deep learning is important because it provides a theoretical background and because deep learning isn't always the correct solution. Geoffrey Hinton's Coursera class "Neural Networks for Machine Learn... covers a lot of different topics, and so does Hugo Larochelle's "Neural Networks Class".


Humanoid Robot Kengoro 'Sweats' To Cool Down, Power Through Push-Ups

#artificialintelligence

Robots are hailed for their intelligence and work efficiency, but excessive heating from prolonged hours of work often affects their performance. To address the heating problem faced by humanoid robots, Japanese researchers have devised an out-of-the-box solution. Using the analogy of sweating that happens in the human body as a result of continuous activity that cools the heated muscles, researchers at the University of Tokyo's JSK Lab presented a novel method at the IEEE/RSJ International Conference on Intelligent Robots and Systems held in South Korea. Their cooling solution addresses the heating problem of a musculoskeletal humanoid robot called Kengoro, which stands at 1.7 meters (5.6 feet) tall and weighs 56 kilograms (123.5 pounds). The Japanese researchers' cooling solution involves tinkering to make the robot "sweat" water straight out of its frame.


Machine Learning A-Z : Hands-On Python & R In Data Science

#artificialintelligence

My name is Kirill Eremenko and I am super-psyched that you are reading this! I teach courses in two distinct Business areas on Udemy: Data Science and Forex Trading. I want you to be confident that I can deliver the best training there is, so below is some of my background in both these fields. Professionally, I am a Data Science management consultant with over five years of experience in finance, retail, transport and other industries. I was trained by the best analytics mentors at Deloitte Australia and today I leverage Big Data to drive business strategy, revamp customer experience and revolutionize existing operational processes.


Bayesian Statistics: MCMC – EFavDB

#artificialintelligence

We review the Metropolis algorithm -- a simple Markov Chain Monte Carlo (MCMC) sampling method -- and its application to estimating posteriors in Bayesian statistics. A simple python example is provided. Follow @efavdb Follow us on twitter for new submission alerts! One of the central aims of statistics is to identify good methods for fitting models to data. Notice that if we could solve for this function, we would be able to identify which parameter values are most likely -- those that are good candidates for a fit.