gutfreund
Can you teach AI common sense?
All the sessions from Transform 2021 are available on-demand now. Even before they speak their first words, human babies develop mental models about objects and people. This is one of the key capabilities that allows us humans to learn to live socially and cooperate (or compete) with each other. But for artificial intelligence, even the most basic behavioral reasoning tasks remain a challenge. Advanced deep learning models can do complicated tasks such as detect people and objects in images, sometimes even better than humans.
Can AI learn to reason about the world like children?
This article is part of our reviews of AI research papers, a series of posts that explore the latest findings in artificial intelligence. Even before they speak their first words, human babies develop mental models about objects and people. This is one of the key capabilities that allows us humans to learn to live socially and cooperate (or compete) with each other. But for artificial intelligence, even the most basic behavioral reasoning tasks remain a challenge. Advanced deep learning models can do complicated tasks such as detect people and objects in images, sometimes even better than humans.
Artificial intelligence in action
By Meg Murphy A person watching videos that show things opening -- a door, a book, curtains, a blooming flower, a yawning dog -- easily understands the same type of action is depicted in each clip. "Computer models fail miserably to identify these things. How do humans do it so effortlessly?" asks Dan Gutfreund, a principal investigator at the MIT-IBM Watson AI Laboratory and a staff member at IBM Research. "We process information as it happens in space and time. How can we teach computer models to do that?"
Watch the weird videos used to train AI what different actions look like
Teaching computers how to understand actions in videos is tougher than getting them to understand images. "Videos are harder because the problem that we are dealing with is one step higher in terms of complexity if we compare it to object recognition," says Dan Gutfreund, a researcher at a joint IBM-MIT laboratory. "Because objects are objects; a hot dog is a hot dog." Meanwhile, understanding the verb "opening" is tricky, he says, because a dog opening its mouth, or a person opening a door, are going to look different. The dataset is not the first one out there that researchers have created to help machines understand images or videos.