Goto

Collaborating Authors

 Deep Learning


An End-to-End Baseline for Video Captioning

arXiv.org Artificial Intelligence

Building correspondences across different modalities, such as video and language, has recently become critical in many visual recognition applications, such as video captioning. Inspired by machine translation, recent models tackle this task using an encoder-decoder strategy. The (video) encoder is traditionally a Convolutional Neural Network (CNN), while the decoding (for language generation) is done using a Recurrent Neural Network (RNN). Current state-of-the-art methods, however, train encoder and decoder separately. CNNs are pretrained on object and/or action recognition tasks and used to encode video-level features. The decoder is then optimised on such static features to generate the video's description. This disjoint setup is arguably sub-optimal for input (video) to output (description) mapping. In this work, we propose to optimise both encoder and decoder simultaneously in an end-to-end fashion. In a two-stage training setting, we first initialise our architecture using pre-trained encoders and decoders -- then, the entire network is trained end-to-end in a fine-tuning stage to learn the most relevant features for video caption generation. In our experiments, we use GoogLeNet and Inception-ResNet-v2 as encoders and an original Soft-Attention (SA-) LSTM as a decoder. Analogously to gains observed in other computer vision problems, we show that end-to-end training significantly improves over the traditional, disjoint training process. We evaluate our End-to-End (EtENet) Networks on the Microsoft Research Video Description (MSVD) and the MSR Video to Text (MSR-VTT) benchmark datasets, showing how EtENet achieves state-of-the-art performance across the board.


A Categorisation of Post-hoc Explanations for Predictive Models

arXiv.org Artificial Intelligence

The ubiquity of machine learning based predictive models in modern society naturally leads people to ask how trustworthy those models are? In predictive modeling, it is quite common to induce a trade-off between accuracy and interpretability. For instance, doctors would like to know how effective some treatment will be for a patient or why the model suggested a particular medication for a patient exhibiting those symptoms? We acknowledge that the necessity for interpretability is a consequence of an incomplete formalisation of the problem, or more precisely of multiple meanings adhered to a particular concept. For certain problems, it is not enough to get the answer (what), the model also has to provide an explanation of how it came to that conclusion (why), because a correct prediction, only partially solves the original problem. In this article we extend existing categorisation of techniques to aid model interpretability and test this categorisation.


Learning to Decipher Hate Symbols

arXiv.org Artificial Intelligence

Existing computational models to understand hate speech typically frame the problem as a simple classification task, bypassing the understanding of hate symbols (e.g., 14 words, kigy) and their secret connotations. In this paper, we propose a novel task of deciphering hate symbols. To do this, we leverage the Urban Dictionary and collected a new, symbol-rich Twitter corpus of hate speech. We investigate neural network latent context models for deciphering hate symbols. More specifically, we study Sequence-to-Sequence models and show how they are able to crack the ciphers based on context. Furthermore, we propose a novel Variational Decipher and show how it can generalize better to unseen hate symbols in a more challenging testing setting.


7 Indicators Of The State-Of-Artificial Intelligence (AI), March 2019

#artificialintelligence

Turing Award winners (from left to right) Yoshua Bengio, Yann LeCun, and Geoffrey Hinton at the ReWork Deep Learning Summit, Montreal, October 2017. AI "Sputnik moment" (say it in Chinese*) is at hand China is overtaking the US not just in the sheer volume of AI research papers submitted and published, but also in the production of high-impact papers as measured by the top 50%, top 10%, and top 1% most-cited papers. "By projecting current trends, we see that China is likely to have more top-10% papers by 2020 and more top-1% papers by 2025" (Allen Institute for Artificial Intelligence). Cisco attributes the decline to their increased confidence that "migrating to the cloud will improve protection efforts, while apparently decreasing reliance on less proven technologies such as artificial intelligence" (Cisco). Nearly 90% of IT leaders see their use of AI/ML increasing in the future and 41% look for technology that is powered by AI, a top factor in their purchasing decisions.


DeepMind has made a prototype product that can diagnose eye diseases

#artificialintelligence

It's a device that scans a patient's retina to diagnose potential issues in real time. How it works: After the retina is scanned, the images are then analyzed by DeepMind's algorithms, which return a detailed diagnosis and an "urgency score." It all takes roughly 30 seconds. The prototype system can detect a range of diseases, including diabetic retinopathy, glaucoma, and age-related macular degeneration. Most notably, it can do this as accurately as top eye specialists, DeepMind claims.


Training A Computer To Read Mammograms As Well As A Doctor

#artificialintelligence

"I was really surprised how primitive information technology is in the hospitals," says Regina Barzilay, a professor at the Massachusetts Institute of Technology who is working on improving mammography with artificial intelligence. "I was really surprised how primitive information technology is in the hospitals," says Regina Barzilay, a professor at the Massachusetts Institute of Technology who is working on improving mammography with artificial intelligence. Regina Barzilay teaches one of the most popular computer science classes at the Massachusetts Institute of Technology. And in her research -- at least until five years ago -- she looked at how a computer could use machine learning to read and decipher obscure ancient texts. "This is clearly of no practical use," she says with a laugh.


Boston Dynamics Enters Warehouse Robots Market, Acquires Kinema Systems

IEEE Spectrum Robotics

If you haven't seen the latest Boston Dynamics video, released last week, it shows an upgraded version of the company's Handle robot moving boxes in a warehouse. Handle is a mobile manipulator that integrates both legs and wheels, and the new version features a swinging "tail" that serves as a counterweight and allows the robot to balance and move in a dynamic fashion--just as you'd expect from the company that created such nimble machines as Atlas, Spot, and BigDog. Boston Dynamics, which SoftBank bought from Google in 2017, is showing off Handle toiling in a warehouse for a reason: The company is officially entering the logistics market, with plans to offer robots for material-handling applications. As part of that strategy, it is announcing today the acquisition of Kinema Systems, a startup based in Menlo Park, Calif., that develops vision sensors and deep-learning software to enable industrial robot arms to locate and move boxes. Boston Dynamics founder and CEO Marc Raibert says the two Handle robots seen in the video aren't moving as fast as they could, and one of the factors limiting their performance is their vision systems.


Four Ways Evolutionary AI Can Extend AI's Problem-Solving Capacity - Digitally Cognizant

#artificialintelligence

Deep neural networks (DNN) have produced groundbreaking results in many complex applications of AI, such as natural language processing, facial recognition, sentiment analytics and object recognition. For instance, the accuracy of Google's machine translation system improved 60% using a DNN approach. Finding the right network architecture – that is, the components of the network and how they are instantiated and connected – is essential to this process. If the architecture is chosen based on history and convenience, the network will not reach its full potential. Much of the recent research in DNNs has focused on designing specialized architectures that excel with specific tasks.


Which Deep Learning Framework is Growing Fastest?

#artificialintelligence

In September 2018, I compared all the major deep learning frameworks in terms of demand, usage, and popularity in this article. TensorFlow was the undisputed heavyweight champion of deep learning frameworks. PyTorch was the young rookie with lots of buzz. How has the landscape changed for the leading deep learning frameworks in the past six months? To answer that question, I looked at the number of job listings on Indeed, Monster, LinkedIn, and SimplyHired.


'Godfathers of AI' Receive Turing Award, the Nobel Prize of Computing - AI Trends

#artificialintelligence

The 2018 Turing Award, known as the "Nobel Prize of computing," has been given to a trio of researchers who laid the foundations for the current boom in artificial intelligence. Yoshua Bengio, Geoffrey Hinton, and Yann LeCun -- sometimes called the'godfathers of AI' -- have been recognized with the $1 million annual prize for their work developing the AI subfield of deep learning. The techniques the trio developed in the 1990s and 2000s enabled huge breakthroughs in tasks like computer vision and speech recognition. Their work underpins the current proliferation of AI technologies, from self-driving cars to automated medical diagnoses. In fact, you probably interacted with the descendants of Bengio, Hinton, and LeCun's algorithms today -- whether that was the facial recognition system that unlocked your phone, or the AI language model that suggested what to write in your last email.