Goto

Collaborating Authors

 Question Answering


Learning to Specialize with Knowledge Distillation for Visual Question Answering

Neural Information Processing Systems

Visual Question Answering (VQA) is a notoriously challenging problem because it involves various heterogeneous tasks defined by questions within a unified framework. Learning specialized models for individual types of tasks is intuitively attracting but surprisingly difficult; it is not straightforward to outperform naive independent ensemble approach. We present a principled algorithm to learn specialized models with knowledge distillation under a multiple choice learning (MCL) framework, where training examples are assigned dynamically to a subset of models for updating network parameters. The assigned and non-assigned models are learned to predict ground-truth answers and imitate their own base models before specialization, respectively. Our approach alleviates the limitation of data deficiency in existing MCL frameworks, and allows each model to learn its own specialized expertise without forgetting general knowledge. The proposed framework is model-agnostic and applicable to any tasks other than VQA, e.g., image classification with a large number of labels but few per-class examples, which is known to be difficult under existing MCL schemes. Our experimental results indeed demonstrate that our method outperforms other baselines for VQA and image classification.


Learning Conditioned Graph Structures for Interpretable Visual Question Answering

Neural Information Processing Systems

Visual Question answering is a challenging problem requiring a combination of concepts from Computer Vision and Natural Language Processing. Most existing approaches use a two streams strategy, computing image and question features that are consequently merged using a variety of techniques. Nonetheless, very few rely on higher level image representations, which can capture semantic and spatial relationships. In this paper, we propose a novel graph-based approach for Visual Question Answering. Our method combines a graph learner module, which learns a question specific graph representation of the input image, with the recent concept of graph convolutions, aiming to learn image representations that capture question specific interactions. We test our approach on the VQA v2 dataset using a simple baseline architecture enhanced by the proposed graph learner module. We obtain promising results with 66.18% accuracy and demonstrate the interpretability of the proposed method.


Dialog-to-Action: Conversational Question Answering Over a Large-Scale Knowledge Base

Neural Information Processing Systems

We present an approach to map utterances in conversation to logical forms, which will be executed on a large-scale knowledge base. To handle enormous ellipsis phenomena in conversation, we introduce dialog memory management to manipulate historical entities, predicates, and logical forms when inferring the logical form of current utterances. Dialog memory management is embodied in a generative model, in which a logical form is interpreted in a top-down manner following a small and flexible grammar. We learn the model from denotations without explicit annotation of logical forms, and evaluate it on a large-scale dataset consisting of 200K dialogs over 12.8M entities. Results verify the benefits of modeling dialog memory, and show that our semantic parsing-based approach outperforms a memory network based encoder-decoder model by a huge margin.


Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering

Neural Information Processing Systems

Accurately answering a question about a given image requires combining observations with general knowledge. While this is effortless for humans, reasoning with general knowledge remains an algorithmic challenge. To advance research in this direction a novel `fact-based' visual question answering (FVQA) task has been introduced recently along with a large set of curated facts which link two entities, i.e., two possible answers, via a relation. Given a question-image pair, deep network techniques have been employed to successively reduce the large set of facts until one of the two entities of the final remaining fact is predicted as the answer. We observe that a successive process which considers one fact at a time to form a local decision is sub-optimal. Instead, we develop an entity graph and use a graph convolutional network to `reason' about the correct answer by jointly considering all entities. We show on the challenging FVQA dataset that this leads to an improvement in accuracy of around 7% compared to the state-of-the-art.


Chain of Reasoning for Visual Question Answering

Neural Information Processing Systems

Reasoning plays an essential role in Visual Question Answering (VQA). Multi-step and dynamic reasoning is often necessary for answering complex questions. For example, a question "What is placed next to the bus on the right of the picture?" talks about a compound object "bus on the right," which is generated by the relation . Furthermore, a new relation including this compound object is then required to infer the answer. However, previous methods support either one-step or static reasoning, without updating relations or generating compound objects. This paper proposes a novel reasoning model for addressing these problems. A chain of reasoning (CoR) is constructed for supporting multi-step and dynamic reasoning on changed relations and objects. In detail, iteratively, the relational reasoning operations form new relations between objects, and the object refining operations generate new compound objects from relations. We achieve new state-of-the-art results on four publicly available datasets. The visualization of the chain of reasoning illustrates the progress that the CoR generates new compound objects that lead to the answer of the question step by step.


Cognitive Bias in Machine Learning โ€“ The Data Lab โ€“ Medium

#artificialintelligence

Companies from a wide range of industries use machine learning data to do everyday business. From consumer marketing and workforce management to healthcare treatment decision solutions and public safety and policing solutions, whether you realize it or not your life is increasingly more affected by the outcomes of machine learning algorithms. Machine learning algorithms make decisions like who gets a bonus, a job interview, whether or not your credit card limit (or interest) is raised, and who gets into a clinical trial. Machine learning algorithms even help make decisions about who gets parole and who languishes in prison. The result is that people's lives and livelihood are affected by the decisions made by machines.


Why Voice Search Will Dominate SEO In 2019 -- And How You Can Capitalize On It

#artificialintelligence

By 2020, 30% of all website sessions will be conducted without a screen. Now, you may be asking yourself, how is that possible? It turns out that voice-only search allows users to browse the web the Internet and consumer information without actually having to scroll through sites on desktops and mobile devices. And this new technology may be the key to successful brands in the future. Voice search essentially allows users to speak into a device as opposed to typing keywords into a search query to generate results.


5 IBM Watson sessions to add to your Think 2019 schedule - Watson

#artificialintelligence

Do you want to learn how you can accelerate your AI strategy or get ahead of the latest AI trends? Or are you more curious to learn what results businesses are achieving by adopting AI? Either way, make sure you attend Think 2019 and experience Watson AI technology first-hand. Here's a sneak peek at five sessions you can't miss: Being able to explain the decisions your AI makes and have trust in them is crucial to accelerating adoption of AI in your business. In these sessions, you'll learn how AI OpenScale provides businesses with confidence in AI decisions and infuses AI throughout its full lifecycle with trust and transparency, explains outcomes, and automatically mitigates bias. However, there are still a variety of hurdles businesses need to overcome to scale and automate their AI.


TED Talks: World's youngest IBM programmer Tanmay talks about artificial intelligence at Sharda University

#artificialintelligence

Exploring the future of artificial intelligence (AI) in our day to day lives, computer whiz kid Tanmay Bakshi said at an event in Greater Noida that instances of fake news, hate speech and harassment on social media can be dealt with the use of AI. Fifteen-year-old Bakshi, the world's youngest IBM Watson programmer, was at Sharda University in Greater Noida on Friday for a TED Talk with students on computer programming and the future of artificial intelligence. He said AI can be monumental in curbing fake news and hate speech. "Fake news is huge and I myself have been a victim of it where one of my TED Talk videos was uploaded on Facebook with the caption that I work for Google and I make billions of dollars a year. I believe social media giants have started using AI to clamp down on fake news and hate speech. For example, Facebook is using machine learning (alternatively known as AI) to understand the content being put up, match it with trusted sources, understand the different point of views which people can have, and when they are absolutely sure that it is fake news then it will be automatically flagged for deletion," Bakshi said.


IBM Watson Suite Aims to Meld AI with HR

#artificialintelligence

IBM has launched a unit designed for human resources to better find talent and recruit using artificial intelligence. The company's HR effort, dubbed IBM Talent & Transformation, includes select Watson AI-based services that can help HR become a growth engine to enable digital transformation. AI can be used to revamp workflow, employee engagement, recruitment and retention while providing a more diverse workforce, the company says. The Watson Talent Suite rolls up behavioral science, AI, and psychology and applies it to HR. Components include Watson Career Coach, a virtual coach that provides advice for career paths, and Watson Candidate Assistant, which looks through the history of job seekers and matches them with openings. These services were developed for IBM's internal HR team and the company claims it drove $107 billion in benefits in 2017 with better employee satisfaction.