Question Answering
Improving Recommendation of Tail Tags for Questions in Community Question Answering
Wu, Yu (Beihang University) | Wu, Wei (Microsoft Research) | Li, Zhoujun (Beihang University) | Zhou, Ming (Microsoft Research)
We study tag recommendation for questions in community question answering (CQA). Tags represent the semantic summarization of questions are useful for navigation and expert finding in CQA and can facilitate content consumption such as searching and mining in these web sites. The task is challenging, as both questions and tags are short and a large fraction of tags are tail tags which occur very infrequently. To solve these problems, we propose matching questions and tags not only by themselves, but also by similar questions and similar tags. The idea is then formalized as a model in which we calculate question-tag similarity using a linear combination of similarity with similar questions and tags weighted by tag importance.Question similarity, tag similarity, and tag importance are learned in a supervised random walk framework by fusing multiple features. Our model thus can not only accurately identify question-tag similarity for head tags, but also improve the accuracy of recommendation of tail tags. Experimental results show that the proposed method significantly outperforms state-of-the-art methods on tag recommendation for questions. Particularly, it improves tail tag recommendation accuracy by a large margin.
Community-Based Question Answering via Heterogeneous Social Network Learning
Fang, Hanyin (Zhejiang University) | Wu, Fei (Zhejiang University) | Zhao, Zhou (Zhejiang University) | Duan, Xinyu (Zhejiang University) | Zhuang, Yueting (Zhejiang University) | Ester, Martin (Simon Fraser University)
Community-based question answering (cQA) sites have accumulated vast amount of questions and corresponding crowdsourced answers over time. How to efficiently share the underlying information and knowledge from reliable (usually highly-reputable) answerers has become an increasingly popular research topic. A major challenge in cQA tasks is the accurate matching of high-quality answers w.r.t given questions. Many of traditional approaches likely recommend corresponding answers merely depending on the content similarity between questions and answers, therefore suffer from the sparsity bottleneck of cQA data. In this paper, we propose a novel framework which encodes not only the contents of question-answer(Q-A) but also the social interaction cues in the community to boost the cQA tasks. More specifically, our framework collaboratively utilizes the rich interaction among questions, answers and answerers to learn the relative quality rank of different answers w.r.t a same question. Moreover, the information in heterogeneous social networks is comprehensively employed to enhance the quality of question-answering (QA) matching by our deep random walk learning framework. Extensive experiments on a large-scale dataset from a real world cQA site show that leveraging the heterogeneous social information indeed achieves better performance than other state-of-the-art cQA methods.
Startup junkie advice for both entrepreneurs and enterprises - IBM Watson
Not every startup CEO can say they were able to grow their business to a point where they were acquired. Even fewer can say they did it twice. But that is exactly the case for AlchemyAPI Founder and CEO Elliot Turner. Turner launched his first startup, MimeStar, a software development company focused on network intrusion detection, while a sophomore in high school. Inc. acquired it by the time he was twenty-one. He quickly saw the shift in the market to the need to democratize artificial intelligence (A.I.), and decided to venture out on his own to start AlchemyAPI.
Should IBM Watson issue USPTO first office actions? I think yes...
I would like to propose that Watson could solve one of the biggest challenges facing anyone trying to innovate and product their innovation with a US patent - the USPTO first office action. While everyone is working hard and I know the patent office is overloaded, here how three problems I've seen over my years of working that perhaps Watson could address: 1) Speed - it can take 6-12 months to get a first office action 2) Almost any patent application is nowadays first rejected due to obviousness. But the patents cited to create this argument are often taken out of context. To me, these seem like challenges that Watson would be perfectly designed to addressed. And all the literature to be reviewed is, by definition, in the public domain.
The Computer That Could Be Smarter Than Us [IBM Watson]
This is the direction of the future. Useful AI that can do the research of a thoudand men instantly. It's definitely worth noting that Watson is capable of learning (a point I didn't touch on in this video), so what you see here is the "baby phase" so to speak. I tried to leave out the technical jargon in this video but for those who want to know more, a wiki dump on Watson is below: According to John Rennie, Watson can process 500 gigabytes, the equivalent of a million books, per second. Software Watson uses IBM's DeepQA software and the Apache UIMA (Unstructured Information Management Architecture) framework.
Cleveland Clinic to use IBM Watson for Genomic Research - Decide Software
Cleveland Clinic to use IBM Watson for Genomic Research: Researchers at Cleveland Clinic will use IBM Watson technology in the area of genomic research to help oncologists deliver personalized medicine by uncovering new cancer treatment options for patients. The Lerner Research Institute's Genomic Medicine Institute at Cleveland Clinic plans to evaluate Watson's ability to help oncologists develop more personalized care to patients for a variety of cancers. Clinicians lack the tools and time required to bring DNA-based treatment options to their patients and to do so, they must correlate data from genome sequencing to reams of medical journals, new studies and clinical records. At a time when medical information is doubling every five years, a faster option is needed. This use of Watson aims to find the "needle in the haystack" through identifying patterns in genome sequencing and medical data to unlock insights that will help clinicians bring the promise of genomic medicine to their patients.
Understanding IBM Watson
Watson has it's visualisation tool called WatsonPaths to show how it has derived answers logically. AI needs large amounts of data and Google, Facebook and Amazon are sitting very pretty in this space. IBM will probably be unable to match either of the 3 – but if it becomes an industry expert – it will mint money in the more expensive and much needed business vertical. To be fair, IBM seems to be transparent on this topic. In this case – they're rolling it out for free!!! Well – upto a point IBM Bluemix services helps with development of a rapid prototype solution.
Measuring Machine Intelligence Through Visual Question Answering
Zitnick, C. Lawrence (Facebook AI Research) | Agrawal, Aishwarya (Virginia Institute of Technology) | Antol, Stanislaw (Virginia Institute of Technology) | Mitchell, Margaret (Microsoft Research) | Batra, Dhruv (Virginia Institute of Technology) | Parikh, Devi (Virginia Institute of Technology)
We begin with a case study exploring the recently popular task of image captioning and its limitations as a task for measuring machine intelligence. An alternative and more promising task is Visual Question Answering that tests a machine's ability to reason about language and vision. We describe a dataset unprecedented in size created for the task that contains over 760,000 human generated questions about images. Using around 10 million human generated answers, machines may be easily evaluated.
Measuring Machine Intelligence Through Visual Question Answering
Zitnick, C. Lawrence (Facebook AI Research) | Agrawal, Aishwarya (Virginia Institute of Technology) | Antol, Stanislaw (Virginia Institute of Technology) | Mitchell, Margaret (Microsoft Research) | Batra, Dhruv (Virginia Institute of Technology) | Parikh, Devi (Virginia Institute of Technology)
As machines have become more intelligent, there has been a renewed interest in methods for measuring their intelligence. A common approach is to propose tasks for which a human excels, but one which machines find difficult. However, an ideal task should also be easy to evaluate and not be easily gameable. We begin with a case study exploring the recently popular task of image captioning and its limitations as a task for measuring machine intelligence. An alternative and more promising task is Visual Question Answering that tests a machine’s ability to reason about language and vision. We describe a dataset unprecedented in size created for the task that contains over 760,000 human generated questions about images. Using around 10 million human generated answers, machines may be easily evaluated.