Question Answering
IBM's Watson Takes On Risk And Regulation In Finance
IBM is deploying sophisticated technology under its Watson Financial Services brand to improve risk management across financial firms. The initiative has several components, said Michael Curry, vice president of engineering for Watson Financial Services at IBM. "One piece is more focused on financial risk -- market, credit and liquidity risk that are associated with portfolios. The Armanta acquisition we announced a few months ago sits in portfolio management but we are applying it in other places too. "The second piece would be more operational risk, and there you have two components. One is some of the areas of financial crimes and fraud where we have our Financial Crimes Insight Engine, a platform for finding patterns of fraud and market abuse.
Quantifying the Impact of Cognitive Biases in Question-Answering Systems
Burghardt, Keith (University of California, Davis) | Hogg, Tad (Institute for Molecular Manufacturing) | Lerman, Kristina (University of Southern California Information Sciences Institute)
Crowdsourcing can identify high-quality solutions to problems; however, individual decisions are constrained by cognitive biases. We investigate some of these biases in an experimental model of a question-answering system. We observe a strong position bias in favor of answers appearing earlier in a list of choices. This effect is enhanced by three cognitive factors: the attention an answer receives, its perceived popularity, and cognitive load, measured by the number of choices a user has to process. While separately weak, these effects synergistically amplify position bias and decouple user choices of best answers from their intrinsic quality. We end our paper by discussing the novel ways we can apply these findings to substantially improve how high-quality answers are found in question-answering systems.
Detecting Misflagged Duplicate Questions in Community Question-Answering Archives
Hoogeveen, Doris (The University of Melbourne, Data61) | Bennett, Andrew (The University of Melbourne) | Li, Yitong (The University of Melbourne) | Verspoor, Karin M. (The University of Melbourne) | Baldwin, Timothy (The University of Melbourne)
In this paper we introduce the task of misflagged duplicate question detection for question pairs in community question-answer (cQA) archives and compare it to the more standard task of detecting valid duplicate questions. A misflagged duplicate is a question that has been erroneously hand-flagged by the community as a duplicate of an archived one, where the two questions are not actually the same. We find that form is flagged duplicate detection, meta data features that capture user authority, question quality, and relational data between questions, outperform pure text-based methods, while for regular duplicate detection a combination of meta data features and semantic features gives the best results. We show that misflagged duplicate questions are even more challenging to model than regular duplicate question detection, but that good results can still be obtained.
LearningQ: A Large-Scale Dataset for Educational Question Generation
Chen, Guanliang (Delft University of Technology) | Yang, Jie (University of Fribourg) | Hauff, Claudia (Delft University of Technology) | Houben, Geert-Jan (Delft University of Technology)
We present LearningQ, a challenging educational question generation dataset containing over 230K document-question pairs. It includes 7K instructor-designed questions assessing knowledge concepts being taught and 223K learner-generated questions seeking in-depth understanding of the taught concepts. We show that, compared to existing datasets that can be used to generate educational questions, LearningQ (i) covers a wide range of educational topics and (ii) contains long and cognitively demanding documents for which question generation requires reasoning over the relationships between sentences and paragraphs. As a result, a significant percentage of LearningQ questions (~30%) require higher-order cognitive skills to solve (such as applying, analyzing), in contrast to existing question-generation datasets that are designed mostly for the lowest cognitive skill level (i.e. remembering). To understand the effectiveness of existing question generation methods in producing educational questions, we evaluate both rule-based and deep neural network based methods on LearningQ. Extensive experiments show that state-of-the-art methods which perform well on existing datasets cannot generate useful educational questions. This implies that LearningQ is a challenging test bed for the generation of high-quality educational questions and worth further investigation. We open-source the dataset and our codes at https://dataverse.mpi-sws.org/dataverse/icwsm18.
QDEE: Question Difficulty and Expertise Estimation in Community Question Answering Sites
Sun, Jiankai (The Ohio State University) | Moosavi, Sobhan (The Ohio State University) | Ramnath, Rajiv (The Ohio State University) | Parthasarathy, Srinivasan (The Ohio State University)
In this paper, we present a framework for Question Difficulty and Expertise Estimation (QDEE) in Community Question Answering sites (CQAs) such as Yahoo! Answers and Stack Overflow, which tackles a fundamental challenge in crowdsourcing: how to appropriately route and assign questions to users with the suitable expertise. This problem domain has been the subject of much research and includes both language-agnostic as well as language conscious solutions. We bring to bear a key language-agnostic insight: that users gain expertise and therefore tend to ask as well as answer more difficult questions over time. We use this insight within the popular competition (directed) graph model to estimate question difficulty and user expertise by identifying key hierarchical structure within said model. An important and novel contribution here is the application of ``social agony'' to this problem domain. Difficulty levels of newly posted questions (the cold-start problem) are estimated by using our QDEE framework and additional textual features. We also propose a model to route newly posted questions to appropriate users based on the difficulty level of the question and the expertise of the user. Extensive experiments on real world CQAs such as Yahoo! Answers and Stack Overflow data demonstrate the improved efficacy of our approach over contemporary state-of-the-art models.
The Natural Language Decathlon: Multitask Learning as Question Answering
McCann, Bryan, Keskar, Nitish Shirish, Xiong, Caiming, Socher, Richard
Deep learning has improved performance on many natural language processing (NLP) tasks individually. However, general NLP models cannot emerge within a paradigm that focuses on the particularities of a single metric, dataset, and task. We introduce the Natural Language Decathlon (decaNLP), a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation extraction, goal-oriented dialogue, semantic parsing, and commonsense pronoun resolution. We cast all tasks as question answering over a context. Furthermore, we present a new Multitask Question Answering Network (MQAN) jointly learns all tasks in decaNLP without any task-specific modules or parameters in the multitask setting. MQAN shows improvements in transfer learning for machine translation and named entity recognition, domain adaptation for sentiment analysis and natural language inference, and zero-shot capabilities for text classification. We demonstrate that the MQAN's multi-pointer-generator decoder is key to this success and performance further improves with an anti-curriculum training strategy. Though designed for decaNLP, MQAN also achieves state of the art results on the WikiSQL semantic parsing task in the single-task setting. We also release code for procuring and processing data, training and evaluating models, and reproducing all experiments for decaNLP.
Fox Sports Teams With IBM Watson to Use Artificial Intelligence for FIFA World Cup
Fox Sports has tapped into the potential of artificial intelligence and machine learning to deliver the innovative FIFA World Cup Highlight Machine, available through its Fox Sports App and FoxSports.com. A collaboration with IBM, the Highlight Machine is enabled by the IBM Watson computer technology. It analyzes video from the FIFA World Cup archive, as well as 2018 footage, and extracts data, allowing users to search for goals, red cards, players by name and the like. It is about creating a "compelling user experience around highlights, not just in 2018 but also previously, at leat 50 years," says David Mowrey, head of product and development at IBM Watson Media. More typically, gathering this sort of data would be a task done manually by employees, but considering the scope of the World Cup, that would be impractical, and arguably impossible, Mowrey explains, because of the enormous volume of video that is involved.
Is "IBM Watson Health Imaging" the Future of Healthcare? - Nanalyze
A track record of prior competency that is above and beyond the norm is what hiring managers look for when they recruit "top talent", as recruiters like to say. Usually "top talents" can command a premium in the market place because everybody wants to employ them. We can equate these "top talents" to top quality stocks. You often hear dividend investors talk about how top dividend growth stocks are "always too expensive". That comment usually refers to the yield for the stock being lower than average, which in the case of a quality stock just represents a greater anticipation of future growth.
Being Negative but Constructively: Lessons Learnt from Creating Better Visual Question Answering Datasets
Chao, Wei-Lun, Hu, Hexiang, Sha, Fei
Visual question answering (Visual QA) has attracted a lot of attention lately, seen essentially as a form of (visual) Turing test that artificial intelligence should strive to achieve. In this paper, we study a crucial component of this task: how can we design good datasets for the task? We focus on the design of multiple-choice based datasets where the learner has to select the right answer from a set of candidate ones including the target (\ie the correct one) and the decoys (\ie the incorrect ones). Through careful analysis of the results attained by state-of-the-art learning models and human annotators on existing datasets, we show that the design of the decoy answers has a significant impact on how and what the learning models learn from the datasets. In particular, the resulting learner can ignore the visual information, the question, or both while still doing well on the task. Inspired by this, we propose automatic procedures to remedy such design deficiencies. We apply the procedures to re-construct decoy answers for two popular Visual QA datasets as well as to create a new Visual QA dataset from the Visual Genome project, resulting in the largest dataset for this task. Extensive empirical studies show that the design deficiencies have been alleviated in the remedied datasets and the performance on them is likely a more faithful indicator of the difference among learning models. The datasets are released and publicly available via http://www.teds.usc.edu/website_vqa/.
Fox Sports' World Cup highlight machine is powered by IBM's Watson
And for soccer (er, football) fans in the US, Fox Sports will be the TV network responsible for bringing them all 64 games from Russia, at least if they want to watch them in English. But, beyond its broadcast offerings, Fox Sports wants to keep people engaged in the competition in different ways. Aside from its partnership with Twitter, which comes in the form of a show that'll stream live from Russia, Fox Sports has teamed up with IBM to build the ultimate World Cup highlight machine. Powered by Watson artificial intelligence, this video hub lets you create on-demand clips from every FIFA World Cup tournament dating back to 1958. Fox Sports says there are 300 archived matches that Watson is capable of analyzing, which you can filter out by World Cup year, team, player, game, play type or any combination of these.