Genre
A Feasibility Study of an Approach to Extend Research Footprints
Osuna, Francisco (University of Texas at El Paso) | Gurijala, Bhanukiran (University of Texas at El Paso) | Esparza, Patricia (University of Texas at El Paso) | Akbar, Monika (University of Texas at El Paso) | Gates, Ann (University of Texas at El Paso)
Funding agencies and the National Academies of Science, Engineering, and Medicine have been promoting the importance of interdisciplinary research (IDR). Supporting team-based IDR requires the ability to discover the expertise needed to solve complex problems. Many universities have adopted expertise systems, which includes the presentation of keywords or concepts to identify experts. The efforts at University of Texas at El Paso (UTEP) have focused on building โcommunities of practiceโ that support diverse faculty who have an affinity for a particular topic and facilitate the ability to identify researchers with diverse expertise, knowledge, and skills who can contribute to new initiatives on campus. Our premise is that the university can facilitate the identification of potential contributors to communities of practice by correlating their associated ontologies to the concepts associated with researchersโ publications and proposal submissions. This paper presents the results of a preliminary study to examine the feasibility of the approach.
Identifying and Tracking Switching, Non-Stationary Opponents: A Bayesian Approach
Hernandez-Leal, Pablo (Instituto Nacional de Astrofisica, Optica y Electronica (INAOE)) | Taylor, Matthew E. (Washington State University) | Rosman, Benjamin (University of the Witwatersrand) | Sucar, L. Enrique (Instituto Nacional de Astrofisica, Optica y Electronica (INAOE)) | Cote, Enrique Munoz de (Instituto Nacional de Astrofisica, Optica y Electronica (INAOE))
In many situations, agents are required to use a set of strategies (behaviors) and switch among them during the course of an interaction. This work focuses on the problem of recognizing the strategy used by an agent within a small number of interactions. We propose using a Bayesian framework to address this problem. Bayesian policy reuse (BPR) has been empirically shown to be efficient at correctly detecting the best policy to use from a library in sequential decision tasks. In this paper we extend BPR to adversarial settings, in particular, to opponents that switch from one stationary strategy to another. Our proposed extension enables learning new models in an online fashion when the learning agent detects that the current policies are not performing optimally. Experiments presented in repeated games show that our approach is capable of efficiently detecting opponent strategies and reacting quickly to behavior switches, thereby yielding better performance than state-of-the-art approaches in terms of average rewards.
Towards Learning From Stories: An Approach to Interactive Machine Learning
Harrison, Brent (Georgia Institute of Technology) | Riedl, Mark O. (Georgia Institute of Technology)
In this work, we introduce a technique that uses stories totrain virtual agents to exhibit believable behavior. This technique uses a compact representation of a story to define the space of acceptable behaviors and then uses this space to assign rewards to certain world states. We show the effectiveness of our technique with a case study in a modified gridworld environment called Pharmacy World. The results show that a reinforcement learning agent using Q-learning was able to learn a policy that results in believable behavior.
Modeling Topic-Level Academic Influence in Scientific Literatures
Shen, Jiaming (Shanghai Jiao Tong University) | Song, Zhenyu (Shanghai Jiao Tong University) | Li, Shitao (Shanghai Jiao Tong University) | Tan, Zhaowei (Shanghai Jiao Tong University) | Mao, Yuning (Shanghai Jiao Tong University) | Fu, Luoyi (Shanghai Jiao Tong University) | Song, Li (Shanghai Jiao Tong University) | Wang, Xinbing (Shanghai Jiao Tong University)
Scientific articles are not born equal. Some generate an entire discipline while others make relatively fewer contributions. When reviewing scientific literatures, it would be useful to identify those important articles and understand how they influence others. In this paper, we introduce J-Index, a quantitative metric modeling topic-level academic influence. J-Index is calculated based on the novelty of each article as well as its contributions to the articles where it is cited. We devise a generative model named Reference Topic Model (RefTM) which jointly utilizes the textual content and citation information in scientific literatures. We show how to learn RefTM to discover both the novelty of each paper and the strength of each citation. Experiments on a collection of more than 420,000 research papers demonstrate that RefTM outperforms the state-of-the-art approaches in terms of topic coherence as well as prediction performance, and validate J-Index's effectiveness of capturing topic-level academic influence in scientific literatures.
From a Scholarly Big Dataset to a Test Collection for Bibliographic Citation Recommendation
Roy, Dwaipayan (Indian Statistical Institute) | Ray, Kunal (Microsoft IDC Bangalore) | Mitra, Mandar (Indian Statistical Institute)
The problem of designing recommender systems for scholarly article citations has been actively researched with more than 200 publications appearing in the last two decades. In spite of this, no definitive results are available about what approaches work best. Arguably the most important reason for this lack of consensus is the dearth of standardised test collections and evaluation protocols, such as those provided by TREC-like forums. CiteSeerX, a "scholarly big dataset" has recently become available. However, this collection provides only the raw material that is yet to be moulded into Cranfield style test collections. In this paper, we discuss the limitations of test collections used in earlier work, and describe how we used CiteSeerX to design a test collection with a well-defined evaluation protocol. The collection consists of over 600,000 research papers and over 2,500 queries. We report some preliminary experimental results using this collection, which are indicative of the performance of elementary content-based techniques. These experiments also made us aware of some shortcomings of CiteSeerX itself.
Analyzing NIH Funding Patterns over Time with Statistical Text Analysis
Park, Jihyun (University of California, Irvine) | Blume-Kohout, Margaret (New Mexico Consortium) | Krestel, Ralf (Hasso Plattner Institut) | Nalisnick, Eric (University of California, Irvine) | Smyth, Padhraic (University of California, Irvine)
In the past few years various government funding organizations such as the U.S. National Institutes of Health and the U.S.\ National Science Foundation have provided access to large publicly-available online databases documenting the grants that they have funded over the past few decades. These databases provide an excellent opportunity for the application of statistical text analysis techniques to infer useful quantitative information about how funding patterns have changed over time. In this paper we analyze data from the National Cancer Institute (part of National Institutes of Health) and show how text classification techniques provide a useful starting point for analyzing how funding for cancer research has evolved over the past 20 years in the United States.
Encoding Lineage in Scholarly Articles
Naim, Sheikh Motahar (University of Texas at El Paso) | Kader, Md Abdul (University of Texas at El Paso) | Boedihardjo, Arnold P. (US Army Corps of Engineers) | Hossain, M. Shahriar (University of Texas at El Paso)
The development of new scientific concepts today is an outcome of the accumulated knowledge built over time. Every scientific domain requires understanding of the trends of the dependencies between its subdomains. Analyses of trends to capture such dependencies using conventional document modeling techniques is a challenging task due to two reasons: (1) conventional vector-space modeling based representation of documents does not realize the history of the content, and (2) neither feature-level nor document-level causality is provided with any digital library metadata or citation network. In this paper, we propose an intuitive temporal representation of a scientific article that encodes inherent historic characteristics of the content. This intuitive representation of each document is then leveraged to discover causal relationships between scientific articles. In addition, we provide a mechanism to explore the lineage of each document in terms of other previously published documents, which illustrates how the theme of the document under analysis evolved over time. Empirical studies reported in the paper show that the proposed technique identifies meaningful causal relationships and discovers meaningful lineage in the scientific literature that could not be discovered through the citation network of the articles.
Automatic Construction of Evaluation Sets and Evaluation of Document Similarity Models in Large Scholarly Retrieval Systems
Krstovski, Kriste (Harvard-Smithsonian Center for Astrophysics) | Smith, David A. (Northeastern University) | Kurtz, Michael J. (Harvard-Smithsonian Center for Astrophysics)
Retrieval systems for scholarly literature offer the ability for the scientific community to search, explore and download scholarly articles across various scientific disciplines. Mostly used by the experts in the particular field, these systems contain user community logs including information on user specific downloaded articles. In this paper we present a novel approach for automatically evaluating document similarity models in large collections of scholarly publications. Unlike typical evaluation settings that use test collections consisting of query documents and human annotated relevance judgments, we use download logs to automatically generate pseudo-relevant set of similar document pairs. More specifically we show that consecutively downloaded document pairs, extracted from a scholarly information retrieval (IR) system, could be utilized as a test collection for evaluating document similarity models. Another novel aspect of our approach lies in the method that we employ for evaluating the performance of the model by comparing the distribution of consecutively downloaded document pairs and random document pairs in log space. Across two families of similarity models, that represent documents in the term vector and topic spaces, we show that our evaluation approach achieves very high correlation with traditional performance metrics such as Mean Average Precision (MAP), while being more efficient to compute.
Creating a Mars Target Encyclopedia by Extracting Information from the Planetary Science Literature
Wagstaff, Kiri L. (Jet Propulsion Laboratory) | Riloff, Ellen (University of Utah) | Lanza, Nina L. (Los Alamos National Laboratory) | Mattmann, Chris A. (Jet Propulsion Laboratory) | Ramirez, Paul M. (Jet Propulsion Laboratory)
Staying up to date with the latest discoveries is a challenge in any scientific field. In planetary science, new observation targets on the surface of Mars are identified and named every day, and new publications announcing new discoveries and conclusions provide frequent updates about these targets. We are constructing a system that uses information extraction and retrieval methods to mine the steadily growing body of planetary science publications about Mars surface targets and automatically construct a concise summary of what is known about each target. The Mars Target Encyclopedia will provide a central, continually updated resource for use by planetary scientists and the interested public. We describe our use of Tika, Sundance, and AutoSlog to extract and summarize information, some of the challenges associated with this domain, and our plans for maturing the system.
EmoGram: An Open-Source Time Sequence-Based Emotion Tracker and Its Innovative Applications
Joshi, Aditya (Monash Research Academy) | Tripathi, Vaibhav (Indian Institute of Technology Bombay) | Soni, Ravindra (Indian Institute of Technology Bombay) | Bhattacharyya, Pushpak (Indian Institute of Technology Bombay) | Carman, Mark James (Monash University)
In this paper, we present an open-source emotion tracker and its innovative applications. Our tracker, EmoGram, tracks emotion changes for a sequence of textual units. It is versatile in terms of the textual unit (tweets, sentences in discourse, etc.) and also what constitutes the time sequence (timestamps of tweets, discourse nature of text, etc.). We demonstrate the utility of our system through our applications: a sequence of commentaries in cricket matches, a sequence of dialogues in a play, and a sequence of tweets related to the Maggi controversy in India in 2015. That one system can be used for these applications is the merit of EmoGram.