Genre
Chinese Relation Extraction by Multiple Instance Learning
Chen, Yu-Ju (National Taiwan University) | Hsu, Jane Yung-jen (National Taiwan University)
Relation extraction, which learns semantic relations of concept pairs from text, is an approach for mining commonsense knowledge. This paper investigates an approach for relation extraction, which helps expand a commonsense knowledge base with little labor work. We proposed a framework that learns new pairs from Chinese corpora by adopting concept pairs in Chinese commonsense knowledge base as seeds. Multiple instance learning is utilized as the learning algorithm for predicting relation for unseen pairs. The performance of our system could be improved by learning multiple iterations. The results in each iteration are manually evaluated and processed to next iteration as seeds. Our experiments extracted new pairs for relations “AtLocation”, “CapableOf”, and “HasProperty”. This study showed that new pairs could be extracted from text without huge humans work.
Effect of Part-of-Speech and Lemmatization Filtering in Email Classification for Automatic Reply
Bonatti, Rogerio (Universidade de Sao Paulo) | Paula, Arthur G. de (Universidade de Sao Paulo) | Lamarca, Victor S. (Universidade de Sao Paulo) | Cozman, Fabio G. (Universidade de Sao Paulo)
We study the automatic reply of email business messages in Brazilian Portuguese. We present a novel corpus containing messages from a real application, and baseline categorization experiments using Naive Bayes and Support Vector Machines. We then discuss the effect of lemmatization and the role of part-of-speech tagging filtering on precision and recall. Support Vector Machines classification coupled with non-lemmatized selection of verbs and nouns, adjectives and adverbs was the best approach, with 87.3% maximum accuracy. Straightforward lemmatization in Portuguese led to the lowest classification results in the group, with 85.3% and 81.7% precision in SVM and Naive Bayes respectively. Thus, while lemmatization reduced precision and recall, part-of-speech filtering improved overall results.
An Analysis of Trimming in Digital Social Networks
Murimi, Renita Margaret (Oklahoma Baptist University)
The study of network sizes in digital social networks is a research question of significant interest. Here, we explore the phenomenon of trimming, which is the decrease in the size of one’s network, and analyze if the rules of social exchange theory – namely, status consistency and reciprocity- can affect trimming. To this end, we use a Hidden Markov Model to investigate the relationship between the frequency of interaction and one’s network size, in which we are able to control for the current size of one’s digital social network. We find that there are significant patterns in sharing tendencies in digital social networks. One is that users who do not share enough are the group that is most likely to be trimmed from a network. Another is that users prefer to have moderate sized networks, i.e. networks with 500 – 1000 friends and prefer friends with moderate sharing tendencies (sharing approximately once a week). We also find that one’s sharing preferences over time tend to align with moderate sharing.
User Participation and Honesty in Online Rating Systems: What a Social Network Can Do
Davoust, Alan (Carleton University) | Esfandiari, Babak (Carleton University)
An important problem with online communities in general, and online rating systems in particular, is uncooperative behavior: lack of user participation, dishonest contributions. This may be due to an incentive structure akin to a Prisoners' Dilemma (PD). We show that introducing an explicit social network to PD games fosters cooperative behavior, and use this insight to design a new aggregation technique for online rating systems. Using a dataset of ratings from Yelp, we show that our aggregation technique outperforms Yelp's proprietary filter, as well as baseline techniques from recommender systems.
Simultaneous Influencing and Mapping for Health Interventions
Marcolino, Leandro Soriano (University of Southern California) | Lakshminarayanan, Aravind (Indian Institute of Technology, Madras) | Yadav, Amulya (University of Southern California) | Tambe, Milind (University of Southern California)
Influence Maximization is an active topic, but it was always assumed full knowledge of the social network graph. However, the graph may actually be unknown beforehand. For example, when selecting a subset of a homeless population to attend interventions concerning health, we deal with a network that is not fully known. Hence, we introduce the novel problem of simultaneously influencing and mapping (i.e., learning) the graph. We study a class of algorithms, where we show that: (i) traditional algorithms may have arbitrarily low performance; (ii) we can effectively influence and map when the independence of objectives hypothesis holds; (iii) when it does not hold, the upper bound for the influence loss converges to 0. We run extensive experiments over four real-life social networks, where we study two alternative models, and obtain significantly better results in both than traditional approaches.
Protecting Wildlife under Imperfect Observation
Nguyen, Thanh Hong (University of Southern California) | Sinha, Arunesh (University of Southern California) | Gholami, Shahrzad (University of Southern California) | Plumptre, Andrew ( Wildlife Conservation Society ) | Joppa, Lucas ( Microsoft Research ) | Tambe, Milind (University of Southern California) | Driciru, Margaret ( Uganda Wildlife Authority ) | Wanyama, Fred ( Uganda Wildlife Authority ) | Rwetsiba, Aggrey ( Uganda Wildlife Authority ) | Critchlow, Rob ( The University of York ) | Beale, Colin ( The University of York )
Wildlife poaching presents a serious extinction threat to many animal species. In order to save wildlife in designated wildlife parks, park rangers conduct patrols over the park area to combat such illegal activities. An important aspect of the patrolling activity of the rangers is to anticipate where the poachers are likely to catch animals and then respond accordingly. Previous work has applied defender-attacker Stackelberg Security Games (SSGs) to solve the problem of wildlife protection, wherein attacker behavioral models are used to predict the behaviors of the poachers. However, these behavioral models have several limitations which limit their accuracy in predicting poachers' behavior. First, existing models fail to account for the rangers' imperfect observations w.r.t poaching activities (due to the limited capability of rangers to patrol thoroughly over a vast geographical area). Second, these models are built upon discrete choice models that assume a single agent choosing targets, while it is infeasible to obtain information about every single attacker in wildlife protection. Third, these models do not consider the effect of past poachers' actions on the current poachers' activities, one of the key factors affecting the poachers' behaviors. In this work, we attempt to address these limitations while providing three main contributions. First, we propose a novel hierarchical behavioral model, HiBRID, to predict the poachers' behaviors wherein the rangers' imperfect detection of poaching signs is taken into account --- a significant advance towards existing behavioral models in security games. Furthermore, HiBRID incorporates the temporal effect on the poachers' behaviors. The model also does not require a known number of attackers. Second, we provide two new heuristics: \textit{parameter separation} and \textit{target abstraction} to reduce the computational complexity in learning the model parameters. Finally, we use the real-world data collected in Queen Elizabeth National Park (QENP) in Uganda over 12 years to evaluate the prediction accuracy of our new model.
Preface: The Beyond NP Workshop
Darwiche, Adnan (University of California, Los Angeles) | Marquest-Silva, Joao (University of Lisbon) | Marquis, Pierre (Université d’Artois)
A new computational paradigm has emerged in computer both Renault and Toyota have deployed online configuration science over the past few decades, which is exemplified by systems based on knowledge compilation). QBF solvers the use of SAT solvers to tackle problems in the complexity have been used in model checking, verification, debugging, class NP. Finally, function problem solvers have and engineering investment is made towards developing been used in model-based diagnosis, design debugging, highly efficient solvers for a prototypical problem CAD and bioinformatics. The cost of this investment is then on a variety of topics, including algorithms; descriptions amortized as these solvers are applied to a broader class of of implementations and/or evaluations of beyond NP problems via reductions (in contrast to developing dedicated solvers; their applications (including encodings); the complexity algorithms for each encountered problem). SAT solvers, classes they reach; and their connections to one for example, are now routinely used to solve problems in another.
Cost-Effective Feature Selection and Ordering for Personalized Energy Estimates
Early, Kirstin (Carnegie Mellon University) | Fienberg, Stephen (Carnegie Mellon University) | Mankoff, Jennifer (Carnegie Mellon University)
Selecting homes with energy-efficient infrastructure is important for renters, because infrastructure influences energy consumption more than in-home behavior.Personalized energy estimates can guide prospective tenants toward energy-efficient homes, but this information is not readily available. Utility estimates are not typically offered to house-hunters, and existing technologies like carbon calculators require users to answer (prohibitively) many questions that may require considerable research to answer. For the task of providing personalized utility estimates to prospective tenants, we present a cost-based model for feature selection at training time, where all features are available and costs assigned to each feature reflect the difficulty of acquisition. At test time, we have immediate access to some features but others are difficult to acquire (costly). In this limited-information setting, we strategically order questions we ask each user, tailored to previous information provided, to give the most accurate predictions while minimizing the cost to users. During the critical first 10 questions that our approach selects, prediction accuracy improves equally to fixed order approaches, but prediction certainty is higher.
Activity Recognition Through Complex Event Processing: First Findings
Hallé, Sylvain (Université du Québec à Chicoutimi) | Gaboury, Sébastien (Université du Québec à Chicoutimi) | Bouchard, Bruno (Université du Québec à Chicoutimi)
The activities of daily living of a patient in a smart home environment can be detected to a large extent by the real-time analysis of characteristics of the habitat's electrical consumption. However, reasoning over the conduct of these activities occurs at a much higher level of abstraction than what the sensors generally produce. In this paper, we leverage the concept of Complex Event Processing (CEP), in which low-level data streams are progressively transformed into higher-level ones, to the task of activity recognition. We show how the use of an appropriate representation for each level of abstraction can greatly simplify the process. We also report on the use of an existing event stream processor to successfully implement the complete chain, from low-level sensor data up to a sequence of discrete and high-level actions.