Technology
Computing Text Semantic Relatedness Using the Contents and Links of a Hypertext Encyclopedia: Extended Abstract
Yazdani, Majid (EPFL/Idiap Research Institute) | Popescu-Belis, Andrei (Idiap Research Institute)
We propose methods for computing semantic relatedness between words or texts by using knowledge from hypertext encyclopedias such as Wikipedia. A network of concepts is built by filtering the encyclopedia's articles, each concept corresponding to an article. A random walk model based on the notion of Visiting Probability (VP) is employed to compute the distance between nodes, and then between sets of nodes.To transfer learning from the network of concepts to text analysis tasks, we develop two common representation approaches. In the first approach, the shared representation space is the set of concepts in the network and every text is represented in this space. In the second approach, a latent space is used as the shared representation, and a transformation from words to the latent space is trained over VP scores.We applied our methods to four important tasks in natural language processing: word similarity, document similarity, document clustering and classification, and ranking in information retrieval. The performance is state-of-the-art or close to it for each task, thus demonstrating the generality of the proposed knowledge resource and the associated methods.
Revisiting Centrality-as-Relevance: Support Sets and Similarity as Geometric Proximity: Extended abstract
Ribeiro, Ricardo (L2F/INESC-ID Lisboa and ISCTE-IUL) | Matos, David Martins de (L2F/INESC-ID Lisboa and Instituto Superior Técnico)
In automatic summarization, centrality-as- relevance means that the most important content of an information source, or of a collection of information sources, corresponds to the most central passages, considering a representation where such notion makes sense (graph, spatial, etc.). We assess the main paradigms and intro- duce a new centrality-based relevance model for automatic summarization that relies on the use of support sets to better estimate the relevant content. Geometric proximity is used to compute semantic relatedness. Centrality (relevance) is determined by considering the whole input source (and not only local information), and by taking into account the existence of minor topics or lateral subjects in the information sources to be summarized. The method consists in creating, for each passage of the input source, a support set consisting only of the most semantically related passages. Then, the determination of the most relevant content is achieved by selecting the passages that oc- cur in the largest number of support sets. This model produces extractive summaries that are generic, and language- and domain-independent. Thorough automatic evaluation shows that the method achieves state-of-the-art performance, both in written text, and automatically transcribed speech summarization, even when compared to considerably more complex approaches.
The Extended Global Cardinality Constraint: An Empirical Survey: Extended Abstract
Nightingale, Peter (University of St. Andrews)
The Extended Global Cardinality Constraint (EGCC) is an important component of constraint solving systems, since it is very widely used to model diverse problems. The literature contains many different versions of this constraint, which trade strength of inference against computational cost. In this paper, I focus on the highest strength of inference usually considered, enforcing generalized arc consistency (GAC) on the target variables. This work is an extensive empirical survey of algorithms and optimizations, considering both GAC on the target variables, and tightening the bounds of the cardinality variables. I evaluate a number of key techniques from the literature, and report important implementation details of those techniques, which have often not been described in published papers. Two new optimizations are proposed for EGCC. One of the novel optimizations (dynamic partitioning, generalized from AllDifferent) was found to speed up search by 5.6 times in the best case and 1.56 times on average, while exploring the same search tree. The empirical work represents by far the most extensive set of experiments on variants of algorithms for EGCC. Overall, the best combination of optimizations gives a mean speedup of 4.11 times compared to the same implementation without the optimizations. This paper is an extended abstract of the publication in Artificial Intelligence [Nightingale, 2011].
Modeling Social Causality and Responsibility Judgment in Multi-Agent Interactions: Extended Abstract
Mao, Wenji (Chinese Academy of Sciences) | Gratch, Jonathan (University of Southern California)
Based on psychological attribution theory, this paper presents a domain-independent computational model to automate social causality and responsibility judgment according to an agent’s causal knowledge and observations of interaction. The proposed model is also empirically validated via experimental study.
YAGO2: A Spatially and Temporally Enhanced Knowledge Base from Wikipedia: Extended Abstract
Hoffart, Johannes (Max Planck Institute for Informatics) | Suchanek, Fabian M (Max Planck Institute for Informatics) | Berberich, Klaus (Max Planck Institute for Informatics) | Weikum, Gerhard (Max Planck Institute for Informatics)
We present YAGO2, an extension of the YAGO knowledge base, in which entities, facts, and events are anchored in both time and space. YAGO2 is built automatically from Wikipedia, GeoNames, and WordNet. It contains 447 million facts about 9.8 million entities. Human evaluation confirmed an accuracy of 95% of the facts in YAGO2. In this paper, we present the extraction methodology and the integration of the spatio-temporal dimension.
Algorithms for Generating Ordered Solutions for Explicit AND/OR Structures : Extended Abstract
Ghosh, Priyankar (Indian Institute of Technology Kharagpur) | Sharma, Amit (Cornell University) | Chakrabarti, Partha Pratim (Indian Institute of Technology Kharagpur) | Dasgupta, Pallab (Indian Institute of Technology Kharagpur)
We present algorithms for generating alternative solutions for explicit acyclic AND/OR structures in non-decreasing order of cost. Our algorithms use a best first search technique and report the solutions using an implicit representation ordered by cost. Experiments on randomly constructed AND/OR DAGs and problem domains including matrix chain multiplication, finding the secondary structure of RNA, etc, show that the proposed algorithms perform favorably to the existing approach in terms of time and space.
The CQC Algorithm: Cycling in Graphs to Semantically Enrich and Enhance a Bilingual Dictionary: Extended abstract
Flati, Tiziano (La Sapienza University of Rome) | Navigli, Roberto (La Sapienza University of Rome)
Bilingual machine-readable dictionaries are knowledge resources useful in many automatic tasks.However, compared to monolingual computational lexicons like WordNet, bilingual dictionariestypically provide a lower amount of structured information such as lexical and semantic relations, and often do not cover the entire range of possible translations for a word of interest. In this paper we present Cycles and Quasi-Cycles (CQC), a novel algorithm for the automated disambiguation of ambiguous translations in the lexical entries of a bilingual machine-readable dictionary.
Communicating Open Systems: Extended Abstract
d' (University of London) | Inverno, Mark (King’s College London) | Luck, Michael (IIIA, Artificial Intelligence Research Institute / CSIC, Spanish National Research Council) | Noriega, Pablo (IIIA, Artificial Intelligence Research Institute / CSIC, Spanish National Research Council) | Rodriguez-Aguilar, Juan A (IIIA, Artificial Intelligence Research Institute / CSIC, Spanish National Research Council) | Sierra, Carles
Just as conventional institutions are organisationalstructures for coordinating the activities of multipleinteracting individuals, electronic institutions providea computational analogue for coordinating theactivities of multiple interacting software agents.In this paper, we argue that open multi-agent systemscan be effectively designed and implementedas electronic institutions, for which we provide acomprehensive computational model. More specifically,the paper provides an operational semanticsfor electronic institutions, specifying the essentialdata structures, the state representation and the keyoperations necessary to implement them.
Evaluating Indirect Strategies for Chinese — Spanish Statistical Machine Translation: Extended Abstract
Costa-jussà, Marta R. (Institute for Infocomm Research) | Henríquez, Carlos (Universitat Politècnica de Catalunya) | Banchs, Rafael E. (Institute for Infocomm Research)
Although, Chinese and Spanish are two of the most spoken languages in the world, not much research has been done in machine translation for this language pair. This paper focuses on investigating the state-of-the-art of Chinese-to-Spanish statistical machine translation (SMT), which nowadays is one of the most popular approaches to machine translation. We conduct experimental work with the largest of these three corpora to explore alternative SMT strategies by means of using a pivot language. Three alternatives are considered for pivoting: cascading, pseudo-corpus and triangulation. As pivot language, we use either English, Arabic or French. Results show that, for a phrase-based SMT system, English is the best pivot language between Chinese and Spanish. We propose a system output combination using the pivot strategies which is capable of outperforming the direct translation strategy. The main objective of this work is motivating and involving the research community to work in this important pair of languages given their demographic impact.
Detecting and Tracking Disease Outbreaks by Mining Social Media Data
Xie, Yusheng (Northwestern University) | Chen, Zhengzhang (Northwestern University) | Cheng, Yu (Northwestern University) | Zhang, Kunpeng (Northwestern University) | Agrawal, Ankit (Northwestern University) | Liao, Wei-keng (Northwestern University) | Choudhary, Alok (Northwestern University)
The emergence and ubiquity of online social networks have enriched web data with evolving interactions and communities both at mega-scale and in real-time. This data offers an unprecedented opportunity for studying the interaction between society and disease outbreaks. The challenge we describe in this data paper is how to extract and leverage epidemic outbreak insights from massive amounts of social media data and how this exercise can benefit medical professionals, patients, and policymakers alike. We attempt to prepare the research community for this challenge with four datasets. Publishing the four datasets will commoditize the data infrastructure to allow a higher and more efficient focal point for the research community.