Asia
Parallel Depth First Proof Number Search
Kaneko, Tomoyuki (The University of Tokyo)
The depth first proof number search (df-pn) is an effective and popular algorithm for solving and-or tree problems by using proof and disproof numbers. This paper presents a simple but effective parallelization of the df-pn search algorithm for a shared-memory system. In this parallelization, multiple agents autonomously conduct the df-pn with a shared transposition table. For effective cooperation of agents, virtual proof and disproof numbers are introduced for each node, which is an estimation of future proof and disproof numbers by using the number of agents working on the node's descendants as a possible increase. Experimental results on large checkmate problems in shogi, which is a popular chess variant in Japan, show that reasonable increases in speed were achieved with small overheads in memory.
Extracting Ontological Selectional Preferences for Non-Pertainym Adjectives from the Google Corpus
Tanner, John J. (University of Central Florida) | Gomez, Fernando (University of Central Florida)
While there has been much research into using selectional preferences for word sense disambiguation (WSD), much difficulty has been encountered. To facilitate study into this difficulty and aid in WSD in general, a database of the selectional preferences of non-pertainym prenomial adjectives extracted from the Google Web 1T 5-gram Corpus is proposed. A variety of methods for computing the preferences of each adjective over a set of noun categories from WordNet have been evaluated via simulated disambiguation of pseudohomonyms. The best method of these involves computing for each noun category the ratio of single-word common (i.e. not proper) noun lemma types which can co-occur with a given adjective to the number of single-word common noun lemmata whose estimated frequency is greater than a threshold based on the frequency of the adjective. The database produced by this procedure will be made available to the public.
Gaussian Mixture Model with Local Consistency
Liu, Jialu (Zhejiang University) | Cai, Deng (Zhejiang University) | He, Xiaofei (Zhejiang University)
Gaussian Mixture Model (GMM) is one of the most popular data clustering methods which can be viewed as a linear combination of different Gaussian components. In GMM, each cluster obeys Gaussian distribution and the task of clustering is to group observations into different components through estimating each cluster's own parameters. The Expectation-Maximization algorithm is always involved in such estimation problem. However, many previous studies have shown naturally occurring data may reside on or close to an underlying submanifold. In this paper, we consider the case where the probability distribution is supported on a submanifold of the ambient space. We take into account the smoothness of the conditional probability distribution along the geodesics of data manifold. That is, if two observations are close in intrinsic geometry, their distributions over different Gaussian components are similar. Simply speaking, we introduce a novel method based on manifold structure for data clustering, called Locally Consistent Gaussian Mixture Model (LCGMM). Specifically, we construct a nearest neighbor graph and adopt Kullback-Leibler Divergence as the distance measurement to regularize the objective function of GMM. Experiments on several data sets demonstrate the effectiveness of such regularization.
Two-Stage Sparse Representation for Robust Recognition on Large-Scale Database
He, Ran (Dalian University of Technology) | Hu, BaoGang (Chinese Academy of Sciences) | Zheng, Wei-Shi (Queen Mary University of London) | Guo, YanQing (Dalian University of Technology)
This paper proposes a novel robust sparse representation method, called the two-stage sparse representation (TSR), for robust recognition on a large-scale database. Based on the divide and conquer strategy, TSR divides the procedure of robust recognition into outlier detection stage and recognition stage. In the first stage, a weighted linear regression is used to learn a metric in which noise and outliers in image pixels are detected. In the second stage, based on the learnt metric, the large-scale dataset is firstly filtered into a small set according to the nearest neighbor criterion. Then a sparse representation is computed by the non-negative least squares technique. The sparse solution is unique and can be optimized efficiently. The extensive numerical experiments on several public databases demonstrate that the proposed TSR approach generally obtains better classification accuracy than the state of the art Sparse Representation Classification (SRC). At the same time, by using the TSR, a significant reduction of computational cost is reached by over fifty times in comparison with the SRC, which enables the TSR to be deployed more suitably for large-scale dataset.
Learning to Surface Deep Web Content
Wu, Zhaohui (Xi'an Jiaotong University) | Jiang, Lu (Xi'an Jiaotong University) | Zheng, Qinghua (Xi'an Jiaotong University) | Liu, Jun (Xi'an Jiaotong University)
We propose a novel deep web crawling framework based on reinforcement learning. The crawler is regarded as an agent and deep web database as the environment. The agent perceives its current state and submits a selected action (query) to the environment according to Q-value. Based on the framework we develop an adaptive crawling method. Experimental results show that it outperforms the state of art methods in crawling capability and breaks through the assumption of full-text search implied by existing methods.
UserRec: A User Recommendation Framework in Social Tagging Systems
Zhou, Tom Chao (The Chinese University of Hong Kong) | Ma, Hao (The Chinese University of Hong Kong) | Lyu, Michael R. (The Chinese University of Hong Kong) | King, Irwin (The Chinese University of Hong Kong)
Social tagging systems have emerged as an effective way for users to annotate and share objects on the Web. However, with the growth of social tagging systems, users are easily overwhelmed by the large amount of data and it is very difficult for users to dig out information that he/she is interested in. Though the tagging system has provided interest-based social network features to enable the user to keep track of other users' tagging activities, there is still no automatic and effective way for the user to discover other users with common interests. In this paper, we propose a User Recommendation (UserRec) framework for user interest modeling and interest-based user recommendation, aiming to boost information sharing among users with similar interests. Our work brings three major contributions to the research community: (1) we propose a tag-graph based community detection method to model the users' personal interests, which are further represented by discrete topic distributions; (2) the similarity values between users' topic distributions are measured by Kullback-Leibler divergence (KL-divergence), and the similarity values are further used to perform interest-based user recommendation; and (3) by analyzing users' roles in a tagging system, we find users' roles in a tagging system are similar to Web pages in the Internet. Experiments on tagging dataset of Web pages (Yahoo!~Delicious) show that UserRec outperforms other state-of-the-art recommender system approaches.
Preferences and Learning in Multi-Agent Negotiation
Aydogan, Reyhan (Bogazici University)
In online, dynamic environments, the service requested by consumers may not be readily served by the producers. This requires the consumers and producers to negotiate on the content of the service. To automate this process, agents play a key role in e-commerce. As far as the agents' negotiation strategies are concerned, understanding and reasoning on their users' preferences are important to generate the right offers on behalf of their users. Besides taking other participant's needs into account is important to be able to negotiate effectively. However, preferences of participants are almost always private. The best that can happen is that participants may learn each other's preferences through interactions over time. As agents learn each other's preferences, they can provide better-targeted offers and thus enable faster negotiation. My research direction involves representing and reasoning on preferences, and learning preferences though interaction in automated negotiation.
Appliance Recognition and Unattended Appliance Detection for Energy Conservation
Lee, Shih-Chiang (National Taiwan University) | Lin, Gu-Yuan (National Taiwan University) | Jih, Wan-Rong (National Taiwan University) | Hsu, Jane Yung-Jen (National Taiwan University)
Providing energy conservation services becomes a hot research topic because more and more people attach importance to environmental protection. This research proposes a framework that consists of four process models: appliance recognition, activity-appliances model, unattended appliances detection, and energy conservation service. Appliance recognition model can recognizes the operating states of appliances from raw sensing data of electric power. An activity-appliances model has been built to associate activities with appliances according to the data of Open Mind Common Sense Project. Using the relationship between activities can help to detect unattended appliances, which are consuming electric power but not take part in the residentโs activities. After obtain information of appliance operating states and unattended appliances, residents can receive energy conservation services for notifying the energy consumption information. Finally, the experimental results show that dynamic Baysian network approach can achieve higher than 92% accuracy for appliance recognition. Data of activity-appliances model shows most appliances are strong activity-related.
Bridging Common Sense Knowledge Bases with Analogy by Graph Similarity
Kuo, Yen-Ling (National Taiwan University) | Hsu, Jane Yung-jen (National Taiwan University)
Present-day programs are brittle as computers are notoriously lacking in common sense. While significant progress has been made in building large common sense knowledge bases, they are intrinsically incomplete and inconsistent. This paper presents a novel approach to bridging the gaps between multiple knowledge bases, making it possible to answer queries based on knowledge collected from multiple sources without a common ontology. New assertions are found by computing graph similarity with principle component analysis to draw analogies across multiple knowledge bases. Experiments are designed to find new assertions for a Chinese commonsense knowledge base using the OMCS ConceptNet and similarly for WordNet. The assertions are voted by online users to verify that 75.77% / 77.59% for Chinese ConceptNet / WordNet respectively are good, despite the low overlap in coverage among the knowledge bases.
A unified view of Automata-based algorithms for Frequent Episode Discovery
Achar, Avinash, Laxman, Srivatsan, Sastry, P. S.
Frequent Episode Discovery framework is a popular framework in Temporal Data Mining with many applications. Over the years many different notions of frequencies of episodes have been proposed along with different algorithms for episode discovery. In this paper we present a unified view of all such frequency counting algorithms. We present a generic algorithm such that all current algorithms are special cases of it. This unified view allows one to gain insights into different frequencies and we present quantitative relationships among different frequencies. Our unified view also helps in obtaining correctness proofs for various algorithms as we show here. We also point out how this unified view helps us to consider generalization of the algorithm so that they can discover episodes with general partial orders.