Goto

Collaborating Authors

 Statistical Learning


Cost-Sensitive Risk Stratification in the Diagnosis of Heart Disease

AAAI Conferences

We investigate machine learning methods for diagnostic screening of heart disease. Coronary heart disease is the leading cause of death in the US, causing more deaths than all types of cancers combined. Early diagnosis of heart disease in women is harder than it is in men and typically requires the administration of several clinical tests on the patient. Most risk stratification methods aggregate the results of such tests, including the risky, invasive procedures that cannot be administered on all patients. In this paper, our goal is to identify patients who are under high-risk of having heart disease and related adverse events, using a minimal number of diagnostic tests, especially less invasive ones. The low frequency of patients with severe heart disease in the dataset is challenging for most conventional machine learning methods. To overcome this problem, we develop and apply a cost-sensitive k nearest neighbor algorithm. Our contributions are two fold: First, we compare the predictive value of several diagnostic procedures for heart disease, including electrocardiography, angiography, radionuclide perfusion and conclude that in womens heart disease, certain combinations of non-invasive techniques are more predictive than some of the widely used invasive procedures. Then, we evaluate held out data and achieve an AUROC over 0.70, signifying valuable clinical utility, using only the least costly and least invasive tests.


Resource Management for Public Sensing

AAAI Conferences

Public sensing is a new research area in the fields of wireless sensor networks and mobile computing. It leverages the mobile sensors and system resources readily available in mobile phones to execute sensing tasks. In order to plan, execute and adapt large-scale sensing tasks, applications need to query for the available resources, e.g. the density of certain sensors. We investigate how such information can be provided, and we propose a resource manager for public sensing. Our primary goal is to minimize the energy consumed by the mobile devices to make public sensing feasible without disturbing users. We propose a cluster-based protocol for collecting local views of the resource state using local ad-hoc communication since this is much more energy-efficient than long-range (e.g. cellular) communication. We compare our solution to a standard approach where mobile devices communicate their resource states using the cellular phone network. We show that 65% of the energy is saved and the communication load on the infrastructure is reduced by 90% while an average delivery ratio of 93% is retained.


Recognizing Continuous Social Engagement Level in Dyadic Conversation by Using Turn-taking and Speech Emotion Patterns

AAAI Conferences

Recognizing social interests plays an important role of aiding human-computer interaction and human collaborative works. The recognition of social interest could be of great help to determine the smoothness of the interaction, which could be an indicator for group work performance and relationship. From socio-psychological theories, social engagement is the observable form of inner social interest, and represented as patterns of turn-taking and speech emotion during a face-to-face conversation. With these two kinds of features, a multi-layer learning structure is proposed to model the continuous trend of engagement. The level of engagement is classified into “high” and “low” two levels according to human-annotated score. In the result of assessing two-level engagemet, the highest accuracy of our model can reach 79.1%.


Twenty-Five Years of Combining Symbolic and Numeric Learning

AAAI Conferences

For nearly 25 years my research group has investigated the use of domain knowledge, expressed in some version of mathematical logic, that is refined or exploited by numeric-based learning algorithms. These include what we called knowledge-based neural networks and knowledge-based support vector machines. I will cover the key ideas of these methods, as well as the behind-the-scenes motivations that lead to them. I will also describe why we switched from using the phrase 'prior knowledge' to using 'advice.' Finally, I will cover some of our recent work on fast learning and inference for Markov Logic Networks (which can be viewed as a knowledge-based graphical model).


Capturing Browsing Interests of Users into Web Usage Profiles

AAAI Conferences

We present a new weighted session similarity measure to capture the browsing interests of users in web usage profiles discovered from web log data. We base our similarity measure on the reasonable assumption that when users spend longer times on pages or revisit pages in the same session, then very likely, such pages are of greater interest to the user. The proposed similarity measure combines structural similarity with session-wise page significance. The latter, representing the degree of user interest, is computed using frequency and duration of a page access. Web usage profiles are generated using this similarity measure by applying a fuzzy clustering algorithm to web log data. For evaluating the effectiveness of the proposed measure, we adapt two model-based collaborative filtering algorithms for recommending pages. Experimental results show considerable improvement in overall performance of recommender systems as compared to use of other existing similarity measures.


Predicting Crowd-Based Translation Quality with Language-Independent Feature Vectors

AAAI Conferences

Research over the past years has shown that machine translation results can be greatly enhanced with the help of mono- or bilingual human contributors, e.g. by asking hu- mans to proofread or correct outputs of machine translation systems. However, it remains difficult to determine the quality of individual revisions. This paper proposes a meth- od to determine the quality of individual contributions by analyzing task-independent data. Examples of such data are completion time, number of keystrokes, etc. An initial evaluation showed promising F-measure values larger than 0.8 for support vector machine and decision tree based classifications of a combined test set of Vietnamese and German translations.


Learning Sociocultural Knowledge via Crowdsourced Examples

AAAI Conferences

Computational systems can use sociocultural knowledge to understand human behavior and interact with humans in more natural ways. However, such systems are limited by their reliance on hand-authored sociocultural knowledge and models. We introduce an approach to automatically learn robust, script-like sociocultural knowledge from crowdsourced narratives. Crowdsourcing, the use of anonymous human workers, provides an opportunity for rapidly acquir­ing a corpus of examples of situations that are highly specialized for our purpose yet sufficiently varied, from which we can learn a versatile script. We describe a semi-automated process by which we query human workers to write natural language narrative examples of a given situation and learn the set of events that can occur and the typical even ordering.


Learning from Crowds and Experts

AAAI Conferences

Crowdsourcing services are often used to collect a large amount of labeled data for machine learning. Although they provide us an easy way to get labels at very low cost in a short period, they have serious limitations. One of them is the variable quality of the crowd-generated data. There have been many attempts to increase the reliability of crowd-generated data and the quality of classifiers obtained from such data. However, in these problem settings, relatively few researchers have tried using expert-generated data to achieve further improvements. In this paper, we extend three models that deal with the problem of learning from crowds to utilize ground truths: a latent class model, a personal classifier model, and a data-dependent error model. We evaluate the proposed methods against two baseline methods on a real data set to demonstrate the effectiveness of combining crowd-generated data and expert-generated data.


Crowdclustering with Sparse Pairwise Labels: A Matrix Completion Approach

AAAI Conferences

Crowdsourcing utilizes human ability by distributing tasks to a large number of workers. It is especially suitable for solving data clustering problems because it provides a way to obtain a similarity measure between objects based on manual annotations, which capture the human perception of similarity among objects.This is in contrast to most clustering algorithms that face the challenge of finding an appropriate similarity measure for the given dataset. Several algorithms have been developed for crowdclustering that combine partial clustering results, each obtained by annotations provided by a different worker, into a single data partition. However, existing crowd-clustering approaches require a large number of annotations, due to the noisy nature of human annotations, leading to a high computational cost in addition to the large cost associated with annotation. We address this problem by developing a novel approach for crowclustering that exploits the technique of matrix completion. Instead of using all the annotations, the proposed algorithm constructs a partially observed similarity matrix based on a subset of pairwise annotation labels that are agreed upon by most annotators. It then deploys the matrix completion algorithm to complete the similarity matrix and obtains the final data partition by applying a spectral clustering algorithm to the completed similarity matrix. We show, both theoretically and empirically, that the proposed approach needs only a small number of manual annotations to obtain an accurate data partition. In effect, we highlight the trade-off between a large number of noisy crowdsourced labels and a small number of high quality labels.


Towards Action Representation within the Framework of Conceptual Spaces: Preliminary Results

AAAI Conferences

We propose an approach for the representation of actions based on the conceptual spaces framework developed by Gärdenfors (2004). Action categories are regarded as properties in the sense of Gärdenfors (2011) and are understood as convex regions in action space. Action categories are mainly described by a force signature that represents the forces that act upon a main trajector involved in the action. This force signature is approximated via a representation that specifies the time-indexed position of the trajector relative to several landmarks. We also present a computational approach to extract such representations from video data. We present results on the Motionese dataset consisting of videos of parents demonstrating actions on objects to their children. We evaluate the representations on a clustering and a classification task showing that, while our representations seems to be reasonable, only a handful of actions can be discriminated reliably.