Goto

Collaborating Authors

 Country


The News that Matters to You: Design and Deployment of a Personalized News Service

AAAI Conferences

With the growth of online information, many people are challenged in finding and reading the information most important for their interests. From 2008-2010 we built an experimental personalized news system where readers can subscribe to organized channels of information that are curated by experts. AI technology was employed to radically reduce the work load of curators and to efficiently present information to readers. The system has gone through three implementation cycles and processed over 16 million news stories from about 12,000 RSS feeds on over 8000 topics organized by 160 curators for over 600 registered readers. This paper describes the approach, engineering and AI technology of the system.


The Stock Sonar — Sentiment Analysis of Stocks Based on a Hybrid Approach

AAAI Conferences

The Stock Sonar (TSS) is a stock sentiment analysis application based on a novel hybrid approach. While previous work focused on document level sentiment classification, or extracted only generic sentiment at the phrase level, TSS integrates sentiment dictionaries, phrase-level compositional patterns, and predicate-level semantic events. TSS generates precise in text sentiment tagging as well as sentiment-oriented event summaries for a given stock, which are also aggregated into sentiment scores. Hence, TSS allows investors to get the essence of thousands of articles every day and may help them to make timely, informed trading decisions. The extracted sentiment is also shown to improve the accuracy of an existing document-level sentiment classifier.


The Glass Infrastructure: Using Common Sense to Create a Dynamic, Place-Based Social Information System

AAAI Conferences

Most organizations have a wealth of knowledge about themselves available online, but little for a visitor to interact with on-site. At the MIT Media Lab, we have designed and deployed a novel intelligent signage system, the Glass Infrastructure (GI) that enables small groups of users to physically interact with this data and to discover the latent connections between people, projects, and ideas. The displays are built on an adaptive, unsupervised model of the organization developed using dimensionality reduction and common sense knowledge which automatically classifies and organizes the information. The GI is currently in daily use at the lab. We discuss the AI model’s development, the integration of AI into an HCI interface, and the use of the GI during the lab’s peak visitor periods. We show that the GI is used repeatedly by lab visitors and provides a window into the workings of the organization.


Abductive Inference for Combat: Using SCARE-S2 to Find High-Value Targets in Afghanistan

AAAI Conferences

Recently, geospatial abduction was introduced by the authors in [Shakarian et. al. 2010] as a way to infer unobserved geographic phenomena from a set of known observations and constraints between the two. In this paper, we introduce the SCARE-S2 software tool which applies geospatial abduction to the environment of Afghanistan. Unlike previous work, where we looked for small weapon caches supporting local attacks, here we look for insurgent high-value targets (HVT's), supporting insurgent operations in two provinces. These HVT's include the locations of insurgent leaders and major supply depots. Applying this method of inference to Afghanistan introduces several practical issues not addressed in previous work. Namely, we are conducting inference in a much larger area (24,940 sq km as compared to 675 sq km in previous work), on more varied terrain, and must consider the influence of many local tribes. We address all of these problems and evaluate our software on 6 months of real-world counter-insurgency data. We show that we are able to abduce regions of a relatively small area (on average, under 100 sq km and each containing, on average, 4.8 villages) that are more dense with HVT's (35 X more than the overall area considered).


A Machine Learning Based System for Semi-Automatically Redacting Documents

AAAI Conferences

Redacting text documents has traditionally been a mostly manual activity, making it expensive and prone to disclosure risks. This paper describes a semi-automated system to ensure a specified level of privacy in text data sets. Recent work has attempted to quantify the likelihood of privacy breaches for text data. We build on these notions to provide a means of obstructing such breaches by framing it as a multi-class classification problem. Our system gives users fine-grained control over the level of privacy needed to obstruct sensitive concepts present in that data. Additionally, our system is designed to respect a user-defined utility metric on the data (such as disclosure of a particular concept), which our methods try to maximize while anonymizing. We describe our redaction framework, algorithms, as well as a prototype tool built in to Microsoft Word that allows enterprise users to redact documents before sharing them internally and obscure client specific information. In addition we show experimental evaluation using publicly available data sets that show the effectiveness of our approach against both automated attackers and human subjects.The results show that we are able to preserve the utility of a text corpus while reducing disclosure risk of the sensitive concept.


Monitoring Entities in an Uncertain World: Entity Resolution and Referential Integrity

AAAI Conferences

This paper describes a system to help intelligence analysts track and analyze information being published in multiple sources, particularly open sources on the Web. The system integrates technology for Web harvesting, natural language extraction, and network analytics, and allows analysts to view and explore the results via a Web application. One of the difficult problems we address is the entity resolution problem, which occurs when there are multiple, differing ways to refer to the same entity. The problem is particularly complex when noisy data is being aggregated over time, there is no clean master list of entities, and the entities under investigation are intentionally being deceptive. Our system must not only perform entity resolution with noisy data, but must also gracefully recover when entity resolution mistakes are subsequently corrected. We present a case study in arms trafficking that illustrates the issues, and describe how they are addressed.


Detecting Falls with Location Sensors and Accelerometers

AAAI Conferences

Due to the rapid aging of the population, many technical solutions for the care of the elderly are being developed, often involving fall detection with accelerometers. We present a novel approach to fall detection with location sensors. In our application, a user wears up to four tags on the body whose locations are detected with radio sensors. This makes it possible to recognize the user’s activity, including falling any lying afterwards, and the context in terms of the location in the apartment. We compared fall detection using location sensors, accelerometers and accelerometers combined with the context. A scenario consisting of events difficult to recognize as falls or non-falls was used for the comparison. The accuracy of the methods that utilized the context was almost 40 percentage points higher compared to the methods without the context. The accuracy of pure location-based methods was around 10 percentage points higher than the accuracy of accelerometers combined with the context.


Accelerating the Discovery of Data Quality Rules: A Case Study

AAAI Conferences

Poor quality data is a growing and costly problem that affects many enterprises across all aspects of their business ranging from operational efficiency to revenue protection. In this paper, we present an application -- Data Quality Rules Accelerator (DQRA) -- that accelerates Data Quality (DQ) efforts (e.g. data profiling and cleansing) by automatically discovering DQ rules for detecting inconsistencies in data. We then present two evaluations. The first evaluation compares DQRA to existing solutions; and shows that DQRA either outperformed or achieved performance comparable with these solutions on metrics such as precision, recall, and runtime. The second evaluation is a case study where DQRA was piloted at a large utilities company to improve data quality as part of a legacy migration effort. DQRA was able to discover rules that detected data inconsistencies directly impacting revenue and operational efficiency. Moreover, DQRA was able to significantly reduce the amount of effort required to develop these rules compared to the state of the practice. Finally, we describe ongoing efforts to deploy DQRA.


Learning by Demonstration Technology for Military Planning and Decision Making: A Deployment Story

AAAI Conferences

Learning by demonstration technology has long held the promise to empower non-programmers to customize and extend software. We describe the deployment of a learning by demonstration capability to support user creation of automated procedures in a collaborative planning environment that is used widely by the U.S. Army. This technology, which has been in operational use since the summer of 2010, has helped to reduce user workloads by automating repetitive and time-consuming tasks. The technology has also provided the unexpected benefit of enabling standardization of products and processes. 


Modeling Player Retention in Madden NFL 11

AAAI Conferences

Video games are increasingly producing huge datasets available for analysis resulting from players engaging in interactive environments. These datasets enable investigation of individual player behavior at a massive scale, which can lead to reduced production costs and improved player retention. We present an approach for modeling player retention in Madden NFL 11, a commercial football game. Our approach encodes gameplay patterns of specific players as feature vectors and models player retention as a regression problem. By building an accurate model of player retention, we are able to identify which gameplay elements are most influential in maintaining active players. The outcome of our tool is recommendations which will be used to influence the design of future titles in the Madden NFL series.