Goto

Collaborating Authors

 Asia


Handwritten Bangla Alphabet Recognition using an MLP Based Classifier

arXiv.org Artificial Intelligence

The work presented here involves the design of a Multi Layer Perceptron (MLP) based classifier for recognition of handwritten Bangla alphabet using a 76 element feature set Bangla is the second most popular script and language in the Indian subcontinent and the fifth most popular language in the world. The feature set developed for representing handwritten characters of Bangla alphabet includes 24 shadow features, 16 centroid features and 36 longest-run features. Recognition performances of the MLP designed to work with this feature set are experimentally observed as 86.46% and 75.05% on the samples of the training and the test sets respectively. The work has useful application in the development of a complete OCR system for handwritten Bangla text.


An MLP based Approach for Recognition of Handwritten `Bangla' Numerals

arXiv.org Artificial Intelligence

The work presented here involves the design of a Multi Layer Perceptron (MLP) based pattern classifier for recognition of handwritten Bangla digits using a 76 element feature vector. Bangla is the second most popular script and language in the Indian subcontinent and the fifth most popular language in the world. The feature set developed for representing handwritten Bangla numerals here includes 24 shadow features, 16 centroid features and 36 longest-run features. On experimentation with a database of 6000 samples, the technique yields an average recognition rate of 96.67% evaluated after three-fold cross validation of results. It is useful for applications related to OCR of handwritten Bangla Digit and can also be extended to include OCR of handwritten characters of Bangla alphabet.


Consistency Techniques for Flow-Based Projection-Safe Global Cost Functions in Weighted Constraint Satisfaction

Journal of Artificial Intelligence Research

Many combinatorial problems deal with preferences and violations, the goal of which is to find solutions with the minimum cost. Weighted constraint satisfaction is a framework for modeling such problems, which consists of a set of cost functions to measure the degree of violation or preferences of different combinations of variable assignments. Typical solution methods for weighted constraint satisfaction problems (WCSPs) are based on branch-and-bound search, which are made practical through the use of powerful consistency techniques such as AC*, FDAC*, EDAC* to deduce hidden cost information and value pruning during search. These techniques, however, are designed to be efficient only on binary and ternary cost functions which are represented in table form. In tackling many real-life problems, high arity (or global) cost functions are required. We investigate efficient representation scheme and algorithms to bring the benefits of the consistency techniques to also high arity cost functions, which are often derived from hard global constraints from classical constraint satisfaction. The literature suggests some global cost functions can be represented as flow networks, and the minimum cost flow algorithm can be used to compute the minimum costs of such networks in polynomial time. We show that naive adoption of this flow-based algorithmic method for global cost functions can result in a stronger form of null-inverse consistency. We further show how the method can be modified to handle cost projections and extensions to maintain generalized versions of AC* and FDAC* for cost functions with more than two variables. Similar generalization for the stronger EDAC* is less straightforward. We reveal the oscillation problem when enforcing EDAC* on cost functions sharing more than one variable. To avoid oscillation, we propose a weak version of EDAC* and generalize it to weak EDGAC* for non-binary cost functions. Using various benchmarks involving the soft variants of hard global constraints ALLDIFFERENT, GCC, SAME, and REGULAR, empirical results demonstrate that our proposal gives improvements of up to an order of magnitude when compared with the traditional constraint optimization approach, both in terms of time and pruning.


Global Dynamics of Online Group Conversations

AAAI Conferences

Public online groups allow individuals to carry out conver- sations of common interests. Study of such group conversa- tions provides a unique opportunity to study patterns of hu- man conversations without violating individual privacy. The observational studies conducted in this paper are an attempt to identify the main correlates of continued growth of con- versations, thereby clearing the path to developing predictive models user participation. We study temporal evolution of online group discussions. Surprisingly, we find that individual discussion groups dis- play distinctively q-exponential shaped inter-message times to reply distributions, unlike the power law distributions seen in email conversations. We show, using simulations, that the heavy-tailed distribution of time to reply, which we also ob- serve when all data is combined, originate from mixtures of q-exponentials. We also find that popular threads come to be so from the very beginning as opposed to evolving to be more popular as they grow. This raises new possibilities for devel- oping generative models of thread growth.


Tutorials

AAAI Conferences

The ICWSM 2012 conference tutorials will be How to Analyze Massive Social Network Datasets without a Cluster, presented by Derek Ruths; Charting Collections of Connections in Social Media: Creating Maps and Measures with NodeXL, presented by Marc Smith; Evidenced-Based Social Design of Online Communities: Getting to Critical Mass and Encouraging Contributions, presented by Paul Resnick and Robert Kraut; Sentiment Mining from User Generated Content, presented by Lyle Ungar and Ronen Feldman; and Information Extraction for Social Media Anaylsis, presented by Denilson Barbosa.


Identifying Microblogs for Targeted Contextual Advertising

AAAI Conferences

Micro-blogging sites such as Facebook, Twitter, Google+ present a nice opportunity for targeting advertisements that are contextually related to the microblog content. By virtue of the sparse and noisy text makes identifying the microblogs suitable for advertising a very hard problem. In this work, we approach the problem of identifying the microblogs that could be targeted for advertisements as a two-step classification approach. In the first pass, microblogs suitable for advertising are identified. Next, in the second pass, we build a model to find the sentiment of the advertisable microblog. The systems use features derived from the Part-of-speech tags, the tweet content and uses external resources such as query logs and n-gram dictionaries from previously labeled data.This work aims at providing a thorough insight into the problem and analyzing various features to assess which features contribute the most towards identifying the tweets that can be targeted for advertisements.


Finding Influential Authors in Brand-Page Communities

AAAI Conferences

Enterprises are increasingly using social media forums to engage with their customer online- a phenomenon known as Social Customer Relation Management (Social CRM) . In this context, it is important for an enterprise to identify “influential authors” and engage with them on a priority basis. We present a study towards finding influential authors on Twitter forums where an implicit network based on user interactions is created and analyzed. Furthermore, author profile features and user interaction features are combined in a decision tree classification model for finding influential authors. A novel objective evaluation criterion is used for evaluating various features and modeling techniques. We compare our methods with other approaches that use either only the formal connections or only the author profile features and show a significant improvement in the classification accuracy over these baselines as well as over using Klout score.


More or Less: Amount of Personal Information Displayed in Social Network Site Profiles and Its Impact on Viewers’ Intentions to Socialize with the Profile Owner

AAAI Conferences

This paper presents the results of an experiment that employed a 2 (low vs. high information) by 2 (male vs. female profile) design to investigate the relationship between amount of information displayed in a Social Network Site (SNS) profile and profile viewers’ intentions to engage in further social interactions (communicate online, add to SNS profile, and meet face-to-face) with the profile owner. The results indicate that more information increases the likelihood of relationship initiation for male profiles but decreases it for female profiles. Also, viewers are inclined to initiate an interaction when less information is presented in an SNS profile of a person from the opposite sex; but require more information from their own sex.


Tracking Sentiment and Topic Dynamics from Social Media

AAAI Conferences

We propose a dynamic joint sentiment-topic model (dJST) which allows the detection and tracking of views of current and recurrent interests and shifts in topic and sentiment. Both topic and sentiment dynamics are captured by assuming that the current sentiment-topic specific word distributions are generated according to the word distributions at previous epochs. We derive efficient online inference procedures to sequentially update the model with newly arrived data and show the effectiveness of our proposed model on the Mozilla add-on reviews crawled between 2007 and 2011.


Feasibility Study on Detection of Transportation Information Exploiting Twitter as a Sensor

AAAI Conferences

The concept of a smart community has recently been attracting great attention as a means of utilizing energy effectively. One of the modules constituting the smart community is an intelligent transportation system, in which various sensors track movements of people and vehicles in real time to optimize migration pathways or means. Social media have the potential to serve as sensors, since people often post transportation information on such media. This paper presents a feasibility study on detecting information, focusing on train status information, by exploiting Twitter as a sensor. We dealt with two issues: (1) for the ambiguity of textual information expressed in tweets, we utilized heuristic rules in text manipulation, and (2) for the differences in the numbers of tweets among train lines, we optimized parameter values in statistical analysis for each train line. The experimental results show that the F-measure of detecting the information was more than 0.85 and the time taken to detect the information was less than 4 minutes. As a result we confirmed the high potential of detecting transportation information through Twitter.