Goto

Collaborating Authors

 South America


A Convex Formulation for Learning from Crowds

AAAI Conferences

Recently crowdsourcing services are often used to collect a large amount of labeled data for machine learning, since they provide us an easy way to get labels at very low cost and in a short period. The use of crowdsourcing has introduced a new challenge in machine learning, that is, coping with the variable quality of crowd-generated data. Although there have been many recent attempts to address the quality problem of multiple workers, only a few of the existing methods consider the problem of learning classifiers directly from such noisy data. All these methods modeled the true labels as latent variables, which resulted in non-convex optimization problems. In this paper, we propose a convex optimization formulation for learning from crowds without estimating the true labels by introducing personal models of the individual crowd workers. We also devise an efficient iterative method for solving the convex optimization problems by exploiting conditional independence structures in multiple classifiers. We evaluate the proposed method against three competing methods on synthetic data sets and a real crowdsourced data set and demonstrate that the proposed method outperforms the other three methods.


Polarimetric SAR Image Segmentation with B-Splines and a New Statistical Model

arXiv.org Machine Learning

SAR sensors work on the microwaves spectrum, so they are almost immune to adverse weather conditions and they are able to penetrate, to some extent, the surface of certain targets. The first civilian SAR satellite was launched in 1978, and it was followed by a constellation of other similar sensors, mostly devoted to specific applications and in all cases operated at a single frequency and polarization. The Shuttle Imaging Radar-C/X-band SAR (SIRC/XSAR), launched in 1994, could be operated simultaneously at three frequencies, with two of them able to transmit and receive at both horizontal and vertical polarization. This polarimetric capability provides a more complete description of the target [46]. Polarimetric images are multiple complex-valued data sets requiring, thus, specialized models and algorithms.


Hypothesis Testing in Speckled Data with Stochastic Distances

arXiv.org Machine Learning

Images obtained with coherent illumination, as is the case of sonar, ultrasound-B, laser and Synthetic Aperture Radar -- SAR, are affected by speckle noise which reduces the ability to extract information from the data. Specialized techniques are required to deal with such imagery, which has been modeled by the G0 distribution and under which regions with different degrees of roughness and mean brightness can be characterized by two parameters; a third parameter, the number of looks, is related to the overall signal-to-noise ratio. Assessing distances between samples is an important step in image analysis; they provide grounds of the separability and, therefore, of the performance of classification procedures. This work derives and compares eight stochastic distances and assesses the performance of hypothesis tests that employ them and maximum likelihood estimation. We conclude that tests based on the triangular distance have the closest empirical size to the theoretical one, while those based on the arithmetic-geometric distances have the best power. Since the power of tests based on the triangular distance is close to optimum, we conclude that the safest choice is using this distance for hypothesis testing, even when compared with classical distances as Kullback-Leibler and Bhattacharyya.


Nonparametric Edge Detection in Speckled Imagery

arXiv.org Machine Learning

We address the issue of edge detection in Synthetic Aperture Radar imagery. In particular, we propose nonparametric methods for edge detection, and numerically compare them to an alternative method that has been recently proposed in the literature. Our results show that some of the proposed methods display superior results and are computationally simpler than the existing method. An application to real (not simulated) data is presented and discussed.


Generalized Statistical Complexity of SAR Imagery

arXiv.org Machine Learning

A new generalized Statistical Complexity Measure (SCM) was proposed by Rosso et al in 2010. It is a functional that captures the notions of order/disorder and of distance to an equilibrium distribution. The former is computed by a measure of entropy, while the latter depends on the definition of a stochastic divergence. When the scene is illuminated by coherent radiation, image data is corrupted by speckle noise, as is the case of ultrasound-B, sonar, laser and Synthetic Aperture Radar (SAR) sensors. In the amplitude and intensity formats, this noise is multiplicative and non-Gaussian requiring, thus, specialized techniques for image processing and understanding. One of the most successful family of models for describing these images is the Multiplicative Model which leads, among other probability distributions, to the G0 law. This distribution has been validated in the literature as an expressive and tractable model, deserving the "universal" denomination for its ability to describe most types of targets. In order to compute the statistical complexity of a site in an image corrupted by speckle noise, we assume that the equilibrium distribution is that of fully developed speckle, namely the Gamma law in intensity format, which appears in areas with little or no texture. We use the Shannon entropy along with the Hellinger distance to measure the statistical complexity of intensity SAR images, and we show that it is an expressive feature capable of identifying many types of targets.


Aggregating Content and Network Information to Curate Twitter User Lists

arXiv.org Artificial Intelligence

Twitter introduced user lists in late 2009, allowing users to be grouped according to meaningful topics or themes. Lists have since been adopted by media outlets as a means of organising content around news stories. Thus the curation of these lists is important - they should contain the key information gatekeepers and present a balanced perspective on a story. Here we address this list curation process from a recommender systems perspective. We propose a variety of criteria for generating user list recommendations, based on content analysis, network analysis, and the "crowdsourcing" of existing user lists. We demonstrate that these types of criteria are often only successful for datasets with certain characteristics. To resolve this issue, we propose the aggregation of these different "views" of a news story on Twitter to produce more accurate user recommendations to support the curation process.


Critical behavior in a cross-situational lexicon learning scenario

arXiv.org Artificial Intelligence

The problem of early word-learning has been subject of philosophical controversy for centuries [1]. The always visionary Augustine argued that the child makes the connections between words and their referents by understanding the referential intentions of others, thus anticipating the modern theory of mind in about fifteen centuries [2]. In the 17th century, Locke's empiricism supported the associationist viewpoint, which contends that the mechanism of word learning is sensitivity to covariation, i.e., if two events occur at the same time, they become associated. Here we examine a radical offshoot of the associationist approach to lexicon acquisition termed crosssituational or observational learning [3], which asserts that the meaning of a word can be determined by looking for something in common across all observed uses of that word [4]. In other words, learning takes place through the statistical sampling of the contexts in which a word appears.


Preface

AAAI Conferences

From this excellent collection of papers, three for presentation at ICAPS 2012, the were selected for special recognition. ICAPS continues Nguyen, Vien Tran, Tran Cao Son and Enrico the traditional high standards of AIPS and ECP Pontelli were selected for Best Student Paper as an archival forum for new research in the Award. In addition to the oral presentation of these e 45 papers included in this volume, consisting papers, the technical program of this year's of 37 long papers and 8 short papers, are ICAPS conference includes invited talks by those selected for plenary presentation at three distinguished speakers: Robert O. Ambrose ICAPS 2012 from a total of 132 submissions. Topics under various constraints and assumptions, included real-time planning, planning in mixed to empirical evaluation of planning and discrete-continuous domains, planning for systems scheduling techniques in practical applications. Papers in the subareas of optimal planning, probabilistic were encouraged from a range of neighboring and non-deterministic planning, planning disciplines, including model-based and scheduling for transportation, robot path reasoning, hybrid systems, run-time verification, planning, and new developments in heuristics control and robotics.


On Modeling the Tactical Planning of Oil Pipeline Networks

AAAI Conferences

This paper aims at incorporating tactical aspects of oil pipeline networks to the supply chain planning model. The strategic design of supply chains is covered in literature by well understood and recurring patterns such as multi-commodity networks, dynamic parameters over time, capacity on facilities, transportation capacity or facilities with demand, production and inventory. We consider the following characteristics: capacity for in-transit inventory, transit time and flow reversal. Our objective is a better estimate for resources required by the network and therewith allow a more precise optimization of their use. All aspects are modeled to be efficiently solved by linear programming algorithms.


Automated Planning for Liner Shipping Fleet Repositioning

AAAI Conferences

The Liner Shipping Fleet Repositioning Problem (LSFRP) poses a large financial burden on liner shipping firms. During repositioning, vessels are moved between services in a liner shipping network. The LSFRP is characterized by chains of interacting activities, many of which have costs that are a function of their duration; for example, sailing slowly between two ports is cheaper than sailing quickly. Despite its great industrial importance, the LSFRP has received little attention in the literature. We show how the LSFRP can be solved sub-optimally using the planner POPF and optimally with a mixed-integer program (MIP) and a novel method called Temporal Optimization Planning (TOP). We evaluate the performance of each of these techniques on a dataset of real-world instances from our industrial collaborator, and show that automated planning scales to the size of problems faced by industry.