Technology
Crowdsourcing Backdoor Identification for Combinatorial Optimization
Bras, Ronan Le (Cornell University) | Bernstein, Richard (Cornell University) | Gomes, Carla P (Cornell University) | Selman, Bart (Cornell University) | Dover, R. Bruce van (Cornell University)
We will show how human computation insights can be key to identifying so-called backdoor variables in combinatorial optimization problems. Backdoor variables can be used to obtain dramatic speed-ups in combinatorial search. Our approach leverages the complementary strength of human input, based on a visual identification of problem structure, crowdsourcing, and the power of combinatorial solvers to exploit complex constraints. We describe our work in the context of the domain of materials discovery. The motivation for considering the materials discovery domain comes from the fact that new materials can provide solutions for key challenges in sustainability, e.g., in energy, new catalysts for more efficient fuel cell technology.
A Multi-Objective Memetic Algorithm for Vehicle Resource Allocation in Sustainable Transportation Planning
Lau, Hoong Chuin (Singapore Management University) | Agussurja, Lucas (Singapore Management University) | Cheng, Shih-Fen (Singapore Management University) | Tan, Pang Jin (DHL Supply Chain Singapore)
Sustainable supply chain management has been an increasingly important topic of research in recent years. At the strategic level, there are computational models which study supply and distribution networks with environmental considerations. At the operational level, there are, for example, routing and scheduling models which are constrained by carbon emissions. Our paper explores work in tactical planning with regards to vehicle resource allocation from distribution centers to customer locations in a multi-echelon logistics network. We formulate the bi-objective optimization problem exactly and design a memetic algorithm to efficiently derive an approximate Pareto front. We illustrate the applicability of our approach with a large real-world dataset.
Information Fusion Based Learning for Frugal Traffic State Sensing
Joshi, Vikas (IBM India Research Labs) | Rajamani, Nithya (IBM India Research Labs) | Katsuki, Takayuki (IBM Tokyo Research Labs) | Prathapaneni, Naveen (IBM India Research Labs,) | Subramaniam, L. V. (IBM India Research Labs)
Traffic sensing is a key baseline input for sustainablecities to plan and administer demand-supplymanagement through better road networks, publictransportation, urban policies etc., Humans sensethe environment frugally using a combination ofcomplementary information signals from differentsensors. For example, by viewing and/or hearingtraffic one could identify the state of traffic on theroad. In this paper, we demonstrate a fusion basedlearning approach to classify the traffic states usinglow cost audio and image data analysis using realworld dataset. Road side collected traffic acousticsignals and traffic image snapshots obtained fromfixed camera are used to classify the traffic conditioninto three broad classes viz., Jam, Mediumand Free. The classification is done on f10sec audio,image snapshot in that 10secg data tuple. Weextract traffic relevant features from audio and imagedata to form a composite feature vector. Inparticular, we extract the audio features comprisingMFCC (Mel-Frequency Cepstral Coefficients)classifier based features, honk events and energypeaks. A simple heuristic based image classifier isused, where vehicular density and number of cornerpoints within the road segment are estimated andare used as features for traffic sensing. Finally thecomposite vector is tested for its ability to discriminatethe traffic classes using Decision tree classifier,SVM classifier, Discriminant classifier and Logisticregression based classifier. Information fusion atmultiple levels (audio, image, overall) shows consistentlybetter performance than individual leveldecision making. Low cost sensor fusion based oncomplementary weak classifiers and noisy featuresstill generates high quality results with an overallaccuracy of 93 - 96%.
Estimating Reference Evapotranspiration for Irrigation Management in the Texas High Plains
Holman, Daniel Ellis (Texas Tech University and Texas A&M AgriLife Research) | Sridharan, Mohan (Texas Tech University) | Gowda, Prasanna (United States Department of Agriculture - Agricultural Research Service) | Porter, Dana (Texas A&M AgriLife Extension Service) | Marek, Thomas (Texas A&M AgriLife Research) | Howell, Terry (United States Department of Agriculture - Agricultural Research Service) | Moorhead, Jerry (United States Department of Agriculture - Agricultural Research Service)
Accurate estimates of daily crop evapotranspiration (ET) are needed for efficient irrigation management in regions where crop water demand exceeds rainfall. Daily grass or alfalfa reference ET values and crop coefficients are widely used to estimate crop water demand. Inaccurate reference ET estimates can hence have a tremendous impact on irrigation costs and the demands on freshwater resources. ET networks calculate reference ET using precise measurements of meteorological data. These networks are typically characterized by gaps in spatial coverage and lack of sufficient funding, creating an immediate need for alternative sources that can fill data gaps without high costs. Although non-agricultural weather stations provide publicly accessible meteorological data, there are concerns that the data may be unsuitable for estimating reference ET due to factors such as weather station siting, data formats and quality control issues. The objective of our research is to enable the use of alternative data sources, adapting sophisticated machine learning algorithms such as Gaussian process models and neural networks to discover and model the nonlinear relationships between non-ET weather station data and the reference ET computed by ET networks. Using data from the Texas High Plains region in the U.S., we demonstrate significant improvement in estimation accuracy in comparison with baseline regression models typically used for irrigation management applications.
Semi-Supervised Learning for Integration of Aerosol Predictions from Multiple Satellite Instruments
Djuric, Nemanja (Temple University) | Kansakar, Lakesh (Temple University) | Vucetic, Slobodan (Temple University)
Aerosol Optical Depth (AOD), recognized as one of the most important quantities in understanding and predicting the Earth's climate, is estimated daily on a global scale by several Earth-observing satellite instruments. Each instrument has different coverage and sensitivity to atmospheric and surface conditions, and, as a result, the quality of AOD estimated by different instruments varies across the globe. We present a method for learning how to aggregate AOD estimations from multiple satellite instruments into a more accurate estimation. The proposed method is semi-supervised, as it is able to learn from a small number of labeled data, where labels come from a few accurate and expensive ground-based instruments, and a large number of unlabeled data. The method uses a latent variable to partition the data, so that in each partition the expert AOD estimations are aggregated in a different, optimal way. We applied the method to combine AOD estimations from 5 instruments aboard 4 satellites, and the results indicate that it can successfully exploit labeled and unlabeled data to produce accurate aggregated AOD estimations.
Short-Term Wind Power Forecasting Using Gaussian Processes
Chen, Niya (Beihang University) | Qian, Zheng (Beihang University) | Nabney, Ian T. (Aston University) | Meng, Xiaofeng (Beihang University)
Since wind has an intrinsically complex and stochastic nature, accurate wind power forecasts are necessary for the safety and economics of wind energy utilization. In this paper, we investigate a combination of numeric and probabilistic models: one-day-ahead wind power forecasts were made with Gaussian Processes (GPs) applied to the outputs of a Numerical Weather Prediction (NWP) model. Firstly the wind speed data from NWP was corrected by a GP. Then, as there is always a defined limit on power generated in a wind turbine due the turbine controlling strategy, a Censored GP was used to model the relationship between the corrected wind speed and power output. To validate the proposed approach, two real world datasets were used for model construction and testing. The simulation results were compared with the persistence method and Artificial Neural Networks (ANNs); the proposed model achieves about 11% improvement in forecasting accuracy (Mean Absolute Error) compared to the ANN model on one dataset, and nearly 5% improvement on another.
Automatic Name-Face Alignment to Enable Cross-Media News Retrieval
Zhang, Yuejie (Fudan University) | Wu, Wei (Fudan University) | Li, Yang (Fudan University) | Jin, Cheng (Fudan University) | Xue, Xiangyang (Fudan University) | Fan, Jianping (The University of North Carolina at Charlotte)
A new algorithm is developed in this paper to support automatic name-face alignment for achieving more accurate cross-media news retrieval. We focus on extracting valuable information from large amounts of news images and their captions, where multi-level image-caption pairs are constructed for characterizing both significant names with higher salience and their cohesion with human faces extracted from news images. To remedy the issue of lacking enough related information for rare name, Web mining is introduced to acquire the extra multimodal information. We also emphasize on an optimization mechanism by our Improved Self-Adaptive Simulated Annealing Genetic Algorithm to verify the feasibility of alignment combinations. Our experiments have obtained very positive results.
Social Influence Locality for Modeling Retweeting Behaviors
Zhang, Jing (Tsinghua University) | Liu, Biao (Tsinghua University) | Tang, Jie (Tsinghua University) | Chen, Ting (Tsinghua University) | Li, Juanzi (Tsinghua University)
We study an interesting phenomenon of social influence locality in a large microblogging network, which suggests that users' behaviors are mainly influenced by close friends in their ego networks. We provide a formal definition for the notion of social influence locality and develop two instantiation functions based on pairwise influence and structural diversity. The defined influence locality functions have strong predictive power. Without any additional features, we can obtain a F1-score of 71.65% for predicting users' retweet behaviors by training a logistic regression classifier based on the defined functions. Our analysis also reveals several intriguing discoveries. For example, though the probability of a user retweeting a microblog is positively correlated with the number of friends who have retweeted the microblog, it is surprisingly negatively correlated with the number of connected circles that are formed by those friends.
Social Collaborative Filtering by Trust
Yang, Bo (Jilin University) | Lei, Yu (Jilin University) | Liu, Dayou (Jilin University) | Liu, Jiming (Hong Kong Baptist University)
To accurately and actively provide users with their potentially interested information or services is the main task of a recommender system. Collaborative filtering is one of the most widely adopted recommender algorithms, whereas it is suffering the issues of data sparsity and cold start that will severely degrade quality of recommendations. To address such issues, this article proposes a novel method, trying to improve the performance of collaborative filtering recommendation by means of elaborately integrating twofold sparse information, the conventional rating data given by users and the social trust network among the same users. It is a model-based method adopting matrix factorization technique to map users into low-dimensional latent feature spaces in terms of their trust relationship, aiming to reflect users’ reciprocal influence on their own opinions more reasonably. The validations against a real-world dataset show that the proposed method performs much better than state-of-the-art recommendation algorithms for social collaborative filtering by trust.
Promoting Diversity in Recommendation by Entropy Regularizer
Qin, Lijing (Tsinghua University) | Zhu, Xiaoyan (Tsinghua University)
We study the problem of diverse promoting recommendation task: selecting a subset of diverse items that can better predict a given user's preference. Recommendation techniques primarily based on user or item similarity can suffer from the risk that users cannot get expected information from the over-specified recommendation lists. In this paper, we propose an entropy regularizer to capture the notion of diversity. The entropy regularizer has good properties in that it satisfies monotonicity and submodularity, such that when we combine it with a modular rating set function, we get submodular objective function, which can be maximized approximately by efficient greedy algorithm, with provable constant factor guarantee of optimality. We apply our approach on the top-$K$ prediction problem and evaluate its performance on MovieLens data set, which is a standard database containing movie rating data collected from a popular online movie recommender system. We compare our model with the state-of-the-art recommendation algorithms. Our experiments show that the entropy regularizer effectively captures diversity and hence improves the performance of recommendation task.