Statistical Learning
Function Space Particle Optimization for Bayesian Neural Networks
Wang, Ziyu, Ren, Tongzheng, Zhu, Jun, Zhang, Bo
While Bayesian neural networks (BNNs) have drawn increasing attention, their posterior inference remains challenging, due to the high-dimensional and over-parameterized nature. To address this issue, several highly flexible and scalable variational inference procedures based on the idea of particle optimization have been proposed. These methods directly optimize a set of particles to approximate the target posterior. However, their application to BNNs often yields sub-optimal performance, as such methods have a particular failure mode on over-parameterized models. In this paper, we propose to solve this issue by performing particle optimization directly in the space of regression functions. We demonstrate through extensive experiments that our method successfully overcomes this issue, and outperforms strong baselines in a variety of tasks including prediction, defense against adversarial examples, and reinforcement learning.
Efficient online learning with kernels for adversarial large scale problems
Jézéquel, Rémi, Gaillard, Pierre, Rudi, Alessandro
We are interested in a framework of online learning with kernels for low-dimensional but large-scale and potentially adversarial datasets. Considering the Gaussian kernel, we study the computational and theoretical performance of online variations of kernel Ridge regression. The resulting algorithm is based on approximations of the Gaussian kernel through Taylor expansion. It achieves for $d$-dimensional inputs a (close to) optimal regret of order $O((\log n)^{d+1})$ with per-round time complexity and space complexity $O((\log n)^{2d})$. This makes the algorithm a suitable choice as soon as $n \gg e^d$ which is likely to happen in a scenario with small dimensional and large-scale dataset.
Interaction-aware Factorization Machines for Recommender Systems
Hong, Fuxing, Huang, Dongbo, Chen, Ge
Factorization Machine (FM) is a widely used supervised learning approach by effectively modeling of feature interactions. Despite the successful application of FM and its many deep learning variants, treating every feature interaction fairly may degrade the performance. For example, the interactions of a useless feature may introduce noises; the importance of a feature may also differ when interacting with different features. In this work, we propose a novel model named \emph{Interaction-aware Factorization Machine} (IFM) by introducing Interaction-Aware Mechanism (IAM), which comprises the \emph{feature aspect} and the \emph{field aspect}, to learn flexible interactions on two levels. The feature aspect learns feature interaction importance via an attention network while the field aspect learns the feature interaction effect as a parametric similarity of the feature interaction vector and the corresponding field interaction prototype. IFM introduces more structured control and learns feature interaction importance in a stratified manner, which allows for more leverage in tweaking the interactions on both feature-wise and field-wise levels. Besides, we give a more generalized architecture and propose Interaction-aware Neural Network (INN) and DeepIFM to capture higher-order interactions. To further improve both the performance and efficiency of IFM, a sampling scheme is developed to select interactions based on the field aspect importance. The experimental results from two well-known datasets show the superiority of the proposed models over the state-of-the-art methods.
TensorMap: Lidar-Based Topological Mapping and Localization via Tensor Decompositions
Rambhatla, Sirisha, Sidiropoulos, Nikos D., Haupt, Jarvis
We propose a technique to develop (and localize in) topological maps from light detection and ranging (Lidar) data. Localizing an autonomous vehicle with respect to a reference map in real-time is crucial for its safe operation. Owing to the rich information provided by Lidar sensors, these are emerging as a promising choice for this task. However, since a Lidar outputs a large amount of data every fraction of a second, it is progressively harder to process the information in real-time. Consequently, current systems have migrated towards faster alternatives at the expense of accuracy. To overcome this inherent trade-off between latency and accuracy, we propose a technique to develop topological maps from Lidar data using the orthogonal Tucker3 tensor decomposition. Our experimental evaluations demonstrate that in addition to achieving a high compression ratio as compared to full data, the proposed technique, $\textit{TensorMap}$, also accurately detects the position of the vehicle in a graph-based representation of a map. We also analyze the robustness of the proposed technique to Gaussian and translational noise, thus initiating explorations into potential applications of tensor decompositions in Lidar data analysis.
On the well-posedness of Bayesian inverse problems
The subject of this article is the introduction of a weaker concept of well-posedness of Bayesian inverse problems. The conventional concept of (`Lipschitz') well-posedness in [Stuart 2010, Acta Numerica 19, pp. 451-559] is difficult to verify in practice, especially when considering blackbox models, and probably too strong in many contexts. Our concept replaces the Lipschitz continuity of the posterior measure in the Hellinger distance by just continuity. This weakening is tolerable, since the continuity is in general only used as a stability criterion. The main result of this article is a proof of well-posedness for a large class of Bayesian inverse problems, where very little or no information about the underlying model is available. It includes any Bayesian inverse problem arising when observing finite-dimensional data perturbed by additive, non-degenerate Gaussian noise. Moreover, well-posedness with respect to other probability metrics is investigated, including weak convergence, total variation, Wasserstein, and also the Kullback-Leibler divergence.
Day-Ahead Hourly Forecasting of Power Generation from Photovoltaic Plants
Gigoni, Lorenzo, Betti, Alessandro, Crisostomi, Emanuele, Franco, Alessandro, Tucci, Mauro, Bizzarri, Fabrizio, Mucci, Debora
The ability to accurately forecast power generation from renewable sources is nowadays recognised as a fundamental skill to improve the operation of power systems. Despite the general interest of the power community in this topic, it is not always simple to compare different forecasting methodologies, and infer the impact of single components in providing accurate predictions. In this paper we extensively compare simple forecasting methodologies with more sophisticated ones over 32 photovoltaic plants of different size and technology over a whole year. Also, we try to evaluate the impact of weather conditions and weather forecasts on the prediction of PV power generation. I. INTRODUCTION High penetration levels of Distributed Energy Resources (DERs), typically based on renewable generation, introduce several challenges in power system operation, due to the intrinsic intermittent and uncertain nature of such DERs. In this context, it is fundamental to develop the ability to accurately forecast energy production from renewable sources, like solar photovoltaic (PV), wind power and river hydro, to obtain short-and midterm forecasts. Dispatchability: secure power systems' daily operation mainly relies upon day-ahead dispatches of power plants [1]. Accordingly, meaningful day-ahead plans can be performed only if accurate day-ahead predictions of power generation from renewable sources, together with reliable predictions of the day-ahead load consumption forecasts (e.g., see [2]) are available; Efficiency: as output power fluctuations from intermittent sources may cause frequency and voltage fluctuations in the system (see [3]), some countries have introduced penalties for power generators that fail to accurately predict their power generation for the next day; thus, some energy producers prefer to underestimate their day-ahead power generation forecasts to avoid to incur in penalties in the next day. Monitoring: mismatches between power forecasts and the actually generated power may be also used by energy producers to monitor the plant operation, to evaluate the natural degradation of the efficiency of the plant due to the aging of some components (see [4]) or for early detection of incipient faults.
Towards Efficient Data Valuation Based on the Shapley Value
Jia, Ruoxi, Dao, David, Wang, Boxin, Hubis, Frances Ann, Hynes, Nick, Gurel, Nezihe Merve, Li, Bo, Zhang, Ce, Song, Dawn, Spanos, Costas
"How much is my data worth?" is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. In this paper, we study the problem of data valuation by utilizing the Shapley value, a popular notion of value which originated in coopoerative game theory. The Shapley value defines a unique payoff scheme that satisfies many desiderata for the notion of data value. However, the Shapley value often requires exponential time to compute. To meet this challenge, we propose a repertoire of efficient algorithms for approximating the Shapley value. We also demonstrate the value of each training instance for various benchmark datasets.
Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
Lapuschkin, Sebastian, Wäldchen, Stephan, Binder, Alexander, Montavon, Grégoire, Samek, Wojciech, Müller, Klaus-Robert
Current learning machines have successfully solved hard application problems, reaching high accuracy and displaying seemingly "intelligent" behavior. Here we apply recent techniques for explaining decisions of state-of-the-art learning machines and analyze various tasks from computer vision and arcade games. This showcases a spectrum of problem-solving behaviors ranging from naive and short-sighted, to well-informed and strategic. We observe that standard performance evaluation metrics can be oblivious to distinguishing these diverse problem solving behaviors. Furthermore, we propose our semi-automated Spectral Relevance Analysis that provides a practically effective way of characterizing and validating the behavior of nonlinear learning machines. This helps to assess whether a learned model indeed delivers reliably for the problem that it was conceived for. Furthermore, our work intends to add a voice of caution to the ongoing excitement about machine intelligence and pledges to evaluate and judge some of these recent successes in a more nuanced manner.
Integrated analysis of the urban water-electricity demand nexus in the Midwestern United States
Obringer, Renee, Kumar, Rohini, Nateghi, Roshanak
Considering the interdependencies between water and electricity use is critical for ensuring conservation measures are successful in lowering the net water and electricity use in a city. This water-electricity demand nexus will become even more important as cities continue to grow, causing water and electricity utilities additional stress, especially given the likely impacts of future global climatic and socioeconomic changes. Here, we propose a modeling framework based in statistical learning theory for predicting the climate-sensitive portion of the coupled water-electricity demand nexus. The predictive models were built and tested on six Midwestern cities. The results showed that water use was better predicted than electricity use, indicating that water use is slightly more sensitive to climate than electricity use. Additionally, the results demonstrated the importance of the variability in the El Nino/Southern Oscillation index, which explained the majority of the covariance in the water-electricity nexus. Our modeling results suggest that stronger El Ninos lead to an overall increase in water and electricity use in these cities. The integrated modeling framework presented here can be used to characterize the climate-related sensitivity of the water-electricity demand nexus, accounting for the coupled water and electricity use rather than modeling them separately, as independent variables.
Ordinal Distance Metric Learning with MDS for Image Ranking
Image ranking is to rank images based on some known ranked images. In this paper, we propose an improved linear ordinal distance metric learning approach based on the linear distance metric learning model. By decomposing the distance metric $A$ as $L^TL$, the problem can be cast as looking for a linear map between two sets of points in different spaces, meanwhile maintaining some data structures. The ordinal relation of the labels can be maintained via classical multidimensional scaling, a popular tool for dimension reduction in statistics. A least squares fitting term is then introduced to the cost function, which can also maintain the local data structure. The resulting model is an unconstrained problem, and can better fit the data structure. Extensive numerical results demonstrate the improvement of the new approach over the linear distance metric learning model both in speed and ranking performance.