Country
Tech companies in China and U.S. are vying to sell facial recognition software for UAE spy program
As lawmakers, citizens, and company's debate the use of facial recognition software in the U.S., tech giants in America and China have been busy hawking products to eager surveillance states abroad. Among the burgeoning markets, according to a report by Buzzfeed News, are monarchies in the United Arab Emirates (UAE), particularly in Dubai, where political leaders have often jailed citizens and journalists that they deem to be political dissidents. Critics of the UAE include Human Rights Watch (HRW) who has frequently derided the country for its authoritarian tendencies. Private companies like IBM are looking to governments accused of violating human rights as a market for facial recognition software. 'UAE authorities have launched a sustained assault on freedom of expression and association since 2011,' says HRW in its analysis.
US Army is developing 'Google Earth on steroids' that will be able to simulate INSIDE of buildings
An initiative by the U.S. military looks to develop what one researcher is calling'Google Earth on steroids' that maps entire landscapes, helping to simulate environments and train soldiers. In a report from National Defense, one researcher working on the project revealed that the system will be granular enough to map the inside of buildings and eventually entire cities which can then be used in simulated training exercises. The military hopes to inform the creation of these realistic simulations, what they call Simulated Training Environments (STE), by building a comprehensive and highly detailed 3D map of locations around the globe -- an initiative dubbed One World Terrain. An initiative by the U.S. military looks to develop what one researcher is calling ' Google Earth on steroids' that maps entire landscapes. Simulations could help the U.S. military train soldiers and glean useful data in the field While the project may sound like a developer's nightmare, recent advances in drone technology and databases of satellite imagery have brought the project firmly into reality.
Deterministic PAC-Bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, Vaishnavh, Kolter, J. Zico
The ability of overparameterized deep networks to generalize well has been linked to the fact that stochastic gradient descent (SGD) finds solutions that lie in flat, wide minima in the training loss -- minima where the output of the network is resilient to small random noise added to its parameters. So far this observation has been used to provide generalization guarantees only for neural networks whose parameters are either \textit{stochastic} or \textit{compressed}. In this work, we present a general PAC-Bayesian framework that leverages this observation to provide a bound on the original network learned -- a network that is deterministic and uncompressed. What enables us to do this is a key novelty in our approach: our framework allows us to show that if on training data, the interactions between the weight matrices satisfy certain conditions that imply a wide training loss minimum, these conditions themselves {\em generalize} to the interactions between the matrices on test data, thereby implying a wide test loss minimum. We then apply our general framework in a setup where we assume that the pre-activation values of the network are not too small (although we assume this only on the training data). In this setup, we provide a generalization guarantee for the original (deterministic, uncompressed) network, that does not scale with product of the spectral norms of the weight matrices -- a guarantee that would not have been possible with prior approaches.
Interior-point Methods Strike Back: Solving the Wasserstein Barycenter Problem
Ge, Dongdong, Wang, Haoyue, Xiong, Zikai, Ye, Yinyu
Computing the Wasserstein barycenter of a set of probability measures under the optimal transport metric can quickly become prohibitive for traditional second-order algorithms, such as interior-point methods, as the support size of the measures increases. In this paper, we overcome the difficulty by developing a new adapted interior-point method that fully exploits the problem's special matrix structure to reduce the iteration complexity and speed up the Newton procedure. Different from regularization approaches, our method achieves a well-balanced tradeoff between accuracy and speed. A numerical comparison on various distributions with existing algorithms exhibits the computational advantages of our approach. Moreover, we demonstrate the practicality of our algorithm on image benchmark problems including MNIST and Fashion-MNIST.
Gaussian Differential Privacy
Dong, Jinshuo, Roth, Aaron, Su, Weijie J.
Differential privacy has seen remarkable success as a rigorous and practical formalization of data privacy in the past decade. This privacy definition and its divergence based relaxations, however, have several acknowledged weaknesses, either in handling composition of private algorithms or in analyzing important primitives like privacy amplification by subsampling. Inspired by the hypothesis testing formulation of privacy, this paper proposes a new relaxation, which we term `$f$-differential privacy' ($f$-DP). This notion of privacy has a number of appealing properties and, in particular, avoids difficulties associated with divergence based relaxations. First, $f$-DP preserves the hypothesis testing interpretation. In addition, $f$-DP allows for lossless reasoning about composition in an algebraic fashion. Moreover, we provide a powerful technique to import existing results proven for original DP to $f$-DP and, as an application, obtain a simple subsampling theorem for $f$-DP. In addition to the above findings, we introduce a canonical single-parameter family of privacy notions within the $f$-DP class that is referred to as `Gaussian differential privacy' (GDP), defined based on testing two shifted Gaussians. GDP is focal among the $f$-DP class because of a central limit theorem we prove. More precisely, the privacy guarantees of \emph{any} hypothesis testing based definition of privacy (including original DP) converges to GDP in the limit under composition. The CLT also yields a computationally inexpensive tool for analyzing the exact composition of private algorithms. Taken together, this collection of attractive properties render $f$-DP a mathematically coherent, analytically tractable, and versatile framework for private data analysis. Finally, we demonstrate the use of the tools we develop by giving an improved privacy analysis of noisy stochastic gradient descent.
Factorized Inference in Deep Markov Models for Incomplete Multimodal Time Series
Tan, Zhi-Xuan, Soh, Harold, Ong, Desmond C.
Integrating deep learning with latent state space models has the potential to yield temporal models that are powerful, yet tractable and interpretable. Unfortunately, current models are not designed to handle missing data or multiple data modalities, which are both prevalent in real-world data. In this work, we introduce a factorized inference method for Multimodal Deep Markov Models (MDMMs), allowing us to filter and smooth in the presence of missing data, while also performing uncertainty-aware multimodal fusion. We derive this method by factorizing the posterior p(z|x) for non-linear state space models, and develop a variational backward-forward algorithm for inference. Because our method handles incompleteness over both time and modalities, it is capable of interpolation, extrapolation, conditional generation, and label prediction in multimodal time series. We demonstrate these capabilities on both synthetic and real-world multimodal data under high levels of data deletion. Our method performs well even with more than 50% missing data, and outperforms existing deep approaches to inference in latent time series.
AlignFlow: Cycle Consistent Learning from Multiple Domains via Normalizing Flows
Grover, Aditya, Chute, Christopher, Shu, Rui, Cao, Zhangjie, Ermon, Stefano
Given unpaired data from multiple domains, a key challenge is to efficiently exploit these data sources for modeling a target domain. Variants of this problem have been studied in many contexts, such as cross-domain translation and domain adaptation. We propose AlignFlow, a generative modeling framework for learning from multiple domains via normalizing flows. The use of normalizing flows in AlignFlow allows for a) flexibility in specifying learning objectives via adversarial training, maximum likelihood estimation, or a hybrid of the two methods; and b) exact inference of the shared latent factors across domains at test time. We derive theoretical results for the conditions under which AlignFlow guarantees marginal consistency for the different learning objectives. Furthermore, we show that AlignFlow guarantees exact cycle consistency in mapping datapoints from one domain to another. Empirically, AlignFlow can be used for data-efficient density estimation given multiple data sources and shows significant improvements over relevant baselines on unsupervised domain adaptation.
Clustered Gaussian Graphical Model via Symmetric Convex Clustering
Yao, Tianyi, Allen, Genevera I.
However, accurately determining connectivity is not directly observable, numerous techniques which neurons carry out similar neurological tasks via controlled such as correlations and partial correlations have been proposed experiments is both labor-intensive and prohibitively to estimate such functional connectivity from neural expensive on a large scale. Thus, it is of great interest to recording data (see [3] for a comprehensive review). In this cluster neurons that have similar connectivity profiles into work, we define functional connectivity between each pair of functionally coherent groups in a data-driven manner. In this recorded neurons to be their pairwise partial correlation or work, we propose the clustered Gaussian graphical model edges in an undirected GGM in high dimensions. Because (GGM) and a novel symmetric convex clustering penalty the pairwise partial correlation between two neurons takes in an unified convex optimization framework for inferring activities of all the other recorded neurons into account, it functional clusters among neurons from neural activity data.
Interpretable Adversarial Training for Text
Generating high-quality and interpretable adversarial examples in the text domain is a much more daunting task than it is in the image domain. This is due partly to the discrete nature of text, partly to the problem of ensuring that the adversarial examples are still probable and interpretable, and partly to the problem of maintaining label invariance under input perturbations. In order to address some of these challenges, we introduce sparse projected gradient descent (SPGD), a new approach to crafting interpretable adversarial examples for text. SPGD imposes a directional regularization constraint on input perturbations by projecting them onto the directions to nearby word embeddings with highest cosine similarities. This constraint ensures that perturbations move each word embedding in an interpretable direction (i.e., towards another nearby word embedding). Moreover, SPGD imposes a sparsity constraint on perturbations at the sentence level by ignoring word-embedding perturbations whose norms are below a certain threshold. This constraint ensures that our method changes only a few words per sequence, leading to higher quality adversarial examples. Our experiments with the IMDB movie review dataset show that the proposed SPGD method improves adversarial example interpretability and likelihood (evaluated by average per-word perplexity) compared to state-of-the-art methods, while suffering little to no loss in training performance.
Deep multi-class learning from label proportions
Dulac-Arnold, Gabriel, Zeghidour, Neil, Cuturi, Marco, Beyer, Lucas, Vert, Jean-Philippe
The standard setting of supervised classification in machine learning assumes that we have access to a training set of samples and to their labels; our goal is then to estimate a classifier able to predict the label of new samples. In many real-world situations, however, collecting training sets of labeled examples is not possible, and alternative learning scenarios must be considered. We focus in this paper on a particular setting where one has access to bags of examples, and where for each bag only the proportions of the labels in the bag are available; the task is still to learn a classifier to predict the label of individual samples. This setting, which following Yu et al. [2013] we refer to as learning from label proportions (LLP), is relevant in many situations where labeling of individual samples is time-consuming, difficult, or just not possible, while side-channel information can be used to reconstruct the proportions of label within a given bag. For example, Musicant et al. [2007] explain how LLP is a natural setting to analyze single particle mass spectrometry data, while Quadrianto et al. [2009] discuss applications in e-commerce, politics or spam filtering.