Statistical Learning
Interpretable agent communication from scratch(with a generic visual processor emerging on the side)
Dessì, Roberto, Kharitonov, Eugene, Baroni, Marco
As deep networks begin to be deployed as autonomous agents, the issue of how they can communicate with each other becomes important. Here, we train two deep nets from scratch to perform realistic referent identification through unsupervised emergent communication. We show that the largely interpretable emergent protocol allows the nets to successfully communicate even about object types they did not see at training time. The visual representations induced as a by-product of our training regime, moreover, show comparable quality, when re-used as generic visual features, to a recent self-supervised learning model. Our results provide concrete evidence of the viability of (interpretable) emergent deep net communication in a more realistic scenario than previously considered, as well as establishing an intriguing link between this field and self-supervised visual learning.
Transient Chaos in BERT
Inoue, Katsuma, Ohara, Soh, Kuniyoshi, Yasuo, Nakajima, Kohei
Language is an outcome of our complex and dynamic human-interactions and the technique of natural language processing (NLP) is hence built on human linguistic activities. Bidirectional Encoder Representations from Transformers (BERT) has recently gained its popularity by establishing the state-of-the-art scores in several NLP benchmarks. A Lite BERT (ALBERT) is literally characterized as a lightweight version of BERT, in which the number of BERT parameters is reduced by repeatedly applying the same neural network called Transformer's encoder layer. By pre-training the parameters with a massive amount of natural language data, ALBERT can convert input sentences into versatile high-dimensional vectors potentially capable of solving multiple NLP tasks. In that sense, ALBERT can be regarded as a well-designed high-dimensional dynamical system whose operator is the Transformer's encoder, and essential structures of human language are thus expected to be encapsulated in its dynamics. In this study, we investigated the embedded properties of ALBERT to reveal how NLP tasks are effectively solved by exploiting its dynamics. We thereby aimed to explore the nature of human language from the dynamical expressions of the NLP model. Our short-term analysis clarified that the pre-trained model stably yields trajectories with higher dimensionality, which would enhance the expressive capacity required for NLP tasks. Also, our long-term analysis revealed that ALBERT intrinsically shows transient chaos, a typical nonlinear phenomenon showing chaotic dynamics only in its transient, and the pre-trained ALBERT model tends to produce the chaotic trajectory for a significantly longer time period compared to a randomly-initialized one. Our results imply that local chaoticity would contribute to improving NLP performance, uncovering a novel aspect in the role of chaotic dynamics in human language behaviors.
Provably Faster Algorithms for Bilevel Optimization
Yang, Junjie, Ji, Kaiyi, Liang, Yingbin
Bilevel optimization has been widely applied in many important machine learning applications such as hyperparameter optimization and meta-learning. Recently, several momentum-based algorithms have been proposed to solve bilevel optimization problems faster. However, those momentum-based algorithms do not achieve provably better computational complexity than $\mathcal{O}(\epsilon^{-2})$ of the SGD-based algorithm. In this paper, we propose two new algorithms for bilevel optimization, where the first algorithm adopts momentum-based recursive iterations, and the second algorithm adopts recursive gradient estimations in nested loops to decrease the variance. We show that both algorithms achieve the complexity of $\mathcal{O}(\epsilon^{-1.5})$, which outperforms all existing algorithms by the order of magnitude. Our experiments validate our theoretical results and demonstrate the superior empirical performance of our algorithms in hyperparameter applications. Our codes for MRBO, VRBO and other benchmarks are available $\text{online}^1$.
Bayesian Optimization over Hybrid Spaces
Deshwal, Aryan, Belakaria, Syrine, Doppa, Janardhan Rao
We consider the problem of optimizing hybrid structures (mixture of discrete and continuous input variables) via expensive black-box function evaluations. This problem arises in many real-world applications. For example, in materials design optimization via lab experiments, discrete and continuous variables correspond to the presence/absence of primitive elements and their relative concentrations respectively. The key challenge is to accurately model the complex interactions between discrete and continuous variables. In this paper, we propose a novel approach referred as Hybrid Bayesian Optimization (HyBO) by utilizing diffusion kernels, which are naturally defined over continuous and discrete variables. We develop a principled approach for constructing diffusion kernels over hybrid spaces by utilizing the additive kernel formulation, which allows additive interactions of all orders in a tractable manner. We theoretically analyze the modeling strength of additive hybrid kernels and prove that it has the universal approximation property. Our experiments on synthetic and six diverse real-world benchmarks show that HyBO significantly outperforms the state-of-the-art methods.
Automatically Differentiable Random Coefficient Logistic Demand Estimation
The random coefficient logistic demand model of Berry et al. (1995) (henceforth BLP) has been a workhorse of the New Empirical Industrial Organization literature, allowing for varied substitution patterns across products, and accounting for endogeneity of price. The reliability of its estimation has been the subject of rigorous debate (Nevo, 2000; Conlon and Gortmaker, 2020; Knittel and Metaxoglou, 2014), and the estimator itself has been the study of many proposed advances in econometric techniques as a sophisticated yet widely used structural model (Hong et al., 2020; Forneron and Ng, 2020). The most common implementation of the BLP estimator involves the use of a nested fixed point (NFP) as an inner loop within an outer loop of GMM estimation, although we acknowledge the Mathematical Programming with Equilibrium Constraints (MPEC) approach of Dubé et al. (2012), which is beyond the scope of this paper. Dubé et al. (2012) and Conlon and Gortmaker (2020) find that derivative-free optimization algorithms such as the Nelder-Meade or simplex algorithms often fail to converge or converge to the wrong solution. As such, the literature has settled on the use of analytical derivatives with a derivative-based optimization algorithm such as L-BFGS. Nevo (2000) provides the analytical derivative for demand-only (DO) BLP in detail, and Conlon and Gortmaker (2020) indicate that the same is possible for demand-and-supply (DS) BLP, although it involves tensor products.
Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
von Kügelgen, Julius, Sharma, Yash, Gresele, Luigi, Brendel, Wieland, Schölkopf, Bernhard, Besserve, Michel, Locatello, Francesco
Self-supervised representation learning has shown remarkable success in a number of domains. A common practice is to perform data augmentation via hand-crafted transformations intended to leave the semantics of the data invariant. We seek to understand the empirical success of this approach from a theoretical perspective. We formulate the augmentation process as a latent variable model by postulating a partition of the latent representation into a content component, which is assumed invariant to augmentation, and a style component, which is allowed to change. Unlike prior work on disentanglement and independent component analysis, we allow for both nontrivial statistical and causal dependencies in the latent space. We study the identifiability of the latent representation based on pairs of views of the observations and prove sufficient conditions that allow us to identify the invariant content partition up to an invertible mapping in both generative and discriminative settings. We find numerical simulations with dependent latent variables are consistent with our theory. Lastly, we introduce Causal3DIdent, a dataset of high-dimensional, visually complex images with rich causal dependencies, which we use to study the effect of data augmentations performed in practice.
Inference for Network Regression Models with Community Structure
Pan, Mengjie, McCormick, Tyler H., Fosdick, Bailey K.
Network regression models, where the outcome comprises the valued edge in a network and the predictors are actor or dyad-level covariates, are used extensively in the social and biological sciences. Valid inference relies on accurately modeling the residual dependencies among the relations. Frequently homogeneity assumptions are placed on the errors which are commonly incorrect and ignore critical, natural clustering of the actors. In this work, we present a novel regression modeling framework that models the errors as resulting from a community-based dependence structure and exploits the subsequent exchangeability properties of the error distribution to obtain parsimonious standard errors for regression parameters.
Diving Deep into Linear Regression and Polynomial Regression
I'm almost certain that now you might want to learn about these branches in greater detail. Worry not, I'll surely open the gates to these subsets in the posts to come. If you missed my post, you can find it at the following link: Branches of Artificial Intelligence. Previously, we discussed Machine Learning. We also discussed its subsets -- Supervised Learning, Unsupervised Learning, and Reinforcement Learning.
Supervised Learning -- K Nearest Neighbors Algorithm (KNN)
This article explains one of the simplest machine learning algorithm K Nearest Neighbors(KNN). KNN classifier and KNN regression are explained with examples in this article. K nearest neighbors algorithm basically predicts on the principle that the data is in the same class as the nearest data. According to name of the algorithm, "nearest neighbors" represents the closest data and "k" represents how many closest data is chosen. K value is a hyper parameter so it is tuned by the user and each trial usually gives different results.
Node2vec explained graphically
Node2vec is an embedding method that transforms graphs (or networks) into numerical representations [1]. For example, given a social network where people (nodes) interact via relations (edges), node2vec generates numerical representation, i.e., a list of numbers, to represent each person. This representation preserves the structure of the original network in some sense so that people closely related have similar representations and vice versa. In this article, we'll go through the intuition of the node2vec method and, in particular, how second order random walk on graph works via a series of animations. A random walk on a graph can be thought of as a "walker" traversing the graph along the edges of the graph.