Learning Graphical Models
Markov Blanket Discovery using Minimum Message Length
Li, Yang, Korb, Kevin B, Allison, Lloyd
Causal discovery automates the learning of causal Bayesian networks from data and has been of active interest from their beginning. With the sourcing of large data sets off the internet, interest in scaling up to very large data sets has grown. One approach to this is to parallelize search using Markov Blanket (MB) discovery as a first step, followed by a process of combining MBs in a global causal model. We develop and explore three new methods of MB discovery using Minimum Message Length (MML) and compare them empirically to the best existing methods, whether developed specifically as MB discovery or as feature selection. Our best MML method is consistently competitive and has some advantageous features.
Efficient Bayesian Sampling Using Normalizing Flows to Assist Markov Chain Monte Carlo Methods
Gabrié, Marylou, Rotskoff, Grant M., Vanden-Eijnden, Eric
Markov Chain Monte Carlo (MCMC) algorithms (Liu, Since no data set from the target posterior distribution 2008) are nowadays the methods of choice to sample complex is available beforehand, the flow is typically posterior distributions. MCMC methods generate a trained using the reverse Kullback-Leibler (KL) sequence of configurations over which the time average of divergence that only requires samples from a base any suitable observable converges towards its ensemble average distribution. This strategy may perform poorly over some target distribution, here the posterior. This when the posterior is complicated and hard to is achieved by proposing new samples from a proposal density sample with an untrained normalizing flow. Here that is easy to sample, then accepting or rejecting them we explore a distinct training strategy, using the using a criterion that guarantees that the transition kernel of direct KL divergence as loss, in which samples the chain is in detailed balance with respect to the posterior from the posterior are generated by (i) assisting density: a popular choice is Metropolis-Hastings criterion.
Architectures of Meaning, A Systematic Corpus Analysis of NLP Systems
Wysocki, Oskar, Florea, Malina, Landers, Donal, Freitas, Andre
Natural Language Processing (NLP) systems have been subjected to a Cambrian explosion of architectural paradigms in the past few years. The scale on the number of contributions and its exponential growth, bring challenges in understanding how NLP architectural patterns evolve and consolidate in different sub-areas and tasks. This paper aims to provide the methodological support for the interpretation of NLP architectural patterns at scale by applying statistical corpus analysis methods over large-scale NLP corpora. We analyse the use of corpus statistics to compute large-scale collocation patterns jointly with graph visualisation methods as a device to interpret architectural patterns at scale. The proposed methods aims to address questions such as: - What is the complete list of architectural patterns present in NLP? - What are the prevailing architectural patterns (classifiers, layers, regularisation, linguistic resources) for each NLP task? - How these patterns are evolving over time and what are the emerging consolidated/canonical architectural motifs?
Difference Between Algorithm and Artificial Intelligence
By 2035 AI could boost average profitability rates by 38 percent and lead to an economic increase of $14 Trillion. The words Artificial Intelligence (AI), and algorithms are most often misused and misunderstood. There are often used interchangeably when they shouldn't be. This leads to unnecessary confusion. In this article, let's understand what AI and algorithms are, and what the difference between them is.
The World of Reality, Causality and Real Artificial Intelligence: Exposing the Great Unknown Unknowns
"All men by nature desire to know." - Aristotle "He who does not know what the world is does not know where he is." - Marcus Aurelius "If I have seen further, it is by standing on the shoulders of giants." "The universe is a giant causal machine. The world is "at the bottom" governed by causal algorithms. Our bodies are causal machines. Our brains and minds are causal AI computers". The 3 biggest unknown unknowns are described and analyzed in terms of human intelligence and machine intelligence. A deep understanding of reality and its causality is to revolutionize the world, its science and technology, AI machines including. The content is the intro of Real AI Project Confidential Report: How to Engineer Man-Machine Superintelligence 2025: AI for Everything and Everyone (AI4EE). It is all a power set of {known, unknown; known unknown}, known knowns, known unknowns, unknown knowns, and unknown unknowns, like as the material universe's material parts: about 4.6% of baryonic matter, about 26.8% of dark matter, and about 68.3% of dark energy. There are a big number of sciences, all sorts and kinds, hard sciences and soft sciences. But what we are still missing is the science of all sciences, the Science of the World as a Whole, thus making it the biggest unknown unknowns. It is what man/AI does not know what it does not know, neither understand, nor aware of its scope and scale, sense and extent. "the universe consists of objects having various qualities and standing in various relationships" (Whitehead, Russell), "the world is the totality of states of affairs" (D. "World of physical objects and events, including, in particular, biological beings; World of mental objects and events; World of objective contents of thought" (K. How the world is still an unknown unknown one could see from the most popular lexical ontology, WordNet,see supplement. The construct of the world is typically missing its essential meaning, "the world as a whole", the world of reality, the ultimate totality of all worlds, universes, and realities, beings, things, and entities, the unified totalities. The world or reality or being or existence is "all that is, has been and will be". Of which the physical universe and cosmos is a key part, as "the totality of space and times and matter and energy, with all causative fundamental interactions".
High-level Decisions from a Safe Maneuver Catalog with Reinforcement Learning for Safe and Cooperative Automated Merging
Kamran, Danial, Ren, Yu, Lauer, Martin
Reinforcement learning (RL) has recently been used for solving challenging decision-making problems in the context of automated driving. However, one of the main drawbacks of the presented RL-based policies is the lack of safety guarantees, since they strive to reduce the expected number of collisions but still tolerate them. In this paper, we propose an efficient RL-based decision-making pipeline for safe and cooperative automated driving in merging scenarios. The RL agent is able to predict the current situation and provide high-level decisions, specifying the operation mode of the low level planner which is responsible for safety. In order to learn a more generic policy, we propose a scalable RL architecture for the merging scenario that is not sensitive to changes in the environment configurations. According to our experiments, the proposed RL agent can efficiently identify cooperative drivers from their vehicle state history and generate interactive maneuvers, resulting in faster and more comfortable automated driving. At the same time, thanks to the safety constraints inside the planner, all of the maneuvers are collision free and safe.
Obtaining Causal Information by Merging Datasets with MAXENT
Mejia, Sergio Hernan Garrido, Kirschbaum, Elke, Janzing, Dominik
The investigation of the question "which treatment has a causal effect on a target variable?" is of particular relevance in a large number of scientific disciplines. This challenging task becomes even more difficult if not all treatment variables were or even cannot be observed jointly with the target variable. Another, similarly important and challenging task is to quantify the causal influence of a treatment on a target in the presence of confounders. In this paper, we discuss how causal knowledge can be obtained without having observed all variables jointly, but by merging the statistical information from different datasets. We first show how the maximum entropy principle can be used to identify edges among random variables when assuming causal sufficiency and an extended version of faithfulness. Additionally, we derive bounds on the interventional distribution and the average causal effect of a treatment on a target variable in the presence of confounders. In both cases we assume that only subsets of the variables have been observed jointly.
Adversarial Attack for Uncertainty Estimation: Identifying Critical Regions in Neural Networks
Alarab, Ismail, Prakoonwit, Simant
We propose a novel method to capture data points near decision boundary in neural network that are often referred to a specific type of uncertainty. In our approach, we sought to perform uncertainty estimation based on the idea of adversarial attack method. In this paper, uncertainty estimates are derived from the input perturbations, unlike previous studies that provide perturbations on the model's parameters as in Bayesian approach. We are able to produce uncertainty with couple of perturbations on the inputs. Interestingly, we apply the proposed method to datasets derived from blockchain. We compare the performance of model uncertainty with the most recent uncertainty methods. We show that the proposed method has revealed a significant outperformance over other methods and provided less risk to capture model uncertainty in machine learning.
Temporal-aware Language Representation Learning From Crowdsourced Labels
Hao, Yang, Zhai, Xiao, Ding, Wenbiao, Liu, Zitao
Learning effective language representations from crowdsourced labels is crucial for many real-world machine learning tasks. A challenging aspect of this problem is that the quality of crowdsourced labels suffer high intra- and inter-observer variability. Since the high-capacity deep neural networks can easily memorize all disagreements among crowdsourced labels, directly applying existing supervised language representation learning algorithms may yield suboptimal solutions. In this paper, we propose \emph{TACMA}, a \underline{t}emporal-\underline{a}ware language representation learning heuristic for \underline{c}rowdsourced labels with \underline{m}ultiple \underline{a}nnotators. The proposed approach (1) explicitly models the intra-observer variability with attention mechanism; (2) computes and aggregates per-sample confidence scores from multiple workers to address the inter-observer disagreements. The proposed heuristic is extremely easy to implement in around 5 lines of code. The proposed heuristic is evaluated on four synthetic and four real-world data sets. The results show that our approach outperforms a wide range of state-of-the-art baselines in terms of prediction accuracy and AUC. To encourage the reproducible results, we make our code publicly available at \url{https://github.com/CrowdsourcingMining/TACMA}.
Input Dependent Sparse Gaussian Processes
Jafrasteh, Bahram, Villacampa-Calvo, Carlos, Hernández-Lobato, Daniel
Gaussian Processes (GPs) are Bayesian models that provide uncertainty estimates associated to the predictions made. They are also very flexible due to their non-parametric nature. Nevertheless, GPs suffer from poor scalability as the number of training instances N increases. More precisely, they have a cubic cost with respect to $N$. To overcome this problem, sparse GP approximations are often used, where a set of $M \ll N$ inducing points is introduced during training. The location of the inducing points is learned by considering them as parameters of an approximate posterior distribution $q$. Sparse GPs, combined with variational inference for inferring $q$, reduce the training cost of GPs to $\mathcal{O}(M^3)$. Critically, the inducing points determine the flexibility of the model and they are often located in regions of the input space where the latent function changes. A limitation is, however, that for some learning tasks a large number of inducing points may be required to obtain a good prediction performance. To address this limitation, we propose here to amortize the computation of the inducing points locations, as well as the parameters of the variational posterior approximation q. For this, we use a neural network that receives the observed data as an input and outputs the inducing points locations and the parameters of $q$. We evaluate our method in several experiments, showing that it performs similar or better than other state-of-the-art sparse variational GP approaches. However, with our method the number of inducing points is reduced drastically due to their dependency on the input data. This makes our method scale to larger datasets and have faster training and prediction times.