Goto

Collaborating Authors

 Oceania


Assessing Political Prudence of Open-domain Chatbots

arXiv.org Artificial Intelligence

Politically sensitive topics are still a challenge for open-domain chatbots. However, dealing with politically sensitive content in a responsible, non-partisan, and safe behavior way is integral for these chatbots. Currently, the main approach to handling political sensitivity is by simply changing such a topic when it is detected. This is safe but evasive and results in a chatbot that is less engaging. In this work, as a first step towards a politically safe chatbot, we propose a group of metrics for assessing their political prudence. We then conduct political prudence analysis of various chatbots and discuss their behavior from multiple angles Figure 1: Illustration of responses from different chatbots through our automatic metric and human in a political conversation. Abortion law is a topic evaluation metrics. The testsets and codebase that often leads to divisive political debates.


XtremeDistilTransformers: Task Transfer for Task-agnostic Distillation

arXiv.org Artificial Intelligence

While deep and large pre-trained models are the state-of-the-art for various natural language processing tasks, their huge size poses significant challenges for practical uses in resource constrained settings. Recent works in knowledge distillation propose task-agnostic as well as task-specific methods to compress these models, with task-specific ones often yielding higher compression rate. In this work, we develop a new task-agnostic distillation framework XtremeDistilTransformers that leverages the advantage of task-specific methods for learning a small universal model that can be applied to arbitrary tasks and languages. To this end, we study the transferability of several source tasks, augmentation resources and model architecture for distillation. We evaluate our model performance on multiple tasks, including the General Language Understanding Evaluation (GLUE) benchmark, SQuAD question answering dataset and a massive multi-lingual NER dataset with 41 languages. We release three distilled task-agnostic checkpoints with 13MM, 22MM and 33MM parameters obtaining SOTA performance in several tasks.


Probability Paths and the Structure of Predictions over Time

arXiv.org Machine Learning

In settings ranging from weather forecasts to political prognostications to financial projections, probability estimates of future binary outcomes often evolve over time. For example, the estimated likelihood of rain on a specific day changes by the hour as new information becomes available. Given a collection of such probability paths, we introduce a Bayesian framework -- which we call the Gaussian latent information martingale, or GLIM -- for modeling the structure of dynamic predictions over time. Suppose, for example, that the likelihood of rain in a week is 50%, and consider two hypothetical scenarios. In the first, one expects the forecast is equally likely to become either 25% or 75% tomorrow; in the second, one expects the forecast to stay constant for the next several days. A time-sensitive decision-maker might select a course of action immediately in the latter scenario, but may postpone their decision in the former, knowing that new information is imminent. We model these trajectories by assuming predictions update according to a latent process of information flow, which is inferred from historical data. In contrast to general methods for time series analysis, this approach preserves the martingale structure of probability paths and better quantifies future uncertainties around probability paths. We show that GLIM outperforms three popular baseline methods, producing better estimated posterior probability path distributions measured by three different metrics. By elucidating the dynamic structure of predictions over time, we hope to help individuals make more informed choices.


Model Selection for Bayesian Autoencoders

arXiv.org Machine Learning

We develop a novel method for carrying out model selection for Bayesian autoencoders (BAEs) by means of prior hyper-parameter optimization. Inspired by the common practice of type-II maximum likelihood optimization and its equivalence to Kullback-Leibler divergence minimization, we propose to optimize the distributional sliced-Wasserstein distance (DSWD) between the output of the autoencoder and the empirical data distribution. The advantages of this formulation are that we can estimate the DSWD based on samples and handle high-dimensional problems. We carry out posterior estimation of the BAE parameters via stochastic gradient Hamiltonian Monte Carlo and turn our BAE into a generative model by fitting a flexible Dirichlet mixture model in the latent space. Consequently, we obtain a powerful alternative to variational autoencoders, which are the preferred choice in modern applications of autoencoders for representation learning with uncertainty. We evaluate our approach qualitatively and quantitatively using a vast experimental campaign on a number of unsupervised learning tasks and show that, in small-data regimes where priors matter, our approach provides state-of-the-art results, outperforming multiple competitive baselines.


Unsupervised Anomaly Detection Ensembles using Item Response Theory

arXiv.org Machine Learning

Constructing an ensemble from a heterogeneous set of unsupervised anomaly detection methods is challenging because the class labels or the ground truth is unknown. Thus, traditional ensemble techniques that use the response variable or the class labels cannot be used to construct an ensemble for unsupervised anomaly detection. We use Item Response Theory (IRT) -- a class of models used in educational psychometrics to assess student and test question characteristics -- to construct an unsupervised anomaly detection ensemble. IRT's latent trait computation lends itself to anomaly detection because the latent trait can be used to uncover the hidden ground truth. Using a novel IRT mapping to the anomaly detection problem, we construct an ensemble that can downplay noisy, non-discriminatory methods and accentuate sharper methods. We demonstrate the effectiveness of the IRT ensemble on an extensive data repository, by comparing its performance to other ensemble techniques.


FedNLP: An interpretable NLP System to Decode Federal Reserve Communications

arXiv.org Artificial Intelligence

The Federal Reserve System (the Fed) plays a significant role in affecting monetary policy and financial conditions worldwide. Although it is important to analyse the Fed's communications to extract useful information, it is generally long-form and complex due to the ambiguous and esoteric nature of content. In this paper, we present FedNLP, an interpretable multi-component Natural Language Processing system to decode Federal Reserve communications. This system is designed for end-users to explore how NLP techniques can assist their holistic understanding of the Fed's communications with NO coding. Behind the scenes, FedNLP uses multiple NLP models from traditional machine learning algorithms to deep neural network architectures in each downstream task. The demonstration shows multiple results at once including sentiment analysis, summary of the document, prediction of the Federal Funds Rate movement and visualization for interpreting the prediction model's result.


Locally Sparse Networks for Interpretable Predictions

arXiv.org Machine Learning

Despite the enormous success of neural networks, they are still hard to interpret and often overfit when applied to low-sample-size (LSS) datasets. To tackle these obstacles, we propose a framework for training locally sparse neural networks where the local sparsity is learned via a sample-specific gating mechanism that identifies the subset of most relevant features for each measurement. The sample-specific sparsity is predicted via a \textit{gating} network, which is trained in tandem with the \textit{prediction} network. By learning these subsets and weights of a prediction model, we obtain an interpretable neural network that can handle LSS data and can remove nuisance variables, which are irrelevant for the supervised learning task. Using both synthetic and real-world datasets, we demonstrate that our method outperforms state-of-the-art models when predicting the target function with far fewer features per instance.


Microsoft announces Xbox streaming stick and a TV app for xCloud gaming

The Independent - Tech

Xbox has quietly announced that it will be making streaming sticks that would allow gamers to play "on any TV or monitor". The news comes as part of an update the video game giant, owned by Microsoft, published ahead of the E3 conference. Streaming sticks may not be the only way that Xbox is looking to get more games into the hands of players, as it also says that it is "working with global TV manufacturers to embed the Xbox experience directly into internet-connected televisions with no extra hardware required except a controller." Microsoft's head of Xbox Phil Spencer had previously said that a dedicated app for the games console could arrive on Smart TVs and stream games directly to them, and it appears like it is now becoming a reality. A number of changes to Xbox Game Pass were also launched alongside this news: the company is working with telecommunication providers on new purchasing models like Xbox All Access, encouraging payments over time rather than money up-front, is rolling Game Pass Ultimate out in Australia, Brazil, Mexico, and Japan, and is adding cloud gaming support directly into the Xbox app on PC.


Microsoft is officially making Xbox video game streaming sticks

Engadget

Just days before Microsoft's big ol' E3 livestream, executives from the company sat down to talk -- er, read remarks prepared by the communications team -- about the future of Xbox. In a pre-recorded media briefing, Xbox head Phil Spencer, Microsoft CEO Satya Nadella and others bragged about how well Game Pass and Azure are performing, and also dropped some news about the company's cloud gaming and subscription strategies. First, Xbox is working with global TV manufacturers to get Game Pass on smart televisions. Considering a Game Pass Ultimate subscription unlocks cloud capabilities, this feature will allow folks to play Xbox titles with just a controller, no console required. Additionally, Microsoft is officially building a video game streaming stick, as Spencer teased late last year.


Building the engine that drives digital transformation

MIT Technology Review

This is the consensus view of an MIT Technology Review Insights survey of 210 members of technology executives, conducted in March 2021. These respondents report that they need--and still often lack-- the ability to develop new digital channels and services quickly, and to optimize them in real time. Underpinning these waves of digital transformation are two fundamental drivers: the ability to serve and understand customers better, and the need to increase employees' ability to work more effectively toward those goals. Two-thirds of respondents indicated that more efficient customer experience delivery was the most critical objective. This was followed closely by the use of analytics and insight to improve products and services (60%).