Goto

Collaborating Authors

 Government


Can we still protect our data in the artificial intelligence era? - VoxEurop

#artificialintelligence

Donald Trump has won the United States presidency and Brexit has promised to take the UK out of the European Union. Both campaigns employ Cambridge Analytica, which harvested the data of millions of Facebook users to personalise electoral messaging to them and sway their voting intentions. Millions of people begin to ask themselves whether, in the digital era, they have lost something deeply valuable: their privacy. Two years later, countless European email inboxes would be filling up with messages from companies, asking people for permission to continue processing their data โ€“ the aim was compliance with the new General Data Protection Regulation (GDPR). Despite its imperfections, this law has served as a point of reference for laws in Brazil and Japan, and inaugurated the modern era of data protection.


What Happens When AI Fighter Pilots Take to the Skies?

#artificialintelligence

In 2022, the pilot of an F-16 fighter jet will jink hard to the right and flick over into a roll, struggling to evade the plane behind them. Years of training and experience will suddenly become redundant. The AI algorithm controlling the chasing plane will have changed the face of war forever. AI first demonstrated the sorts of aerobatic skills needed for dogfighting back in 2008. Andrew Ng's team at Stanford University developed an AI-piloted helicopter that learned how to perform stunts simply by watching human pilots.


Evaluating the Construct Validity of Text Embeddings with Application to Survey Questions

arXiv.org Artificial Intelligence

Text embedding models from Natural Language Processing can map text data (e.g. words, sentences, documents) to supposedly meaningful numerical representations (a.k.a. text embeddings). While such models are increasingly applied in social science research, one important issue is often not addressed: the extent to which these embeddings are valid representations of constructs relevant for social science research. We therefore propose the use of the classic construct validity framework to evaluate the validity of text embeddings. We show how this framework can be adapted to the opaque and high-dimensional nature of text embeddings, with application to survey questions. We include several popular text embedding methods (e.g. fastText, GloVe, BERT, Sentence-BERT, Universal Sentence Encoder) in our construct validity analyses. We find evidence of convergent and discriminant validity in some cases. We also show that embeddings can be used to predict respondent's answers to completely new survey questions. Furthermore, BERT-based embedding techniques and the Universal Sentence Encoder provide more valid representations of survey questions than do others. Our results thus highlight the necessity to examine the construct validity of text embeddings before deploying them in social science research.


Interpolation and Regularization for Causal Learning

arXiv.org Machine Learning

We study the problem of learning causal models from observational data through the lens of interpolation and its counterpart -- regularization. A large volume of recent theoretical, as well as empirical work, suggests that, in highly complex model classes, interpolating estimators can have good statistical generalization properties and can even be optimal for statistical learning. Motivated by an analogy between statistical and causal learning recently highlighted by Janzing (2019), we investigate whether interpolating estimators can also learn good causal models. To this end, we consider a simple linearly confounded model and derive precise asymptotics for the *causal risk* of the min-norm interpolator and ridge-regularized regressors in the high-dimensional regime. Under the principle of independent causal mechanisms, a standard assumption in causal learning, we find that interpolators cannot be optimal and causal learning requires stronger regularization than statistical learning. This resolves a recent conjecture in Janzing (2019). Beyond this assumption, we find a larger range of behavior that can be precisely characterized with a new measure of *confounding strength*. If the confounding strength is negative, causal learning requires weaker regularization than statistical learning, interpolators can be optimal, and the optimal regularization can even be negative. If the confounding strength is large, the optimal regularization is infinite, and learning from observational data is actively harmful.


A new LDA formulation with covariates

arXiv.org Machine Learning

The Latent Dirichlet Allocation (LDA) model is a popular method for creating mixed-membership clusters. Despite having been originally developed for text analysis, LDA has been used for a wide range of other applications. We propose a new formulation for the LDA model which incorporates covariates. In this model, a negative binomial regression is embedded within LDA, enabling straight-forward interpretation of the regression coefficients and the analysis of the quantity of cluster-specific elements in each sampling units (instead of the analysis being focused on modeling the proportion of each cluster, as in Structural Topic Models). We use slice sampling within a Gibbs sampling algorithm to estimate model parameters. We rely on simulations to show how our algorithm is able to successfully retrieve the true parameter values and the ability to make predictions for the abundance matrix using the information given by the covariates. The model is illustrated using real data sets from three different areas: text-mining of Coronavirus articles, analysis of grocery shopping baskets, and ecology of tree species on Barro Colorado Island (Panama). This model allows the identification of mixed-membership clusters in discrete data and provides inference on the relationship between covariates and the abundance of these clusters.


Fine-grained Prediction of Political Leaning on Social Media with Unsupervised Deep Learning

Journal of Artificial Intelligence Research

Predicting the political leaning of social media users is an increasingly popular task, given its usefulness for electoral forecasts, opinion dynamics models and for studying the political dimension of polarization and disinformation. Here, we propose a novel unsupervised technique for learning fine-grained political leaning from the textual content of social media posts. Our technique leverages a deep neural network for learning latent political ideologies in a representation learning task. Then, users are projected in a low-dimensional ideology space where they are subsequently clustered. The political leaning of a user is automatically derived from the cluster to which the user is assigned. We evaluated our technique in two challenging classification tasks and we compared it to baselines and other state-of-the-art approaches. Our technique obtains the best results among all unsupervised techniques, with micro F1 = 0.426 in the 8-class task and micro F1 = 0.772 in the 3-class task. Other than being interesting on their own, our results also pave the way for the development of new and better unsupervised approaches for the detection of fine-grained political leaning.


Artificial Intelligence: Status of Developing and Acquiring Capabilities for Weapon Systems

#artificialintelligence

The Department of Defense (DOD) is actively pursuing artificial intelligence (AI) capabilities. AI refers to computer systems designed to replicate a range of human functions and continually get better at their assigned tasks. GAO previously identified three waves or types of AI, shown below. DOD recognizes that developing and using AI differs from traditional software. Traditional software is programmed to perform tasks based on static instructions, whereas AI is programmed to learn to improve at its given tasks.


Ex-Google CEO slams 'dithering' on 5G and claims US is 'well behind' China's progress

The Independent - Tech

Former Google CEO Eric Schmidt has slammed the US government for its slow 5G rollout, arguing that the government's "dithering" has left America "well behind" China. Dr Schmidt penned an op-ed in the Wall Street Journal alongside Harvard government professor Graham Allison, saying that the US is "far behind in almost every dimension of 5G while other nations โ€“ including China โ€“ race ahead". The authors said the Biden administration must make 5G a "national priority". If not, "China will own the 5G future", they said. "The step up to real 5G speeds will lead to analogous breakthroughs in autonomous vehicles, virtual-reality applications like the metaverse, and other areas that have yet to be invented," Dr Schmidt and Dr Allison wrote.


Where AI Falls Down in Cybersecurity

#artificialintelligence

Artificial intelligence (AI) burst onto the scene like a superhero trying out his first cape, accompanied by a loud and sustained PR ruckus. But the cybersecurity crowd's oohs and ahs soon began to fade into ums and uhs as the scoreboard ticked off more fails than wins for AI. Don't expect any sympathy from the spectator seats. Industry observers are pulling back their cheers for the beleaguered tech star. For example, Gartner describes the state of cyber AI as immature and advises security analysts to "treat AI offerings as experimental, complementary controls."


Elon Musk-owned Neuralink confirms monkeys died during tests but rejects abuse claim

Daily Mail - Science & tech

Elon Musk's brain-chip firm Neuralink has admitted monkeys died during tests, but denied claims of animal abuse put forward by an animal rights group. The biotech firm is developing a brain-computer interface, that it claims could one day make humans hyper-intelligent, and allow paralyzed people to walk again. Last week the Physicians Committee for Responsible Medicine (PCRM) lodged a complaint with the US Department of Agriculture, alleging several counts of animal abuse between 2017 and 2020, involving test monkeys owned by Neuralink. They claimed the macaque monkeys, housed at a University of California Davis research facility, were subject to experiments that amounted to torture, with evidence of rashes, self-mutilation and brain hemorrhages seen in documentation. Neuralink has hit back at the claims of abuse, calling out the PCRM as a group that oppose any use of animals in research.