Goto

Collaborating Authors

 Government


Asymptotic theory of in-context learning by linear attention

arXiv.org Machine Learning

Transformers have a remarkable ability to learn and execute tasks based on examples provided within the input itself, without explicit prior training. It has been argued that this capability, known as in-context learning (ICL), is a cornerstone of Transformers' success, yet questions about the necessary sample complexity, pretraining task diversity, and context length for successful ICL remain unresolved. Here, we provide a precise answer to these questions in an exactly solvable model of ICL of a linear regression task by linear attention. We derive sharp asymptotics for the learning curve in a phenomenologically-rich scaling regime where the token dimension is taken to infinity; the context length and pretraining task diversity scale proportionally with the token dimension; and the number of pretraining examples scales quadratically. We demonstrate a double-descent learning curve with increasing pretraining examples, and uncover a phase transition in the model's behavior between low and high task diversity regimes: In the low diversity regime, the model tends toward memorization of training tasks, whereas in the high diversity regime, it achieves genuine in-context learning and generalization beyond the scope of pretrained tasks. These theoretical insights are empirically validated through experiments with both linear attention and full nonlinear Transformer architectures.


Transfer Learning for Spatial Autoregressive Models

arXiv.org Machine Learning

The spatial autoregressive (SAR) model has been widely applied in various empirical economic studies to characterize the spatial dependence among subjects. However, the precision of estimating the SAR model diminishes when the sample size of the target data is limited. In this paper, we propose a new transfer learning framework for the SAR model to borrow the information from similar source data to improve both estimation and prediction. When the informative source data sets are known, we introduce a two-stage algorithm, including a transferring stage and a debiasing stage, to estimate the unknown parameters and also establish the theoretical convergence rates for the resulting estimators. If we do not know which sources to transfer, a transferable source detection algorithm is proposed to detect informative sources data based on spatial residual bootstrap to retain the necessary spatial dependence. Its detection consistency is also derived. Simulation studies demonstrate that using informative source data, our transfer learning algorithm significantly enhances the performance of the classical two-stage least squares estimator. In the empirical application, we apply our method to the election prediction in swing states in the 2020 U.S. presidential election, utilizing polling data from the 2016 U.S. presidential election along with other demographic and geographical data. The empirical results show that our method outperforms traditional estimation methods.


Are seed-sowing drones the answer to global deforestation?

Al Jazeera

Santa Cruz Cabralia, Bahia, Brazil โ€“ With a loud whir, the drone takes flight. Minutes later, the humming sound gives way to a distinctive rattling as the machine, hovering about 20 metres above the ground, begins unloading its precious cargo and a cocktail of seeds rains down onto the land below. Given time, these seeds will grow into trees and, eventually, it is hoped, a thriving forest will stand where there was once just sparse vegetation. That is what the startup which operates this drone, a large contraption that looks a bit like a Pokemon ball with antennae, hopes. The 54 hectares (133 acres) here which have been badly degraded by agriculture and cattle farming in the Brazilian state of Bahia are just the start.


Fox News AI Newsletter: How artificial intelligence is reshaping modern warfare

FOX News

NEXT-GEN BATTLE: Modern warfare is changing rapidly, and harnessing artificial intelligence is key to staying ahead of America's adversaries. Modern warfare is rapidly changing -- and artificial intelligence may only speed up that process. FUNNY BOT: A team of university researchers in the Netherlands says they've developed an artificial intelligence (AI) platform that can recognize sarcasm, according to a new report. AI (artificial intelligence) letters are placed on a computer motherboard in this illustration taken on June 23, 2023. 'OUTCOMPETE CHINA': A bipartisan group of U.S. senators on Wednesday joined in a call to boost American funding of artificial intelligence research.


US Official Warns a Cell Network Flaw Is Being Exploited for Spying

WIRED

Laser warfare, among all the long-unfulfilled imaginings of science fiction writers, is right up there with flying cars. After decades of research, the US military is actively deploying laser defense systems in the Middle East to shoot down drones launched by adversaries like Yemen's Houthi rebels, one of several recent deployments of laser tech in actual combat situations. In less pew-pew-oriented security news, the debate continues over the extension of Section 702 of the Foreign Intelligence Surveillance Act, signed by President Biden last month, as 20 civil liberties organizations sent a letter to the Justice Department demanding more clarity on when the NSA can demand US tech companies cooperate in its wiretaps. Elsewhere, WIRED obtained emails showing how New York City decided to deploy a gun-detection system called Evolv in subways despite false-positive rates as high as 85 percent. At the Google I/O developer conference, meanwhile, the search giant debuted a new AI-based feature in Android that's designed to detect if a phone has been stolen and automatically lock it down.


How China is using AI news anchors to deliver its propaganda

The Guardian

The news presenter has a deeply uncanny air as he delivers a partisan and pejorative message in Mandarin: Taiwan's outgoing president, Tsai Ing-wen, is as effective as limp spinach, her period in office beset by economic under performance, social problems and protests. "Water spinach looks at water spinach. Turns out that water spinach isn't just a name," says the presenter, in an extended metaphor about Tsai being "Hollow Tsai" โ€“ a pun related to the Mandarin word for water spinach. This is not a conventional broadcast journalist, even if the lack of impartiality is no longer a shock. The anchor is generated by an artificial intelligence programme, and the segment is trying, albeit clumsily, to influence the Taiwanese presidential election. The source and creator of the video are unknown, but the clip is designed to make voters doubt politicians who want Taiwan to remain at arm's length from China, which claims that the self-governing island is part of its territory.


On Robust Reinforcement Learning with Lipschitz-Bounded Policy Networks

arXiv.org Artificial Intelligence

This paper presents a study of robust policy networks in deep reinforcement learning. We investigate the benefits of policy parameterizations that naturally satisfy constraints on their Lipschitz bound, analyzing their empirical performance and robustness on two representative problems: pendulum swing-up and Atari Pong. We illustrate that policy networks with small Lipschitz bounds are significantly more robust to disturbances, random noise, and targeted adversarial attacks than unconstrained policies composed of vanilla multi-layer perceptrons or convolutional neural networks. Moreover, we find that choosing a policy parameterization with a non-conservative Lipschitz bound and an expressive, nonlinear layer architecture gives the user much finer control over the performance-robustness trade-off than existing state-of-the-art methods based on spectral normalization.


Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses

arXiv.org Artificial Intelligence

Offline reinforcement learning (RL) addresses the challenge of expensive and high-risk data exploration inherent in RL by pre-training policies on vast amounts of offline data, enabling direct deployment or fine-tuning in real-world environments. However, this training paradigm can compromise policy robustness, leading to degraded performance in practical conditions due to observation perturbations or intentional attacks. While adversarial attacks and defenses have been extensively studied in deep learning, their application in offline RL is limited. This paper proposes a framework to enhance the robustness of offline RL models by leveraging advanced adversarial attacks and defenses. The framework attacks the actor and critic components by perturbing observations during training and using adversarial defenses as regularization to enhance the learned policy. Four attacks and two defenses are introduced and evaluated on the D4RL benchmark. The results show the vulnerability of both the actor and critic to attacks and the effectiveness of the defenses in improving policy robustness. This framework holds promise for enhancing the reliability of offline RL models in practical scenarios.


Real-Time Go-Around Prediction: A case study of JFK airport

arXiv.org Artificial Intelligence

In this paper, we employ the long-short-term memory model (LSTM) to predict the real-time go-around probability as an arrival flight is approaching JFK airport and within 10 nm of the landing runway threshold. We further develop methods to examine the causes to go-around occurrences both from a global view and an individual flight perspective. According to our results, in-trail spacing, and simultaneous runway operation appear to be the top factors that contribute to overall go-around occurrences. We then integrate these pre-trained models and analyses with real-time data streaming, and finally develop a demo web-based user interface that integrates the different components designed previously into a real-time tool that can eventually be used by flight crews and other line personnel to identify situations in which there is a high risk of a go-around.


MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing

arXiv.org Artificial Intelligence

Memes have evolved as a prevalent medium for diverse communication, ranging from humour to propaganda. With the rising popularity of image-focused content, there is a growing need to explore its potential harm from different aspects. Previous studies have analyzed memes in closed settings - detecting harm, applying semantic labels, and offering natural language explanations. To extend this research, we introduce MemeMQA, a multimodal question-answering framework aiming to solicit accurate responses to structured questions while providing coherent explanations. We curate MemeMQACorpus, a new dataset featuring 1,880 questions related to 1,122 memes with corresponding answer-explanation pairs. We further propose ARSENAL, a novel two-stage multimodal framework that leverages the reasoning capabilities of LLMs to address MemeMQA. We benchmark MemeMQA using competitive baselines and demonstrate its superiority - ~18% enhanced answer prediction accuracy and distinct text generation lead across various metrics measuring lexical and semantic alignment over the best baseline. We analyze ARSENAL's robustness through diversification of question-set, confounder-based evaluation regarding MemeMQA's generalizability, and modality-specific assessment, enhancing our understanding of meme interpretation in the multimodal communication landscape.