Government
Submodular Information Selection for Hypothesis Testing with Misclassification Penalties
Bhargav, Jayanth, Ghasemi, Mahsa, Sundaram, Shreyas
We consider the problem of selecting an optimal subset of information sources for a hypothesis testing/classification task where the goal is to identify the true state of the world from a finite set of hypotheses, based on finite observation samples from the sources. In order to characterize the learning performance, we propose a misclassification penalty framework, which enables nonuniform treatment of different misclassification errors. In a centralized Bayesian learning setting, we study two variants of the subset selection problem: (i) selecting a minimum cost information set to ensure that the maximum penalty of misclassifying the true hypothesis is below a desired bound and (ii) selecting an optimal information set under a limited budget to minimize the maximum penalty of misclassifying the true hypothesis. Under certain assumptions, we prove that the objective (or constraints) of these combinatorial optimization problems are weak (or approximate) submodular, and establish high-probability performance guarantees for greedy algorithms. Further, we propose an alternate metric for information set selection which is based on the total penalty of misclassification. We prove that this metric is submodular and establish near-optimal guarantees for the greedy algorithms for both the information set selection problems. Finally, we present numerical simulations to validate our theoretical results over several randomly generated instances.
Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation
Sanyal, Amartya, Hu, Yaxi, Yu, Yaodong, Ma, Yian, Wang, Yixin, Schรถlkopf, Bernhard
"Accuracy-on-the-line" is a widely observed phenomenon in machine learning, where a model's accuracy on in-distribution (ID) and out-of-distribution (OOD) data is positively correlated across different hyperparameters and data configurations. But when does this useful relationship break down? In this work, we explore its robustness. The key observation is that noisy data and the presence of nuisance features can be sufficient to shatter the Accuracy-on-the-line phenomenon. In these cases, ID and OOD accuracy can become negatively correlated, leading to "Accuracy-on-the-wrong-line". This phenomenon can also occur in the presence of spurious (shortcut) features, which tend to overshadow the more complex signal (core, non-spurious) features, resulting in a large nuisance feature space. Moreover, scaling to larger datasets does not mitigate this undesirable behavior and may even exacerbate it. We formally prove a lower bound on Out-of-distribution (OOD) error in a linear classification model, characterizing the conditions on the noise and nuisance features for a large OOD error. We finally demonstrate this phenomenon across both synthetic and real datasets with noisy data and nuisance features.
A Russian Propaganda Network Is Promoting an AI-Manipulated Biden Video
In recent weeks, as so-called cheapfake video clips suggesting President Joe Biden is unfit for office have gone viral on social media, a Kremlin-affiliated disinformation network has been promoting a parody music video featuring Biden wearing a diaper and being pushed around in a wheelchair. The video is called "Bye, Bye Biden" and has been viewed more than 5 million times on X since it was first promoted in the middle of May. It depicts Biden as senile, wearing a hearing aid, and taking a lot of medication. It also shows him giving money to a character who seems to represent illegal migrants while denying money to US citizens until they change their costume to mimic the Ukrainian flag. Another scene shows Biden opening the front door of a family home that features a Confederate flag on the wall and allowing migrants to come in and take over. Finally, the video contains references to stolen election conspiracies pushed by former president Donald Trump.
Silicon Valley wants unfettered control of the tech market. That's why it's cosying up to Trump Evgeny Morozov
Hardly a week passes without another billionaire endorsing Donald Trump. With Joe Biden proposing a 25% tax on those with assets over 100m ( 80m), this is no shock. The pro-Trump multimillionaire club now includes a growing number of venture capitalists. Unlike hedge funders or private equity barons, venture capitalists have traditionally held progressive credentials. They've styled themselves as the heroes of innovation, and the Democrats have done more to polish their progressive image than anyone else.
Why China's dominance in commercial drones has become a global security matter
But on June 14, the US House of Representatives passed a bill that would completely ban DJI's drones from being sold in the US. The bill is now being discussed in the Senate as part of the annual defense budget negotiations. While its market dominance has attracted scrutiny for years, it's increasingly clear that DJI's commercial products are so good and affordable they are also being used on active battlefields to scout out the enemy or carry bombs. As the US worries about the potential for conflict between China and Taiwan, the military implications of DJI's commercial drones are becoming a top policy concern. DJI has managed to set the gold standard for commercial drones because it is built on decades of electronic manufacturing prowess and policy support in Shenzhen.
How OpenAI's Decision Not to Operate in China Will Reshape the Chinese AI Scene
OpenAI's abrupt move to ban access to its services in China is setting the scene for an industry shakeup, as local AI leaders from Baidu Inc. to Alibaba Group Holding Ltd. move to grab more of the field. The ChatGPT creator this week sent memos to Chinese users warning it will cut off access to its widely used AI development software and tools from July, triggering a scramble to fill the void. Since Tuesday, at least a half-dozen companies and startups including Tencent Holdings Ltd. and Zhipu AI began offering incentives to developers making the switch. OpenAI's shift will accentuate the divide between China and the U.S., which is trying to curb Beijing's AI and chip efforts. While the startup's exit offers an opportunity for sector leaders to grow their user base, it also deprives entrepreneurs and cash-strapped startups of some of the best tools available to fine-tune or get their AI applications off the ground.
UK needs system for recording AI misuse and malfunctions, thinktank says
The UK needs a system for recording misuse and malfunctions in artificial intelligence or ministers risk being unaware of alarming incidents involving the technology, according to a report. The next government should create a system for logging incidents involving AI in public services and should consider building a central hub for collating AI-related episodes across the UK, said the Centre for Long-Term Resilience (CLTR), a thinktank. CLTR, which focuses on government responses to unforeseen crises and extreme risks, said an incident reporting regime such as the system operated by the Air Accidents Investigation Branch (AAIB) was vital for using the technology successfully. The report cites 10,000 AI "safety incidents" recorded by news outlets since 2014, listed in a database compiled by the Organisation for Economic Co-operation and Development, an international research body. Examples logged on the OECD's AI safety incident monitor include a deepfake of the Labour leader, Keir Starmer, purportedly being abusive to party staff, Google's Gemini model portraying German second world war soldiers as people of colour, incidents involving self-driving cars and a man who planned to assassinate the late queen drawing encouragement from a chatbot.
India exports rockets, explosives to Israel amid Gaza war, documents reveal
In the early morning hours of May 15, the cargo vessel Borkum stopped off the Spanish coast, lingering in the waters a short distance from Cartagena. At the port, protesters waved Palestinian flags and called on authorities to inspect the ship based on suspicions that it carried weapons bound for Israel. Leftist members of the European Parliament sent a letter to Spanish President Pedro Sรกnchez requesting that the ship be prevented from docking. "Allowing a ship loaded with weapons destined for Israel is to allow the transit of arms to a country currently under investigation for genocide against the Palestinian people," the group of nine MEPs warned. Before the Spanish government could take a stand, the Borkum cancelled its planned stopover and continued to the Slovenian port of Koper.
WV-Net: A foundation model for SAR WV-mode satellite imagery trained using contrastive self-supervised learning on 10 million images
Glaser, Yannik, Stopa, Justin E., Wolniewicz, Linnea M., Foster, Ralph, Vandemark, Doug, Mouche, Alexis, Chapron, Bertrand, Sadowski, Peter
The European Space Agency's Copernicus Sentinel-1 (S-1) mission is a constellation of C-band synthetic aperture radar (SAR) satellites that provide unprecedented monitoring of the world's oceans. S-1's wave mode (WV) captures 20x20 km image patches at 5 m pixel resolution and is unaffected by cloud cover or time-of-day. The mission's open data policy has made SAR data easily accessible for a range of applications, but the need for manual image annotations is a bottleneck that hinders the use of machine learning methods. This study uses nearly 10 million WV-mode images and contrastive self-supervised learning to train a semantic embedding model called WV-Net. In multiple downstream tasks, WV-Net outperforms a comparable model that was pre-trained on natural images (ImageNet) with supervised learning. Experiments show improvements for estimating wave height (0.50 vs 0.60 RMSE using linear probing), estimating near-surface air temperature (0.90 vs 0.97 RMSE), and performing multilabel-classification of geophysical and atmospheric phenomena (0.96 vs 0.95 micro-averaged AUROC). WV-Net embeddings are also superior in an unsupervised image-retrieval task and scale better in data-sparse settings. Together, these results demonstrate that WV-Net embeddings can support geophysical research by providing a convenient foundation model for a variety of data analysis and exploration tasks.
Generative Discrimination: What Happens When Generative AI Exhibits Bias, and What Can Be Done About It
Hacker, Philipp, Mittelstadt, Brent, Borgesius, Frederik Zuiderveen, Wachter, Sandra
As generative Artificial Intelligence (genAI) technologies proliferate across sectors, they offer significant benefits but also risk exacerbating discrimination. This chapter explores how genAI intersects with non-discrimination laws, identifying shortcomings and suggesting improvements. It highlights two main types of discriminatory outputs: (i) demeaning and abusive content and (ii) subtler biases due to inadequate representation of protected groups, which may not be overtly discriminatory in individual cases but have cumulative discriminatory effects. For example, genAI systems may predominantly depict white men when asked for images of people in important jobs. This chapter examines these issues, categorizing problematic outputs into three legal categories: discriminatory content; harassment; and legally hard cases like unbalanced content, harmful stereotypes or misclassification. It argues for holding genAI providers and deployers liable for discriminatory outputs and highlights the inadequacy of traditional legal frameworks to address genAI-specific issues. The chapter suggests updating EU laws, including the AI Act, to mitigate biases in training and input data, mandating testing and auditing, and evolving legislation to enforce standards for bias mitigation and inclusivity as technology advances.