Government
Ethical Considerations for Responsible Data Curation
Andrews, Jerone T. A., Zhao, Dora, Thong, William, Modas, Apostolos, Papakyriakopoulos, Orestis, Xiang, Alice
Human-centric computer vision (HCCV) data curation practices often neglect privacy and bias concerns, leading to dataset retractions and unfair models. HCCV datasets constructed through nonconsensual web scraping lack crucial metadata for comprehensive fairness and robustness evaluations. Current remedies are post hoc, lack persuasive justification for adoption, or fail to provide proper contextualization for appropriate application. Our research focuses on proactive, domain-specific recommendations, covering purpose, privacy and consent, and diversity, for curating HCCV evaluation datasets, addressing privacy and bias concerns. We adopt an ante hoc reflective perspective, drawing from current practices, guidelines, dataset withdrawals, and audits, to inform our considerations and recommendations.
Impact of Adversarial Training on Robustness and Generalizability of Language Models
Altinisik, Enes, Sajjad, Hassan, Sencar, Husrev Taha, Messaoud, Safa, Chawla, Sanjay
Adversarial training is widely acknowledged as the most effective defense against adversarial attacks. However, it is also well established that achieving both robustness and generalization in adversarially trained models involves a trade-off. The goal of this work is to provide an in depth comparison of different approaches for adversarial training in language models. Specifically, we study the effect of pre-training data augmentation as well as training time input perturbations vs. embedding space perturbations on the robustness and generalization of transformer-based language models. Our findings suggest that better robustness can be achieved by pre-training data augmentation or by training with input space perturbation. However, training with embedding space perturbation significantly improves generalization. A linguistic correlation analysis of neurons of the learned models reveals that the improved generalization is due to 'more specialized' neurons. To the best of our knowledge, this is the first work to carry out a deep qualitative analysis of different methods of generating adversarial examples in adversarial training of language models.
Reinforcement Learning in Non-Markovian Environments
Chandak, Siddharth, Shah, Pratik, Borkar, Vivek S, Dodhia, Parth
Motivated by the novel paradigm developed by Van Roy and coauthors for reinforcement learning in arbitrary non-Markovian environments, we propose a related formulation and explicitly pin down the error caused by non-Markovianity of observations when the Q-learning algorithm is applied on this formulation. Based on this observation, we propose that the criterion for agent design should be to seek good approximations for certain conditional laws. Inspired by classical stochastic control, we show that our problem reduces to that of recursive computation of approximate sufficient statistics. This leads to an autoencoder-based scheme for agent design which is then numerically tested on partially observed reinforcement learning environments.
ChatGPT exploded into public life a year ago. Now we know what went on behind the scenes John Naughton
If a week is a long time in politics, a year is an eternity in tech. Just over 12 months ago, the industry was humming along in its usual way. The big platforms were deep into what Cory Doctorow calls "enshittification" โ the process in which platforms go from being initially good to their users, to abusing them to make things better for their business customers and finally to abusing those customers in order to claw back all the value for themselves. Elon Musk was ramping up his efforts to alienate advertisers on Twitter/X and accelerate the death spiral of his expensive toy. TikTok was monopolising every waking hour of teenagers.
Did Israel's overreliance on tech cause October 7 intelligence failure?
An overreliance on technology by Israel's intelligence agencies and military has continued to shape the current conflict in Gaza, analysts say, while also being partially responsible for the failure to detect the Hamas attack on October 7. Hamas's surprise attack on army outposts and surrounding villages in southern Israel, which resulted in the deaths of 1,200 Israeli and foreign nationals, mostly civilians, took the Israeli intelligence agencies by surprise. Hamas fighters also took about 240 people captive. Israel, in its brutal military response, has killed more than 17,000 Palestinians in Gaza since then. Within both Israel and the wider Arab region, many have asked how Shin Bet, one of the world's most respected and feared intelligence agencies, which is responsible for Israel's domestic security, could have been outmatched by Hamas using bulldozers and paragliders. The world's disbelief has sparked a bounty of conspiracy theories in some quarters.
The Gospel: Israel turns to a new AI system in the Gaza war
More than 60 days into the Israel-Gaza war, two Israeli news outlets โ 972 magazine and Local Call โ published a report on The Gospel, a new artificial intelligence system deployed in Gaza. The AI helps generate new targets at an unprecedented rate, allowing the Israeli military to loosen its already permissive constraints on the killing of civilians. The exchange of hostages between Israel and Hamas late last month created some challenges for the Netanyahu government โ and its messaging. Producer Meenakshi Ravi looks at how Israeli media has been reporting on the story. As the world is focused on the events unfolding in Gaza, Israel has also escalated its attacks on Palestinians in the occupied West Bank, where Hamas has no authority or military presence.
European Union reaches agreement on landmark legislation to regulate AI
European Union policymakers have agreed on landmark legislation to regulate artificial intelligence (AI), paving the way for the most ambitious set of standards yet to control the use of the game-changing technology. The agreement to support the "AI Act" on Friday came after nearly 38 hours of negotiations between lawmakers and policymakers. "The AI Act is a global first. A unique legal framework for the development of AI you can trust," EU chief Ursula von der Leyen said. A commitment we took in our political guidelines โ and we delivered.
EU agrees 'historic' deal with world's first laws to regulate AI
The world's first comprehensive laws to regulate artificial intelligence have been agreed in a landmark deal after a marathon 37-hour negotiation between the European Parliament and EU member states. The agreement was described as "historic" by Thierry Breton, the European Commissioner responsible for a suite of laws in Europe that will also govern social media and search engines, covering giants such as X, TikTok and Google. Breton said 100 people had been in a room for almost three days to seal the deal. He said it was "worth the few hours of sleep" to make the "historic" deal. Carme Artigas, Spain's secretary of state for AI, who facilitated the negotiations, said France and Germany supported the text, amid reports that tech companies in those countries were fighting for a lighter touch approach to foster innovation among small companies.
EU strikes deal to regulate AI tech in landmark act
The European Union reached a hard-fought deal on what is poised to become the most comprehensive regulation of artificial intelligence in the Western world. Thierry Breton, the bloc's internal market chief, said the deal strikes a balance between fostering innovation and protecting the rights of people and companies. "We spent a lot of time on finding the right balance between making the most of AI potential to support law enforcement while protecting our citizens' fundamental rights," he said early Saturday in a statement. "We do not want any mass surveillance in Europe."
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
Yuan, Xin, Guo, Jie, Qiu, Weidong, Huang, Zheng, Li, Shujun
Mis- and disinformation online have become a major societal problem as major sources of online harms of different kinds. One common form of mis- and disinformation is out-of-context (OOC) information, where different pieces of information are falsely associated, e.g., a real image combined with a false textual caption or a misleading textual description. Although some past studies have attempted to defend against OOC mis- and disinformation through external evidence, they tend to disregard the role of different pieces of evidence with different stances. Motivated by the intuition that the stance of evidence represents a bias towards different detection results, we propose a stance extraction network (SEN) that can extract the stances of different pieces of multi-modal evidence in a unified framework. Moreover, we introduce a support-refutation score calculated based on the co-occurrence relations of named entities into the textual SEN. Extensive experiments on a public large-scale dataset demonstrated that our proposed method outperformed the state-of-the-art baselines, with the best model achieving a performance gain of 3.2% in accuracy.