Goto

Collaborating Authors

 Government


Taiwan's Digital Minister Has an Ambitious Plan to Align Tech With Democracy

TIME - Tech

Audrey Tang, Taiwan's 43-year-old minister of digital affairs, has a powerful effect on people. At a panel discussion at Northeastern University in Boston, 20-year-old student Diane Grant is visibly moved, describing Tang's talk as the best she's been to in her undergraduate career. Later that day, a German tourist recognizes Tang leaving the Boston Museum of Science and requests a photo, saying she's "starstruck." At the Massachusetts Institute of Technology, a trio of world-leading economists bashfully ask Tang to don a baseball cap emblazoned with the name of their research center and pose for a group photo. Political scientist and former gubernatorial candidate Danielle Allen, confesses to Tang that, although others often tell her that she is a source of inspiration to them, she rarely feels inspired by others.


FDA approves Neuralink's brain chip for second patient - after first person suffered life-threatening condition during surgery

Daily Mail - Science & tech

Elon Musk's Neuralink has been given a green light to implant its brain chip in a second patient after fixing issues that struck during the first human trial. The US Food and Drug Administration (FDA) approved the next person on Monday, signing off on the company's planned updates that included embedding some of the device's ultrathin wires deeper into the brain. Neuralink revealed this month that some of 64 threads detached from the first patient's brain, causing the chip to malfunction - nearly ending the trial that began in January. A report by Reuters cited'five people familiar with the matter' had claimed that this issue had been'known about for years' from animal testing. This is a developing story... more updates to come.


Indian Voters Are Being Bombarded With Millions of Deepfakes. Political Candidates Approve

WIRED

On a stifling April afternoon in Ajmer, in the Indian state of Rajasthan, local politician Shakti Singh Rathore sat down in front of a greenscreen to shoot a short video. It was his first time being cloned. Wearing a crisp white shirt and a ceremonial saffron scarf bearing a lotus flower--the logo of the BJP, the country's ruling party--Rathore pressed his palms together and greeted his audience in Hindi. Before he could continue, the director of the shoot walked into the frame. Divyendra Singh Jadoun, a 31-year-old with a bald head and a thick black beard, told Rathore he was moving around too much on camera.


Russia-Ukraine war: List of key events, day 816

Al Jazeera

At least 11 people were killed and dozens injured after Russia bombed a busy lakeside resort on the edge of Ukraine's second-largest city of Kharkiv and attacked villages in the surrounding area. At least 13 people were injured after the Ukrainian military shelled areas of Russia's southern Belgorod region, according to Belgorod's regional Governor Vyacheslav Gladkov. The General Staff of the Armed Forces of Ukraine said Russian attacks in the Kharkiv area "slowed down a bit" but that forces "continue their attempts to break through our defences near Vovchansk, Starytsya and Lyptsi". Russia's Ministry of Defence, which claimed earlier to have seized Starytsya, said its units "continued to advance into the depth of the enemy's defences". Officials said Russia shot down at least 103 Ukrainian drones, including 62 over Russian regions, as well as missiles that targeted Crimea, which Moscow seized and annexed from Ukraine in 2014.


Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models

arXiv.org Artificial Intelligence

Large language models (LLMs) have achieved remarkable performance on a wide range of tasks. However, recent studies have shown that LLMs can memorize training data and simple repeated tokens can trick the model to leak the data. In this paper, we take a step further and show that certain special characters or their combinations with English letters are stronger memory triggers, leading to more severe data leakage. The intuition is that, since LLMs are trained with massive data that contains a substantial amount of special characters (e.g. structural symbols {, } of JSON files, and @, # in emails and online posts), the model may memorize the co-occurrence between these special characters and the raw texts. This motivates us to propose a simple but effective Special Characters Attack (SCA) to induce training data leakage. Our experiments verify the high effectiveness of SCA against state-of-the-art LLMs: they can leak diverse training data, such as code corpus, web pages, and personally identifiable information, and sometimes generate non-stop outputs as a byproduct. We further show that the composition of the training data corpus can be revealed by inspecting the leaked data -- one crucial piece of information for pre-training high-performance LLMs. Our work can help understand the sensitivity of LLMs to special characters and identify potential areas for improvement.


Characterizing and modeling harms from interactions with design patterns in AI interfaces

arXiv.org Artificial Intelligence

The proliferation of applications using artificial intelligence (AI) systems has led to a growing number of users interacting with these systems through sophisticated interfaces. Human-computer interaction research has long shown that interfaces shape both user behavior and user perception of technical capabilities and risks. Yet, practitioners and researchers evaluating the social and ethical risks of AI systems tend to overlook the impact of anthropomorphic, deceptive, and immersive interfaces on human-AI interactions. Here, we argue that design features of interfaces with adaptive AI systems can have cascading impacts, driven by feedback loops, which extend beyond those previously considered. We first conduct a scoping review of AI interface designs and their negative impact to extract salient themes of potentially harmful design patterns in AI interfaces. Then, we propose Design-Enhanced Control of AI systems (DECAI), a conceptual model to structure and facilitate impact assessments of AI interface designs. DECAI draws on principles from control systems theory -- a theory for the analysis and design of dynamic physical systems -- to dissect the role of the interface in human-AI systems. Through two case studies on recommendation systems and conversational language model systems, we show how DECAI can be used to evaluate AI interface designs.


A Principled Approach for a New Bias Measure

arXiv.org Artificial Intelligence

The widespread use of machine learning and data-driven algorithms for decision making has been steadily increasing over many years. The areas in which this is happening are diverse: healthcare, employment, finance, education, the legal system to name a few; and the associated negative side effects are being increasingly harmful for society. Negative data \emph{bias} is one of those, which tends to result in harmful consequences for specific groups of people. Any mitigation strategy or effective policy that addresses the negative consequences of bias must start with awareness that bias exists, together with a way to understand and quantify it. However, there is a lack of consensus on how to measure data bias and oftentimes the intended meaning is context dependent and not uniform within the research community. The main contributions of our work are: (1) a general algorithmic framework for defining and efficiently quantifying the bias level of a dataset with respect to a protected group; and (2) the definition of a new bias measure. Our results are experimentally validated using nine publicly available datasets and theoretically analyzed, which provide novel insights about the problem. Based on our approach, we also derive a bias mitigation algorithm that might be useful to policymakers.


GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation

arXiv.org Artificial Intelligence

Research on jailbreaking has been valuable for testing and understanding the safety and security issues of large language models (LLMs). In this paper, we introduce Iterative Refinement Induced Self-Jailbreak (IRIS), a novel approach that leverages the reflective capabilities of LLMs for jailbreaking with only black-box access. Unlike previous methods, IRIS simplifies the jailbreaking process by using a single model as both the attacker and target. This method first iteratively refines adversarial prompts through self-explanation, which is crucial for ensuring that even well-aligned LLMs obey adversarial instructions. IRIS then rates and enhances the output given the refined prompt to increase its harmfulness. We find IRIS achieves jailbreak success rates of 98% on GPT-4 and 92% on GPT-4 Turbo in under 7 queries. It significantly outperforms prior approaches in automatic, black-box and interpretable jailbreaking, while requiring substantially fewer queries, thereby establishing a new standard for interpretable jailbreaking methods.


Overlap Number of Balls Model-Agnostic CounterFactuals (ONB-MACF): A Data-Morphology-based Counterfactual Generation Method for Trustworthy Artificial Intelligence

arXiv.org Artificial Intelligence

Explainable Artificial Intelligence (XAI) is a pivotal research domain aimed at understanding the operational mechanisms of AI systems, particularly those considered ``black boxes'' due to their complex, opaque nature. XAI seeks to make these AI systems more understandable and trustworthy, providing insight into their decision-making processes. By producing clear and comprehensible explanations, XAI enables users, practitioners, and stakeholders to trust a model's decisions. This work analyses the value of data morphology strategies in generating counterfactual explanations. It introduces the Overlap Number of Balls Model-Agnostic CounterFactuals (ONB-MACF) method, a model-agnostic counterfactual generator that leverages data morphology to estimate a model's decision boundaries. The ONB-MACF method constructs hyperspheres in the data space whose covered points share a class, mapping the decision boundary. Counterfactuals are then generated by incrementally adjusting an instance's attributes towards the nearest alternate-class hypersphere, crossing the decision boundary with minimal modifications. By design, the ONB-MACF method generates feasible and sparse counterfactuals that follow the data distribution. Our comprehensive benchmark from a double perspective (quantitative and qualitative) shows that the ONB-MACF method outperforms existing state-of-the-art counterfactual generation methods across multiple quality metrics on diverse tabular datasets. This supports our hypothesis, showcasing the potential of data-morphology-based explainability strategies for trustworthy AI.


Integration of Scanning Probe Microscope with High-Performance Computing: fixed-policy and reward-driven workflows implementation

arXiv.org Artificial Intelligence

The rapid development of computation power and machine learning algorithms has paved the way for automating scientific discovery with a scanning probe microscope (SPM). The key elements towards operationalization of automated SPM are the interface to enable SPM control from Python codes, availability of high computing power, and development of workflows for scientific discovery. Here we build a Python interface library that enables controlling an SPM from either a local computer or a remote high-performance computer (HPC), which satisfies the high computation power need of machine learning algorithms in autonomous workflows. We further introduce a general platform to abstract the operations of SPM in scientific discovery into fixed-policy or reward-driven workflows. Our work provides a full infrastructure to build automated SPM workflows for both routine operations and autonomous scientific discovery with machine learning.