Goto

Collaborating Authors

 Government


A Probabilistic Perspective on Unlearning and Alignment for Large Language Models

arXiv.org Artificial Intelligence

Comprehensive evaluation of Large Language Models (LLMs) is an open research problem. Existing evaluations rely on deterministic point estimates generated via greedy decoding. However, we find that deterministic evaluations fail to capture the whole output distribution of a model, yielding inaccurate estimations of model capabilities. This is particularly problematic in critical contexts such as unlearning and alignment, where precise model evaluations are crucial. To remedy this, we introduce the first formal probabilistic evaluation framework in LLMs. Namely, we derive novel metrics with high-probability guarantees concerning the output distribution of a model. Our metrics are application-independent and allow practitioners to make more reliable estimates about model capabilities before deployment. Through a case study focused on unlearning, we reveal that deterministic evaluations falsely indicate successful unlearning, whereas our probabilistic evaluations demonstrate that most if not all of the supposedly unlearned information remains accessible in these models. Additionally, we propose a novel unlearning loss based on entropy optimization and adaptive temperature scaling, which significantly improves unlearning in probabilistic settings on recent benchmarks. Our proposed shift from point estimates to probabilistic evaluations of output distributions represents an important step toward comprehensive evaluations of LLMs. Large Language Models (LLMs) are widely employed across various applications, from chatbots to code generation, relying on outputs generated through probabilistic decoding methods such as beam-search and multinominal sampling.


The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

arXiv.org Artificial Intelligence

Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a dataset that maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide what alignment data.


CultureLLM: Incorporating Cultural Differences into Large Language Models

arXiv.org Artificial Intelligence

Large language models (LLMs) are reported to be partial to certain cultures owing to the training data dominance from the English corpora. Since multilingual cultural data are often expensive to collect, existing efforts handle this by prompt engineering or culture-specific pre-training. However, they might overlook the knowledge deficiency of low-resource culture and require extensive computing resources. In this paper, we propose CultureLLM, a cost-effective solution to incorporate cultural differences into LLMs. CultureLLM adopts World Value Survey (WVS) as seed data and generates semantically equivalent training data via the proposed semantic data augmentation. Using only 50 seed samples from WVS with augmented data, we fine-tune culture-specific LLMs and one unified model (CultureLLM-One) for 9 cultures covering rich and low-resource languages. Extensive experiments on 60 culture-related datasets demonstrate that CultureLLM significantly outperforms various counterparts such as GPT-3.5 (by 8.1%) and Gemini Pro (by 9.5%) with comparable performance to GPT-4 or even better. Our human study shows that the generated samples are semantically equivalent to the original samples, providing an effective solution for LLMs augmentation. Code is released at https://github.com/Scarelette/CultureLLM.


Guardian of the Ensembles: Introducing Pairwise Adversarially Robust Loss for Resisting Adversarial Attacks in DNN Ensembles

arXiv.org Artificial Intelligence

Adversarial attacks rely on transferability, where an adversarial example (AE) crafted on a surrogate classifier tends to mislead a target classifier. Recent ensemble methods demonstrate that AEs are less likely to mislead multiple classifiers in an ensemble. This paper proposes a new ensemble training using a Pairwise Adversarially Robust Loss (PARL) that by construction produces an ensemble of classifiers with diverse decision boundaries. PARL utilizes outputs and gradients of each layer with respect to network parameters in every classifier within the ensemble simultaneously. PARL is demonstrated to achieve higher robustness against black-box transfer attacks than previous ensemble methods as well as adversarial training without adversely affecting clean example accuracy. Extensive experiments using standard Resnet20, WideResnet28-10 classifiers demonstrate the robustness of PARL against state-of-the-art adversarial attacks. While maintaining similar clean accuracy and lesser training time, the proposed architecture has a 24.8% increase in robust accuracy ($\epsilon$ = 0.07) from the state-of-the art method.


U.S. tightens curbs on China's access to AI memory and chip tools

The Japan Times

The U.S. unveiled new restrictions on China's access to vital components for chips and AI, escalating a campaign to contain Beijing's technological ambitions but stopping short of earlier proposals that would have sanctioned more key Chinese firms. The Department of Commerce slapped fresh curbs on the sale of high-bandwidth memory chips made by U.S. and foreign companies, likely affecting South Korea's SK Hynix and Samsung Electronics as well as Idaho-based Micron Technology. Those components handle data storage and are essential for AI applications. The agency also expanded existing controls on chipmaking gear, including products made by U.S. firms at foreign facilities, but with carveouts for key allies, such as Japan and the Netherlands. That comes after months of negotiations between Washington, Tokyo and the Hague, during which Biden officials floated -- but ultimately did not pursue -- applying U.S. controls to companies like Tokyo Electron and ASML Holding NV.


US unleashes another crackdown on China's chip industry

Al Jazeera

The United States has launched its third crackdown in three years on China's semiconductor industry, curbing exports to 140 companies, including chip equipment maker Naura Technology Group, among other moves. The latest effort on Monday to hobble Beijing's chipmaking ambitions also hits Chinese chip toolmakers Piotech, ACM Research and SiCarrier Technology with new export restrictions as part of the package, which also takes aim at shipments of advanced memory chips and more chipmaking tools to China. The move is one of President Joe Biden's last large-scale efforts to stymie China's ability to access and produce chips that can help advance artificial intelligence for military applications, or otherwise threaten US national security. It comes just weeks before the swearing-in of Republican President-elect Donald Trump, who is expected to retain many of Biden's tough-on-China measures. The package includes curbs on China-bound shipments of high bandwidth memory (HBM) chips, critical for high-end applications like AI training; curbs on 24 additional chipmaking tools and three software tools; and export curbs on chipmaking equipment made in countries such as Singapore and Malaysia.


Israeli attacks kill two people in Lebanon; Hezbollah responds

Al Jazeera

Israel has killed two people, including a State Security officer, in separate attacks in Lebanon as it continues its assaults on the country since the ceasefire with Hezbollah came into effect last week. For its part, the Lebanese group said on Monday that it carried out a "preliminary defensive response" to the "repeated violations" of the ceasefire by attacking an Israeli military base in the hills of Kfar Chouba, a disputed area that Lebanon claims as its own. Hezbollah said Israeli breaches of the truce that went into effect on Wednesday include deadly air raids across Lebanon, shooting at civilians in the south, and flying drones and jets in Lebanese airspace, including over the capital, Beirut. The group said it launched its "warning" attack because "appeals by the relevant authorities to stop these violations did not succeed". The renewed violence highlights the fragility of the ceasefire, which ended a devastating war that killed nearly 4,000 people in Lebanon and saw Hezbollah fire rockets daily at Israel.


The Download: words of wisdom from the departing White House tech advisor, and controversial AI manga translation

MIT Technology Review

President Biden's administration will end within two months, and likely to depart with him is Arati Prabhakar, the top mind for science and technology in his cabinet. She has served as Director of the White House Office of Science and Technology Policy since 2022 and was the first to demonstrate ChatGPT to the president in the Oval Office. Prabhakar was instrumental in passing the president's executive order on AI in 2023, which sets guidelines for tech companies to make AI safer and more transparent (though it relies on voluntary participation). As she prepares for the end of the administration, MIT Technology Review sat down with Prabhakar and asked her to reflect on President Biden's AI accomplishments, and how the approach to AI risks, immigration policies, the CHIPS Act and more could change under Trump. This manga publisher is using Anthropic's AI to translate Japanese comics into English A Japanese publishing startup is using Anthropic's flagship large language model Claude to help translate manga into English, allowing the company to churn out a new title for a Western audience in just a few days rather than the 2-3 months it would take a team of humans.


If You're Going to Make Something, Here's How to Make It Robust

WIRED

Christopher Tidy was 10 years old the first time he took apart an engine. The carburetor--the block of machinery that supplies a gas engine with fuel and air and helps to spark ignition--was a mess. It was blocked with thick layers of congealed fuel and dust. Tidy saw the problem and just happened to have some tools nearby and a burning curiosity about how exactly this thing worked and what he could do to fix it. That quickly turned into an attempt "to assemble a kind of Frankenstein engine" out of the parts of many discarded petrol engines. He disassembled the rumbling machine piece by piece until he found the offending parts, then doused the carburetor in gasoline, followed by water and dish soap, then scrubbed it clean with a toothbrush.


This Website Shows How Much Google's AI Can Glean From Your Photos

WIRED

Software engineer Vishnu Mohandas decided he would quit Google in more ways than one when he learned the tech giant had briefly helped the US military develop AI to study drone footage. In 2020, he left his job working on Google Assistant and also stopped backing up all of his images to Google Photos. He feared that his content could be used to train AI systems, even if they weren't specifically ones tied to the Pentagon project. "I don't control any of the future outcomes that this will enable," Mohandas thought. "So now, shouldn't I be more responsible?" Mohandas, who taught himself programming and is based in Bengaluru, India, decided he wanted to develop an alternative service for storing and sharing photos that is open source and end-to-end encrypted.