Goto

Collaborating Authors

 Law


Scientist wants to implant prisoners with 'memories' of their crimes that show the victim's perspective

Daily Mail - Science & tech

A scientist has unveiled a concept for a prison of the future that he has claimed would fast-track a criminal's release to minutes, instead of years or decades. Called Cognify, the design would implant synthetic memories of a person's crime into their brain, but showing their victim's perspective. The system could feature a VR-like device that displays AI-generated footage of the offence, coupled with a brain implant that induces emotional states like remorse or regret - feelings some individuals may not produce on their own. The concept, developed by Hashem Al-Ghaili, would ensure the long-term effects of the therapy session by making the memories permanent. Called Cognify, the design would implant synthetic memories of a person's crime into their brain, but showing their victim's perspective There are more than 1.7 million people currently incarcerated in the US.


AI: World's biggest music labels sue over copyright

BBC News

Supporters have compared machine learning by AI tools to the way humans learn by reading, hearing and seeing previous works. But in the complaints, which were filed in federal court in Massachusetts and New York, the record labels say the AI firms are simply making money from having copied the songs. The complaints say Suno and Udio produce works like "Prancing Queen" that even devoted ABBA fans would struggle to distinguish from an authentic recording from the band. Songs cited in the Udio lawsuit include Mariah Carey's "All I Want for Christmas is You" and "My Girl" by The Temptations. They said there was nothing about AI that excused the firms from "playing by the rules" and warned that the "wholesale theft" of the recordings threatened "the entire music ecosystem". The lawsuits come just months after roughly 200 artists including Billie Eilish and Nicki Minaj signed a letter calling for the "predatory" use of artificial intelligence (AI) in the music industry to be stopped.


Record labels sue AI music generators for 'massive infringement of recorded music'

Engadget

Major music labels are taking on AI startups that they believe trained on their songs without paying. The filings against the AI companies reportedly demand injunctions against future use and damages of up to 150,000 per infringed work. The suits appear aimed at establishing licensed training as the only acceptable industry framework for AI moving forward -- while instilling fear in companies that train their models without consent. Suno AI and Udio AI (Uncharted Labs run the latter) are startups with software that generates music based on text inputs. The former is a partner of Microsoft for its CoPilot music generation tool.


US Record Labels Sue AI Music Generators Suno and Udio for Copyright Infringement

WIRED

The music industry has officially declared war on Suno and Udio, two of the most prominent AI music generators. The plaintiffs seek damages up to 150,000 per work infringed. The lawsuit against Suno is filed in Massachusetts, while the case against Udio's parent company Uncharted Inc. was filed in New York. Suno and Udio did not immediately respond to a request to comment. "Unlicensed services like Suno and Udio that claim it's'fair' to copy an artist's life's work and exploit it for their own profit without consent or pay set back the promise of genuinely innovative AI for us all," Recording Industry Association of America chairman and CEO Mitch Glazier said in a press release.


Geologists raise concerns over possible censorship and bias in Chinese chatbot

The Guardian

Geologists have raised concerns about potential Chinese censorship and bias in a chatbot being developed with the backing of the International Union of Geological Sciences (IUGS), one of the world's largest scientific organisations and a Unesco partner. The GeoGPT chatbot is aimed at geoscientists and researchers, particularly in the global south, to help them develop their understanding of earth sciences by drawing on swaths of data and research on billions of years of the planet's history. It is an initiative from Deep-time Digital Earth (DDE), a largely Chinese-funded programme founded in 2019 to enhance international scientific cooperation and help countries to realise the UN's sustainable development goals. Part of the underlying AI for GeoGPT is Qwen, a large language model built by the Chinese tech company Alibaba. Responding to the article, DDE representatives Michael Stephenson, Hans Thybo, Chengshan Wang and Ishwaran Natarajan said the chatbot also used Meta's Llama, another large language model, and that during testing they had not noticed any state censorship, which they said was "unlikely" given that the system was "based entirely in geoscience information".


LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content

arXiv.org Artificial Intelligence

As large language models (LLMs) become increasingly prevalent in a wide variety of applications, concerns about the safety of their outputs have become more significant. Most efforts at safety-tuning or moderation today take on a predominantly Western-centric view of safety, especially for toxic, hateful, or violent speech. In this paper, we describe LionGuard, a Singapore-contextualized moderation classifier that can serve as guardrails against unsafe LLM outputs. When assessed on Singlish data, LionGuard outperforms existing widely-used moderation APIs, which are not finetuned for the Singapore context, by 14% (binary) and up to 51% (multi-label). Our work highlights the benefits of localization for moderation classifiers and presents a practical and scalable approach for low-resource languages.


GIEBench: Towards Holistic Evaluation of Group Identity-based Empathy for Large Language Models

arXiv.org Artificial Intelligence

As large language models (LLMs) continue to develop and gain widespread application, the ability of LLMs to exhibit empathy towards diverse group identities and understand their perspectives is increasingly recognized as critical. Most existing benchmarks for empathy evaluation of LLMs focus primarily on universal human emotions, such as sadness and pain, often overlooking the context of individuals' group identities. To address this gap, we introduce GIEBench, a comprehensive benchmark that includes 11 identity dimensions, covering 97 group identities with a total of 999 single-choice questions related to specific group identities. GIEBench is designed to evaluate the empathy of LLMs when presented with specific group identities such as gender, age, occupation, and race, emphasizing their ability to respond from the standpoint of the identified group. This supports the ongoing development of empathetic LLM applications tailored to users with different identities. Our evaluation of 23 LLMs revealed that while these LLMs understand different identity standpoints, they fail to consistently exhibit equal empathy across these identities without explicit instructions to adopt those perspectives. This highlights the need for improved alignment of LLMs with diverse values to better accommodate the multifaceted nature of human identities. Our datasets are available at https://github.com/GIEBench/GIEBench.


SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have highlighted the necessity of effective unlearning mechanisms to comply with data regulations and ethical AI practices. LLM unlearning aims at removing undesired data influences and associated model capabilities without compromising utility beyond the scope of unlearning. While interest in studying LLM unlearning is growing, the impact of the optimizer choice for LLM unlearning remains unexplored. In this work, we shed light on the significance of optimizer selection in LLM unlearning for the first time, establishing a clear connection between second-order optimization and influence unlearning (a classical approach using influence functions to update the model for data influence removal). This insight propels us to develop a second-order optimization-based LLM unlearning framework, termed Second-Order UnLearning (SOUL), which extends the static, one-shot model update using influence unlearning to a dynamic, iterative unlearning process. Our extensive experiments show that SOUL consistently outperforms conventional first-order methods across various unlearning tasks, models, and metrics, indicating that second-order optimization offers an effective and broadly applicable solution for LLM unlearning. Codes are available at https://github.com/OPTML-Group/SOUL.


Modeling the Sacred: Considerations when Using Religious Texts in Natural Language Processing

arXiv.org Artificial Intelligence

This position paper concerns the use of religious texts in Natural Language Processing (NLP), which is of special interest to the Ethics of NLP. Religious texts are expressions of culturally important values, and machine learned models have a propensity to reproduce cultural values encoded in their training data. Furthermore, translations of religious texts are frequently used by NLP researchers when language data is scarce. This repurposes the translations from their original uses and motivations, which often involve attracting new followers. This paper argues that NLP's use of such texts raises considerations that go beyond model biases, including data provenance, cultural contexts, and their use in proselytism. We argue for more consideration of researcher positionality, and of the perspectives of marginalized linguistic and religious communities.


eagerlearners at SemEval2024 Task 5: The Legal Argument Reasoning Task in Civil Procedure

arXiv.org Artificial Intelligence

This study investigates the performance of the zero-shot method in classifying data using three large language models, alongside two models with large input token sizes and the two pre-trained models on legal data. Our main dataset comes from the domain of U.S. civil procedure. It includes summaries of legal cases, specific questions, potential answers, and detailed explanations for why each solution is relevant, all sourced from a book aimed at law students. By comparing different methods, we aimed to understand how effectively they handle the complexities found in legal datasets. Our findings show how well the zero-shot method of large language models can understand complicated data. We achieved our highest F1 score of 64% in these experiments.