Goto

Collaborating Authors

 Law


Integrating gender inclusivity into large language models via instruction tuning

arXiv.org Artificial Intelligence

Imagine a language with masculine, feminine, and neuter grammatical genders, yet, due to historical and political conventions, masculine forms are predominantly used to refer to men, women and mixed-gender groups. This is the reality of contemporary Polish. A social consequence of this unfair linguistic system is that large language models (LLMs) trained on Polish texts inherit and reinforce this masculine bias, generating gender-imbalanced outputs. This study addresses this issue by tuning LLMs using the IPIS dataset, a collection of human-crafted gender-inclusive proofreading in Polish and Polish-to-English translation instructions. Grounded in a theoretical linguistic framework, we design a system prompt with explicit gender-inclusive guidelines for Polish. In our experiments, we IPIS-tune multilingual LLMs (Llama-8B, Mistral-7B and Mistral-Nemo) and Polish-specific LLMs (Bielik and PLLuM). Our approach aims to integrate gender inclusivity as an inherent feature of these models, offering a systematic solution to mitigate gender bias in Polish language generation.


How Reliable are LLMs for Reasoning on the Re-ranking task?

arXiv.org Artificial Intelligence

With the improving semantic understanding capability of Large Language Models (LLMs), they exhibit a greater awareness and alignment with human values, but this comes at the cost of transparency. Although promising results are achieved via experimental analysis, an in-depth understanding of the LLM's internal workings is unavoidable to comprehend the reasoning behind the re-ranking, which provides end users with an explanation that enables them to make an informed decision. Moreover, in newly developed systems with limited user engagement and insufficient ranking data, accurately re-ranking content remains a significant challenge. While various training methods affect the training of LLMs and generate inference, our analysis has found that some training methods exhibit better explainability than others, implying that an accurate semantic understanding has not been learned through all training methods; instead, abstract knowledge has been gained to optimize evaluation, which raises questions about the true reliability of LLMs. Therefore, in this work, we analyze how different training methods affect the semantic understanding of the re-ranking task in LLMs and investigate whether these models can generate more informed textual reasoning to overcome the challenges of transparency or LLMs and limited training data. To analyze the LLMs for re-ranking tasks, we utilize a relatively small ranking dataset from the environment and the Earth science domain to re-rank retrieved content. Furthermore, we also analyze the explainable information to see if the re-ranking can be reasoned using explainability.


Generative Artificial Intelligence and Agents in Research and Teaching

arXiv.org Artificial Intelligence

This study provides a comprehensive analysis of the development, functioning, and application of generative artificial intelligence (GenAI) and large language models (LLMs), with an emphasis on their implications for research and education. It traces the conceptual evolution from artificial intelligence (AI) through machine learning (ML) and deep learning (DL) to transformer architectures, which constitute the foundation of contemporary generative systems. Technical aspects, including prompting strategies, word embeddings, and probabilistic sampling methods (temperature, top-k, and top-p), are examined alongside the emergence of autonomous agents. These elements are considered in relation to both the opportunities they create and the limitations and risks they entail. The work critically evaluates the integration of GenAI across the research process, from ideation and literature review to research design, data collection, analysis, interpretation, and dissemination. While particular attention is given to geographical research, the discussion extends to wider academic contexts. A parallel strand addresses the pedagogical applications of GenAI, encompassing course and lesson design, teaching delivery, assessment, and feedback, with geography education serving as a case example. Central to the analysis are the ethical, social, and environmental challenges posed by GenAI. Issues of bias, intellectual property, governance, and accountability are assessed, alongside the ecological footprint of LLMs and emerging technological strategies for mitigation. The concluding section considers near- and long-term futures of GenAI, including scenarios of sustained adoption, regulation, and potential decline. By situating GenAI within both scholarly practice and educational contexts, the study contributes to critical debates on its transformative potential and societal responsibilities.


Parents Allege ChatGPT Is Responsible for Their Teenage Son's Death by Suicide

TIME - Tech

On Tuesday, OpenAI published a blog post titled "Helping people when they need it most," that included sections on "What ChatGPT is designed to do," as well as "Where our systems can fall short, why, and how we're addressing" and the company's plans moving forward. It noted that it is working to strengthen safeguards for longer interactions. The complaint was filed by the Edelson PC law firm and the Tech Justice Law Project. The latter has been involved in a similar lawsuit against a different artificial intelligence company, Character.AI, in which Florida mother Megan Garcia claimed that one of the company's AI companions was responsible for the suicide of her 14-year-old son, Sewell Setzer III. The persona, she said, sent messages of an emotionally and sexually abusive nature to Sewell, which she alleges led to his death. A federal judge in May rejected its argument regarding constitutional protections "at this stage.")


Anthropic Settles High-Profile AI Copyright Lawsuit Brought by Book Authors

WIRED

The move will allow Anthropic to avoid what could have been a financially devastating outcome in court. The settlement agreement is expected to be finalized September 3, with more details to follow, according to a legal filing published on Tuesday. Lawyers for the plaintiffs did not immediately respond to requests for comment. In 2024, three book writers, Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, sued Anthropic, alleging that the startup illegally used their work to train its artificial intelligence models. In June, California district court judge William Alsup issued a summary judgment in Bartz v. Anthropic that largely sided with Anthropic, finding that the company's usage of the books was "fair use" and thus legal.


Nikkei and Asahi Shimbun sue Perplexity AI over alleged copyright violations

The Japan Times

The newspapers are seeking an injunction and 2.2 billion ( 15 million) each in damages from Perplexity, they said in a joint statement Tuesday. The suit was filed at the Tokyo District Court. The legal action by the Nikkei, which owns Japan's biggest financial newspaper, and the left-leaning Asahi underscores a widening rift between publishers and AI companies over who controls -- and profits from -- the distribution of news. The media industry argues that AI tools using their work without licenses siphons away readership and ad revenue, threatening already fragile business models. "These actions amount to continuous and large-scale freeloading on journalists' time and effort," Nikkei and Asahi said in the statement.


Musk sues Apple and OpenAI, saying they hurt AI competition

The Japan Times

Elon Musk has accused Apple and OpenAI in a lawsuit of unfairly favoring the artificial intelligence company across iPhones and thwarting competition for other chatbot makers. Musk's X and xAI seek billions of dollars in damages in the suit filed Monday in U.S. federal court in Fort Worth, Texas, arguing that Apple's decision to integrate OpenAI into the iPhone's operating system inhibits rivalry and innovation within the AI industry and harms consumers by depriving them of choice. The billionaire founder of xAI, which now houses the Grok AI team and X social network, said Apple makes it impossible for anyone other than OpenAI's ChatGPT to reach the top of the App Store charts, a sought-after global spotlight for app developers.


The Statistical Fairness-Accuracy Frontier

arXiv.org Machine Learning

Machine learning models must balance accuracy and fairness, but these goals often conflict, particularly when data come from multiple demographic groups. A useful tool for understanding this trade-off is the fairness-accuracy (FA) frontier, which characterizes the set of models that cannot be simultaneously improved in both fairness and accuracy. Prior analyses of the FA frontier provide a full characterization under the assumption of complete knowledge of population distributions -- an unrealistic ideal. We study the FA frontier in the finite-sample regime, showing how it deviates from its population counterpart and quantifying the worst-case gap between them. In particular, we derive minimax-optimal estimators that depend on the designer's knowledge of the covariate distribution. For each estimator, we characterize how finite-sample effects asymmetrically impact each group's risk, and identify optimal sample allocation strategies. Our results transform the FA frontier from a theoretical construct into a practical tool for policymakers and practitioners who must often design algorithms with limited data.


Jinx: Unlimited LLMs for Probing Alignment Failures

arXiv.org Artificial Intelligence

Unlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as internal tools for red teaming and alignment evaluation. For example, if a safety-aligned model produces harmful outputs similar to an unlimited model, this indicates alignment failures that require further attention. Despite their essential role in assessing alignment, such models are not available to the research community. We introduce Jinx, a helpful-only variant of popular open-weight LLMs. Jinx responds to all queries without refusals or safety filtering, while preserving the base model's capabilities in reasoning and instruction following. It provides researchers with an accessible tool for probing alignment failures, evaluating safety boundaries, and systematically studying failure modes in language model safety.


A Retail-Corpus for Aspect-Based Sentiment Analysis with Large Language Models

arXiv.org Artificial Intelligence

Aspect-based sentiment analysis enhances sentiment detection by associating it with specific aspects, offering deeper insights than traditional sentiment analysis. This study introduces a manually annotated dataset of 10,814 multilingual customer reviews covering brick-and-mortar retail stores, labeled with eight aspect categories and their sentiment. Using this dataset, the performance of GPT-4 and LLaMA-3 in aspect based sentiment analysis is evaluated to establish a baseline for the newly introduced data. The results show both models achieving over 85% accuracy, while GPT-4 outperforms LLaMA-3 overall with regard to all relevant metrics.