Goto

Collaborating Authors

 Law


Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance

arXiv.org Artificial Intelligence

Calls for engagement with the public in Artificial Intelligence (AI) research, development, and governance are increasing, leading to the use of surveys to capture people's values, perceptions, and experiences related to AI. In this paper, we critically examine the state of human participant surveys associated with these topics. Through both a reflexive analysis of a survey pilot spanning six countries and a systematic literature review of 44 papers featuring public surveys related to AI, we explore prominent perspectives and methodological nuances associated with surveys to date. We find that public surveys on AI topics are vulnerable to specific Western knowledge, values, and assumptions in their design, including in their positioning of ethical concepts and societal values, lack sufficient critical discourse surrounding deployment strategies, and demonstrate inconsistent forms of transparency in their reporting. Based on our findings, we distill provocations and heuristic questions for our community, to recognize the limitations of surveys for meeting the goals of engagement, and to cultivate shared principles to design, deploy, and interpret surveys cautiously and responsibly.


Online Planning in POMDPs with State-Requests

arXiv.org Artificial Intelligence

In key real-world problems, full state information is sometimes available but only at a high cost, like activating precise yet energy-intensive sensors or consulting humans, thereby compelling the agent to operate under partial observability. For this scenario, we propose AEMS-SR (Anytime Error Minimization Search with State Requests), a principled online planning algorithm tailored for POMDPs with state requests. By representing the search space as a graph instead of a tree, AEMS-SR avoids the exponential growth of the search space originating from state requests. Theoretical analysis demonstrates AEMS-SR's $\varepsilon$-optimality, ensuring solution quality, while empirical evaluations illustrate its effectiveness compared with AEMS and POMCP, two SOTA online planning algorithms. AEMS-SR enables efficient planning in domains characterized by partial observability and costly state requests offering practical benefits across various applications.


Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models

arXiv.org Artificial Intelligence

The legal landscape encompasses a wide array of lawsuit types, presenting lawyers with challenges in delivering timely and accurate information to clients, particularly concerning critical aspects like potential imprisonment duration or financial repercussions. Compounded by the scarcity of legal experts, there's an urgent need to enhance the efficiency of traditional legal workflows. Recent advances in deep learning, especially Large Language Models (LLMs), offer promising solutions to this challenge. Leveraging LLMs' mathematical reasoning capabilities, we propose a novel approach integrating LLM-based methodologies with specially designed prompts to address precision requirements in legal Artificial Intelligence (LegalAI) applications. The proposed work seeks to bridge the gap between traditional legal practices and modern technological advancements, paving the way for a more accessible, efficient, and equitable legal system. To validate this method, we introduce a curated dataset tailored to precision-oriented LegalAI tasks, serving as a benchmark for evaluating LLM-based approaches. Extensive experimentation confirms the efficacy of our methodology in generating accurate numerical estimates within the legal domain, emphasizing the role of LLMs in streamlining legal processes and meeting the evolving demands of LegalAI.


Matching Input and Output Devices and Physical Disabilities for Human-Robot Workstations

arXiv.org Artificial Intelligence

Matching Input and Output Devices and Physical Disabilities for Human-Robot Workstations Carlo Weidemann 1,, Nils Mandischer 2,, and Burkhard Corves 1 Abstract -- As labor shortage is rising at an alarming rate, it is imperative to enable all people to work, particularly people with disabilities and elderly people. Robots are often used as universal tool to assist people with disabilities. However, for such human-robot workstations universal design fails. We mitigate the challenges of selecting an individualized set of input and output devices by matching devices required by the work process and individual disabilities adhering to the Convention on the Rights of Persons with Disabilities passed by the United Nations. The objective is to facilitate economically viable workstations with just the required devices, hence, lowering overall cost of corporate inclusion and during redesign of workplaces. Our work focuses on developing an efficient approach to filter input and output devices based on a person's disabilities, resulting in a tailored list of usable devices. The methodology enables an automated assessment of devices compatible with specific disabilities defined in International Classification of Functioning, Disability and Health. In many countries, companies are obliged by law to include people with disabilities (PwD). Meanwhile, the labor shortage is ever-present. Due to over-aging demographics and the trend towards less immigration, the gap between open positions and skilled laborers is growing and there is no turning point in sight. However, enabling skilled people to participate who would otherwise not be able to work due to congenital (PwD) or acquired (elderly, accident victims) disabilities, can become this exact turning point.


Large Language Models as Co-Pilots for Causal Inference in Medical Studies

arXiv.org Artificial Intelligence

The validity of medical studies based on real-world clinical data, such as observational studies, depends on critical assumptions necessary for drawing causal conclusions about medical interventions. Many published studies are flawed because they violate these assumptions and entail biases such as residual confounding, selection bias, and misalignment between treatment and measurement times. Although researchers are aware of these pitfalls, they continue to occur because anticipating and addressing them in the context of a specific study can be challenging without a large, often unwieldy, interdisciplinary team with extensive expertise. To address this expertise gap, we explore the use of large language models (LLMs) as co-pilot tools to assist researchers in identifying study design flaws that undermine the validity of causal inferences. We propose a conceptual framework for LLMs as causal co-pilots that encode domain knowledge across various fields, engaging with researchers in natural language interactions to provide contextualized assistance in study design. We provide illustrative examples of how LLMs can function as causal co-pilots, propose a structured framework for their grounding in existing causal inference frameworks, and highlight the unique challenges and opportunities in adapting LLMs for reliable use in epidemiological research.


FairAIED: Navigating Fairness, Bias, and Ethics in Educational AI Applications

arXiv.org Artificial Intelligence

The integration of Artificial Intelligence (AI) into education has transformative potential, providing tailored learning experiences and creative instructional approaches. However, the inherent biases in AI algorithms hinder this improvement by unintentionally perpetuating prejudice against specific demographics, especially in human-centered applications like education. This survey delves deeply into the developing topic of algorithmic fairness in educational contexts, providing a comprehensive evaluation of the diverse literature on fairness, bias, and ethics in AI-driven educational applications. It identifies the common forms of biases, such as data-related, algorithmic, and user-interaction, that fundamentally undermine the accomplishment of fairness in AI teaching aids. By outlining existing techniques for mitigating these biases, ranging from varied data gathering to algorithmic fairness interventions, the survey emphasizes the critical role of ethical considerations and legal frameworks in shaping a more equitable educational environment. Furthermore, it guides readers through the complexities of fairness measurements, methods, and datasets, shedding light on the way to bias reduction. Despite these gains, this survey highlights long-standing issues, such as achieving a balance between fairness and accuracy, as well as the need for diverse datasets. Overcoming these challenges and ensuring the ethical and fair use of AI's promise in education call for a collaborative, interdisciplinary approach.


Embedding And Clustering Your Data Can Improve Contrastive Pretraining

arXiv.org Artificial Intelligence

Recent studies of large-scale contrastive pretraining in the text embedding domain show that using single-source minibatches, rather than mixed-source minibatches, can substantially improve overall model accuracy. In this work, we explore extending training data stratification beyond source granularity by leveraging a pretrained text embedding model and the classic k-means clustering algorithm to further split training data apart by the semantic clusters within each source. Experimentally, we observe a notable increase in NDCG@10 when pretraining a BERT-based text embedding model on query-passage pairs from the MSMARCO passage retrieval dataset. Additionally, we conceptually connect our clustering approach to both the Topic Aware Sampling (TAS) aspect of the TAS-B methodology and the nearest-neighbor-based hard-negative mining aspect of the ANCE methodology and discuss how this unified view motivates future lines of research on the organization of contrastive pretraining data.


New Jersey's 500 Million Bid to Become an AI Epicenter

WIRED

New Jersey has a new plan to become the US hub for AI innovation. The state's governor signed a law on Thursday that will offer up to 500 million in tax credits for artificial intelligence companies to set up shop in the state. "We want New Jerseyans to stand at the forefront of the AI revolution--and build a more prosperous world in the process," New Jersey's Governor, Phil Murphy, a Democrat, said in a statement. "And in so doing, we are going to establish New Jersey as the home-base for R&D in generative AI." AI companies and data centers that power AI that operate at large scales in New Jersey can qualify for the tax credits, which divert unspent funds from two other state tax credit programs for job creation and real estate development enacted in response to the Covid-19 pandemic. Critics of the plan fear that it could be a win for profitable AI companies, but a loss for the state.


OpenAI tests new search engine called SearchGPT amid AI arms race

The Guardian

OpenAI is testing a new search engine that uses generative artificial intelligence to produce results, raising the prospect of a significant challenge to Google's dominance of the online search market. SearchGPT will launch with a small group of users and publishers before a potential wider rollout, the company announced on Thursday. OpenAI ultimately intends to incorporate the search features into ChatGPT, rather offer a standalone product. OpenAI said SearchGPT is a temporary prototype that will combine the company's AI models, such as ChatGPT, with the ability to search the internet. It will respond conversationally to searches, while providing up-to-date information with "clear links to relevant sources".


A new tool for copyright holders can show if their work is in AI training data

MIT Technology Review

A number of publishers and writers are in the middle of litigation against tech companies, claiming their intellectual property has been scraped into AI training data sets without their permission. The New York Times' ongoing case against OpenAI is probably the most high-profile of these. "There is a complete lack of transparency in terms of which content is used to train models, and we think this is preventing finding the right balance [between AI companies and content creators]," says Yves-Alexandre de Montjoye, an associate professor of applied mathematics and computer science at Imperial College London, who led the research. It was presented at the International Conference on Machine Learning, a top AI conference being held in Vienna this week. To create the traps, the team used a word generator to create thousands of synthetic sentences.