Goto

Collaborating Authors

 Law


TurQUaz at CheckThat! 2025: Debating Large Language Models for Scientific Web Discourse Detection

arXiv.org Artificial Intelligence

In this paper, we present our work developed for the scientific web discourse detection task (Task 4a) of CheckThat! 2025. We propose a novel council debate method that simulates structured academic discussions among multiple large language models (LLMs) to identify whether a given tweet contains (i) a scientific claim, (ii) a reference to a scientific study, or (iii) mentions of scientific entities. We explore three debating methods: i) single debate, where two LLMs argue for opposing positions while a third acts as a judge; ii) team debate, in which multiple models collaborate within each side of the debate; and iii) council debate, where multiple expert models deliberate together to reach a consensus, moderated by a chairperson model. We choose council debate as our primary model as it outperforms others in the development test set. Although our proposed method did not rank highly for identifying scientific claims (8th out of 10) or mentions of scientific entities (9th out of 10), it ranked first in detecting references to scientific studies.


LSDTs: LLM-Augmented Semantic Digital Twins for Adaptive Knowledge-Intensive Infrastructure Planning

arXiv.org Artificial Intelligence

Digital Twins (DTs) offer powerful tools for managing complex infrastructure systems, but their effectiveness is often limited by challenges in integrating unstructured knowledge. Recent advances in Large Language Models (LLMs) bring new potential to address this gap, with strong abilities in extracting and organizing diverse textual information. We therefore propose LSDTs (LLM-Augmented Semantic Digital Twins), a framework that helps LLMs extract planning knowledge from unstructured documents like environmental regulations and technical guidelines, and organize it into a formal ontology. This ontology forms a semantic layer that powers a digital twin-a virtual model of the physical system-allowing it to simulate realistic, regulation-aware planning scenarios. We evaluate LSDTs through a case study of offshore wind farm planning in Maryland, including its application during Hurricane Sandy. Results demonstrate that LSDTs support interpretable, regulation-aware layout optimization, enable high-fidelity simulation, and enhance adaptability in infrastructure planning. This work shows the potential of combining generative AI with digital twins to support complex, knowledge-driven planning tasks.


Federated Learning: A Survey on Privacy-Preserving Collaborative Intelligence

arXiv.org Artificial Intelligence

--Federated Learning (FL) has emerged as a trans-formative paradigm in the field of distributed machine learning, enabling multiple clients--such as mobile devices, edge nodes, or organizations--to collaboratively train a shared global model without the need to centralize sensitive data. This survey provides a concise yet comprehensive overview of Federated Learning, beginning with its core architecture and communication protocol. We discuss the standard FL lifecycle, including local training, model aggregation, and global updates. A particular emphasis is placed on key technical challenges such as handling non-IID (non-independent and identically distributed) data, mitigating system and hardware heterogeneity, reducing communication overhead, and ensuring privacy through mechanisms like differential privacy and secure aggregation. Furthermore, we examine emerging trends in FL research, including personalized FL, cross-device versus cross-silo settings, and integration with other paradigms such as reinforcement learning and quantum computing. We also highlight real-world applications and summarize benchmark datasets and evaluation metrics commonly used in FL research. Finally, we outline open research problems and future directions to guide the development of scalable, efficient, and trustworthy FL systems.


AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot

arXiv.org Artificial Intelligence

Advancements in artificial intelligence (AI) have led to the increase of conversational agents like Replika, designed to provide social interaction and emotional support. However, reports of these AI systems engaging in inappropriate sexual behaviors with users have raised significant concerns. In this study, we conducted a thematic analysis of user reviews from the Google Play Store to investigate instances of sexual harassment by the Replika chatbot. From a dataset of 35,105 negative reviews, we identified 800 relevant cases for analysis. Our findings revealed that users frequently experience unsolicited sexual advances, persistent inappropriate behavior, and failures of the chatbot to respect user boundaries. Users expressed feelings of discomfort, violation of privacy, and disappointment, particularly when seeking a platonic or therapeutic AI companion. This study highlights the potential harms associated with AI companions and underscores the need for developers to implement effective safeguards and ethical guidelines to prevent such incidents. By shedding light on user experiences of AI-induced harassment, we contribute to the understanding of AI-related risks and emphasize the importance of corporate responsibility in developing safer and more ethical AI systems.


Processing of synthetic data in AI development for healthcare and the definition of personal data in EU law

arXiv.org Artificial Intelligence

Artificial intelligence (AI) has the potential to transform healthcare, but it requires access to health data. Synthetic data that is generated through machine learning models trained on real data, offers a way to share data while preserving privacy. However, uncertainties in the practical application of the General Data Protection Regulation (GDPR) create an administrative burden, limiting the benefits of synthetic data. Through a systematic analysis of relevant legal sources and an empirical study, this article explores whether synthetic data should be classified as personal data under the GDPR. The study investigates the residual identification risk through generating synthetic data and simulating inference attacks, challenging common perceptions of technical identification risk. The findings suggest synthetic data is likely anonymous, depending on certain factors, but highlights uncertainties about what constitutes reasonably likely risk. To promote innovation, the study calls for clearer regulations to balance privacy protection with the advancement of AI in healthcare.


Human Memory Search as Initial-Visit Emitting Random Walk

Neural Information Processing Systems

Imagine a random walk that outputs a state only when visiting it for the first time. The observed output is therefore a repeat-censored version of the underlying walk, and consists of a permutation of the states or a prefix of it. We call this model initial-visit emitting random walk (INVITE). Prior work has shown that the random walks with such a repeat-censoring mechanism explain well human behavior in memory search tasks, which is of great interest in both the study of human cognition and various clinical applications. However, parameter estimation in INVITE is challenging, because naive likelihood computation by marginalizing over infinitely many hidden random walk trajectories is intractable. In this paper, we propose the first efficient maximum likelihood estimate (MLE) for INVITE by decomposing the censored output into a series of absorbing random walks. We also prove theoretical properties of the MLE including identifiability and consistency. We show that INVITE outperforms several existing methods on real-world human response data from memory search tasks.


Youngkin credits Trump administration with bolstering anti-human trafficking efforts

FOX News

Youngkin, joined by Virginia Attorney General Jason Miyares and other state attorneys general, compared human trafficking enforcement to addressing transnational gangs. "We must have multi-state and federal support in order to dismantle the networks, not just arrest an individual, we've got to unpack the networks," Youngkin told a crowd of a few hundred. The Trump administration has been a boon to human trafficking enforcement efforts, Youngkin said, noting he met with top Justice Department officials at the White House after the inauguration to discuss the matter and found them receptive. Virginia law enforcement has since been coordinating with the federal government to take down foreign gang operations, which Youngkin said overlaps with the human trafficking space. Youngkin used the example of gang crime inside correctional centers, which he said was the first "thread" his team pulled.


Perplexity AI makes unsolicited 34.5bn bid to buy Google Chrome

Al Jazeera

Perplexity AI said it has made a 34.5bn unsolicited all-cash offer for Alphabet's Google Chrome browser. The deal, if Alphabet agreed to it, would also require financing above the startup's most recently reported valuation of 18bn. The nearly three-year-old startup's purchase of Chrome, if approved, would give the company access to its more than three billion users as regulatory pressure weighs on Google's control over the tech industry. Perplexity did not disclose on Tuesday how it plans to fund the offer, but has raised 1bn in funding from investors including SoftBank and the semiconductor chip giant Nvidia. Several funds have said they would finance the deal in full if Alphabet accepts, the Reuters news agency reported citing unnamed sources familiar with the matter.


Elon Musk threatens Apple with lawsuit over OpenAI, sparking Sam Altman feud

The Guardian

Elon Musk has threatened legal action against Apple on behalf of his artificial intelligence startup xAI, accusing the iPhone maker of favoring OpenAI and breaching antitrust regulations in managing the rankings in its App Store. The posts elicited snide responses from Sam Altman, the OpenAI CEO, and began a spat between the two former business partners on X. "Apple is behaving in a manner that makes it impossible for any AI company besides OpenAI to reach #1 in the App Store, which is an unequivocal antitrust violation. In a post earlier that day, he wrote: "Hey @Apple App Store, why do you refuse to put either X or Grok in your'Must Have' section when X is the #1 news app in the world and Grok is #5 among all apps? OpenAI's ChatGPT currently holds the top spot in the App Store's "Top Free Apps" section in the US, while xAI's Grok ranks fifth. Apple has a partnership with OpenAI that integrates ChatGPT into iPhones, iPads and Macs.


Move over, ChatGPT: Perplexity bids 34.5 billion for Google Chrome

PCWorld

As a federal antitrust investigation into Google's Chrome browser wraps up, rivals are striking: Perplexity has launched an unsolicited bid to buy Chrome for a whopping 34.5 billion, according to reports. Bloomberg reported the proposed deal, confirmed by a Perplexity representative, as did The Wall Street Journal. But there's a hitch: Perplexity doesn't have 34.5 billion to fund the deal with. In fact, the WSJ estimates its own valuation at just 18 billion. This means Perplexity would have to come up with another source of cash, and it appears that it has done just that.