Goto

Collaborating Authors

 Law


JULI: Jailbreak Large Language Models by Self-Introspection

arXiv.org Artificial Intelligence

Large Language Models (LLMs) are trained with safety alignment to prevent generating malicious content. Although some attacks have highlighted vulnerabilities in these safety-aligned LLMs, they typically have limitations, such as necessitating access to the model weights or the generation process. Since proprietary models through API-calling do not grant users such permissions, these attacks find it challenging to compromise them. In this paper, we propose Jailbreaking Using LLM Introspection (JULI), which jailbreaks LLMs by manipulating the token log probabilities, using a tiny plug-in block, BiasNet. JULI relies solely on the knowledge of the target LLM's predicted token log probabilities. It can effectively jailbreak API-calling LLMs under a black-box setting and knowing only top-$5$ token log probabilities. Our approach demonstrates superior effectiveness, outperforming existing state-of-the-art (SOTA) approaches across multiple metrics.


Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets

arXiv.org Artificial Intelligence

Can LLMs Be Trusted for Evaluating RAG Systems? Abstract--Retrieval-Augmented Generation (RAG) has advanced significantly in recent years. The complexity of RAG systems, which involve multiple components--such as indexi ng, retrieval, and generation--along with numerous other param e-ters, poses substantial challenges for systematic evaluat ion and quality enhancement. Previous research highlights that ev aluating RAG systems is essential for documenting advancements, com - paring configurations, and identifying effective approach es for domain-specific applications. This study systematically r eviews 63 academic articles to provide a comprehensive overview of state-of-the-art RAG evaluation methodologies, focusing on four key areas: datasets, retrievers, indexing and databases, a nd the generator component. We observe the feasibility of an automated evaluation approach for each component of a RAG system, leveraging an LLM capable of both generating evalua tion datasets and conducting evaluations. In addition, we found that further practical research is essential to provide compani es with clear guidance on the do's and don'ts of implementing and evaluating RAG systems. By synthesizing evaluation approa ches for key RAG components and emphasizing the creation and adaptation of domain-specific datasets for benchmarking, w e contribute to the advancement of systematic evaluation met hods and the improvement of evaluation rigor for RAG systems. Furthermore, by examining the interplay between automated approaches leveraging LLMs and human judgment, we contribute to the ongoing discourse on balancing automation and human input, clarifying their respective contributions, limita tions, and challenges in achieving robust and reliable evaluations. In recent years, Large Language Models (LLMs) have made significant progress in research and have grown increasingl y popular [1]. However, LLMs face several challenges, includ - ing issues with hallucinations caused by insufficient conte xt [2], as well as limitations in their learned content, which prevent them from addressing questions requiring specific or proprietary information [1].


Leveraging Deep Learning for Physical Model Bias of Global Air Quality Estimates

arXiv.org Artificial Intelligence

Air pollution is the world's largest environmental risk factor for human disease and premature death, resulting in more than 6 million permature deaths in 2019. Currently, there is still a challenge to model one of the most important air pollutants, surface ozone, particularly at scales relevant for human health impacts, with the drivers of global ozone trends at these scales largely unknown, limiting the practical use of physics-based models. We employ a 2D Convolutional Neural Network based architecture that estimate surface ozone MOMO-Chem model residuals, referred to as model bias. We demonstrate the potential of this technique in North America and Europe, highlighting its ability better to capture physical model residuals compared to a traditional machine learning method. We assess the impact of incorporating land use information from high-resolution satellite imagery to improve model estimates. Importantly, we discuss how our results can improve our scientific understanding of the factors impacting ozone bias at urban scales that can be used to improve environmental policy.


Uncertainty Quantification for Surface Ozone Emulators using Deep Learning

arXiv.org Artificial Intelligence

Air pollution is a global hazard, and as of 2023, 94\% of the world's population is exposed to unsafe pollution levels. Surface Ozone (O3), an important pollutant, and the drivers of its trends are difficult to model, and traditional physics-based models fall short in their practical use for scales relevant to human-health impacts. Deep Learning-based emulators have shown promise in capturing complex climate patterns, but overall lack the interpretability necessary to support critical decision making for policy changes and public health measures. We implement an uncertainty-aware U-Net architecture to predict the Multi-mOdel Multi-cOnstituent Chemical data assimilation (MOMO-Chem) model's surface ozone residuals (bias) using Bayesian and quantile regression methods. We demonstrate the capability of our techniques in regional estimation of bias in North America and Europe for June 2019. We highlight the uncertainty quantification (UQ) scores between our two UQ methodologies and discern which ground stations are optimal and sub-optimal candidates for MOMO-Chem bias correction, and evaluate the impact of land-use information in surface ozone residual modeling.


Learning to Reason for Factuality

arXiv.org Artificial Intelligence

Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than their non-reasoning counterparts on long-form factuality benchmarks. However, extending online Reinforcement Learning (RL), a key component in recent R-LLM advancements, to the long-form factuality setting poses several unique challenges due to the lack of reliable verification methods. Previous work has utilized automatic factuality evaluation frameworks such as FActScore to curate preference data in the offline RL setting, yet we find that directly leveraging such methods as the reward in online RL leads to reward hacking in multiple ways, such as producing less detailed or relevant responses. We propose a novel reward function that simultaneously considers the factual precision, response detail level, and answer relevance, and applies online RL to learn high quality factual reasoning. Evaluated on six long-form factuality benchmarks, our factual reasoning model achieves an average reduction of 23.1 percentage points in hallucination rate, a 23% increase in answer detail level, and no degradation in the overall response helpfulness.


Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI

arXiv.org Artificial Intelligence

AI (super) alignment describes the challenge of ensuring (future) AI systems behave in accordance with societal norms and goals. While a quickly evolving literature is addressing biases and inequalities, the geographic variability of alignment remains underexplored. Simply put, what is considered appropriate, truthful, or legal can differ widely across regions due to cultural norms, political realities, and legislation. Alignment measures applied to AI/ML workflows can sometimes produce outcomes that diverge from statistical realities, such as text-to-image models depicting balanced gender ratios in company leadership despite existing imbalances. Crucially, some model outputs are globally acceptable, while others, e.g., questions about Kashmir, depend on knowing the user's location and their context. This geographic sensitivity is not new. For instance, Google Maps renders Kashmir's borders differently based on user location. What is new is the unprecedented scale and automation with which AI now mediates knowledge, expresses opinions, and represents geographic reality to millions of users worldwide, often with little transparency about how context is managed. As we approach Agentic AI, the need for spatio-temporally aware alignment, rather than one-size-fits-all approaches, is increasingly urgent. This paper reviews key geographic research problems, suggests topics for future work, and outlines methods for assessing alignment sensitivity.


Trump calls on CEO of tech firm Intel to resign over China investments

Al Jazeera

United States President Donald Trump has fired off a social media message calling on the head of the US technology firm Intel to resign from his post as chief executive officer. Trump's decision to denounce Intel CEO Lip-Bu Tan on Thursday morning sent the company's stocks tumbling, amid the uncertainty about the future of its leadership. "The CEO of INTEL is highly CONFLICTED and must resign, immediately," Trump wrote. "There is no other solution to this problem. Thank you for your attention to this problem!" Trump's post appeared to be a response to reports that Tan has invested nearly 200m in Chinese technology manufacturing and chip firms, including some with links to the country's military.


Tokyo Electron fires worker suspected of stealing TSMC tech

The Japan Times

Tokyo Electron said Thursday it's fired an employee at its Taipei unit, making its first public statement since the island's government arrested six people suspected of stealing trade secrets from Taiwan Semiconductor Manufacturing Company. The Japanese chip gear maker said in a statement it's cooperating with the ongoing investigation, though it remains unclear if any data had been shared with third parties. Taiwan prosecutors arrested the six suspected of intellectual property theft at TSMC this week, including an individual that local media identified as a former Tokyo Electron staffer. The Japanese company is one of the largest suppliers of semiconductor-fabrication tools and gear to TSMC, which in turn uses the equipment to make Nvidia AI accelerators and Apple iPhone processors. Investigators haven't disclosed many more details of the case, which coincides with a quickening race by the likes of Meta and DeepSeek to develop artificial intelligence in the post-ChatGPT era.