Goto

Collaborating Authors

 Deep Learning




Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Neural Information Processing Systems

While Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed jailbreaks .


TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs

Neural Information Processing Systems

This lack of rich textual edge annotations significantly limits the exploration of contextual relationships between entities, hindering deeper insights into graph-structured data.



RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning

Neural Information Processing Systems

Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At its core, RLHF can be viewed as a specialized instance of Preference-based Reinforcement Learning (PbRL), where the preferences specifically originate from human judgments rather than arbitrary evaluators. Despite this connection, most existing approaches in both RLHF and PbRL primarily focus on optimizing a mean reward objective, neglecting scenarios that necessitate risk-awareness, such as AI safety, healthcare, and autonomous driving. These scenarios often operate under a one-episode-reward setting, which makes conventional risk-sensitive objectives inapplicable.




Scaling Laws in Linear Regression: Compute, Parameters, and Data

Neural Information Processing Systems

From the perspective of statistical learning theory, (1) is rather intriguing. Moreover, they do not provide instance-wise matching lower bounds to verify the tightness of the upper bounds.