Law
Towards trustworthy AI in materials mechanics through domain-guided attention
Talies, Jesco, Breitbarth, Eric, Melching, David
Ensuring the trustworthiness and robustness of deep learning models remains a fundamental challenge, particularly in high-stakes scientific applications. In this study, we present a framework called attention-guided training that combines explainable artificial intelligence techniques with quantitative evaluation and domain-specific priors to guide model attention. We demonstrate that domain specific feedback on model explanations during training can enhance the model's generalization capabilities. We validate our approach on the task of semantic crack tip segmentation in digital image correlation data which is a key application in the fracture mechanical characterization of materials. By aligning model attention with physically meaningful stress fields, such as those described by Williams' analytical solution, attention-guided training ensures that the model focuses on physically relevant regions. This finally leads to improved generalization and more faithful explanations.
Before the Outrage: Challenges and Advances in Predicting Online Antisocial Behavior
Antisocial behavior (ASB) on social media-including hate speech, harassment, and trolling-poses growing challenges for platform safety and societal wellbeing. While prior work has primarily focused on detecting harmful content after it appears, predictive approaches aim to forecast future harmful behaviors-such as hate speech propagation, conversation derailment, or user recidivism-before they fully unfold. Despite increasing interest, the field remains fragmented, lacking a unified taxonomy or clear synthesis of existing methods. This paper presents a systematic review of over 49 studies on ASB prediction, offering a structured taxonomy of five core task types: early harm detection, harm emergence prediction, harm propagation prediction, behavioral risk prediction, and proactive moderation support. We analyze how these tasks differ by temporal framing, prediction granularity, and operational goals. In addition, we examine trends in modeling techniques-from classical machine learning to pre-trained language models-and assess the influence of dataset characteristics on task feasibility and generalization. Our review highlights methodological challenges, such as dataset scarcity, temporal drift, and limited benchmarks, while outlining emerging research directions including multilingual modeling, cross-platform generalization, and human-in-the-loop systems. By organizing the field around a coherent framework, this survey aims to guide future work toward more robust and socially responsible ASB prediction.
A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
Feng, Xiaohua, Zhang, Jiaming, Yu, Fengyuan, Wang, Chengye, Zhang, Li, Li, Kaixiang, Li, Yuyuan, Chen, Chaochao, Yin, Jianwei
With the rapid advancement of generative models, associated privacy concerns have attracted growing attention. To address this, researchers have begun adapting machine unlearning techniques from traditional classification models to generative settings. Although notable progress has been made in this area, a unified framework for systematically organizing and integrating existing work is still lacking. The substantial differences among current studies in terms of unlearning objectives and evaluation protocols hinder the objective and fair comparison of various approaches. While some studies focus on specific types of generative models, they often overlook the commonalities and systematic characteristics inherent in Generative Model Unlearning (GenMU). To bridge this gap, we provide a comprehensive review of current research on GenMU and propose a unified analytical framework for categorizing unlearning objectives, methodological strategies, and evaluation metrics. In addition, we explore the connections between GenMU and related techniques, including model editing, reinforcement learning from human feedback, and controllable generation. We further highlight the potential practical value of unlearning techniques in real-world applications. Finally, we identify key challenges and outline future research directions aimed at laying a solid foundation for further advancements in this field. We consistently maintain the related open-source materials at https://github.com/caxLee/Generative-model-unlearning-survey.
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
Anzenberg, Eitan, Samajpati, Arunava, Chandrasekar, Sivasankaran, Kacholia, Varun
The use of large language models (LLMs) in hiring promises to streamline candidate screening, but it also raises serious concerns regarding accuracy and algorithmic bias where sufficient safeguards are not in place. In this work, we benchmark several state-of-the-art foundational LLMs - including models from OpenAI, Anthropic, Google, Meta, and Deepseek, and compare them with our proprietary domain-specific hiring model (Match Score) for job candidate matching. We evaluate each model's predictive accuracy (ROC AUC, Precision-Recall AUC, F1-score) and fairness (impact ratio of cut-off analysis across declared gender, race, and intersectional subgroups). Our experiments on a dataset of roughly 10,000 real-world recent candidate-job pairs show that Match Score outperforms the general-purpose LLMs on accuracy (ROC AUC 0.85 vs 0.77) and achieves significantly more equitable outcomes across demographic groups. Notably, Match Score attains a minimum race-wise impact ratio of 0.957 (near-parity), versus 0.809 or lower for the best LLMs, (0.906 vs 0.773 for the intersectionals, respectively). We discuss why pretraining biases may cause LLMs with insufficient safeguards to propagate societal biases in hiring scenarios, whereas a bespoke supervised model can more effectively mitigate these biases. Our findings highlight the importance of domain-specific modeling and bias auditing when deploying AI in high-stakes domains such as hiring, and caution against relying on off-the-shelf LLMs for such tasks without extensive fairness safeguards. Furthermore, we show with empirical evidence that there shouldn't be a dichotomy between choosing accuracy and fairness in hiring: a well-designed algorithm can achieve both accuracy in hiring and fairness in outcomes.
Video Forgery Detection for Surveillance Cameras: A Review
Tayfor, Noor B., Rashid, Tarik A., Qader, Shko M., Hassan, Bryar A., Abdalla, Mohammed H., Majidpour, Jafar, Ahmed, Aram M., Ali, Hussein M., Aladdin, Aso M., Abdullah, Abdulhady A., Shamsaldin, Ahmed S., Sidqi, Haval M., Salih, Abdulrahman, Yaseen, Zaher M., Ameen, Azad A., Nayak, Janmenjoy, Hamza, Mahmood Yashar
The widespread availability of video recording through smartphones and digital devices has made video-based evidence more accessible than ever. Surveillance footage plays a crucial role in security, law enforcement, and judicial processes. However, with the rise of advanced video editing tools, tampering with digital recordings has become increasingly easy, raising concerns about their authenticity. Ensuring the integrity of surveillance videos is essential, as manipulated footage can lead to misinformation and undermine judicial decisions. This paper provides a comprehensive review of existing forensic techniques used to detect video forgery, focusing on their effectiveness in verifying the authenticity of surveillance recordings. Various methods, including compression-based analysis, frame duplication detection, and machine learning-based approaches, are explored. The findings highlight the growing necessity for more robust forensic techniques to counteract evolving forgery methods. Strengthening video forensic capabilities will ensure that surveillance recordings remain credible and admissible as legal evidence.
Kimi K2: Open Agentic Intelligence
Kimi Team, null, Bai, Yifan, Bao, Yiping, Chen, Guanduo, Chen, Jiahao, Chen, Ningxin, Chen, Ruijue, Chen, Yanru, Chen, Yuankun, Chen, Yutian, Chen, Zhuofu, Cui, Jialei, Ding, Hao, Dong, Mengnan, Du, Angang, Du, Chenzhuang, Du, Dikang, Du, Yulun, Fan, Yu, Feng, Yichen, Fu, Kelin, Gao, Bofei, Gao, Hongcheng, Gao, Peizhong, Gao, Tong, Gu, Xinran, Guan, Longyu, Guo, Haiqing, Guo, Jianhang, Hu, Hao, Hao, Xiaoru, He, Tianhong, He, Weiran, He, Wenyang, Hong, Chao, Hu, Yangyang, Hu, Zhenxing, Huang, Weixiao, Huang, Zhiqi, Huang, Zihao, Jiang, Tao, Jiang, Zhejun, Jin, Xinyi, Kang, Yongsheng, Lai, Guokun, Li, Cheng, Li, Fang, Li, Haoyang, Li, Ming, Li, Wentao, Li, Yanhao, Li, Yiwei, Li, Zhaowei, Li, Zheming, Lin, Hongzhan, Lin, Xiaohan, Lin, Zongyu, Liu, Chengyin, Liu, Chenyu, Liu, Hongzhang, Liu, Jingyuan, Liu, Junqi, Liu, Liang, Liu, Shaowei, Liu, T. Y., Liu, Tianwei, Liu, Weizhou, Liu, Yangyang, Liu, Yibo, Liu, Yiping, Liu, Yue, Liu, Zhengying, Lu, Enzhe, Lu, Lijun, Ma, Shengling, Ma, Xinyu, Ma, Yingwei, Mao, Shaoguang, Mei, Jie, Men, Xin, Miao, Yibo, Pan, Siyuan, Peng, Yebo, Qin, Ruoyu, Qu, Bowen, Shang, Zeyu, Shi, Lidong, Shi, Shengyuan, Song, Feifan, Su, Jianlin, Su, Zhengyuan, Sun, Xinjie, Sung, Flood, Tang, Heyi, Tao, Jiawen, Teng, Qifeng, Wang, Chensi, Wang, Dinglu, Wang, Feng, Wang, Haiming, Wang, Jianzhou, Wang, Jiaxing, Wang, Jinhong, Wang, Shengjie, Wang, Shuyi, Wang, Yao, Wang, Yejie, Wang, Yiqin, Wang, Yuxin, Wang, Yuzhi, Wang, Zhaoji, Wang, Zhengtao, Wang, Zhexu, Wei, Chu, Wei, Qianqian, Wu, Wenhao, Wu, Xingzhe, Wu, Yuxin, Xiao, Chenjun, Xie, Xiaotong, Xiong, Weimin, Xu, Boyu, Xu, Jing, Xu, Jinjing, Xu, L. H., Xu, Lin, Xu, Suting, Xu, Weixin, Xu, Xinran, Xu, Yangchuan, Xu, Ziyao, Yan, Junjie, Yan, Yuzi, Yang, Xiaofei, Yang, Ying, Yang, Zhen, Yang, Zhilin, Yang, Zonghan, Yao, Haotian, Yao, Xingcheng, Ye, Wenjie, Ye, Zhuorui, Yin, Bohong, Yu, Longhui, Yuan, Enming, Yuan, Hongbang, Yuan, Mengjie, Zhan, Haobing, Zhang, Dehao, Zhang, Hao, Zhang, Wanlu, Zhang, Xiaobin, Zhang, Yangkun, Zhang, Yizhi, Zhang, Yongting, Zhang, Yu, Zhang, Yutao, Zhang, Yutong, Zhang, Zheng, Zhao, Haotian, Zhao, Yikai, Zheng, Huabin, Zheng, Shaojie, Zhou, Jianren, Zhou, Xinyu, Zhou, Zaida, Zhu, Zhen, Zhuang, Weiyu, Zu, Xinxing
We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.
Customize Multi-modal RAI Guardrails with Precedent-based predictions
Yang, Cheng-Fu, Tran, Thanh, Christodoulopoulos, Christos, Ruan, Weitong, Gupta, Rahul, Chang, Kai-Wei
A multi-modal guardrail must effectively filter image content based on user-defined policies, identifying material that may be hateful, reinforce harmful stereotypes, contain explicit material, or spread misinformation. Deploying such guardrails in real-world applications, however, poses significant challenges. Users often require varied and highly customizable policies and typically cannot provide abundant examples for each custom policy. Consequently, an ideal guardrail should be scalable to the multiple policies and adaptable to evolving user standards with minimal retraining. Existing fine-tuning methods typically condition predictions on pre-defined policies, restricting their generalizability to new policies or necessitating extensive retraining to adapt. Conversely, training-free methods struggle with limited context lengths, making it difficult to incorporate all the policies comprehensively. To overcome these limitations, we propose to condition model's judgment on "precedents", which are the reasoning processes of prior data points similar to the given input. By leveraging precedents instead of fixed policies, our approach greatly enhances the flexibility and adaptability of the guardrail. In this paper, we introduce a critique-revise mechanism for collecting high-quality precedents and two strategies that utilize precedents for robust prediction. Experimental results demonstrate that our approach outperforms previous methods across both few-shot and full-dataset scenarios and exhibits superior generalization to novel policies.
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
Kim, Taesoo, Kim, Jinju, Kim, Dongchan, Ko, Jong Hwan, Park, Gyeong-Moon
The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individual voices from pre-trained model parameters has not been explored. In this paper, we address the new challenge of speaker identity unlearning for ZS-TTS systems. To meet this goal, we propose the first machine unlearning frameworks for ZS-TTS, especially Teacher-Guided Unlearning (TGU), designed to ensure the model forgets designated speaker identities while retaining its ability to generate accurate speech for other speakers. Our proposed methods incorporate randomness to prevent consistent replication of forget speakers' voices, assuring unlearned identities remain untraceable. Additionally, we propose a new evaluation metric, speaker-Zero Retrain Forgetting (spk-ZRF). This assesses the model's ability to disregard prompts associated with forgotten speakers, effectively neutralizing its knowledge of these voices. The experiments conducted on the state-of-the-art model demonstrate that TGU prevents the model from replicating forget speakers' voices while maintaining high quality for other speakers. The demo is available at https://speechunlearn.github.io/
Strategic Filtering for Content Moderation: Free Speech or Free of Distortion?
Ahmadi, Saba, Blum, Avrim, Xu, Haifeng, Yao, Fan
User-generated content (UGC) on social media platforms is vulnerable to incitements and manipulations, necessitating effective regulations. To address these challenges, those platforms often deploy automated content moderators tasked with evaluating the harmfulness of UGC and filtering out content that violates established guidelines. However, such moderation inevitably gives rise to strategic responses from users, who strive to express themselves within the confines of guidelines. Such phenomena call for a careful balance between: 1. ensuring freedom of speech -- by minimizing the restriction of expression; and 2. reducing social distortion -- measured by the total amount of content manipulation. We tackle the problem of optimizing this balance through the lens of mechanism design, aiming at optimizing the trade-off between minimizing social distortion and maximizing free speech. Although determining the optimal trade-off is NP-hard, we propose practical methods to approximate the optimal solution. Additionally, we provide generalization guarantees determining the amount of finite offline data required to approximate the optimal moderator effectively.
Matching Game Preferences Through Dialogical Large Language Models: A Perspective
Fabre, Renaud, Egret, Daniel, Bellot, Patrice
This perspective paper explores the future potential of "conversational intelligence" by examining how Large Language Models (LLMs) could be combined with GRAPHYP's network system to better understand human conversations and preferences. Using recent research and case studies, we propose a conceptual framework that could make AI rea-soning transparent and traceable, allowing humans to see and understand how AI reaches its conclusions. We present the conceptual perspective of "Matching Game Preferences through Dialogical Large Language Models (D-LLMs)," a proposed system that would allow multiple users to share their different preferences through structured conversations. This approach envisions personalizing LLMs by embedding individual user preferences directly into how the model makes decisions. The proposed D-LLM framework would require three main components: (1) reasoning processes that could analyze different search experiences and guide performance, (2) classification systems that would identify user preference patterns, and (3) dialogue approaches that could help humans resolve conflicting information. This perspective framework aims to create an interpretable AI system where users could examine, understand, and combine the different human preferences that influence AI responses, detected through GRAPHYP's search experience networks. The goal of this perspective is to envision AI systems that would not only provide answers but also show users how those answers were reached, making artificial intelligence more transparent and trustworthy for human decision-making.