Education
Social Identity in Human-Agent Interaction: A Primer
Social identity theory (SIT) and social categorization theory (SCT) are two facets of the social identity approach (SIA) to understanding social phenomena. SIT and SCT are models that describe and explain how people interact with one another socially, connecting the individual to the group through an understanding of underlying psychological mechanisms and intergroup behaviour. SIT, originally developed in the 1970s, and SCT, a later, more general offshoot, have been broadly applied to a range of social phenomena among people. The rise of increasingly social machines embedded in daily life has spurned efforts on understanding whether and how artificial agents can and do participate in SIA activities. As agents like social robots and chatbots powered by sophisticated large language models (LLMs) advance, understanding the real and potential roles of these technologies as social entities is crucial. Here, I provide a primer on SIA and extrapolate, through case studies and imagined examples, how SIT and SCT can apply to artificial social agents. I emphasize that not all human models and sub-theories will apply. I further argue that, given the emerging competence of these machines and our tendency to be taken in by them, we experts may need to don the hat of the uncanny killjoy, for our own good.
Intern-S1: A Scientific Multimodal Foundation Model
Bai, Lei, Cai, Zhongrui, Cao, Yuhang, Cao, Maosong, Cao, Weihan, Chen, Chiyu, Chen, Haojiong, Chen, Kai, Chen, Pengcheng, Chen, Ying, Chen, Yongkang, Cheng, Yu, Chu, Pei, Chu, Tao, Cui, Erfei, Cui, Ganqu, Cui, Long, Cui, Ziyun, Deng, Nianchen, Ding, Ning, Dong, Nanqing, Dong, Peijie, Dou, Shihan, Du, Sinan, Duan, Haodong, Fan, Caihua, Gao, Ben, Gao, Changjiang, Gao, Jianfei, Gao, Songyang, Gao, Yang, Gao, Zhangwei, Ge, Jiaye, Ge, Qiming, Gu, Lixin, Gu, Yuzhe, Guo, Aijia, Guo, Qipeng, Guo, Xu, He, Conghui, He, Junjun, Hong, Yili, Hou, Siyuan, Hu, Caiyu, Hu, Hanglei, Hu, Jucheng, Hu, Ming, Hua, Zhouqi, Huang, Haian, Huang, Junhao, Huang, Xu, Huang, Zixian, Jiang, Zhe, Kong, Lingkai, Li, Linyang, Li, Peiji, Li, Pengze, Li, Shuaibin, Li, Tianbin, Li, Wei, Li, Yuqiang, Lin, Dahua, Lin, Junyao, Lin, Tianyi, Lin, Zhishan, Liu, Hongwei, Liu, Jiangning, Liu, Jiyao, Liu, Junnan, Liu, Kai, Liu, Kaiwen, Liu, Kuikun, Liu, Shichun, Liu, Shudong, Liu, Wei, Liu, Xinyao, Liu, Yuhong, Liu, Zhan, Lu, Yinquan, Lv, Haijun, Lv, Hongxia, Lv, Huijie, Lv, Qitan, Lv, Ying, Lyu, Chengqi, Ma, Chenglong, Ma, Jianpeng, Ma, Ren, Ma, Runmin, Ma, Runyuan, Ma, Xinzhu, Ma, Yichuan, Ma, Zihan, Mi, Sixuan, Ning, Junzhi, Ning, Wenchang, Pang, Xinle, Peng, Jiahui, Peng, Runyu, Qiao, Yu, Qiu, Jiantao, Qu, Xiaoye, Qu, Yuan, Ren, Yuchen, Shang, Fukai, Shao, Wenqi, Shen, Junhao, Shen, Shuaike, Song, Chunfeng, Song, Demin, Song, Diping, Su, Chenlin, Su, Weijie, Sun, Weigao, Sun, Yu, Tan, Qian, Tang, Cheng, Tang, Huanze, Tang, Kexian, Tang, Shixiang, Tong, Jian, Wang, Aoran, Wang, Bin, Wang, Dong, Wang, Lintao, Wang, Rui, Wang, Weiyun, Wang, Wenhai, Wang, Jiaqi, Wang, Yi, Wang, Ziyi, Wu, Ling-I, Wu, Wen, Wu, Yue, Wu, Zijian, Xiao, Linchen, Xing, Shuhao, Xu, Chao, Xu, Huihui, Xu, Jun, Xu, Ruiliang, Xu, Wanghan, Yang, GanLin, Yang, Yuming, Ye, Haochen, Ye, Jin, Ye, Shenglong, Yu, Jia, Yu, Jiashuo, Yu, Jing, Yuan, Fei, Zang, Yuhang, Zhang, Bo, Zhang, Chao, Zhang, Chen, Zhang, Hongjie, Zhang, Jin, Zhang, Qiaosheng, Zhang, Qiuyinzhe, Zhang, Songyang, Zhang, Taolin, Zhang, Wenlong, Zhang, Wenwei, Zhang, Yechen, Zhang, Ziyang, Zhao, Haiteng, Zhao, Qian, Zhao, Xiangyu, Zhao, Xiangyu, Zhou, Bowen, Zhou, Dongzhan, Zhou, Peiheng, Zhou, Yuhao, Zhou, Yunhua, Zhu, Dongsheng, Zhu, Lin, Zou, Yicheng
In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that of closed-source models. However, in high-value but more challenging scientific professional fields, either the fields still rely on expert models, or the progress of general foundation models lags significantly compared to those in popular areas, far from sufficient for transforming scientific research and leaving substantial gap between open-source models and closed-source models in these scientific domains. To mitigate this gap and explore a step further toward Artificial General Intelligence (AGI), we introduce Intern-S1, a specialized generalist equipped with general understanding and reasoning capabilities with expertise to analyze multiple science modal data. Intern-S1 is a multimodal Mixture-of-Experts (MoE) model with 28 billion activated parameters and 241 billion total parameters, continually pre-trained on 5T tokens, including over 2.5T tokens from scientific domains. In the post-training stage, Intern-S1 undergoes offline and then online reinforcement learning (RL) in InternBootCamp, where we propose Mixture-of-Rewards (MoR) to synergize the RL training on more than 1000 tasks simultaneously. Through integrated innovations in algorithms, data, and training systems, Intern-S1 achieved top-tier performance in online RL training. On comprehensive evaluation benchmarks, Intern-S1 demonstrates competitive performance on general reasoning tasks among open-source models and significantly outperforms open-source models in scientific domains, surpassing closed-source state-of-the-art models in professional tasks, such as molecular synthesis planning, reaction condition prediction, predicting thermodynamic stabilities for crystals. Our models are available at https://huggingface.co/internlm/Intern-S1.
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
Dong, Wenhan, Sun, Zhen, Zhao, Yuemeng, Peng, Zifan, Wu, Jun, Zheng, Jingyi, Liu, Yule, He, Xinlei, Wang, Yu, Wang, Ruiming, Huang, Xinyi, Mo, Lei
Large language models (LLMs) have demonstrated potential in educational applications, yet their capacity to accurately assess the cognitive alignment of reading materials with students' developmental stages remains insufficiently explored. This gap is particularly critical given the foundational educational principle of the Zone of Proximal Development (ZPD), which emphasizes the need to match learning resources with Students' Cognitive Abilities (SCA). Despite the importance of this alignment, there is a notable absence of comprehensive studies investigating LLMs' ability to evaluate reading comprehension difficulty across different student age groups, especially in the context of Chinese language education. To fill this gap, we introduce ZPD-SCA, a novel benchmark specifically designed to assess stage-level Chinese reading comprehension difficulty. The benchmark is annotated by 60 Special Grade teachers, a group that represents the top 0.15% of all in-service teachers nationwide. Experimental results reveal that LLMs perform poorly in zero-shot learning scenarios, with Qwen-max and GLM even falling below the probability of random guessing. When provided with in-context examples, LLMs performance improves substantially, with some models achieving nearly double the accuracy of their zero-shot baselines. These results reveal that LLMs possess emerging abilities to assess reading difficulty, while also exposing limitations in their current training for educationally aligned judgment. Notably, even the best-performing models display systematic directional biases, suggesting difficulties in accurately aligning material difficulty with SCA. Furthermore, significant variations in model performance across different genres underscore the complexity of task. We envision that ZPD-SCA can provide a foundation for evaluating and improving LLMs in cognitively aligned educational applications.
SEA-BED: Southeast Asia Embedding Benchmark
Ponwitayarat, Wuttikorn, Ng, Raymond, Montalan, Jann Railey, Aung, Thura, Ngui, Jian Gang, Susanto, Yosephine, Tjhi, William, Tasawong, Panuthep, Cambria, Erik, Chuangsuwanich, Ekapol, Nutanong, Sarana, Limkonchotiwat, Peerat
Sentence embeddings are essential for NLP tasks such as semantic search, re-ranking, and textual similarity. Although multilingual benchmarks like MMTEB broaden coverage, Southeast Asia (SEA) datasets are scarce and often machine-translated, missing native linguistic properties. With nearly 700 million speakers, the SEA region lacks a region-specific embedding benchmark. We introduce SEA-BED, the first large-scale SEA embedding benchmark with 169 datasets across 9 tasks and 10 languages, where 71% are formulated by humans, not machine generation or translation. We address three research questions: (1) which SEA languages and tasks are challenging, (2) whether SEA languages show unique performance gaps globally, and (3) how human vs. machine translations affect evaluation. We evaluate 17 embedding models across six studies, analyzing task and language challenges, cross-benchmark comparisons, and translation trade-offs. Results show sharp ranking shifts, inconsistent model performance among SEA languages, and the importance of human-curated datasets for low-resource languages like Burmese.
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
Guo, Haiyang, Zeng, Fanhu, Zhu, Fei, Wang, Jiayi, Wang, Xukai, Zhou, Jingang, Zhao, Hongbo, Liu, Wenzhuo, Ma, Shijie, Wang, Da-Han, Zhang, Xu-Yao, Liu, Cheng-Lin
The rapid advancement of generative models has empowered modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specific domains. However, these models are fundamentally constrained by \emph{catastrophic forgetting}, \ie~a persistent challenge where models experience performance degradation on previously learned tasks when adapting to new tasks. To address this practical limitation, numerous approaches have been proposed to enhance the adaptability and scalability of generative AI in real-world applications. In this work, we present a comprehensive survey of continual learning methods for mainstream generative AI models, encompassing large language models, multimodal large language models, vision-language-action models, and diffusion models. Drawing inspiration from the memory mechanisms of the human brain, we systematically categorize these approaches into three paradigms: architecture-based, regularization-based, and replay-based methods, while elucidating their underlying methodologies and motivations. We further analyze continual learning setups for different generative models, including training objectives, benchmarks, and core backbones, thereby providing deeper insights into the field. The project page of this paper is available at https://github.com/Ghy0501/Awesome-Continual-Learning-in-Generative-Models.
Self-Correcting Code Generation Using Small Language Models
Cho, Jeonghun, Kang, Deokhyung, Kim, Hyounghun, Lee, Gary Geunbae
Self-correction has demonstrated potential in code generation by allowing language models to revise and improve their outputs through successive refinement. Recent studies have explored prompting-based strategies that incorporate verification or feedback loops using proprietary models, as well as training-based methods that leverage their strong reasoning capabilities. However, whether smaller models possess the capacity to effectively guide their outputs through self-reflection remains unexplored. Our findings reveal that smaller models struggle to exhibit reflective revision behavior across both self-correction paradigms. In response, we introduce CoCoS, an approach designed to enhance the ability of small language models for multi-turn code correction. Specifically, we propose an online reinforcement learning objective that trains the model to confidently maintain correct outputs while progressively correcting incorrect outputs as turns proceed. Our approach features an accumulated reward function that aggregates rewards across the entire trajectory and a fine-grained reward better suited to multi-turn correction scenarios. This facilitates the model in enhancing initial response quality while achieving substantial improvements through self-correction. With 1B-scale models, CoCoS achieves improvements of 35.8% on the MBPP and 27.7% on HumanEval compared to the baselines.
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
Rahman, Salman, Jiang, Liwei, Shiffer, James, Liu, Genglin, Issaka, Sheriff, Parvez, Md Rizwan, Palangi, Hamid, Chang, Kai-Wei, Choi, Yejin, Gabriel, Saadia
Multi-turn interactions with language models (LMs) pose critical safety risks, as harmful intent can be strategically spread across exchanges. Yet, the vast majority of prior work has focused on single-turn safety, while adaptability and diversity remain among the key challenges of multi-turn red-teaming. To address these challenges, we present X-Teaming, a scalable framework that systematically explores how seemingly harmless interactions escalate into harmful outcomes and generates corresponding attack scenarios. X-Teaming employs collaborative agents for planning, attack optimization, and verification, achieving state-of-the-art multi-turn jailbreak effectiveness and diversity with success rates up to 98.1% across representative leading open-weight and closed-source models. In particular, X-Teaming achieves a 96.2% attack success rate against the latest Claude 3.7 Sonnet model, which has been considered nearly immune to single-turn attacks. Building on X-Teaming, we introduce XGuard-Train, an open-source multi-turn safety training dataset that is 20x larger than the previous best resource, comprising 30K interactive jailbreaks, designed to enable robust multi-turn safety alignment for LMs. Our work offers essential tools and insights for mitigating sophisticated conversational attacks, advancing the multi-turn safety of LMs.
Schools' safety tools are spying on kids -- even at home
A new system called Scanary uses AI and radar to scan up to 25,000 people an hour. School is back in session, but here's something no one told you at orientation: Your kids may have more eyes on them than just their teachers'. Even if you don't have kids in school, you really need to know about this. A new study from UC San Diego uncovered what's really going on with those student safety tools schools buy. You know, the ones that are supposed to stop bullying, flag mental health struggles and prevent school shootings?
Parameter-Free Logit Distillation via Sorting Mechanism
Knowledge distillation (KD) aims to distill the knowledge from the teacher (larger) to the student (smaller) model via soft-label for the efficient neural network. In general, the performance of a model is determined by accuracy, which is measured with labels. However, existing KD approaches usually use the teacher with its original distribution, neglecting the potential of incorrect prediction. This may contradict the motivation of hard-label learning through cross-entropy loss, which may lead to sub-optimal knowledge distillation on certain samples. To address this issue, we propose a novel logit processing scheme via a sorting mechanism. Specifically, our method has a two-fold goal: (1) fixing the incorrect prediction of the teacher based on the labels and (2) reordering the distribution in a natural way according to priority rank at once. As an easy-to-use, plug-and-play pre-processing, our sort method can be effectively applied to existing logit-based KD methods. Extensive experiments on the CIFAR-100 and ImageNet datasets demonstrate the effectiveness of our method.
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
Seegmiller, Parker, Mehta, Kartik, Saha, Soumya, Tao, Chenyang, Oraby, Shereen, Gupta, Arpit, Chung, Tagyoung, Bansal, Mohit, Peng, Nanyun
Recent works improving LLM math reasoning with synthetic data have used unique setups, making comparison of data synthesis strategies impractical. This leaves many unanswered questions about the roles of different factors in the synthetic data pipeline, such as the impact of filtering low-quality problems. To address this gap, we introduce FLAMES, a Framework for LLM Assessment of Math rEasoning Data Synthesis, and perform a systematic study of 10 existing data synthesis strategies and multiple other factors impacting the performance of synthetic math reasoning data. Our FLAMES experiments provide several valuable insights about the optimal balance of difficulty and diversity of synthetic data. First, data agents designed to increase problem complexity lead to best improvements on most math metrics. Second, with a fixed data generation budget, keeping higher problem coverage is more important than keeping only problems with reliable solutions. Third, GSM8K- and MATH-based synthetic data can lead to improvements on competition-level benchmarks, showcasing easy-to-hard generalization. Leveraging insights from our FLAMES experiments, we design two novel data synthesis strategies for improving out-of-domain generalization and robustness. Further, we develop the FLAMES dataset, an effective blend of our novel and existing data synthesis strategies, outperforming public datasets on OlympiadBench (+15.7), CollegeMath (+4.5), GSMPlus (+6.5), and MATH (+3.1). Fine-tuning Qwen2.5-Math-7B on the FLAMES dataset achieves 81.4% on MATH, surpassing larger Llama3 405B, GPT-4o and Claude 3.5 Sonnet.