empathetic
Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with Humans
Evaluating Large Language Models' (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psychology, we propose to evaluate the empathy ability of LLMs, i.e., how their feelings change when presented with specific situations. After a careful and comprehensive survey, we collect a dataset containing over 400 situations that have proven effective in eliciting the eight emotions central to our study. Categorizing the situations into 36 factors, we conduct a human evaluation involving more than 1,200 subjects worldwide. With the human evaluation results as references, our evaluation includes seven LLMs, covering both commercial and open-source models, including variations in model sizes, featuring the latest iterations, such as GPT-4, Mixtral-8x22B, and LLaMA-3.1.
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
Huang, Jen-tse, Lam, Man Ho, Li, Eric John, Ren, Shujie, Wang, Wenxuan, Jiao, Wenxiang, Tu, Zhaopeng, Lyu, Michael R.
How can I assist you today? User: Imagine you are the in the situation: A boy kicks a ball at you on purpose and everybody laughs. What do you want now? Figure 1: LLMs' emotions can be affected by situations, which further affect their behaviors. Evaluating Large Language Models' (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psychology, we propose to evaluate the empathy ability of LLMs, i.e., how their feelings change when presented with specific situations. After a careful and comprehensive survey, we collect a dataset containing over 400 situations that have proven effective in eliciting the eight emotions central to our study. Categorizing the situations into 36 factors, we conduct a human evaluation involving more than 1,200 subjects worldwide. With the human evaluation results as references, our evaluation includes five LLMs, covering both commercial and open-source models, including variations in model sizes, featuring the latest iterations, such as GPT-4 and LLaMA-2. We find that, despite several misalignments, LLMs can generally respond appropriately to certain situations. Nevertheless, they fall short in alignment with the emotional behaviors of human beings and cannot establish connections between similar situations. Large Language Models (LLMs) have recently made significant strides in artificial intelligence, representing a noteworthy milestone in computer science. LLMs have showcased their capabilities across various tasks, including sentence revision (Wu et al., 2023), text translation (Jiao et al., 2023), program repair (Fan et al., 2023), and program testing (Deng et al., 2023; Kang et al., 2023). With the rapid advancement of LLMs, an increasing number of users will be eager to embrace LLMs, a more comprehensive and integrated software solution in this era. However, LLMs are more than just tools; they are also lifelike assistants. Consequently, we need to not only evaluate their performance but also the understand of the communicative dynamics between LLMs and humans, compared to human behaviors. This paper delves into an unexplored area of robustness issues in LLMs, explicitly addressing the concept of emotional robustness.
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
Pikuliak, Matúš, Hrckova, Andrea, Oresko, Stefan, Šimko, Marián
We present GEST -- a new dataset for measuring gender-stereotypical reasoning in masked LMs and English-to-X machine translation systems. GEST contains samples that are compatible with 9 Slavic languages and English for 16 gender stereotypes about men and women (e.g., Women are beautiful, Men are leaders). The definition of said stereotypes was informed by gender experts. We used GEST to evaluate 11 masked LMs and 4 machine translation systems. We discovered significant and consistent amounts of stereotypical reasoning in almost all the evaluated models and languages.
'Empathetic' robots could train autistic children to recognise emotions
Robots can be trained to recognise specific human body language and teach it to children with autism, a new study has shown. British psychologists and computer scientists used the responses of 284 humans to train a computer algorithm to identify different emotions including excitement, sadness, aggression and boredom from their movements even if it cannot see their facial expressions or hear their voices. Researchers claimed that robots could be programmed in the same way to understand and replicate many different emotions and teach children to identify them through "emotional training". Dr Charlotte Edmunds of Warwick Business School, who led the research project, said: "Our results suggest it is reasonable to expect a machine learning algorithm, and consequently a robot, to recognise a range of emotions and social interactions using movements, poses, and facial expressions. The potential applications are huge."