How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
Liu, Ryan, Sumers, Theodore R., Dasgupta, Ishita, Griffiths, Thomas L.
–arXiv.org Artificial Intelligence
In day-to-day communication, people often approximate the truth -- for example, rounding the By treating these objectives as separate, standard approaches time or omitting details -- in order to be maximally implicitly assume that honesty and helpfulness are jointly helpful to the listener. How do large language achievable. In everyday conversation, however, they can be models (LLMs) handle such nuanced tradeoffs? in tension (Wilson & Sperber, 2002). People regularly round To address this question, we use psychological numerical values like time (Van Der Henst et al., 2002), models and experiments designed to characterize distance (Krifka, 2007), and monetary values (Kao et al., human behavior to analyze LLMs. We test a 2014), or endorse false generalizations (van Rooij & Schulz, range of LLMs and explore how optimization for 2020) when they believe these approximations will help the human preferences or inference-time reasoning listener. People also regularly trade literal honesty for other affects these trade-offs. We find that reinforcement communicative goals, as implicit in figures of speech like learning from human feedback improves metaphors (Tendahl & Gibbs Jr, 2008), hyperbole (Carston both honesty and helpfulness, while chain-ofthought & Wearing, 2011) and irony (Popa-Wyatt, 2019).
arXiv.org Artificial Intelligence
Feb-11-2024