Personal
Can you pass that tool?: Implications of Indirect Speech in Physical Human-Robot Collaboration
Zhang, Yan, Ratnayake, Tharaka Sachintha, Sew, Cherie, Knibbe, Jarrod, Goncalves, Jorge, Johal, Wafa
Indirect speech acts (ISAs) are a natural pragmatic feature of human communication, allowing requests to be conveyed implicitly while maintaining subtlety and flexibility. Although advancements in speech recognition have enabled natural language interactions with robots through direct, explicit commands--providing clarity in communication--the rise of large language models presents the potential for robots to interpret ISAs. However, empirical evidence on the effects of ISAs on human-robot collaboration (HRC) remains limited. To address this, we conducted a Wizard-of-Oz study (N=36), engaging a participant and a robot in collaborative physical tasks. Our findings indicate that robots capable of understanding ISAs significantly improve human's perceived robot anthropomorphism, team performance, and trust. However, the effectiveness of ISAs is task- and context-dependent, thus requiring careful use. These results highlight the importance of appropriately integrating direct and indirect requests in HRC to enhance collaborative experiences and task performance.
Personality Structured Interview for Large Language Model Simulation in Personality Research
Wang, Pengda, Zou, Huiqi, Chen, Hanjie, Sun, Tianjun, Xiao, Ziang, Oswald, Frederick L.
Although psychometrics researchers have recently explored the use of large language models (LLMs) as proxies for human participants, LLMs often fail to generate heterogeneous data with human-like diversity, which diminishes their value in advancing social science research. To address these challenges, we explored the potential of the theory-informed Personality Structured Interview (PSI) as a tool for simulating human responses in personality research. In this approach, the simulation is grounded in nuanced real-human interview transcripts that target the personality construct of interest. We have provided a growing set of 357 structured interview transcripts from a representative sample, each containing an individual's response to 32 open-ended questions carefully designed to gather theory-based personality evidence. Additionally, grounded in psychometric research, we have summarized an evaluation framework to systematically validate LLM-generated psychometric data. Results from three experiments demonstrate that well-designed structured interviews could improve human-like heterogeneity in LLM-simulated personality data and predict personality-related behavioral outcomes (i.e., organizational citizenship behaviors and counterproductive work behavior). We further discuss the role of theory-informed structured interviews in LLM-based simulation and outline a general framework for designing structured interviews to simulate human-like data for psychometric research.
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control
Su, Jinyan, Healey, Jennifer, Nakov, Preslav, Cardie, Claire
Retrieval-Augmented Generation (RAG) has emerged as a powerful approach to mitigate large language model (LLM) hallucinations by incorporating external knowledge retrieval. However, existing RAG frameworks often apply retrieval indiscriminately,leading to inefficiencies-over-retrieving when unnecessary or failing to retrieve iteratively when required for complex reasoning. Recent adaptive retrieval strategies, though adaptively navigates these retrieval strategies, predict only based on query complexity and lacks user-driven flexibility, making them infeasible for diverse user application needs. In this paper, we introduce a novel user-controllable RAG framework that enables dynamic adjustment of the accuracy-cost trade-off. Our approach leverages two classifiers: one trained to prioritize accuracy and another to prioritize retrieval efficiency. Via an interpretable control parameter $\alpha$, users can seamlessly navigate between minimal-cost retrieval and high-accuracy retrieval based on their specific requirements. We empirically demonstrate that our approach effectively balances accuracy, retrieval cost, and user controllability, making it a practical and adaptable solution for real-world applications.
Viktor Antonov, art director for Half-Life 2 and Dishonored, has died, according to colleagues
Viktor Antonov, best known for his work as art lead on Half-Life 2 and Dishonored, has reportedly died at age 52. Half-Life writer Marc Laidlaw broke the news in an Instagram Story, and other colleagues have since taken to social media to pay tribute as well. "I didn't want to say much till I felt it was confirmed, but I learned today that Viktor Antonov, our visionary art lead on HL2, has died," Laidla wrote in the now-expired post, which was reshared by LambdaGeneration on Saturday night. Antonov got his start in video games working on Redneck Rampage, and in addition to serving as art director for Half-Life 2 and Dishonored, he went on to consult on titles including Doom (2016) and Fallout 4. The Bulgarian artist just recently appeared in a documentary celebrating the 20th anniversary of Half-life 2 this past November. I wish I told you how much admiration I had for you but we get caught in our lives until a surprise like this hits us," Raphael Colantonio, founder of Arkane Studios and Wolfeye Studios, wrote on Bluesky. "You were instrumental to the success of Arkane Studios and an inspiration to many of us, also a friend with whom I have many fond memories." In another post, game designer Harvey Smith added, "All this about his impact and talent is true, but I will also always remember how much he made me laugh, with his dry, devastating wit.
Why Amazon Web Services CEO Matt Garman Is Playing the Long Game on AI
Matt Garman took the helm at Amazon Web Services (AWS), the cloud computing arm of the U.S. tech giant, in June, but he joined the business around 19 years ago as an intern. He went on to become AWS's first product manager and helped to build and launch many of its core services, before eventually becoming the CEO last year. Like many other tech companies, AWS, which is Amazon's most profitable unit, is betting big on AI. In April 2023, the company launched Amazon Bedrock, which gives cloud customers access to foundation models built by AI companies including Anthropic and Mistral. At its re:Invent conference in Las Vegas in December, the AWS made a series of announcements, including a new generation of foundation AI models, called Nova. It also said that it's building one of the world's most powerful AI supercomputers with Anthropic, which it has a strategic partnership with, using a giant cluster of AWS's Trainium 2 training chips. TIME spoke with Garman a few days after the re:Invent conference, about his AI ambitions, how he's thinking about ensuring the technology is safe, and how the company is balancing its energy needs with its emissions targets.
Crash victims honoured at basketball matches
Four students killed in a car crash were honoured at a university as basketball matches resumed for the first time since the incident. Makyle Bayley, 22, Eva Darold-Tchikaya, 21, Anthony "TJ" Hibbert, 24 and Daljang Wol, 22, died when a car crashed into a building on Magdalen Street, Colchester on 1 February. Mr Hibbert and Mr Wol played for the Essex Rebels, who dedicated Saturday's fixtures to the victims and held an applause in their memory. University of Essex director of sport Dave Parry said: "We've lost four really loved members of our university and sporting community, who gave so much to their friends and others." Mr Bayley was a member of the British Universities and Colleges Sport (BUCS) basketball team, while Ms Darold-Tchikaya was a member of the Essex Blades dance club and other societies.Dawid Wojtowicz/BBCSaturday's basketball fixtures at the University of Essex were dedicated to the victimsDawid Wojtowicz/BBCIt was the first time matches had been played there since the incident Last week, more than 1,000 people including students, staff and relatives of the victims attended a gathering.
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
Rao, Rajath, Ganesan, Adithya, Kjell, Oscar, Luby, Jonah, Raghavan, Akshay, Feltman, Scott, Ringwald, Whitney, Boyd, Ryan L., Luft, Benjamin, Ruggero, Camilo, Ryant, Neville, Kotov, Roman, Schwartz, H. Andrew
Current speech encoding pipelines often rely on an additional text-based LM to get robust representations of human communication, even though SotA speech-to-text models often have a LM within. This work proposes an approach to improve the LM within an audio model such that the subsequent text-LM is unnecessary. We introduce WhiSPA (Whisper with Semantic and Psychological Alignment), which leverages a novel audio training objective: contrastive loss with a language model embedding as a teacher. Using over 500k speech segments from mental health audio interviews, we evaluate the utility of aligning Whisper's latent space with semantic representations from a text autoencoder (SBERT) and lexically derived embeddings of basic psychological dimensions: emotion and personality. Over self-supervised affective tasks and downstream psychological tasks, WhiSPA surpasses current speech encoders, achieving an average error reduction of 73.4% and 83.8%, respectively. WhiSPA demonstrates that it is not always necessary to run a subsequent text LM on speech-to-text output in order to get a rich psychological representation of human communication.
RAS: Retrieval-And-Structuring for Knowledge-Intensive LLM Generation
Jiang, Pengcheng, Cao, Lang, Zhu, Ruike, Jiang, Minhao, Zhang, Yunyi, Sun, Jimeng, Han, Jiawei
Retrieval-augmented language models often struggle with knowledge-intensive tasks due to inefficient retrieval, unstructured knowledge integration, and single-pass architectures. We present Retrieval-And-Structuring (RAS), a novel framework that dynamically constructs and reasons over query-specific knowledge graphs through iterative retrieval and structuring. RAS introduces four key technical innovations: (1) a themescoped retrieval mechanism that efficiently narrows the search space while maintaining retrieval quality, (2) an action planning module that determines knowledge needs and generates focused sub-queries, (3) a dynamic knowledge structuring approach that converts retrieved text into an evolving knowledge graph, and (4) a graph-augmented answering component that leverages the accumulated structured information. Our framework achieves state-of-the-art performance, surpassing leading baselines by 6.4% with open-source language models and 7.0% with proprietary models on seven knowledge-intensive generation datasets across all evaluation metrics. Detailed ablation studies verify the contribution of each technical component to the overall system performance.
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
Lu, Pan, Chen, Bowen, Liu, Sheng, Thapa, Rahul, Boen, Joseph, Zou, James
Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning. Existing methods augment large language models (LLMs) with external tools but are restricted to specialized domains, limited tool types, or require additional training data. In this paper, we introduce OctoTools, a training-free, user-friendly, and easily extensible open-source agentic framework designed to tackle complex reasoning across diverse domains. OctoTools introduces standardized tool cards to encapsulate tool functionality, a planner for both high-level and low-level planning, and an executor to carry out tool usage. We validate OctoTools' generality across 16 diverse tasks (including MathVista, MMLU-Pro, MedQA, and GAIA-Text), achieving substantial average accuracy gains of 9.3% over GPT-4o. Furthermore, OctoTools outperforms AutoGen, GPT-Functions and LangChain by up to 10.6% when given the same set of tools. Through comprehensive analysis and ablations, OctoTools demonstrates advantages in task planning, effective tool usage, and multi-step problem solving.
FairFare: A Tool for Crowdsourcing Rideshare Data to Empower Labor Organizers
Calacci, Dana, Rao, Varun Nagaraj, Dalal, Samantha, Di, Catherine, Pua, Kok-Wei, Schwartz, Andrew, Spitzberg, Danny, Monroy-Hernández, Andrés
In recent years, labor organizers representing rideshare and delivery workers have advocated for regulations to improve working conditions in the rideshare industry that set wage floors and job loss protections [67]. To call for these improvements, organizers need to understand workers' existing conditions [37], a significant data access and social computing challenge in the rideshare industry. Labor organizers representing rideshare workers typically rely on a collage of qualitative anecdotes and screenshots to provide data about existing working conditions [24]. While these qualitative data provide rich, "thick descriptions" [30] of workers' experience, they are often dismissed by platforms as non-representative, cherry-picked examples. Rideshare platforms, on the other hand, have exclusive access to large-scale, comprehensive quantitative datasets of driver, trip, and pay data that they can draw upon to create authoritative narratives about working conditions in their industry [72]. Labor organizers need comprehensive access to large-scale quantitative data describing working conditions to conduct rigorous, independent investigations and contest platform-driven narratives. There are tools and legal frameworks that empower individual rideshare workers to independently access quantitative work data (e.g., Gridwise and Data Subject Access Requests). However, these tools and frameworks do not provide an intuitive way to aggregate individual worker data into a dataset that provides collective insight into overarching working conditions. Algorithmic auditing scholarship provides methods, like crowdsourcing data, to independently investigate black-boxed systems [66].