Personal
Where Did the Royals Go So Wrong With Kate Middleton? It's Been Years in the Making.
This article was originally featured in Foreign Policy, the magazine of global politics and ideas. A family snap of the Princess of Wales with her three children has dominated headlines and group chats since its release on U.K. Mother's Day last weekend. Princess Catherine, whom the palace says is recovering from a January abdominal surgery, is known chiefly for never putting a foot wrong during nearly two decades of intense public scrutiny--first as the girlfriend of Prince William, then as a wife and mother to future kings, and an advocate for uncontroversial but important causes, such as early childhood development. Yet, even for a woman defined by her seeming perfection--Hilary Mantel once wrote that the former duchess appeared to have been designed by a committee and built by craftsmen--the Mother's Day photo of Catherine and her family was judged to be a little too perfect. The uncanny valley of the photo was prime territory for conspiracy theories, already circulating, that the princess is missing or perhaps even dead.
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
Mascarell, Laura, Chalumattu, Ribin, Rios, Annette
The advent of Large Language Models (LLMs) has led to remarkable progress on a wide range of natural language processing tasks. Despite the advances, these large-sized models still suffer from hallucinating information in their output, which poses a major issue in automatic text summarization, as we must guarantee that the generated summary is consistent with the content of the source document. Previous research addresses the challenging task of detecting hallucinations in the output (i.e. inconsistency detection) in order to evaluate the faithfulness of the generated summaries. However, these works primarily focus on English and recent multilingual approaches lack German data. This work presents absinth, a manually annotated dataset for hallucination detection in German news summarization and explores the capabilities of novel open-source LLMs on this task in both fine-tuning and in-context learning settings. We open-source and release the absinth dataset to foster further research on hallucination detection in German.
Evaluating LLMs for Gender Disparities in Notable Persons
Rhue, Lauren, Goethals, Sofie, Sundararajan, Arun
This study examines the use of Large Language Models (LLMs) for retrieving factual information, addressing concerns over their propensity to produce factually incorrect "hallucinated" responses or to altogether decline to even answer prompt at all. Specifically, it investigates the presence of gender-based biases in LLMs' responses to factual inquiries. This paper takes a multi-pronged approach to evaluating GPT models by evaluating fairness across multiple dimensions of recall, hallucinations and declinations. Our findings reveal discernible gender disparities in the responses generated by GPT-3.5. While advancements in GPT-4 have led to improvements in performance, they have not fully eradicated these gender disparities, notably in instances where responses are declined. The study further explores the origins of these disparities by examining the influence of gender associations in prompts and the homogeneity in the responses.
On STPA for Distributed Development of Safe Autonomous Driving: An Interview Study
Nouri, Ali, Berger, Christian, Tรถrner, Fredrik
Safety analysis is used to identify hazards and build knowledge during the design phase of safety-relevant functions. This is especially true for complex AI-enabled and software intensive systems such as Autonomous Drive (AD). System-Theoretic Process Analysis (STPA) is a novel method applied in safety-related fields like defense and aerospace, which is also becoming popular in the automotive industry. However, STPA assumes prerequisites that are not fully valid in the automotive system engineering with distributed system development and multi-abstraction design levels. This would inhibit software developers from using STPA to analyze their software as part of a bigger system, resulting in a lack of traceability. This can be seen as a maintainability challenge in continuous development and deployment (DevOps). In this paper, STPA's different guidelines for the automotive industry, e.g. J31887/ISO21448/STPA handbook, are firstly compared to assess their applicability to the distributed development of complex AI-enabled systems like AD. Further, an approach to overcome the challenges of using STPA in a multi-level design context is proposed. By conducting an interview study with automotive industry experts for the development of AD, the challenges are validated and the effectiveness of the proposed approach is evaluated.
A Conversational Brain-Artificial Intelligence Interface
Meunier, Anja, ลฝรกk, Michal Robert, Munz, Lucas, Garkot, Sofiya, Eder, Manuel, Xu, Jiachen, Grosse-Wentrup, Moritz
We introduce Brain-Artificial Intelligence Interfaces (BAIs) as a new class of Brain-Computer Interfaces (BCIs). Unlike conventional BCIs, which rely on intact cognitive capabilities, BAIs leverage the power of artificial intelligence to replace parts of the neuro-cognitive processing pipeline. BAIs allow users to accomplish complex tasks by providing high-level intentions, while a pre-trained AI agent determines low-level details. This approach enlarges the target audience of BCIs to individuals with cognitive impairments, a population often excluded from the benefits of conventional BCIs. We present the general concept of BAIs and illustrate the potential of this new approach with a Conversational BAI based on EEG. In particular, we show in an experiment with simulated phone conversations that the Conversational BAI enables complex communication without the need to generate language. Our work thus demonstrates, for the first time, the ability of a speech neuroprosthesis to enable fluent communication in realistic scenarios with non-invasive technologies.
FARPLS: A Feature-Augmented Robot Trajectory Preference Labeling System to Assist Human Labelers' Preference Elicitation
Lyu, Hanfang, Bai, Yuanchen, Liang, Xin, Das, Ujaan, Shi, Chuhan, Gong, Leiliang, Li, Yingchi, Sun, Mingfei, Ge, Ming, Ma, Xiaojuan
Preference-based learning aims to align robot task objectives with human values. One of the most common methods to infer human preferences is by pairwise comparisons of robot task trajectories. Traditional comparison-based preference labeling systems seldom support labelers to digest and identify critical differences between complex trajectories recorded in videos. Our formative study (N = 12) suggests that individuals may overlook non-salient task features and establish biased preference criteria during their preference elicitation process because of partial observations. In addition, they may experience mental fatigue when given many pairs to compare, causing their label quality to deteriorate. To mitigate these issues, we propose FARPLS, a Feature-Augmented Robot trajectory Preference Labeling System. FARPLS highlights potential outliers in a wide variety of task features that matter to humans and extracts the corresponding video keyframes for easy review and comparison. It also dynamically adjusts the labeling order according to users' familiarities, difficulties of the trajectory pair, and level of disagreements. At the same time, the system monitors labelers' consistency and provides feedback on labeling progress to keep labelers engaged. A between-subjects study (N = 42, 105 pairs of robot pick-and-place trajectories per person) shows that FARPLS can help users establish preference criteria more easily and notice more relevant details in the presented trajectories than the conventional interface. FARPLS also improves labeling consistency and engagement, mitigating challenges in preference elicitation without raising cognitive loads significantly
From Text to Self: Users' Perceptions of Potential of AI on Interpersonal Communication and Self
Fu, Yue, Foell, Sami, Xu, Xuhai, Hiniker, Alexis
In the rapidly evolving landscape of AI-mediated communication (AIMC), tools powered by Large Language Models (LLMs) are becoming integral to interpersonal communication. Employing a mixed-methods approach, we conducted a one-week diary and interview study to explore users' perceptions of these tools' ability to: 1) support interpersonal communication in the short-term, and 2) lead to potential long-term effects. Our findings indicate that participants view AIMC support favorably, citing benefits such as increased communication confidence, and finding precise language to express their thoughts, navigating linguistic and cultural barriers. However, the study also uncovers current limitations of AIMC tools, including verbosity, unnatural responses, and excessive emotional intensity. These shortcomings are further exacerbated by user concerns about inauthenticity and potential overreliance on the technology. Furthermore, we identified four key communication spaces delineated by communication stakes (high or low) and relationship dynamics (formal or informal) that differentially predict users' attitudes toward AIMC tools. Specifically, participants found the tool is more suitable for communicating in formal relationships than informal ones and more beneficial in high-stakes than low-stakes communication.
Stephen Salter obituary
Stephen Salter, who has died aged 85, was the inventor of the Salter's Duck, a wave-power device that was the first of its kind and promised to provide a new source of renewable energy for the world โ until it was effectively killed off by the nuclear industry. In 1982, after eight years of development under Salter's direction at Edinburgh University, the United Kingdom Atomic Energy Authority (UKAEA) was asked by the government to see if the duck might be a cost-effective way of making large quantities of electricity. To the great surprise of Salter, and others, the UKAEA came to the conclusion that it was uneconomic, and that no further government funding should be given to the project. A decade later it emerged that thanks to a misplaced decimal point, the review had made Salter's duck look 10 times more expensive than the experiments showed it was likely to be. The UKAEA claimed this was just a mistake, but Salter, who had never been allowed to see the results of the secret evaluation, put it another way: asking the nuclear industry to evaluate an alternative source of energy was like putting King Herod in charge of a children's home, he suggested.
Happy International Women's Day!
To celebrate International Women's Day, we take a look back over the past 12 months and highlight some of the women we've interviewed and featured, and who've written about their research on AIhub. Elizabeth Ondula is an Electrical Engineer from the Technical University of Kenya and is currently a PhD student of Computer Science at USC. She is a member of the Autonomous Networks Research Group, and co-organizes a bi-weekly reinforcement learning group, SUITERS-RL. Prior to academia, she had roles as a Software Engineer at IBM Research in Kenya, Head of Product Development of Brave Venture Labs and Co-lead of Hardware Research at iHub Nairobi. We interviewed Elizabeth as part of our series featuring the AAAI Doctoral Consortium participants.
Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive Learning
Lin, Bingqian, Long, Yanxin, Zhu, Yi, Zhu, Fengda, Liang, Xiaodan, Ye, Qixiang, Lin, Liang
Vision-and-language navigation (VLN) asks an agent to follow a given language instruction to navigate through a real 3D environment. Despite significant advances, conventional VLN agents are trained typically under disturbance-free environments and may easily fail in real-world scenarios, since they are unaware of how to deal with various possible disturbances, such as sudden obstacles or human interruptions, which widely exist and may usually cause an unexpected route deviation. In this paper, we present a model-agnostic training paradigm, called Progressive Perturbation-aware Contrastive Learning (PROPER) to enhance the generalization ability of existing VLN agents, by requiring them to learn towards deviation-robust navigation. Specifically, a simple yet effective path perturbation scheme is introduced to implement the route deviation, with which the agent is required to still navigate successfully following the original instruction. Since directly enforcing the agent to learn perturbed trajectories may lead to inefficient training, a progressively perturbed trajectory augmentation strategy is designed, where the agent can self-adaptively learn to navigate under perturbation with the improvement of its navigation performance for each specific trajectory. For encouraging the agent to well capture the difference brought by perturbation, a perturbation-aware contrastive learning mechanism is further developed by contrasting perturbation-free trajectory encodings and perturbation-based counterparts. Extensive experiments on R2R show that PROPER can benefit multiple VLN baselines in perturbation-free scenarios. We further collect the perturbed path data to construct an introspection subset based on the R2R, called Path-Perturbed R2R (PP-R2R). The results on PP-R2R show unsatisfying robustness of popular VLN agents and the capability of PROPER in improving the navigation robustness.