Goto

Collaborating Authors

 Personal


The Robotability Score: Enabling Harmonious Robot Navigation on Urban Streets

arXiv.org Artificial Intelligence

This paper introduces the Robotability Score ($R$), a novel metric that quantifies the suitability of urban environments for autonomous robot navigation. Through expert interviews and surveys, we identify and weigh key features contributing to R for wheeled robots on urban streets. Our findings reveal that pedestrian density, crowd dynamics and pedestrian flow are the most critical factors, collectively accounting for 28% of the total score. Computing robotability across New York City yields significant variation; the area of highest R is 3.0 times more "robotable" than the area of lowest R. Deployments of a physical robot on high and low robotability areas show the adequacy of the score in anticipating the ease of robot navigation. This new framework for evaluating urban landscapes aims to reduce uncertainty in robot deployment while respecting established mobility patterns and urban planning principles, contributing to the discourse on harmonious human-robot environments.


"All Roads Lead to ChatGPT": How Generative AI is Eroding Social Interactions and Student Learning Communities

arXiv.org Artificial Intelligence

The widespread adoption of generative AI is already impacti ng learning and help-seeking. While the benefits of generative AI are well-understood, recent studies have also raised concernsabout increased potential for cheating and negative impacts on stud ents' metacognition and critical thinking. However, the potenti al impacts on social interactions, peer learning, and classroom dynamics are not yet well understood. To investigate these aspect s, we conducted 17 semi-structured interviews with undergraduate computing students across seven R1 universities in NorthAmerica. Our findings suggest that help-seeking requests are now often me di-ated by generative AI. For example, students often redirected questions from their peers to generative AI instead of providing assistance themselves, undermining peer interaction. Students also reported feeling increasingly isolated and demotivated as th e social support systems they rely on begin to break down. These findings are concerning given the important role that social interac tions play in students' learning and sense of belonging.


Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

arXiv.org Artificial Intelligence

Children can acquire language from less than 100 million words of input. Large language models are far less data-efficient: they typically require 3 or 4 orders of magnitude more data and still do not perform as well as humans on many evaluations. These intensive resource demands limit the ability of researchers to train new models and use existing models as developmentally plausible cognitive models. The BabyLM Challenge is a communal effort in which participants compete to optimize language model training on a fixed data budget. Submissions are compared on various evaluation tasks targeting grammatical ability, downstream task performance, and generalization. Participants can submit to up to three tracks with progressively looser data restrictions. From over 30 submissions, we extract concrete recommendations on how best to train data-efficient language models, and on where future efforts should (and perhaps should not) focus. The winning submissions using the LTG-BERT architecture (Samuel et al., 2023) outperformed models trained on trillions of words. Other submissions achieved strong results through training on shorter input sequences or training a student model on a pretrained teacher. Curriculum learning attempts, which accounted for a large number of submissions, were largely unsuccessful, though some showed modest improvements.


How em The Last of Us /em Fans Turned Against Its Breakout Star

Slate

By pretty much every objective measure, HBO's adaptation of the hit postapocalyptic video game The Last of Us has been a roaring success. Never before has a video game narrative been molded into Emmy nominations and such warm reception among respectable critics, industry darlings, and people who have no idea what the term "one-shotting" means. You'd think that the devotees who first fell in love with the game back when it was originally released in 2013 would be toasting the cultural ascendance of their favorite medium--and especially how the story's complicated morality has impacted those who've never picked up a controller. And yet, for as long as the show has been on television, its most dogmatic fans have been caught up in a controversy of much inferior consequence: Specifically, they're furious that Bella Ramsey doesn't look much like Ellie. On the most basic level, this observation is correct.


Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning

arXiv.org Artificial Intelligence

Large language models (LLMs) have shown substantial capacity for generating fluent, contextually appropriate responses. However, they can produce hallucinated outputs, especially when a user query includes one or more false premises-claims that contradict established facts. Such premises can mislead LLMs into offering fabricated or misleading details. Existing approaches include pretraining, fine-tuning, and inference-time techniques that often rely on access to logits or address hallucinations after they occur. These methods tend to be computationally expensive, require extensive training data, or lack proactive mechanisms to prevent hallucination before generation, limiting their efficiency in real-time applications. We propose a retrieval-based framework that identifies and addresses false premises before generation. Our method first transforms a user's query into a logical representation, then applies retrieval-augmented generation (RAG) to assess the validity of each premise using factual sources. Finally, we incorporate the verification results into the LLM's prompt to maintain factual consistency in the final output. Experiments show that this approach effectively reduces hallucinations, improves factual accuracy, and does not require access to model logits or large-scale fine-tuning.


Dynamic Evaluation Framework for Personalized and Trustworthy Agents: A Multi-Session Approach to Preference Adaptability

arXiv.org Artificial Intelligence

Recent advancements in generative AI have significantly increased interest in personalized agents. With increased personalization, there is also a greater need for being able to trust decision-making and action taking capabilities of these agents. However, the evaluation methods for these agents remain outdated and inadequate, often failing to capture the dynamic and evolving nature of user interactions. In this conceptual article, we argue for a paradigm shift in evaluating personalized and adaptive agents. We propose a comprehensive novel framework that models user personas with unique attributes and preferences. In this framework, agents interact with these simulated users through structured interviews to gather their preferences and offer customized recommendations. These recommendations are then assessed dynamically using simulations driven by Large Language Models (LLMs), enabling an adaptive and iterative evaluation process. Our flexible framework is designed to support a variety of agents and applications, ensuring a comprehensive and versatile evaluation of recommendation strategies that focus on proactive, personalized, and trustworthy aspects.


The Hall of AI Fears and Hopes: Comparing the Views of AI Influencers and those of Members of the U.S. Public Through an Interactive Platform

arXiv.org Artificial Intelligence

AI development is shaped by academics and industry leaders - let us call them ``influencers'' - but it is unclear how their views align with those of the public. To address this gap, we developed an interactive platform that served as a data collection tool for exploring public views on AI, including their fears, hopes, and overall sense of hopefulness. We made the platform available to 330 participants representative of the U.S. population in terms of age, sex, ethnicity, and political leaning, and compared their views with those of 100 AI influencers identified by Time magazine. The public fears AI getting out of control, while influencers emphasize regulation, seemingly to deflect attention from their alleged focus on monetizing AI's potential. Interestingly, the views of AI influencers from underrepresented groups such as women and people of color often differ from the views of underrepresented groups in the public.


Towards Smarter Hiring: Are Zero-Shot and Few-Shot Pre-trained LLMs Ready for HR Spoken Interview Transcript Analysis?

arXiv.org Artificial Intelligence

This research paper presents a comprehensive analysis of the performance of prominent pre-trained large language models (LLMs), including GPT-4 Turbo, GPT-3.5 Turbo, text-davinci-003, text-babbage-001, text-curie-001, text-ada-001, llama-2-7b-chat, llama-2-13b-chat, and llama-2-70b-chat, in comparison to expert human evaluators in providing scores, identifying errors, and offering feedback and improvement suggestions to candidates during mock HR (Human Resources) interviews. We introduce a dataset called HURIT (Human Resource Interview Transcripts), which comprises 3,890 HR interview transcripts sourced from real-world HR interview scenarios. Our findings reveal that pre-trained LLMs, particularly GPT-4 Turbo and GPT-3.5 Turbo, exhibit commendable performance and are capable of producing evaluations comparable to those of expert human evaluators. Although these LLMs demonstrate proficiency in providing scores comparable to human experts in terms of human evaluation metrics, they frequently fail to identify errors and offer specific actionable advice for candidate performance improvement in HR interviews. Our research suggests that the current state-of-the-art pre-trained LLMs are not fully conducive for automatic deployment in an HR interview assessment. Instead, our findings advocate for a human-in-the-loop approach, to incorporate manual checks for inconsistencies and provisions for improving feedback quality as a more suitable strategy.


2025 Hugo Award game finalists include Zelda: Echoes of Wisdom and Dragon Age: The Veilguard

Engadget

The Hugo Awards began honoring video games for the first time back in 2021. This week, the organization revealed the list of six finalists for the 2025 awards ceremony. Let's go over the nominations. Two AAA titles are up for the award. The gameplay involves summoning monsters and items to solve puzzles and do battle.


Humanoid robot stuns with perfect side-flip acrobatics

FOX News

A robotics company has advanced from a backflipping robot to a side-flipping robot. Robots aren't just efficient machines anymore, they are now agile performers that can flip and jog. Take, for instance, Unitree, a Chinese robotics company that has been making headlines with its incredible G1 humanoid robot. You might have seen it dancing alongside humans or remembered its predecessor, the H1, which stunned us with a backflip using electric motors. But now, the G1 has taken things to a whole new level.