Goto

Collaborating Authors

 hesitation


Understanding Cognitive States from Head & Hand Motion Data

arXiv.org Artificial Intelligence

The pipeline illustrates the full workflow from data collection in VR, through self-annotation and human baseline evaluation, to modeling and analysis of cognitive states. As virtual reality (VR) and augmented reality (AR) continue to gain popularity, head and hand motion data captured by consumer VR systems have become ubiquitous. Prior work shows such telemetry can be highly identifying and reflect broad user traits, often aligning with intuitive "folk theories" of body language. However, it remains unclear to what extent motion kinematics encode more nuanced cognitive states, such as confusion, hesitation, and readiness, which lack clear correlates with motion. To investigate this, we introduce a novel dataset of head and hand motion with frame-level annotations of these states collected during structured decision-making tasks. Our findings suggest that deep temporal models can infer subtle cognitive states from motion alone, achieving comparable performance with human observers. This work demonstrates that standard VR telemetry contains strong patterns related to users' internal cognitive processes, which opens the door for a new gener- To enhance reproducibility and support future work, we will make our dataset and modeling framework publicly available. Virtual Reality (VR) is rapidly evolving from a specialized tool for simulation and entertainment into a mainstream computing platform for work, education, and social interaction. As users spend more time in these immersive environments, the quality of human-computer interaction becomes paramount. The next generation of VR systems must move beyond explicit, command-based interfaces and develop the capacity for implicit, nuanced understanding. This requires an ability to perceive and adapt to a user's cognitive state in real-time, creating experiences that are more intuitive, supportive, and effective. The key to unlocking this capability lies in decoding the rich, continuous, and often subconscious stream of motion data generated by every user.


Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems

arXiv.org Artificial Intelligence

Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitations, losing important acoustic details. We fine-tune Whisper models on the Speak & Improve 2025 corpus using low-rank adaptation (LoRA), without recourse to external audio training data. We compare three annotation schemes: removing hesitations (Pure), generic tags (Rich), and acoustically precise fillers inferred by Gemini 2.0 Flash from existing audio-transcript pairs (Extra). Our challenge system achieved 6.47% WER (Pure) and 5.81% WER (Extra). Post-challenge experiments reveal that fine-tuning Whisper Large V3 Turbo with the "Extra" scheme yielded a 5.5% WER, an 11.3% relative improvement over the "Pure" scheme (6.2% WER). This demonstrates that explicit, realistic filled-pause labeling significantly enhances ASR accuracy for verbatim L2 speech transcription.


Leveraging Cognitive States for Adaptive Scaffolding of Understanding in Explanatory Tasks in HRI

arXiv.org Artificial Intelligence

-- Understanding how scaffolding strategies influence human understanding in human-robot interaction is important for developing effective assistive systems. This empirical study investigates linguistic scaffolding strategies based on negation as an important means that de-biases the user from potential errors but increases processing costs and hesitations as a means to ameliorate processing costs. In an adaptive strategy, the user state with respect to the current state of understanding and processing capacity was estimated via a scoring scheme based on task performance, prior scaffolding strategy, and current eye gaze behavior . In the study, the adaptive strategy of providing negations and hesitations was compared with a nonadaptive strategy of providing only affirmations. The adaptive scaffolding strategy was generated using the computational model SHIFT . Our findings indicate that using adaptive scaffolding strategies with SHIFT tends to (1) increased processing costs, as reflected in longer reaction times, but (2) improved task understanding, evidenced by a lower error rate of almost 23%. We assessed the efficiency of SHIFT's selected scaffolding strategies across different cognitive states, finding that in three out of five states, the error rate was lower compared to the baseline condition. We discuss how these results align with the assumptions of the SHIFT model and highlight areas for refinement. Moreover, we demonstrate how scaffolding strategies, such as negation and hesitation, contribute to more effective human-robot explanatory dialogues. In the growing field of social robotics, robots are increasingly being designed to assist people in their everyday lives.


SHIFT: An Interdisciplinary Framework for Scaffolding Human Attention and Understanding in Explanatory Tasks

arXiv.org Artificial Intelligence

In this work, we present a domain-independent approach for adaptive scaffolding in robotic explanation generation to guide tasks in human-robot interaction. We present a method for incorporating interdisciplinary research results into a computational model as a pre-configured scoring system implemented in a framework called SHIFT. This involves outlining a procedure for integrating concepts from disciplines outside traditional computer science into a robotics computational framework. Our approach allows us to model the human cognitive state into six observable states within the human partner model. To study the pre-configuration of the system, we implement a reinforcement learning approach on top of our model. This approach allows adaptation to individuals who deviate from the configuration of the scoring system. Therefore, in our proof-of-concept evaluation, the model's adaptability on four different user types shows that the models' adaptation performs better, i.e., recouped faster after exploration and has a higher accumulated reward with our pre-configured scoring system than without it. We discuss further strategies of speeding up the learning phase to enable a realistic adaptation behavior to real users. The system is accessible through docker and supports querying via ROS.


An Active Inference Agent for Simulating Human Translation Processes in a Hierarchical Architecture: Integrating the Task Segment Framework and the HOF taxonomy

arXiv.org Artificial Intelligence

In this paper, we propose modelling human translation production as a hierarchy of three embedded translation processes. The proposed architecture replicates the temporal dynamics of keystroke production across sensorimotor, cognitive, and phenomenal layers. Utilizing data from the CRITT TPR-DB, the Task Segment Framework, and the HOF taxonomy, we demonstrate the temporal breakdown of the typing flow on distinct timelines within these three layers.


SH2: Self-Highlighted Hesitation Helps You Decode More Truthfully

arXiv.org Artificial Intelligence

Large language models (LLMs) demonstrate great performance in text generation. However, LLMs are still suffering from hallucinations. In this work, we propose an inference-time method, Self-Highlighted Hesitation (SH2), to help LLMs decode more truthfully. SH2 is based on a simple fact rooted in information theory that for an LLM, the tokens predicted with lower probabilities are prone to be more informative than others. Our analysis shows that the tokens assigned with lower probabilities by an LLM are more likely to be closely related to factual information, such as nouns, proper nouns, and adjectives. Therefore, we propose to ''highlight'' the factual information by selecting the tokens with the lowest probabilities and concatenating them to the original context, thus forcing the model to repeatedly read and hesitate on these tokens before generation. During decoding, we also adopt contrastive decoding to emphasize the difference in the output probabilities brought by the hesitation. Experimental results demonstrate that our SH2, requiring no additional data or models, can effectively help LLMs elicit factual knowledge and distinguish hallucinated contexts. Significant and consistent improvements are achieved by SH2 for LLaMA-7b and LLaMA2-7b on multiple hallucination tasks.


Adapting an ASR Foundation Model for Spoken Language Assessment

arXiv.org Artificial Intelligence

A crucial part of an accurate and reliable spoken language assessment system is the underlying ASR model. Recently, large-scale pre-trained ASR foundation models such as Whisper have been made available. As the output of these models is designed to be human readable, punctuation is added, numbers are presented in Arabic numeric form and abbreviations are included. Additionally, these models have a tendency to skip disfluencies and hesitations in the output. Though useful for readability, these attributes are not helpful for assessing the ability of a candidate and providing feedback. Here a precise transcription of what a candidate said is needed. In this paper, we give a detailed analysis of Whisper outputs and propose two solutions: fine-tuning and soft prompt tuning. Experiments are conducted on both public speech corpora and an English learner dataset. Results show that we can effectively alter the decoding behaviour of Whisper to generate the exact words spoken in the response.


A.I. Bots Can't Report This Column. But They Can Improve It. - The New York Times

#artificialintelligence

For several days, I tested Wordtune Spices and Rytr, two A.I. writing assistants released by start-ups in the last two years, and compared them with ChatGPT. I'll go over some examples that highlight the strengths and weaknesses of each of these three tools. To use Wordtune Spices, which the Israeli start-up AI21 Labs released last month, you insert text into a box, highlight the sentences you want edited and then click on options to make improvements. Among its best uses during my recent test were quick rewrites of sentences. When you browse the web, an increasing number of sites and apps are asking for a piece of basic information that you probably hand over without hesitation: your email address.


The Perks and Obstacles of AI Adoption in Insurance

#artificialintelligence

Imagine that you are a leader at an insurance company. You know that artificial intelligence (AI) will give you a competitive edge and have decided to invest. You hired two brilliant data scientists, Juana and Yash. Juana develops an AI solution that scans digitized customer files, mines them for relevant information, and calculates accurate pay-outs. You project savings of over $1 million in the next 2 years, and 30% increased staff productivity.


What is Artificial Intelligence?

#artificialintelligence

Constantly use of these products leads us to be dependent on them more and more for our every simple and complex tasks. Life now doesn't seem easier without them because of their easy access, speed of machines, working capacity and solving complex problem in just few mere seconds. Artificial Intelligence (AI), The developed branch of computer Science. Their job is to make the machines so advanced that the machine will itself able to take decisions further on their own without the interference of an external force. Hence, It becomes necessary for us to understand What is (AI) Artificial Intelligence?