Goto

Collaborating Authors

 Personal


Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization

arXiv.org Artificial Intelligence

In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where each event proposal consists of a set of short video segments, we extract visual feature (e.g., C3D feature) from each segment and use the existing image/video captioning approach to generate one sentence description for this segment. Considering that the generated sentences contain rich semantic descriptions about the whole event proposal, we formulate the dense video captioning task as a visual cue aided sentence summarization problem and propose a new two stage Long Short Term Memory (LSTM) approach equipped with a new hierarchical attention mechanism to summarize all generated sentences as one descriptive sentence with the aid of visual features. Specifically, the first-stage LSTM network takes all semantic words from the generated sentences and the visual features from all segments within one event proposal as the input, and acts as the encoder to effectively summarize both semantic and visual information related to this event proposal. The second-stage LSTM network takes the output from the first-stage LSTM network and the visual features from all video segments within one event proposal as the input, and acts as the decoder to generate one descriptive sentence for this event proposal. Our comprehensive experiments on the ActivityNet Captions dataset demonstrate the effectiveness of our newly proposed DaS framework for dense video captioning.


RoboCupRescue: an interview with Adam Jacoff

AIHub

RoboCup is an international scientific initiative with the goal of advancing the state of the science of intelligent robots, AI and automation. The annual RoboCup event will take place from 15-21 July in Salvador, Brazil. The RoboCupRescue League is an important element of the competition and focuses on the challenges involved in search and rescue applications. We caught up with Adam Jacoff, co-founder of the RoboCupRescue league, former RoboCup Trustee, and chair of the organising committee, to find out more. The RoboCupRescue League is now in its 25th year hosting competitions and workshops all around the world.


From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking

arXiv.org Artificial Intelligence

Recent progress in autonomous code generation has fueled excitement around AI agents capable of accelerating scientific discovery by running experiments. However, there is currently no benchmark that evaluates whether such agents can implement scientific ideas when given varied amounts of code as a starting point, interpolating between reproduction (running code) and from-scratch replication (fully re-implementing and running code). We introduce AutoExperiment, a benchmark that evaluates AI agents' ability to implement and run machine learning experiments based on natural language descriptions in research papers. In each task, agents are given a research paper, a codebase with key functions masked out, and a command to run the experiment. The goal is to generate the missing code, execute the experiment in a sandboxed environment, and reproduce the results. AutoExperiment scales in difficulty by varying the number of missing functions $n$, ranging from partial reproduction to full replication. We evaluate state-of-the-art agents and find that performance degrades rapidly as $n$ increases. Agents that can dynamically interact with the environment (e.g. to debug their code) can outperform agents in fixed "agentless" harnesses, and there exists a significant gap between single-shot and multi-trial success rates (Pass@1 vs. Pass@5), motivating verifier approaches to our benchmark. Our findings highlight critical challenges in long-horizon code generation, context retrieval, and autonomous experiment execution, establishing AutoExperiment as a new benchmark for evaluating progress in AI-driven scientific experimentation. Our data and code are open-sourced at https://github.com/j1mk1m/AutoExperiment .


Making optimal decisions without having all the cards in hand

AIHub

The article "Revelations: A Decidable Class of POMDP with Omega-Regular Objectives" won an Outstanding Paper Award at the AAAI 2025 conference, a prestigious international conference about artificial intelligence. This year, only three papers received such an award out of 3,000 accepted and 12,000 submitted! This recognition crowns the results of research initiated in Bordeaux (France) within the Synthรจse team at the Bordeaux Computer Science Research Laboratory (LaBRI), where four of the authors work: Marius Belly, Nathanaรซl Fijalkow, Hugo Gimbert, and Pierre Vandenhove. The work also involved researchers from Paris (Florian Horn) and Antwerp (Guillermo A. Pรฉrez). The article is freely available on arXiv, and this post outlines its main ideas.


StereoTacTip: Vision-based Tactile Sensing with Biomimetic Skin-Marker Arrangements

arXiv.org Artificial Intelligence

Chenghua Lu received the B.S. degree in Mechanical Engineering from Northeastern University, Shenyang, China, in 2017, and the M.S. degree in Mechanical Manufacturing and Automation from the University of Chinese Academy of Sciences, Beijing, China, in 2021. She is currently working toward the Ph.D. degree majoring in Engineering Mathematics with the School of Mathematics Engineering and Technology and Bristol Robotics Laboratory, University of Bristol, Bristol, UK. Her research interests include tactile sensing and soft robotics. Kailuan T ang received a B.S. degree in Communication Engineering from the Southern University of Science and Technology (SUSTech), Shenzhen, China in 2017. He is currently working towards a Ph.D. degree majoring in Mechanics with the School of Mechatronics Engineering, Harbin Institute of Technology.


Former Scale AI CEO Alexandr Wang on AI's Potential and Its 'Deficiencies'

TIME - Tech

On June 12, Alexandr Wang stepped down as Scale's CEO to chase his most ambitious moonshot yet: building smarter-than-human AI as head of Meta's new "superintelligence" division. As part of his move, Meta will invest 14.3 billion for a minority stake in Scale AI, but the real prize isn't his company--it's Wang himself. Wang, 28, is expected to bring a sense of urgency to Meta's AI efforts, which this year have been plagued by delays and underwhelming performance. Once the undisputed leader of open-weight AI, the U.S. tech giant has been overtaken by Chinese rivals like DeepSeek on popular benchmarks. Although Wang, who dropped out of MIT at 19, lacks the academic chops of some of his peers, he offers both insight into the types of data Meta's rivals use to improve their AI systems, and unrivaled ambition.


This May Be Trump's Most Consequential Decision Yet

Slate

This week, Emily Bazelon, John Dickerson, and David Plotz discuss whether the US should join Israel's war on Iran, the tragic Minnesota assassinations and why US political violence is surging now, and the Supreme Court's unsurprising but willfully obtuse decision to uphold Tennessee's youth transgender care ban. Here are some notes and references from this week's show: Alexander Ward, Lara Seligman, and Dustin Volz for The Wall Street Journal (Exclusive): Israel Built Its Case for War With Iran on New Intelligence. The U.S. Didn't Buy It. Thomas L. Friedman for The New York Times (Opinion): The Smart Way for Trump to End the Israel-Iran War Oren Cass for Understanding America (Substack): Is Israel the Ideal "America First" Ally? Warren P. Strobel, Alex Horton, and Abigail Hauslohner for the Washington Post: Navigating Iran crisis, Trump relies on experience over star power Amy Howe for SCOTUSblog: Court upholds Tennessee's ban on certain medical treatments for transgender minors Abbie VanSickle for The New York Times: Sotomayor Writes the Court'Abandons' Transgender Children to'Political Whims' Ella Lee for The Hill: Clarence Thomas urges courts to end deferring to'experts' on gender-affirming care Ian Millhiser for Vox: The Supreme Court's incoherent new attack on trans rights, explained Here are this week's chatters: Emily: A Family Matter by Claire Lynch; The Fall of Affirmative Action: Race, the Supreme Court, and the Future of Higher Education by Justin Driver; A Flower Traveled in My Blood: The Incredible True Story of the Grandmothers Who Fought to Find a Stolen Generation of Children by Haley Cohen Gilliland. John: Mary Cunningham for CBS News: Federal Reserve holds its benchmark interest rate steady at today's FOMC meeting; ABA Banking Journal: Fed's Powell says some areas of U.S. may be'uninsurable' in next decade David: Trip Gabriel for the New York Times: William Langewiesche, the'Steve McQueen of Journalism,' Dies at 70 For this week's Slate Plus bonus episode, Emily, John, and David discuss the exciting possibilities and likely limitations of using AI tools for historical research and writing.


The Hardness of Achieving Impact in AI for Social Impact Research: A Ground-Level View of Challenges & Opportunities

arXiv.org Artificial Intelligence

In an attempt to tackle the UN SDGs, AI for Social Impact (AI4SI) projects focus on harnessing AI to address societal issues in areas such as healthcare, social justice, etc. Unfortunately, despite growing interest in AI4SI, achieving tangible, on-the-ground impact remains a significant challenge. For example, identifying and engaging motivated collaborators who are willing to co-design and deploy AI based solutions in real-world settings is often difficult. Even when such partnerships are established, many AI4SI projects "fail" to progress beyond the proof-of-concept stage, and hence, are unable to transition to at-scale production-level solutions. Furthermore, the unique challenges faced by AI4SI researchers are not always fully recognized within the broader AI community, where such work is sometimes viewed as primarily applied and not aligning with the traditional criteria for novelty emphasized in core AI venues. This paper attempts to shine a light on the diverse challenges faced in AI4SI research by diagnosing a multitude of factors that prevent AI4SI partnerships from achieving real-world impact on the ground. Drawing on semi-structured interviews with six leading AI4SI researchers - complemented by the authors' own lived experiences in conducting AI4SI research - this paper attempts to understand the day-to-day difficulties faced in developing and deploying socially impactful AI solutions. Through thematic analysis, we identify structural and organizational, communication, collaboration, and operational challenges as key barriers to deployment. While there are no easy fixes, we synthesize best practices and actionable strategies drawn from these interviews and our own work in this space. In doing so, we hope this paper serves as a practical reference guide for AI4SI researchers and partner organizations seeking to engage more effectively in socially impactful AI collaborations.


Mapping Caregiver Needs to AI Chatbot Design: Strengths and Gaps in Mental Health Support for Alzheimer's and Dementia Caregivers

arXiv.org Artificial Intelligence

Family caregivers of individuals with Alzheimer's Disease and Related Dementia (AD/ADRD) face significant emotional and logistical challenges that place them at heightened risk for stress, anxiety, and depression. Although recent advances in generative AI -- particularly large language models (LLMs) -- offer new opportunities to support mental health, little is known about how caregivers perceive and engage with such technologies. To address this gap, we developed Carey, a GPT-4o-based chatbot designed to provide informational and emotional support to AD/ADRD caregivers. Using Carey as a technology probe, we conducted semi-structured interviews with 16 family caregivers following scenario-driven interactions grounded in common caregiving stressors. Through inductive coding and reflexive thematic analysis, we surface a systemic understanding of caregiver needs and expectations across six themes -- on-demand information access, emotional support, safe space for disclosure, crisis management, personalization, and data privacy. For each of these themes, we also identified the nuanced tensions in the caregivers' desires and concerns. We present a mapping of caregiver needs, AI chatbot's strengths, gaps, and design recommendations. Our findings offer theoretical and practical insights to inform the design of proactive, trustworthy, and caregiver-centered AI systems that better support the evolving mental health needs of AD/ADRD caregivers.


Amazon Rebuilt Alexa Using a 'Staggering' Amount of AI Tools

WIRED

Daniel Rausch, Amazon's vice president of Alexa and Echo, is in the midst of a major transition. More than a decade beyond the launch of Amazon's Alexa, he's been tasked with creating a new version of the marquee voice assistant, one that's powered by large language models. As he put it in my interview with him, this new assistant, dubbed Alexa, is "a complete rebuild of the architecture." How did his team approach Amazon's largest ever revamp of its voice assistant? They used AI to build AI, of course.