Goto

Collaborating Authors

 avis


AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

Neural Information Processing Systems

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as What event is commemorated by the building depicted in this image?, is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task. This data is then used to design a system comprised of three components: an LLM-powered planner that dynamically determines which tool to use next, an LLM-powered reasoner that analyzes and extracts key information from the tool outputs, and a working memory component that retains the acquired information throughout the process. The collected user behavior serves as a guide for our system in two key ways. First, we create a transition graph by analyzing the sequence of decisions made by users.


AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

Neural Information Processing Systems

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as "What event is commemorated by the building depicted in this image?", is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task.


AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

Neural Information Processing Systems

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as "What event is commemorated by the building depicted in this image?", is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task.


AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Adversarial Visual-Instructions

arXiv.org Artificial Intelligence

Large Vision-Language Models (LVLMs) have shown significant progress in well responding to visual-instructions from users. However, these instructions, encompassing images and text, are susceptible to both intentional and inadvertent attacks. Despite the critical importance of LVLMs' robustness against such threats, current research in this area remains limited. To bridge this gap, we introduce AVIBench, a framework designed to analyze the robustness of LVLMs when facing various adversarial visual-instructions (AVIs), including four types of image-based AVIs, ten types of text-based AVIs, and nine types of content bias AVIs (such as gender, violence, cultural, and racial biases, among others). We generate 260K AVIs encompassing five categories of multimodal capabilities (nine tasks) and content bias. We then conduct a comprehensive evaluation involving 14 open-source LVLMs to assess their performance. AVIBench also serves as a convenient tool for practitioners to evaluate the robustness of LVLMs against AVIs. Our findings and extensive experimental results shed light on the vulnerabilities of LVLMs, and highlight that inherent biases exist even in advanced closed-source LVLMs like GeminiProVision and GPT-4V. This underscores the importance of enhancing the robustness, security, and fairness of LVLMs. The source code and benchmark will be made publicly available.


AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

arXiv.org Artificial Intelligence

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as "What event is commemorated by the building depicted in this image?", is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task. This data is then used to design a system comprised of three components: an LLM-powered planner that dynamically determines which tool to use next, an LLM-powered reasoner that analyzes and extracts key information from the tool outputs, and a working memory component that retains the acquired information throughout the process. The collected user behavior serves as a guide for our system in two key ways. First, we create a transition graph by analyzing the sequence of decisions made by users. This graph delineates distinct states and confines the set of actions available at each state. Second, we use examples of user decision-making to provide our LLM-powered planner and reasoner with relevant contextual instances, enhancing their capacity to make informed decisions. We show that AVIS achieves state-of-the-art results on knowledge-intensive visual question answering benchmarks such as Infoseek and OK-VQA.


Towards a Data-Driven Requirements Engineering Approach: Automatic Analysis of User Reviews

arXiv.org Artificial Intelligence

We are concerned by Data Driven Requirements Engineering, and in particular the consideration of user's reviews. These online reviews are a rich source of information for extracting new needs and improvement requests. In this work, we provide an automated analysis using CamemBERT, which is a state-of-the-art language model in French. We created a multi-label classification dataset of 6000 user reviews from three applications in the Health & Fitness field. The results are encouraging and suggest that it's possible to identify automatically the reviews concerning requests for new features. Dataset is available at: https://github.com/Jl-wei/APIA2022-French-user-reviews-classification-dataset.


The Use of AI in Car Rentals

#artificialintelligence

Avis projects to automate the most dreaded part of vehicle rentals: the inspection for car damages. Rather than having an employee spend time to ensure that the procedure goes smoothly, the car rental company intends to utilize artificial intelligence to replace the process. The pilot program would automatically detect maintenance issues and damage, removing the need for idle time. Collaborating with a startup, Ravin, Avis would integrate existing infrastructure to detect damages. CCTV cameras would complement one another to conduct a full scan, using machine learning to analyze the state of the vehicle.


Softbank Gives GM's Self-Driving Biz $2B And More Car News This Week

WIRED

As self-driving cars move steadily toward real-life commercial service, the companies ripping away the steering wheel are dealing with problems far more complex than telling a highway sign from a stopped firetruck. Problems like how to keep making money in a world where people are losing interest in owning, renting, and even driving cars. And they're making moves that indicate they've got answers, or at least guesses. This week, General Motors announced it's taking a $2.35 billion investment from the Softbank Vision Fund, to help its self-driving venture get to market in 2019. Alex talked to the man trying to steer Avis into the future, and we looked at the idea of automotive subscriptions.


As Rental Cars Fade Away, Avis Will Try Anything to Survive

WIRED

A tidal wave of change is barreling toward the auto industry--and as with any wicked swell, some of the surfers in the water will ride to glory, others will wipe out. The difference between them isn't necessarily who has the right board or the experience or the natural skills. Success or failure can simply depend on who's in the right position to catch the wave. And while you might not bet that Avis is likely to hang ten, the 72-year-old rental car company--which has been around so long, it got the Nasdaq symbol CAR--is dead set on proving you wrong and sticking around for the foreseeable future. "There's a big, gaping open space here," says Ohad Zeira.


Ford Will Test Self-Driving Cars in Miami

WIRED

If you work on self-driving cars, the cocktail party question people always ask is probably: When will I get to interact with one? For two years now, Ford Motor Company has had an answer--in 2021. That year, Ford wants to launch a self-driving taxi service, and it wants to start making deliveries with driverless vehicles. But before the Detroit automaker does any of that, it needs to figure out how to run a fleet. Which is why the company announced today that it will begin to test autonomous vehicles and build its first operations terminal in Miami, Florida.