Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
Kim, Joongwon, Paranjape, Bhargavi, Khot, Tushar, Hajishirzi, Hannaneh
–arXiv.org Artificial Intelligence
Language agents perform complex tasks by using tools to execute each step precisely. However, most existing agents are based on proprietary models or designed to target specific tasks, such as mathematics or multi-hop question answering. We introduce Husky, a holistic, open-source language agent that learns to reason over a unified action space to address a diverse set of complex tasks involving numerical, tabular, and knowledge-based reasoning. Husky iterates between two stages: 1) generating the next action to take towards solving a given task and 2) executing the action using expert models and updating the current solution state. We identify a thorough ontology of actions for addressing complex tasks and curate high-quality data to train expert models for executing these actions. Our experiments show that Husky outperforms prior language agents across 14 evaluation datasets. Moreover, we introduce HuskyQA, a new evaluation set which stress tests language agents for mixed-tool reasoning, with a focus on retrieving missing knowledge and performing numerical reasoning. Despite using 7B models, Husky matches or even exceeds frontier LMs such as GPT-4 on these tasks, showcasing the efficacy of our holistic approach in addressing complex reasoning problems. Our code and models are available at https://github.com/agent-husky/Husky-v1.
arXiv.org Artificial Intelligence
Jun-10-2024
- Country:
- Africa > Rwanda
- Antarctica (0.04)
- Asia
- Indonesia > Bali (0.04)
- Middle East
- Israel
- Haifa District > Haifa (0.04)
- Jerusalem District > Jerusalem (0.04)
- Southern District > Beersheba (0.04)
- Jordan (0.04)
- Syria > Damascus Governorate
- Damascus (0.04)
- UAE > Abu Dhabi Emirate
- Abu Dhabi (0.04)
- Israel
- Singapore (0.04)
- South Korea > Incheon
- Incheon (0.04)
- Europe
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Serbia > Central Serbia
- Belgrade (0.04)
- Netherlands > Gelderland
- Nijmegen (0.04)
- Lithuania (0.04)
- France (0.04)
- Norway > Norwegian Sea (0.04)
- United Kingdom > England
- Leicestershire (0.04)
- Nottinghamshire (0.04)
- Germany (0.04)
- Poland (0.04)
- Austria
- Salzburg > Salzburg (0.04)
- Upper Austria (0.04)
- Spain > Galicia
- Madrid (0.04)
- Belgium > Brussels-Capital Region
- North America
- Canada > Ontario
- Toronto (0.04)
- Central America (0.04)
- Dominican Republic (0.04)
- Mexico
- Oaxaca (0.04)
- Veracruz > Coatzacoalcos (0.04)
- Yucatán (0.04)
- United States
- Massachusetts (0.04)
- Wyoming (0.04)
- Pennsylvania (0.04)
- Colorado (0.04)
- Washington > King County
- Seattle (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- Ohio (0.04)
- Tennessee (0.04)
- Illinois > Cook County
- Chicago (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Arizona (0.04)
- New York (0.05)
- California > San Francisco County
- San Francisco (0.04)
- Maryland > Baltimore (0.04)
- Oklahoma > Tulsa County
- Tulsa (0.04)
- Maine (0.04)
- Indiana > Marion County
- Indianapolis (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Texas > Harris County
- Houston (0.04)
- Canada > Ontario
- Oceania
- Australia (0.04)
- New Zealand (0.04)
- South America
- Brazil (0.04)
- Chile > Santiago Metropolitan Region
- Santiago Province > Santiago (0.04)
- Genre:
- Research Report > New Finding (0.67)
- Workflow (0.95)
- Industry:
- Leisure & Entertainment > Sports
- Media
- Transportation
- Ground
- Infrastructure & Services (0.93)
- Passenger (1.00)
- Banking & Finance (1.00)
- Education (1.00)
- Health & Medicine > Therapeutic Area (0.67)
- Government
- Law (1.00)
- Automobiles & Trucks > Manufacturer (1.00)
- Technology: