Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL Queries
Wang, Zhengren, Yao, Dongwen, Li, Bozhou, Ma, Dongsheng, Li, Bo, Li, Zhiyu, Xiong, Feiyu, Cui, Bin, Tang, Linpeng, Zhang, Wentao
–arXiv.org Artificial Intelligence
The proliferation of unstructured data poses a fundamental challenge to traditional database interfaces. While Text-to-SQL has democratized access to structured data, it remains incapable of interpreting semantic or multi-modal queries. Concurrently, vector search has emerged as the de facto standard for querying unstructured data, but its integration with SQL-termed VectorSQL-still relies on manual query crafting and lacks standardized evaluation methodologies, creating a significant gap between its potential and practical application. To bridge this fundamental gap, we introduce and formalize Text2VectorSQL, a novel task to establish a unified natural language interface for seamlessly querying both structured and unstructured data. To catalyze research in this new domain, we present a comprehensive foundational ecosystem, including: (1) A scalable and robust pipeline for synthesizing high-quality Text-to-VectorSQL training data. (2) VectorSQLBench, the first large-scale, multi-faceted benchmark for this task, encompassing 12 distinct combinations across three database backends (SQLite, PostgreSQL, ClickHouse) and four data sources (BIRD, Spider, arXiv, Wikipedia). (3) Several novel evaluation metrics designed for more nuanced performance analysis. Extensive experiments not only confirm strong baseline performance with our trained models, but also reveal the recall degradation challenge: the integration of SQL filters with vector search can lead to more pronounced result omissions than in conventional filtered vector search. By defining the core task, delivering the essential data and evaluation infrastructure, and identifying key research challenges, our work lays the essential groundwork to build the next generation of unified and intelligent data interfaces. Our repository is available at https://github.com/OpenDCAI/Text2VectorSQL.
arXiv.org Artificial Intelligence
Nov-7-2025
- Country:
- Asia
- Afghanistan > Parwan Province
- Charikar (0.04)
- China > Shanghai
- Shanghai (0.04)
- Middle East > UAE
- Abu Dhabi Emirate > Abu Dhabi (0.14)
- Afghanistan > Parwan Province
- Europe
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Spain (0.04)
- Belgium > Brussels-Capital Region
- North America
- Dominican Republic (0.04)
- United States
- Georgia > Fulton County
- Atlanta (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Georgia > Fulton County
- Asia
- Genre:
- Overview (0.46)
- Research Report (0.50)
- Technology:
- Information Technology
- Artificial Intelligence
- Machine Learning > Neural Networks
- Deep Learning (1.00)
- Natural Language
- Chatbot (1.00)
- Information Retrieval > Query Processing (0.94)
- Large Language Model (1.00)
- Text Processing (0.93)
- Representation & Reasoning (1.00)
- Machine Learning > Neural Networks
- Databases (1.00)
- Information Management (1.00)
- Artificial Intelligence
- Information Technology