Education
Domain Specific Question to SQL Conversion with Embedded Data Balancing Technique
Jyothi, null, Murthy, T. Satyanarayana
The rise of deep learning in natural language processing has fostered the creation of text to structured query language models composed of an encoder and a decoder. Researchers have experimented with various intermediate processing like schema linking, table type aware, value extract. To generate accurate SQL results for the user question. However error analysis performed on the failed cases on these systems shows, 29 percentage of the errors would be because the system was unable to understand the values expressed by the user in their question. This challenge affects the generation of accurate SQL queries, especially when dealing with domain-specific terms and specific value conditions, where traditional methods struggle to maintain consistency and precision. To overcome these obstacles, proposed two intermediations like implementing data balancing technique and over sampling domain-specific queries which would refine the model architecture to enhance value recognition and fine tuning the model for domain-specific questions. This proposed solution achieved 10.98 percentage improvement in accuracy of the model performance compared to the state of the art model tested on WikiSQL dataset. to convert the user question accurately to SQL queries. Applying oversampling technique on the domain-specific questions shown a significant improvement as compared with traditional approaches.
A Survey of Multimodal Retrieval-Augmented Generation
Mei, Lang, Mo, Siyu, Yang, Zhihan, Chen, Chong
Multimodal Retrieval-Augmented Generation (MRAG) enhances large language models (LLMs) by integrating multimodal data (text, images, videos) into retrieval and generation processes, overcoming the limitations of text-only Retrieval-Augmented Generation (RAG). While RAG improves response accuracy by incorporating external textual knowledge, MRAG extends this framework to include multimodal retrieval and generation, leveraging contextual information from diverse data types. This approach reduces hallucinations and enhances question-answering systems by grounding responses in factual, multimodal knowledge. Recent studies show MRAG outperforms traditional RAG, especially in scenarios requiring both visual and textual understanding. This survey reviews MRAG's essential components, datasets, evaluation methods, and limitations, providing insights into its construction and improvement. It also identifies challenges and future research directions, highlighting MRAG's potential to revolutionize multimodal information retrieval and generation. By offering a comprehensive perspective, this work encourages further exploration into this promising paradigm.
ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses
ExpertRAG is a novel theoretical framework that integrates Mixture-of-Experts (MoE) architectures with Retrieval Augmented Generation (RAG) to advance the efficiency and accuracy of knowledge-intensive language modeling. We propose a dynamic retrieval gating mechanism coupled with expert routing, enabling the model to selectively consult an external knowledge store or rely on specialized internal experts based on the query's needs. The paper lays out the theoretical foundations of ExpertRAG, including a probabilistic formulation that treats retrieval and expert selection as latent decisions, and mathematical justifications for its efficiency in both computation and knowledge utilization. We derive formulae to quantify the expected computational cost savings from selective retrieval and the capacity gains from sparse expert utilization. A comparative analysis positions ExpertRAG against standard RAG (with always-on retrieval) and pure MoE models (e.g., Switch Transformer, Mixtral) to highlight its unique balance between parametric knowledge and non-parametric retrieval. We also outline an experimental validation strategy, proposing benchmarks and evaluation protocols to test ExpertRAG's performance on factual recall, generalization, and inference efficiency. The proposed framework, although presented theoretically, is supported by insights from prior work in RAG and MoE, and is poised to provide more factual, efficient, and adaptive generation by leveraging the best of both paradigms. In summary, ExpertRAG contributes a new perspective on scaling and augmenting language models, backed by a thorough analysis and a roadmap for empirical validation.
Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization
Wu, Shujin, Qian, Cheng, Fung, Yi R., Liang, Paul Pu, Ji, Heng
The growing capabilities of large language models (LLMs) present a key challenge of maintaining effective human oversight. Weak-to-strong generalization (W2SG) offers a promising framework for supervising increasingly capable LLMs using weaker ones. Traditional W2SG methods rely on passive learning, where a weak teacher provides noisy demonstrations to train a strong student. This hinders students from employing their knowledge during training and reaching their full potential. In this work, we introduce Alice (pro{A}ctive {l}earning w{i}th tea{c}her's D{e}monstrations), a framework that leverages complementary knowledge between teacher and student to enhance the learning process. We probe the knowledge base of the teacher model by eliciting their uncertainty, and then use these insights together with teachers' responses as demonstrations to guide student models in self-generating improved responses for supervision. In addition, for situations with significant capability gaps between teacher and student models, we introduce cascade Alice, which employs a hierarchical training approach where weak teachers initially supervise intermediate models, who then guide stronger models in sequence. Experimental results demonstrate that our method significantly enhances the W2SG performance, yielding substantial improvements in three key tasks compared to the original W2SG: knowledge-based reasoning (+4.0%), mathematical reasoning (+22.62%), and logical reasoning (+12.11%). This highlights the effectiveness of our new W2SG paradigm that enables more robust knowledge transfer and supervision outcome.
Attention-Augmented Inverse Reinforcement Learning with Graph Convolutions for Multi-Agent Task Allocation
Yin, Huilin, Yang, Zhikun, Zhang, Linchuan, Watzenig, Daniel
This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. Multi-agent task allocation (MATA) plays a vital role in cooperative multi-agent systems, with significant implications for applications such as logistics, search and rescue, and robotic coordination. Although traditional deep reinforcement learning (DRL) methods have been shown to be promising, their effectiveness is hindered by a reliance on manually designed reward functions and inefficiencies in dynamic environments. In this paper, an inverse reinforcement learning (IRL)-based framework is proposed, in which multi-head self-attention (MHSA) and graph attention mechanisms are incorporated to enhance reward function learning and task execution efficiency. Expert demonstrations are utilized to infer optimal reward densities, allowing dependence on handcrafted designs to be reduced and adaptability to be improved. Extensive experiments validate the superiority of the proposed method over widely used multi-agent reinforcement learning (MARL) algorithms in terms of both cumulative rewards and task execution efficiency.
Command A: An Enterprise-Ready Large Language Model
Cohere, Team, :, null, Aakanksha, null, Ahmadian, Arash, Ahmed, Marwan, Alammar, Jay, Alizadeh, Milad, Alnumay, Yazeed, Althammer, Sophia, Arkhangorodsky, Arkady, Aryabumi, Viraat, Aumiller, Dennis, Avalos, Raphaรซl, Aviv, Zahara, Bae, Sammie, Baji, Saurabh, Barbet, Alexandre, Bartolo, Max, Bebensee, Bjรถrn, Beladia, Neeral, Beller-Morales, Walter, Bรฉrard, Alexandre, Berneshawi, Andrew, Bialas, Anna, Blunsom, Phil, Bobkin, Matt, Bongale, Adi, Braun, Sam, Brunet, Maxime, Cahyawijaya, Samuel, Cairuz, David, Campos, Jon Ander, Cao, Cassie, Cao, Kris, Castagnรฉ, Roman, Cendrero, Juliรกn, Currie, Leila Chan, Chandak, Yash, Chang, Diane, Chatziveroglou, Giannis, Chen, Hongyu, Cheng, Claire, Chevalier, Alexis, Chiu, Justin T., Cho, Eugene, Choi, Eugene, Choi, Eujeong, Chung, Tim, Cirik, Volkan, Cismaru, Ana, Clavier, Pierre, Conklin, Henry, Crawhall-Stein, Lucas, Crouse, Devon, Cruz-Salinas, Andres Felipe, Cyrus, Ben, D'souza, Daniel, Dalla-Torre, Hugo, Dang, John, Darling, William, Domingues, Omar Darwiche, Dash, Saurabh, Debugne, Antoine, Dehaze, Thรฉo, Desai, Shaan, Devassy, Joan, Dholakia, Rishit, Duffy, Kyle, Edalati, Ali, Eldeib, Ace, Elkady, Abdullah, Elsharkawy, Sarah, Ergรผn, Irem, Ermis, Beyza, Fadaee, Marzieh, Fan, Boyu, Fayoux, Lucas, Flet-Berliac, Yannis, Frosst, Nick, Gallรฉ, Matthias, Galuba, Wojciech, Garg, Utsav, Geist, Matthieu, Azar, Mohammad Gheshlaghi, Gilsenan-McMahon, Ellen, Goldfarb-Tarrant, Seraphina, Goldsack, Tomas, Gomez, Aidan, Gonzaga, Victor Machado, Govindarajan, Nithya, Govindassamy, Manoj, Grinsztajn, Nathan, Gritsch, Nikolas, Gu, Patrick, Guo, Shangmin, Haefeli, Kilian, Hajjar, Rod, Hawes, Tim, He, Jingyi, Hofstรคtter, Sebastian, Hong, Sungjin, Hooker, Sara, Hosking, Tom, Howe, Stephanie, Hu, Eric, Huang, Renjie, Jain, Hemant, Jain, Ritika, Jakobi, Nick, Jenkins, Madeline, Jordan, JJ, Joshi, Dhruti, Jung, Jason, Kalyanpur, Trushant, Kamalakara, Siddhartha Rao, Kedrzycki, Julia, Keskin, Gokce, Kim, Edward, Kim, Joon, Ko, Wei-Yin, Kocmi, Tom, Kozakov, Michael, Kryลciลski, Wojciech, Jain, Arnav Kumar, Teru, Komal Kumar, Land, Sander, Lasby, Michael, Lasche, Olivia, Lee, Justin, Lewis, Patrick, Li, Jeffrey, Li, Jonathan, Lin, Hangyu, Locatelli, Acyr, Luong, Kevin, Ma, Raymond, Mach, Lukรกลก, Machado, Marina, Magbitang, Joanne, Lopez, Brenda Malacara, Mann, Aryan, Marchisio, Kelly, Markham, Olivia, Matton, Alexandre, McKinney, Alex, McLoughlin, Dominic, Mokry, Jozef, Morisot, Adrien, Moulder, Autumn, Moynehan, Harry, Mozes, Maximilian, Muppalla, Vivek, Murakhovska, Lidiya, Nagarajan, Hemangani, Nandula, Alekhya, Nasir, Hisham, Nehra, Shauna, Netto-Rosen, Josh, Ohashi, Daniel, Owers-Bardsley, James, Ozuzu, Jason, Padilla, Dennis, Park, Gloria, Passaglia, Sam, Pekmez, Jeremy, Penstone, Laura, Piktus, Aleksandra, Ploeg, Case, Poulton, Andrew, Qi, Youran, Raghvendra, Shubha, Ramos, Miguel, Ranjan, Ekagra, Richemond, Pierre, Robert-Michon, Cรฉcile, Rodriguez, Aurรฉlien, Roy, Sudip, Ruder, Sebastian, Ruis, Laura, Rust, Louise, Sachan, Anubhav, Salamanca, Alejandro, Saravanakumar, Kailash Karthik, Satyakam, Isha, Sebag, Alice Schoenauer, Sen, Priyanka, Sepehri, Sholeh, Seshadri, Preethi, Shen, Ye, Sherborne, Tom, Shi, Sylvie Shang, Shivaprasad, Sanal, Shmyhlo, Vladyslav, Shrinivason, Anirudh, Shteinbuk, Inna, Shukayev, Amir, Simard, Mathieu, Snyder, Ella, Spataru, Ava, Spooner, Victoria, Starostina, Trisha, Strub, Florian, Su, Yixuan, Sun, Jimin, Talupuru, Dwarak, Tarassov, Eugene, Tommasone, Elena, Tracey, Jennifer, Trend, Billy, Tumer, Evren, รstรผn, Ahmet, Venkitesh, Bharat, Venuto, David, Verga, Pat, Voisin, Maxime, Wang, Alex, Wang, Donglu, Wang, Shijian, Wen, Edmond, White, Naomi, Willman, Jesse, Winkels, Marysia, Xia, Chen, Xie, Jessica, Xu, Minjie, Yang, Bowen, Yi-Chern, Tan, Zhang, Ivan, Zhao, Zhenyu, Zhao, Zhoujie
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised and multilingual-capable model, with support for 23 languages of global business, and a novel hybrid architecture balancing efficiency with top of the range performance. It offers best-in-class Retrieval Augmented Generation (RAG) capabilities with grounding and tool use to automate sophisticated business processes. These abilities are achieved through a decentralised training approach, including self-refinement algorithms and model merging techniques. We also include results for Command R7B which shares capability and architectural similarities to Command A. Weights for both models have been released for research purposes. This technical report details our original training pipeline and presents an extensive evaluation of our models across a suite of enterprise-relevant tasks and public benchmarks, demonstrating excellent performance and efficiency.
GIScience in the Era of Artificial Intelligence: A Research Agenda Towards Autonomous GIS
Li, Zhenlong, Ning, Huan, Gao, Song, Janowicz, Krzysztof, Li, Wenwen, Arundel, Samantha T., Yang, Chaowei, Bhaduri, Budhendra, Wang, Shaowen, Zhu, A-Xing, Gahegan, Mark, Shekhar, Shashi, Ye, Xinyue, McKenzie, Grant, Cervone, Guido, Hodgson, Michael E.
The advent of generative AI exemplified by large language models (LLMs) opens new ways to represent and compute geographic information and transcends the process of geographic knowledge production, driving geographic information systems (GIS) towards autonomous GIS. Leveraging LLMs as the decision core, autonomous GIS can independently generate and execute geoprocessing workflows to perform spatial analysis. In this vision paper, we further elaborate on the concept of autonomous GIS and present a conceptual framework that defines its five autonomous goals, five autonomous levels, five core functions, and three operational scales. We demonstrate how autonomous GIS could perform geospatial data retrieval, spatial analysis, and map making with four proof-of-concept GIS agents. We conclude by identifying critical challenges and future research directions, including fine-tuning and self-growing decision-cores, autonomous modeling, and examining the societal and practical implications of autonomous GIS. By establishing the groundwork for a paradigm shift in GIScience, this paper envisions a future where GIS moves beyond traditional workflows to autonomously reason, derive, innovate, and advance geospatial solutions to pressing global challenges. Meanwhile, as we design and deploy increasingly intelligent geospatial systems, we carry a responsibility to ensure they are developed in socially responsible ways, serve the public good, and support the continued value of human geographic insight in an AI-augmented future.
Generative Data Imputation for Sparse Learner Performance Data Using Generative Adversarial Imputation Networks
Zhang, Liang, Lin, Jionghao, Sabatini, John, Zapata-Rivera, Diego, Forsyth, Carol, Jiang, Yang, Hollander, John, Hu, Xiangen, Graesser, Arthur C.
DV ANCEMENTS in AI-driven technologies have significantly enhanced modern education through personalized tutoring and adaptive learning strategies on online platforms [1], [2]. Intelligent T utoring Systems (ITSs) exemplify this progress by leveraging advanced machine learning and natural language processing models to create interactive learning environments that improve outcomes across domains like literacy [3], mathematics [4], language learning [5], biology [6] and other STEM fields [7]. As human learners interact with ITSs, often through question-and-answer scenarios with immediate responses, their performance data becomes crucial for learner modeling, enabling systems to track progress, predict future performance, and adapt instruction accordingly [8]. Learner models like Bayesian Knowledge Tracing (BKT) and other knowledge tracing variants utilize the learner performance data to uncover learning characteristics, estimate knowledge states and acquisition [9]. However, in real-world scenarios, missing learner performance data is prevalent due to factors, such as learner dropout or disengagement [10], technical issues or incomplete data logging [11], biased sampling within experimental groups [12], and more. These challenges often lead to sparse data, where items (i.e., questions or problems) remain unattempted (e.g., learners may bypass the question, leave it unanswered due to a lack of response initiation, or make no attempt to engage with it), alongside limited learner interactions [13], [14]. As shown in Figure 1, missing performance records can occur along both the attempt and question dimensions during learner-ITS interactions. In the right portion of the figure's two matrices, entries marked with "?
I started 'vibe coding' my own apps with AI and I'm loving it
I've always had an interest in programming, because I've always had an interest in computers. I put together websites in HTML as a teenager (which, yes, were hosted on GeoCities) and have been occasionally dabbling in Python since. Yet none of my projects got very far and, apart from my early websites, I never made anything useful. My efforts all followed a familiar pattern: I'd fixate on a particular resource--like an O'Reilly book or an online course--and get started with great enthusiasm, but as I'd realize I was months or years away from creating anything remotely useful, I'd give up. That changed in late 2024 when my general frustration with WordPress, which I was using for my personal website, got the better of me. In a fit, I threw my website's content plus a screenshot of it into Claude 3.5 Sonnet and asked the AI to replicate my site with HTML, CSS, and JavaScript.
Beyond Worst-Case Online Classification: VC-Based Regret Bounds for Relaxed Benchmarks
Montasser, Omar, Shetty, Abhishek, Zhivotovskiy, Nikita
We revisit online binary classification by shifting the focus from competing with the best-in-class binary loss to competing against relaxed benchmarks that capture smoothed notions of optimality. Instead of measuring regret relative to the exact minimal binary error -- a standard approach that leads to worst-case bounds tied to the Littlestone dimension -- we consider comparing with predictors that are robust to small input perturbations, perform well under Gaussian smoothing, or maintain a prescribed output margin. Previous examples of this were primarily limited to the hinge loss. Our algorithms achieve regret guarantees that depend only on the VC dimension and the complexity of the instance space (e.g., metric entropy), and notably, they incur only an $O(\log(1/\gamma))$ dependence on the generalized margin $\gamma$. This stands in contrast to most existing regret bounds, which typically exhibit a polynomial dependence on $1/\gamma$. We complement this with matching lower bounds. Our analysis connects recent ideas from adversarial robustness and smoothed online learning.