Large Language Model
KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations
Harbola, Chitranshu, Purwar, Anupam
In the evolving landscape of conversational AI, generating concise, context-aware, and human-like dialogue using small and medium-sized language models (LLMs) remains a complex challenge. This study investigates the influence of LoRA rank, dataset scale, and prompt prefix design on both knowledge retention and stylistic alignment. While fine-tuning improves fluency and enables stylistic customization, its ability to integrate unseen knowledge is constrained -- particularly with smaller datasets. Conversely, RAG-augmented models, equipped to incorporate external documents at inference, demonstrated superior factual accuracy on out-of-distribution prompts, though they lacked the stylistic consistency achieved by fine-tuning. Evaluations by LLM-based judges across knowledge accuracy, conversational quality, and conciseness suggest that fine-tuning is best suited for tone adaptation, whereas RAG excels at real-time knowledge augmentation.
Command A: An Enterprise-Ready Large Language Model
Cohere, Team, :, null, Aakanksha, null, Ahmadian, Arash, Ahmed, Marwan, Alammar, Jay, Alizadeh, Milad, Alnumay, Yazeed, Althammer, Sophia, Arkhangorodsky, Arkady, Aryabumi, Viraat, Aumiller, Dennis, Avalos, Raphaรซl, Aviv, Zahara, Bae, Sammie, Baji, Saurabh, Barbet, Alexandre, Bartolo, Max, Bebensee, Bjรถrn, Beladia, Neeral, Beller-Morales, Walter, Bรฉrard, Alexandre, Berneshawi, Andrew, Bialas, Anna, Blunsom, Phil, Bobkin, Matt, Bongale, Adi, Braun, Sam, Brunet, Maxime, Cahyawijaya, Samuel, Cairuz, David, Campos, Jon Ander, Cao, Cassie, Cao, Kris, Castagnรฉ, Roman, Cendrero, Juliรกn, Currie, Leila Chan, Chandak, Yash, Chang, Diane, Chatziveroglou, Giannis, Chen, Hongyu, Cheng, Claire, Chevalier, Alexis, Chiu, Justin T., Cho, Eugene, Choi, Eugene, Choi, Eujeong, Chung, Tim, Cirik, Volkan, Cismaru, Ana, Clavier, Pierre, Conklin, Henry, Crawhall-Stein, Lucas, Crouse, Devon, Cruz-Salinas, Andres Felipe, Cyrus, Ben, D'souza, Daniel, Dalla-Torre, Hugo, Dang, John, Darling, William, Domingues, Omar Darwiche, Dash, Saurabh, Debugne, Antoine, Dehaze, Thรฉo, Desai, Shaan, Devassy, Joan, Dholakia, Rishit, Duffy, Kyle, Edalati, Ali, Eldeib, Ace, Elkady, Abdullah, Elsharkawy, Sarah, Ergรผn, Irem, Ermis, Beyza, Fadaee, Marzieh, Fan, Boyu, Fayoux, Lucas, Flet-Berliac, Yannis, Frosst, Nick, Gallรฉ, Matthias, Galuba, Wojciech, Garg, Utsav, Geist, Matthieu, Azar, Mohammad Gheshlaghi, Gilsenan-McMahon, Ellen, Goldfarb-Tarrant, Seraphina, Goldsack, Tomas, Gomez, Aidan, Gonzaga, Victor Machado, Govindarajan, Nithya, Govindassamy, Manoj, Grinsztajn, Nathan, Gritsch, Nikolas, Gu, Patrick, Guo, Shangmin, Haefeli, Kilian, Hajjar, Rod, Hawes, Tim, He, Jingyi, Hofstรคtter, Sebastian, Hong, Sungjin, Hooker, Sara, Hosking, Tom, Howe, Stephanie, Hu, Eric, Huang, Renjie, Jain, Hemant, Jain, Ritika, Jakobi, Nick, Jenkins, Madeline, Jordan, JJ, Joshi, Dhruti, Jung, Jason, Kalyanpur, Trushant, Kamalakara, Siddhartha Rao, Kedrzycki, Julia, Keskin, Gokce, Kim, Edward, Kim, Joon, Ko, Wei-Yin, Kocmi, Tom, Kozakov, Michael, Kryลciลski, Wojciech, Jain, Arnav Kumar, Teru, Komal Kumar, Land, Sander, Lasby, Michael, Lasche, Olivia, Lee, Justin, Lewis, Patrick, Li, Jeffrey, Li, Jonathan, Lin, Hangyu, Locatelli, Acyr, Luong, Kevin, Ma, Raymond, Mach, Lukรกลก, Machado, Marina, Magbitang, Joanne, Lopez, Brenda Malacara, Mann, Aryan, Marchisio, Kelly, Markham, Olivia, Matton, Alexandre, McKinney, Alex, McLoughlin, Dominic, Mokry, Jozef, Morisot, Adrien, Moulder, Autumn, Moynehan, Harry, Mozes, Maximilian, Muppalla, Vivek, Murakhovska, Lidiya, Nagarajan, Hemangani, Nandula, Alekhya, Nasir, Hisham, Nehra, Shauna, Netto-Rosen, Josh, Ohashi, Daniel, Owers-Bardsley, James, Ozuzu, Jason, Padilla, Dennis, Park, Gloria, Passaglia, Sam, Pekmez, Jeremy, Penstone, Laura, Piktus, Aleksandra, Ploeg, Case, Poulton, Andrew, Qi, Youran, Raghvendra, Shubha, Ramos, Miguel, Ranjan, Ekagra, Richemond, Pierre, Robert-Michon, Cรฉcile, Rodriguez, Aurรฉlien, Roy, Sudip, Ruder, Sebastian, Ruis, Laura, Rust, Louise, Sachan, Anubhav, Salamanca, Alejandro, Saravanakumar, Kailash Karthik, Satyakam, Isha, Sebag, Alice Schoenauer, Sen, Priyanka, Sepehri, Sholeh, Seshadri, Preethi, Shen, Ye, Sherborne, Tom, Shi, Sylvie Shang, Shivaprasad, Sanal, Shmyhlo, Vladyslav, Shrinivason, Anirudh, Shteinbuk, Inna, Shukayev, Amir, Simard, Mathieu, Snyder, Ella, Spataru, Ava, Spooner, Victoria, Starostina, Trisha, Strub, Florian, Su, Yixuan, Sun, Jimin, Talupuru, Dwarak, Tarassov, Eugene, Tommasone, Elena, Tracey, Jennifer, Trend, Billy, Tumer, Evren, รstรผn, Ahmet, Venkitesh, Bharat, Venuto, David, Verga, Pat, Voisin, Maxime, Wang, Alex, Wang, Donglu, Wang, Shijian, Wen, Edmond, White, Naomi, Willman, Jesse, Winkels, Marysia, Xia, Chen, Xie, Jessica, Xu, Minjie, Yang, Bowen, Yi-Chern, Tan, Zhang, Ivan, Zhao, Zhenyu, Zhao, Zhoujie
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised and multilingual-capable model, with support for 23 languages of global business, and a novel hybrid architecture balancing efficiency with top of the range performance. It offers best-in-class Retrieval Augmented Generation (RAG) capabilities with grounding and tool use to automate sophisticated business processes. These abilities are achieved through a decentralised training approach, including self-refinement algorithms and model merging techniques. We also include results for Command R7B which shares capability and architectural similarities to Command A. Weights for both models have been released for research purposes. This technical report details our original training pipeline and presents an extensive evaluation of our models across a suite of enterprise-relevant tasks and public benchmarks, demonstrating excellent performance and efficiency.
GIScience in the Era of Artificial Intelligence: A Research Agenda Towards Autonomous GIS
Li, Zhenlong, Ning, Huan, Gao, Song, Janowicz, Krzysztof, Li, Wenwen, Arundel, Samantha T., Yang, Chaowei, Bhaduri, Budhendra, Wang, Shaowen, Zhu, A-Xing, Gahegan, Mark, Shekhar, Shashi, Ye, Xinyue, McKenzie, Grant, Cervone, Guido, Hodgson, Michael E.
The advent of generative AI exemplified by large language models (LLMs) opens new ways to represent and compute geographic information and transcends the process of geographic knowledge production, driving geographic information systems (GIS) towards autonomous GIS. Leveraging LLMs as the decision core, autonomous GIS can independently generate and execute geoprocessing workflows to perform spatial analysis. In this vision paper, we further elaborate on the concept of autonomous GIS and present a conceptual framework that defines its five autonomous goals, five autonomous levels, five core functions, and three operational scales. We demonstrate how autonomous GIS could perform geospatial data retrieval, spatial analysis, and map making with four proof-of-concept GIS agents. We conclude by identifying critical challenges and future research directions, including fine-tuning and self-growing decision-cores, autonomous modeling, and examining the societal and practical implications of autonomous GIS. By establishing the groundwork for a paradigm shift in GIScience, this paper envisions a future where GIS moves beyond traditional workflows to autonomously reason, derive, innovate, and advance geospatial solutions to pressing global challenges. Meanwhile, as we design and deploy increasingly intelligent geospatial systems, we carry a responsibility to ensure they are developed in socially responsible ways, serve the public good, and support the continued value of human geographic insight in an AI-augmented future.
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
Zhang, Rongyu, Dong, Menghang, Zhang, Yuan, Heng, Liang, Chi, Xiaowei, Dai, Gaole, Du, Li, Du, Yuan, Zhang, Shanghang
Multimodal Large Language Models (MLLMs) excel in understanding complex language and visual data, enabling generalist robotic systems to interpret instructions and perform embodied tasks. Nevertheless, their real-world deployment is hindered by substantial computational and storage demands. Recent insights into the homogeneous patterns in the LLM layer have inspired sparsification techniques to address these challenges, such as early exit and token pruning. However, these methods often neglect the critical role of the final layers that encode the semantic information most relevant to downstream robotic tasks. Aligning with the recent breakthrough of the Shallow Brain Hypothesis (SBH) in neuroscience and the mixture of experts in model sparsification, we conceptualize each LLM layer as an expert and propose a Mixture-of-Layers Vision-Language-Action model (MoLe-VLA, or simply MoLe) architecture for dynamic LLM layer activation. We introduce a Spatial-Temporal Aware Router (STAR) for MoLe to selectively activate only parts of the layers based on the robot's current state, mimicking the brain's distinct signal pathways specialized for cognition and causal reasoning. Additionally, to compensate for the cognitive ability of LLMs lost in MoLe, we devise a Cognition Self-Knowledge Distillation (CogKD) framework. CogKD enhances the understanding of task demands and improves the generation of task-relevant action sequences by leveraging cognitive features. Extensive experiments conducted in both RLBench simulation and real-world environments demonstrate the superiority of MoLe-VLA in both efficiency and performance. Specifically, MoLe-VLA achieves an 8% improvement in the mean success rate across ten tasks while reducing computational costs by up to x5.6 compared to standard LLMs.
Structuring Scientific Innovation: A Framework for Modeling and Discovering Impactful Knowledge Combinations
Chen, Junlan, Zhang, Kexin, Li, Daifeng, Feng, Yangyang, Zhang, Yuxuan, Deng, Bowen
The emergence of large language models (LLMs) offers new possibilities for structured exploration of scientific knowledge. Rather than viewing scientific discovery as isolated ideas or content, we propose a structured approach that emphasizes the role of method combinations in shaping disruptive insights. Specifically, we investigate how knowledge units--especially those tied to methodological design--can be modeled and recombined to yield research breakthroughs. Our proposed framework addresses two key challenges. First, we introduce a contrastive learning-based mechanism to identify distinguishing features of historically disruptive method combinations within problem-driven contexts. Second, we propose a reasoning-guided Monte Carlo search algorithm that leverages the chain-of-thought capability of LLMs to identify promising knowledge recombinations for new problem statements. Empirical studies across multiple domains show that the framework is capable of modeling the structural dynamics of innovation and successfully highlights combinations with high disruptive potential. This research provides a new path for computationally guided scientific ideation grounded in structured reasoning and historical data modeling.
Plato: Plan to Efficiently Decode for Large Language Model Inference
Jin, Shuowei, Liu, Xueshen, Wu, Yongji, Zheng, Haizhong, Zhang, Qingzhao, Prakash, Atul, Lentz, Matthew, Zhuo, Danyang, Qian, Feng, Mao, Z. Morley
Large language models (LLMs) have achieved remarkable success in natural language tasks, but their inference incurs substantial computational and memory overhead. To improve efficiency, parallel decoding methods like Skeleton-of-Thought (SoT) decompose prompts into sub-problems for concurrent processing. However, these methods significantly compromise answer quality by treating semantically linked sub-problems as independent. We propose Plato, a novel approach that co-designs algorithms and systems for semantic-aware parallel decoding. Plato leverages LLMs to organize sub-problems into a dependency graph based on logical and causal relationships, enabling concurrent decoding of non-dependent nodes while preserving answer coherence and quality. To further enhance efficiency, Plato pipelines planning and node decoding stages, implements a global context cache, and carefully structures node inference prompts to maximize key-value cache reuse and minimize overhead. Our evaluations show that Plato improves throughput by 68% over autoregressive decoding while achieving a 40% net win rate in answer quality. Compared to SoT, Plato demonstrates a remarkable 90% quality net-win rate. Ablation studies reveal that our pipeline design improves speedup by 29%, while our KV cache reuse optimization reduces overhead by 75%.
Why is constrained neural language generation particularly challenging?
Garbacea, Cristina, Mei, Qiaozhu
Recent advances in deep neural language models combined wit h the capacity of large scale datasets have accelerated the development of natural langu age generation systems that produce fluent and coherent texts (to various degrees of succ ess) in a multitude of tasks and application contexts. However, controlling the output of t hese models for specific user and task needs is still an open challenge. This is crucial not onl y to customizing the content and style of the generated language, but also to their safe and re liable deployment in the real world. We present an extensive survey on the emerging topic o f constrained neural language generation in which we formally define and categorize the pro blems of natural language generation by distinguishing between conditions and constraints (the latter being testable conditions on the output text instead of the input), present constrained text generation tasks, and review existing methods and evaluation metrics for cons trained text generation. Our aim is to highlight recent progress and trends in this emergi ng field, informing on the most promising directions and limitations towards advancing th e state-of-the-art of constrained neural language generation research.
Google made an AI model to talk to dolphins
A new large language model AI system may soon allow humans to converse with dolphins. Scheduled to debut in the coming months, researchers will test to see if DolphinGemma and its companion Cetacean Hearing Augmentation Telemetry (CHAT) system can translate and mimic some of the mammal's own complex vocalizations. If successful, the breakthrough may represent the culmination of over four decades' worth of work, documentation, and conservation efforts.. Dolphins are some of the Earth's smartest and most communicative animals. Their social interactions are so complex that researchers at the Wild Dolphin Project (WDP) have spent the last 40 years attempting to decipher them. In the process, WDP has amassed decades' worth of underwater audio and video documenting a single community of Atlantic spotted dolphins in the Bahamas.
OpenAI is phasing out GPT-4.5 for developers
OpenAI has announced its phasing out GPT-4.5 from its developer API in favor of its new GPT-4.1 model. When it launched, OpenAI described GPT-4.5 as its best and most capable model so far, in part because it was a more natural conversationalist and could capably mimic some notion of emotional intelligence. Despite what its name suggests, GPT-4.1 is supposed to be better and more efficient. That means that if you won't find it as in option in the public-facing ChatGPT interface, but you could someday interact with an agent that leverages the model's improvements. GPT-4.1 is supposed to be better at coding and "long context understanding," according to OpenAI, with support for "up to one million tokens of context" and knowledge of the world up to June 2024.
OpenAI's New GPT 4.1 Models Excel at Coding
OpenAI announced today that it is releasing a new family of artificial intelligence models optimized to excel at coding, as it ramps up efforts to fend off increasingly stiff competition from companies like Google and Anthropic. The models are available to developers through OpenAI's application programming interface (API). OpenAI is releasing three sizes of models: GPT 4.1, GPT 4.1 Mini, and GPT 4.1 Nano. Kevin Weil, chief product officer at OpenAI, said on a livestream that the new models are better than OpenAI's most widely used model, GPT-4o, and better than its largest and most powerful model, GPT-4.5, in some ways. GPT-4.1 scored 55 percent on SWE-Bench, a widely used benchmark for gauging the prowess of coding models.