Large Language Model
OpenAI just released GPT-4.5 and says it is its biggest and best chat model yet
People with a 200-a-month ChatGPT Pro account can try out GPT-4.5 today. OpenAI says it will begin rolling out to other users next week. With each release of its GPT models, OpenAI has shown that bigger means better. But there has been a lot of talk about how that approach is hitting a wall--including remarks from OpenAI's former chief scientist Ilya Sutskever. The company's claims about GPT-4.5 feel like a thumb in the eye to the naysayers.
OpenAI's new GPT-4.5 model is a better, more natural conversationalist
In what has already been a busy past few days for new model releases, OpenAI is capping off the week with a research preview of GPT-4.5. The company is touting the new system as its largest and best model for chat yet. In early testing, OpenAI says people found GPT-4.5 to be a more natural conversationalist, with the ability to convey warmth and display a kind of emotional intelligence. In one example shared by OpenAI, a person tells ChatGPT they're going through a hard time after failing a test. Where the company's previous models, including GPT-4o and o3-mini, might commiserate with the individual before offering a long list of unsolicited advice, GPT-4.5 takes a different tact. "Want to talk about what happened, or do you just need a distraction?
OpenAI Launches GPT-4.5 for ChatGPT--It's Huge and Compute-Intensive
GPT-4.5 is here, and OpenAI's newest generative AI model is bigger and more compute-intensive than ever--it's supposedly also better at understanding what ChatGPT users mean with their prompts. Users who want to be part of the first wave to try GPT-4.5, labeled as a research preview, will be required to pay for OpenAI's 200-a-month ChatGPT Pro subscription. Prior to this launch, 2025 has already been filled with new AI model releases. Anthropic recently put out a hybrid reasoning model for its Claude chatbot. Before that, Chinese researchers at DeepSeek rocked Silicon Valley with their release of a powerful model trained on a tiny budget, prompting OpenAI to drop a "mini" version of its reasoning model a month ago.
ConvoyLLM: Dynamic Multi-Lane Convoy Control Using LLMs
Lu, Liping, He, Zhican, Chu, Duanfeng, Wang, Rukang, Peng, Saiqian, Zhou, Pan
ConvoyLLM: Dynamic Multi-Lane Convoy Control Using LLMs Liping Lu 1, Zhican He 1, Duanfeng Chu 2, Rukang Wang 2, Saiqian Peng 2, Pan Zhou 3 Abstract -- This paper proposes a novel method for multilane convoy formation control that uses large language models (LLMs) to tackle coordination challenges in dynamic highway environments. Each connected and autonomous vehicle in the convoy uses a knowledge-driven approach to make real-time adaptive decisions based on various scenarios. Our method enables vehicles to dynamically perform tasks, including obstacle avoidance, convoy joining/leaving, and escort formation switching, all while maintaining the overall convoy structure. We design a Interlaced formation control strategy based on locally dynamic distributed graphs, ensuring the convoy remains stable and flexible. We conduct extensive experiments in the SUMO simulation platform across multiple traffic scenarios, and the results demonstrate that the proposed method is effective, robust, and adaptable to dynamic environments. I. INTRODUCTION With the rapid development of Connected and Automated V ehicles (CA Vs) technology, convoy coordination control has shown significant potential in improving traffic flow efficiency, driving safety, and fuel economy.
SeisMoLLM: Advancing Seismic Monitoring via Cross-modal Transfer with Pre-trained Large Language Model
Wang, Xinghao, Liu, Feng, Su, Rui, Wang, Zhihui, Bai, Lei, Ouyang, Wanli
Recent advances in deep learning have revolutionized seismic monitoring, yet developing a foundation model that performs well across multiple complex tasks remains challenging, particularly when dealing with degraded signals or data scarcity. This work presents SeisMoLLM, the first foundation model that utilizes cross-modal transfer for seismic monitoring, to unleash the power of large-scale pre-training from a large language model without requiring direct pre-training on seismic datasets. Through elaborate waveform tokenization and fine-tuning of pre-trained GPT-2 model, SeisMoLLM achieves state-of-the-art performance on the DiTing and STEAD datasets across five critical tasks: back-azimuth estimation, epicentral distance estimation, magnitude estimation, phase picking, and first-motion polarity classification. It attains 36 best results out of 43 task metrics and 12 top scores out of 16 few-shot generalization metrics, with many relative improvements ranging from 10% to 50%. In addition to its superior performance, SeisMoLLM maintains efficiency comparable to or even better than lightweight models in both training and inference. These findings establish SeisMoLLM as a promising foundation model for practical seismic monitoring and highlight cross-modal transfer as an exciting new direction for earthquake studies, showcasing the potential of advanced deep learning techniques to propel seismology research forward.
NANOGPT: A Query-Driven Large Language Model Retrieval-Augmented Generation System for Nanotechnology Research
Chandrasekhar, Achuth, Farimani, Omid Barati, Ajenifujah, Olabode T., Ock, Janghoon, Farimani, Amir Barati
This paper presents the development and application of a Large Language Model Retrieval-Augmented Generation (LLM-RAG) system tailored for nanotechnology research. The system leverages the capabilities of a sophisticated language model to serve as an intelligent research assistant, enhancing the efficiency and comprehensiveness of literature reviews in the nanotechnology domain. Central to this LLM-RAG system is its advanced query backend retrieval mechanism, which integrates data from multiple reputable sources. The system retrieves relevant literature by utilizing Google Scholar's advanced search, and scraping open-access papers from Elsevier, Springer Nature, and ACS Publications. This multifaceted approach ensures a broad and diverse collection of up-to-date scholarly articles and papers. The proposed system demonstrates significant potential in aiding researchers by providing a streamlined, accurate, and exhaustive literature retrieval process, thereby accelerating research advancements in nanotechnology. The effectiveness of the LLM-RAG system is validated through rigorous testing, illustrating its capability to significantly reduce the time and effort required for comprehensive literature reviews, while maintaining high accuracy, query relevance and outperforming standard, publicly available LLMS.
Large Language Models as Attribution Regularizers for Efficient Model Training
Vukadin, Davor, Šilić, Marin, Delač, Goran
Large Language Models (LLMs) have demonstrated remarkable performance across diverse domains. However, effectiv ely leveraging their vast knowledge for training smaller downstream model s remains an open challenge, especially in domains like tabular data lea rning, where simpler models are often preferred due to interpretability and efficiency. In this paper, we introduce a novel yet straightforward meth od for incorporating LLM-generated global task feature attributions i nto the training process of smaller networks. Specifically, we propose an attribution-matching regularization term that aligns the training dyna mics of the smaller model with the insights provided by the LLM. By doing so, our approach yields superior performance in few-shot learn ing scenarios. Notably, our method requires only black-box API access to th e LLM, making it easy to integrate into existing training pipeline s with minimal computational overhead. Furthermore, we demonstrate how this method can be used to ad dress common issues in real-world datasets, such as skewness and b ias. By integrating high-level knowledge from LLMs, our approach i mproves generalization, even when training data is limited or imbal anced. We validate its effectiveness through extensive experiments a cross multiple tasks, demonstrating improved learning efficiency and model robustness.
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
Gong, Zixuan, Teng, Jiaye, Liu, Yong
Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantically correct. However, existing theoretical analyses hardly account for this two-stage phenomenon. In this paper, we theoretically demonstrate how such two-stage training dynamics occur in transformers. Specifically, we analyze the dynamics of transformers using feature learning techniques under in-context learning regimes, based on a disentangled two-type feature structure. Such disentanglement of feature structure is general in practice, e.g., natural languages contain syntax and semantics, and proteins contain primary and secondary structures. To our best known, this is the first rigorous result regarding a two-stage optimization process in transformers. Additionally, a corollary indicates that such a two-stage process is closely related to the spectral properties of the attention weights, which accords well with empirical findings.
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
Li, Lei, Jia, Sen, Wang, Jianhao, An, Zhaochong, Li, Jiaang, Hwang, Jenq-Neng, Belongie, Serge
Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical perspectives. To address these challenges, we introduce ChatMotion, a multimodal multi-agent framework for human motion analysis. ChatMotion dynamically interprets user intent, decomposes complex tasks into meta-tasks, and activates specialized function modules for motion comprehension. It integrates multiple specialized modules, such as the MotionCore, to analyze human motion from various perspectives. Extensive experiments demonstrate ChatMotion's precision, adaptability, and user engagement for human motion understanding.
Stochastic Rounding for LLM Training: Theory and Practice
Ozkara, Kaan, Yu, Tao, Park, Youngsuk
As the parameters of Large Language Models (LLMs) have scaled to hundreds of billions, the demand for efficient training methods -- balancing faster computation and reduced memory usage without sacrificing accuracy -- has become more critical than ever. In recent years, various mixed precision strategies, which involve different precision levels for optimization components, have been proposed to increase training speed with minimal accuracy degradation. However, these strategies often require manual adjustments and lack theoretical justification. In this work, we leverage stochastic rounding (SR) to address numerical errors of training with low-precision representation. We provide theoretical analyses of implicit regularization and convergence under the Adam optimizer when SR is utilized. With the insights from these analyses, we extend previous BF16 + SR strategy to be used in distributed settings, enhancing the stability and performance for large scale training. Empirical results from pre-training models with up to 6.7B parameters, for the first time, demonstrate that our BF16 with SR strategy outperforms (BF16, FP32) mixed precision strategies, achieving better validation perplexity, up to $1.54\times$ higher throughput, and $30\%$ less memory usage.