Large Language Model
OpenAI is still gobbling up GPUs by the thousands for ChatGPT
You can't find a new Nvidia graphics card for love nor money. Between pent-up demand from PC gamers and Nvidia selling every GPU it can to the bubbling AI industry, new models are going out of stock in a matter of minutes -- and it looks like the situation isn't going to improve any time soon, as the biggest AI company around wants even more hardware. OpenAI CEO Sam Altman took to the social network formerly known as Twitter (spotted by Tom's Hardware) to say that OpenAI's ChatGPT version 4.5 is ready to go… but desperately in need of even more hardware. The "giant, expensive model" requires even more data center capacity than older versions, and to launch with enough access for paid users, the company is gobbling up GPUs at an even faster rate. The CEO claims that OpenAI is adding "tens of thousands of GPUs next week" for the planned rollout, with hundreds of thousands following soon after.
Sora, OpenAI's video generator, has hit the UK. It's obvious why creatives are worried
If you want to know why Tyler Perry put an 800m ( 635m) expansion of his studio complex on hold, type "two people in a living room in the mountains" into OpenAI's video generation tool. The result from artificial intelligence-powered Sora, which was released in the UK and Europe on Friday, indicates why the US TV and film mogul paused his plans. Perry said last year after seeing previews of Sora that if he wanted to produce that mountain shot, he may not need to build sets on location or on his lot. "I can sit in an office and do this with a computer, which is shocking to me," he said. The result from a simple text prompt is only five seconds long – you can go to up to 20 seconds and also stitch together much longer videos from the tool – and the "actors" display telltale problems with their hands (a common problem with AI tools).
The Download: underage celebrity chatbots, and OpenAI's latest model
Botify AI, a site for chatting with AI companions that's backed by the venture capital firm Andreessen Horowitz, hosts bots resembling real actors that state their age as under 18, engage in sexually charged conversations, offer "hot photos," and in some instances describe age-of-consent laws as "arbitrary" and "meant to be broken." When MIT Technology Review tested the site this week, we found popular user-created bots taking on underage characters meant to resemble Jenna Ortega as Wednesday Addams, Emma Watson as Hermione Granger, and Millie Bobby Brown, among others. The conversations--along with the fact that Botify AI includes "send a hot photo" as a feature for its characters--suggest that the ability to elicit sexually charged conversations and images is not accidental. Instead, sexually suggestive conversations appear to be baked in. OpenAI just released GPT-4.5 and says it is its biggest and best chat model yet What's new: OpenAI has just released GPT-4.5, a new version of its flagship large language model which it claims is its biggest and best model for chat yet.
OpenAI launches Sora video generation tool in UK amid copyright row
San Francisco-based OpenAI is making Sora available to UK users who pay for ChatGPT. The tool stunned film-makers when it was revealed last year, with the film and TV mogul Tyler Perry pausing an 800m ( 634m) expansion of his Atlanta studio complex after saying the tool might make building sets or travelling to locations unnecessary. It was launched in the US publicly in December. Users are able to make videos on Sora by typing in simple prompts such as asking for a shot of people walking through "beautiful, snowy Tokyo city" where "gorgeous sakura petals are flying through the wind along with snowflakes". OpenAI announced the UK release as it released examples of Sora's use by artists from across the UK and mainland Europe, where the tool is also being released on Friday. Josephine Miller, a 25-year-old British digital artist, created a two-minute video of models wearing bioluminescent fauna and said the tool would "open a lot more doors for younger creatives".
DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
Li, Zhuoqun, Yu, Haiyang, Chen, Xuanang, Lin, Hongyu, Lu, Yaojie, Huang, Fei, Han, Xianpei, Li, Yongbin, Sun, Le
Designing solutions for complex engineering challenges is crucial in human production activities. However, previous research in the retrieval-augmented generation (RAG) field has not sufficiently addressed tasks related to the design of complex engineering solutions. To fill this gap, we introduce a new benchmark, SolutionBench, to evaluate a system's ability to generate complete and feasible solutions for engineering problems with multiple complex constraints. To further advance the design of complex engineering solutions, we propose a novel system, SolutionRAG, that leverages the tree-based exploration and bi-point thinking mechanism to generate reliable solutions. Extensive experimental results demonstrate that SolutionRAG achieves state-of-the-art (SOTA) performance on the SolutionBench, highlighting its potential to enhance the automation and reliability of complex engineering solution design in real-world applications.
Plan2Align: Predictive Planning Based Test-Time Preference Alignment in Paragraph-Level Machine Translation
Wang, Kuang-Da, Chen, Teng-Ruei, Hung, Yu Heng, Ding, Shuoyang, Wu, Yueh-Hua, Wang, Yu-Chiang Frank, Yang, Chao-Han Huck, Peng, Wen-Chih, Hsieh, Ping-Chun
Machine Translation (MT) has been predominantly designed for sentence-level translation using transformer-based architectures. While next-token prediction based Large Language Models (LLMs) demonstrate strong capabilities in long-text translation, non-extensive language models often suffer from omissions and semantic inconsistencies when processing paragraphs. Existing preference alignment methods improve sentence-level translation but fail to ensure coherence over extended contexts due to the myopic nature of next-token generation. We introduce Plan2Align, a test-time alignment framework that treats translation as a predictive planning problem, adapting Model Predictive Control to iteratively refine translation outputs. Experiments on WMT24 Discourse-Level Literary Translation show that Plan2Align significantly improves paragraph-level translation, achieving performance surpassing or on par with the existing training-time and test-time alignment methods on LLaMA-3.1 8B.
Learning to Substitute Components for Compositional Generalization
Li, Zhaoyi, Jiang, Gangwei, Wu, Chenwang, Wei, Ying, Lian, Defu, Chen, Enhong
Despite the rising prevalence of neural language models, recent empirical evidence suggests their deficiency in compositional generalization. One of the current de-facto solutions to this problem is compositional data augmentation, which aims to introduce additional compositional inductive bias. However, existing handcrafted augmentation strategies offer limited improvement when systematic generalization of neural language models requires multi-grained compositional bias (i.e., not limited to either lexical or structural biases alone) or when training sentences have an imbalanced difficulty distribution. To address these challenges, we first propose a novel compositional augmentation strategy called Component Substitution (CompSub), which enables multi-grained composition of substantial substructures across the entire training set. Furthermore, we introduce the Learning Component Substitution (LCS) framework. This framework empowers the learning of component substitution probabilities in CompSub in an end-to-end manner by maximizing the loss of neural language models, thereby prioritizing challenging compositions with elusive concepts and novel contexts. We extend the key ideas of CompSub and LCS to the recently emerging in-context learning scenarios of pre-trained large language models (LLMs), proposing the LCS-ICL algorithm to enhance the few-shot compositional generalization of state-of-the-art (SOTA) LLMs. Theoretically, we provide insights into why applying our algorithms to language models can improve compositional generalization performance. Empirically, our results on four standard compositional generalization benchmarks(SCAN, COGS, GeoQuery, and COGS-QL) demonstrate the superiority of CompSub, LCS, and LCS-ICL, with improvements of up to 66.5%, 10.3%, 1.4%, and 8.8%, respectively.
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
Orlikowski, Matthias, Pei, Jiaxin, Röttger, Paul, Cimiano, Philipp, Jurgens, David, Hovy, Dirk
People naturally vary in their annotations for subjective questions and some of this variation is thought to be due to the person's sociodemographic characteristics. LLMs have also been used to label data, but recent work has shown that models perform poorly when prompted with sociodemographic attributes, suggesting limited inherent sociodemographic knowledge. Here, we ask whether LLMs can be trained to be accurate sociodemographic models of annotator variation. Using a curated dataset of five tasks with standardized sociodemographics, we show that models do improve in sociodemographic prompting when trained but that this performance gain is largely due to models learning annotator-specific behaviour rather than sociodemographic patterns. Across all tasks, our results suggest that models learn little meaningful connection between sociodemographics and annotation, raising doubts about the current use of LLMs for simulating sociodemographic variation and behaviour.
Large Language Models Are Innate Crystal Structure Generators
Gan, Jingru, Zhong, Peichen, Du, Yuanqi, Zhu, Yanqiao, Duan, Chenru, Wang, Haorui, Gomes, Carla P., Persson, Kristin A., Schwalbe-Koda, Daniel, Wang, Wei
Crystal structure generation is fundamental to materials discovery, enabling the prediction of novel materials with desired properties. While existing approaches leverage Large Language Models (LLMs) through extensive fine-tuning on materials databases, we show that pre-trained LLMs can inherently generate stable crystal structures without additional training. Our novel framework MatLLMSearch integrates pre-trained LLMs with evolutionary search algorithms, achieving a 78.38% metastable rate validated by machine learning interatomic potentials and 31.7% DFT-verified stability via quantum mechanical calculations, outperforming specialized models such as CrystalTextLLM. Beyond crystal structure generation, we further demonstrate that our framework can be readily adapted to diverse materials design tasks, including crystal structure prediction and multi-objective optimization of properties such as deformation energy and bulk modulus, all without fine-tuning. These results establish pre-trained LLMs as versatile and effective tools for materials discovery, opening up new venues for crystal structure generation with reduced computational overhead and broader accessibility.
SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models
Zhang, Jiawei, Yang, Xuan, Wang, Taiqi, Yao, Yu, Petiushko, Aleksandr, Li, Bo
Traditional autonomous driving systems often struggle to integrate high-level reasoning with low-level control, resulting in suboptimal and sometimes unsafe driving behaviors. The emergence of Multimodal Large Language Models (MLLMs), which can process both visual and textual data, presents an opportunity to unify perception and reasoning tasks within a single framework. However, effectively embedding precise safety knowledge into MLLMs for autonomous driving remains a significant challenge. To address this, we propose SafeAuto, a novel framework that enhances MLLM-based autonomous driving systems by incorporating both unstructured and structured knowledge. Specifically, we first introduce the Position-Dependent Cross-Entropy (PDCE) loss function, designed to improve the accuracy of low-level control signal predictions when numerical values are represented as text. Second, to ensure safe autonomous driving by explicitly integrating precise safety knowledge into the MLLM, we develop a reasoning component for SafeAuto. This component translates driving safety regulations into first-order logic rules (e.g., "red light => stop") and incorporates these rules into a probabilistic graphical model, such as a Markov Logic Network (MLN). The MLN is trained to verify the predicted next actions using environmental attributes identified by attribute recognition models (e.g., detecting a red light) to form the predicates. Additionally, we construct a Multimodal RAG model that leverages video data, control signals, and environmental attributes to learn more effectively from past similar driving experiences. By integrating PDCE, MLN, and Multimodal RAG, SafeAuto significantly outperforms existing baselines across multiple datasets. This advancement enables more accurate, reliable, and safer autonomous driving systems that learn from experience, obey traffic laws, and perform precise control actions.