Goto

Collaborating Authors

 Large Language Model


OpenAI fires back at Elon Musk in legal fight over breach of contract claims

The Guardian

OpenAI has hit back at Elon Musk's lawsuit accusing it of betraying its altruistic roots, claiming the Tesla chief executive had in fact supported the artificial intelligence company's plans to create a for-profit unit. Executives at the ChatGPT maker released a blogpost containing what it claimed was historical email correspondence with Musk in which the entrepreneur suggested merging the San Francisco-based startup with Tesla and said it should attach to the electric carmaker "as its cash cow". The blog, authored by OpenAI executives including its chief executive, Sam Altman, claims that in 2017 "we and Elon decided the next step for the mission was to create a for-profit entity". Last week Musk filed a lawsuit accusing OpenAI, where he was a founding board member, of deviating from its foundational mission by forming a for-profit unit โ€“ and putting making money before its core aim of producing technology for the benefit of humanity. "We're sad that it's come to this with someone whom we've deeply admired โ€“ someone who inspired us to aim higher, then told us we would fail, started a competitor, and then sued us when we started making meaningful progress towards OpenAI's mission without him," said OpenAI.


What's Going On with Kara Swisher's Book Tour?

Slate

Last week saw the release of Kara Swisher's Burn Book, the highly anticipated career memoir from a titanic, justly celebrated veteran of tech journalism. Considering her unique, outsize stature in Silicon Valley, and her decadeslong record of landing bombshell inside scoops about the single most important industry of the 21st century, Swisher's choice to promote her latest project with the help of famous friends (Don Lemon, Massachusetts Gov. Maura Healey, etc.) certainly makes sense. What makes much less sense, however, is her selection of tech-world executives. The book tour is going to be lit -- with guest moderators like @RobertIger, @laurenepowell, @mcuban, @donlemon, @reidhoffman, @sama and more. Some of the "moderators" on her tour include Laurene Powell Jobs, Disney CEO Bob Iger, OpenAI CEO Sam Altman, LinkedIn co-founder Reid Hoffman, and Lean In board member Adam Grant. Per NPR's Steve Inskeep, she personally requested that these folks "interview her on stage," in a series of conversations she intends to turn into individual podcast episodes.


OpenAI says Elon Musk wanted it to merge with Tesla to create a for-profit entity

Engadget

Elon Musk, who sued OpenAI for violating its non-profit mission and chasing profits, allegedly wanted the organization to merge with Tesla when it was starting to plan its transition into a for-profit entity in order to accomplish its goals. Well, either that or get full control of the company, OpenAI said in a blog post. The organization responded to Musk's lawsuit by publishing old emails from 2015 to 2018 when he was still involved in its operations. When OpenAI introduced itself to the world back in 2015, it announced that it had 1 billion in funding. Apparently, Musk was the one who suggested that figure, even though OpenAI had raised less than 45 million from him and around 90 million from other donors.


Learning to Decode Collaboratively with Multiple Language Models

arXiv.org Artificial Intelligence

We propose a method to teach multiple large language models (LLM) to collaborate by interleaving their generations at the token level. We model the decision of which LLM generates the next token as a latent variable. By optimizing the marginal likelihood of a training set under our latent variable model, the base LLM automatically learns when to generate itself and when to call on one of the ``assistant'' language models to generate, all without direct supervision. Token-level collaboration during decoding allows for a fusion of each model's expertise in a manner tailored to the specific task at hand. Our collaborative decoding is especially useful in cross-domain settings where a generalist base LLM learns to invoke domain expert models. On instruction-following, domain-specific QA, and reasoning tasks, we show that the performance of the joint system exceeds that of the individual models. Through qualitative analysis of the learned latent decisions, we show models trained with our method exhibit several interesting collaboration patterns, e.g., template-filling. Our code is available at https://github.com/clinicalml/co-llm.


Learning with Language-Guided State Abstractions

arXiv.org Artificial Intelligence

We describe a framework for using natural language to design state abstractions for imitation learning. Generalizable policy learning in high-dimensional observation spaces is facilitated by well-designed state representations, which can surface important features of an environment and hide irrelevant ones. These state representations are typically manually specified, or derived from other labor-intensive labeling procedures. Our method, LGA (language-guided abstraction), uses a combination of natural language supervision and background knowledge from language models (LMs) to automatically build state representations tailored to unseen tasks. In LGA, a user first provides a (possibly incomplete) description of a target task in natural language; next, a pre-trained LM translates this task description into a state abstraction function that masks out irrelevant features; finally, an imitation policy is trained using a small number of demonstrations and LGA-generated abstract states. Experiments on simulated robotic tasks show that LGA yields state abstractions similar to those designed by humans, but in a fraction of the time, and that these abstractions improve generalization and robustness in the presence of spurious correlations and ambiguous specifications. We illustrate the utility of the learned abstractions on mobile manipulation tasks with a Spot robot.


Preference optimization of protein language models as a multi-objective binder design paradigm

arXiv.org Artificial Intelligence

We present a multi-objective binder design paradigm based on instruction fine-tuning and direct preference optimization (DPO) of autoregressive protein language models (pLMs). Multiple design objectives are encoded in the language model through direct optimization on expert curated preference sequence datasets comprising preferred and dispreferred distributions. We show the proposed alignment strategy enables ProtGPT2 to effectively design binders conditioned on specified receptors and a drug developability criterion. Generated binder samples demonstrate median isoelectric point (pI) improvements by $17\%-60\%$.


Benchmarking the Text-to-SQL Capability of Large Language Models: A Comprehensive Evaluation

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have emerged as a powerful tool in advancing the Text-to-SQL task, significantly outperforming traditional methods. Nevertheless, as a nascent research field, there is still no consensus on the optimal prompt templates and design frameworks. Additionally, existing benchmarks inadequately explore the performance of LLMs across the various sub-tasks of the Text-to-SQL process, which hinders the assessment of LLMs' cognitive capabilities and the optimization of LLM-based solutions. To address the aforementioned issues, we firstly construct a new dataset designed to mitigate the risk of overfitting in LLMs. Then we formulate five evaluation tasks to comprehensively assess the performance of diverse methods across various LLMs throughout the Text-to-SQL process. Our study highlights the performance disparities among LLMs and proposes optimal in-context learning solutions tailored to each task. These findings offer valuable insights for facilitating the development of LLM-based Text-to-SQL systems.


Levels of AI Agents: from Rules to Large Language Models

arXiv.org Artificial Intelligence

AI agents are defined as artificial entities to perceive the environment, make decisions and take actions. Inspired by the 6 levels of autonomous driving by Society of Automotive Engineers, the AI agents are also categorized based on utilities and strongness, as the following levels: L0, no AI, with tools taking into account perception plus actions; L1, using rule-based AI; L2, making rule-based AI replaced by IL/RL-based AI, with additional reasoning & decision making; L3, applying LLM-based AI instead of IL/RL-based AI, additionally setting up memory & reflection; L4, based on L3, facilitating autonomous learning & generalization; L5, based on L4, appending personality of emotion and character and collaborative behavior with multi-agents.


Prompt Mining for Language-based Human Mobility Forecasting

arXiv.org Artificial Intelligence

With the advancement of large language models, language-based forecasting has recently emerged as an innovative approach for predicting human mobility patterns. The core idea is to use prompts to transform the raw mobility data given as numerical values into natural language sentences so that the language models can be leveraged to generate the description for future observations. However, previous studies have only employed fixed and manually designed templates to transform numerical values into sentences. Since the forecasting performance of language models heavily relies on prompts, using fixed templates for prompting may limit the forecasting capability of language models. In this paper, we propose a novel framework for prompt mining in language-based mobility forecasting, aiming to explore diverse prompt design strategies. Specifically, the framework includes a prompt generation stage based on the information entropy of prompts and a prompt refinement stage to integrate mechanisms such as the chain of thought. Experimental results on real-world large-scale data demonstrate the superiority of generated prompts from our prompt mining pipeline. Additionally, the comparison of different prompt variants shows that the proposed prompt refinement process is effective. Our study presents a promising direction for further advancing language-based mobility forecasting.


Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies

arXiv.org Artificial Intelligence

Neural networks have become a cornerstone of machine learning. As the trend for these to get more and more complex continues, so does the underlying hardware and software infrastructure for training and deployment. In this survey we answer three research questions: "What types of model parallelism exist?", "What are the challenges of model parallelism?", and "What is a modern use-case of model parallelism?" We answer the first question by looking at how neural networks can be parallelised and expressing these as operator graphs while exploring the available dimensions. The dimensions along which neural networks can be parallelised are intra-operator and inter-operator. We answer the second question by collecting and listing both implementation challenges for the types of parallelism, as well as the problem of optimally partitioning the operator graph. We answer the last question by collecting and listing how parallelism is applied in modern multi-billion parameter transformer networks, to the extend that this is possible with the limited information shared about these networks.