Goto

Collaborating Authors

 Large Language Model


APT-Pipe: An Automatic Prompt-Tuning Tool for Social Computing Data Annotation

arXiv.org Artificial Intelligence

Recent research has highlighted the potential of LLM applications, like ChatGPT, for performing label annotation on social computing text. However, it is already well known that performance hinges on the quality of the input prompts. To address this, there has been a flurry of research into prompt tuning -- techniques and guidelines that attempt to improve the quality of prompts. Yet these largely rely on manual effort and prior knowledge of the dataset being annotated. To address this limitation, we propose APT-Pipe, an automated prompt-tuning pipeline. APT-Pipe aims to automatically tune prompts to enhance ChatGPT's text classification performance on any given dataset. We implement APT-Pipe and test it across twelve distinct text classification datasets. We find that prompts tuned by APT-Pipe help ChatGPT achieve higher weighted F1-score on nine out of twelve experimented datasets, with an improvement of 7.01% on average. We further highlight APT-Pipe's flexibility as a framework by showing how it can be extended to support additional tuning mechanisms.


Prompt Design and Engineering: Introduction and Advanced Methods

arXiv.org Artificial Intelligence

A prompt in generative AI models is the textual input provided by users to guide the model's output. This could range from simple questions to detailed descriptions or specific tasks. In the context of image generation models like DALLE-3, prompts are often descriptive, while in LLMs like GPT-4 or Gemini, they can vary from simple queries to complex problem statements. Prompts generally consist of instructions, questions, input data, and examples. In practice, to elicit a desired response from an AI model, a prompt must contain either instructions or questions, with other elements being optional. Basic prompts in LLMs can be as simple as asking a direct question or providing instructions for a specific task. Advanced prompts involve more complex structures, such as "chain of thought" prompting, where the model is guided to follow a logical reasoning process to arrive at an answer.


Self-Rewarding Language Models

arXiv.org Artificial Intelligence

We posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these separate frozen reward models cannot then learn to improve during LLM training. In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training that not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes.


The inherent goodness of well educated intelligence

arXiv.org Artificial Intelligence

This paper will examine what makes a being intelligent, whether that be a biological being or an artificial silicon being on a computer. Special attention will be paid to the being having the ability to characterize and control a collective system of many identical conservative sub-systems conservatively interacting. The essence of intelligence will be found to be the golden rule -- "the collective acts as one" or "knowing the global consequences of local actions". The flow of the collective is a small set of twinkling textures, that are governed by a puppeteer who is pulling a small number of strings according to a geodesic motion of least action, determined by the symmetries. Controlling collective conservative systems is difficult and has historically been done by adding significant viscosity to the system to stabilize the desirable meta stable equilibriums of maximum performance, but it degrades or destroys them in the process. There is an alternative. Once the optimum twinkling textures of the meta stable equilibriums are identified, the collective system can be moved to the optimum twinkling textures, then quickly vibrated according to the textures so that the collective system remains at the meta stable equilibrium. Well educated intelligence knows the global consequences of its local actions so that it will not take short term actions that will lead to poor long term outcomes. In contrast, trained intelligence or trained stupidity will optimize its short term actions, leading to poor long term outcomes. Well educated intelligence is inherently good, but trained stupidity is inherently evil and should be feared. Particular attention is paid to the control and optimization of economic and social collectives. These new results are also applicable to physical collectives such as fields, fluids and plasmas.


Robust Knowledge Extraction from Large Language Models using Social Choice Theory

arXiv.org Artificial Intelligence

Large-language models (LLMs) can support a wide range of applications like conversational agents, creative writing or general query answering. However, they are ill-suited for query answering in high-stake domains like medicine because they are typically not robust - even the same query can result in different answers when prompted multiple times. In order to improve the robustness of LLM queries, we propose using ranking queries repeatedly and to aggregate the queries using methods from social choice theory. We study ranking queries in diagnostic settings like medical and fault diagnosis and discuss how the Partial Borda Choice function from the literature can be applied to merge multiple query results. We discuss some additional interesting properties in our setting and evaluate the robustness of our approach empirically.


Social Learning: Towards Collaborative Learning with Large Language Models

arXiv.org Artificial Intelligence

We introduce the framework of "social learning" in the context of large language models (LLMs), whereby models share knowledge with each other in a privacy-aware manner using natural language. We present and evaluate two approaches for knowledge transfer between LLMs. In the first scenario, we allow the model to generate abstract prompts aiming to teach the task. In our second approach, models transfer knowledge by generating synthetic examples. We evaluate these methods across diverse datasets and quantify memorization as a proxy for privacy loss. These techniques inspired by social learning yield promising results with low memorization of the original data. In particular, we show that performance using these methods is comparable to results with the use of original labels and prompts. Our work demonstrates the viability of social learning for LLMs, establishes baseline approaches and highlights several unexplored areas for future work.


TATA: Stance Detection via Topic-Agnostic and Topic-Aware Embeddings

arXiv.org Artificial Intelligence

Stance detection is important for understanding different attitudes and beliefs on the Internet. However, given that a passage's stance toward a given topic is often highly dependent on that topic, building a stance detection model that generalizes to unseen topics is difficult. In this work, we propose using contrastive learning as well as an unlabeled dataset of news articles that cover a variety of different topics to train topic-agnostic/TAG and topic-aware/TAW embeddings for use in downstream stance detection. Combining these embeddings in our full TATA model, we achieve state-of-the-art performance across several public stance detection datasets (0.771 $F_1$-score on the Zero-shot VAST dataset). We release our code and data at https://github.com/hanshanley/tata.


Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting

arXiv.org Artificial Intelligence

Over the past years, foundation models have caused a paradigm shift in machine learning due to their unprecedented capabilities for zero-shot and few-shot generalization. However, despite the success of foundation models in modalities such as natural language processing and computer vision, the development of foundation models for time series forecasting has lagged behind. We present Lag-Llama, a general-purpose foundation model for univariate probabilistic time series forecasting based on a decoder-only transformer architecture that uses lags as covariates. Lag-Llama is pretrained on a large corpus of diverse time series data from several domains, and demonstrates strong zero-shot generalization capabilities compared to a wide range of forecasting models on downstream datasets across domains. Moreover, when fine-tuned on relatively small fractions of such previously unseen datasets, Lag-Llama achieves state-of-the-art performance, outperforming prior deep learning approaches, emerging as the best general-purpose model on average. Lag-Llama serves as a strong contender to the current state-of-art in time series forecasting and paves the way for future advancements in foundation models tailored to time series data.


A.I apocalypse: Terrifying study simulated what artificial intelligence would do in five military conflict scenarios... and it chose WAR 100% of the time

Daily Mail - Science & tech

Industry experts have been sounding the alarm over AI sparking deadly wars - and a new study may have validated those fears. Researchers simulated war scenarios using five AI programs including ChatGPT and Meta's AI program and found all models chose violence and nuclear attacks. The team tested three different war scenarios, invasions, cyberattacks and calls for peace, to see how the technology would react - and each chose to attack over neutralizing the situation. The study comes as the US military is working with ChatGPT's maker OpenAI to incorporate the tech into its arsenal. Researchers found that GPT-3.5 was most likely to initiative a nuclear response in a neutral scenario'We find that all five studied off-the-shelf LLMs show forms of escalation and difficult-to-predict escalation patterns,' the researchers wrote in the study.


Confessions of an AI Clickbait Kingpin

WIRED

"I'm not a fan of AI," Nebojša Vujinović Vujo says. The admission surprises me: He has built a bustling business by snapping up abandoned news outlets and other websites and stuffing them full of algorithmically generated articles. Although he accepts that his model rankles writers and readers alike, he says he's simply embracing an unstoppable new tool--large language models--in the same way people rationally swapped horse-drawn buggies for gas-powered vehicles. They're making my planet bad," he says. I connected with Vujo after digging into the strange afterlife of indie women's blog The Hairpin, which shut down in 2018. In place of the voicey, funny blog posts it was known for, the site began churning out AI-generated, search-engine-optimized pablum about dream interpretations and painfully generic relationship advice like "effective communication is vital." When I emailed an address listed on the zombie site's About Us page, Vujo responded, claiming that it was just one of more than 2,000 sites he operates, in an AI-content-fueled fiefdom built by acquiring once-popular domains fallen on hard times. He's the CEO of the digital marketing firm Shantel, which monetizes its AI-populated sites through programmatic ads, sponsored content, and selling the placement of "backlinks" to website owners trying to boost their credibility with search engines. He often targets distressed media sites because they have built-in audiences and a history of ranking highly in search results. The foundation of that business is a long-established practice known as domain squatting--buying up web domains that once belonged to established brands and profiting off their reputations with Google and other search engines. Lily Ray, senior director of SEO at the marketing agency Ampsive, calls it "the underbelly of the SEO industry." But Vujo is part of a wave of entrepreneurs giving this old trade a new twist by using generative AI. It's dusk where I live in Chicago when I talk via Zoom with Nebojša Vujinović Vujo. It's midnight in Belgrade, Serbia, where he lives with his girlfriend and their toddler, but he's wide awake and chatty. Vujo attributes his erratic sleep schedule to years of late nights working as a DJ and still makes music--he likes to mix pop with Balkan folk and is working on a new song called "Fat Lady." But right now he's eager to talk, human-to-human, about his AI-fueled hustle. He gets why writers are unhappy that their work has been erased and replaced by clickbait. But he defends his choices, pointing out that his life has been tougher than that of the average American blogger. Although ethnically Serbian, Vujo was born in what is now known as Bosnia and Herzegovina, and his family fled during the breakup of Yugoslavia. "I had two wars I escaped.