Goto

Collaborating Authors

 Large Language Model


ChatGPT's desktop app finally comes to Windows, with features missing

PCWorld

Even though Microsoft is a heavy investor in OpenAI, the company behind ChatGPT chose to first release the AI chatbot's desktop app on macOS before Windows back in May of this year. Now, the wait is over for Windows users. The ChatGPT desktop app is finally available on Windows, but with an important caveat: this is an "early version" for paid subscribers who are part of ChatGPT's Plus, Team, Enterprise, or Edu plans. Once you install the app, all you need to do is press the Alt Space keyboard shortcut to launch a new conversation with ChatGPT. The desktop app has access to OpenAI's latest AI models, and you can perform all the core tasks you'd expect from ChatGPT, including asking it questions, having it analyze images, and uploading files to it.


The Download: AI for debates, and what to know about the Oropouche virus

MIT Technology Review

Reaching a consensus in a democracy is difficult because people hold such different ideological, political, and social views. Perhaps an AI tool could help. Researchers from Google DeepMind trained a system of large language models to operate as a "caucus mediator," generating summaries that outline a group's areas of agreement on complex but important social or political issues. The researchers say their work highlights the potential of AI to help groups of people find common ground when discussing contentious subjects. But it's not going to replace human mediators anytime soon.


Silicon Valley Takes Artificial General Intelligence Seriously--Washington Must Too

TIME - Tech

Artificial General Intelligence--machines that can learn and perform any cognitive task that a human can--has long been relegated to the realm of science fiction. But recent developments show that AGI is no longer a distant speculation; it's an impending reality that demands our immediate attention. On Sept. 17, during a Senate Judiciary Subcommittee hearing titled "Oversight of AI: Insiders' Perspectives," whistleblowers from leading AI companies sounded the alarm on the rapid advancement toward AGI and the glaring lack of oversight. Helen Toner, a former board member of OpenAI and director of strategy at Georgetown University's Center for Security and Emerging Technology, testified that, "The biggest disconnect that I see between AI insider perspectives and public perceptions of AI companies is when it comes to the idea of artificial general intelligence." She continued that leading AI companies such as OpenAI, Google, and Anthropic are "treating building AGI as an entirely serious goal."


Here's the deal: AI giants get to grab all your data unless you say they can't. Fancy that? No, neither do I Chris Stokel-Walker

The Guardian

Imagine someone drives up to a pub in a top-of-the-range sports car โ€“ a 1.5m Koenigsegg Regera, to pick one at random โ€“ parks up and saunters out of the vehicle. They come into the pub you're drinking in and begin walking around its patrons, slipping their hand into your pocket in full view, smiling at you as they take out your wallet and empty it of its cash and cards. The not-so-subtle pickpocket stops if you shout and ask what the hell they're doing. "Sorry for the inconvenience," the pickpocket says. Yet it seems to be the approach the government is pursuing in order to placate AI companies. A consultation is soon to open, the Financial Times reports, that will allow AI companies to scrape content from individuals and organisations unless they explicitly opt out of their data being used.


Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum

arXiv.org Artificial Intelligence

There is increasing interest in looking at dialects in NLP. However, most work to date still treats dialects as discrete categories. For instance, evaluative work in variation-oriented NLP for English often works with Indian English or African-American Venacular English as homogeneous categories (Faisal et al., 2024; Ziems et al., 2023), yet even within one variety there is substantial variation. We examine within-dialect variation and show that performance critically varies within categories. We measure speech-to-text performance on Italian dialects, and empirically observe a geographical performance disparity. This disparity correlates substantially (-0.5) with linguistic similarity to the highest performing dialect variety. We cross-examine our results against dialectometry methods, and interpret the performance disparity to be due to a bias towards dialects that are more similar to the standard variety in the speech-to-text model examined. We additionally leverage geostatistical methods to predict zero-shot performance at unseen sites, and find the incorporation of geographical information to substantially improve prediction performance, indicating there to be geographical structure in the performance distribution.


DFlow: Diverse Dialogue Flow Simulation with Large Language Models

arXiv.org Artificial Intelligence

Developing language model-based dialogue agents requires effective data to train models that can follow specific task logic. However, most existing data augmentation methods focus on increasing diversity in language, topics, or dialogue acts at the utterance level, largely neglecting a critical aspect of task logic diversity at the dialogue level. This paper proposes a novel data augmentation method designed to enhance the diversity of synthetic dialogues by focusing on task execution logic. Our method uses LLMs to generate decision tree-structured task plans, which enables the derivation of diverse dialogue trajectories for a given task. Each trajectory, referred to as a "dialog flow", guides the generation of a multi-turn dialogue that follows a unique trajectory. We apply this method to generate a task-oriented dialogue dataset comprising 3,886 dialogue flows across 15 different domains. We validate the effectiveness of this dataset using the next action prediction task, where models fine-tuned on our dataset outperform strong baselines, including GPT-4. Upon acceptance of this paper, we plan to release the code and data publicly.


Critical Questions Generation: Motivation and Challenges

arXiv.org Artificial Intelligence

The development of Large Language Models (LLMs) has brought impressive performances on mitigation strategies against misinformation, such as counterargument generation. However, LLMs are still seriously hindered by outdated knowledge and by their tendency to generate hallucinated content. In order to circumvent these issues, we propose a new task, namely, Critical Questions Generation, consisting of processing an argumentative text to generate the critical questions (CQs) raised by it. In argumentation theory CQs are tools designed to lay bare the blind spots of an argument by pointing at the information it could be missing. Thus, instead of trying to deploy LLMs to produce knowledgeable and relevant counterarguments, we use them to question arguments, without requiring any external knowledge. Research on CQs Generation using LLMs requires a reference dataset for large scale experimentation. Thus, in this work we investigate two complementary methods to create such a resource: (i) instantiating CQs templates as defined by Walton's argumentation theory and (ii), using LLMs as CQs generators. By doing so, we contribute with a procedure to establish what is a valid CQ and conclude that, while LLMs are reasonable CQ generators, they still have a wide margin for improvement in this task.


Tell me what I need to know: Exploring LLM-based (Personalized) Abstractive Multi-Source Meeting Summarization

arXiv.org Artificial Intelligence

Meeting summarization is crucial in digital communication, but existing solutions struggle with salience identification to generate personalized, workable summaries, and context understanding to fully comprehend the meetings' content. Previous attempts to address these issues by considering related supplementary resources (e.g., presentation slides) alongside transcripts are hindered by models' limited context sizes and handling the additional complexities of the multi-source tasks, such as identifying relevant information in additional files and seamlessly aligning it with the meeting content. This work explores multi-source meeting summarization considering supplementary materials through a three-stage large language model approach: identifying transcript passages needing additional context, inferring relevant details from supplementary materials and inserting them into the transcript, and generating a summary from this enriched transcript. Our multi-source approach enhances model understanding, increasing summary relevance by ~9% and producing more content-rich outputs. We introduce a personalization protocol that extracts participant characteristics and tailors summaries accordingly, improving informativeness by ~10%. This work further provides insights on performance-cost trade-offs across four leading model families, including edge-device capable options. Our approach can be extended to similar complex generative tasks benefitting from additional resources and personalization, such as dialogue systems and action planning.


SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment

arXiv.org Artificial Intelligence

Existing preference alignment is a one-size-fits-all alignment mechanism, where the part of the large language model (LLM) parametric knowledge with non-preferred features is uniformly blocked to all the users. However, this part of knowledge can be useful to advanced users whose expertise qualifies them to handle these information. The one-size-fits-all alignment mechanism undermines LLM's utility for these qualified users. To address this problem, we propose SudoLM, a framework that lets LLMs learn access control over specific parametric knowledge for users with different credentials via authorization alignment. SudoLM allows authorized users to unlock their access to all the parametric knowledge with an assigned SUDO key while blocking access to non-qualified users. Experiments on two application scenarios demonstrate that SudoLM effectively controls the user's access to the parametric knowledge and maintains its general utility.


Large Language Models Are Overparameterized Text Encoders

arXiv.org Artificial Intelligence

Large language models (LLMs) demonstrate strong performance as text embedding models when finetuned with supervised contrastive training. However, their large size balloons inference time and memory requirements. In this paper, we show that by pruning the last $p\%$ layers of an LLM before supervised training for only 1000 steps, we can achieve a proportional reduction in memory and inference time. We evaluate four different state-of-the-art LLMs on text embedding tasks and find that our method can prune up to 30\% of layers with negligible impact on performance and up to 80\% with only a modest drop. With only three lines of code, our method is easily implemented in any pipeline for transforming LLMs to text encoders. We also propose $\text{L}^3 \text{Prune}$, a novel layer-pruning strategy based on the model's initial loss that provides two optimal pruning configurations: a large variant with negligible performance loss and a small variant for resource-constrained settings. On average, the large variant prunes 21\% of the parameters with a $-0.3$ performance drop, and the small variant only suffers from a $-5.1$ decrease while pruning 74\% of the model. We consider these results strong evidence that LLMs are overparameterized for text embedding tasks, and can be easily pruned.