Goto

Collaborating Authors

 Large Language Model


Self-Directed Synthetic Dialogues and Revisions Technical Report

arXiv.org Artificial Intelligence

Synthetic data has become an important tool in the fine-tuning of language models to follow instructions and solve complex problems. Nevertheless, the majority of open data to date is often lacking multi-turn data and collected on closed models, limiting progress on advancing open fine-tuning methods. We introduce Self Directed Synthetic Dialogues (SDSD), an experimental dataset consisting of guided conversations of language models talking to themselves. The dataset consists of multi-turn conversations generated with DBRX, Llama 2 70B, and Mistral Large, all instructed to follow a conversation plan generated prior to the conversation. We also explore including principles from Constitutional AI and other related works to create synthetic preference data via revisions to the final conversation turn. We hope this work encourages further exploration in multi-turn data and the use of open models for expanding the impact of synthetic data.


Mark Zuckerberg Just Intensified the Battle for AI's Future

TIME - Tech

The tech industry is currently embroiled in a heated debate over the future of AI: should powerful systems be open-source and freely accessible, or closed and tightly monitored for dangers? On Tuesday, Meta CEO Mark Zuckerberg fired a salvo into this ongoing battle, publishing not just a new series of powerful AI models, but also a manifesto forcefully advocating for the open-source approach. The document, which was widely praised by venture capitalists and tech leaders like Elon Musk and Jack Dorsey, serves as both a philosophical treatise and a rallying cry for proponents of open-source AI development. It arrives as intensifying global efforts to regulate AI have galvanized resistance from open-source advocates, who see some of those potential laws as threats to innovation and accessibility. At the heart of Meta's announcement on Tuesday was the release of its latest generation of Llama large language models, the company's answer to ChatGPT.


AI trained on AI garbage spits out AI garbage

MIT Technology Review

Current AI models aren't just going to collapse, says Shumailov, but there may still be substantive effects: The improvements will slow down, and performance might suffer. To determine the potential effect on performance, Shumailov and his colleagues fine-tuned a large language model (LLM) on a set of data from Wikipedia, then fine-tuned the new model on its own output over nine generations. The team measured how nonsensical the output was using a "perplexity score," which measures an AI model's confidence in its ability to predict the next part of a sequence; a higher score translates to a less accurate model. The models trained on other models' outputs had higher perplexity scores. "some started before 1360--was typically accomplished by a master mason and a small team of itinerant masons, supplemented by local parish labourers, according to Poyntz Wright. But other authors reject this model, suggesting instead that leading architects designed the parish church towers based on early examples of Perpendicular."


The Download: Chinese LLMs, and transforming heavy-duty trucking

MIT Technology Review

When police departments first started buying and deploying bodycams in the wake of the police killing of Michael Brown in Ferguson, Missouri, a decade ago, activists hoped it would bring about real change. Years later, despite what's become a multibillion-dollar market for these devices, the tech is far from a panacea. Most footage they generate goes unwatched. And if they do finally provide video to the public, it usually doesn't tell the complete story. A handful of AI startups see this problem as an opportunity to create what are essentially bodycam-to-text programs for different players in the legal system, mining this footage for misdeeds. But like the bodycams themselves, the technology still faces procedural, legal, and cultural barriers to success.


The Morning After: Netflix's new gaming boss is a former Epic Games exec

Engadget

Netflix has hired Alain Tascan as its new president of games. Before joining Netflix, Tascan was executive vice president for Epic Games and oversaw first-party development for some of the company's (and gaming's) most successful titles, like Fortnite, Rocket League and Fall Guys. Since launching its games project in 2021, Netflix has acquired notable indie studios Night School, Boss Fight, Next Games and Spry Fox and has brought many great indie games to mobile -- seriously, search the app store, if only for Into The Breach. Netflix recently said it has 80-plus games currently in development. A multiplayer Squid Game project will be part of that, coinciding with the hit show's next season, later this year.


AI Testing Mostly Uses English Right Now. That's Risky

TIME - Tech

Over the last year, governments, academia, and industry have invested considerable resources into investigating the harms of advanced AI. But one massive factor seems to be continuously overlooked: right now, AI's primary tests and models are confined to English. Advanced AI could be used in many languages to cause harm, but focusing primarily on English may leave us with only part of the answer. It also ignores those most vulnerable to its harms. After the release of ChatGPT in November, 2022, AI developers expressed surprise at a capability displayed by the model: It could "speak" at least 80 languages, not just English.


Why Chinese companies are betting on open-source AI

MIT Technology Review

The good news is it's actually not that hard! I recently dug around and realized that many Chinese AI models are much more accessible overseas than I expected. You can access the majority of them either by registering accounts on their websites or using popular open-source AI platforms like Hugging Face. So I published this practical guide today that lists a dozen of the top Chinese LLM chatbots you can use and the methods to easily access them in minutes, from anywhere in the world. During my experiments with these models, one thing soon became clear: While most Chinese AI companies have set a higher bar for access to their products than their Western counterparts, a trend toward open-sourcing AI models is making them ever more accessible to an overseas audience.


Traditional Methods Outperform Generative LLMs at Forecasting Credit Ratings

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have been shown to perform well for many downstream tasks. Transfer learning can enable LLMs to acquire skills that were not targeted during pre-training. In financial contexts, LLMs can sometimes beat well-established benchmarks. This paper investigates how well LLMs perform in the task of forecasting corporate credit ratings. We show that while LLMs are very good at encoding textual information, traditional methods are still very competitive when it comes to encoding numeric and multimodal data. For our task, current LLMs perform worse than a more traditional XGBoost architecture that combines fundamental and macroeconomic data with high-density text-based embedding features.


Using Large Language Models to Compare Explainable Models for Smart Home Human Activity Recognition

arXiv.org Artificial Intelligence

Recognizing daily activities with unobtrusive sensors in smart environments enables various healthcare applications. Monitoring how subjects perform activities at home and their changes over time can reveal early symptoms of health issues, such as cognitive decline. Most approaches in this field use deep learning models, which are often seen as black boxes mapping sensor data to activities. However, non-expert users like clinicians need to trust and understand these models' outputs. Thus, eXplainable AI (XAI) methods for Human Activity Recognition have emerged to provide intuitive natural language explanations from these models. Different XAI methods generate different explanations, and their effectiveness is typically evaluated through user surveys, that are often challenging in terms of costs and fairness. This paper proposes an automatic evaluation method using Large Language Models (LLMs) to identify, in a pool of candidates, the best XAI approach for non-expert users. Our preliminary results suggest that LLM evaluation aligns with user surveys.


Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge

arXiv.org Artificial Intelligence

Human-produced emissions are growing at an alarming rate, causing already observable changes in the climate and environment in general. Each year global carbon dioxide emissions hit a new record, and it is reported that 0.5% of total US greenhouse gas emissions are attributed to data centres as of 2021. The release of ChatGPT in late 2022 sparked social interest in Large Language Models (LLMs), the new generation of Language Models with a large number of parameters and trained on massive amounts of data. Currently, numerous companies are releasing products featuring various LLMs, with many more models in development and awaiting release. Deep Learning research is a competitive field, with only models that reach top performance attracting attention and being utilized. Hence, achieving better accuracy and results is often the first priority, while the model's efficiency and the environmental impact of the study are neglected. However, LLMs demand substantial computational resources and are very costly to train, both financially and environmentally. It becomes essential to raise awareness and promote conscious decisions about algorithmic and hardware choices. Providing information on training time, the approximate carbon dioxide emissions and power consumption would assist future studies in making necessary adjustments and determining the compatibility of available computational resources with model requirements. In this study, we infused T5 LLM with external knowledge and fine-tuned the model for Question-Answering task. Furthermore, we calculated and reported the approximate environmental impact for both steps. The findings demonstrate that the smaller models may not always be sustainable options, and increased training does not always imply better performance. The most optimal outcome is achieved by carefully considering both performance and efficiency factors.