Goto

Collaborating Authors

 Large Language Model


Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing

arXiv.org Artificial Intelligence

With the rapid development of LLMs, it is natural to ask how to harness their capabilities efficiently. In this paper, we explore whether it is feasible to direct each input query to a single most suitable LLM. To this end, we propose LLM routing for challenging reasoning tasks. Our extensive experiments suggest that such routing shows promise but is not feasible in all scenarios, so more robust approaches should be investigated to fill this gap.


Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

arXiv.org Artificial Intelligence

The alignments of reasoning abilities between smaller and larger Language Models are largely conducted via Supervised Fine-Tuning (SFT) using demonstrations generated from robust Large Language Models (LLMs). Although these approaches deliver more performant models, they do not show sufficiently strong generalization ability as the training only relies on the provided demonstrations. In this paper, we propose the Self-refine Instruction-tuning method that elicits Smaller Language Models to self-refine their abilities. Our approach is based on a two-stage process, where reasoning abilities are first transferred between LLMs and Small Language Models (SLMs) via Instruction-tuning on demonstrations provided by LLMs, and then the instructed models Self-refine their abilities through preference optimization strategies. In particular, the second phase operates refinement heuristics based on the Direct Preference Optimization algorithm, where the SLMs are elicited to deliver a series of reasoning paths by automatically sampling the generated responses and providing rewards using ground truths from the LLMs. Results obtained on commonsense and math reasoning tasks show that this approach significantly outperforms Instruction-tuning in both in-domain and out-domain scenarios, aligning the reasoning abilities of Smaller and Larger Language Models.


RST-LoRA: A Discourse-Aware Low-Rank Adaptation for Long Document Abstractive Summarization

arXiv.org Artificial Intelligence

For long document summarization, discourse structure is important to discern the key content of the text and the differences in importance level between sentences. Unfortunately, the integration of rhetorical structure theory (RST) into parameter-efficient fine-tuning strategies for long document summarization remains unexplored. Therefore, this paper introduces RST-LoRA and proposes four RST-aware variants to explicitly incorporate RST into the LoRA model. Our empirical evaluation demonstrates that incorporating the type and uncertainty of rhetorical relations can complementarily enhance the performance of LoRA in summarization tasks. Furthermore, the best-performing variant we introduced outperforms the vanilla LoRA and full-parameter fine-tuning models, as confirmed by multiple automatic and human evaluations, and even surpasses previous state-of-the-art methods.


Eight US newspapers sue OpenAI and Microsoft for copyright infringement

The Guardian

The New York Daily News, Chicago Tribune, Denver Post and other papers filed the lawsuit on Tuesday in a New York federal court. "We've spent billions of dollars gathering information and reporting news at our publications, and we can't allow OpenAI and Microsoft to expand the Big Tech playbook of stealing our work to build their own businesses at our expense," said a written statement from Frank Pine, executive editor for the MediaNews Group and Tribune Publishing. The other newspapers that are part of the lawsuit are MediaNews Group's Mercury News, Denver Post, Orange County Register and St Paul Pioneer-Press, and Tribune Publishing's Orlando Sentinel and South Florida Sun Sentinel. All of the newspapers are owned by Alden Global Capital. Microsoft declined to comment on Tuesday.


8 major newspapers join legal backlash against OpenAI, Microsoft

Washington Post - Technology News

The publications were joined in the suit by South Florida's Sun Sentinel, the Denver Post, Orange County (Calif.) The lawsuit alleges that OpenAI and Microsoft used their news articles to train and run their AI tools, including OpenAI's ChatGPT. All eight newspapers are owned by New York City-based hedge fund Alden Global Capital.


There's an AI Lobbying Frenzy in Washington. Big Tech Is Dominating

TIME - Tech

The number of groups lobbying the U.S. federal government on artificial intelligence nearly tripled from 2022 to 2023, rocketing from 158 to 451 organizations, according to data from OpenSecrets, a nonprofit that tracks and publishes data on campaign finance and lobbying. Data on the total amount spent on lobbying by each organization and interviews with two congressional staffers, two nonprofit advocates familiar with AI lobbying efforts, and two named experts suggest that large technology companies have so far dominated efforts to influence potential AI legislation. And while these companies have publicly been supportive of AI regulation, in closed-door conversations with officials they tend to push for light-touch and voluntary rules, say Congressional staffers and advocates. In November 2022, OpenAI released its wildly popular chatbot, ChatGPT. Six months later, leading AI researchers and industry executives signed a statement warning that "the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."


AI Detectors for ChatGPT: Everything You Need to Know

WIRED

Detecting when text has been generated by tools like ChatGPT is a difficult task. Popular artificial- intelligence-detection tools, like GPTZero, may provide some guidance for users by telling them when something was written by a bot and not a human, but even specialized software is not foolproof and can spit out false positives. As a journalist who started covering AI detection over a year ago, I wanted to curate some of WIRED's best articles on the topic to help readers like you better understand this complicated issue. Have even more questions about spotting outputs from ChatGPT and other chatbot tools? Sign up for my AI Unlocked newsletter, and reach out to me directly with anything AI-related that you would like answered or want WIRED to explore more.


Break down YouTube videos to the gist with this tool -- now 145 off

PCWorld

YouTube can be a great resource to stay up to date on industry developments or learn new skills. But it can also be incredibly time-consuming to watch videos all day. That's where a TubeOnAi Premium Lite Plan can help. This clever tool downloads transcripts of videos from the channel you follow and summarizes the content by leveraging GPT-4 AI. Right now, you can get a lifetime subscription to this time-saving tool for just 78.99.


Correcting misinformation on social media with a large language model

arXiv.org Artificial Intelligence

Real-world misinformation can be partially correct and even factual but misleading. It undermines public trust in science and democracy, particularly on social media, where it can spread rapidly. High-quality and timely correction of misinformation that identifies and explains its (in)accuracies has been shown to effectively reduce false beliefs. Despite the wide acceptance of manual correction, it is difficult to be timely and scalable, a concern as technologies like large language models (LLMs) make misinformation easier to produce. LLMs also have versatile capabilities that could accelerate misinformation correction-however, they struggle due to a lack of recent information, a tendency to produce false content, and limitations in addressing multimodal information. We propose MUSE, an LLM augmented with access to and credibility evaluation of up-to-date information. By retrieving evidence as refutations or contexts, MUSE identifies and explains (in)accuracies in a piece of content-not presupposed to be misinformation-with references. It also describes images and conducts multimodal searches to verify and correct multimodal content. Fact-checking experts evaluate responses to social media content that are not presupposed to be (non-)misinformation but broadly include incorrect, partially correct, and correct posts, that may or may not be misleading. We propose and evaluate 13 dimensions of misinformation correction quality, ranging from the accuracy of identifications and factuality of explanations to the relevance and credibility of references. The results demonstrate MUSE's ability to promptly write high-quality responses to potential misinformation on social media-overall, MUSE outperforms GPT-4 by 37% and even high-quality responses from laypeople by 29%. This work reveals LLMs' potential to help combat real-world misinformation effectively and efficiently.


Artificial intelligence and machine learning applications for cultured meat

arXiv.org Artificial Intelligence

Cultured meat has the potential to provide a complementary meat industry with reduced environmental, ethical, and health impacts. However, major technological challenges remain which require time- and resource-intensive research and development efforts. Machine learning has the potential to accelerate cultured meat technology by streamlining experiments, predicting optimal results, and reducing experimentation time and resources. However, the use of machine learning in cultured meat is in its infancy. This review covers the work available to date on the use of machine learning in cultured meat and explores future possibilities. We address four major areas of cultured meat research and development: establishing cell lines, cell culture media design, microscopy and image analysis, and bioprocessing and food processing optimization. This review aims to provide the foundation necessary for both cultured meat and machine learning scientists to identify research opportunities at the intersection between cultured meat and machine learning.