Large Language Model
Label Alignment and Reassignment with Generalist Large Language Model for Enhanced Cross-Domain Named Entity Recognition
Named entity recognition on the in-domain supervised and few-shot settings have been extensively discussed in the NLP community and made significant progress. However, cross-domain NER, a more common task in practical scenarios, still poses a challenge for most NER methods. Previous research efforts in that area primarily focus on knowledge transfer such as correlate label information from source to target domains but few works pay attention to the problem of label conflict. In this study, we introduce a label alignment and reassignment approach, namely LAR, to address this issue for enhanced cross-domain named entity recognition, which includes two core procedures: label alignment between source and target domains and label reassignment for type inference. The process of label reassignment can significantly be enhanced by integrating with an advanced large-scale language model such as ChatGPT. We conduct an extensive range of experiments on NER datasets involving both supervised and zero-shot scenarios. Empirical experimental results demonstrate the validation of our method with remarkable performance under the supervised and zero-shot out-of-domain settings compared to SOTA methods.
MathViz-E: A Case-study in Domain-Specialized Tool-Using Agents
Bulusu, Arya, Man, Brandon, Jagmohan, Ashish, Vempaty, Aditya, Mari-Wyka, Jennifer, Akkil, Deepak
There has been significant recent interest in harnessing LLMs to control software systems through multi-step reasoning, planning and tool-usage. While some promising results have been obtained, application to specific domains raises several general issues including the control of specialized domain tools, the lack of existing datasets for training and evaluation, and the non-triviality of automated system evaluation and improvement. In this paper, we present a case-study where we examine these issues in the context of a specific domain. Specifically, we present an automated math visualizer and solver system for mathematical pedagogy. The system orchestrates mathematical solvers and math graphing tools to produce accurate visualizations from simple natural language commands. We describe the creation of specialized data-sets, and also develop an auto-evaluator to easily evaluate the outputs of our system by comparing them to ground-truth expressions. We have open sourced the data-sets and code for the proposed system.
AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game
Chi, Yizhou, Mao, Lingjun, Tang, Zineng
Strategic social deduction games serve as valuable testbeds for evaluating the understanding and inference skills of language models, offering crucial insights into social science, artificial intelligence, and strategic gaming. This paper focuses on creating proxies of human behavior in simulated environments, with Among Us utilized as a tool for studying simulated human behavior. The study introduces a text-based game environment, named AmongAgents, that mirrors the dynamics of Among Us. Players act as crew members aboard a spaceship, tasked with identifying impostors who are sabotaging the ship and eliminating the crew. Within this environment, the behavior of simulated language agents is analyzed. The experiments involve diverse game sequences featuring different configurations of Crewmates and Impostor personality archetypes. Our work demonstrates that state-of-the-art large language models (LLMs) can effectively grasp the game rules and make decisions based on the current context. This work aims to promote further exploration of LLMs in goal-oriented games with incomplete information and complex action spaces, as these settings offer valuable opportunities to assess language model performance in socially driven scenarios.
Bailicai: A Domain-Optimized Retrieval-Augmented Generation Framework for Medical Applications
Long, Cui, Liu, Yongbin, Ouyang, Chunping, Yu, Ying
Large Language Models (LLMs) have exhibited remarkable proficiency in natural language understanding, prompting extensive exploration of their potential applications across diverse domains. In the medical domain, open-source LLMs have demonstrated moderate efficacy following domain-specific fine-tuning; however, they remain substantially inferior to proprietary models such as GPT-4 and GPT-3.5. These open-source models encounter limitations in the comprehensiveness of domain-specific knowledge and exhibit a propensity for 'hallucinations' during text generation. To mitigate these issues, researchers have implemented the Retrieval-Augmented Generation (RAG) approach, which augments LLMs with background information from external knowledge bases while preserving the model's internal parameters. However, document noise can adversely affect performance, and the application of RAG in the medical field remains in its nascent stages. This study presents the Bailicai framework: a novel integration of retrieval-augmented generation with large language models optimized for the medical domain. The Bailicai framework augments the performance of LLMs in medicine through the implementation of four sub-modules. Experimental results demonstrate that the Bailicai approach surpasses existing medical domain LLMs across multiple medical benchmarks and exceeds the performance of GPT-3.5. Furthermore, the Bailicai method effectively attenuates the prevalent issue of hallucinations in medical applications of LLMs and ameliorates the noise-related challenges associated with traditional RAG techniques when processing irrelevant or pseudo-relevant documents.
Meta launches open-source AI app 'competitive' with closed rivals
Meta has claimed that its new artificial intelligence model is the first open-source system that will rival products from competitors such as OpenAI and Anthropic. In a blogpost, the company said its new model, with the unwieldy name of Llama 3.1 405B, "is competitive" with others โ including those from OpenAI and Anthropic โ "across a range of tasks". If true, it would mean that for the first time, one of the most powerful AI models in the world is available without an intermediary charging for access โ or controlling what its technology is used for. "Developers can fully customise the models for their needs and applications, train on new datasets, and conduct additional fine-tuning," Meta said. "This enables the broader developer community and the world to more fully realise the power of generative AI. Developers can fully customise for their applications and run in any environment โฆ all without sharing data with Meta."
Meta's New Llama 3.1 AI Model Is Free, Powerful, and Risky
Most tech moguls hope to sell artificial intelligence to the masses. But Mark Zuckerberg is giving away what Meta considers to be one of the world's best AI models for free. Meta released the biggest, most capable version of a large language model called Llama on Monday, free of charge. Meta has not disclosed the cost of developing Llama 3.1 but Zuckerberg recently told investors that his company is spending billions on AI development. Through this latest release, Meta is showing that the closed approach favored by most AI companies is not the only way to develop AI.
Meta AI is now available in Spanish, Portugese, French and more
Meta AI launched in September 2023 using the Llama 2 learning language model. Nearly a year later, Meta has announced a new round of features for its AI assistant and a fresh LLM to support it: Llama 3.1. These updates include an expansion of who can access Meta AI. Thanks to the addition of Argentina, Chile, Colombia, Ecuador, Mexico, Peru and Cameroon, the assistant is now available in 22 countries. However, some of the new features are location or language-specific for the time being.
Llama 3.1 is Meta's latest salvo in the battle for AI dominance
Meta on Tuesday announced the release of Llama 3.1, the latest version of its large language model that the company claims now rivals competitors from OpenAI and Anthropic. The new model comes just three months after Meta launched Llama 3 by integrating it into Meta AI, a chatbot that now lives in Facebook, Messenger, Instagram and WhatsApp and also powers the company's smart glasses. In the interim, OpenAI and Anthropic already released new versions of their own AI models, a sign that Silicon Valley's AI arms race isn't slowing down any time soon. In a blog post, Meta said that the new model, called Llama 3.1 405B, is the first openly available model that can compete available rivals in general knowledge, math skills and translating across multiple languages. The model was trained on more than 16,000 NVIDIA H100 GPUs, currently the fastest available chips that cost roughly 25,000 each, and can beat rivals on over 150 benchmarks, Meta claimed.
How to access Chinese LLM chatbots across the world
For users in the West, finding these Chinese models and trying them out can feel challenging, owing to language barriers and registration requirements. And indeed, there are still hoops to jump through if you don't have a valid Chinese phone number. But in fact, a lot of the chatbots support conversations in English and are surprisingly easy to access. Whether you're just curious to find out how well they perform or want to conduct more serious experiments for work, there are lots of ways to access Chinese LLM-powered chatbots. Here's how anyone can try one out in minutes.
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
Li, Zhuowan, Li, Cheng, Zhang, Mingyang, Mei, Qiaozhu, Bendersky, Michael
Retrieval Augmented Generation (RAG) has been a powerful tool for Large Language Models (LLMs) to efficiently process overly lengthy contexts. However, recent LLMs like Gemini-1.5 and GPT-4 show exceptional capabilities to understand long contexts directly. We conduct a comprehensive comparison between RAG and long-context (LC) LLMs, aiming to leverage the strengths of both. We benchmark RAG and LC across various public datasets using three latest LLMs. Results reveal that when resourced sufficiently, LC consistently outperforms RAG in terms of average performance. However, RAG's significantly lower cost remains a distinct advantage. Based on this observation, we propose Self-Route, a simple yet effective method that routes queries to RAG or LC based on model self-reflection. Self-Route significantly reduces the computation cost while maintaining a comparable performance to LC. Our findings provide a guideline for long-context applications of LLMs using RAG and LC.