Goto

Collaborating Authors

 Large Language Model


MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems

arXiv.org Artificial Intelligence

Recent advancements in large language models, such as GPT-4, have demonstrated remarkable capabilities in processing standard queries. Despite these advancements, their performance substantially declines in \textbf{advanced mathematical problems requiring complex, multi-step logical reasoning}. To enhance their inferential capabilities, current research has delved into \textit{prompting engineering}, exemplified by methodologies such as the Tree of Thought and Graph of Thought. Nonetheless, these existing approaches encounter two significant limitations. Firstly, their effectiveness in tackling complex mathematical problems is somewhat constrained. Secondly, the necessity to design distinct prompts for individual problems hampers their generalizability. In response to these limitations, this paper introduces the \textit{Multi-Agent System for conditional Mining} (\textbf{MACM}) prompting method. It not only resolves intricate mathematical problems but also demonstrates strong generalization capabilities across various mathematical contexts. With the assistance of MACM, the accuracy of GPT-4 Turbo on the most challenging level five mathematical problems in the MATH dataset increase from $\mathbf{54.68\%} \text{ to } \mathbf{76.73\%}$. The code is available in \url{https://github.com/bin123apple/MACM}.


Autonomous Artificial Intelligence Agents for Clinical Decision Making in Oncology

arXiv.org Artificial Intelligence

Multimodal artificial intelligence (AI) systems have the potential to enhance clinical decision-making by interpreting various types of medical data. However, the effectiveness of these models across all medical fields is uncertain. Each discipline presents unique challenges that need to be addressed for optimal performance. This complexity is further increased when attempting to integrate different fields into a single model. Here, we introduce an alternative approach to multimodal medical AI that utilizes the generalist capabilities of a large language model (LLM) as a central reasoning engine. This engine autonomously coordinates and deploys a set of specialized medical AI tools. These tools include text, radiology and histopathology image interpretation, genomic data processing, web searches, and document retrieval from medical guidelines. We validate our system across a series of clinical oncology scenarios that closely resemble typical patient care workflows. We show that the system has a high capability in employing appropriate tools (97%), drawing correct conclusions (93.6%), and providing complete (94%), and helpful (89.2%) recommendations for individual patient cases while consistently referencing relevant literature (82.5%) upon instruction. This work provides evidence that LLMs can effectively plan and execute domain-specific models to retrieve or synthesize new information when used as autonomous agents. This enables them to function as specialist, patient-tailored clinical assistants. It also simplifies regulatory compliance by allowing each component tool to be individually validated and approved. We believe, that our work can serve as a proof-of-concept for more advanced LLM-agents in the medical domain.


PoLLMgraph: Unraveling Hallucinations in Large Language Models via State Transition Dynamics

arXiv.org Artificial Intelligence

Despite tremendous advancements in large language models (LLMs) over recent years, a notably urgent challenge for their practical deployment is the phenomenon of hallucination, where the model fabricates facts and produces non-factual statements. In response, we propose PoLLMgraph, a Polygraph for LLMs, as an effective model-based white-box detection and forecasting approach. PoLLMgraph distinctly differs from the large body of existing research that concentrates on addressing such challenges through black-box evaluations. In particular, we demonstrate that hallucination can be effectively detected by analyzing the LLM's internal state transition dynamics during generation via tractable probabilistic models. Experimental results on various open-source LLMs confirm the efficacy of PoLLMgraph, outperforming state-of-the-art methods by a considerable margin, evidenced by over 20% improvement in AUC-ROC on common benchmarking datasets like TruthfulQA. Our work paves a new way for model-based white-box analysis of LLMs, motivating the research community to further explore, understand, and refine the intricate dynamics of LLM behaviors.


Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only Shallowly

arXiv.org Artificial Intelligence

Despite their strong ability to retrieve knowledge in English, current large language models show imbalance abilities in different languages. Two approaches are proposed to address this, i.e., multilingual pretraining and multilingual instruction tuning. However, whether and how do such methods contribute to the cross-lingual knowledge alignment inside the models is unknown. In this paper, we propose CLiKA, a systematic framework to assess the cross-lingual knowledge alignment of LLMs in the Performance, Consistency and Conductivity levels, and explored the effect of multilingual pretraining and instruction tuning on the degree of alignment. Results show that: while both multilingual pretraining and instruction tuning are beneficial for cross-lingual knowledge alignment, the training strategy needs to be carefully designed. Namely, continued pretraining improves the alignment of the target language at the cost of other languages, while mixed pretraining affect other languages less. Also, the overall cross-lingual knowledge alignment, especially in the conductivity level, is unsatisfactory for all tested LLMs, and neither multilingual pretraining nor instruction tuning can substantially improve the cross-lingual knowledge conductivity.


Multicalibration for Confidence Scoring in LLMs

arXiv.org Machine Learning

This paper proposes the use of "multicalibration" to yield interpretable and reliable confidence scores for outputs generated by large language models (LLMs). Multicalibration asks for calibration not just marginally, but simultaneously across various intersecting groupings of the data. We show how to form groupings for prompt/completion pairs that are correlated with the probability of correctness via two techniques: clustering within an embedding space, and "self-annotation" - querying the LLM by asking it various yes-or-no questions about the prompt. We also develop novel variants of multicalibration algorithms that offer performance improvements by reducing their tendency to overfit. Through systematic benchmarking across various question answering datasets and LLMs, we show how our techniques can yield confidence scores that provide substantial improvements in fine-grained measures of both calibration and accuracy compared to existing methods.


Fox News AI Newsletter: Tech's 'craziest talent war'

FOX News

Elon Musk says Tesla is raising compensation for its AI engineers, saying OpenAI is "aggressively recruiting" them. 'CRAZIEST TALENT WAR': Tesla CEO Elon Musk said the electric vehicle giant is giving its artificial intelligence engineers a raise as the automaker tries to fend off poaching efforts by ChatGPT creator OpenAI. COSTLY GAME: More and more sports bettors appear to be turning to artificial intelligence to help counter the notoriously unpredictable tournament, which is often referred to as March Madness. LEISURE TIME: Billionaire investor and New York Mets owner Steve Cohen said in a Wednesday appearance on CNBC's "Squawk Box," that he believes that the majority of workers will eventually have a four-day work week and three-day weekend, which will expand opportunities for individuals to engage in leisurely pursuits. FIGHT AGAINST AI: Comedian George Carlin's estate has agreed to a settlement with the media company it sued earlier this year over the use of artificial intelligence.


YouTube CEO warns OpenAI that training models on its videos is against the rules

Engadget

AI models using individual's work without permission (or compensation) is nothing new, with entities like The New York Times and Getty Images initiating lawsuits against AI creators alongside artists and writers. In March, OpenAI CTO Mira Murati contributed to the ongoing uncertainty, telling The Wall Street Journal she wasn't sure if Sora, the company's new text-to-video AI tool, takes data from YouTube, Instagram or Facebook posts. Now, YouTube's CEO Neal Mohan has responded with a clear warning to OpenAI that using its videos to teach Sora would be a "clear violation" of the platform's terms of use. In an interview with Bloomberg Originals host Emily Chang, Mohan stated, "From a creator's perspective, when a creator uploads their hard work to our platform, they have certain expectations. One of those expectations is that the terms of service is going to be abided by. It does not allow for things like transcripts or video bits to be downloaded, and that is a clear violation of our terms of service. Those are the rules of the road in terms of content on our platform."


The AI deepfake apocalypse is here. These are the ideas for fighting it.

Washington Post - Technology News

Even before OpenAI released ChatGPT in late 2022 and kicked off the AI boom, camera makers Nikon and Leica began developing ways to imprint special "metadata" that lists when and by whom a photo was taken directly when the image is made by the camera. Canon and Sony have begun similar programs, and Qualcomm, which makes computer chips for smartphones, says it has a similar project to add metadata to images taken on phone cameras.


This AI Startup Wants You to Talk to Houses, Cars, and Factories

WIRED

We've all been astonished at how chatbots seem to understand the world. But what if they were truly connect to the real world? What if the dataset behind the chat interface was physical reality itself, captured in real time by interpreting the input of billions of sensors sprinkled around the globe? As cofounder and CEO Ivan Poupyrev puts it, "Think of ChatGPT, but for physical reality." Archetype's foundational model is called Newton.


Towards Realistic Few-Shot Relation Extraction: A New Meta Dataset and Evaluation

arXiv.org Artificial Intelligence

We introduce a meta dataset for few-shot relation extraction, which includes two datasets derived from existing supervised relation extraction datasets - NYT29 (Takanobu et al., 2019; Nayak and Ng, 2020) and WIKI-DATA (Sorokin and Gurevych, 2017) - as well as a few-shot form of the TACRED dataset (Sabo et al., 2021). Importantly, all these few-shot datasets were generated under realistic assumptions such as: the test relations are different from any relations a model might have seen before, limited training data, and a preponderance of candidate relation mentions that do not correspond to any of the relations of interest. Using this large resource, we conduct a comprehensive evaluation of six recent few-shot relation extraction methods, and observe that no method comes out as a clear winner. Further, the overall performance on this task is low, indicating substantial need for future research. We release all versions of the data, i.e., both supervised and few-shot, for future research.