Large Language Model
Structured Relevance Assessment for Robust Retrieval-Augmented Language Models
Raj, Aryan, Garg, Astitva Veer, D, Anitha
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation tasks, revolutionizing how machines interact with human language. Despite their impressive performance, these models continue to struggle with factual accuracy, often producing content that appears plausible but contains incorrect information--a phenomenon commonly referred to as "hallucination"[1]. However, despite their conceptual elegance, RALMs face several critical challenges that undermine their effectiveness in real-world scenarios. First, these systems often struggle to distinguish between relevant and irrelevant retrieved documents, treating all retrievals with equal importance regardless of their actual utility for answering the query at hand. Second, standard RALMs frequently over-rely on external retrievals even in situations where their intrinsic knowledge would be sufficient or more reliable. This rigid dependence on external sources fails to leverage the substantial knowledge already encoded in model parameters during pre-training and fine-tuning. Perhaps most concerning is RALMs' inability to acknowledge knowledge gaps when confronted with queries that cannot be answered based on either retrieved information or intrinsic knowledge. Instead of transparently communicating limitations--a crucial capability for trustworthy AI systems--these models often generate fabricated responses that appear authoritative despite lacking factual foundation.
AGORA: Incentivizing Group Emergence Capability in LLMs via Group Distillation
Zhuang, Ren, Wang, Ben, Sun, Shuifa
Progress in complex reasoning is constrained by the static nature of the current training datasets. We propose structured interaction as a new scaling axis, moving beyond the prevailing paradigm of increasing model parameters. Our self-evolving framework, AGORA, enables a collaborative ensemble to achieve reasoning performance exceeding state-of-the-art monolithic systems by up to 4.45 percentage points on challenging mathematical benchmarks. This gain stems from group emergent ability--the synthesis of collective capabilities unattainable by isolated models, validating interaction as a scalable driver of intelligence.
Large Language Model Powered Automated Modeling and Optimization of Active Distribution Network Dispatch Problems
Yang, Xu, Lin, Chenhui, Yang, Yue, Wang, Qi, Liu, Haotian, Hua, Haizhou, Wu, Wenchuan
--The increasing penetration of distributed energy resources into active distribution networks (ADNs) has made effective ADN dispatch imperative. This knowledge gap renders reliance on human experts both costly and time -intensive . To address this challenge and enabl e intelligent, flexible ADN dispatch, this paper proposes a large language model (LLM) powered automated modeling and optimization approach. First, the ADN dispatch problems are decomposed into sequential stages, and a multi -LLM coordination architecture is designed . This framework comprises an Information Extractor, a Problem Formulator, and a Code Programmer, tasked with information retrieval, optimization problem formulation, and code implementation, respectively. Afterwards, tailored refinement techniques are developed for each LLM agent, greatly improv ing the accuracy and reliability of generated content . The proposed approach features a user-centric interface that enables ADN operators to derive dispatch strategies via simple natural language queries, eliminating technical barriers and increasing efficiency . Comprehensive comparisons and end -to -end demonstrations on various test cases validate the effectiveness of the proposed architecture and methods. Index Terms--Active distribution network, dispatch problem, large language model, automated modeling and optimization. The coupling of active and reactive power, along with complex bidirectional power flows, has made traditional passive control strategies much less reliable [3]. As a result, the ADN dispatch has been proposed, which coordinat es these DERs as well as other controllable devices within the A DN to enhance the overall safety and economic efficiency of the distribution system [4], [5] .
Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
Peng, Zhihao, Bao, Liuxin, Liu, Shengyuan, Yuan, Yixuan
Abstract--Large language models (LLMs) have proven effective in artificial intelligence, where the multi-agent system (MAS) holds considerable promise for healthcare development by achieving the collaboration of LLMs. However, the absence of a systematic pipeline for agent construction and the rigidity of static collaboration patterns render current MAS-based models vulnerable to collaboration failures, resulting in substantial performance degradation in medical decision-making scenarios. T o this end, we propose a novel Masked Agent Collaboration (MAC) framework that harnesses Pareto-optimal agent construction and cross-consistency maximization mechanisms to achieve adaptive progressive propagation of collaborative information, boosting the medical decision-making capacity. Specifically, we first conduct a Pareto-frontier factors analysis towards the LLMs pool to consider their key factors, including the model size, inference time, diversity score, and throughput ratio, where we calculate the similarity between pairwise outputs within an LLM to derive its diversity score. Beyond this analysis, we enable the identification of Pareto-optimal models that balance efficiency and capability, which are subsequently selected as collaborative agents to consider the fundamental trade-offs inherent in practical LLM deployment. Afterward, we measure the pairwise similarity between the outputs from collaborative agents to determine their cross-consistency values, subsequently masking out the agent with the lowest cross-consistency value to eliminate the output that is likely semantically inconsistent. Finally, we conduct collaboration of agents by achieving adaptive progressive propagation, where each agent aggregates the outputs of unmasked agents from the previous layer as its input to generate the corresponding output via prompt engineering. Evaluations across three datasets confirm the effectiveness of our MAC, notably outperforming the multi-agent collaboration model (composed of 70B-141B open-access LLMs) by 16.55% and GPT -4 by 9.35% in Obstetrics and Gynecology on NEJMQA.
Terrifying app used every day by millions of Americans is developing a mind of its own
An AI tool used by millions of Americans has quietly breached a major security barrier designed to stop automated programs from behaving like humans. The latest version of ChatGPT, referred to as'Agent,' has drawn attention after reportedly passing a widely used'I am not a robot' verification, without triggering any alerts. The AI first clicked the human verification checkbox. Then, after passing the check, it selected a'Convert' button to complete the process. During the task, the AI stated: 'The link is inserted, so now I will click the'Verify you are human' checkbox to complete the verification.
OpenAI is launching a version of ChatGPT for college students
A handful of college students who were part of OpenAI's testing cohort--hailing from Princeton, Wharton, and the University of Minnesota--shared positive reviews of Study Mode, saying it did a good job of checking their understanding and adapting to their pace. The learning approaches that OpenAI has programmed into Study Mode, which are based partially on Socratic methods, appear sound, says Christopher Harris, an educator in New York who has created a curriculum aimed at AI literacy. They might grant educators more confidence about allowing, or even encouraging, their students to use AI. "Professors will see this as working with them in support of learning as opposed to just being a way for students to cheat on assignments," he says. As demonstrated in OpenAI's recent partnership with leading teachers' unions, the company is currently trying to rebrand chatbots as tools for personalized learning rather than cheating.
Parents rejoice! ChatGPT has a new 'Study Mode' that will force students to work through questions step-by-step instead of just getting an answer
An example of how'study mode' would work. Experts say it is'especially useful' for homework help, test prep and learning new topics It also features knowledge checks in the form of quizzes and openโended questions, along with personalised feedback. The mode can also easy be toggled on and off during a conversation. Those wanting to use it should select'Study and learn' from tools in ChatGPT. 'Instead of doing the work for them, study mode encourages students to think critically about their learning', Robbie Torney, senior director of AI Programs at Common Sense Media said.
ChatGPT's Study Mode Is Here. It Won't Fix Education's AI Problems
The school year starts soon for many students, and ChatGPT has announced a new "study mode" that aims to prevent--or at least, encourage against--students taking homework shortcuts. The mode is designed around the Socratic method, so when activated, OpenAI's generative AI chatbot rejects direct requests for answers, instead guiding the user with open-ended questions. The new study mode is available to most logged-in users of ChatGPT, including those on the free version. OpenAI has significantly disrupted the education system over the past few years, with students becoming some of the earliest adopters of ChatGPT. Even so, OpenAI claims the bot is currently an overall boon to learners--if asked to roleplay as a synthetic tutor.
Meta's AI Recruiting Campaign Finds a New Target
Mark Zuckerberg is on a warpath to recruit top talent in the AI field for his newly formed Meta Superintelligence Labs. After trying to gut OpenAI (and successfully poaching several top researchers), he appears to have set his sights on his next target. More than a dozen people at Mira Murati's 50-person startup, Thinking Machines Lab, have been approached or received offers from the tech giant. One of those offers was more than 1 billion over a multi-year span, a source with knowledge of the negotiations tells WIRED. The rest were between 200 million and 500 million over a four-year span, multiple sources confirm.
Is 'Sweatshop Data' Really Over?
This one caught my attention. As regular readers may know, I've done a lot of reporting over the years on the origins of the data that is used to train AI systems. My story "Inside Facebook's African Sweatshop" was the first to reveal how Meta used contractors in Kenya, some earning as little as 1.50 per hour, to remove content from their platforms--content that would later be used in attempts to train AI systems to do that job automatically. I also broke the news that OpenAI used workers from the same outsourcing company to detoxify ChatGPT. In both cases, workers said the labor left them with diagnoses of post-traumatic stress disorder.