Goto

Collaborating Authors

 Large Language Model


Prompt Smart, Pay Less: Cost-Aware APO for Real-World Applications

arXiv.org Artificial Intelligence

Prompt design is a critical factor in the effectiveness of Large Language Models (LLMs), yet remains largely heuristic, manual, and difficult to scale. This paper presents the first comprehensive evaluation of Automatic Prompt Optimization (APO) methods for real-world, high-stakes multiclass classification in a commercial setting, addressing a critical gap in the existing literature where most of the APO frameworks have been validated only on benchmark classification tasks of limited complexity. We introduce APE-OPRO, a novel hybrid framework that combines the complementary strengths of APE and OPRO, achieving notably better cost-efficiency, around $18\%$ improvement over OPRO, without sacrificing performance. We benchmark APE-OPRO alongside both gradient-free (APE, OPRO) and gradient-based (ProTeGi) methods on a dataset of ~2,500 labeled products. Our results highlight key trade-offs: ProTeGi offers the strongest absolute performance at lower API cost but higher computational time as noted in~\cite{protegi}, while APE-OPRO strikes a compelling balance between performance, API efficiency, and scalability. We further conduct ablation studies on depth and breadth hyperparameters, and reveal notable sensitivity to label formatting, indicating implicit sensitivity in LLM behavior. These findings provide actionable insights for implementing APO in commercial applications and establish a foundation for future research in multi-label, vision, and multimodal prompt optimization scenarios.


Why Braking? Scenario Extraction and Reasoning Utilizing LLM

arXiv.org Artificial Intelligence

The growing number of ADAS-equipped vehicles has led to a dramatic increase in driving data, yet most of them capture routine driving behavior. Identifying and understanding safety-critical corner cases within this vast dataset remains a significant challenge. Braking events are particularly indicative of potentially hazardous situations, motivating the central question of our research: Why does a vehicle brake? Existing approaches primarily rely on rule-based heuristics to retrieve target scenarios using predefined condition filters. While effective in simple environments such as highways, these methods lack generalization in complex urban settings. In this paper, we propose a novel framework that leverages Large Language Model (LLM) for scenario understanding and reasoning. Our method bridges the gap between low-level numerical signals and natural language descriptions, enabling LLM to interpret and classify driving scenarios. We propose a dual-path scenario retrieval that supports both category-based search for known scenarios and embedding-based retrieval for unknown Out-of-Distribution (OOD) scenarios. To facilitate evaluation, we curate scenario annotations on the Argoverse 2 Sensor Dataset. Experimental results show that our method outperforms rule-based baselines and generalizes well to OOD scenarios.


Continuously Updating Digital Twins using Large Language Models

arXiv.org Artificial Intelligence

Digital twins are models of real-world systems that can simulate their dynamics in response to potential actions. In complex settings, the state and action variables, and available data and knowledge relevant to a system can constantly change, requiring digital twins to continuously update with these changes to remain relevant. Current approaches struggle in this regard, as they require fixed, well-defined modelling environments, and they cannot adapt to novel variables without re-designs, or incorporate new information without re-training. To address this, we frame digital twinning as an in-context learning problem using large language models, enabling seamless updates to the twin at inference time. We develop CALM-DT, a Context-Adaptive Language Model-based Digital Twin that can accurately simulate across diverse state-action spaces using in-context learning alone by utilising fine-tuned encoders for sample retrieval. We empirically demonstrate CALM-DT's competitive performance with existing digital twin approaches, and its unique ability to adapt to changes in its modelling environment without parameter updates.


DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph

arXiv.org Artificial Intelligence

Text-to-SQL, which translates a natural language question into an SQL query, has advanced with in-context learning of Large Language Models (LLMs). However, existing methods show little improvement in performance compared to randomly chosen demonstrations, and significant performance drops when smaller LLMs (e.g., Llama 3.1-8B) are used. This indicates that these methods heavily rely on the intrinsic capabilities of hyper-scaled LLMs, rather than effectively retrieving useful demonstrations. In this paper, we propose a novel approach for effectively retrieving demonstrations and generating SQL queries. We construct a Deep Contextual Schema Link Graph, which contains key information and semantic relationship between a question and its database schema items. This graph-based structure enables effective representation of Text-to-SQL samples and retrieval of useful demonstrations for in-context learning. Experimental results on the Spider benchmark demonstrate the effectiveness of our approach, showing consistent improvements in SQL generation performance and efficiency across both hyper-scaled LLMs and small LLMs. The code is available at https://github.com/jjklle/DCG-SQL}{https://github.com/jjklle/DCG-SQL.


OpenAI CEO tells Federal Reserve confab that entire job categories will disappear due to AI

The Guardian

During his latest trip to Washington, OpenAI's chief executive, Sam Altman, painted a sweeping vision of an AI-dominated future in which entire job categories disappear, presidents follow ChatGPT's recommendations and hostile nations wield artificial intelligence as a weapon of mass destruction, all while positioning his company as the indispensable architect of humanity's technological destiny. Speaking at the Capital Framework for Large Banks conference at the Federal Reserve board of governors, Altman told the crowd that certain job categories would be completely eliminated by AI advancement. "Some areas, again, I think just like totally, totally gone," he said, singling out customer support roles. "That's a category where I just say, you know what, when you call customer support, you're on target and AI, and that's fine." The OpenAI founder described the transformation of customer service as already complete, telling the Federal Reserve vice-chair for supervision, Michelle Bowman: "Now you call one of these things and AI answers. It can do everything that any customer support agent at that company could do. It does not make mistakes. You call once, the thing just happens, it's done."


OpenAI Seeks Additional Capital From Investors as Part of Its 40 Billion Round

WIRED

OpenAI is seeking capital from new and existing investors, two people familiar with the company's plans tell WIRED. The fundraising effort is part of a 40 billion round announced in March. The round will reopen on Monday, July 28, according to one of the sources, who has direct knowledge of the fundraising effort. The 40 billion round announced earlier this year brought OpenAI's valuation up to 300 billion, making it one of the most highly valued private startups in history. The round was led by Japanese investment conglomerate SoftBank, which committed to contributing 75 percent of the total funding.


DeepMind and OpenAI claim gold in International Mathematical Olympiad

New Scientist

Experimental AI models from Google DeepMind and OpenAI have achieved a gold-level performance in the International Mathematical Olympiad (IMO) for the first time. The companies are hailing the moment as an important milestone for AIs that might one day solve hard scientific or mathematical problems, but mathematicians are more cautious because details of the models' results and how they work haven't been made public. The IMO, one of the world's most prestigious competitions for young mathematicians, has long been seen by AI researchers as a litmus test for mathematical reasoning that AI systems tend to struggle with. After last year's competition held in Bath, UK, Google DeepMindannounced that AI systems it had developed, called AlphaProof and AlphaGeometry, had together achieved a silver medal-level performance, but its entries weren't graded by the competition's official markers. Before this year's contest, which was held in Queensland, Australia, companies including Google, Huawei and TikTok-owner ByteDance, as well as academic researchers, approached the organisers to ask whether they could have their AI models' performance officially graded, says Gregor Dolinar, the IMO's president.


UK government urged to offer more transparency over OpenAI deal

The Guardian

Ministers are facing calls for greater transparency about public data that may be shared with the US tech company OpenAI after the government signed a wide-ranging agreement with the 300m ( 222m) company that critics compared to letting a fox into a henhouse. Chi Onwurah, the chair of the House of Commons select committee on science, innovation and technology, warned that Monday's sweeping memorandum of understanding between OpenAI's chief executive, Sam Altman, and the technology secretary, Peter Kyle, was "very thin on detail" and called for guarantees that public data would remain in the UK and clarity about how much of it OpenAI would have access to. The deal paves the way for the Silicon Valley firm behind ChatGPT to explore deploying advanced AI technology in areas including justice, defence and security, and education. It includes OpenAI and the government "partnering to develop safeguards that protect the public and uphold democratic values". Kyle said he wanted Britain to be "front and centre when it comes to developing and deploying AI" and "this can't be achieved without companies like OpenAI".


Human teens beat AI at an international math competition

Popular Science

Breakthroughs, discoveries, and DIY tips sent every weekday. For the first time ever, AI models achieved prestigious gold-level scores at the International Mathematics Olympiad, one of the world's premiere math competitions. Their success is an undeniable bragging right for the technology's biggest supporters. But as it stands, Google and OpenAI's most cutting-edge, experimental AI programs still can't beat an extremely smart teenager. It may seem ironic, but complex mathematics is still one of AI's biggest hurdles.


AI Slop Might Finally Cure Our Internet Addiction

The Atlantic - Technology

For a while, dating apps seemed to make it easier, putting a city's worth of single people in the palm of your hand. But AI has cast a paranoid pall over what can already be a suboptimal experience. If you get a message that feels a little off, it is hard to know whether you are flirting with a bot--or just someone insecure enough to use ChatGPT as their own Cyrano de Bergerac. In frustration, my friend Lonni has started picking up women at the nail salon like it's 1997. Or, in the midst of an emotionally fraught conversation with a friend or family member, a text might read strangely.