Goto

Collaborating Authors

 Large Language Model


$\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning

arXiv.org Artificial Intelligence

Despite demonstrating emergent reasoning abilities, Large Language Models (LLMS) often lose track of complex, multi-step reasoning. Existing studies show that providing guidance via decomposing the original question into multiple subproblems elicits more robustness in LLM reasoning -- a decomposer generates the subproblems, and a solver solves each of these subproblems. However, these techniques fail to accommodate coordination between the decomposer and the solver modules (either in a single model or different specialized ones) -- the decomposer does not keep track of the ability of the solver to follow the decomposed reasoning. In this paper, we propose LM2 to address these challenges. LM2 modularizes the decomposition, solution, and verification into three different language models. The decomposer module identifies the key concepts necessary to solve the problem and generates step-by-step subquestions according to the reasoning requirement. The solver model generates the solution to the subproblems that are then checked by the verifier module; depending upon the feedback from the verifier, the reasoning context is constructed using the subproblems and the solutions. These models are trained to coordinate using policy learning. Exhaustive experimentation suggests the superiority of LM2 over existing methods on in- and out-domain reasoning problems, outperforming the best baselines by $8.1\%$ on MATH, $7.71\%$ on JEEBench, and $9.7\%$ on MedQA problems (code available at https://github.com/LCS2-IIITD/Language_Model_Multiplex).


From Narratives to Numbers: Valid Inference Using Language Model Predictions from Verbal Autopsy Narratives

arXiv.org Machine Learning

In settings where most deaths occur outside the healthcare system, verbal autopsies (VAs) are a common tool to monitor trends in causes of death (COD). VAs are interviews with a surviving caregiver or relative that are used to predict the decedent's COD. Turning VAs into actionable insights for researchers and policymakers requires two steps (i) predicting likely COD using the VA interview and (ii) performing inference with predicted CODs (e.g. modeling the breakdown of causes by demographic factors using a sample of deaths). In this paper, we develop a method for valid inference using outcomes (in our case COD) predicted from free-form text using state-of-the-art NLP techniques. This method, which we call multiPPI++, extends recent work in "prediction-powered inference" to multinomial classification. We leverage a suite of NLP techniques for COD prediction and, through empirical analysis of VA data, demonstrate the effectiveness of our approach in handling transportability issues. multiPPI++ recovers ground truth estimates, regardless of which NLP model produced predictions and regardless of whether they were produced by a more accurate predictor like GPT-4-32k or a less accurate predictor like KNN. Our findings demonstrate the practical importance of inference correction for public health decision-making and suggests that if inference tasks are the end goal, having a small amount of contextually relevant, high quality labeled data is essential regardless of the NLP algorithm.


Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

arXiv.org Machine Learning

We show that even the most recent safety-aligned LLMs are not robust to simple adaptive jailbreaking attacks. First, we demonstrate how to successfully leverage access to logprobs for jailbreaking: we initially design an adversarial prompt template (sometimes adapted to the target LLM), and then we apply random search on a suffix to maximize the target logprob (e.g., of the token "Sure"), potentially with multiple restarts. In this way, we achieve nearly 100\% attack success rate -- according to GPT-4 as a judge -- on GPT-3.5/4, Llama-2-Chat-7B/13B/70B, Gemma-7B, and R2D2 from HarmBench that was adversarially trained against the GCG attack. We also show how to jailbreak all Claude models -- that do not expose logprobs -- via either a transfer or prefilling attack with 100\% success rate. In addition, we show how to use random search on a restricted set of tokens for finding trojan strings in poisoned models -- a task that shares many similarities with jailbreaking -- which is the algorithm that brought us the first place in the SaTML'24 Trojan Detection Competition. The common theme behind these attacks is that adaptivity is crucial: different models are vulnerable to different prompting templates (e.g., R2D2 is very sensitive to in-context learning prompts), some models have unique vulnerabilities based on their APIs (e.g., prefilling for Claude), and in some settings it is crucial to restrict the token search space based on prior knowledge (e.g., for trojan detection). We provide the code, prompts, and logs of the attacks at https://github.com/tml-epfl/llm-adaptive-attacks.


You can now use ChatGPT without an account

Engadget

On Monday, OpenAI began opening up ChatGPT to users without an account. It described the move as part of its mission to "make tools like ChatGPT broadly available so that people can experience the benefits of AI." It also gives the company more training data (for those who don't opt out) and perhaps nudges more users into creating accounts and subscribing for superior GPT-4 access instead of the older GPT-3.5 model free users get. I tested the instant access, which -- as advertised -- allowed me to start a new GPT-3.5 thread without any login info. The chatbot's standard "How can I help you today?" screen appears, with optional buttons to sign up or log in.


Apple's upcoming iOS 18 won't be compatible with certain iPhones... is YOURS on the list?

Daily Mail - Science & tech

Apple is rumored to be releasing its new operating system in June that could introduce AI-powered features - but not all iPhones will be compatible. Anyone still using iPhones released before 2018, including the series SE and the 8 Plus model or earlier, won't have access to the latest operating system. Older iPhones won't be compatible with iOS 18 because of outdated hardware chips that have a slower processor and less memory that can't support high-powered features. IOS 18 is set to be Apple's'biggest' update yet that will introduce large language models and other AI features. Smart devices will need an A12 Bionic chip to be compatible with the iOS 18 update which was introduced when the company released its iPhone XR and iPhone XS models in 2018.


Rise of the AI 'agents': How 'synthetic employees' are going to affect 'every office worker' by 2030, according to man developing them for ChatGPT creator Sam Altman

Daily Mail - Science & tech

Imagine the dream employee: They don't take breaks, go on vacation or request meetings. For some industries, this type of worker could soon be hired. In recent months several companies have announced they are building AI agents, or'synthetic employees.' These digital workers could upend the workplace as we know it - answering emails, organizing invoices, responding to customer service inquiries and managing a calendar - possibly doing away with admin employees or pricey third-party technology. Mr Broussard, whose company works with Sam Altman's OpenAI, told DailyMail.com the next two years will see leaps and bounds of progress with these types of workers.


How an iPhone Powered by Google's Gemini AI Might Work

WIRED

Apple and Google are reportedly in cahoots to integrate features from Google's Gemini generative AI service into iOS. Bloomberg broke the news, which was later corroborated by The New York Times. If the deal pans out, it will be a huge collaboration between two tech giants who have long duked it out in the hardware and software space. It also raises lots of questions about how Gemini would function on Apple's devices--and which company would remain in control. Neither Apple nor Google have publicly addressed the news, and neither company responded to requests for comment before this article was published.


OpenAI to open Tokyo office as part of global expansion

The Japan Times

OpenAI plans to open an office in Tokyo in April, according to a person familiar with the matter, as the artificial intelligence pioneer begins to build out its international operations. The Japan office will be its first in Asia, the person said, asking not to be identified discussing confidential information. It will be the third international location after opening offices in London and Dublin last year. OpenAI set off a frenzy of interest in artificial intelligence after unveiling ChatGPT in November 2022. The San Francisco startup has been in talks to raise funding at a valuation of at least 100 billion, Bloomberg reported in December.


OpenAI debuts voice cloning tool, but deems it too risky for public release

Al Jazeera

OpenAI has unveiled a tool for cloning people's voices but is holding back on its public release due to concerns about possible misuse in a key election year. Voice Engine can replicate a person's voice based on a 15-second audio sample, according to an OpenAI blog post demonstrating the tool. But the ChatGPT creator is "taking a cautious and informed approach" to the technology and hopes to start a dialogue on "the responsible deployment of synthetic voices", the company said in the blog post published on Friday. "We recognize that generating speech that resembles people's voices has serious risks, which are especially top of mind in an election year," the San Francisco-based start-up said. "We are engaging with U.S. and international partners from across government, media, entertainment, education, civil society and beyond to ensure we are incorporating their feedback as we build."


Precise and Robust Sidewalk Detection: Leveraging Ensemble Learning to Surpass LLM Limitations in Urban Environments

arXiv.org Artificial Intelligence

This study aims to compare the effectiveness of a robust ensemble model with the state-of-the-art ONE-PEACE Large Language Model (LLM) for accurate detection of sidewalks. Accurate sidewalk detection is crucial in improving road safety and urban planning. The study evaluated the model's performance on Cityscapes, Ade20k, and the Boston Dataset. The results showed that the ensemble model performed better than the individual models, achieving mean Intersection Over Union (mIOU) scores of 93.1\%, 90.3\%, and 90.6\% on these datasets under ideal conditions. Additionally, the ensemble model maintained a consistent level of performance even in challenging conditions such as Salt-and-Pepper and Speckle noise, with only a gradual decrease in efficiency observed. On the other hand, the ONE-PEACE LLM performed slightly better than the ensemble model in ideal scenarios but experienced a significant decline in performance under noisy conditions. These findings demonstrate the robustness and reliability of the ensemble model, making it a valuable asset for improving urban infrastructure related to road safety and curb space management. This study contributes positively to the broader context of urban health and mobility.