Government
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
Dugan, Owen, Beneto, Donato Manuel Jimenez, Loh, Charlotte, Chen, Zhuo, Dangovski, Rumen, Soljačić, Marin
To achieve accurate calculations, language model systems often enable LLMs to generate code for arithmetic operations. However, this approach compromises speed and security and, if finetuning is involved, risks the language model losing prior capabilities. We propose a framework that enables exact arithmetic in a single autoregressive step, providing faster, more secure, and more interpretable LLM systems with arithmetic capabilities. We use the hidden states of an LLM to control a symbolic architecture which performs arithmetic. Our implementation using Llama 3 8B Instruct with OccamNet as a symbolic model (OccamLlama) achieves 100% accuracy on single arithmetic operations (+,,,, sin, cos, log, exp,), outperforming GPT 4o and on par with GPT 4o using a code interpreter. OccamLlama also outperforms GPT 4o both with and without a code interpreter on mathematical problem solving benchmarks involving challenging arithmetic, thus enabling small LLMs to match the arithmetic performance of even much larger models. We will make our code public shortly.
A Grassroots Architecture to Supplant Global Digital Platforms by a Global Digital Democracy
We present an architectural alternative to global digital platforms termed grassroots, designed to serve the social, economic, civic, and political needs of local digital communities, as well as their federation. Grassroots platforms may offer local communities an alternative to global digital platforms while operating solely on the smartphones of their members, forsaking any global resources other than the network itself. Such communities may form digital economies without initial capital or external credit, exercise sovereign democratic governance, and federate, ultimately resulting in the grassroots formation of a global digital democracy.
A Holistic Indicator of Polarization to Measure Online Sexism
Ghafouri, Vahid, Such, Jose, Suarez-Tangil, Guillermo
The online trend of the manosphere and feminist discourse on social networks requires a holistic measure of the level of sexism in an online community. This indicator is important for policymakers and moderators of online communities (e.g., subreddits) and computational social scientists, either to revise moderation strategies based on the degree of sexism or to match and compare the temporal sexism across different platforms and communities with real-time events and infer social scientific insights. In this paper, we build a model that can provide a comparable holistic indicator of toxicity targeted toward male and female identity and male and female individuals. Despite previous supervised NLP methods that require annotation of toxic comments at the target level (e.g. annotating comments that are specifically toxic toward women) to detect targeted toxic comments, our indicator uses supervised NLP to detect the presence of toxicity and unsupervised word embedding association test to detect the target automatically. We apply our model to gender discourse communities (e.g., r/TheRedPill, r/MGTOW, r/FemaleDatingStrategy) to detect the level of toxicity toward genders (i.e., sexism). Our results show that our framework accurately and consistently (93% correlation) measures the level of sexism in a community. We finally discuss how our framework can be generalized in the future to measure qualities other than toxicity (e.g. sentiment, humor) toward general-purpose targets and turn into an indicator of different sorts of polarizations.
Continual Learning of Large Language Models: A Comprehensive Survey
Shi, Haizhou, Xu, Zihao, Wang, Hengyi, Qin, Weiyi, Wang, Wenyuan, Wang, Yibin, Wang, Zifeng, Ebrahimi, Sayna, Wang, Hao
The recent success of large language models (LLMs) trained on static, pre-collected, general datasets has sparked numerous research directions and applications. One such direction addresses the non-trivial challenge of integrating pre-trained LLMs into dynamic data distributions, task structures, and user preferences. Pre-trained LLMs, when tailored for specific needs, often experience significant performance degradation in previous knowledge domains -- a phenomenon known as "catastrophic forgetting". While extensively studied in the continual learning (CL) community, it presents new manifestations in the realm of LLMs. In this survey, we provide a comprehensive overview of the current research progress on LLMs within the context of CL. This survey is structured into four main sections: we first describe an overview of continually learning LLMs, consisting of two directions of continuity: vertical continuity (or vertical continual learning), i.e., continual adaptation from general to specific capabilities, and horizontal continuity (or horizontal continual learning), i.e., continual adaptation across time and domains (Section 3). We then summarize three stages of learning LLMs in the context of modern CL: Continual Pre-Training (CPT), Domain-Adaptive Pre-training (DAP), and Continual Fine-Tuning (CFT) (Section 4). Then we provide an overview of evaluation protocols for continual learning with LLMs, along with the current available data sources (Section 5). Finally, we discuss intriguing questions pertaining to continual learning for LLMs (Section 6). The full list of papers examined in this survey is available at https://github.com/Wang-ML-Lab/llm-continual-learning-survey.
Uncertainty estimation in satellite precipitation spatial prediction by combining distributional regression algorithms
Papacharalampous, Georgia, Tyralis, Hristos, Doulamis, Nikolaos, Doulamis, Anastasios
To facilitate effective decision-making, gridded satellite precipitation products should include uncertainty estimates. Machine learning has been proposed for issuing such estimates. However, most existing algorithms for this purpose rely on quantile regression. Distributional regression offers distinct advantages over quantile regression, including the ability to model intermittency as well as a stronger ability to extrapolate beyond the training data, which is critical for predicting extreme precipitation. In this work, we introduce the concept of distributional regression for the engineering task of creating precipitation datasets through data merging. Building upon this concept, we propose new ensemble learning methods that can be valuable not only for spatial prediction but also for prediction problems in general. These methods exploit conditional zero-adjusted probability distributions estimated with generalized additive models for location, scale, and shape (GAMLSS), spline-based GAMLSS and distributional regression forests as well as their ensembles (stacking based on quantile regression, and equal-weight averaging). To identify the most effective methods for our specific problem, we compared them to benchmarks using a large, multi-source precipitation dataset. Stacking emerged as the most successful strategy. Three specific stacking methods achieved the best performance based on the quantile scoring rule, although the ranking of these methods varied across quantile levels. This suggests that a task-specific combination of multiple algorithms could yield significant benefits.
Data Shapley in One Training Run
Wang, Jiachen T., Mittal, Prateek, Song, Dawn, Jia, Ruoxi
Data Shapley provides a principled framework for attributing data's contribution within machine learning contexts. However, existing approaches require re-training models on different data subsets, which is computationally intensive, foreclosing their application to large-scale models. Furthermore, they produce the same attribution score for any models produced by running the learning algorithm, meaning they cannot perform targeted attribution towards a specific model obtained from a single run of the algorithm. This paper introduces In-Run Data Shapley, which addresses these limitations by offering scalable data attribution for a target model of interest. In its most efficient implementation, our technique incurs negligible additional runtime compared to standard model training. This dramatic efficiency improvement makes it possible to perform data attribution for the foundation model pretraining stage for the first time. We present several case studies that offer fresh insights into pretraining data's contribution and discuss their implications for copyright in generative AI and pretraining data curation.
One Big Topic Didn't Come Up at the Debate. Thank God.
The first 2024 presidential debate unfolded Thursday night between President Joe Biden and former President Donald Trump and it was, as fellow Slate writer Jill Filipovic put it, the "most painful two hours of television in living memory." If you missed it, good for you. If you did watch it, you'd be forgiven for not remembering the topics that were actually debated vaguely shouted about. In between Biden's excruciatingly painful and "nightmarishly confused" performance and Trump's boorish firehose of lies and falsehoods, the debate also ostensibly featured a wide range of topics like abortion, the economy, climate change, foreign policy in Ukraine and Israel, election integrity, immigration, veterans, race, crime, health care, and even which of the two candidates has a better golf game. However, there was one major issue that the debate failed to broach--despite the fact that it's one of the most (if not the most) consequential developments since the last election cycle: artificial intelligence.
WATCH: Fox News Digital focus group reacts to Biden, Trump sparring on cognitive ability, golf games
A group of voters polled by Fox News Digital in real-time react to former President Trump's defense of his cognitive abilities. Independent and Republican voters in Fox News Digital's focus group appeared to have mixed reactions to President Biden and former President Trump's sparring over their respective cognitive abilities and golf handicaps, while Democrats generally disapproved. During the CNN Presidential Debate on Thursday night, CNN moderator Dana Bash presented the ages Biden and Trump would be at the end of a potential second four-year term. Biden would be 86, while Trump would be 82. Former President Trump, left, and President Biden squared off in their high-stakes 2024 election debate on Thursday, and the contrast between the pair could not have been starker, a body language expert tells Fox News.
FCC chair asks telecoms companies to prove they're actually trying to stop political AI robocalls
FCC Chairwoman Jessica Rosenworcel has drafted a series of letters to nine major telecom companies, including AT&T and Comcast, to ask if they're actually doing anything about AI political robocalls. AI-generated voices are getting pretty good at mimicking humans and we've already seen this technology in action, when an audio deepfake urged voters to skip the New Hampshire Democratic primary. "We know that AI technologies will make it cheap and easy to flood our networks with deepfakes used to mislead and betray trust. It is especially chilling to see AI voice cloning used to impersonate candidates during elections. As AI tools become more accessible to bad actors and scammers, we need to do everything we can to keep this junk off our networks," wrote Rosenworcel. It's worth noting that all AI robocalls were banned back in February, political or not, but the big telecom companies have yet to announce any enforcement plans.
The Download: AI video games' research potential, and US government website redesigns
Beyond just gaming however, it's a development that raises a tantalizing prospect: might AI video games allow neuroscientists and psychologists to probe more deeply, and unravel enduring mysteries about our brains and behavior? Our senior reporter Jessica Hamzelou decided to find out. This story is from The Checkup, our weekly newsletter all about biotech and health. Sign up to receive it in your inbox every Thursday. Before the internet, Americans may have interacted with the federal government by stepping into grand buildings adorned with impressive stone columns and gleaming marble floors.