Government
Data-adaptive Safety Rules for Training Reward Models
Li, Xiaomin, Gao, Mingye, Zhang, Zhiwei, Fan, Jingxuan, Li, Weiyu
Reinforcement Learning from Human Feedback (RLHF) is commonly employed to tailor models to human preferences, especially to improve the safety of outputs from large language models (LLMs). Traditionally, this method depends on selecting preferred responses from pairs. However, due to the variability in human opinions and the challenges in directly comparing two responses, there is an increasing trend towards fine-grained annotation approaches that evaluate responses using multiple targeted metrics or rules. The challenge lies in efficiently choosing and applying these rules to handle the diverse range of preference data. In this paper, we propose a dynamic method that adaptively selects the most important rules for each response pair. We introduce a mathematical framework that utilizes the maximum discrepancy across paired responses and demonstrate theoretically that this approach maximizes the mutual information between the rule-based annotations and the underlying true preferences. We then train an 8B reward model using this adaptively labeled preference dataset and assess its efficacy using RewardBench. As of January 25, 2025, our model achieved the highest safety performance on the leaderboard, surpassing various larger models.
DebiasPI: Inference-time Debiasing by Prompt Iteration of a Text-to-Image Generative Model
Bonna, Sarah, Huang, Yu-Cheng, Novozhilova, Ekaterina, Paik, Sejin, Shan, Zhengyang, Feng, Michelle Yilin, Gao, Ge, Tayal, Yonish, Kulkarni, Rushil, Yu, Jialin, Divekar, Nupur, Ghadiyaram, Deepti, Wijaya, Derry, Betke, Margrit
Ethical intervention prompting has emerged as a tool to counter demographic biases of text-to-image generative AI models. Existing solutions either require to retrain the model or struggle to generate images that reflect desired distributions on gender and race. We propose an inference-time process called DebiasPI for Debiasing-by-Prompt-Iteration that provides prompt intervention by enabling the user to control the distributions of individuals' demographic attributes in image generation. DebiasPI keeps track of which attributes have been generated either by probing the internal state of the model or by using external attribute classifiers. Its control loop guides the text-to-image model to select not yet sufficiently represented attributes, With DebiasPI, we were able to create images with equal representations of race and gender that visualize challenging concepts of news headlines. We also experimented with the attributes age, body type, profession, and skin tone, and measured how attributes change when our intervention prompt targets the distribution of an unrelated attribute type. We found, for example, if the text-to-image model is asked to balance racial representation, gender representation improves but the skin tone becomes less diverse. Attempts to cover a wide range of skin colors with various intervention prompts showed that the model struggles to generate the palest skin tones. We conducted various ablation studies, in which we removed DebiasPI's attribute control, that reveal the model's propensity to generate young, male characters.
China's DeepSeek Surprise
One week ago, a new and formidable challenger for OpenAI's throne emerged. A Chinese AI start-up, DeepSeek, launched a model that appeared to match the most powerful version of ChatGPT but, at least according to its creator, was a fraction of the cost to build. The program, called DeepSeek-R1, has incited plenty of concern: Ultrapowerful Chinese AI models are exactly what many leaders of American AI companies feared when they, and more recently President Donald Trump, have sounded alarms about a technological race between the United States and the People's Republic of China. This is a "wake up call for America," Alexandr Wang, the CEO of Scale AI, commented on social media. But at the same time, many Americans--including much of the tech industry--appear to be lauding this Chinese AI.
DeepSeek hit with 'large-scale' cyber-attack after AI chatbot tops app stores
DeepSeek said its newly popular app was hit with a cyber-attack on Monday, which forced the Chinese company to temporarily limit registrations. The attack came after the DeepSeek AI assistant app skyrocketed to the top of Apple's App Store, becoming the highest rated free app in the US, and climbed high in Google's Play Store. On its status page, DeepSeek said it started to investigate the issue late Monday night Beijing time. After about two hours of monitoring, the company said it was the victim of a "large-scale malicious attack". While DeekSeek limited registrations, existing users were still able to log on as usual.
What to Know About DeepSeek, the Chinese AI Company Causing Stock Market Chaos
A new Chinese AI model, created by the Hangzhou-based startup DeepSeek, has stunned the American AI industry by outperforming some of OpenAI's leading models, displacing ChatGPT at the top of the iOS app store, and usurping Meta as the leading purveyor of so-called open source AI tools. All of which has raised a critical question: despite American sanctions on Beijing's ability to access advanced semiconductors, is China catching up with the U.S. in the global AI race? At a supposed cost of just 6 million to train, DeepSeek's new R1 model, released last week, was able to match the performance on several math and reasoning metrics by OpenAI's o1 model โ the outcome of tens of billions of dollars in investment by OpenAI and its patron Microsoft. The Chinese model is also cheaper for users. The upshot: the U.S. tech industry is suddenly faced with a potentially cheaper and more powerful challenger, unnerving investors, who sold off American tech stocks on Monday morning.
Coca-Cola announces new Orange Cream flavor: 'Iconic and nostalgic taste'
Coca-Cola's new futuristic flavor was co-created using artificial intelligence. Some Americans said it tasted better than the original recipe, but others couldn't stomach a whole can. Coca-Cola is debuting a new flavor โ and it's got a hint of citrus in it. Coca-Cola Orange Cream will be available nationwide starting Feb. 10, the Atlanta-based soda company announced on Monday morning. Described as "the delicious taste of Coca-Cola infused with refreshing orange and smooth, creamy vanilla flavors," Coca-Cola Orange Cream will also be available in a Zero Sugar version.
Why We're in Love with Apocalypse
It's a mite soon to start grieving, but scientists now project that life on Earth will probably end in about a billion years. A Monday in February, 1,000,002,025, would be my guess. On that inhospitable day, give or take a few million years, the sun will become so hot that the oceans will boil, Earth's oxygen will disappear, and photosynthesis will cease, as will all living things. We should be so lucky. There's a pretty fair chance that life could be wiped out well before then--say, in early June, 2034, or on a cloudy Sunday in November, 3633. Plenty of people do, as it turns out, and, if you want to know who they are, Dorian Lynskey's "Everything Must Go: The Stories We Tell About the End of the World" (Pantheon) is a good place to start. Lynskey, a British journalist and podcaster, has assembled biological, geological, archeological, literary, and cinematic permutations of existential finales, leaving no stone unturned, be it meteor, comet, or asteroid. If a book, a song, a story, a film, a headline, a title, or a study has "world" and "end" in it, Lynskey has unearthed it.
'Serious concerns' about DWP's use of AI to read correspondence from benefit claimants
When your mailbag brims with 25,000 letters and emails every day, deciding which to answer first is daunting. When lurking within are pleas for help from some of the country's most vulnerable people, the stakes only get higher. That is the challenge facing the Department for Work and Pensions (DWP) as correspondence floods in from benefit applicants and claimants โ of which there are more than 20 million, including pensioners, in the UK. The DWP thinks it may have found a solution in using artificial intelligence to read it all first โ including handwritten missives. Human reading used to take weeks and could leave the most vulnerable people waiting for too long for help.
AI prototypes for UK welfare system dropped as officials lament 'false starts'
Ministers have shut down or dropped at least half a dozen artificial intelligence prototypes intended for the welfare system, the Guardian has learned, in a sign of the headwinds facing Keir Starmer's effort to increase government efficiency. Pilots of AI technology to enhance staff training, improve the service in jobcentres, speed up disability benefit payments and modernise communication systems are not being taken forward, freedom of information (FoI) requests reveal. Officials have internally admitted that ensuring AI systems are "scalable, reliable [and] thoroughly tested" are key challenges and say there have been many "frustrations and false starts". Not all trials would be expected to make it into regular use, but two of those now scrapped had been highlighted by the Department for Work and Pensions (DWP) in its latest annual report as examples of how it had "successfully tested multiple generative AI proofs of concept". A-cubed was intended to help staff steer jobseekers into work.
Google pushes global agenda to educate workers and lawmakers about AI
Alphabet's Google, already facing an unprecedented regulatory onslaught, is looking to shape public perception and policies on artificial intelligence ahead of a global wave of AI regulation. A key priority, one executive said, comes in building out educational programs to train the workforce on AI. "Getting more people and organizations, including governments, familiar with AI and using AI tools, makes for better AI policy and opens up new opportunities -- it's a virtuous cycle," said Kent Walker, Alphabet's president of global affairs.