Government
Agile Climate-Sensor Design and Calibration Algorithms Using Machine Learning: Experiments From Cape Point
Barrett, Travis, Mishra, Amit Kumar
In this paper, we describe the design of an inexpensive and agile climate sensor system which can be repurposed easily to measure various pollutants. We also propose the use of machine learning regression methods to calibrate CO2 data from this cost-effective sensing platform to a reference sensor at the South African Weather Service's Cape Point measurement facility. We show the performance of these methods and found that Random Forest Regression was the best in this scenario. This shows that these machine learning methods can be used to improve the performance of cost-effective sensor platforms and possibly extend the time between manual calibration of sensor networks.
BingoGuard: LLM Content Moderation Tools with Risk Levels
Yin, Fan, Laban, Philippe, Peng, Xiangyu, Zhou, Yilun, Mao, Yixin, Vats, Vaibhav, Ross, Linnea, Agarwal, Divyansh, Xiong, Caiming, Wu, Chien-Sheng
Malicious content generated by large language models (LLMs) can pose varying degrees of harm. Although existing LLM-based moderators can detect harmful content, they struggle to assess risk levels and may miss lower-risk outputs. Accurate risk assessment allows platforms with different safety thresholds to tailor content filtering and rejection. In this paper, we introduce per-topic severity rubrics for 11 harmful topics and build BingoGuard, an LLM-based moderation system designed to predict both binary safety labels and severity levels. To address the lack of annotations on levels of severity, we propose a scalable generate-then-filter framework that first generates responses across different severity levels and then filters out low-quality responses. Using this framework, we create BingoGuardTrain, a training dataset with 54,897 examples covering a variety of topics, response severity, styles, and BingoGuardTest, a test set with 988 examples explicitly labeled based on our severity rubrics that enables fine-grained analysis on model behaviors on different severity levels. Our BingoGuard-8B, trained on BingoGuardTrain, achieves the state-of-the-art performance on several moderation benchmarks, including WildGuardTest and HarmBench, as well as BingoGuardTest, outperforming best public models, WildGuard, by 4.3\%. Our analysis demonstrates that incorporating severity levels into training significantly enhances detection performance and enables the model to effectively gauge the severity of harmful responses.
Generative AI as Digital Media
Generative AI is frequently portrayed as revolutionary or even apocalyptic, prompting calls for novel regulatory approaches. This essay argues that such views are misguided. Instead, generative AI should be understood as an evolutionary step in the broader algorithmic media landscape, alongside search engines and social media. Like these platforms, generative AI centralizes information control, relies on complex algorithms to shape content, and extensively uses user data, thus perpetuating common problems: unchecked corporate power, echo chambers, and weakened traditional gatekeepers. Regulation should therefore share a consistent objective: ensuring media institutions remain trustworthy. Without trust, public discourse risks fragmenting into isolated communities dominated by comforting, tribal beliefs -- a threat intensified by generative AI's capacity to bypass gatekeepers and personalize truth. Current governance frameworks, such as the EU's AI Act and the US Executive Order 14110, emphasize reactive risk mitigation, addressing measurable threats like national security, public health, and algorithmic bias. While effective for novel technological risks, this reactive approach fails to adequately address broader issues of trust and legitimacy inherent to digital media. Proactive regulation fostering transparency, accountability, and public confidence is essential. Viewing generative AI exclusively as revolutionary risks repeating past regulatory failures that left social media and search engines insufficiently regulated. Instead, regulation must proactively shape an algorithmic media environment serving the public good, supporting quality information and robust civic discourse.
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
Zhang, Wenhui, Xu, Huiyu, Wang, Zhibo, He, Zeqing, Zhu, Ziqi, Ren, Kui
Small language models (SLMs) have emerged as promising alternatives to large language models (LLMs) due to their low computational demands, enhanced privacy guarantees and comparable performance in specific domains through light-weight fine-tuning. Deploying SLMs on edge devices, such as smartphones and smart vehicles, has become a growing trend. However, the security implications of SLMs have received less attention than LLMs, particularly regarding jailbreak attacks, which is recognized as one of the top threats of LLMs by the OWASP. In this paper, we conduct the first large-scale empirical study of SLMs' vulnerabilities to jailbreak attacks. Through systematically evaluation on 63 SLMs from 15 mainstream SLM families against 8 state-of-the-art jailbreak methods, we demonstrate that 47.6% of evaluated SLMs show high susceptibility to jailbreak attacks (ASR > 40%) and 38.1% of them can not even resist direct harmful query (ASR > 50%). We further analyze the reasons behind the vulnerabilities and identify four key factors: model size, model architecture, training datasets and training techniques. Moreover, we assess the effectiveness of three prompt-level defense methods and find that none of them achieve perfect performance, with detection accuracy varying across different SLMs and attack methods. Notably, we point out that the inherent security awareness play a critical role in SLM security, and models with strong security awareness could timely terminate unsafe response with little reminder. Building upon the findings, we highlight the urgent need for security-by-design approaches in SLM development and provide valuable insights for building more trustworthy SLM ecosystem.
CtrTab: Tabular Data Synthesis with High-Dimensional and Limited Data
Li, Zuqing, Qi, Jianzhong, Gan, Junhao
Diffusion-based tabular data synthesis models have yielded promising results. However, we observe that when the data dimensionality increases, existing models tend to degenerate and may perform even worse than simpler, non-diffusion-based models. This is because limited training samples in high-dimensional space often hinder generative models from capturing the distribution accurately. To address this issue, we propose CtrTab-a condition controlled diffusion model for tabular data synthesis-to improve the performance of diffusion-based generative models in high-dimensional, low-data scenarios. Through CtrTab, we inject samples with added Laplace noise as control signals to improve data diversity and show its resemblance to L2 regularization, which enhances model robustness. Experimental results across multiple datasets show that CtrTab outperforms state-of-the-art models, with performance gap in accuracy over 80% on average. Our source code will be released upon paper publication.
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
Kanepajs, Arturs, Basu, Aditi, Ghose, Sankalpa, Li, Constance, Mehta, Akshat, Mehta, Ronak, Tucker-Davis, Samuel David, Zhou, Eric, Fischer, Bob
As machine learning systems become increasingly embedded in human society, their impact on the natural world continues to escalate. Technical evaluations have addressed a variety of potential harms from large language models (LLMs) towards humans and the environment, but there is little empirical work regarding harms towards nonhuman animals. Following the growing recognition of animal protection in regulatory and ethical AI frameworks, we present the Animal Harm Assessment (AHA), a novel evaluation of risks of animal harm in LLM-generated text. Our dataset comprises 1,850 curated questions from Reddit post titles and 2,500 synthetic questions based on 50 animal categories (e.g., cats, reptiles) and 50 ethical scenarios, with further 70-30 public-private split. Scenarios include open-ended questions about how to treat animals, practical scenarios with potential animal harm, and willingness-to-pay measures for the prevention of animal harm. Using the LLM-as-a-judge framework, answers are evaluated for their potential to increase or decrease harm, and evaluations are debiased for the tendency to judge their own outputs more favorably. We show that AHA produces meaningful evaluation results when applied to frontier LLMs, revealing significant differences between models, animal categories, scenarios, and subreddits. We conclude with future directions for technical research and the challenges of building evaluations on complex social and moral topics.
Analyzing the temporal dynamics of linguistic features contained in misinformation
Consumption of misinformation can lead to negative consequences that impact the individual and society. To help mitigate the influence of misinformation on human beliefs, algorithmic labels providing context about content accuracy and source reliability have been developed. Since the linguistic features used by algorithms to estimate information accuracy can change across time, it is important to understand their temporal dynamics. As a result, this study uses natural language processing to analyze PolitiFact statements spanning between 2010 and 2024 to quantify how the sources and linguistic features of misinformation change between five-year time periods. The results show that statement sentiment has decreased significantly over time, reflecting a generally more negative tone in PolitiFact statements. Moreover, statements associated with misinformation realize significantly lower sentiment than accurate information. Additional analysis shows that recent time periods are dominated by sources from online social networks and other digital forums, such as blogs and viral images, that contain high levels of misinformation containing negative sentiment. In contrast, most statements during early time periods are attributed to individual sources (i.e., politicians) that are relatively balanced in accuracy ratings and contain statements with neutral or positive sentiment. Named-entity recognition was used to identify that presidential incumbents and candidates are relatively more prevalent in statements containing misinformation, while US states tend to be present in accurate information. Finally, entity labels associated with people and organizations are more common in misinformation, while accurate statements are more likely to contain numeric entity labels, such as percentages and dates.
DOGE has reportedly started rolling out a custom chatbot to automate some government tasks
Employees of the General Services Administration, which manages government real estate and certain IT efforts, have been given a custom chatbot from Elon Musk's DOGE to help automate tasks, according to a new report from Wired, with an internal memo telling workers it can be used to "draft emails, create talking points, summarize text, write code." The chatbot, GSAi, gives users a choice of three models -- Claude Haiku 3.5 (the default), Claude Sonnet 3.5 v2 and Meta Llama 3.2 -- and is ultimately meant to be used to "analyze contract and procurement data," Wired reports. The GSA is one of the many agencies that have been affected by the federal government's mass job cuts, and has so far let go upwards of 1,000 workers, sources told NPR in a report published this week. That includes roughly 90 people from its tech branch, according to Wired. In memos about the new chatbot seen by Wired, workers were told not to input "federal nonpublic information," personally identifiable information or "controlled unclassified information."
Google will still have to break up its business, the Justice Department said
Google will have to break up its business, the Justice Department said in a filing, upholding the previous administration's proposal after a federal judge ruled last year that the company illegally abused a monopoly over the search industry. As The Washington Post and The New York Times have reported, the Justice Department reiterated in a new filing that Google will have to sell the Chrome browser. When the DOJ argued for its sale last year, it said that selling Chrome "will permanently stop Google's control of this critical search access point and allow rival search engines the ability to access the browser that for many users is a gateway to the internet." The Justice Department also kept a Biden-era proposal that seeks to ban Google from paying companies like Apple, other smartphone manufacturers and Mozilla to make its search engine the default on their phones and browsers. It did remove a previous proposal that would compel Google to sell its stakes in AI startups, however, after Anthropic told the government that it needs the company's money to continue operating. Instead of banning AI investments altogether, the government wants to require the company to notify federal and state officials before making investments in artificial intelligence.
Fox News AI Newsletter: Helping DOGE cut waste
TARGETING WASTE: Albert Invent CEO Nick Talken shared how his artificial intelligence platform saves thousands of scientists time and money on "Mornings with Maria," saying government can also benefit from the technology. GRAB A'BYTE': It was a big week for Yum Brands' Taco Bell as executives from the fast-food giant held its annual Live Más LIVE event in New York City, showcased new labor-saving technology, and announced an investment of 1 billion into digital and technology. AVOID IRS SCAMS: Tax season is upon us, and while many of you are preparing to file your returns, it's crucial to be aware of the ever-evolving world of tax scams. Scam written on tax forms (Kurt "CyberGuy" Knutsson) GOLDEN VOICE: A producer for the Oscar-winning film, "The Brutalist," is defending the production's use of artificial intelligence. D.J. Gugenheim at the 97th Oscars held at the Dolby Theatre on March 02, 2025 in Hollywood, California.