Government
Towards Sustainable Web Agents: A Plea for Transparency and Dedicated Metrics for Energy Consumption
Krupp, Lars, Geißler, Daniel, Lukowicz, Paul, Karolus, Jakob
Improvements in the area of large language models have shifted towards the construction of models capable of using external tools and interpreting their outputs. These so-called web agents have the ability to interact autonomously with the internet. This allows them to become powerful daily assistants handling time-consuming, repetitive tasks while supporting users in their daily activities. While web agent research is thriving, the sustainability aspect of this research direction remains largely unexplored. We provide an initial exploration of the energy and CO2 cost associated with web agents. Our results show how different philosophies in web agent creation can severely impact the associated expended energy. We highlight lacking transparency regarding the disclosure of model parameters and processes used for some web agents as a limiting factor when estimating energy consumption. As such, our work advocates a change in thinking when evaluating web agents, warranting dedicated metrics for energy consumption and sustainability.
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
Lee, Christine, Porfirio, David, Wang, Xinyu Jessica, Zhao, Kevin, Mutlu, Bilge
Automated planning is traditionally the domain of experts, utilized in fields like manufacturing and healthcare with the aid of expert planning tools. Recent advancements in LLMs have made planning more accessible to everyday users due to their potential to assist users with complex planning tasks. However, LLMs face several application challenges within end-user planning, including consistency, accuracy, and user trust issues. This paper introduces VeriPlan, a system that applies formal verification techniques, specifically model checking, to enhance the reliability and flexibility of LLMs for end-user planning. In addition to the LLM planner, VeriPlan includes three additional core features -- a rule translator, flexibility sliders, and a model checker -- that engage users in the verification process. Through a user study (n=12), we evaluate VeriPlan, demonstrating improvements in the perceived quality, usability, and user satisfaction of LLMs. Our work shows the effective integration of formal verification and user-control features with LLMs for end-user planning tasks.
Generalization is not a universal guarantee: Estimating similarity to training data with an ensemble out-of-distribution metric
Schreyer, W. Max, Anderson, Christopher, Thompson, Reid F.
Failure of machine learning models to generalize to new data is a core problem limiting the reliability of AI systems, partly due to the lack of simple and robust methods for comparing new data to the original training dataset. We propose a standardized approach for assessing data similarity in a model-agnostic manner by constructing a supervised autoencoder for generalizability estimation (SAGE). We compare points in a low-dimensional embedded latent space, defining empirical probability measures for k -Nearest Neighbors (kNN) distance, reconstruction of inputs and task-based performance. As proof of concept for classification tasks, we use MNIST and CIFAR-10 to demonstrate how an ensemble output probability score can separate deformed images from a mixture of typical test examples, and how this SAGE score is robust to transformations of increasing severity. As further proof of concept, we extend this approach to a regression task using non-imaging data (UCI Abalone). In all cases, we show that out-of-the-box model performance increases after SAGE score filtering, even when applied to data from the model's own training and test datasets. Our out-of-distribution scoring method can be introduced during several steps of model construction and assessment, leading to future improvements in responsible deep learning implementation. 1 Background The presence of generalization gaps, where machine learning performance degrades when a trained model encounters previously-unseen data, represents a critical ongoing challenge in the implementation of AI systems.
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
Xing, Wenpeng, Li, Minghao, Li, Mohan, Han, Meng
Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming from both environmental and system-level factors. These vulnerabilities manifest through sensor spoofing, adversarial attacks, and failures in task and motion planning, posing significant challenges to robustness and safety. Despite the growing body of research, existing reviews rarely focus specifically on the unique safety and security challenges of embodied AI systems. Most prior work either addresses general AI vulnerabilities or focuses on isolated aspects, lacking a dedicated and unified framework tailored to embodied AI. This survey fills this critical gap by: (1) categorizing vulnerabilities specific to embodied AI into exogenous (e.g., physical attacks, cybersecurity threats) and endogenous (e.g., sensor failures, software flaws) origins; (2) systematically analyzing adversarial attack paradigms unique to embodied AI, with a focus on their impact on perception, decision-making, and embodied interaction; (3) investigating attack vectors targeting large vision-language models (LVLMs) and large language models (LLMs) within embodied systems, such as jailbreak attacks and instruction misinterpretation; (4) evaluating robustness challenges in algorithms for embodied perception, decision-making, and task planning; and (5) proposing targeted strategies to enhance the safety and reliability of embodied AI systems. By integrating these dimensions, we provide a comprehensive framework for understanding the interplay between vulnerabilities and safety in embodied AI.
Why Grimes No Longer Believes That Art Is Dead
A couple of years ago, Grimes thought art might be dying. She worried that TikTok was overwhelming attention spans; that transgressive artists were becoming more sanitized; that gimmicky NFTs like the Bored Ape Yacht Club--digital cartoon monkeys which were selling for millions of dollars--were warping value systems. "I just went through this whole big'art isn't worth anything' internal existential crisis," the Canadian singer-songwriter says. "But I've come out the other end thinking, actually, maybe it's the main thing that matters. In the last year, I feel like things became way more about artists again." The rise of AI, Grimes believes, has played a role in that shift, perhaps paradoxically. Earlier this month, Grimes was honored at the TIME100 AI Impact Awards in Dubai for her role in shaping the present and future of the technology. While many other artists are terrified of AI and its potential to replace them, Grimes has embraced the technology, even releasing an AI tool allowing people to sing through her voice. Grimes' penchant for seriously engaging with what others fear or distrust makes her one of pop culture's most singular--and at times divisive--figures. But Grimes wears her contrarianism as a badge of honor, and doesn't hesitate to offer insights and perspectives on a variety of issues. "I'm so canceled that I basically have nothing left to lose," she says. She argues that hyper-partisan hysteria has consumed social media, and wishes people would have more measured, nuanced conversations, even with people that they disagree with. "A lot of people think I'm one way or the other, but my whole vibe is just like, I just want people to think well," she says.
DeSantis announces Florida 'DOGE task force'
Florida is creating a "state DOGE task force" to "eliminate unnecessary bureaucracy," Gov. Ron DeSantis announced Monday. Florida is creating a "DOGE task force" to "eliminate unnecessary bureaucracy and to continue to ensure tax dollars are used in the most efficient way possible," Gov. Ron DeSantis announced Monday. The Republican said the Sunshine State "has never been in better fiscal health," but "we always want to get better, and so we looked to see what [Elon] Musk is doing with the [Department of Government Efficiency] in Washington, D.C." "And the one thing I think that they are doing that we need to incorporate is to utilize and leverage technology like artificial intelligence to be able to police the payments and the operations and the contracts that are done in government," DeSantis continued, speaking behind a lectern with the message "Keeping Florida Efficient." "For example, we have people that review these contracts and if there is DEI, they nix it and things like that. But this is some high-powered stuff and I think would be able to provide us with some good information," he added. "We have already been doing this stuff.
Apple announces 500bn in US investments over next four years
Apple said on Monday it would spend 500bn in US investments in the next four years that will include a giant factory in Texas for artificial intelligence servers and add about 20,000 research and development jobs across the country in that time. That 500bn in expected spending includes everything from purchases from US suppliers to US filming of television shows and movies for its Apple TV service. The company declined to say how much of the figure it was already planning to spend with its US supply base, which includes firms such as Corning that makes glass for iPhones in Kentucky. The move comes after media reports that the Apple CEO, Tim Cook, met President Donald Trump last week. Many of Apple's products that are assembled in China could face 10% tariffs imposed by Trump earlier this month, though the iPhone maker had secured some waivers from China tariffs in the first Trump administration.
Warmongers and authoritarians suffocating global human rights, warns UN
Warmongers and authoritarians are "suffocating" human rights across the world, the chief of the United Nations has warned. Speaking at the UN Human Rights Council in Geneva on Monday, Secretary-General Antonio Guterres depicted a world where human rights were "on the ropes and being pummelled hard". Highlighting the devastating effects of conflicts, including in the Middle East, Ukraine and Congo, Guterres noted abuses linked to economics, technology, climate change, migration, and gender. Guterres called out a "morally bankrupt global financial system" that favours profits over planet protections. He also spoke of those who might exploit artificial intelligence to harm people, and leaders who seek to demonise migrants or restrict women's rights.
UK delays plans to regulate AI as ministers seek to align with Trump administration
Ministers have delayed plans to regulate artificial intelligence as the UK government seeks to align itself with Donald Trump's administration on the technology, the Guardian has learned. A long-awaited AI bill, which ministers had originally intended to publish before Christmas, is not expected to appear in parliament before the summer, according to three Labour sources briefed on the plans. Ministers had intended to publish a short bill within months of entering office that would have required companies to hand over large AI models such as ChatGPT for testing by the UK's AI Security Institute. Trump's election has led to a rethink, however. A senior Labour source said the bill was "properly in the background" and that there were still "no hard proposals in terms of what the legislation looks like".
Future of AI in focus at Web Summit Qatar 2025
The future of artificial intelligence (AI) has been the focus of tech entrepreneurs and financial backers gathered in Doha for the second annual Web Summit hosted by Qatar. The four-day digital technology and emerging innovation summit kicked off its second day on Monday, with attendees eyeing an AI environment being transformed rapidly. Leading entrepreneurs from around the world, including Alexander Wang, founder and CEO of Scale AI, and Alexis Ohanian, co-founder of Reddit and general partner at Seven Seven Six, took centre stage at the event on the opening day. Reporting from Doha, Al Jazeera's Colin Baker said the summit is grappling with questions over the future of AI amid "companies and investors that are changing that landscape more rapidly than we expected". The United States and China are leading in preparedness for AI, said Wang of US company Scale AI.