Government
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
Fernandes, Leandro Carísio, Ribeiro, Leandro dos Santos, de Castro, Marcos Vinícius Borela, Pacheco, Leonardo Augusto da Silva, Sandes, Edans Flávius de Oliveira
This paper introduces JurisTCU, a Brazilian Portuguese dataset for legal information retrieval (LIR). The dataset is freely available and consists of 16,045 jurisprudential documents from the Brazilian Federal Court of Accounts, along with 150 queries annotated with relevance judgments. It addresses the scarcity of Portuguese-language LIR datasets with query relevance annotations. The queries are organized into three groups: real user keyword-based queries, synthetic keyword-based queries, and synthetic question-based queries. Relevance judgments were produced through a hybrid approach combining LLM-based scoring with expert domain validation. We used JurisTCU in 14 experiments using lexical search (document expansion methods) and semantic search (BERT-based and OpenAI embeddings). We show that the document expansion methods significantly improve the performance of standard BM25 search on this dataset, with improvements exceeding 45% in P@10, R@10, and nDCG@10 metrics when evaluating short keyword-based queries. Among the embedding models, the OpenAI models produced the best results, with improvements of approximately 70% in P@10, R@10, and nDCG@10 metrics for short keyword-based queries, suggesting that these dense embeddings capture semantic relationships in this domain, surpassing the reliance on lexical terms. Besides offering a dataset for the Portuguese-language IR research community, suitable for evaluating search systems, the results also contribute to enhancing a search system highly relevant to Brazilian citizens.
Investigating Execution-Aware Language Models for Code Optimization
Di Menna, Federico, Traini, Luca, Bavota, Gabriele, Cortellessa, Vittorio
Code optimization is the process of enhancing code efficiency, while preserving its intended functionality. This process often requires a deep understanding of the code execution behavior at run-time to identify and address inefficiencies effectively. Recent studies have shown that language models can play a significant role in automating code optimization. However, these models may have insufficient knowledge of how code execute at run-time. To address this limitation, researchers have developed strategies that integrate code execution information into language models. These strategies have shown promise, enhancing the effectiveness of language models in various software engineering tasks. However, despite the close relationship between code execution behavior and efficiency, the specific impact of these strategies on code optimization remains largely unexplored. This study investigates how incorporating code execution information into language models affects their ability to optimize code. Specifically, we apply three different training strategies to incorporate four code execution aspects -- line executions, line coverage, branch coverage, and variable states -- into CodeT5+, a well-known language model for code. Our results indicate that execution-aware models provide limited benefits compared to the standard CodeT5+ model in optimizing code.
A Grey-box Text Attack Framework using Explainable AI
Chiramal, Esther, Kai, Kelvin Soh Boon
Explainable AI is a strong strategy implemented to understand complex black-box model predictions in a human interpretable language. It provides the evidence required to execute the use of trustworthy and reliable AI systems. On the other hand, however, it also opens the door to locating possible vulnerabilities in an AI model. Traditional adversarial text attack uses word substitution, data augmentation techniques and gradient-based attacks on powerful pre-trained Bidirectional Encoder Representations from Transformers (BERT) variants to generate adversarial sentences. These attacks are generally whitebox in nature and not practical as they can be easily detected by humans E.g. Changing the word from "Poor" to "Rich". We proposed a simple yet effective Grey-box cum Black-box approach that does not require the knowledge of the model while using a set of surrogate Transformer/BERT models to perform the attack using Explainable AI techniques. As Transformers are the current state-of-the-art models for almost all Natural Language Processing (NLP) tasks, an attack generated from BERT1 is transferable to BERT2. This transferability is made possible due to the attention mechanism in the transformer that allows the model to capture long-range dependencies in a sequence. Using the power of BERT generalisation via attention, we attempt to exploit how transformers learn by attacking a few surrogate transformer variants which are all based on a different architecture. We demonstrate that this approach is highly effective to generate semantically good sentences by changing as little as one word that is not detectable by humans while still fooling other BERT models.
Route Sparse Autoencoder to Interpret Large Language Models
Shi, Wei, Li, Sihang, Liang, Tao, Wan, Mingyang, Ma, Gojun, Wang, Xiang, He, Xiangnan
Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning. Sparse autoencoders (SAEs) have demonstrated promise in this domain by extracting interpretable and monosemantic features. However, prior works primarily focus on feature extraction from a single layer, failing to effectively capture activations that span multiple layers. In this paper, we introduce Route Sparse Autoencoder (RouteSAE), a new framework that integrates a routing mechanism with a shared SAE to efficiently extract features from multiple layers. It dynamically assigns weights to activations from different layers, incurring minimal parameter overhead while achieving high interpretability and flexibility for targeted feature manipulation. We evaluate RouteSAE through extensive experiments on Llama-3.2-1B-Instruct. Specifically, under the same sparsity constraint of 64, RouteSAE extracts 22.5% more features than baseline SAEs while achieving a 22.3% higher interpretability score. These results underscore the potential of RouteSAE as a scalable and effective method for LLM interpretability, with applications in feature discovery and model intervention. Our codes are available at https://github.com/swei2001/RouteSAEs.
Automating Violence Detection and Categorization from Ancient Texts
Abdelhalim, Alhassan, Regneri, Michaela
Violence descriptions in literature offer valuable insights for a wide range of research in the humanities. For historians, depictions of violence are of special interest for analyzing the societal dynamics surrounding large wars and individual conflicts of influential people. Harvesting data for violence research manually is laborious and time-consuming. This study is the first one to evaluate the effectiveness of large language models (LLMs) in identifying violence in ancient texts and categorizing it across multiple dimensions. Our experiments identify LLMs as a valuable tool to scale up the accurate analysis of historical texts and show the effect of fine-tuning and data augmentation, yielding an F1-score of up to 0.93 for violence detection and 0.86 for fine-grained violence categorization.
ToolFuzz -- Automated Agent Tool Testing
Milev, Ivan, Balunović, Mislav, Baader, Maximilian, Vechev, Martin
Large Language Model (LLM) Agents leverage the advanced reasoning capabilities of LLMs in real-world applications. To interface with an environment, these agents often rely on tools, such as web search or database APIs. As the agent provides the LLM with tool documentation along the user query, the completeness and correctness of this documentation is critical. However, tool documentation is often over-, under-, or ill-specified, impeding the agent's accuracy. Standard software testing approaches struggle to identify these errors as they are expressed in natural language. Thus, despite its importance, there currently exists no automated method to test the tool documentation for agents. To address this issue, we present ToolFuzz, the first method for automated testing of tool documentations. ToolFuzz is designed to discover two types of errors: (1) user queries leading to tool runtime errors and (2) user queries that lead to incorrect agent responses. ToolFuzz can generate a large and diverse set of natural inputs, effectively finding tool description errors at a low false positive rate. Further, we present two straightforward prompt-engineering approaches. We evaluate all three tool testing approaches on 32 common LangChain tools and 35 newly created custom tools and 2 novel benchmarks to further strengthen the assessment. We find that many publicly available tools suffer from underspecification. Specifically, we show that ToolFuzz identifies 20x more erroneous inputs compared to the prompt-engineering approaches, making it a key component for building reliable AI agents.
Tesla's stock defied gravity for years. Is Elon Musk's EV party over?
Tesla's stock has dropped by nearly half in three months. Even so, investors are still debating whether Elon Musk's electric-vehicle maker remains overpriced. The company's market capitalization has dropped 45% since hitting an all-time high of 1.5 trillion on Dec. 17, erasing most of the gains the stock made after CEO Musk helped finance the election victory of U.S. President Donald Trump. And yet Tesla continues to fetch a valuation far above those of the world's biggest automotive and technology firms, judging by standard financial metrics. That's because most investors and analysts have bought Musk's pitch that the world's most-valuable automaker isn't really a car company at all, but rather an artificial-intelligence pioneer that will soon unleash a revolution in robotaxis and humanoid robots.
DOGE's Plans to Replace Humans With AI Are Already Under Way
If you have tips about the remaking of the federal government, you can contact Matteo Wong on Signal at @matteowong.52. A new phase of the president and the Department of Government Efficiency's attempts to downsize and remake the civil service is under way. The idea is simple: use generative AI to automate work that was previously done by people. The Trump administration is testing a new chatbot with 1,500 federal employees at the General Services Administration and may release it to the entire agency as soon as this Friday--meaning it could be used by more than 10,000 workers who are responsible for more than 100 billion in contracts and services. This article is based in part on conversations with several current and former GSA employees with knowledge of the technology, all of whom requested anonymity to speak about confidential information; it is also based on internal GSA documents that I reviewed, as well as the software's code base, which is visible on GitHub.
The Download: supercharging the power grid, and a new Chinese AI agent
Rob Gramlich is founder and president of Grid Strategies and was economic advisor to the chairman of the Federal Energy Regulatory Commission during the George W. Bush administration. US electricity consumption is rising faster than it has in decades. Accommodating that growth will require building wind turbines, solar farms, and other power plants faster than we ever have before--and expanding the network of wires needed to connect those facilities to the grid. But one major problem is that it's expensive and slow to secure permits for new transmission lines and build them across the country. Fortunately, there are some shortcuts that could expand the capacity of the existing system without requiring completely new infrastructure: a suite of hardware and software tools known as advanced transmission technologies (ATTs), which can increase both the capacity and the efficiency of the power sector.
With drones and North Korean troops, Russia pushes back Ukraine's offensive
Russian and North Korean forces have made significant battlefield advances in recent days in the Kursk region of Russia, threatening Ukraine's supply lines and its hold on a patch of land it hopes to use as a bargaining chip in future negotiations, according to Ukrainian soldiers, Russian military bloggers and military analysts. Working together, a new influx of North Korean soldiers and well-trained Russian drone units, advancing under the cover of ferocious artillery fire and aerial bombardment, have been able to overwhelm important Ukrainian positions, Ukrainian soldiers said. "It's true; we can't stop them," said Oleksii, commander of a Ukrainian communications unit fighting in the area, when reached by phone. "They just sweep us away, advancing in groups of 50 North Koreans while we have only six men on our positions.