Government
Can AI expose tax loopholes? Towards a new generation of legal policy assistants
Fratrič, Peter, Holzenberger, Nils, Amariles, David Restrepo
The legislative process is the backbone of a state built on solid institutions. Yet, due to the complexity of laws -- particularly tax law -- policies may lead to inequality and social tensions. In this study, we introduce a novel prototype system designed to address the issues of tax loopholes and tax avoidance. Our hybrid solution integrates a natural language interface with a domain-specific language tailored for planning. We demonstrate on a case study how tax loopholes and avoidance schemes can be exposed. We conclude that our prototype can help enhance social welfare by systematically identifying and addressing tax gaps stemming from loopholes.
A Study into Investigating Temporal Robustness of LLMs
Wallat, Jonas, Abdallah, Abdelrahman, Jatowt, Adam, Anand, Avishek
Large Language Models (LLMs) encapsulate a surprising amount of factual world knowledge. However, their performance on temporal questions and historical knowledge is limited because they often cannot understand temporal scope and orientation or neglect the temporal aspect altogether. In this study, we aim to measure precisely how robust LLMs are for question answering based on their ability to process temporal information and perform tasks requiring temporal reasoning and temporal factual knowledge. Specifically, we design eight time-sensitive robustness tests for factual information to check the sensitivity of six popular LLMs in the zero-shot setting. Overall, we find LLMs lacking temporal robustness, especially to temporal reformulations and the use of different granularities of temporal references. We show how a selection of these eight tests can be used automatically to judge a model's temporal robustness for user questions on the fly. Finally, we apply the findings of this study to improve the temporal QA performance by up to 55 percent.
An Iterative Feedback Mechanism for Improving Natural Language Class Descriptions in Open-Vocabulary Object Detection
Kim, Louis Y., Karker, Michelle, Valledor, Victoria, Lee, Seiyoung C., Brzoska, Karl F., Duff, Margaret, Palladino, Anthony
Recent advances in open-vocabulary object detection models will enable Automatic Target Recognition systems to be sustainable and repurposed by non-technical end-users for a variety of applications or missions. New, and potentially nuanced, classes can be defined with natural language text descriptions in the field, immediately before runtime, without needing to retrain the model. We present an approach for improving non-technical users' natural language text descriptions of their desired targets of interest, using a combination of analysis techniques on the text embeddings, and proper combinations of embeddings for contrastive examples. We quantify the improvement that our feedback mechanism provides by demonstrating performance with multiple publicly-available open-vocabulary object detection models.
KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications
Bommarito, Michael J, Katz, Daniel Martin, Bommarito, Jillian
We present the KL3M tokenizers, a family of specialized tokenizers for legal, financial, and governmental text. Despite established work on tokenization, specialized tokenizers for professional domains remain understudied. Our paper offers two main contributions to this area. First, we introduce domain-specific BPE tokenizers for legal, financial, and governmental text. Our kl3m-004-128k-cased tokenizer uses 9-17% fewer tokens than GPT-4o and Llama3 for domain-specific documents, despite having a smaller vocabulary. For specialized terminology, our cased tokenizer is even more efficient, using up to 83% fewer tokens for legal terms and 39% fewer tokens for financial terms. Second, we develop character-level BPE tokenizers (4K, 8K, and 16K vocabulary sizes) for text correction tasks like OCR post-processing. These tokenizers keep consistent token boundaries between error-containing and correct text, making it easier for models to learn correction patterns. These tokenizers help professional applications by fitting more text in context windows, reducing computational needs, and preserving the meaning of domain-specific terms. Our analysis shows these efficiency gains directly benefit the processing of long legal and financial documents. We release all tokenizers and code through GitHub and Hugging Face to support further research in specialized tokenization.
TAET: Two-Stage Adversarial Equalization Training on Long-Tailed Distributions
YuHang, Wang, Guo, Junkang, Liu, Aolei, Wang, Kaihao, Wu, Zaitong, Liu, Zhenyu, Yin, Wenfei, Liu, Jian
Adversarial robustness is a critical challenge in deploying deep neural networks for real-world applications. While adversarial training is a widely recognized defense strategy, most existing studies focus on balanced datasets, overlooking the prevalence of long-tailed distributions in real-world data, which significantly complicates robustness. This paper provides a comprehensive analysis of adversarial training under long-tailed distributions and identifies limitations in the current state-of-the-art method, AT-BSL, in achieving robust performance under such conditions. To address these challenges, we propose a novel training framework, TAET, which integrates an initial stabilization phase followed by a stratified equalization adversarial training phase. Additionally, prior work on long-tailed robustness has largely ignored the crucial evaluation metric of balanced accuracy. To bridge this gap, we introduce the concept of balanced robustness, a comprehensive metric tailored for assessing robustness under long-tailed distributions. Extensive experiments demonstrate that our method surpasses existing advanced defenses, achieving significant improvements in both memory and computational efficiency. This work represents a substantial advancement in addressing robustness challenges in real-world applications. Our code is available at: https://github.com/BuhuiOK/TAET-Two-Stage-Adversarial-Equalization-Training-on-Long-Tailed-Distributions.
JD Vance takes shot at Harris as he jokes that drinking led to her 'word salads'
Tech expert Kurt'CyberGuy' Knutsson joins'Fox & Friends' to discuss the future of AI development in the United States. Vice President JD Vance took a shot at former Vice President Kamala Harris, suggesting her alcohol habits were responsible for her "word salads." Vance's remarks came as he described the difference between how he and Harris have handled the role as vice president, and he speculated about the relationship dynamic between Harris and former President Joe Biden. "Well, I don't have four shots of vodka before every meeting," Vance said in an interview with radio host and Daily Caller editor Vince Coglianese in an interview that aired Thursday. "That's one way I think that Kamala really tried to bring herself into the role, is these word salads. I think I would need the help of a lot of alcohol to answer a question the way that Kamala Harris answered questions."
The illusory reality of WWI dazzle camouflage, re-examined
During World War I, Allied navies started implementing shocking, cubist-inspired "dazzle" paint jobs on ships. The now-iconic geometric designs were intended to throw off the visual perception of German U-boats crews and prevent them from accurately targeting ships with torpedoes. Conventional wisdom claims the bizarre camouflage pattern worked and helped turn the tide of Great War naval battles. But new research reevaluating one of the only rigorous studies testing that hypothesis suggests those conclusions were probably overblown. Researchers now claim another phenomena known as the "horizon effect" may have actually done more to throw off submarine gunners than the wacky aesthetic.
The Download: the future of energy, and chatting about chatbots
Where can you find lasers, electric guitars, and racks full of novel batteries, all in the same giant room? This week, the answer was the 2025 ARPA-E Energy Innovation Summit just outside Washington, DC. Energy innovation can take many forms, and the variety in energy research was on display at the summit. ARPA-E, part of the US Department of Energy, provides funding for high-risk, high-reward research projects. The summit gathers projects the agency has funded, along with investors, policymakers, and journalists.
Russia, Ukraine ramp up drone attacks overnight despite truce talks
Russian bombardments in eastern Ukraine ramped up overnight, killing two people, as Ukraine hit Russia's Engels military airfield in the country's southwest region of Saratov with drones. Both Russia and Ukraine stepped up aerial attacks in the early hours of Thursday as United States President Donald Trump pushes both sides to agree to a ceasefire after more than three years of fighting. Ukrainian officials in the northeastern Sumy and Kharkiv regions said two people were killed and several others injured after Russia dropped more than three dozen glide bombs on the towns in the border regions. Russian drone attacks on the town of Kropyvnytskyi, hundreds of kilometres from the front line, wounded 14 people and damaged rail infrastructure. "Kropyvnytskyi underwent the most massive enemy attack. Peaceful residential buildings were destroyed," regional governor Andriy Raikovych said.
Meta AI is coming to Europe this week
Meta is rolling out its AI assistant across 41 European countries, including to members of the European Union, starting this week. It will also extend its access to 21 overseas European territories. In its announcement, Meta said that it has taken the company longer to bring its AI technology to European users as it continues to "navigate its complex regulatory system." The company was planning to make its AI technology available in the region last year, but it had to put its plans on pause after the Irish Data Protection Commission asked it to delay training its Large Language Models on content posted by adult European users on Facebook and Instagram. A month after the Irish regulator's request, Meta said that it wasn't going to release its new multimodal Llama models in the region "due to the unpredictable nature of the European regulatory environment."