Law
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
Contestabile, Matilde, Ferrara, Chiara, Giovannetti, Alberto, Parrillo, Giovanni, Vandin, Andrea
Process Mining (PM), initially developed for industrial and business contexts, has recently been applied to social systems, including legal ones. However, PM's efficacy in the legal domain is limited by the accessibility and quality of datasets. We introduce ProLiFIC (Procedural Lawmaking Flow in Italian Chambers), a comprehensive event log of the Italian lawmaking process from 1987 to 2022. Created from unstructured data from the Normattiva portal and structured using large language models (LLMs), ProLiFIC aligns with recent efforts in integrating PM with LLMs. We exemplify preliminary analyses and propose ProLiFIC as a benchmark for legal PM, fostering new developments.
Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
Subramanian, Seganrasan, Verma, Abhigya
The ability of large language models (LLMs) to process and reason over long textual inputs is critical for a wide range of real-world applications. However, progress in this area is significantly constrained by the absence of high-quality, diverse, and verifiable long-context datasets suitable for both training and evaluation. This work introduces a modular, extensible framework for synthetic long-context data generation via prompt-based interaction with LLMs. The framework supports multiple training and alignment objectives, including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO). It encompasses four core generation paradigms: multi-turn conversational dialogues, document-grounded input-output pairs, verifiable instruction-response tasks, and long-context reasoning examples. Through templated prompting, a model-agnostic architecture, and metadata-enriched outputs, the proposed approach facilitates scalable, controllable, and purpose-aligned dataset creation for advancing long-context capabilities in LLMs.
Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
Pokharel, Gaurab, Farabi, Shafkat, Fowler, Patrick J., Das, Sanmay
A surge of recent work explores the ethical and societal implications of large-scale AI models that make "moral" judgments. Much of this literature focuses either on alignment with human judgments through various thought experiments or on the group fairness implications of AI judgments. However, the most immediate and likely use of AI is to help or fully replace the so-called street-level bureaucrats, the individuals deciding to allocate scarce social resources or approve benefits. There is a rich history underlying how principles of local justice determine how society decides on prioritization mechanisms in such domains. In this paper, we examine how well LLM judgments align with human judgments, as well as with socially and politically determined vulnerability scoring systems currently used in the domain of homelessness resource allocation. Crucially, we use real data on those needing services (maintaining strict confidentiality by only using local large models) to perform our analyses. We find that LLM prioritizations are extremely inconsistent in several ways: internally on different runs, between different LLMs, and between LLMs and the vulnerability scoring systems. At the same time, LLMs demonstrate qualitative consistency with lay human judgments in pairwise testing.
DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross domain patent retrieval
Ayaou, Iliass, Cavallucci, Denis, Chibane, Hicham
Patent prior-art retrieval becomes especially challenging when relevant disclosures cross technological boundaries. Existing benchmarks lack explicit domain partitions, making it difficult to assess how retrieval systems cope with such shifts. We introduce DAPFAM, a family-level benchmark with explicit IN-domain and OUT-domain partitions defined by a new IPC3 overlap scheme. The dataset contains 1,247 query families and 45,336 target families aggregated at the family level to reduce international redundancy, with citation based relevance judgments. We conduct 249 controlled experiments spanning lexical (BM25) and dense (transformer) backends, document and passage level retrieval, multiple query and document representations, aggregation strategies, and hybrid fusion via Reciprocal Rank Fusion (RRF). Results reveal a pronounced domain gap: OUT-domain performance remains roughly five times lower than IN-domain across all configurations. Passage-level retrieval consistently outperforms document-level, and dense methods provide modest gains over BM25, but none close the OUT-domain gap. Document-level RRF yields strong effectiveness efficiency trade-offs with minimal overhead. By exposing the persistent challenge of cross-domain retrieval, DAPFAM provides a reproducible, compute-aware testbed for developing more robust patent IR systems. The dataset is publicly available on huggingface at https://huggingface.co/datasets/datalyes/DAPFAM_patent.
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
Hu, Wenbin, Li, Haoran, Jing, Huihao, Hu, Qi, Zeng, Ziqian, Han, Sirui, Xu, Heli, Chu, Tianshu, Hu, Peizhao, Song, Yangqiu
While Large Language Models (LLMs) exhibit remarkable capabilities, they also introduce significant safety and privacy risks. Current mitigation strategies often fail to preserve contextual reasoning capabilities in risky scenarios. Instead, they rely heavily on sensitive pattern matching to protect LLMs, which limits the scope. Furthermore, they overlook established safety and privacy standards, leading to systemic risks for legal compliance. To address these gaps, we formulate safety and privacy issues into contextualized compliance problems following the Contextual Integrity (CI) theory. Under the CI framework, we align our model with three critical regulatory standards: GDPR, EU AI Act, and HIPAA. Specifically, we employ reinforcement learning (RL) with a rule-based reward to incentivize contextual reasoning capabilities while enhancing compliance with safety and privacy norms. Through extensive experiments, we demonstrate that our method not only significantly enhances legal compliance (achieving a +8.58% accuracy improvement in safety/privacy benchmarks) but also further improves general reasoning capability. For OpenThinker-7B, a strong reasoning model that significantly outperforms its base model Qwen2.5-7B-Instruct across diverse subjects, our method enhances its general reasoning capabilities, with +2.05% and +8.98% accuracy improvement on the MMLU and LegalBench benchmark, respectively.
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
Liu, Sheng, Sheng, Qiang, Wang, Danding, Li, Yang, Yang, Guang, Cao, Juan
Despite advances in improving large language model (LLM) to refuse to answer malicious instructions, widely used LLMs remain vulnerable to jailbreak attacks where attackers generate instructions with distributions differing from safety alignment corpora. New attacks expose LLMs' inability to recognize unseen malicious instructions, highlighting a critical distributional mismatch between training data and real-world attacks that forces developers into reactive patching cycles. To tackle this challenge, we propose IMAGINE, a synthesis framework that leverages embedding space distribution analysis to generate jailbreak-like instructions. This approach effectively fills the distributional gap between authentic jailbreak patterns and safety alignment corpora. IMAGINE follows an iterative optimization process that dynamically evolves text generation distributions across iterations, thereby augmenting the coverage of safety alignment data distributions through synthesized data examples. Based on the safety-aligned corpus enhanced through IMAGINE, our framework demonstrates significant decreases in attack success rate on Qwen2.5, Llama3.1, and Llama3.2 without compromising their utility.
Should AI Get Legal Rights?
In the often strange world of AI research, some people are exploring whether the machines should be able to unionize. In Silicon Valley, there's a small but growing field called model welfare, which is working to figure out whether AI models are conscious and deserving of moral considerations, such as legal rights. Within the past year, two research organizations studying model welfare have popped up: Conscium and Eleos AI Research. Anthropic also hired its first AI welfare researcher last year. Earlier this month, Anthropic said it gave its Claude chatbot the ability to terminate "persistently harmful or abusive user interactions" that could be "potentially distressing."
Neuralink's Bid to Trademark 'Telepathy' and 'Telekinesis' Faces Legal Issues
The United States Patent and Trademark Office has rejected Neuralink's attempt to trademark the product names Telepathy and Telekinesis, citing pending applications by another person for the same trademarks. Neuralink, the brain implant company co-founded by Elon Musk, filed to trademark the names in March. But in letters sent to Neuralink in August, the trademark office is refusing to allow the applications to move forward. It says Wesley Berry, a computer scientist and co-founder of tech startup Prophetic, previously filed trademark applications for Telepathy in May 2023 and Telekinesis in August 2024. Prophetic is building a wearable headset to induce lucid dreaming, but only Berry is the author of the trademark applications, not Prophetic.
DAVID MARCUS: Forgive me, but I was wrong about school prayer
Fox News contributor Jonathan Morris and Pastor Robert Jeffress react to the president unveiling new guidance on public school prayer. The battle over prayer in school is raging in Texas right now, with Attorney General Ken Paxton vowing to defend any school district that introduces the controversial practice under a recent state law expanding religious expression in education. For the entirety of my life, and I'm old, the prohibition on public school-sponsored prayer seemed like settled Constitutional science, owing to a 1962 Supreme Court decision barring what had previously been a widespread and normal practice. In the past, I agreed with this form of separation of church and state. For me it was almost a question of better safe than sorry regarding the rights of minority religions, and importantly, I believed that Christian moral values were so ingrained in our culture that 30 seconds a day of praying could be forsaken.
Valve trademarks the 'Steam Frame,' but what the heck is it?
After the smash hit that is the Steam Deck, all eyes are on Valve for its next hardware move. A console to take on Sony and Nintendo? A new trademark filing for the "Steam Frame" has gamers and press alike turning the speculation up to 11. And yeah, I couldn't resist doing some of my own. The United States Patent and Trademark Office has a public filing for the Steam Frame name, assigned to Valve Corporation and its corporate office in Bellevue, Washington, and began on September 2nd.