Goto

Collaborating Authors

 Government


Compliance of AI Systems

arXiv.org Artificial Intelligence

The increasing integration of artificial intelligence (AI) systems in various fields requires solid concepts to ensure compliance with upcoming legislation. This paper systematically examines the compliance of AI systems with relevant legislation, focusing on the EU's AI Act and the compliance of data sets. The analysis highlighted many challenges associated with edge devices, which are increasingly being used to deploy AI applications closer and closer to the data sources. Such devices often face unique issues due to their decentralized nature and limited computing resources for implementing sophisticated compliance mechanisms. By analyzing AI implementations, the paper identifies challenges and proposes the first best practices for legal compliance when developing, deploying, and running AI. The importance of data set compliance is highlighted as a cornerstone for ensuring the trustworthiness, transparency, and explainability of AI systems, which must be aligned with ethical standards set forth in regulatory frameworks such as the AI Act. The insights gained should contribute to the ongoing discourse on the responsible development and deployment of embedded AI systems.


Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance

arXiv.org Artificial Intelligence

Existing benchmarks are becoming saturated and struggle to separate model performances due to factors like data contamination and advancing LLM capabilities. This paper introduces EMDM (Enhanced Model Differentiation Metric), a novel weighted metric that revitalizes benchmarks by enhancing model separation. EMDM integrates final answer and Chain-of-Thought (CoT) reasoning correctness, assigning weights based on the complexity and reasoning depth required to solve a given sample in the evaluation data. Using a baseline LLM in two setups-Unguided, where the model has no prior exposure to test samples, and Guided, where the model has prior knowledge of the desired answer-EMDM distinguishes instances of varying difficulty. The CoT and answer correctness from these setups inform an optimization objective for weight assignment, resulting in a more nuanced evaluation of model performance. Compared to the exact match (EM) metric, which achieves 17% separation on ARC-Challenge, EMDM achieves 46%, demonstrating its effectiveness in differentiating models based on reasoning and knowledge requirements.


Dynamic Knowledge Integration for Evidence-Driven Counter-Argument Generation with Large Language Models

arXiv.org Artificial Intelligence

This paper investigates the role of dynamic external knowledge integration in improving counter-argument generation using Large Language Models (LLMs). While LLMs have shown promise in argumentative tasks, their tendency to generate lengthy, potentially unfactual responses highlights the need for more controlled and evidence-based approaches. We introduce a new manually curated dataset of argument and counter-argument pairs specifically designed to balance argumentative complexity with evaluative feasibility. We also propose a new LLM-as-a-Judge evaluation methodology that shows a stronger correlation with human judgments compared to traditional reference-based metrics. Our experimental results demonstrate that integrating dynamic external knowledge from the web significantly improves the quality of generated counter-arguments, particularly in terms of relatedness, persuasiveness, and factuality. The findings suggest that combining LLMs with real-time external knowledge retrieval offers a promising direction for developing more effective and reliable counter-argumentation systems.


Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning

arXiv.org Artificial Intelligence

Large language models (LLMs) have achieved remarkable performance on knowledge graph question answering (KGQA) tasks by planning and interacting with knowledge graphs. However, existing methods often confuse tool utilization with knowledge reasoning, harming readability of model outputs and giving rise to hallucinatory tool invocations, which hinder the advancement of KGQA. To address this issue, we propose Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning (MemQ) to decouple LLM from tool invocation tasks using LLM-built query memory. By establishing a memory module with explicit descriptions of query statements, the proposed MemQ facilitates the KGQA process with natural language reasoning and memory-augmented query reconstruction. Meanwhile, we design an effective and readable reasoning to enhance the LLM's reasoning capability in KGQA. Experimental results that MemQ achieves state-of-the-art performance on widely used benchmarks WebQSP and CWQ.


Transformer Meets Twicing: Harnessing Unattended Residual Information

arXiv.org Artificial Intelligence

Transformer-based deep learning models have achieved state-of-the-art performance across numerous language and vision tasks. While the self-attention mechanism, a core component of transformers, has proven capable of handling complex data patterns, it has been observed that the representational capacity of the attention matrix degrades significantly across transformer layers, thereby hurting its overall performance. In this work, we leverage the connection between selfattention computations and low-pass non-local means (NLM) smoothing filters and propose the Twicing Attention, a novel attention mechanism that uses kernel twicing procedure in nonparametric regression to alleviate the low-pass behavior of associated NLM smoothing with compelling theoretical guarantees and enhanced adversarial robustness. This approach enables the extraction and reuse of meaningful information retained in the residuals following the imperfect smoothing operation at each layer. Our proposed method offers two key advantages over standard self-attention: 1) a provably slower decay of representational capacity and 2) improved robustness and accuracy across various data modalities and tasks. We empirically demonstrate the performance gains of our model over baseline transformers on multiple tasks and benchmarks, including image classification and language modeling, on both clean and corrupted data. They have also demonstrated strong performance in knowledge transfer from pretraining tasks to various downstream tasks with weak or no supervision (Radford et al., 2018; 2019; Devlin et al., 2018). At the core of these models is the dot-product self-attention mechanism, which learns self-alignment between tokens in an input sequence by estimating the relative importance of each token with respect to all others. The mechanism then transforms each token into a weighted average of the feature representations of the other tokens with weights proportional to the learned importance scores.


Llamarine: Open-source Maritime Industry-specific Large Language Model

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have demonstrated substantial potential in addressing complex reasoning tasks, yet their general-purpose nature often limits their effectiveness in specialized domains such as maritime navigation. To bridge this gap, we introduce Llamarine, the first open-source LLM designed specifically for maritime navigation. Llamarine 1.0 is developed through continued pretraining and fine-tuning on a high-quality corpus comprising maritime textbooks, research publications, and web text from Wikipedia. This domain-specific training enables the model to acquire expert-level knowledge in navigational principles, collision avoidance, route optimization, and regulatory compliance. Our key contributions include (a) the curation of a comprehensive maritime dataset from authoritative sources, ensuring depth and reliability in the model's knowledge base; (b) the development of a foundational model capable of reasoning about complex navigational challenges with greater accuracy than general-purpose LLMs; and (c) the establishment of a benchmark to evaluate performance in maritime-specific decision-making tasks. Experimental results demonstrate that Llamarine outperforms both general-purpose and commercial LLMs in critical navigation-related tasks, such as trajectory planning, risk assessment, and compliance with maritime regulations. By providing an open-source foundation model trained exclusively on high-quality maritime literature, Llamarine paves the way for AI-driven advancements in maritime safety, efficiency, and operational decision-making.


The arrogant ex-soldier who turned into a triple killer

BBC News

Former soldier Kyle Clifford raped and murdered Louise Hunt, and killed her sister Hannah and mother Carol in attacks described by police as "barbaric". What happened and what has emerged since? Days before the attacks, Louise had ended an 18-month relationship with Clifford. She told Clifford, who she had met through a dating app, it was "sucking the life out of me". They did not like the way Clifford treated Louise, finding him disrespectful, arrogant, rude and "odd". He had hidden relationships with other women from Louise, and went on a dating site moments after receiving the message ending theirs.


Trump blasts Rep. Al Green as 'an embarrassment' to Democrats, says he 'should be forced to take an IQ test'

FOX News

Rep. Al Green, D-Texas, was removed from President Donald Trump's speech to a joint session of Congress after disrupting the event. EXCLUSIVE: President Donald Trump told Fox News Digital Thursday that Rep. Al Green "should be forced to pass an IQ test because he is a low IQ individual and we don't need low IQ individuals in Congress," after the Democrat disrupted his joint session address. The House of Representatives Thursday, in a bipartisan vote, censured Green, D-Texas, for interrupting the president's Tuesday joint session address to Congress. President Donald Trump attends a joint session of Congress at the U.S. Capitol in Washington, D.C., March 4, 2025. In an exclusive interview with Fox News Digital, the president reacted.


It's time to ban Chinese AI app DeepSeek from 'government devices,' state AGs urge Congress

FOX News

Trump counselor Alina Habba responds to concerns of China buying up American real estate on'The Ingraham Angle.' State attorneys general have joined the growing calls from elected officials urging Congress to pass a law banning the Chinese-owned DeepSeek AI app on all government devices, saying "China is a clear and present danger" to the U.S. "DeepSeek appears to be another tool for Chinese spies to attack America's national security," the letter, signed by 21 attorneys general to House and Senate leaders, said. "Given the Chinese desire to steal America's secrets and the ability of DeepSeek to carry out this theft, Congress should quickly pass legislation to ban DeepSeek on government devices," the letter read. "Congress passed similar legislation two years ago to prevent TikTok from stealing information from our government." Montana AG Austin Knudsen, who drafted the letter, wrote that "China is trying to steal America's secrets. Congress should shut down China's latest Trojan horse by passing the No DeepSeek on Government Devices Act."


Spaceraft carrying hopping robot to attempt Moon landing

BBC News

Nasa is partnering with a range of private companies that transport spacecraft and instruments to the Moon. It says this is cheaper than developing and blasting off their own missions. Intuitive Machines successfully landed a craft called Odysseus on the Moon in February last year, but it tipped over during the descent, meaning not all the scientific work could be carried out. Space agencies globally are competing to build human settlements on the Moon in a race to exploit resources and advance scientific understanding of other worlds. In the US, the Moon mission is seen as a stepping stone for the longer-term and much more ambitious goal of human settlement on Mars.