Goto

Collaborating Authors

 Government


RegNLP in Action: Facilitating Compliance Through Automated Information Retrieval and Answer Generation

arXiv.org Artificial Intelligence

Regulatory documents, issued by governmental regulatory bodies, establish rules, guidelines, and standards that organizations must adhere to for legal compliance. These documents, characterized by their length, complexity and frequent updates, are challenging to interpret, requiring significant allocation of time and expertise on the part of organizations to ensure ongoing compliance.Regulatory Natural Language Processing (RegNLP) is a multidisciplinary subfield aimed at simplifying access to and interpretation of regulatory rules and obligations. We define an Automated Question-Passage Generation task for RegNLP, create the ObliQA dataset containing 27,869 questions derived from the Abu Dhabi Global Markets (ADGM) financial regulation document collection, design a baseline Regulatory Information Retrieval and Answer Generation system, and evaluate it with RePASs, a novel evaluation metric that tests whether generated answers accurately capture all relevant obligations and avoid contradictions.


Categorical data clustering: 25 years beyond K-modes

arXiv.org Artificial Intelligence

The clustering of categorical data is a common and important task in computer science, offering profound implications across a spectrum of applications. Unlike purely numerical data, categorical data often lack inherent ordering as in nominal data, or have varying levels of order as in ordinal data, thus requiring specialized methodologies for efficient organization and analysis. This review provides a comprehensive synthesis of categorical data clustering in the past twenty-five years, starting from the introduction of K-modes. It elucidates the pivotal role of categorical data clustering in diverse fields such as health sciences, natural sciences, social sciences, education, engineering and economics. Practical comparisons are conducted for algorithms having public implementations, highlighting distinguishing clustering methodologies and revealing the performance of recent algorithms on several benchmark categorical datasets. Finally, challenges and opportunities in the field are discussed.


Using machine learning for fault detection in lighthouse light sensors

arXiv.org Artificial Intelligence

Lighthouses play a crucial role in ensuring maritime safety by signaling hazardous areas such as dangerous coastlines, shoals, reefs, and rocks, along with aiding harbor entries and aerial navigation. This is achieved through the use of photoresistor sensors that activate or deactivate based on the time of day. However, a significant issue is the potential malfunction of these sensors, leading to the gradual misalignment of the light's operational timing. This paper introduces an innovative machine learning-based approach for automatically detecting such malfunctions. We evaluate four distinct algorithms: decision trees, random forest, extreme gradient boosting, and multi-layer perceptron. Our findings indicate that the multi-layer perceptron is the most effective, capable of detecting timing discrepancies as small as 10-15 minutes. This accuracy makes it a highly efficient tool for automating the detection of faults in lighthouse light sensors.


Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review

arXiv.org Artificial Intelligence

Large Language Model (LLM) pretraining traditionally relies on autoregressive language modeling on randomly sampled data blocks from web-scale datasets. We take inspiration from human learning techniques like spaced repetition to hypothesize that random data sampling for LLMs leads to high training cost and low quality models which tend to forget data. In order to effectively commit web-scale information to long-term memory, we propose the LFR (Learn, Focus, and Review) pedagogy, a new dynamic training paradigm which focuses and repeatedly reviews complex data blocks at systematic intervals based on the model's learning pace and progress. LFR records the model perplexities for different data blocks and frequently revisits blocks with higher perplexity which are more likely to be forgotten. We pretrain the GPT-2 models (124M - 1.5B) from scratch on the OpenWebText dataset using LFR. We test on downstream tasks from the language modeling, question answering, translation, and problem solving domains to achieve consistently lower perplexity and higher accuracy than the baseline OpenAI models, while obtaining a 20x pretraining speed-up.


OneEdit: A Neural-Symbolic Collaboratively Knowledge Editing System

arXiv.org Artificial Intelligence

Knowledge representation has been a central aim of AI since its inception. Symbolic Knowledge Graphs (KGs) and neural Large Language Models (LLMs) can both represent knowledge. KGs provide highly accurate and explicit knowledge representation, but face scalability issue; while LLMs offer expansive coverage of knowledge, but incur significant training costs and struggle with precise and reliable knowledge manipulation. To this end, we introduce OneEdit, a neural-symbolic prototype system for collaborative knowledge editing using natural language, which facilitates easy-to-use knowledge management with KG and LLM. OneEdit consists of three modules: 1) The Interpreter serves for user interaction with natural language; 2) The Controller manages editing requests from various users, leveraging the KG with rollbacks to handle knowledge conflicts and prevent toxic knowledge attacks; 3) The Editor utilizes the knowledge from the Controller to edit KG and LLM. We conduct experiments on two new datasets with KGs which demonstrate that OneEdit can achieve superior performance.


Explainable AI: Definition and attributes of a good explanation for health AI

arXiv.org Artificial Intelligence

Proposals of artificial intelligence (AI) solutions based on increasingly complex and accurate predictive models are becoming ubiquitous across many disciplines. As the complexity of these models grows, transparency and users' understanding often diminish. This suggests that accurate prediction alone is insufficient for making an AI-based solution truly useful. In the development of healthcare systems, this introduces new issues related to accountability and safety. Understanding how and why an AI system makes a recommendation may require complex explanations of its inner workings and reasoning processes. Although research on explainable AI (XAI) has significantly increased in recent years and there is high demand for XAI in medicine, defining what constitutes a good explanation remains ad hoc, and providing adequate explanations continues to be challenging. To fully realize the potential of AI, it is critical to address two fundamental questions about explanations for safety-critical AI applications, such as health-AI: (1) What is an explanation in health-AI? and (2) What are the attributes of a good explanation in health-AI? In this study, we examined published literature and gathered expert opinions through a two-round Delphi study. The research outputs include (1) a definition of what constitutes an explanation in health-AI and (2) a comprehensive list of attributes that characterize a good explanation in health-AI.


Explainable Malware Analysis: Concepts, Approaches and Challenges

arXiv.org Artificial Intelligence

Machine learning (ML) has seen exponential growth in recent years, finding applications in various domains such as finance, medicine, and cybersecurity. Malware remains a significant threat to modern computing, frequently used by attackers to compromise systems. While numerous machine learning-based approaches for malware detection achieve high performance, they often lack transparency and fail to explain their predictions. This is a critical drawback in malware analysis, where understanding the rationale behind detections is essential for security analysts to verify and disseminate information. Explainable AI (XAI) addresses this issue by maintaining high accuracy while producing models that provide clear, understandable explanations for their decisions. In this survey, we comprehensively review the current state-of-the-art ML-based malware detection techniques and popular XAI approaches. Additionally, we discuss research implementations and the challenges of explainable malware analysis. This theoretical survey serves as an entry point for researchers interested in XAI applications in malware detection. By analyzing recent advancements in explainable malware analysis, we offer a broad overview of the progress in this field, positioning our work as the first to extensively cover XAI methods for malware classification and detection.


The BBC breached editorial guidelines over 1,500 times in Israel-Hamas conflict, report claims

FOX News

A new report found the British Broadcasting Corporation (BBC) guilty of violating its own editorial guidelines over a thousand times in its coverage of the Israel-Hamas war. According to The Telegraph, the report analyzed four months of BBC output on television, radio, online, podcasts and on social media during the height of the conflict and found a "deeply worrying pattern of bias" against Israel. British lawyer Trevor Asserson and a team of about 20 lawyers and 20 data scientists used artificial intelligence to analyze nine million words from the news outlet, starting the day of the October 7, 2023, terror attack. The researchers allegedly identified 1,553 instances where the BBC violated its own editorial guidelines on impartiality, accuracy, editorial values and public interest. Hundreds attend a protest called by the National Jewish Assembly, The Campaign Against Antisemitsim and the UK Lawyers for Israel at the BBC Broadcasting House on October 16, 2023, in London, England.


Houthis claim downing another US MQ-9 Reaper drone over Yemen

Al Jazeera

The Houthis have claimed to have shot down a United States military drone over Yemen, in the latest attack by the group, which has disrupted shipping trade through the crucial Bab al-Mandeb Strait, drawing US strikes. The Yemeni group has carried out dozens of attacks on ships with links to Israel in a show of solidarity with Palestinians amid Israel's 11-month-old war on Gaza. Yahya Saree, the military spokesman of the Houthi group, said in a prerecorded video message released early on Sunday that the MQ-9 Reaper was shot down by air defences over Marib as "it was carrying out hostile activities". This is the eighth drone of this type to be shot down since the start of the war on Gaza, he said. The group has not so far released footage of the downed attack and surveillance aircraft that costs about 30m.


Japanese bank seeks to help regional economy with bus business

The Japan Times

Japanese regional banking group Senshu Ikeda Holdings' entry into the reservation-based transit bus business is aimed at stimulating the regional economy, President and CEO Atsushi Ukawa said in a recent interview. "Even regional banks in urban areas must think about serving the local community," Ukawa said of the first reservation bus operations by a regional bank in Japan. He said that the Osaka-based company will work with local governments to expand the operation area to complement public transport. Senshu Ikeda operates an "on-demand bus," which uses artificial intelligence to run according to users' desired dates, times and locations. It partnered with companies, including auto parts maker Aisin, to launch the bus operations on a trial basis in four municipalities in Osaka Prefecture in January 2023.