Goto

Collaborating Authors

 Government


TrIM: Transformed Iterative Mondrian Forests for Gradient-based Dimension Reduction and High-Dimensional Regression

arXiv.org Machine Learning

We propose a computationally efficient algorithm for gradient-based linear dimension reduction and high-dimensional regression. The algorithm initially computes a Mondrian forest and uses this estimator to identify a relevant feature subspace of the inputs from an estimate of the expected gradient outer product (EGOP) of the regression function. In addition, we introduce an iterative approach known as Transformed Iterative Mondrian (TrIM) forest to improve the Mondrian forest estimator by using the EGOP estimate to update the set of features and weights used by the Mondrian partitioning mechanism. We obtain consistency guarantees and convergence rates for the estimation of the EGOP matrix and the random forest estimator obtained from one iteration of the TrIM algorithm. Lastly, we demonstrate the effectiveness of our proposed algorithm for learning the relevant feature subspace across a variety of settings with both simulated and real data.


Three senators introduce bill to protect artists and journalists from unauthorized AI use

Engadget

Three US Senators introduced a bill that aims to rein in the rise and use of AI generated content and deepfakes by protecting the work of artists, songwriters and journalists. The Content Original Protection and Integrity from Edited and Deepfaked Media (COPIED) Act was introduced to the Senate Friday morning. The bill is a bipartisan effort authorized by Sen. Marsha Blackburn (R-Tenn.), Sen. Maria Cantwell (D-Wash.) and Sen. Martin Heinrich (D-N.M.), according to a press alert issued by Blackburn's office. The COPIED ACT would, if enacted, create transparency standards through the National Institutes of Standards and Technology (NIST) to set guidelines for "content provenance information, watermarking, and synthetic content detection," according to the press release.


What We Know About the New U.K. Government's Approach to AI

TIME - Tech

When the U.K. hosted the world's first AI Safety Summit last November, Rishi Sunak, the then Prime Minister, said the achievements at the event would "tip the balance in favor of humanity." At the two-day event, held in the cradle of modern computing, Bletchley Park, AI labs committed to share their models with governments before public release, and 29 countries pledged to collaborate on mitigating risks from artificial intelligence. It was part of the Sunak-led Conservative government's effort to position the U.K. as a leader in artificial intelligence governance, which also involved establishing the world's first AI Safety Institute--a government body tasked with evaluating models for potentially dangerous capabilities. While the U.S. and other allied nations subsequently set up their own similar institutes, the U.K. institute boasts 10 times the funding of its American counterpart. Eight months later, on July 5, after a landslide loss to the Labour Party, Sunak left office and the newly elected Prime Minister Keir Starmer began forming his new government.


The EU will start enforcing its new AI regulations on August 1

Engadget

The European Union has published the full and final text for the EU AI Act in its Official Journal, as reported by TechCrunch. Since the new law will come into force 20 days after its publication, that means it will be enforceable starting on August 1. All its provisions will be fully applicable in two years' time, but some of them will be implemented much earlier than that. Six months from now, the bloc will start implementing bans on prohibited applications for AI, such as the use of social credit ranking systems, the collection and compilation of facial recognition information for databases, as well the use of real time emotion recognition systems in schools and workplaces. In nine months, the EU will start implementing codes of practice on AI developers.


Japan and Britain agree on wide-ranging cooperation, including in AI

The Japan Times

Prime Minister Fumio Kishida and new British Prime Minister Keir Starmer agreed on Thursday to boost cooperation between their countries in a wide range of fields including artificial intelligence. In their first in-person talks, held in Washington, the two leaders also confirmed that the security of the Euro-Atlantic and Indo-Pacific regions are inseparable. Kishida on July 6 had telephone talks with Starmer, who led his party to a landslide general election victory on July 4. They also agreed to promote the joint development of the next-generation fighter jet involving their countries plus Italy and discussed the situations in the Middle East and East Asia, including North Korea, and affirmed their close collaboration. Separately, Kishida met with new Dutch Prime Minister Dick Schoof and congratulated him on his inauguration.


U.S. lawmakers raise concerns over Microsoft deal with Emirati AI firm

The Japan Times

U.S. Republican lawmakers asked the administration of President Joe Biden for an intelligence assessment of Microsoft's 1.5 billion investment in UAE-based artificial intelligence firm G42 over concerns about the transfer of sensitive technology and G42's historic ties to China. Rep. Michael McCaul, chair of the House Foreign Affairs Committee, and John Moolenaar, leader of the Select Committee on China, made the request for a briefing in a letter dated Wednesday to White House national security adviser Jake Sullivan, the committees said. The Republicans said they want the briefing on the deal, announced in April, before it advances to a second phase involving the transfer of export-restricted semiconductor chips and model weights, sophisticated data that improves an AI model's ability to emulate human reasoning.


Beyond static AI evaluations: advancing human interaction evaluations for LLM harms and risks

arXiv.org Artificial Intelligence

Model evaluations are central to understanding the safety, risks, and societal impacts of AI systems. While most real-world AI applications involve human-AI interaction, most current evaluations (e.g., common benchmarks) of AI models do not. Instead, they incorporate human factors in limited ways, assessing the safety of models in isolation, thereby falling short of capturing the complexity of human-model interactions. In this paper, we discuss and operationalize a definition of an emerging category of evaluations -- "human interaction evaluations" (HIEs) -- which focus on the assessment of human-model interactions or the process and the outcomes of humans using models. First, we argue that HIEs can be used to increase the validity of safety evaluations, assess direct human impact and interaction-specific harms, and guide future assessments of models' societal impact. Second, we propose a safety-focused HIE design framework -- containing a human-LLM interaction taxonomy -- with three stages: (1) identifying the risk or harm area, (2) characterizing the use context, and (3) choosing the evaluation parameters. Third, we apply our framework to two potential evaluations for overreliance and persuasion risks. Finally, we conclude with tangible recommendations for addressing concerns over costs, replicability, and unrepresentativeness of HIEs.


AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities

arXiv.org Artificial Intelligence

In this survey we are focusing on utilizing drone-based systems for the detection of individuals, particularly by identifying human screams and other distress signals. This study has significant relevance in post-disaster scenarios, including events such as earthquakes, hurricanes, military conflicts, wildfires, and more. These drones are capable of hovering over disaster-stricken areas that may be challenging for rescue teams to access directly. Unmanned aerial vehicles (UAVs), commonly referred to as drones, are frequently deployed for search-and-rescue missions during disaster situations. Typically, drones capture aerial images to assess structural damage and identify the extent of the disaster. They also employ thermal imaging technology to detect body heat signatures, which can help locate individuals. In some cases, larger drones are used to deliver essential supplies to people stranded in isolated disaster-stricken areas. In our discussions, we delve into the unique challenges associated with locating humans through aerial acoustics. The auditory system must distinguish between human cries and sounds that occur naturally, such as animal calls and wind. Additionally, it should be capable of recognizing distinct patterns related to signals like shouting, clapping, or other ways in which people attempt to signal rescue teams. To tackle this challenge, one solution involves harnessing artificial intelligence (AI) to analyze sound frequencies and identify common audio signatures. Deep learning-based networks, such as convolutional neural networks (CNNs), can be trained using these signatures to filter out noise generated by drone motors and other environmental factors. Furthermore, employing signal processing techniques like the direction of arrival (DOA) based on microphone array signals can enhance the precision of tracking the source of human noises.


OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

arXiv.org Artificial Intelligence

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-level image-text interleaved dataset. Using an efficient data engine, we filter and extract large-scale high-quality documents, which contain 8.6 billion images and 1,696 billion text tokens. Compared to counterparts (e.g., MMC4, OBELICS), our dataset 1) has 15 times larger scales while maintaining good data quality; 2) features more diverse sources, including both English and non-English websites as well as video-centric websites; 3) is more flexible, easily degradable from an image-text interleaved format to pure text corpus and image-text pairs. Through comprehensive analysis and experiments, we validate the quality, usability, and effectiveness of the proposed dataset. We hope this could provide a solid data foundation for future multimodal model research.


A Survey on Symbolic Knowledge Distillation of Large Language Models

arXiv.org Artificial Intelligence

This survey paper delves into the emerging and critical area of symbolic knowledge distillation in Large Language Models (LLMs). As LLMs like Generative Pre-trained Transformer-3 (GPT-3) and Bidirectional Encoder Representations from Transformers (BERT) continue to expand in scale and complexity, the challenge of effectively harnessing their extensive knowledge becomes paramount. This survey concentrates on the process of distilling the intricate, often implicit knowledge contained within these models into a more symbolic, explicit form. This transformation is crucial for enhancing the interpretability, efficiency, and applicability of LLMs. We categorize the existing research based on methodologies and applications, focusing on how symbolic knowledge distillation can be used to improve the transparency and functionality of smaller, more efficient Artificial Intelligence (AI) models. The survey discusses the core challenges, including maintaining the depth of knowledge in a comprehensible format, and explores the various approaches and techniques that have been developed in this field. We identify gaps in current research and potential opportunities for future advancements. This survey aims to provide a comprehensive overview of symbolic knowledge distillation in LLMs, spotlighting its significance in the progression towards more accessible and efficient AI systems.