Goto

Collaborating Authors

 Government


On Event Individuation for Document-Level Information Extraction

arXiv.org Artificial Intelligence

As information extraction (IE) systems have grown more adept at processing whole documents, the classic task of template filling has seen renewed interest as benchmark for document-level IE. In this position paper, we call into question the suitability of template filling for this purpose. We argue that the task demands definitive answers to thorny questions of event individuation -- the problem of distinguishing distinct events -- about which even human experts disagree. Through an annotation study and error analysis, we show that this raises concerns about the usefulness of template filling metrics, the quality of datasets for the task, and the ability of models to learn it. Finally, we consider possible solutions.


Federated Unlearning: How to Efficiently Erase a Client in FL?

arXiv.org Artificial Intelligence

With privacy legislation empowering the users with the right to be forgotten, it has become essential to make a model amenable for forgetting some of its training data. However, existing unlearning methods in the machine learning context can not be directly applied in the context of distributed settings like federated learning due to the differences in learning protocol and the presence of multiple actors. In this paper, we tackle the problem of federated unlearning for the case of erasing a client by removing the influence of their entire local data from the trained global model. To erase a client, we propose to first perform local unlearning at the client to be erased, and then use the locally unlearned model as the initialization to run very few rounds of federated learning between the server and the remaining clients to obtain the unlearned global model. We empirically evaluate our unlearning method by employing multiple performance measures on three datasets, and demonstrate that our unlearning method achieves comparable performance as the gold standard unlearning method of federated retraining from scratch, while being significantly efficient. Unlike prior works, our unlearning method neither requires global access to the data used for training nor the history of the parameter updates to be stored by the server or any of the clients.


A framework for spatial heat risk assessment using a generalized similarity measure

arXiv.org Artificial Intelligence

In this study, we develop a novel framework to assess health risks As it was noted by Intergovernmental Panel on Climate Change due to heat hazards across various localities (zip codes) across the (IPCC) (2014), impacts from extreme climate-related events emerge state of Maryland with the help of two commonly used indicators: from risk that are not only related to a specific hazard (e.g., heat exposure and vulnerability. Our approach quantifies each of the waves), but also directly depends on the two other elements; exposure two aforementioned indicators by developing their corresponding and vulnerability. Exposure addresses the population and assets feature vectors and subsequently computes indicator-specific reference at risk while vulnerability indicates the susceptibility of human and vectors that signify a high risk environment by clustering the natural systems during an extreme event[16].


Show, Write, and Retrieve: Entity-aware Article Generation and Retrieval

arXiv.org Artificial Intelligence

Article comprehension is an important challenge in natural language processing with many applications such as article generation or image-to-article retrieval. Prior work typically encodes all tokens in articles uniformly using pretrained language models. However, in many applications, such as understanding news stories, these articles are based on real-world events and may reference many named entities that are difficult to accurately recognize and predict by language models. To address this challenge, we propose an ENtity-aware article GeneratIoN and rEtrieval (ENGINE) framework, to explicitly incorporate named entities into language models. ENGINE has two main components: a named-entity extraction module to extract named entities from both metadata and embedded images associated with articles, and an entity-aware mechanism that enhances the model's ability to recognize and predict entity names. We conducted experiments on three public datasets: GoodNews, VisualNews, and WikiText, where our results demonstrate that our model can boost both article generation and article retrieval performance, with a 4-5 perplexity improvement in article generation and a 3-4% boost in recall@1 in article retrieval. We release our implementation at https://github.com/Zhongping-Zhang/ENGINE .


Modeling Supply and Demand in Public Transportation Systems

arXiv.org Machine Learning

We propose two neural network based and data-driven supply and demand models to analyze the efficiency, identify service gaps, and determine the significant predictors of demand, in the bus system for the Department of Public Transportation (HDPT) in Harrisonburg City, Virginia, which is the home to James Madison University (JMU). The supply and demand models, one temporal and one spatial, take many variables into account, including the demographic data surrounding the bus stops, the metrics that the HDPT reports to the federal government, and the drastic change in population between when JMU is on or off session. These direct and data-driven models to quantify supply and demand and identify service gaps can generalize to other cities' bus systems. Keywords-- transportation systems, bus systems, public transportation, direct ridership models, data driven models, mathematical modeling, neural networks, machine learning, supply models, demand models, machine learning, service gaps, social vulnerability, public transportation access, GIS data, data science, data quality.


Predicting Battery Lifetime Under Varying Usage Conditions from Early Aging Data

arXiv.org Machine Learning

Accurate battery lifetime prediction is important for preventative maintenance, warranties, and improved cell design and manufacturing. However, manufacturing variability and usage-dependent degradation make life prediction challenging. Here, we investigate new features derived from capacity-voltage data in early life to predict the lifetime of cells cycled under widely varying charge rates, discharge rates, and depths of discharge. Features were extracted from regularly scheduled reference performance tests (i.e., low rate full cycles) during cycling. The early-life features capture a cell's state of health and the rate of change of component-level degradation modes, some of which correlate strongly with cell lifetime. Using a newly generated dataset from 225 nickel-manganese-cobalt/graphite Li-ion cells aged under a wide range of conditions, we demonstrate a lifetime prediction of in-distribution cells with 15.1% mean absolute percentage error using no more than the first 15% of data, for most cells. Further testing using a hierarchical Bayesian regression model shows improved performance on extrapolation, achieving 21.8% mean absolute percentage error for out-of-distribution cells. Our approach highlights the importance of using domain knowledge of lithium-ion battery degradation modes to inform feature engineering. Further, we provide the community with a new publicly available battery aging dataset with cells cycled beyond 80% of their rated capacity.


Polar Ducks and Where to Find Them: Enhancing Entity Linking with Duck Typing and Polar Box Embeddings

arXiv.org Artificial Intelligence

Entity linking methods based on dense retrieval are an efficient and widely used solution in large-scale applications, but they fall short of the performance of generative models, as they are sensitive to the structure of the embedding space. In order to address this issue, this paper introduces DUCK, an approach to infusing structural information in the space of entity representations, using prior knowledge of entity types. Inspired by duck typing in programming languages, we propose to define the type of an entity based on the relations that it has with other entities in a knowledge graph. Then, porting the concept of box embeddings to spherical polar coordinates, we propose to represent relations as boxes on the hypersphere. We optimize the model to cluster entities of similar type by placing them inside the boxes corresponding to their relations. Our experiments show that our method sets new state-of-the-art results on standard entity-disambiguation benchmarks, it improves the performance of the model by up to 7.9 F1 points, outperforms other type-aware approaches, and matches the results of generative models with 18 times more parameters.


Families of Israeli hostages held by Hamas cling to digital clues

Washington Post - Technology News

Refael Franco, the former deputy head of the National Cyber Directorate, is overseeing the construction of the group's platform, which is based on cutting-edge AI and facial recognition technology. It is a sophisticated system that cross-checks images posted by militants posted on social media against photos of the hostages provided by families. When a match occurs, the system can geolocate -- within seconds -- the approximate location of a missing person. Open-source intelligence experts then try to zero in further, relying on contextual clues like mosques, local shops or the angle of the sun.


The US Has Failed to Pass AI Regulation. New York City Is Stepping Up

WIRED

As the US federal government struggles to meaningfully regulate AI--or even function--New York City is stepping into the governance gap. The city introduced an AI Action Plan this week that mayor Eric Adams calls a first of its kind in the nation. The set of roughly 40 policy initiatives is designed to protect residents against harm like bias or discrimination from AI. It includes development of standards for AI purchased by city agencies and new mechanisms to gauge the risk of AI used by city departments. New York's AI regulation could soon expand still further.


Drone strikes target US military bases in Syria, Iraq as regional tensions from Israel-Hamas War escalate

FOX News

Drone expert Brett Velicovich joined'FOX & Friends First' to discuss attacks against U.S. military bases in the Middle East as war rages between Israel and Hamas. A drone strike targeted a U.S. base in Syria on Wednesday, the same day as the attempted drone attacks in Iraq, Fox News has learned. A U.S. defense official told Fox News that an undisclosed number of drones targeted the U.S. Al-Tanf base, located near Syria's shared border with Iraq and Jordan. Lebanon's Iran-aligned Al Mayadeen TV reported on Thursday that two U.S. military bases in Syria came under attack, Reuters reported. In addition to the drone attack on the Al-Tanf base, Al Mayadeen TV reported a missile targeted the Conoco base in the countryside of the northern Deir al-Zor region.