Goto

Collaborating Authors

 Government


Learning Homographic Disambiguation Representation for Neural Machine Translation

arXiv.org Artificial Intelligence

Homographs, words with the same spelling but different meanings, remain challenging in Neural Machine Translation (NMT). While recent works leverage various word embedding approaches to differentiate word sense in NMT, they do not focus on the pivotal components in resolving ambiguities of homographs in NMT: the hidden states of an encoder. In this paper, we propose a novel approach to tackle homographic issues of NMT in the latent space. We first train an encoder (aka "HDR-encoder") to learn universal sentence representations in a natural language inference (NLI) task. We further fine-tune the encoder using homograph-based synset sentences from WordNet, enabling it to learn word-level homographic disambiguation representations (HDR). The pre-trained HDR-encoder is subsequently integrated with a transformer-based NMT in various schemes to improve translation accuracy. Experiments on four translation directions demonstrate the effectiveness of the proposed method in enhancing the performance of NMT systems in the BLEU scores (up to +2.3 compared to a solid baseline). The effects can be verified by other metrics (F1, precision, and recall) of translation accuracy in an additional disambiguation task. Visualization methods like heatmaps, T-SNE and translation examples are also utilized to demonstrate the effects of the proposed method.


The growclusters Package for R

arXiv.org Artificial Intelligence

The growclusters package for R implements an enhanced version of k-means clustering that allows discovery of local clusterings or partitions for a collection of data sets that each draw their cluster means from a single, global partition. The package contains functions to estimate a partition structure for multivariate data. Estimation is performed under a penalized optimization derived from Bayesian non-parametric formulations. This paper describes some of the functions and capabilities of the growclusters package, including the creation of R Shiny applications designed to visually illustrate the operation and functionality of the growclusters package.


FollowMe: Vehicle Behaviour Prediction in Autonomous Vehicle Settings

arXiv.org Artificial Intelligence

An ego vehicle following a virtual lead vehicle planned route is an essential component when autonomous and non-autonomous vehicles interact. Yet, there is a question about the driver's ability to follow the planned lead vehicle route. Thus, predicting the trajectory of the ego vehicle route given a lead vehicle route is of interest. We introduce a new dataset, the FollowMe dataset, which offers a motion and behavior prediction problem by answering the latter question of the driver's ability to follow a lead vehicle. We also introduce a deep spatio-temporal graph model FollowMe-STGCNN as a baseline for the dataset. In our experiments and analysis, we show the design benefits of FollowMe-STGCNN in capturing the interactions that lie within the dataset. We contrast the performance of FollowMe-STGCNN with prior motion prediction models showing the need to have a different design mechanism to address the lead vehicle following settings.


When the Curious Abandon Honesty: Federated Learning Is Not Private

arXiv.org Artificial Intelligence

In federated learning (FL), data does not leave personal devices when they are jointly training a machine learning model. Instead, these devices share gradients, parameters, or other model updates, with a central party (e.g., a company) coordinating the training. Because data never "leaves" personal devices, FL is often presented as privacy-preserving. Yet, recently it was shown that this protection is but a thin facade, as even a passive, honest-but-curious attacker observing gradients can reconstruct data of individual users contributing to the protocol. In this work, we show a novel data reconstruction attack which allows an active and dishonest central party to efficiently extract user data from the received gradients. While prior work on data reconstruction in FL relies on solving computationally expensive optimization problems or on making easily detectable modifications to the shared model's architecture or parameters, in our attack the central party makes inconspicuous changes to the shared model's weights before sending them out to the users. We call the modified weights of our attack trap weights. Our active attacker is able to recover user data perfectly, i.e., with zero error, even when this data stems from the same class. Recovery comes with near-zero costs: the attack requires no complex optimization objectives. Instead, our attacker exploits inherent data leakage from model gradients and simply amplifies this effect by maliciously altering the weights of the shared model through the trap weights. These specificities enable our attack to scale to fully-connected and convolutional deep neural networks trained with large mini-batches of data. For example, for the high-dimensional vision dataset ImageNet, we perfectly reconstruct more than 50% of the training data points from mini-batches as large as 100 data points.


Data efficiency and extrapolation trends in neural network interatomic potentials

arXiv.org Artificial Intelligence

Over the last few years, key architectural advances have been proposed for neural network interatomic potentials (NNIPs), such as incorporating message-passing networks, equivariance, or many-body expansion terms. Although modern NNIP models exhibit small differences in energy/forces errors, improvements in accuracy are still considered the main target when developing new NNIP architectures. In this work, we show how architectural and optimization choices influence the generalization of NNIPs, revealing trends in molecular dynamics (MD) stability, data efficiency, and loss landscapes. Using the 3BPA dataset, we show that test errors in NNIP follow a scaling relation and can be robust to noise, but cannot predict MD stability in the high-accuracy regime. To circumvent this problem, we propose the use of loss landscape visualizations and a metric of loss entropy for predicting the generalization power of NNIPs. With a large-scale study on NequIP and MACE, we show that the loss entropy predicts out-of-distribution error and MD stability despite being computed only on the training set. Using this probe, we demonstrate how the choice of optimizers, loss function weighting, data normalization, and other architectural decisions influence the extrapolation behavior of NNIPs. Finally, we relate loss entropy to data efficiency, demonstrating that flatter landscapes also predict learning curve slopes. Our work provides a deep learning justification for the extrapolation performance of many common NNIPs, and introduces tools beyond accuracy metrics that can be used to inform the development of next-generation models.


Recent Advances in Modeling and Control of Epidemics using a Mean Field Approach

arXiv.org Artificial Intelligence

Modeling and control of epidemics such as the novel Corona virus have assumed paramount importance at a global level. A natural and powerful dynamical modeling framework to use in this context is a continuous time Markov decision process (CTMDP) that encompasses classical compartmental paradigms such as the Susceptible-Infected-Recovered (SIR) model. The challenges with CTMDP based models motivate the need for a more efficient approach and the mean field approach offers an effective alternative. The mean field approach computes the collective behavior of a dynamical system comprising numerous interacting nodes (where nodes represent individuals in the population). This paper (a) presents an overview of the mean field approach to epidemic modeling and control and (b) provides a state-of-the-art update on recent advances on this topic. Our discussion in this paper proceeds along two specific threads. The first thread assumes that the individual nodes faithfully follow a socially optimal control policy prescribed by a regulatory authority. The second thread allows the individual nodes to exhibit independent, strategic behavior. In this case, the strategic interaction is modeled as a mean field game and the control is based on the associated mean field Nash equilibria. In this paper, we start with a discussion of modeling of epidemics using an extended compartmental model - SIVR and provide an illustrative example. We next provide a review of relevant literature, using a mean field approach, on optimal control of epidemics, dealing with how a regulatory authority may optimally contain epidemic spread in a population. Following this, we provide an update on the literature on the use of the mean field game based approach in the study of epidemic spread and control. We conclude the paper with relevant future research directions.


The problems with a moratorium on training large AI systems

#artificialintelligence

In late March, the Future of Life Institute released an open letter (and a related FAQ) calling "on all AI labs to immediately pause for at least six months the training of AI systems more powerful than GPT-4. This pause should be public and verifiable, and include all key actors. If such a pause cannot be enacted quickly, governments should step in and institute a moratorium." The letter, which also stated that "Powerful AI systems should be developed only once we are confident that their effects will be positive and their risks will be manageable," was initially signed by over a thousand people, including many notable technology leaders. Many thousands more added their signatures after its publication.


Biden administration wants your input on rules for AI models like ChatGPT

Engadget

American officials are taking further steps to set rules for AI systems like ChatGPT. The National Telecommunications and Information Administration (NTIA) is asking for public comments on possible regulations that hold AI creators accountable. The measures will ideally help the Biden administration ensure that these models work as promised "without causing harm," the NTIA says. While the request is open-ended, the NTIA suggests input on areas like incentives for trustworthy AI, safety testing methods and the amount of data access needed to assess systems. The agency is also wondering if different strategies might be necessary for certain fields, such as healthcare.


'We have to move fast': US looks to establish rules for artificial intelligence

The Guardian

The US government is taking its first tentative steps toward establishing rules for artificial intelligence tools, as the frenzy over generative AI and chatbots reach a fever pitch. The US commerce department on Tuesday announced it is officially requesting public comment on how to create accountability measures for AI, seeking help on how to advise US policymakers to approach the technology. "In the same way that financial audits created trust in the accuracy of financial statements for businesses, accountability mechanisms for AI can help assure that an AI system is trustworthy," said Alan Davidson, the head of the National Telecommunications and Information Administration (NTIA), at a press conference at the University of Pittsburgh. Davidson said that the NTIA is seeking feedback from the public, including from researchers, industry groups, and privacy and digital rights organizations on the development of audits and assessments of AI tools created by private industry. He also said that the NTIA looking to establish guardrails that would allow the government to determine whether AI systems perform the way companies claim they do, whether they are safe and effective, whether they have discriminatory outcomes or "reflect unacceptable levels of bias", whether they spread or perpetuate misinformation, and whether they respect individuals' privacy.


Biden administration asks public for help regulating AI systems like ChatGPT

FOX News

Artificial Intelligence poses both risks and rewards, but developers should be weary of technologies that could threaten "scary" outcomes, AI technologist says. Federal regulators are asking the public for input on policies that would hold artificial intelligence (AI) systems accountable and help manage risks from the rapidly growing and powerful technology. As programs like ChatGPT gain popularity for their astounding ability to answer written questions with human-like responses, policymakers and tech experts are increasingly concerned with their potential for misuse, including how artificially-generated news reports can rapidly spread fabricated and false information. Now that ChatGPT has more than 100 million monthly active users, the government is beginning to study how these programs should be regulated. The National Telecommunications and Information Administration, a Commerce Department agency that advises the White House on telecommunications and information policy, solicited public feedback Tuesday as it works to develop policies to "ensure artificial intelligence (AI) systems work as claimed – and without causing harm."