Goto

Collaborating Authors

 Government


Sharded Bayesian Additive Regression Trees

arXiv.org Artificial Intelligence

In this paper we develop the randomized Sharded Bayesian Additive Regression Trees (SBT) model. We introduce a randomization auxiliary variable and a sharding tree to decide partitioning of data, and fit each partition component to a sub-model using Bayesian Additive Regression Tree (BART). By observing that the optimal design of a sharding tree can determine optimal sharding for sub-models on a product space, we introduce an intersection tree structure to completely specify both the sharding and modeling using only tree structures. In addition to experiments, we also derive the theoretical optimal weights for minimizing posterior contractions and prove the worst-case complexity of SBT.


Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery

arXiv.org Artificial Intelligence

The strength of modern generative models lies in their ability to be controlled through text-based prompts. Typical "hard" prompts are made from interpretable words and tokens, and must be hand-crafted by humans. There are also "soft" prompts, which consist of continuous feature vectors. These can be discovered using powerful optimization methods, but they cannot be easily interpreted, re-used across models, or plugged into a text-based interface. We describe an approach to robustly optimize hard text prompts through efficient gradient-based optimization. Our approach automatically generates hard text-based prompts for both text-to-image and text-to-text applications. In the text-to-image setting, the method creates hard prompts for diffusion models, allowing API users to easily generate, discover, and mix and match image concepts without prior knowledge on how to prompt the model. In the text-to-text setting, we show that hard prompts can be automatically discovered that are effective in tuning LMs for classification.


Retrosynthetic Planning with Dual Value Networks

arXiv.org Artificial Intelligence

Retrosynthesis, which aims to find a route to synthesize a target molecule from commercially available starting materials, is a critical task in drug discovery and materials design. Recently, the combination of ML-based single-step reaction predictors with multi-step planners has led to promising results. However, the single-step predictors are mostly trained offline to optimize the single-step accuracy, without considering complete routes. Here, we leverage reinforcement learning (RL) to improve the single-step predictor, by using a tree-shaped MDP to optimize complete routes. Specifically, we propose a novel online training algorithm, called Planning with Dual Value Networks (PDVN), which alternates between the planning phase and updating phase. In PDVN, we construct two separate value networks to predict the synthesizability and cost of molecules, respectively. To maintain the single-step accuracy, we design a two-branch network structure for the single-step predictor. On the widely-used USPTO dataset, our PDVN algorithm improves the search success rate of existing multi-step planners (e.g., increasing the success rate from 85.79% to 98.95% for Retro*, and reducing the number of model calls by half while solving 99.47% molecules for RetroGraph). Additionally, PDVN helps find shorter synthesis routes (e.g., reducing the average route length from 5.76 to 4.83 for Retro*, and from 5.63 to 4.78 for RetroGraph).


Reward Gaming in Conditional Text Generation

arXiv.org Artificial Intelligence

To align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations. Under this framework, we identify three common cases where high rewards are incorrectly assigned to undesirable patterns: noise-induced spurious correlation, naturally occurring spurious correlation, and covariate shift. We show that even though learned metrics achieve high performance on the distribution of the data used to train the reward function, the undesirable patterns may be amplified during RL training of the text generation model. While there has been discussion about reward gaming in the RL or safety community, in this discussion piece, we would like to highlight reward gaming in the natural language generation (NLG) community using concrete conditional text generation examples and discuss potential fixes and areas for future work.


Speaking Multiple Languages Affects the Moral Bias of Language Models

arXiv.org Artificial Intelligence

Pre-trained multilingual language models (PMLMs) are commonly used when dealing with data from multiple languages and cross-lingual transfer. However, PMLMs are trained on varying amounts of data for each language. In practice this means their performance is often much better on English than many other languages. We explore to what extent this also applies to moral norms. Do the models capture moral norms from English and impose them on other languages? Do the models exhibit random and thus potentially harmful beliefs in certain languages? Both these issues could negatively impact cross-lingual transfer and potentially lead to harmful outcomes. In this paper, we (1) apply the MoralDirection framework to multilingual models, comparing results in German, Czech, Arabic, Chinese, and English, (2) analyse model behaviour on filtered parallel subtitles corpora, and (3) apply the models to a Moral Foundations Questionnaire, comparing with human responses from different countries. Our experiments demonstrate that, indeed, PMLMs encode differing moral biases, but these do not necessarily correspond to cultural differences or commonalities in human opinions. We release our code and models.


Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations

arXiv.org Artificial Intelligence

Most weakly supervised named entity recognition (NER) models rely on domain-specific dictionaries provided by experts. This approach is infeasible in many domains where dictionaries do not exist. While a phrase retrieval model was used to construct pseudo-dictionaries with entities retrieved from Wikipedia automatically in a recent study, these dictionaries often have limited coverage because the retriever is likely to retrieve popular entities rather than rare ones. In this study, we present a novel framework, HighGEN, that generates NER datasets with high-coverage pseudo-dictionaries. Specifically, we create entity-rich dictionaries with a novel search method, called phrase embedding search, which encourages the retriever to search a space densely populated with various entities. In addition, we use a new verification process based on the embedding distance between candidate entity mentions and entity types to reduce the false-positive noise in weak labels generated by high-coverage dictionaries. We demonstrate that HighGEN outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets.


Federal regulators fine Amazon $25 million over child privacy issues

Washington Post - Technology News

The U.S. government alleges that Amazon violated the Children's Online Privacy Protection Act, a 1998 law that has recently been enforced against other popular tech companies including Fortnite-maker Epic Games and YouTube. More than 800,000 children under the age of 13 have their own Alexa profiles, according to the lawsuit filed by the Department of Justice on behalf of the Federal Trade Commission. About five years ago, the company began offering a number of products specifically aimed at children, including the "Echo Dot Kids Edition" smart speaker and parental controls called "FreeTime on Alexa."


US, allies prep voluntary AI code of conduct, Blinken says

FOX News

Center for A.I. Safety Director Dan Hendrycks explains concerns about how the rapid growth of artificial intelligence could impact society. Secretary of State Antony Blinken said Wednesday that the United States is working with its European allies to develop a conduct code for artificial intelligence. Blinken is in Sweden for a meeting of the EU-U.S. Trade and Technology Council, which is jointly led by American and European officials. "We need accountable artificial intelligence. Generative AI is a complete game changer," European Commission Vice President Margrethe Vestager said at a press conference after the meeting, saying a draft of a voluntary code of conduct for artificial intelligence would be ready within a matter of weeks.


GM is developing a drone-killing off-road pickup for the US Army

FOX News

A General Motors pickup has never hauled something like this. GM Defense is collaborating with military contractor Black Sage Technologies to integrate a drone defense system into the Infantry Squad Vehicle (ISV) that GM Defense recently began supplying to the US Army. The ISV is based on the last-generation Chevrolet Colorado ZR2 midsize pickup and manufactured in Concord, N.C., using frames supplied by NASCAR's Hendrick Motorsports. The midsize truck was engineered for high-speed off-road driving and designed to fit inside a CH-47 Chinook helicopter, slung from a UH-60 Blackhawk helicopter, or air-dropped from a cargo plane by parachute for quick deployment into the field. The vehicle can be outfitted to fit nine troops, but there are several configurations that mix passenger, cargo and arms carrying capabilities.


Putin says drone attacks on Moscow are attempt by Ukraine 'to intimidate Russia'

FOX News

Senior foreign affairs correspondent Greg Palkot reports the latest from London. Russian President Vladimir Putin is speaking out following a drone attack on Moscow, calling the strikes an attempt by Ukraine to "intimidate" his country. The remarks come after eight drones targeted Russia's capital early Tuesday before being shot down or diverted with electronic jammers. Moscow Mayor Sergei Sobyanin said the attack caused "insignificant damage" to several buildings and that two people received treatment for unspecified injuries but did not need hospitalization. Residents of two high-rise buildings damaged in the attack were evacuated.