Government
Data Augmentations for Improved (Large) Language Model Generalization
Feder, Amir, Wald, Yoav, Shi, Claudia, Saria, Suchi, Blei, David
The reliance of text classifiers on spurious correlations can lead to poor generalization at deployment, raising concerns about their use in safety-critical domains such as healthcare. In this work, we propose to use counterfactual data augmentation, guided by knowledge of the causal structure of the data, to simulate interventions on spurious features and to learn more robust text classifiers. We show that this strategy is appropriate in prediction problems where the label is spuriously correlated with an attribute. Under the assumptions of such problems, we discuss the favorable sample complexity of counterfactual data augmentation, compared to importance re-weighting. Pragmatically, we match examples using auxiliary data, based on diff-in-diff methodology, and use a large language model (LLM) to represent a conditional probability of text. Through extensive experimentation on learning caregiver-invariant predictors of clinical diagnoses from medical narratives and on semi-synthetic data, we demonstrate that our method for simulating interventions improves out-of-distribution (OOD) accuracy compared to baseline invariant learning algorithms.
The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
Salinas, Abel, Shah, Parth Vipul, Huang, Yuzhong, McCormack, Robert, Morstatter, Fred
Large Language Models (LLMs) have seen widespread deployment in various real-world applications. Understanding these biases is crucial to comprehend the potential downstream consequences when using LLMs to make decisions, particularly for historically disadvantaged groups. In this work, we propose a simple method for analyzing and comparing demographic bias in LLMs, through the lens of job recommendations. We demonstrate the effectiveness of our method by measuring intersectional biases within ChatGPT and LLaMA, two cutting-edge LLMs. Our experiments primarily focus on uncovering gender identity and nationality bias; however, our method can be extended to examine biases associated with any intersection of demographic identities. We identify distinct biases in both models toward various demographic identities, such as both models consistently suggesting low-paying jobs for Mexican workers or preferring to recommend secretarial roles to women. Our study highlights the importance of measuring the bias of LLMs in downstream applications to understand the potential for harm and inequitable outcomes.
Stable generative modeling using diffusion maps
Gottwald, Georg, Li, Fengyi, Marzouk, Youssef, Reich, Sebastian
We consider the problem of sampling from an unknown distribution for which only a sufficiently large number of training samples are available. Such settings have recently drawn considerable interest in the context of generative modelling. In this paper, we propose a generative model combining diffusion maps and Langevin dynamics. Diffusion maps are used to approximate the drift term from the available training samples, which is then implemented in a discrete-time Langevin sampler to generate new samples. By setting the kernel bandwidth to match the time step size used in the unadjusted Langevin algorithm, our method effectively circumvents any stability issues typically associated with time-stepping stiff stochastic differential equations. More precisely, we introduce a novel split-step scheme, ensuring that the generated samples remain within the convex hull of the training samples. Our framework can be naturally extended to generate conditional samples. We demonstrate the performance of our proposed scheme through experiments on synthetic datasets with increasing dimensions and on a stochastic subgrid-scale parametrization conditional sampling problem.
Auditing and Generating Synthetic Data with Controllable Trust Trade-offs
Belgodere, Brian, Dognin, Pierre, Ivankay, Adam, Melnyk, Igor, Mroueh, Youssef, Mojsilovic, Aleksandra, Navratil, Jiri, Nitsure, Apoorva, Padhi, Inkit, Rigotti, Mattia, Ross, Jerret, Schiff, Yair, Vedpathak, Radhika, Young, Richard A.
Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate unbiased, privacy-preserving data while maintaining fidelity to the original data. However, assessing the trustworthiness of synthetic datasets and models is a critical challenge. We introduce a holistic auditing framework that comprehensively evaluates synthetic datasets and AI models. It focuses on preventing bias and discrimination, ensures fidelity to the source data, assesses utility, robustness, and privacy preservation. We demonstrate the framework's effectiveness by auditing various generative models across diverse use cases like education, healthcare, banking, and human resources, spanning different data modalities such as tabular, time-series, vision, and natural language. This holistic assessment is essential for compliance with regulatory safeguards. We introduce a trustworthiness index to rank synthetic datasets based on their safeguards trade-offs. Furthermore, we present a trustworthiness-driven model selection and cross-validation process during training, exemplified with "TrustFormers" across various data types. This approach allows for controllable trustworthiness trade-offs in synthetic data creation. Our auditing framework fosters collaboration among stakeholders, including data scientists, governance experts, internal reviewers, external certifiers, and regulators. This transparent reporting should become a standard practice to prevent bias, discrimination, and privacy violations, ensuring compliance with policies and providing accountability, safety, and performance guarantees.
California wants to reduce traffic. The Newsom administration thinks AI can help
Being stuck in traffic is a familiar problem for many Californians, but state officials want to harness the power of artificial intelligence to discover new solutions. The California Department of Transportation, teaming up with other state agencies, is asking technology companies by Jan. 25 to propose generative AI tools that could help California reduce traffic and make roads safer, especially for pedestrians, cyclists and scooter riders. Generative AI tools such as ChatGPT can quickly produce text, images and other content, but the technology can also help workers brainstorm ideas. The request shows how California is trying to tap into AI to improve government services at a time when lawmakers seek to safeguard against the technology's potential risks. California politicians set the stage for more AI regulation in 2024, but they'll also face challenges as they try to place more guardrails around AI's impact on jobs, safety and discrimination.
California wants to reduce traffic. The Newsom administration thinks AI can help
Being stuck in traffic is a familiar problem for many Californians, but state officials want to harness the power of artificial intelligence to discover new solutions. The California Department of Transportation, teaming up with other state agencies, is asking technology companies by Jan. 25 to propose generative AI tools that could help California reduce traffic and make roads safer, especially for pedestrians, cyclists and scooter riders. Generative AI tools such as ChatGPT can quickly produce text, images and other content, but the technology can also help workers brainstorm ideas. The request shows how California is trying to tap into AI to improve government services at a time when lawmakers seek to safeguard against the technology's potential risks. California politicians set the stage for more AI regulation in 2024, but they'll also face challenges as they try to place more guardrails around AI's impact on jobs, safety and discrimination.
Judges in England and Wales Given Cautious Approval to Use AI in Writing Legal Opinions
England's 1,000-year-old legal system -- still steeped in traditions that include wearing wigs and robes -- has taken a cautious step into the future by giving judges permission to use artificial intelligence to help produce rulings. The Courts and Tribunals Judiciary last month said AI could help write opinions but stressed it shouldn't be used for research or legal analyses because the technology can fabricate information and provide misleading, inaccurate and biased information. "Judges do not need to shun the careful use of AI," said Master of the Rolls Geoffrey Vos, the second-highest ranking judge in England and Wales. "But they must ensure that they protect confidence and take full personal responsibility for everything they produce." At a time when scholars and legal experts are pondering a future when AI could replace lawyers, help select jurors or even decide cases, the approach spelled out Dec. 11 by the judiciary is restrained. But for a profession slow to embrace technological change, it's a proactive step as government and industry -- and society in general -- react to a rapidly advancing technology alternately portrayed as a panacea and a menace.
The top tech trends to watch in 2024
Projects like Chat2024, developed by a Miami start-up called Delphi, let you pose questions to chatbots based on presidential candidates, trained on their published statements and transcripts of video appearances. And at least one congressional candidate -- Shamaine Daniels, a Pennsylvania Democrat -- has begun to use an AI robocaller developed by a company called Civox to engage thousands of potential voters in customized conversations.
Vulcan launch: Why is NASA going back to the moon?
NASA's first mission to the moon's surface since the Apollo missions in the 1970s has begun with the launch of a new Vulcan rocket carrying a robotic lander with seven scientific instruments. The mission, which launched from Cape Canaveral in Florida at 7.18am GMT (2.18am EST) on 8 January, forms the first part of NASA's ambitious Commercial Lunar Payload Services (CLPS) programme, with six more launches planned for this year. Unlike previous NASA missions, which were carried out almost entirely in-house, these efforts will be public-private partnerships, aided by space companies. The Vulcan rocket was built by both Lockheed Martin and Boeing as part of the United Launch Alliance (ULA), and the Peregrine lander was built by space robotics company Astrobotic. The lander will take 46 days to reach the moon, before attempting to land on 23 February.
Four lessons from 2023 that tell us where AI regulation is going
Most broadly, we are likely to see the strategies that emerged last year continue, expand, and begin to be implemented. For example, following President Biden's executive order, various US government agencies may outline new best practices but empower AI companies to police themselves. And across the pond, companies and regulators will begin to grapple with Europe's AI Act and its risk-based approach. It certainly won't be seamless, and there's bound to be a lot of discussion about how these new laws and policies actually work in practice. While writing this piece, I took some time to reflect on how we got here.