Goto

Collaborating Authors

 Government


Robin Hood and Matthew Effects -- Differential Privacy Has Disparate Impact on Synthetic Data

arXiv.org Artificial Intelligence

Generative models trained using Differential Privacy (DP) are increasingly used to produce and share synthetic data in a privacy-friendly manner. In this paper, we set out to analyze the impact of DP on these models vis-a-vis underrepresented classes and subgroups of data. We do so from two angles: 1) the size of classes and subgroups in the synthetic data, and 2) classification accuracy on them. We also evaluate the effect of various levels of imbalance and privacy budgets. Our experiments, conducted using three state-of-the-art DP models (PrivBayes, DP-WGAN, and PATE-GAN), show that DP results in opposite size distributions in the generated synthetic data. More precisely, it affects the gap between the majority and minority classes and subgroups, either reducing it (a "Robin Hood" effect) or increasing it ("Matthew" effect). However, both of these size shifts lead to similar disparate impacts on a classifier's accuracy, affecting disproportionately more the underrepresented subparts of the data. As a result, we call for caution when analyzing or training a model on synthetic data, or risk treating different subpopulations unevenly, which might also lead to unreliable conclusions.


Zero-Shot Information Extraction as a Unified Text-to-Triple Translation

arXiv.org Artificial Intelligence

We cast a suite of information extraction tasks into a text-to-triple translation framework. Instead of solving each task relying on task-specific datasets and models, we formalize the task as a translation between task-specific input text and output triples. By taking the task-specific input, we enable a task-agnostic translation by leveraging the latent knowledge that a pre-trained language model has about the task. We further demonstrate that a simple pre-training task of predicting which relational information corresponds to which input text is an effective way to produce task-specific outputs. This enables the zero-shot transfer of our framework to downstream tasks. We study the zero-shot performance of this framework on open information extraction (OIE2016, NYT, WEB, PENN), relation classification (FewRel and TACRED), and factual probe (Google-RE and T-REx). The model transfers non-trivially to most tasks and is often competitive with a fully supervised method without the need for any task-specific training. For instance, we significantly outperform the F1 score of the supervised open information extraction without needing to use its training set.


Inequality Constrained Stochastic Nonlinear Optimization via Active-Set Sequential Quadratic Programming

arXiv.org Machine Learning

We study nonlinear optimization problems with stochastic objective and deterministic equality and inequality constraints, which emerge in numerous applications including finance, manufacturing, power systems and, recently, deep neural networks. We propose an active-set stochastic sequential quadratic programming algorithm, using a differentiable exact augmented Lagrangian as the merit function. The algorithm adaptively selects the penalty parameters of augmented Lagrangian and performs stochastic line search to decide the stepsize. The global convergence is established: for any initialization, the "liminf" of the KKT residuals converges to zero almost surely. Our algorithm and analysis further develop the prior work \cite{Na2021Adaptive} by allowing nonlinear inequality constraints. We demonstrate the performance of the algorithm on a subset of nonlinear problems collected in the CUTEst test set.


WRENCH: A Comprehensive Benchmark for Weak Supervision

arXiv.org Machine Learning

Recent \emph{Weak Supervision (WS)} approaches have had widespread success in easing the bottleneck of labeling training data for machine learning by synthesizing labels from multiple potentially noisy supervision sources. However, proper measurement and analysis of these approaches remain a challenge. First, datasets used in existing works are often private and/or custom, limiting standardization. Second, WS datasets with the same name and base data often vary in terms of the labels and weak supervision sources used, a significant "hidden" source of evaluation variance. Finally, WS studies often diverge in terms of the evaluation protocol and ablations used. To address these problems, we introduce a benchmark platform, \benchmark, for a thorough and standardized evaluation of WS approaches. It consists of 22 varied real-world datasets for classification and sequence tagging; a range of real, synthetic, and procedurally-generated weak supervision sources; and a modular, extensible framework for WS evaluation, including implementations for popular WS methods. We use \benchmark to conduct extensive comparisons over more than 100 method variants to demonstrate its efficacy as a benchmark platform. The code is available at \url{https://github.com/JieyuZ2/wrench}.


British court disagrees with Australia, rules that AIs cannot be patent inventors

#artificialintelligence

The UK and Australia may have made a historic pact last week, but one thing they can't agree on is whether AIs can be patent inventors. AIs are increasingly being used to come up with new ideas and there's an argument they should therefore be listed as the inventor by patent agencies. However, opponents say that patents are a statutory right and can only be granted to a person. US-based Dr Stephen Thaler, the founder of Imagination Engines, has been leading the fight to give credit to machines for their creations. Dr Thaler's AI device, DABUS, consists of neural networks and was used to invent an emergency warning light, a food container that improves grip and heat transfer, and more.


Paige Receives First Ever FDA Approval for AI Product in Digital Pathology

#artificialintelligence

NEW YORK--(BUSINESS WIRE)--Paige, the global leader in AI-based diagnostic software in pathology, today announced that the U.S. Food and Drug Administration (FDA) has granted de novo marketing authorization for Paige Prostate, a clinical-grade AI solution for prostate cancer detection. As a novel technology, Paige Prostate is the first AI-based pathology product to receive de novo approval from the FDA, allowing in vitro diagnostic (IVD) use via Paige's FDA-cleared FullFocus digital pathology viewer. With a projected 60 percent increase in the number of cancer cases globally in the next two decades and a decrease in the number of pathologists relative to this diagnostic demand, there is a significant need to provide new technologies for the practice of pathology. Paige Prostate is a cancer detection solution that identifies foci suspicious for cancer and provides this information to the pathologist. Paige Prostate is designed to assist pathologists in finding small foci of cancer and enable pathologists to work efficiently and confidently in their diagnostic process.


Greece used AI to curb COVID: what other nations can learn

#artificialintelligence

Greece's decision to deploy machine learning in pandemic surveillance will be much-studied around the world.Credit: Konstantinos Tsakalidis/Bloomberg/Getty A few months into the COVID-19 pandemic, operations researcher Kimon Drakopoulos e-mailed both the Greek prime minister and the head of the country's COVID-19 scientific task force to ask if they needed any extra advice. Drakopoulos works in data science at the University of Southern California in Los Angeles, and is originally from Greece. To his surprise, he received a reply from Prime Minister Kyriakos Mitsotakis within hours. The European Union was asking member states, many of which had implemented widespread lockdowns in March, to allow non-essential travel to recommence from July 2020, and the Greek government needed help in deciding when and how to reopen borders. Greece, like many other countries, lacked the capacity to test all travellers, particularly those not displaying symptoms.


AI cannot be regulated by technical measures alone

#artificialintelligence

Any attempt to regulate artificial intelligence (AI) must not rely solely on technical measures to mitigate potential harms, and should instead move to address the fundamental power imbalances between those who develop or deploy the technology and those who are subject to it, says a report commissioned by European Digital Rights (EDRi). Published on 21 September 2021, the 155-page report Beyond debiasing: regulating AI and its inequalities specifically criticised the European Union's (EU) "technocratic" approach to AI regulation, which it said was too narrowly focused on implementing technical bias mitigation measures, otherwise known as "debiasing", to be effective at preventing the full range of AI-related harms. The European Commission's (EC) proposed Artificial Intelligence Act (AIA) was published in April 2021 and sought to create a risk-based, market-led approach to regulating AI through the establishment of self-assessments, transparency procedures and various technical standards. Digital civil rights experts and organisations have previously told Computer Weekly that although the regulation is a step in the right direction, it will ultimately fail to protect people's fundamental rights and mitigate the technology's worst abuses because it does not address the fundamental power imbalances between tech firms and those who are subject to their systems. The EDRi-commissioned report said that while European policymakers have publicly recognised that AI can produce a broad range of harms across different domains โ€“ including employment, housing, education, health and policing โ€“ their laser focus on algorithmic debiasing stems from a misunderstanding of the existing techniques and their effectiveness.


Royal Air Force is testing self-driving cars in Oxfordshire

Daily Mail - Science & tech

The Royal Air Force is testing its own autonomous vehicle to deliver supplies around a base in Oxfordshire to'free up personnel from mundane tasks'. Its specially-designed self-driving car, called Kar-Go, is a zero-emissions delivery vehicle capable of travelling at speeds of up to 60 miles/hour. It's been zipping around the Royal Air Force base of Brize Norton in Oxfordshire, delivering tools, equipment and supplies to personnel as part of a trial. When arriving at its destination on the base, RAF personnel meet Kar-Go and a hatch is automatically released enabling them to collect the cargo. Slightly odd in appearance, Kar-Go looks a bit like a gigantic green computer mouse with protruding wheels, complete with flashing lights and a spacious boot.


UK launches National AI Strategy

AIHub

On 22 September 2021, the UK government released its National AI Strategy. According to the document, the new strategy "represents the start of a step-change for AI in the UK, recognising the power of AI to increase resilience, productivity, growth and innovation across the private and public sectors. This is how we will prepare the UK for the next ten years."