Government
Bias Mitigation for Machine Learning Classifiers: A Comprehensive Survey
Hort, Max, Chen, Zhenpeng, Zhang, Jie M., Harman, Mark, Sarro, Federica
This paper provides a comprehensive survey of bias mitigation methods for achieving fairness in Machine Learning (ML) models. We collect a total of 341 publications concerning bias mitigation for ML classifiers. These methods can be distinguished based on their intervention procedure (i.e., pre-processing, in-processing, post-processing) and the technique they apply. We investigate how existing bias mitigation methods are evaluated in the literature. In particular, we consider datasets, metrics and benchmarking. Based on the gathered insights (e.g., What is the most popular fairness metric? How many datasets are used for evaluating bias mitigation methods?), we hope to support practitioners in making informed choices when developing and evaluating new bias mitigation methods.
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
Moniri, Behrad, Lee, Donghwan, Hassani, Hamed, Dobriban, Edgar
Feature learning is thought to be one of the fundamental reasons for the success of deep neural networks. It is rigorously known that in two-layer fully-connected neural networks under certain conditions, one step of gradient descent on the first layer followed by ridge regression on the second layer can lead to feature learning; characterized by the appearance of a separated rank-one component -- spike -- in the spectrum of the feature matrix. However, with a constant gradient descent step size, this spike only carries information from the linear component of the target function and therefore learning non-linear components is impossible. We show that with a learning rate that grows with the sample size, such training in fact introduces multiple rank-one components, each corresponding to a specific polynomial feature. We further prove that the limiting large-dimensional and large sample training and test errors of the updated neural networks are fully characterized by these spikes. By precisely analyzing the improvement in the loss, we demonstrate that these non-linear features can enhance learning.
Feature Learning and Generalization in Deep Networks with Orthogonal Weights
Day, Hannah, Kahn, Yonatan, Roberts, Daniel A.
Fully-connected deep neural networks with weights initialized from independent Gaussian distributions can be tuned to criticality, which prevents the exponential growth or decay of signals propagating through the network. However, such networks still exhibit fluctuations that grow linearly with the depth of the network, which may impair the training of networks with width comparable to depth. We show analytically that rectangular networks with tanh activations and weights initialized from the ensemble of orthogonal matrices have corresponding preactivation fluctuations which are independent of depth, to leading order in inverse width. Moreover, we demonstrate numerically that, at initialization, all correlators involving the neural tangent kernel (NTK) and its descendants at leading order in inverse width -- which govern the evolution of observables during training -- saturate at a depth of $\sim 20$, rather than growing without bound as in the case of Gaussian initializations. We speculate that this structure preserves finite-width feature learning while reducing overall noise, thus improving both generalization and training speed. We provide some experimental justification by relating empirical measurements of the NTK to the superior performance of deep nonlinear orthogonal networks trained under full-batch gradient descent on the MNIST and CIFAR-10 classification tasks.
Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes
Learning the distribution of data on Riemannian manifolds is crucial for modeling data from non-Euclidean space, which is required by many applications from diverse scientific fields. Yet, existing generative models on manifolds suffer from expensive divergence computation or rely on approximations of heat kernel. These limitations restrict their applicability to simple geometries and hinder scalability to high dimensions. In this work, we introduce the Riemannian Diffusion Mixture, a principled framework for building a generative process on manifolds as a mixture of endpoint-conditioned diffusion processes instead of relying on the denoising approach of previous diffusion models, for which the generative process is characterized by its drift guiding toward the most probable endpoint with respect to the geometry of the manifold. We further propose a simple yet efficient training objective for learning the mixture process, that is readily applicable to general manifolds. Our method outperforms previous generative models on various manifolds while scaling to high dimensions and requires a dramatically reduced number of in-training simulation steps for general manifolds. Deep generative models have shown great success in learning the distributions of the data represented in Euclidean space, e.g., images and text. While the focus of the previous works has been biased toward data in the Euclidean space, modeling the distribution of the data that naturally resides in the Riemannian manifold with specific geometry has been underexplored, while they are required for wide application: For example, the earth and climate science data (Karpatne et al., 2018; Mathieu & Nickel, 2020) lives in the sphere, whereas the protein structures (Jumper et al., 2021; Watson et al., 2022) and the robotic movements (Simeonov et al., 2022) are best represented by the group SE(3). Moreover, 3D computer graphics shapes (Hoppe et al., 1992) can be identified as a general closed manifold. However, previous generative methods are ill-suited for modeling these data as they do not take into consideration the specific geometry describing the data space and may assign a non-zero probability to regions outside the desired space.
DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecasting
Cachay, Salva Rรผhling, Zhao, Bo, Joren, Hailey, Yu, Rose
While diffusion models can successfully generate data and make predictions, they are predominantly designed for static images. We propose an approach for efficiently training diffusion models for probabilistic spatiotemporal forecasting, where generating stable and accurate rollout forecasts remains challenging, Our method, DYffusion, leverages the temporal dynamics in the data, directly coupling it with the diffusion steps in the model. We train a stochastic, time-conditioned interpolator and a forecaster network that mimic the forward and reverse processes of standard diffusion models, respectively. DYffusion naturally facilitates multi-step and long-range forecasting, allowing for highly flexible, continuous-time sampling trajectories and the ability to trade-off performance with accelerated sampling at inference time. In addition, the dynamics-informed diffusion process in DYffusion imposes a strong inductive bias and significantly improves computational efficiency compared to traditional Gaussian noise-based diffusion models. Our approach performs competitively on probabilistic forecasting of complex dynamics in sea surface temperatures, Navier-Stokes flows, and spring mesh systems.
A Doctored Biden Video Is a Test Case for Facebook's Deepfake Policies
During the 2022 US midterm elections, a manipulated video of President Joe Biden circulated on Facebook. The original footage showed Biden placing an "I voted" sticker on his granddaughter's chest and kissing her on the cheek. The doctored version looped the footage to make it appear he was repeatedly touching the girl, with a caption that labeled him a "pedophile." Meta left the video up. Today, the company's Oversight Board--an independent body that looks into the platform's content moderation--announced that it will review that decision, in an attempt to push Meta to address how it will handle manipulated media and election disinformation ahead of the 2024 US presidential election and more than 50 other votes to be held around the world next year.
Are we ready to trust AI with our bodies?
That's why I was intrigued when I read my colleague Rhiannon Williams' latest piece about AI gym trainers. Lumin Fitness is a gym in Texas staffed pretty much entirely by virtual AI coaches designed to guide gym goers through workouts (there's one human employee on hand--to switch everything off and on, perhaps.) Patrons can complete a solo workout program with the help of a virtual coach in their own designated station, or participate in a high-intensity functional training class with others. Sensors in both the equipment and the floor-to-ceiling LED screens that line the walls of the gym track users' movements, and Lumin uses machine learning models to tailor advice. The gym owners are confident that these new AI trainers will encourage people like me who feel intimidated or unmotivated to work out. Over the next few years, artificial intelligence is going to have a bigger and bigger effect on us and the way we live.
Russia launches dozens of drones into Ukraine in latest air raid: Kyiv
Russia launched 36 drone attacks overnight on Ukraine, according to Kyiv's air force, in Moscow's latest air raid targeting the country. Ukraine's air force said in a statement on Tuesday that its defence systems had destroyed 27 of the drones. The attacks using Iran-made Shahed drones targeted the Odesa, Mykolaiv and Kherson regions of Ukraine, the air force said on the Telegram messaging app. Moscow launched a total of 36 Iranian-made drones from the Russia-annexed Crimean peninsula, it added. The air force did not say which targets, if any, the nine other drones may have hit.
America's secret asset against AI workforce takeover
Kara Frederick, tech director at the Heritage Foundation, discusses the need for regulations on artificial intelligence as lawmakers and tech titans discuss the potential risks. Two significant shifts are changing America's workforce as we've known it. First, artificial intelligence (AI) continues to transform everything about work. AI technologies-related job displacement presents a major challenge to the American worker and it continues to disrupt our economy. Equally disruptive is our rapidly aging workforce.
Downing Street trying to agree statement about AI risks with world leaders
Rishi Sunak's advisers are trying to thrash out an agreement among world leaders on a statement warning about the risks of artificial intelligence as they finalise the agenda for the AI safety summit next month. Downing Street officials have been touring the world talking to their counterparts from China to the EU and the US as they work to agree on words to be used in a communique at the two-day conference. But they are unlikely to agree a new international organisation to scrutinise cutting-edge AI, despite interest from the UK in giving the government's AI taskforce a global role. Sunak's AI summit will produce a communique on the risks of AI models, provide an update on White House-brokered safety guidelines and end with "like-minded" countries debating how national security agencies can scrutinise the most dangerous versions of the technology. The possibility of some form of international cooperation on cutting-edge AI that can pose a threat to human life will also be discussed on the final day of the summit on 1 and 2 November at Bletchley Park, according to a draft agenda seen by the Guardian.