Africa
An Infinite-Width Analysis on the Jacobian-Regularised Training of a Neural Network
The recent theoretical analysis of deep neural networks in their infinite-width limits has deepened our understanding of initialisation, feature learning, and training of those networks, and brought new practical techniques for finding appropriate hyperparameters, learning network weights, and performing inference. In this paper, we broaden this line of research by showing that this infinite-width analysis can be extended to the Jacobian of a deep neural network. We show that a multilayer perceptron (MLP) and its Jacobian at initialisation jointly converge to a Gaussian process (GP) as the widths of the MLP's hidden layers go to infinity and characterise this GP. We also prove that in the infinite-width limit, the evolution of the MLP under the so-called robust training (i.e., training with a regulariser on the Jacobian) is described by a linear first-order ordinary differential equation that is determined by a variant of the Neural Tangent Kernel. We experimentally show the relevance of our theoretical claims to wide finite networks, and empirically analyse the properties of kernel regression solution to obtain an insight into Jacobian regularisation.
Experimenting with generative AI in the classroom
As artificial intelligence (AI) challenges us to reimagine new ways of doing and being, Dr Marcel O'Gorman, professor of English Language and Literature, embraces emerging technologies and applies them to his pedagogy in the classroom. O'Gorman has published widely about the impacts of technology, and his most recent research focuses on how critical and inclusive design methods might help tackle some of the moral and ethical issues faced by contemporary technoculture. O'Gorman recently wrapped up teaching a fourth-year undergraduate course on techno-critical writing and design that focused on key issues around responsible innovation, such as algorithmic bias, conflict minerals and the colonial practices of big tech on the global stage. Students applied what they learned by writing and designing projects throughout the course. "They wrote stories in ChatGPT that tested the AI for gender bias. They generated images in DALL-E 2 that traced a racist history in the AI's training data," O'Gorman says.
At least 85 civilians, including women and children, dead after 'mistaken' army drone attack
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Emergency response officials said at least 85 people have been confirmed dead after a "mistaken" army drone attack on a religious gathering in northwest Nigeria. The victims were killed Sunday night by drones "targeting terrorists and bandits" in Kaduna state's Tudun Biri village, according to government and security officials. They were observing a Muslim holiday.
Nigerian military drone attack kills 85 civilians in error
A Nigerian military attack that used drones to target rebels instead killed at least 85 civilians gathered for a religious celebration, authorities said Monday. The attack was the latest in recent errant bombings of residents in Nigeria's troubled regions; between February 2014 when a Nigerian military aircraft dropped a bomb on Daglun in Borno state killing 20 civilians and September 2022, there were at least 14 documented incidences of such bombings in residential areas. The attack on Sunday night in Tudun Biri village of Kaduna state's Igabi council area took place as Muslims gathered there to observe the holiday celebrating the birthday of the Prophet Muhammad. Kaduna Governor Uba Sani said civilians were "mistakenly killed and many others were wounded" by a drone "targeting terrorists and bandits". The National Emergency Management Agency said in a statement on Tuesday that "85 dead bodies have so far been buried while search is still ongoing".
RESIN-EDITOR: A Schema-guided Hierarchical Event Graph Visualizer and Editor
Nguyen, Khanh Duy, Zhang, Zixuan, Suchocki, Reece, Li, Sha, Palmer, Martha, Brown, Susan, Han, Jiawei, Ji, Heng
In this paper, we present RESIN-EDITOR, an interactive event graph visualizer and editor designed for analyzing complex events. Our RESIN-EDITOR system allows users to render and freely edit hierarchical event graphs extracted from multimedia and multi-document news clusters with guidance from human-curated event schemas. RESIN-EDITOR's unique features include hierarchical graph visualization, comprehensive source tracing, and interactive user editing, which is more powerful and versatile than existing Information Extraction (IE) visualization tools. In our evaluation of RESIN-EDITOR, we demonstrate ways in which our tool is effective in understanding complex events and enhancing system performance. The source code, a video demonstration, and a live website for RESIN-EDITOR have been made publicly available.
D-Bot: Database Diagnosis System using Large Language Models
Zhou, Xuanhe, Li, Guoliang, Sun, Zhaoyan, Liu, Zhiyuan, Chen, Weize, Wu, Jianming, Liu, Jiesi, Feng, Ruohang, Zeng, Guoyang
Database administrators (DBAs) play an important role in managing, maintaining and optimizing database systems. However, it is hard and tedious for DBAs to manage a large number of databases and give timely response (waiting for hours is intolerable in many online cases). In addition, existing empirical methods only support limited diagnosis scenarios, which are also labor-intensive to update the diagnosis rules for database version updates. Recently large language models (LLMs) have shown great potential in various fields. Thus, we propose D-Bot, an LLM-based database diagnosis system that can automatically acquire knowledge from diagnosis documents, and generate reasonable and well-founded diagnosis report (i.e., identifying the root causes and solutions) within acceptable time (e.g., under 10 minutes compared to hours by a DBA). The techniques in D-Bot include (i) offline knowledge extraction from documents, (ii) automatic prompt generation (e.g., knowledge matching, tool retrieval), (iii) root cause analysis using tree search algorithm, and (iv) collaborative mechanism for complex anomalies with multiple root causes. We verify D-Bot on real benchmarks (including 539 anomalies of six typical applications), and the results show that D-Bot can effectively analyze the root causes of unseen anomalies and significantly outperforms traditional methods and vanilla models like GPT-4.
Delegated Classification
Saig, Eden, Talgam-Cohen, Inbal, Rosenfeld, Nir
When machine learning is outsourced to a rational agent, conflicts of interest might arise and severely impact predictive performance. In this work, we propose a theoretical framework for incentive-aware delegation of machine learning tasks. We model delegation as a principal-agent game, in which accurate learning can be incentivized by the principal using performance-based contracts. Adapting the economic theory of contract design to this setting, we define budget-optimal contracts and prove they take a simple threshold form under reasonable assumptions. In the binary-action case, the optimality of such contracts is shown to be equivalent to the classic Neyman-Pearson lemma, establishing a formal connection between contract design and statistical hypothesis testing. Empirically, we demonstrate that budget-optimal contracts can be constructed using small-scale data, leveraging recent advances in the study of learning curves and scaling laws. Performance and economic outcomes are evaluated using synthetic and real-world classification tasks.
Feature-Learning Networks Are Consistent Across Widths At Realistic Scales
Vyas, Nikhil, Atanasov, Alexander, Bordelon, Blake, Morwani, Depen, Sainathan, Sabarish, Pehlevan, Cengiz
We study the effect of width on the dynamics of feature-learning neural networks across a variety of architectures and datasets. Early in training, wide neural networks trained on online data have not only identical loss curves but also agree in their point-wise test predictions throughout training. For simple tasks such as CIFAR-5m this holds throughout training for networks of realistic widths. We also show that structural properties of the models, including internal representations, preactivation distributions, edge of stability phenomena, and large learning rate effects are consistent across large widths. This motivates the hypothesis that phenomena seen in realistic models can be captured by infinite-width, feature-learning limits. For harder tasks (such as ImageNet and language modeling), and later training times, finite-width deviations grow systematically. Two distinct effects cause these deviations across widths. First, the network output has initialization-dependent variance scaling inversely with width, which can be removed by ensembling networks. We observe, however, that ensembles of narrower networks perform worse than a single wide network. We call this the bias of narrower width. We conclude with a spectral perspective on the origin of this finite-width bias.
Learning Robust Output Control Barrier Functions from Safe Expert Demonstrations
Lindemann, Lars, Robey, Alexander, Jiang, Lejun, Das, Satyajeet, Tu, Stephen, Matni, Nikolai
We assume that a model of the system dynamics and a state estimator are available along with corresponding error bounds, e.g., estimated from data in practice. We first propose robust output control barrier functions (ROCBFs) as a means to guarantee safety, as defined through controlled forward invariance of a safe set. We then formulate an optimization problem to learn ROCBFs from expert demonstrations that exhibit safe system behavior, e.g., data collected from a human operator or an expert controller. When the parametrization of the ROCBF is linear, then we show that, under mild assumptions, the optimization problem is convex. Along with the optimization problem, we provide verifiable conditions in terms of the density of the data, smoothness of the system model and state estimator, and the size of the error bounds that guarantee validity of the obtained ROCBF. Towards obtaining a practical control algorithm, we propose an algorithmic implementation of our theoretical framework that accounts for assumptions made in our framework in practice. We empirically validate our algorithm in the autonomous driving simulator CARLA and demonstrate how to learn safe control laws from RGB camera images.