Government
Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural Networks
Kalan, Seyed Mohammadreza Mousavi, Fabian, Zalan, Avestimehr, A. Salman, Soltanolkotabi, Mahdi
Transfer learning has emerged as a powerful technique for improving the performance of machine learning models on new domains where labeled training data may be scarce. In this approach a model trained for a source task, where plenty of labeled training data is available, is used as a starting point for training a model on a related target task with only few labeled training data. Despite recent empirical success of transfer learning approaches, the benefits and fundamental limits of transfer learning are poorly understood. In this paper we develop a statistical minimax framework to characterize the fundamental limits of transfer learning in the context of regression with linear and one-hidden layer neural network models. Specifically, we derive a lower-bound for the target generalization error achievable by any algorithm as a function of the number of labeled source and target data as well as appropriate notions of similarity between the source and target tasks. Our lower bound provides new insights into the benefits and limitations of transfer learning. We further corroborate our theoretical finding with various experiments.
Understanding and mitigating exploding inverses in invertible neural networks
Behrmann, Jens, Vicol, Paul, Wang, Kuan-Chieh, Grosse, Roger, Jacobsen, Jörn-Henrik
Invertible neural networks (INNs) have been used to design generative models, implement memory-saving gradient computation, and solve inverse problems. In this work, we show that commonly-used INN architectures suffer from exploding inverses and are thus prone to becoming numerically non-invertible. Across a wide range of INN use-cases, we reveal failures including the non-applicability of the change-of-variables formula on in- and out-of-distribution (OOD) data, incorrect gradients for memory-saving backprop, and the inability to sample from normalizing flow models. We further derive bi-Lipschitz properties of atomic building blocks of common architectures. These insights into the stability of INNs then provide ways forward to remedy these failures. For tasks where local invertibility is sufficient, like memory-saving backprop, we propose a flexible and efficient regularizer. For problems where global invertibility is necessary, such as applying normalizing flows on OOD data, we show the importance of designing stable INN building blocks.
A Survey of Constrained Gaussian Process Regression: Approaches and Implementation Challenges
Swiler, Laura, Gulian, Mamikon, Frankel, Ari, Safta, Cosmin, Jakeman, John
Gaussian process regression is a popular Bayesian framework for surrogate modeling of expensive data sources. As part of a broader effort in scientific machine learning, many recent works have incorporated physical constraints or other a priori information within Gaussian process regression to supplement limited data and regularize the behavior of the model. We provide an overview and survey of several classes of Gaussian process constraints, including positivity or bound constraints, monotonicity and convexity constraints, differential equation constraints provided by linear PDEs, and boundary condition constraints. We compare the strategies behind each approach as well as the differences in implementation, concluding with a discussion of the computational challenges introduced by constraints.
The limits of min-max optimization algorithms: convergence to spurious non-critical sets
Hsieh, Ya-Ping, Mertikopoulos, Panayotis, Cevher, Volkan
Compared to minimization problems, the min-max landscape in machine learning applications is considerably more convoluted because of the existence of cycles and similar phenomena. Such oscillatory behaviors are well-understood in the convex-concave regime, and many algorithms are known to overcome them. In this paper, we go beyond the convex-concave setting and we characterize the convergence properties of a wide class of zeroth-, first-, and (scalable) second-order methods in non-convex/non-concave problems. In particular, we show that these state-of-the-art min-max optimization algorithms may converge with arbitrarily high probability to attractors that are in no way min-max optimal or even stationary. Spurious convergence phenomena of this type can arise even in two-dimensional problems, a fact which corroborates the empirical evidence surrounding the formidable difficulty of training GANs.
The Technology 202: Amazon's move to temporarily bar police from using its facial recognition software could have long-term consequences
Law enforcement's use of facial recognition technology was always controversial. Amazon's surprise announcement that it would put a moratorium on police use of its facial recognition software for the next year underscores the big questions surrounding the technology as protests spark a nationwide debate about police brutality and surveillance tactics. Amazon's brief news release never mentioned the words George Floyd, but my Post colleague Jay Greene notes the company hinted that recent events drove this decision. "We've advocated that governments should put in place stronger regulations to govern the ethical use of facial recognition technology, and in recent days, Congress appears ready to take on this challenge," the company said in a statement. "We hope this one-year moratorium might give Congress enough time to implement appropriate rules, and we stand ready to help if requested."
Protecting privacy in an AI-driven world
Our world is undergoing an information Big Bang, in which the universe of data doubles every two years and quintillions of bytes of data are generated every day.1 For decades, Moore's Law on the doubling of computing power every 18-24 months has driven the growth of information technology. Now–as billions of smartphones and other devices collect and transmit data over high-speed global networks, store data in ever-larger data centers, and analyze it using increasingly powerful and sophisticated software–Metcalfe's Law comes into play. It treats the value of networks as a function of the square of the number of nodes, meaning that network effects exponentially compound this historical growth in information. As 5G networks and eventually quantum computing deploy, this data explosion will grow even faster and bigger. The impact of big data is commonly described in terms of three "Vs": volume, variety, and velocity.2
NASA's Mars rover drivers need your help
You may be able to help NASA's Curiosity rover drivers better navigate Mars. Using the online tool AI4Mars to label terrain features in pictures downloaded from the Red Planet, you can train an artificial intelligence algorithm to automatically read the landscape. Is that a big rock to the left? AI4Mars, which is hosted on the citizen science website Zooniverse, lets you draw boundaries around terrain and choose one of four labels. Those labels are key to sharpening the Martian terrain-classification algorithm called SPOC (Soil Property and Object Classification).
Dowden: AI and data science conversion path open and diverse
The government and the Office for Students have announced that 2,500 places on artificial intelligence (AI) and data science conversion courses are now open to applicants. Some 1,000 scholarships will be open to students from "under-represented backgrounds", according to a statement from the Department for Digital, Culture, Media and Sport (DCMS). Oliver Dowden, DCMS secretary, said: "It is vital we increase diversity across our tech sector and give everyone with the aptitude and talent the opportunity to build a successful career. "This will help make sure artificial intelligence developed in the UK reflects the needs and make-up of society as a whole, which will also help mitigate the risk of biased technologies being developed." Funding has been allocated to 18 English universities, which, according to a DCMS statement, will deliver courses to a further 10 universities. In the venture, the government is working with the Office for Students, an independent regulator that reports to the Department for Education and was established in 2017. Together, they have created a fund of up to £24m to support the scholarships. These are open to non-STEM (science, technology, engineering and maths) graduates, as well as those with degrees in STEM subjects. A year ago, in June 2019, the department announced a similar tranche of £13.5m funding for up to 2,500 AI and data science conversion courses for professionals who have degrees in other disciplines, as well as 1,000 scholarships. A DCMS spokesperson confirmed: "People can now apply for places on the courses.