Oceania
STANLEY: Stochastic Gradient Anisotropic Langevin Dynamics for Learning Energy-Based Models
Karimi, Belhal, Xie, Jianwen, Li, Ping
We propose in this paper, STANLEY, a STochastic gradient ANisotropic LangEvin dYnamics, for sampling high dimensional data. With the growing efficacy and potential of Energy-Based modeling, also known as non-normalized probabilistic modeling, for modeling a generative process of different natures of high dimensional data observations, we present an end-to-end learning algorithm for Energy-Based models (EBM) with the purpose of improving the quality of the resulting sampled data points. While the unknown normalizing constant of EBMs makes the training procedure intractable, resorting to Markov Chain Monte Carlo (MCMC) is in general a viable option. Realizing what MCMC entails for the EBM training, we propose in this paper, a novel high dimensional sampling method, based on an anisotropic stepsize and a gradient-informed covariance matrix, embedded into a discretized Langevin diffusion. We motivate the necessity for an anisotropic update of the negative samples in the Markov Chain by the nonlinearity of the backbone of the EBM, here a Convolutional Neural Network. Our resulting method, namely STANLEY, is an optimization algorithm for training Energy-Based models via our newly introduced MCMC method. We provide a theoretical understanding of our sampling scheme by proving that the sampler leads to a geometrically uniformly ergodic Markov Chain. Several image generation experiments are provided in our paper to show the effectiveness of our method.
Piecewise Deterministic Markov Processes for Bayesian Neural Networks
Goan, Ethan, Perrin, Dimitri, Mengersen, Kerrie, Fookes, Clinton
Inference on modern Bayesian Neural Networks (BNNs) often relies on a variational inference treatment, imposing violated assumptions of independence and the form of the posterior. Traditional MCMC approaches avoid these assumptions at the cost of increased computation due to its incompatibility to subsampling of the likelihood. New Piecewise Deterministic Markov Process (PDMP) samplers permit subsampling, though introduce a model specific inhomogenous Poisson Process (IPPs) which is difficult to sample from. This work introduces a new generic and adaptive thinning scheme for sampling from these IPPs, and demonstrates how this approach can accelerate the application of PDMPs for inference in BNNs. Experimentation illustrates how inference with these methods is computationally feasible, can improve predictive accuracy, MCMC mixing performance, and provide informative uncertainty measurements when compared against other approximate inference schemes.
Not US 'job' to keep allies on 'cutting edge' of AI development, former CIA chief says
Former CIA director Gen. David Petraeus spoke with Fox News Digital about the "breathtaking" pace of AI development and what responsibility the US has to keep allies on the "cutting edge." The U.S. has no responsibility to keep its allies on the "cutting edge" of artificial intelligence (AI) development – unless it is a matter of national security, a former CIA director and retired Army officer tells Fox News Digital. "First and foremost is to ensure that we are on the cutting edge," retired four-star Gen. David Petraeus said during a recent Zoom interview. "Certainly, we should share to varying degrees with our closest partners, and they share with us – again, no one of us is smarter than all of us together in these kinds of endeavors – but it's not our job entirely to ensure that they are proceeding along these lines, unless it's, of course, in our national interest to do so, and it is in a number of different cases," he explained. The pace of AI development has dominated conversation since public access to ChatGPT in November 2022, particularly with concerns over who will stay at the top of the game – a race that drove countries to reassess their investments in the burgeoning field.
Five Eyes intelligence chiefs warn on China's 'theft' of intellectual property
The Five Eyes countries' intelligence chiefs came together on Tuesday to accuse China of intellectual property theft and using artificial intelligence for hacking and spying against the nations, in a rare joint statement by the allies. Officials from the United States, Britain, Canada, Australia and New Zealand -- known as the Five Eyes intelligence sharing network -- made the comments following meetings with private companies in the U.S. innovation hub Silicon Valley. U.S. FBI Director Christopher Wray said the "unprecedented" joint call was meant to confront the "unprecedented threat" China poses to innovation across the world.
Constrained Reweighting of Distributions: an Optimal Transport Approach
Chakraborty, Abhisek, Bhattacharya, Anirban, Pati, Debdeep
We commonly encounter the problem of identifying an optimally weight adjusted version of the empirical distribution of observed data, adhering to predefined constraints on the weights. Such constraints often manifest as restrictions on the moments, tail behaviour, shapes, number of modes, etc., of the resulting weight adjusted empirical distribution. In this article, we substantially enhance the flexibility of such methodology by introducing a nonparametrically imbued distributional constraints on the weights, and developing a general framework leveraging the maximum entropy principle and tools from optimal transport. The key idea is to ensure that the maximum entropy weight adjusted empirical distribution of the observed data is close to a pre-specified probability distribution in terms of the optimal transport metric while allowing for subtle departures. The versatility of the framework is demonstrated in the context of three disparate applications where data re-weighting is warranted to satisfy side constraints on the optimization problem at the heart of the statistical task: namely, portfolio allocation, semi-parametric inference for complex surveys, and ensuring algorithmic fairness in machine learning algorithms.
A Surrogate-Assisted Extended Generative Adversarial Network for Parameter Optimization in Free-Form Metasurface Design
Dai, Manna, Jiang, Yang, Yang, Feng, Chattoraj, Joyjit, Xia, Yingzhi, Xu, Xinxing, Zhao, Weijiang, Dao, My Ha, Liu, Yong
Metasurfaces have widespread applications in fifth-generation (5G) microwave communication. Among the metasurface family, free-form metasurfaces excel in achieving intricate spectral responses compared to regular-shape counterparts. However, conventional numerical methods for free-form metasurfaces are time-consuming and demand specialized expertise. Alternatively, recent studies demonstrate that deep learning has great potential to accelerate and refine metasurface designs. Here, we present XGAN, an extended generative adversarial network (GAN) with a surrogate for high-quality free-form metasurface designs. The proposed surrogate provides a physical constraint to XGAN so that XGAN can accurately generate metasurfaces monolithically from input spectral responses. In comparative experiments involving 20000 free-form metasurface designs, XGAN achieves 0.9734 average accuracy and is 500 times faster than the conventional methodology. This method facilitates the metasurface library building for specific spectral responses and can be extended to various inverse design problems, including optical metamaterials, nanophotonic devices, and drug discovery.
MWE as WSD: Solving Multiword Expression Identification with Word Sense Disambiguation
Tanner, Joshua, Hoffman, Jacob
Recent approaches to word sense disambiguation (WSD) utilize encodings of the sense gloss (definition), in addition to the input context, to improve performance. In this work we demonstrate that this approach can be adapted for use in multiword expression (MWE) identification by training models which use gloss and context information to filter MWE candidates produced by a rule-based extraction pipeline. Our approach substantially improves precision, outperforming the state-of-the-art in MWE identification on the DiMSUM dataset by up to 1.9 F1 points and achieving competitive results on the PARSEME 1.1 English dataset. Our models also retain most of their WSD performance, showing that a single model can be used for both tasks. Finally, building on similar approaches using Bi-encoders for WSD, we introduce a novel Poly-encoder architecture which improves MWE identification performance.
A Survey of Requirements for COVID-19 Mitigation Strategies. Part I: Newspaper Clips
Jamroga, Wojciech, Mestel, David, Roenne, Peter B., Ryan, Peter Y. A., Skrobot, Marjan
The COVID-19 pandemic has influenced virtually all aspects of our lives. Across the world, countries have applied various mitigation strategies for the epidemic, based on social, political, and technological instruments. We postulate that one should {identify the relevant requirements} before committing to a particular mitigation strategy. One way to achieve it is through an overview of what is considered relevant by the general public, and referred to in the media. To this end, we have collected a number of news clips that mention the possible goals and requirements for a mitigation strategy. The snippets are sorted thematically into several categories, such as health-related goals, social and political impact, civil rights, ethical requirements, and so on. In a forthcoming companion paper, we will present a digest of the requirements, derived from the news clips, and a preliminary take on their formal specification.
Strategic Abilities of Asynchronous Agents: Semantic Side Effects and How to Tame Them
Jamroga, Wojciech, Penczek, Wojciech, Sidoruk, Teofil
Recently, we have proposed a framework for verification of agents' abilities in asynchronous multi-agent systems, together with an algorithm for automated reduction of models. The semantics was built on the modeling tradition of distributed systems. As we show here, this can sometimes lead to counterintuitive interpretation of formulas when reasoning about the outcome of strategies. First, the semantics disregards finite paths, and thus yields unnatural evaluation of strategies with deadlocks. Secondly, the semantic representations do not allow to capture the asymmetry between proactive agents and the recipients of their choices. We propose how to avoid the problems by a suitable extension of the representations and change of the execution semantics for asynchronous MAS. We also prove that the model reduction scheme still works in the modified framework.
Facebook whistleblower Frances Haugen issues chilling warning about AI and says it could soon have 'civilisation-altering impacts'
Advances in artificial intelligence could have'civilisation-altering impacts' and rapidly increase the amount of dangerous misinformation being spread online, a former Facebook employee has warned. Whistleblower Frances Haugen said as AI became bigger and economies relied more on software running on data centres the world would start to see an'era of opacity' creep in. The former engineer and product manager - who quit Facebook in 2021 after leaking thousands of documents showing toxic content was being spread knowingly by the platform - said without stronger regulation there would be'a repeat of what we saw with social media' on a far greater scale. 'When we start getting into scalable systems that run on data centres, a very small number of people can have civilisation-impacting levels of power,' Ms Haugen told the National Press Club on Tuesday. 'At Facebook, there's a very small number of people who really understand how these algorithms work and yet it impacts what everyone sees in the news.