Asia
The Monitoring Game: China's Artificial Intelligence Push
It's all keen and mean on the artificial intelligence (AI) front in China, which is now vying with the United States as the top dog in the field. US companies can still boast the big cheese operators, but China is making strides in other areas. The UN World Intellectual Property Organisation's Thursday report found that IBM had, with 8,920 patents in the field, the largest AI portfolio, followed by Microsoft with 5,930. China, however, was found dominant in 17 of 20 academic institutions involved in the business of patenting AI. The scramble has been a bitter one. The Trump administration has been inflicting various punitive measures through tariffs, accusing Beijing of being the lead thief in global intellectual property matters.
7 amazing robots based on animals
When it comes to robots, science fiction has conditioned us to think of androids – bipedal machines approximating the human form. But the next generation of robots may be based on very different types of animals: snakes, flies, locusts and even the multi-tentacle octopus. Israeli scientists are hard at work on just such contraptions. Here's a look at seven of the most fascinating designs that can help with everything from exploring our insides to cleaning up the mess we make on the planet. Medrobotics' signature product, the Flex Robotic System, allows physicians to reach deep into the body with minimal risk.
Artificial Intelligence: A New Reality for Chemical Engineers - Chemical Engineering Page 1
As in many other sectors, artificial intelligence (AI) technologies are beginning to emerge in the chemical process industries (CPI). While AI-assisted solutions, and other associated technologies, such as robotic process automation (RPA), Internet of Things (IoT), automated drones and quantum computing, are still relatively new for many CPI applications, developers and users alike are realizing their potential benefits for expediting research and development (R&D), predictive maintenance, process optimization and more. Within its Smart Operations initiative, Henkel AG & Co. KGaA (Düsseldorf, Germany; www.henkel.com) is utilizing AI capabilities in its global process operations and supply chain. "We use AI to run efficient analyses of complex data arrays for achieving higher production performance, quick product innovation and scaleup for our self-adjusting production systems," explains Sandeep Sreekumar, global head of Adhesive Digital Operations at Henkel. "Our focus is not only on collecting internal manufacturing data, but also on actively working with customers on data collection opportunities during product usage to make improvements and adjust to changing customer needs," says Sreekumar.
Recognizing the right tool for the job
A*STAR researchers working with colleagues in Japan have developed a method by which robots can automatically recognize an object as a potential tool and use it, despite never having seen it before. For humans, the ability to recognize and use tools is almost instinctive. There are also many examples in which tool use seems hardwired into the brain of animals: some birds and primates use sticks or stones to obtain food, for example. One proposed reason for this neurologically embedded ability to use tools is that the animal's brain perceives the external object as an extension of its own body. Inspired by this idea, Keng Peng Tee and his colleagues from the A*STAR Institute for Infocomm Research, along with Gowrishankar Ganesh from the CNRS-AIST Joint Robotics Laboratory located in Tsukuba, Japan developed an algorithm that enables robots to recognize, and immediately use tools that they have never seen before.
Extracting Multiple-Relations in One-Pass with Pre-Trained Transformers
Wang, Haoyu, Tan, Ming, Yu, Mo, Chang, Shiyu, Wang, Dakuo, Xu, Kun, Guo, Xiaoxiao, Potdar, Saloni
Most approaches to extraction multiple relations from a paragraph require multiple passes over the paragraph. In practice, multiple passes are computationally expensive and this makes difficult to scale to longer paragraphs and larger text corpora. In this work, we focus on the task of multiple relation extraction by encoding the paragraph only once (one-pass). We build our solution on the pre-trained self-attentive (Transformer) models, where we first add a structured prediction layer to handle extraction between multiple entity pairs, then enhance the paragraph embedding to capture multiple relational information associated with each entity with an entity-aware attention technique. We show that our approach is not only scalable but can also perform state-of-the-art on the standard benchmark ACE 2005.
Global Fitting of the Response Surface via Estimating Multiple Contours of a Simulator
Yang, Feng, Lin, C. Devon, Ranjan, Pritam
Computer simulators are nowadays widely used to understand complex physical systems in many areas such as aerospace, renewable energy, climate modeling, and manufacturing. One fundamental issue in the study of computer simulators is known as experimental design, that is, how to select the input settings where the computer simulator is run and the corresponding response is collected. Extra care should be taken in the selection process because computer simulators can be computationally expensive to run. The selection shall acknowledge and achieve the goal of the analysis. This article focuses on the goal of producing more accurate prediction which is important for risk assessment and decision making. We propose two new methods of design approaches that sequentially select input settings to achieve this goal. The approaches make novel applications of simultaneous and sequential contour estimations. Numerical examples are employed to demonstrate the effectiveness of the proposed approaches.
Study of Robust Distributed Diffusion RLS Algorithms with Side Information for Adaptive Networks
Yu, Y., Zhao, H., de Lamare, R. C., Zakharov, Y., Lu, L.
This work develops robust diffusion recursive least squares algorithms to mitigate the performance degradation often experienced in networks of agents in the presence of impulsive noise. The first algorithm minimizes an exponentially weighted least-squares cost function subject to a time-dependent constraint on the squared norm of the intermediate update at each node. A recursive strategy for computing the constraint is proposed using side information from the neighboring nodes to further improve the robustness. We also analyze the mean-square convergence behavior of the proposed algorithm. The second proposed algorithm is a modification of the first one based on the dichotomous coordinate descent iterations. It has a performance similar to that of the former, however its complexity is significantly lower especially when input regressors of agents have a shift structure and it is well suited to practical implementation. Simulations show the superiority of the proposed algorithms over previously reported techniques in various impulsive noise scenarios.
Is There an Analog of Nesterov Acceleration for MCMC?
Ma, Yi-An, Chatterji, Niladri, Cheng, Xiang, Flammarion, Nicolas, Bartlett, Peter, Jordan, Michael I.
While optimization methodology has provided much of the underlying algorithmic machinery that has driven the theory and practice of machine learning in recent years, sampling-based methodology, in particular Markov chain Monte Carlo (MCMC), remains of critical importance, given its role in linking algorithms to statistical inference and, in particular, its ability to provide notions of confidence that are lacking in optimization-based methodology. However, the classical theory of MCMC is largely asymptotic and the theory has not developed as rapidly in recent years as the theory of optimization. Recently, however, a literature has emerged that derives nonasymptotic rates for MCMC algorithms [see, e.g., 9, 12, 10, 8, 6, 14, 21, 22, 2, 5]. This work has explicitly aimed at making use of ideas from optimization; in particular, whereas the classical literature on MCMC focused on reversible Markov chains, the recent literature has focused on nonreversible stochastic processes that are built on gradients [see, e.g., 18, 20, 3, 1]. In particular, the gradient-based Langevin algorithm [33, 32, 13] has been shown to be a form of gradient descent on the space of probabilities [see, e.g., 36]. What has not yet emerged is an analog of acceleration. Recall that the notion of acceleration has played a key role in gradient-based optimization methods [26]. In particular, the Nesterov accelerated gradient descent (AGD) method, an instance of the general family of "momentum methods," provably achieves faster convergence rate than gradient descent (GD) in a variety of settings [25]. Moreover, it achieves the optimal convergence rate under an oracle model of optimization complexity in the convex setting [24].
Minimax experimental design: Bridging the gap between statistical and worst-case approaches to least squares regression
Dereziński, Michał, Clarkson, Kenneth L., Mahoney, Michael W., Warmuth, Manfred K.
In experimental design, we are given a large collection of vectors, each with a hidden response value that we assume derives from an underlying linear model, and we wish to pick a small subset of the vectors such that querying the corresponding responses will lead to a good estimator of the model. A classical approach in statistics is to assume the responses are linear, plus zero-mean i.i.d. Gaussian noise, in which case the goal is to provide an unbiased estimator with smallest mean squared error (A-optimal design). A related approach, more common in computer science, is to assume the responses are arbitrary but fixed, in which case the goal is to estimate the least squares solution using few responses, as quickly as possible, for worst-case inputs. Despite many attempts, characterizing the relationship between these two approaches has proven elusive. We address this by proposing a framework for experimental design where the responses are produced by an arbitrary unknown distribution. We show that there is an efficient randomized experimental design procedure that achieves strong variance bounds for an unbiased estimator using few responses in this general model. Nearly tight bounds for the classical A-optimality criterion, as well as improved bounds for worst-case responses, emerge as special cases of this result. In the process, we develop a new algorithm for a joint sampling distribution called volume sampling, and we propose a new i.i.d. importance sampling method: inverse score sampling. A key novelty of our analysis is in developing new expected error bounds for worst-case regression by controlling the tail behavior of i.i.d. sampling via the jointness of volume sampling. Our result motivates a new minimax-optimality criterion for experimental design which can be viewed as an extension of both A-optimal design and sampling for worst-case regression.
Adversarial Networks and Autoencoders: The Primal-Dual Relationship and Generalization Bounds
Husain, Hisham, Nock, Richard, Williamson, Robert C.
Since the introduction of Generative Adversarial Networks (GANs) and Variational Autoencoders (VAE), the literature on generative modelling has witnessed an overwhelming resurgence. The impressive, yet elusive empirical performance of GANs has lead to the rise of many GAN-VAE hybrids, with the hopes of GAN level performance and additional benefits of VAE, such as an encoder for feature reduction, which is not offered by GANs. Recently, the Wasserstein Autoencoder (WAE) was proposed, achieving performance similar to that of GANs, yet it is still unclear whether the two are fundamentally different or can be further improved into a unified model. In this work, we study the $f$-GAN and WAE models and make two main discoveries. First, we find that the $f$-GAN objective is equivalent to an autoencoder-like objective, which has close links, and is in some cases equivalent to the WAE objective - we refer to this as the $f$-WAE. This equivalence allows us to explicate the success of WAE. Second, the equivalence result allows us to, for the first time, prove generalization bounds for Autoencoder models (WAE and $f$-WAE), which is a pertinent problem when it comes to theoretical analyses of generative models. Furthermore, we show that the $f$-WAE objective is related to other statistical quantities such as the $f$-divergence and in particular, upper bounded by the Wasserstein distance, which then allows us to tap into existing efficient (regularized) OT solvers to minimize $f$-WAE. Our findings thus recommend the $f$-WAE as a tighter alternative to WAE, comment on generalization abilities and make a step towards unifying these models.