Europe
On Adversarial Examples for Character-Level Neural Machine Translation
Ebrahimi, Javid, Lowd, Daniel, Dou, Dejing
Evaluating on adversarial examples has become a standard procedure to measure robustness of deep learning models. Due to the difficulty of creating white-box adversarial examples for discrete text input, most analyses of the robustness of NLP models have been done through black-box adversarial examples. We investigate adversarial examples for character-level neural machine translation (NMT), and contrast black-box adversaries with a novel white-box adversary, which employs differentiable string-edit operations to rank adversarial changes. We propose two novel types of attacks which aim to remove or change a word in a translation, rather than simply break the NMT. We demonstrate that white-box adversarial examples are significantly stronger than their black-box counterparts in different attack scenarios, which show more serious vulnerabilities than previously known. In addition, after performing adversarial training, which takes only 3 times longer than regular training, we can improve the model's robustness significantly.
Improving Text-to-SQL Evaluation Methodology
Finegan-Dollak, Catherine, Kummerfeld, Jonathan K., Zhang, Li, Ramanathan, Karthik, Sadasivam, Sesh, Zhang, Rui, Radev, Dragomir
To be informative, an evaluation must measure how well systems generalize to realistic unseen data. We identify limitations of and propose improvements to current evaluations of text-to-SQL systems. First, we compare human-generated and automatically generated questions, characterizing properties of queries necessary for real-world applications. To facilitate evaluation on multiple datasets, we release standardized and improved versions of seven existing datasets and one new text-to-SQL dataset. Second, we show that the current division of data into training and test sets measures robustness to variations in the way questions are asked, but only partially tests how well systems generalize to new queries; therefore, we propose a complementary dataset split for evaluation of future work. Finally, we demonstrate how the common practice of anonymizing variables during evaluation removes an important challenge of the task. Our observations highlight key difficulties, and our methodology enables effective measurement of future development.
DALEX: explainers for complex predictive models
Predictive modeling is invaded by elastic, yet complex methods such as neural networks or ensembles (model stacking, boosting or bagging). Such methods are usually described by a large number of parameters or hyper parameters - a price that one needs to pay for elasticity. The very number of parameters makes models hard to understand. This paper describes a consistent collection of explainers for predictive models, a.k.a. black boxes. Each explainer is a technique for exploration of a black box model. Presented approaches are model-agnostic, what means that they extract useful information from any predictive method despite its internal structure. Each explainer is linked with a specific aspect of a model. Some are useful in decomposing predictions, some serve better in understanding performance, while others are useful in understanding importance and conditional responses of a particular variable. Every explainer presented in this paper works for a single model or for a collection of models. In the latter case, models can be compared against each other. Such comparison helps to find strengths and weaknesses of different approaches and gives additional possibilities for model validation. Presented explainers are implemented in the DALEX package for R. They are based on a uniform standardized grammar of model exploration which may be easily extended. The current implementation supports the most popular frameworks for classification and regression.
signSGD: Compressed Optimisation for Non-Convex Problems
Bernstein, Jeremy, Wang, Yu-Xiang, Azizzadenesheli, Kamyar, Anandkumar, Anima
Training large neural networks requires distributing learning across multiple workers, where the cost of communicating gradients can be a significant bottleneck. signSGD alleviates this problem by transmitting just the sign of each minibatch stochastic gradient. We prove that it can get the best of both worlds: compressed gradients and SGD-level convergence rate. The relative $\ell_1/\ell_2$ geometry of gradients, noise and curvature informs whether signSGD or SGD is theoretically better suited to a particular problem. On the practical side we find that the momentum counterpart of signSGD is able to match the accuracy and convergence speed of Adam on deep Imagenet models. We extend our theory to the distributed setting, where the parameter server uses majority vote to aggregate gradient signs from each worker enabling 1-bit compression of worker-server communication in both directions. Using a theorem by Gauss we prove that majority vote can achieve the same reduction in variance as full precision distributed SGD. Thus, there is great promise for sign-based optimisation schemes to achieve fast communication and fast convergence. Code to reproduce experiments is to be found at https://github.com/jxbz/signSGD.
An Inductive Formalization of Self Reproduction in Dynamical Hierarchies
Formalizing self reproduction in dynamical hierarchies is one of the important problems in Artificial Life (AL) studies. We study, in this paper, an inductively defined algebraic framework for self reproduction on macroscopic organizational levels under dynamical system setting for simulated AL models and explore some existential results. Starting with defining self reproduction for atomic entities we define self reproduction with possible mutations on higher organizational levels in terms of hierarchical sets and the corresponding inductively defined `meta' - reactions. We introduce constraints to distinguish a collection of entities from genuine cases of emergent organizational structures.
GraphRNN: Generating Realistic Graphs with Deep Auto-regressive Models
You, Jiaxuan, Ying, Rex, Ren, Xiang, Hamilton, William L., Leskovec, Jure
Modeling and generating graphs is fundamental for studying networks in biology, engineering, and social sciences. However, modeling complex distributions over graphs and then efficiently sampling from these distributions is challenging due to the non-unique, high-dimensional nature of graphs and the complex, non-local dependencies that exist between edges in a given graph. Here we propose GraphRNN, a deep autoregressive model that addresses the above challenges and approximates any distribution of graphs with minimal assumptions about their structure. GraphRNN learns to generate graphs by training on a representative set of graphs and decomposes the graph generation process into a sequence of node and edge formations, conditioned on the graph structure generated so far. In order to quantitatively evaluate the performance of GraphRNN, we introduce a benchmark suite of datasets, baselines and novel evaluation metrics based on Maximum Mean Discrepancy, which measure distances between sets of graphs. Our experiments show that GraphRNN significantly outperforms all baselines, learning to generate diverse graphs that match the structural characteristics of a target set, while also scaling to graphs 50 times larger than previous deep models.
An Improved Generic Bet-and-Run Strategy for Speeding Up Stochastic Local Search
Weise, Thomas, Wu, Zijun, Wagner, Markus
A commonly used strategy for improving optimization algorithms is to restart the algorithm when it is believed to be trapped in an inferior part of the search space. Building on the recent success of Bet-and-Run approaches for restarted local search solvers, we introduce an improved generic Bet-and-Run strategy. The goal is to obtain the best possible results within a given time budget t using a given black-box optimization algorithm. If no prior knowledge about problem features and algorithm behavior is available, the question about how to use the time budget most efficiently arises. We propose to first start k>=1 independent runs of the algorithm during an initialization budget t1
Switzerland's Xherdan Shaqiri Is a Roomba Made of Lead
Switzerland beat Serbia 2โ1 on Friday in a terrifically exciting game that was decided at the brink of regulation when Swiss star Xherdan Shaqiri sprinted 60 yards and slotted the ball just beyond the goalkeeper's reach. Shaqiri is a mercurial player, but when he's in good form he can seemingly do anything he wants on the pitch. He also happens to be built like a rotary phone. Squat and stout, he has no business being as quick as he is, and yet it shouldn't surprise anyone that he was able to leave the Serbian defense in his dust. Given Shaqiri's unique dimensions, his goal looked less like a goal and more like, well, a whole bunch of other bizarre things.
Global Bigdata Conference
A Chinese aphorism says that "the fire burns highest when everyone adds wood to it." It's an apt way to describe the way that industrial design and product development are becoming a collaborative undertaking. Cities like Shenzhen, long known as factory towns that churn out low-end toys and shoes, are embracing a new identity as creative meccas for design. This trend is gathering steam worldwide, for one main reason: design tools are starting to function less like inanimate objects and more like colleagues or assistants. As people and machines begin working together in new ways, the field of design will turn into a team sport, one where human ingenuity combines with artificial intelligence and automation to broaden the possibilities of how society shapes the world.
Experts Bet on First Deepfakes Political Scandal
A quiet wager has taken hold among researchers who study artificial intelligence techniques and the societal impacts of such technologies. They're betting whether or not someone will create a so-called Deepfake video about a political candidate that receives more than 2 million views before getting debunked by the end of 2018. The actual stakes in the bet are fairly small: Manhattan cocktails as a reward for the "yes" camp and tropical tiki drinks for the "no" camp. But the implications of the technology behind the bet's premise could potentially reshape governments and undermine societal trust in the idea of having shared facts. It all comes down to when the technology may mature enough to digitally create fake but believable videos of politicians and celebrities saying or doing things that never actually happened in real life.