Goto

Collaborating Authors

 Government


Structured Content Preservation for Unsupervised Text Style Transfer

arXiv.org Machine Learning

Text style transfer aims to modify the style of a sentence while keeping its content unchanged. Recent style transfer systems often fail to faithfully preserve the content after changing the style. In particular, we achieve the goal by devising rich model objectives based on both the sentence's lexical information and a language model that conditions on content. The resulting model therefore is encouraged to retain the semantic meaning of the target sentences. We perform extensive experiments that compare our model to other existing approaches in the tasks of sentiment and political slant transfer. Our model achieves significant improvement in terms of both content preservation and style transfer in automatic and human evaluation. Text style transfer is an important task in designing sophisticated and controllable natural language generation (NLG) systems. The goal of this task is to convert a sentence from one style (e.g., negative sentiment) to another (e.g., positive sentiment), while preserving the style-independent content (e.g., the name of the food being discussed). Typically, it is difficult to find parallel data with different styles. So we must learn to disentangle the representations of the style from the content. However, it is impossible to separate the two components by simply adding or dropping certain words.


Scalable End-to-End Autonomous Vehicle Testing via Rare-event Simulation

arXiv.org Machine Learning

Recent breakthroughs in deep learning have accelerated the development of autonomous vehicles (AVs); many research prototypes now operate on real roads alongside human drivers. While advances in computer-vision techniques have made human-level performance possible on narrow perception tasks such as object recognition, several fatal accidents involving AVs underscore the importance of testing whether the perception and control pipeline--when considered as a whole system--can safely interact with humans. Unfortunately, testing AVs in real environments, the most straightforward validation framework for system-level input-output behavior, requires prohibitive amounts of time due to the rare nature of serious accidents [49]. Concretely, a recent study [29] argues that AVs need to drive "hundreds of millions of miles and, under some scenarios, hundreds of billions of miles to create enough data to clearly demonstrate their safety." Alteratively, formally verifying an AV algorithm's "correctness" [34, 2, 47, 37] is difficult since all driving policies are subject to crashes caused by other drivers [49]. It is unreasonable to ask that the policy be safe under all scenarios. Unfortunately, ruling out scenarios where the AV should not be blamed is a task subject to logical inconsistency, combinatorial growth in specification complexity, and subjective assignment of fault. Motivated by the challenges underlying real-world testing and formal verification, we consider a probabilistic paradigm--which we call a risk-based framework--where our goal is to evaluate the probability of an accident under a base distribution representing standard traffic behavior.


Understanding Deep Neural Networks through Input Uncertainties

arXiv.org Machine Learning

Techniques for understanding the functioning of complex machine learning models are becoming increasingly popular, not only to improve the validation process, but also to extract new insights about the data via exploratory analysis. Though a large class of such tools currently exists, most assume that predictions are point estimates and use a sensitivity analysis of these estimates to interpret the model. Using lightweight probabilistic networks we show how including prediction uncertainties in the sensitivity analysis leads to: (i) more robust and generalizable models; and (ii) a new approach for model interpretation through uncertainty decomposition. In particular, we introduce a new regularization that takes both the mean and variance of a prediction into account and demonstrate that the resulting networks provide improved generalization to unseen data. Furthermore, we propose a new technique to explain prediction uncertainties through uncertainties in the input domain, thus providing new ways to validate and interpret deep learning models.


Escaping the Curse of Dimensionality in Similarity Learning: Efficient Frank-Wolfe Algorithm and Generalization Bounds

arXiv.org Machine Learning

High-dimensional and sparse data are commonly encountered in many applications of machine learning, such as computer vision, bioinformatics, text mining and behavioral targeting. To classify, cluster or rank data points, it is important to be able to compute semantically meaningful similarities between them. However, defining an appropriate similarity measure for a given task is often difficult as only a small and unknown subset of all features are actually relevant. For instance, in drug discovery studies, chemical compounds are typically represented by a large number of sparse features describing their 2D and 3D properties, and only a few of them play in role in determining whether the compound will bind to a particular target receptor (Leach and Gillet, 2007). In text classification and clustering, a document is often represented as a sparse bag of words, and only a small subset of the dictionary is generally useful to discriminate between documents about different topics. Another example is targeted advertising, where ads are selected based on fine-grained user history (Chen et al., 2009). Similarity and metric learning (Bellet et al., 2015) offers principled approaches to construct a taskspecific similarity measure by learning it from weakly supervised data, and has been used in many application domains. The main theme in these methods is to learn the parameters of a similarity (or distance) function such that it agrees with task-specific similarity judgments (e.g., of the form "data point x should


Low-Precision Random Fourier Features for Memory-Constrained Kernel Approximation

arXiv.org Artificial Intelligence

We investigate how to train kernel approximation methods that generalize well under a memory budget. Building on recent theoretical work, we define a measure of kernel approximation error which we find to be much more predictive of the empirical generalization performance of kernel approximation methods than conventional metrics. An important consequence of this definition is that a kernel approximation matrix must be high-rank to attain close approximation. Because storing a high-rank approximation is memory-intensive, we propose using a low-precision quantization of random Fourier features (LP-RFFs) to build a high-rank approximation under a memory budget. Theoretically, we show quantization has a negligible effect on generalization performance in important settings. Empirically, we demonstrate across four benchmark datasets that LP-RFFs can match the performance of full-precision RFFs and the Nystr\"{o}m method, with 3x-10x and 50x-460x less memory, respectively.


Towards a more efficient use of process and product traceability data for continuous improvement of industrial performances

arXiv.org Artificial Intelligence

Nowadays all industrial sectors are increasingly faced with the explosion in the amount of data. Therefore, it raises the question of the efficient use of this large amount of data. In this research work, we are concerned with process and product traceability data. In some sectors (e.g. pharmaceutical and agro-food), the collection and storage of these data are required. Beyond this constraint (regulatory and / or contractual), we are interested in the use of these data for continuous improvements of industrial performances. Two research axes were identified: product recall and responsiveness towards production hazards. For the first axis, a procedure for product recall exploiting traceability data will be propose. The development of detection and prognosis functions combining process and product data is envisaged for the second axis.


Some DJI Matrice 200 drones are falling out of the sky

Engadget

Some DJI drones are falling from the sky and no one is sure why. The United Kingdom Civil Aviation Authority (CAA) issued a safety notice Friday warning that some DJI Matrice 200 model drones have lost power mid-flight without warning and dropped straight down. The Chinese drone maker acknowledged the issue and said that it is working to address the matter. According to the CAA, there have been a small number of incidents reported in which the Matrice 200 has completely lost power during flight. The problem occurs even when there appears to be charge remaining in the drone's battery.


How The UK Government Uses Artificial Intelligence To Identify Welfare And State Benefits Fraud

#artificialintelligence

Investment in data strategy, technologies that support machine learning and artificial intelligence, and hiring skilled data professionals is a top priority for the UK government. Ministers of the Department for Work and Pensions (DWP) have rolled out and tested AI to automate claims processing and fight fraud within their department. Over the last year, the department unleashed artificial intelligence algorithms to track down large-scale corruption of the benefit and welfare program to stop criminal gangs who are responsible for extremely large losses. Ultimately, this effort protects taxpayers' money and gets the benefits to those who they are intended for. Benefit fraud at the hands of criminal gangs cost the Department of Work and Pensions nearly £2.1 billion in 2016, a rise of £200 million in just one year.


How the UK Could Leverage AI to Lead The 4th Industrial Revolution

#artificialintelligence

The growth of AI is much more than a technological advancement. AI will shape the future of the entire world. The impact of AI will be so huge that governments and companies who dominate AI will define the way our world will operate in the future. The economic impact of AI is estimated to be $15 trillion over the next 10 years, and a dramatic shift is currently underway that will determine which countries will have the advantage. Some countries have made AI a core element of their economic and geopolitical agenda.


EU roaming charges will 'probably' return after Brexit, MPs warn

The Independent - Tech

The UK Government will not enforce the ban on mobile roaming charges for British citizens travelling in the EU after Brexit, a cross-party committee of MPs has warned. The EU's ban on mobile roaming charges for voice, text messages and data has meant using a smartphone on the continent is the same cost as using it at home. But the House of Commons EU Scrutiny Committee said in a report that this is unlikely to remain in place for the UK following withdrawal. It will mean that network operators will be the ones to decide whether to re-introduce roaming charges on UK customers, following their abolition last year across the whole EU. The I.F.O. is fuelled by eight electric engines, which is able to push the flying object to an estimated top speed of about 120mph.