Oceania
XL-Editor: Post-editing Sentences with XLNet
Shih, Yong-Siang, Chang, Wei-Cheng, Yang, Yiming
While neural sequence generation models achieve initial su c-cess for many NLP applications, the canonical decoding procedure with left-to-right generation order (i.e., autoreg res-sive) in one-pass can not reflect the true nature of human revising a sentence to obtain a refined result. In this work, we propose XL-Editor, a novel training framework that enables state-of-the-art generalized autoregressive pretrainin g methods, XLNet specifically, to revise a given sentence by the variable-length insertion probability. Concretely, XL-E ditor can (1) estimate the probability of inserting a variable-le ngth sequence into a specific position of a given sentence; (2) execute post-editing operations such as insertion, deletion, and replacement based on the estimated variable-length insert ion probability; (3) complement existing sequence-to-sequen ce models to refine the generated sequences. Empirically, we first demonstrate better post-editing capabilities of XL-E ditor over XLNet on the text insertion and deletion tasks, which validates the effectiveness of our proposed framework. Fur - thermore, we extend XL-Editor to the unpaired text style transfer task, where transferring the target style onto a gi ven sentence can be naturally viewed as post-editing the senten ce into the target style. XL-Editor achieves significant impro ve-ment in style transfer accuracy and also maintains coherent semantic of the original sentence, showing the broad applic ability of our method.
LSTM-Assisted Evolutionary Self-Expressive Subspace Clustering
Xu, Di, Long, Tianhang, Gao, Junbin
Massive volumes of high-dimensional data that evolves over time is continuously collected by contemporary information processing systems, which brings up the problem of organizing this data into clusters, i.e. achieve the purpose of dimensional deduction, and meanwhile learning its temporal evolution patterns. In this paper, a framework for evolutionary subspace clustering, referred to as LSTM-ESCM, is introduced, which aims at clustering a set of evolving high-dimensional data points that lie in a union of low-dimensional evolving subspaces. In order to obtain the parsimonious data representation at each time step, we propose to exploit the so-called self-expressive trait of the data at each time point. At the same time, LSTM networks are implemented to extract the inherited temporal patterns behind data in an overall time frame. An efficient algorithm has been proposed based on MATLAB. Next, experiments are carried out on real-world datasets to demonstrate the effectiveness of our proposed approach. And the results show that the suggested algorithm dramatically outperforms other known similar approaches in terms of both run time and accuracy.
Kernels of Mallows Models under the Hamming Distance for solving the Quadratic Assignment Problem
Arza, Etor, Perez, Aritz, Irurozki, Ekhine, Ceberio, Josu
The Quadratic Assignment Problem (QAP) is a well-known permutation-based combinatorial optimization problem with real applications in industrial and logistics environments. Motivated by the challenge that this NPhard problem represents, it has captured the attention of the evolutionary computation community for decades. As a result, a large number of algorithms have been proposed to optimize this algorithm. Among these, exact methods are only able to solve instances of size n 40, and thus, many heuristic and metaheuristic methods have been applied to the QAP. In this work, we follow this direction by approaching the QAP through Estimation of Distribution Algorithms (EDAs). Particularly, a nonparametric distance-based exponential probabilistic model is used. Based on the analysis of the characteristics of the QAP, and previous work in the area, we introduce Kernels of Mallows Model under the Hamming distance to the context of EDAs. Conducted experiments point out that the performance of the proposed algorithm in the QAP is superior to (i) the classical EDAs adapted to deal with the QAP, and also (ii) to the specific EDAs proposed in the literature to deal with permutation problems.1. Introduction The Quadratic Assignment Problem (QAP) [30] is a well known combinatorial optimization problem. Along with other problems, such as the traveling salesman problem, the linear ordering problem and the flowshop scheduling problem, it belongs to the family of permutation-based (a permutation is a bijection of the set {1,...,n } onto itself) problems [10]. The QAP has been applied in many different environments over the years, to name but a few notable examples, selecting optimal hospital layouts [24], optimally placing components on circuit boards [44], assigning gates at airports [23] or optimizing data transmission [38]. Sahni and Gonzalez [45] proved that the QAP is an NPhard optimization problem, and as such, no polynomial-time exact algorithm can solve this problem unless P NP. M: The size of the set of selected solutions. S: The number of new solutions generated at each iteration.
Sky's the limit! Alphabet's drone delivery service Wing officially begins service in Virginia
After years of development, Alphabet's drone delivery service Wing is officially open for business. The company announced the beginning of service for residents of Christiansburg, Virginia, who will be able to order over-the-counter medication, snacks, and other small items and have them airlifted straight to their homes by a drone. Initially, Wing will deliver goods on behalf of three partner companies with FedEx, Walgreens, and Super Magnolia, a local Virginia grocery store chain. After years of preparation, Alphabet's drone delivery service Wing has officially begun operations in Christiansburg, Virginia The company made the announcement via a blog post on Medium and included a video showing how the delivery service will work. The FAA approved Alphabet's drone delivery program in March, and the company announced it's plans for'store to door' of more than 100 products in Virginia last September.
Wing launches drone delivery in Christiansburg, Virginia
As of this afternoon, select residents of Christiansburg, Virginia can tap drones operated by Google parent company Alphabet's Wing for quick and easy deliveries of packages, over-the-counter medications, snacks, and gifts. The company today revealed that it's become the first to operate a commercial air delivery service directly to homes in the U.S., with the launch of a previously announced pilot involving FedEx Express, Walgreens, and local Virginia retailer Sugar Magnolia. From Walgreens, the first retail pharmacy to partner with Wing in the U.S., drone fleets will ferry over-the-counter medicines and other wellness items to folks' homes. And on the FedEx Express side, recipients living within designated Christiansburg zones who opt in will receive some shipments via drone, in customized boxes. Most orders within the four-mile radius of Wing's distribution facility are fulfilled within about 10 minutes, according to the company.
IDC Expects Asia/Pacific* Artificial Intelligence Systems Spending to Reach USD 6.2 Billion in 2019
SINGAPORE, October 18th, 2019 – Asia/Pacific* spending on artificial intelligence (AI) systems will reach USD 6.2 billion in 2019, recording an increase of almost 54% when compared to 2018, according to the latest IDC Worldwide Semiannual Artificial Intelligence Systems Spending Guide. Evidently, as industries invest aggressively in projects that utilize AI software capabilities, IDC expects spending on AI systems will increase to USD 21.4 billion by 2023 with a compound annual growth rate (CAGR) of 39.6% over the 2018-23 forecast period. From providing chat bots for better customer service to improve the efficiency of operations and tasks for their business models, industries like Banking, Retail and professional services are spending in this technology at scale says," Ritika Srivastava, Associate Market Analyst at IDC Asia/Pacific. In 2019, Asia/Pacific* spending on AI systems will be led by the Banking industry with 10.7% share of the total, followed by retail with a 10.2% share.
Group scours Pacific for sunken WWII battleships, lost war graves
FILE - In this June 4, 1942 file photo provided by the U.S. Navy shows the USS Yorktown listing heavily to port after being struck by Japanese bombers and torpedo planes in the Battle of Midway. Researchers scouring the world's oceans for sunken World War II ships are honing in on debris fields deep in the Pacific.(AP MIDWAY ATOLL, Northwestern Hawaiian Islands (AP) -- Deep-sea explorers scouring the world's oceans for sunken World War II ships are focusing on debris fields deep in the Pacific, in an area where one of the most decisive battles of the time took place. Hundreds of miles off Midway Atoll, nearly halfway between the United States and Japan, a research vessel is launching underwater robots miles into the abyss to look for warships from the famed Battle of Midway. Weeks of grid searches around the Northwestern Hawaiian Islands have already led the crew of the Petrel to one sunken warship, the Japanese ship the Kaga.
Your Next Boss Could Be a Computer
At its core, technology exists to make our lives easier. Thanks to artificial intelligence, our tools have gotten smarter, and we're more productive as a result. According to a study released earlier today, workers around the world not only recognize AI's importance in the modern workplace – they embrace it. Conducted over the summer in partnership between Oracle and Future Workspace, the second annual AI at Work study asked 8,370 employees, managers and HR leaders from 10 countries about AI and its place in their work. Researchers found that AI is rapidly changing not only how we conduct business, but the very relationship between people and the tech they use every day.
Context-Driven Data Mining through Bias Removal and Data Incompleteness Mitigation
Batarseh, Feras A., Kulkarni, Ajay
The results of data mining endeavors are majorly driven by data quality. Throughout these deployments, serious show-stopper problems are still unresolved, such as: data collection ambiguities, data imbalance, hidden biases in data, the lack of domain information, and data incompleteness. This paper is based on the premise that context can aid in mitigating these issues. In a traditional data science lifecycle, context is not considered. Context-driven Data Science Lifecycle (C-DSL); the main contribution of this paper, is developed to address these challenges. Two case studies (using data-sets from sports events) are developed to test C-DSL. Results from both case studies are evaluated using common data mining metrics such as: coefficient of determination (R2 value) and confusion matrices. The work presented in this paper aims to re-define the lifecycle and introduce tangible improvements to its outcomes.
Machine Learning Systems for Highly-Distributed and Rapidly-Growing Data
The usability and practicality of any machine learning (ML) applications are largely influenced by two critical but hard-to-attain factors: low latency and low cost. Unfortunately, achieving low latency and low cost is very challenging when ML depends on real-world data that are highly distributed and rapidly growing (e.g., data collected by mobile phones and video cameras all over the world). Such real-world data pose many challenges in communication and computation. For example, when training data are distributed across data centers that span multiple continents, communication among data centers can easily overwhelm the limited wide-area network bandwidth, leading to prohibitively high latency and high cost. In this dissertation, we demonstrate that the latency and cost of ML on highly-distributed and rapidly-growing data can be improved by one to two orders of magnitude by designing ML systems that exploit the characteristics of ML algorithms, ML model structures, and ML training/serving data. We support this thesis statement with three contributions. First, we design a system that provides both low-latency and low-cost ML serving (inferencing) over large-scale and continuously-growing datasets, such as videos. Second, we build a system that makes ML training over geo-distributed datasets as fast as training within a single data center. Third, we present a first detailed study and a system-level solution on a fundamental and largely overlooked problem: ML training over non-IID (i.e., not independent and identically distributed) data partitions (e.g., facial images collected by cameras varies according to the demographics of each camera's location).